One Ollama Endpoint, Two Very Different Backends

Introduction Nine namespaces in my cluster call a local language model. The SDR stack tags transcriptions, the politics dashboard summarizes feeds, the congressional trade tracker writes daily summaries, and several agents run against the service continuously. All of them were designed to use one hostname on port 11434. That hostname fronts two active Ollama deployments: an RTX 5090 in a desktop tower that is powered off some of the time, and an NVIDIA GB10 Spark board with unified memory where GPU allocations count against the pod’s memory limit. The repository contains a CPU-only deployment manifest, but the kustomization leaves it disabled under normal circumstances. The active fallback is the Spark deployment. ...

August 14, 2026 · 7 min · Robert D. White
Available as a Tor onion service