Soba Docs

Models

Adding an endpoint

Ollama, Together, Groq, vLLM and anything self-hosted speak the same API, so adding one is configuration rather than code, and the cost class is part of that configuration.

~/.soba/models.json:

JSON
[
  { "id": "together", "baseUrl": "https://api.together.xyz/v1",
    "apiKeyEnv": "TOGETHER_API_KEY", "costClass": "user-key",
    "pricing": { "inputMicrosPerMTok": 150000, "outputMicrosPerMTok": 600000 } }
]

Ollama, OpenRelay, Together, Groq, DeepInfra, Fireworks, vLLM and anything self-hosted all speak the OpenAI-compatible API, so there is no adapter to write.

Fields#

Field
id The runtime id this endpoint is advertised under
baseUrl The OpenAI-compatible base, including /v1
apiKeyEnv Preferred. The name of an environment variable holding the key
apiKey The key inline. Works, but see below
costClass Required: user-hardware or user-key. No default; see below
pricing inputMicrosPerMTok / outputMicrosPerMTok. Required for metered classes

Prefer apiKeyEnv over apiKey#

A key written into a config file outlives the reason it was put there and gets copied around with it.

A configured key that isn't set is treated exactly like a signed-out CLI: the runtime is withheld rather than advertised, and reported to you with the fix, rather than failing every run.

Put the key in ~/.soba/service.env (mode 0600) so it survives a service install without appearing in a world-readable unit file. See Keeping it running.

There is no default cost class#

A row without a valid one is skipped.

Guessing who pays is the mistake the field exists to prevent. Declaring it is a claim about whose money this is, and the only money reachable from this file is the machine owner's. Nothing here can spend yours, which is why the file is trusted to say it, and why the only classes it accepts are user-hardware and user-key.

Prices are integer micros#

Throughout: inputMicrosPerMTok is micros per million input tokens. No floats, no currency strings.

route.maxCostMicros stops a run that reaches its ceiling, and that check happens between turns, because nobody knows a turn's cost before it happens. The guarantee is "stops as soon as it knows", not "never exceeds".

Ollama needs no entry#

Ollama is detected by probing its port, and reports the models you have actually pulled. It is user-hardware, so it carries no pricing. Add an entry only when you want to point at a non-default host. SOBA_OLLAMA_URL also works.

Nothing here downloads a model. ollama pull does, and the next probe finds it — the list an endpoint advertises is the list it answered with, never a guess.

If Ollama is running but has nothing pulled, the machine reports it as present and unusable rather than missing, with the one command that fixes it. "Not installed" and "installed with nothing in it" are different problems.

Which model a run gets#

A run may name one. If it does not — or names one this machine has not pulled — the model comes from defaultModels in ~/.soba/policy.json, keyed by runtime id:

JSON
{ "defaultModels": { "ollama": "qwen2.5-coder:7b" } }

The Soba app writes this for you: open the runtime's page, go to Models and press one. Without it the answer is whichever model the endpoint happened to list first, which is nobody's decision.

The same key holds an agent CLI's choice, where the value is one of the aliases it accepts rather than a model it has pulled — { "defaultModels": { "claude-code": "opus" } }. Removing the key, which the Auto row on that page does, puts the decision back where it was: with the CLI's own settings.

A run that names a model you do not have still runs, on the default — and the app is told which model answered, in a status event. It used to substitute in silence, which turned "you never pulled that model" into "the agent behaved strangely" with nothing connecting the two.

Choose the model deliberately

Small models are unreliable at tool calling, and Ollama's advertised capabilities do not tell you which. See Agents and models.

© 2026 Soba resolved = machine grant ∩ broker request