Soba Docs

Integrate

The endpoint

Point your existing OpenAI client at Soba. One base URL, one key, and the user's id on every request. Prompts, tools, streaming and error handling stay exactly as they are.

Soba speaks the OpenAI API. There is no SDK to install and no provider to import.

JavaScript
// No SDK to install. Your OpenAI client already speaks to Soba.
import OpenAI from "openai";

const soba = new OpenAI({
  baseURL: "https://api.soba.so/v1", // the only line that changes
  apiKey: process.env.SOBA_KEY,
});

const response = await soba.chat.completions.create({
  messages: [
    { role: "system", content: "You are a helpful assistant" },
    { role: "user", content: "What is Soba?" },
  ],
  model: "auto",
  user: "usr_123",
});
Python
import os
from openai import OpenAI

soba = OpenAI(
    base_url="https://api.soba.so/v1",  # the only line that changes
    api_key=os.environ["SOBA_KEY"],
)

response = soba.chat.completions.create(
    messages=[
        {"role": "system", "content": "You are a helpful assistant"},
        {"role": "user", "content": "What is Soba?"},
    ],
    model="auto",
    user="usr_123",
)
Terminal
curl https://api.soba.so/v1/chat/completions \
  -H "Authorization: Bearer $SOBA_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto",
    "user": "usr_123",
    "messages": [
      { "role": "user", "content": "What is Soba?" }
    ]
  }'

The three things that change#

baseURL https://api.soba.so/v1
apiKey A key from your dashboard, shaped sk_soba_<id>.<secret>. Not the token a paired machine holds, which is a different credential entirely
user The signed-in user's id, on every request

Nothing else moves. Prompts, tools, streaming and error handling stay exactly as they are, because this is the same request you already send, at a different host.

The `user` field is not optional here

It is a hint in the OpenAI API and load-bearing in Soba: it is how a run is attributed to a person, and therefore how Soba knows whose machine the run may reach. A request without it cannot be routed to anyone's own compute, and cannot be metered against anyone's plan. See Provider terms.

Prove it in one request#

Terminal
curl https://api.soba.so/v1/chat/completions \
  -H "Authorization: Bearer $SOBA_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "soba/echo", "user": "usr_123",
        "messages": [{ "role": "user", "content": "ping" }] }'

soba/echo is answered by Soba itself. It needs no paired machine and no provider key, nothing is billed for it, and it still proves the three things that actually break: the key, the base URL, and the user field.

That matters on day one, when nobody has connected a machine yet. Soba has no provider to import, so at integration time there is nothing to connect, and without echo a first request would have nowhere to land and every integration would begin with a failure.

Models#

Model
auto Routes each run to the best compute that user has and this endpoint can reach: a model running on their own machine first, and a handback to your app when there is nothing to route to
soba/echo Answered by Soba. No machine, no provider key, nothing billed. The health check

Provider-specific aliases work too. auto is the one to reach for, because naming a model forces your app to know what a given user has connected. See Routing a run.

What `auto` will not pick here

Their Claude Pro or Max plan, or their ChatGPT plan. Those runs are served by Claude Code and Codex, which drive their own agent loop and cannot be operated through a stateless chat/completions exchange; the reason is at the end of this page. The same auto on soba.run() picks from everything, subscription tiers included.

What you get immediately#

Pointing the call site at Soba, before anyone connects anything, already gives you:

  • Metering per user, stamped with its cost class
  • Spend ceilings, checked between turns
  • A handback for anyone with nothing connected, so your existing code serves them instead of the run failing

And it gives you placement. Once the call site points at Soba, turning on bring-your-own compute later needs no code change from you: the same request starts landing on a model the user runs themselves as soon as they connect one. Reaching the plan they already pay for is the one thing that does need a change at the call site, and that change is soba.run().

When nothing is reachable#

Soba never holds your provider keys, so it never serves a run on them. When a user has nothing reachable and the route allows app, the request comes back as a 409 with the code serve_in_app, and your code serves it with the client it already uses:

JavaScript
async function complete(params, userId) {
  try {
    return await soba.chat.completions.create({ ...params, user: userId });
  } catch (err) {
    if (err.code !== "serve_in_app") throw err;
    return await openai.chat.completions.create(params); // your key, your server, as today
  }
}

The run has already been counted against the user's plan by then, so a plan that stops at its allowance returns a 402 instead and nothing is handed back. Report what your own call used with POST /v1/runs/:id/usage, using the id in the Soba-Run-Id header, and token-counted plans, the dashboard and conversion data stay complete. soba.run() does all of this for you.

What this endpoint cannot do#

chat/completions is stateless and the caller owns the loop. That is fine for a model, and it is the reason this endpoint can only reach model runtimes.

Claude Code and Codex run their own loop and never hand a tool call back to wait on an HTTP round trip. Reaching those means letting Soba own the loop, which is what soba.run() is for.

© 2026 Soba resolved = machine grant ∩ broker request