Point your existing OpenAI client at Soba. One base URL, one key, and the user's id on every request. Prompts, tools, streaming and error handling stay exactly as they are.
Soba speaks the OpenAI API. There is no SDK to install and no provider to import.
JavaScript
// No SDK to install. Your OpenAI client already speaks to Soba.
import OpenAI from "openai";
const soba = new OpenAI({
baseURL: "https://api.soba.so/v1", // the only line that changes
apiKey: process.env.SOBA_KEY,
});
const response = await soba.chat.completions.create({
messages: [
{ role: "system", content: "You are a helpful assistant" },
{ role: "user", content: "What is Soba?" },
],
model: "auto",
user: "usr_123",
});
Python
import os
from openai import OpenAI
soba = OpenAI(
base_url="https://api.soba.so/v1", # the only line that changes
api_key=os.environ["SOBA_KEY"],
)
response = soba.chat.completions.create(
messages=[
{"role": "system", "content": "You are a helpful assistant"},
{"role": "user", "content": "What is Soba?"},
],
model="auto",
user="usr_123",
)
Terminal
curl https://api.soba.so/v1/chat/completions \
-H "Authorization: Bearer $SOBA_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"user": "usr_123",
"messages": [
{ "role": "user", "content": "What is Soba?" }
]
}'
The three things that change#
|
|
baseURL |
https://api.soba.so/v1 |
apiKey |
A key from your dashboard, shaped sk_soba_<id>.<secret>. Not the token a paired machine holds, which is a different credential entirely |
user |
The signed-in user's id, on every request |
Nothing else moves. Prompts, tools, streaming and error handling stay exactly as they
are, because this is the same request you already send, at a different host.
The `user` field is not optional here
It is a hint in the OpenAI API and load-bearing in Soba: it is how a run is attributed
to a person, and therefore how Soba knows whose machine the run may reach. A
request without it cannot be routed to anyone's own compute, and cannot be metered
against anyone's plan. See Provider terms.
Prove it in one request#
Terminal
curl https://api.soba.so/v1/chat/completions \
-H "Authorization: Bearer $SOBA_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "soba/echo", "user": "usr_123",
"messages": [{ "role": "user", "content": "ping" }] }'
soba/echo is answered by Soba itself. It needs no paired machine and no provider
key, nothing is billed for it, and it still proves the three things that actually
break: the key, the base URL, and the user field.
That matters on day one, when nobody has connected a machine yet. Soba has no provider
to import, so at integration time there is nothing to connect, and without echo a
first request would have nowhere to land and every integration would begin with a
failure.
Models#
| Model |
|
auto |
Routes each run to the best compute that user has and this endpoint can reach: a model running on their own machine first, and a handback to your app when there is nothing to route to |
soba/echo |
Answered by Soba. No machine, no provider key, nothing billed. The health check |
Provider-specific aliases work too. auto is the one to reach for, because naming a
model forces your app to know what a given user has connected. See
Routing a run.
What `auto` will not pick here
Their Claude Pro or Max plan, or their ChatGPT plan. Those runs are served by Claude
Code and Codex, which drive their own agent loop and cannot be operated through a
stateless chat/completions exchange; the reason is
at the end of this page. The same auto on
soba.run() picks from everything, subscription tiers included.
Pointing the call site at Soba, before anyone connects anything, already gives you:
- Metering per user, stamped with its cost class
- Spend ceilings, checked between turns
- A handback for anyone with nothing connected, so your existing code serves them
instead of the run failing
And it gives you placement. Once the call site points at Soba, turning on
bring-your-own compute later needs no code change from you: the same request starts
landing on a model the user runs themselves as soon as they connect one. Reaching the
plan they already pay for is the one thing that does need a change at the call site,
and that change is soba.run().
When nothing is reachable#
Soba never holds your provider keys, so it never serves a run on them. When a user has
nothing reachable and the route allows app, the request comes back as a 409 with the
code serve_in_app, and your code serves it with the client it already uses:
JavaScript
async function complete(params, userId) {
try {
return await soba.chat.completions.create({ ...params, user: userId });
} catch (err) {
if (err.code !== "serve_in_app") throw err;
return await openai.chat.completions.create(params); // your key, your server, as today
}
}
The run has already been counted against the user's plan by then, so a plan that stops
at its allowance returns a 402 instead and nothing is handed back. Report what your
own call used with POST /v1/runs/:id/usage, using the id in the
Soba-Run-Id header, and token-counted plans, the dashboard and conversion data stay
complete. soba.run() does all of this for you.
What this endpoint cannot do#
chat/completions is stateless and the caller owns the loop. That is fine for a
model, and it is the reason this endpoint can only reach model runtimes.
Claude Code and Codex run their own loop and never hand a tool call back to wait on an
HTTP round trip. Reaching those means letting Soba own the loop, which is what
soba.run() is for.