Every endpoint your app calls, in one place: the two ways to start a run, the two that pair a machine, and the reads that answer what a user has and what they used.
https://api.soba.so/v1
One base, one credential. Your server sends the app key
(sk_soba_<id>.<secret>, not the token a paired machine holds):
Authorization: Bearer sk_soba_…
How a paired machine talks to Soba is a different surface entirely, and not one your
app ever sends to. This page is the half your app touches.
Running#
POST /v1/chat/completions#
The OpenAI-compatible endpoint. The request you already send, at a different host, with
user required. Reaches model runtimes; cannot drive Claude Code or Codex. Full
treatment on The endpoint.
POST /v1/runs#
One request that comes back finished, with Soba owning the agent loop.
JSON
{
"user": "usr_123",
"model": "auto",
"prompt": "Summarise everything I saved this week",
"tools": [],
"route": { "prefer": ["user-hardware", "user-subscription"] }
}
Streams events as text/event-stream. The SDK is a
convenience over exactly this shape, not a different product. See
Runs and the SDK.
The run's id comes back in a Soba-Run-Id header, ahead of the first event,
because the three routes below are addressed by it.
Answering a run#
A stream goes one way. A tool call, an approval and a cancellation all need a
return path, so each is a request of its own, keyed to the run:
|
|
POST /v1/runs/:id/tool_result |
{ callId, result, isError? }. Your app's answer to a tool_call frame |
POST /v1/runs/:id/approval |
{ approvalId, decision, reason? }. Anything that is not "allow" is a denial |
POST /v1/runs/:id/cancel |
Stop the run. Best effort: tokens already produced are already produced |
POST /v1/runs/:id/usage |
A RunUsage. What a run your app served (app) used, reported against the id Soba handed back |
The SDK sends all three for you. They are written down because an app driving a run
over raw HTTP has to send them itself.
Neither run endpoint is idempotent
A retry that crosses a boundary you cannot observe starts a second run. See
Errors and retries.
Pairing a machine#
The device authorization grant. The first two are the only endpoints
a worker calls before it has a token, and they take no credential of their own: the
pk_soba_ connect key in the path names the app, and it is public by design.
|
|
POST /c/<pk>/v1/device/code |
Issue a device code and a user code; return the verification URI and a poll interval |
POST /c/<pk>/v1/device/token |
Polled until the user answers; returns the pairing token and where the machine should connect |
slow_down widens the poll interval permanently by five seconds, per RFC 8628.
Your server: minting a session#
One call, with your sk_soba_ key. Everything the browser then does is scoped to the
one user you named, for ten minutes.
|
|
POST /v1/end_users/session |
{ user, email? }. Returns { session, expires_in, url }; send them to url |
DELETE /v1/end_users/session?user=<id> |
End every live session for that user, e.g. when they sign out of your app |
The browser: what a session may do#
All four take Authorization: Bearer est_… and act for exactly one end user.
|
|
GET /v1/connect |
The app's name and the pairing command, so the page can render itself |
GET /v1/device/lookup?code= |
What is asking to be paired: hostname, platform, arch, client |
POST /v1/device/approve |
{ user_code }. Binds the pairing token to this user |
POST /v1/device/deny |
{ user_code }. The worker is told access_denied rather than left polling |
GET /v1/machines |
This user's machines and their live state. What <ComputeStatus /> reads |
The approval is the attribution
POST /v1/device/approve is where the pairing token is bound to a person, and it is
bound to the one the session names, never to anything the machine claimed. A run
may only ever reach a worker owned by the user it is attributed to.
Reading#
The ledger is authoritative, and these are how you read it rather than inferring it
from webhooks.
|
|
GET /v1/users/:id |
What this user has: their plan, allowance, usage this period, and what they have connected |
GET /v1/users/:id/compute |
Each connected machine, its runtimes, cost classes and live state |
GET /v1/usage |
Usage rows, filterable by user, period and environment. What the dashboard reads |
GET /v1/plans |
The plans you configured, as your pricing page renders them |
GET /v1/usage returns the same integer micros everything else in Soba uses. No
floats, no currency strings, and costClass on every row so a total never has to
remember what it routed to.
Plans and payment#
Checkout, entitlement and the customer portal are driven through
<SobaPlans /> and <SobaCheckout />, or the hosted fallback page,
rather than assembled from raw endpoints. Both are thin over this API, and the reason
to prefer them is that they create every Checkout Session in your connected Stripe
account tagged with the user's id, which is what links a payment to the trial that led
to it. See Plans and billing and Stripe.
Errors#
Every endpoint returns the same envelope, and the status codes mean the same thing on
each. See Errors and retries.
JSON
{ "error": { "type": "no_compute", "message": "…", "code": "no_compute" } }