Soba Docs

Integrate

Errors and retries

What comes back when a run cannot land, which failures are worth retrying, and why a run that stops on a ceiling is not a crash.

Most of this product is about a run finding somewhere to go. This page is about the times it does not, and about telling the three kinds of failure apart, because they want three different responses:

Example What to do
Your request is wrong No user, a model nobody has, a cwd outside every granted root Fix it. Retrying repeats it
Nowhere to run it Asleep laptop, signed-out CLI, every reachable class outside allow Retry later, or widen what you accept
The run ran and stopped A ceiling reached, an approval refused, a timeout Not an error. Read the terminal event

The envelope#

The endpoint is OpenAI-compatible in failure too, so your existing error handling already reads it:

JSON
{
  "error": {
    "type": "no_compute",
    "message": "No compute this run may use is reachable for usr_123.",
    "code": "no_compute"
  }
}
Status Means
400 Malformed. A run with neither prompt nor a non-empty messages, an unparseable route
401 The key is missing, malformed, revoked, or belongs to another app
402 The user's plan is exhausted and its overage is stop. See what does not spend an allowance
403 The request asked for something this key may not have: a cost class outside the environment ceiling, most often a metered class on a development key
404 No such model or runtime alias
409 Not served by Soba. no_compute: nothing this run may use is connected and awake. serve_in_app: nothing of the user's is reachable and the route allows app, so the run is yours to serve
422 The user id is missing. It cannot be defaulted, because a run with no user cannot be attributed or metered
429 Rate limited by Soba. Back off; the Retry-After header says how long. Not to be confused with a delivered error carrying reason: "subscription_exhausted", which is a user's own plan running out and is not retryable
5xx Soba's problem. Retry with backoff

409 is the one worth handling deliberately, and the route you sent decides which one you get. An app that allows app gets serve_in_app and serves the run with its existing client; the SDK's fallback does that for you. An app that allows only user-hardware and user-subscription gets no_compute: it has explicitly chosen a run that can fail rather than a run that costs money, and this is that choice arriving.

Streaming fails differently#

An error before the first byte is an HTTP status you can catch. After it, the response is already 200 and streaming, so a failure arrives as a terminal error event instead:

JSON
{"type":"error","message":"The machine serving this run disconnected."}

error and done are the only terminal types. A stream that ends without either is a transport failure, not an outcome. Treat a truncated stream as retryable and a delivered error as final.

Do not retry a stream from the beginning by reflex

A run that produced 4,000 tokens on someone's own machine and then dropped costs you nothing to retry, and a run on a metered class costs the same again. costClass on the opening status tells you which one you are about to repeat.

What is retryable#

Situation Retry?
429, 5xx, a truncated stream Yes, with exponential backoff
409 no_compute Yes, but not immediately: the machine is asleep, not busy. Minutes, not milliseconds
409 serve_in_app No. Serve it from your app now, then report its usage
400, 401, 403, 404, 422 No. The same request fails the same way
402 plan exhausted No. It is a billing state, and clears when the user acts. <RunLimit /> is what asks them to
A delivered error event Only if the message says the machine dropped. A policy refusal repeats
reason: "subscription_exhausted" on a delivered error No. The user's own plan is spent and refills on the provider's clock — see the ChatGPT plan channel. retryAfterSeconds says when, when we were told

POST /v1/runs is not idempotent. Retrying a run that already started is a second run, billed as one. If you retry across a network boundary you cannot observe, key your own request id and hold the result yourself.

What does not spend an allowance#

A run served by the user's own machine does not count against a plan whose connectedReward is unlimited. The allowance bounds what you buy, and a run you did not buy spent nothing. It is the same reasoning that already keeps an exhausted monthly cap from touching a machine-served run.

This is what makes the promise on a pricing card mechanically true rather than a line of copy. It also means the reward is earned rather than switched on: only compute that actually serves a run lifts the cap, so a machine that was paired and never wakes lifts nothing.

A connectedReward of included works the other way round, and deliberately: a larger allowance is a promise about the month, so it applies while a machine is connected even on a run that machine is not serving. credit changes no allowance at all; it is money back, assessed at renewal.

A 402 therefore says one of two different things, and <RunLimit /> tells them apart: connect something and carry on, or this plan offers nothing for connecting, so the way on is a bigger one.

A ceiling is not a crash#

Three things stop a healthy run, and all three are ordinary outcomes:

A cost ceiling. route.maxCostMicros stops a run between turns, because nobody knows a turn's cost before it happens. The guarantee is stops as soon as it knows, not never exceeds. You get a terminal event and the usage up to that point, which is billable.

A refused approval. A tool inside the machine's askable set was attempted and the answer was no. In deny mode the call fails and the run continues: the model is told, and usually works around it. Do not treat one refused tool as a failed run.

A timeout. The machine's maxTimeoutMs clamps whatever you asked for, so a run can end earlier than your own timeout suggests. That is the machine owner's ceiling, not a bug.

Failures that never reach you#

Some things you might expect to see as errors are resolved before a run exists, on purpose:

  • A signed-out CLI is withheld rather than advertised and failed at spawn. It cannot be routed to, so it cannot fail a run. The owner is told, with authHint; you are not.
  • A tool a policy does not grant is never offered to the model, so the model does not call it and you do not get a failed call.
  • A request for more than a machine grants is refused on the worker before anything spawns.

The pattern is deliberate: a failure the machine's owner can fix is reported to the owner, and a failure your app can fix is reported to you. See Troubleshooting a machine for the owner's half.

© 2026 Soba resolved = machine grant ∩ broker request