What comes back when a run cannot land, which failures are worth retrying, and why a run that stops on a ceiling is not a crash.
Most of this product is about a run finding somewhere to go. This page is about the
times it does not, and about telling the three kinds of failure apart, because they
want three different responses:
|
Example |
What to do |
| Your request is wrong |
No user, a model nobody has, a cwd outside every granted root |
Fix it. Retrying repeats it |
| Nowhere to run it |
Asleep laptop, signed-out CLI, every reachable class outside allow |
Retry later, or widen what you accept |
| The run ran and stopped |
A ceiling reached, an approval refused, a timeout |
Not an error. Read the terminal event |
The envelope#
The endpoint is OpenAI-compatible in failure too, so your existing error handling
already reads it:
JSON
{
"error": {
"type": "no_compute",
"message": "No compute this run may use is reachable for usr_123.",
"code": "no_compute"
}
}
| Status |
Means |
400 |
Malformed. A run with neither prompt nor a non-empty messages, an unparseable route |
401 |
The key is missing, malformed, revoked, or belongs to another app |
402 |
The user's plan is exhausted and its overage is stop. See what does not spend an allowance |
403 |
The request asked for something this key may not have: a cost class outside the environment ceiling, most often a metered class on a development key |
404 |
No such model or runtime alias |
409 |
Not served by Soba. no_compute: nothing this run may use is connected and awake. serve_in_app: nothing of the user's is reachable and the route allows app, so the run is yours to serve |
422 |
The user id is missing. It cannot be defaulted, because a run with no user cannot be attributed or metered |
429 |
Rate limited by Soba. Back off; the Retry-After header says how long. Not to be confused with a delivered error carrying reason: "subscription_exhausted", which is a user's own plan running out and is not retryable |
5xx |
Soba's problem. Retry with backoff |
409 is the one worth handling deliberately, and the route you sent
decides which one you get. An app that allows app gets serve_in_app and serves the
run with its existing client; the SDK's
fallback does that for you. An app that allows
only user-hardware and user-subscription gets no_compute: it has explicitly chosen
a run that can fail rather than a run that costs money, and this is that choice
arriving.
Streaming fails differently#
An error before the first byte is an HTTP status you can catch. After it, the response
is already 200 and streaming, so a failure arrives as a terminal
error event instead:
JSON
{"type":"error","message":"The machine serving this run disconnected."}
error and done are the only terminal types. A stream that ends without either is a
transport failure, not an outcome. Treat a truncated stream as retryable and a
delivered error as final.
Do not retry a stream from the beginning by reflex
A run that produced 4,000 tokens on someone's own machine and then dropped costs you
nothing to retry, and a run on a metered class costs the same again. costClass on
the opening status tells you which one you are about to repeat.
What is retryable#
| Situation |
Retry? |
429, 5xx, a truncated stream |
Yes, with exponential backoff |
409 no_compute |
Yes, but not immediately: the machine is asleep, not busy. Minutes, not milliseconds |
409 serve_in_app |
No. Serve it from your app now, then report its usage |
400, 401, 403, 404, 422 |
No. The same request fails the same way |
402 plan exhausted |
No. It is a billing state, and clears when the user acts. <RunLimit /> is what asks them to |
A delivered error event |
Only if the message says the machine dropped. A policy refusal repeats |
reason: "subscription_exhausted" on a delivered error |
No. The user's own plan is spent and refills on the provider's clock — see the ChatGPT plan channel. retryAfterSeconds says when, when we were told |
POST /v1/runs is not idempotent. Retrying a run that already started is a second
run, billed as one. If you retry across a network boundary you cannot observe, key
your own request id and hold the result yourself.
What does not spend an allowance#
A run served by the user's own machine does not count against a plan whose
connectedReward is unlimited. The allowance bounds what you buy, and a
run you did not buy spent nothing. It is the same reasoning that already keeps
an exhausted monthly cap from touching a machine-served run.
This is what makes the promise on a pricing card mechanically true rather than a
line of copy. It also means the reward is earned rather than switched on: only
compute that actually serves a run lifts the cap, so a machine that was paired
and never wakes lifts nothing.
A connectedReward of included works the other way round, and deliberately: a
larger allowance is a promise about the month, so it applies while a machine is
connected even on a run that machine is not serving. credit changes no
allowance at all; it is money back, assessed at renewal.
A 402 therefore says one of two different things, and
<RunLimit /> tells them apart: connect something and carry
on, or this plan offers nothing for connecting, so the way on is a bigger one.
A ceiling is not a crash#
Three things stop a healthy run, and all three are ordinary outcomes:
A cost ceiling. route.maxCostMicros stops a run between turns,
because nobody knows a turn's cost before it happens. The guarantee is stops as soon
as it knows, not never exceeds. You get a terminal event and the usage up to that
point, which is billable.
A refused approval. A tool inside the machine's askable set was
attempted and the answer was no. In deny mode the call fails and the run
continues: the model is told, and usually works around it. Do not treat one
refused tool as a failed run.
A timeout. The machine's maxTimeoutMs clamps whatever you asked for, so a run can
end earlier than your own timeout suggests. That is the machine owner's ceiling, not a
bug.
Failures that never reach you#
Some things you might expect to see as errors are resolved before a run exists, on
purpose:
- A signed-out CLI is withheld rather than
advertised and failed at spawn. It cannot be routed to, so it cannot fail a run. The
owner is told, with
authHint; you are not.
- A tool a policy does not grant is never offered to the model, so the model does
not call it and you do not get a failed call.
- A request for more than a machine grants is refused
on the worker before anything spawns.
The pattern is deliberate: a failure the machine's owner can fix is reported to the
owner, and a failure your app can fix is reported to you. See
Troubleshooting a machine for the owner's half.