Soba Docs

Get started

How a run works

The path a run takes from your app to a subprocess on someone else's laptop and back, and why the machine dials out rather than listening.

A run passes through four things, and the whole design turns on one connection between two of them.

Your app calls the broker; the broker and the worker share one WebSocket carrying frames in both directions; the worker spawns the runtime. The socket is opened by the worker, dialling out to the broker. Your app defines and runs the tools Broker routes the run, meters, bills Worker on the user’s own machine Runtime claude, codex, or a model API one WebSocket dialled out by the worker
Every frame in the sequence below crosses that one socket, and the end that opens it is the one sitting behind a home router.

Your app is your product: your prompt, your tools, your users.

Soba decides where the run is served, meters what it used, and bills your users for your plan through your own Stripe account.

The worker is a small program running on your user's own machine, which is the one placement the product cannot trade away. The next section says why.

The runtime is what actually does the thinking: their Claude Code, their Codex, or a model endpoint.

Your app never talks to a worker, and Soba never runs a model. Each one hands off to the next.

Why it runs on their machine#

The compute Soba is reaching for is a Claude or ChatGPT plan your user already pays for. Plans like that are sold to a person, at a price that only works because it is sold to a person, and they are licensed accordingly. Anthropic does not permit a third party to offer claude.ai logins or rate limits to its own users, so there is no version of this that runs in your cloud on your users' behalf.

What is left is the honest version: the run happens where the plan already lives. On their machine, under their login, billed to their account. Nothing is pooled, nothing is resold, and no credential is ever handed to you or to Soba.

Move that same run onto a server you operate and it stops being a person using the plan they bought. That is the prohibited thing, and it is the line the worker exists to stay on the right side of. Provider terms has the invariant and what enforces it.

The other half of the answer is economic. A subsidised consumer plan gives your user more compute than you could afford to buy for them, so a run served on their machine costs you nothing and outperforms what your budget would have stretched to. The compliance argument says it must be there; the economics say you would want it there anyway.

Why the worker calls Soba, and not the other way round#

Soba cannot call a laptop. A machine on home wifi has no address the internet can reach; it sits behind a router. To dial in, every one of your users would have to open a port on that router: a support burden on you and a hole in their network.

So the worker places the call, the way a browser calls a website, and that connection stays open. When work arrives, Soba sends it down a pipe that already exists. It is the difference between an integration your users have to configure and one they install.

What that buys:

  • no inbound port, so nothing on the machine is reachable from outside
  • no public address and no port forwarding
  • NAT traversal, which is what makes a laptop on a home network a viable place to run anything at all

What happens during a run#

1. The machine says what it can do. The moment it connects, long before any run exists, it reports every runtime it can actually serve: its models, its native tools, its cost class. A runtime the user is signed out of is left off that list rather than advertised as broken, so it cannot be routed to by mistake.

2. Work arrives. The prompt or the structured messages, an optional system prompt, and two optional requests: what this run should be allowed to do, and which cost classes you will accept. Both are requests, and the next step is why that word is doing work.

3. The machine decides what this run may do. Nothing is spawned until two questions are settled, and each is an intersection. Read ∩ as only what is in both:

Two resolutions, each an intersection. Permissions are the machine grant intersected with the broker request; the tier is what the machine can serve intersected with what the app declared it would accept. Neither result can be wider than its inputs. The machine grant what the owner allows The broker request what your app asks for permissions never wider than either ∩ What this machine can serve the tiers it actually has What the app will accept route.allow tier one of those two lists ∩

Neither line can grant anything. The machine's owner has already said what a run may do there, your app asks for something, and the run gets the overlap, never more than the owner allowed, never more than your app asked for. The second line settles which compute serves it the same way.

A run that comes out empty on either line is refused before a process starts. This is the whole security model, and it is decided on the machine rather than by Soba for one reason: it has to hold even if Soba is lying.

4. Output streams back. Every runtime is normalised onto five event types. The first event names the tier that won, before any output. A cheaper tier covering for a sleeping laptop will visibly underperform, and a visible downgrade beats a silent one.

5. Your tools, if you defined any. The model's tool call comes back to your app, your app executes it, and the result goes the other way. The code and the data stay on your side throughout. See App-defined tools.

6. A question, if a tool needs consent. Only for tools inside the machine's askable set. Anything outside that set is refused without asking anyone, because asking would imply it could be allowed. See Approvals.

7. The run ends. Carrying its usage: the model, the provider, the token counts, and for metered classes the cost in micros and the cost class that produced it. A usage record is self-describing, so a meter never has to remember what it routed to in order to know whose money was spent.

What this leaves out#

Two of those steps stand in for a page each, and neither is summarised well by one paragraph:

  • What the machine spawns at step 3. An agent runtime (claude-code, codex) runs its own loop, so the resolved policy is passed to it and is advisory; a model runtime is driven by Soba's own loop, so the same policy is enforced in-process, immediately before the tool runs. Agents and models.
  • How the machine was picked at step 2. The workers that user has connected, a preference order across tiers, and the cost classes your app declared it would accept. Routing a run.

Your app reaches all of this over the endpoint or /v1/runs.

© 2026 Soba resolved = machine grant ∩ broker request