Why the swarm sends a whole request to one device
Published
The popular picture of a compute swarm is one large model split across many machines. That picture comes from datacentres, where the links between accelerators were built for exactly this traffic. A swarm of devices people already own has no such links. When the design for this project was written down, the first decision was to stop pretending it did.
The arithmetic that killed model parallelism here
Splitting one forward pass across devices means exchanging activations between the layers that live on different machines. The architecture document works through that exchange on its own numbers and calls them what they are: a design calculation, not a measurement. No real model has run in this project, and no weights are loaded, so none of it is a benchmark.
- For a 7B model at hidden size 4096 in bf16, the architecture document of the swarm specification estimates roughly 8 KB of activation exchange per token at every device boundary, over the 20 to 60 ms round-trip time it states for a consumer link. That is the spec estimate, not a measurement, and the waiting rather than the compute becomes the cost.
- The same spec estimate for a 70B model lands above about 32 KB per layer step, and the architecture document calls that unusable over consumer links.
- Splitting one model therefore buys latency that gets worse as more devices are added. The document is blunt about it: a swarm that promises one model across thousands of phones, only faster, is describing a datacentre interconnect on hardware that does not have one.
What the unit of work actually is
The unit of work is a whole request: a completion, a classification, an embedding, dispatched to one device that already holds that model. The coordinator routes an encrypted envelope and relays the encrypted result, and the architecture document is careful about what that costs. The coordinator sees routing metadata such as the job id, the recipient, the size, a TTL and the model class. It does not see the prompt, the model output or the result, because those travel in an envelope it cannot open.
The routing fields are inside the authenticated data of the envelope rather than beside it. That is a design choice with a nameable consequence: a coordinator that edits the recipient or the expiry to help itself breaks the job instead of quietly succeeding.
Why one device can take a whole request
- A job goes only to a device that already holds the model, so a scheduling decision cannot start a multi-gigabyte download on a machine that may leave the network first.
- Free memory is a hard filter rather than a score. A model that does not fit fails at load time, after the job was accepted, so the memory floor is checked before dispatch.
- Consumer devices sleep, throttle and disconnect, so leases are short, expiry causes a re-dispatch to another device holding the same model, and the job id makes a retry distinguishable from a replay.
- Capability epochs revoke older claims and key material without needing a certificate authority.
The three things the design picked instead
Once the forward pass is off the table, three patterns remain that survive a consumer link. They are worth naming separately, because they are not three flavours of the same idea.
- Independent jobs on independent devices: the unit this project actually specifies.
- Replication for reliability and verification: run the same whole request elsewhere and compare, which is what the settlement plane pays for.
- Verify-then-trust on selected results: spend the second run only where the answer is worth checking.
The device tiers make the same point from the other side. Phones and tablets are treated as small-model workers that are expected to leave mid-job, laptops as the mainstream worker handing a small model to an integrated GPU, and workstations as the reliable tier that usually does the verifying. A design that depends on any single one of those behaving like a datacentre would be broken on all three.
What the swarm does buy
Capacity, jurisdiction and resilience, in that order, and not latency on a single request. Those are three different products from the one the popular picture describes, and saying so early is what keeps the rest of the design honest.
What has actually been run is a mock backend over real HTTP, and the status notes it plainly: no real model inference and no real GPU ran anywhere. The phase after this one is an llmeu.com page that states what was measured, an entry point for contributing a device, and a content surface generated from the model registry rather than from hand-written claims.
Surse
Each link goes to the source the post cites. Where a source does not state a fact, the post shows it as unverified instead of filling it in.