Article ·
Handling 218,000 requests on 2 vCPUs: backend architecture under constraints
When hardware is limited, architecture can no longer hide behind extra servers. A look back at a real-time voice platform that held the load on a very small server.
The context
The project: a real-time multiplayer game organised in rooms, with built-in voice chat and voice permissions that change with each player's role. The infrastructure: a VPS with 2 vCPUs and 2 GB of RAM. The result: 218,000 requests handled in 3 hours.
On average that is about twenty requests per second — but the average hides what matters. In a real-time game, load comes in waves: everyone joins a room at the same time, everyone speaks at once. Those peaks are what you design for.
Decision 1: separate signalling from media transport
A web voice system does two very different things. Signalling — who joins which room, who may speak, who leaves — is made of small, frequent messages. Media — the audio streams themselves — is a heavy, continuous volume.
Mixing them forces every component to be good at both. I separated them: signalling goes through Socket.io, and audio streams through WebRTC, relayed by a mediasoup SFU (Selective Forwarding Unit). Each layer can then be sized for what it actually does.
Decision 2: forward rather than mix
An SFU does not decode or mix audio: it selectively forwards streams to the relevant participants. The server therefore does far less computation than an architecture mixing every voice server-side. On 2 vCPUs, that difference is decisive.
Decision 3: the room as the unit of isolation
Each room is an independent perimeter: its participants, its streams, its permissions. An activity spike in one room does not spread to the others, and the state to manage stays small and local.
Voice permissions (RBAC) also play a performance role: only players allowed to speak produce an audio stream. Fewer streams produced means fewer streams to forward.
Decision 4: frugality everywhere else
On 2 GB of RAM, every connection and every byte counts. Nginx in front, Node.js processes supervised by PM2, and one simple rule: do nothing server-side that can be avoided — no needless processing, no superfluous state, no message sent to those who do not need it.
What carries over to other systems
- Constraints are the specification. Available hardware is not an obstacle to work around; it is an input of the architecture.
- Separate what has different load profiles. Short messages and heavy streams, reads and writes, sync and async.
- Isolate to contain. Splitting into independent units limits the spread of spikes and failures.
- Design for peaks, not for the average.
- Measure. An architecture is judged by what it survives in real life, not on a diagram.
These principles apply as much to a multi-tenant SaaS or an AI inference service as to a game. Modest infrastructure simply forces you to apply them earlier — which is often a good thing.
Go further: my work as a backend architect and the diagram of this architecture.