Reliable backends in practice
Concrete failure models, solutions, tradeoffs and reproducible examples. No employer names, no unsupported scale claims. Just the engineering.
Articles are published every couple of weeks. RSS
Coming first
Transactions and consistency
The request timed out. Did the operation fail?
Why a timeout is uncertainty, not rejection, and how stable identities and explicit outcomes make retries safe.
Messaging
Acknowledged is not completed
Delivery guarantees and business-effect guarantees are different things. Where the acknowledgment boundary belongs.
Concurrency
More workers can make a queue slower
Bounded pools, admission control and why recovery traffic must share the same budget.
Real-time systems
A scalable WebSocket service still has state
Socket ownership, discovery, slow clients and cross-node cleanup when you add replicas.
Consistency
Version numbers protect state, not every side effect
Stale-write protection does not deduplicate notifications or guarantee downstream delivery.
Integration
Put provider differences at the boundary
A common interface only helps if it preserves identity, routing, authentication and error semantics.
Working on one of these problems right now?
The articles explain the technique. On a call we can apply it to your system.
Book a call