Nexorith
News and analysis from the world of production systems.
Multi-Region Failover Planning
July 29, 2026
Multi-region failover is mostly decided before the incident. The two questions that matter - how fresh does the standby data need to be, and who is allowed to press the button - sound managerial, but they drive every technical choice downstream from replication topology to health-check placement.
Data is the long pole. Application tiers scale horizontally and redeploy anywhere in minutes; a two-hundred-gigabyte database does not. Asynchronous replication buys you availability at the price of a recovery-point gap, and knowing your actual tolerance for that gap - in minutes, in euros - changes which fancy technologies are even admissible.
Structuring DNS for Reliability
September 22, 2026
Domain resolution outages produce an awkward paradox: server instances remain perfectly healthy while clients perceive complete downtime. Insulating against platform-level registrar outages demands secondary authoritative DNS providers operating over separate routing infrastructure.…
A Field Guide to Graceful Degradation
July 4, 2026
Every system has a sequence in which its features should die. Recommendations fail before checkout; search suggestions fail before search; thumbnails fail before the image. Writing that order down - and enforcing it with dependency-aware timeouts and bulkheads - is what separates a partial outage fr…
Cache Invalidation Patterns That Survive Traffic
May 13, 2026
While everyone chuckles at the difficulty of cache invalidation, systems routinely fall victim to stampedes when popular keys expire. Request collapsing represents the most effective safeguard: a single worker refreshes the expired entry while other concurrent threads consume slightly outdated value…
What Good Observability Actually Looks Like
August 11, 2026
Monitoring consoles sprawl uncontrollably while offering little insight during live incidents. True observability operates under inverted priorities: an on-call engineer gets paged, and telemetry systems must identify the root diff within sixty seconds.…
More reading
- A Practical Guide to API Rate Limiting — Engineering, May 18, 2026
- HTTP/3 and QUIC: What Changed for Operators — Networking, April 27, 2026
- Managing Secrets Without Losing Sleep — Security, June 7, 2026
About us
Our contributors have spent years on-call for large platforms. This site collects the playbooks, postmortems and reference material we wish someone had handed us earlier.