Nexorith

Nexorith

News and analysis from the world of production systems.

Operations

Multi-Region Failover Planning

July 29, 2026

Multi-region failover is mostly decided before the incident. The two questions that matter - how fresh does the standby data need to be, and who is allowed to press the button - sound managerial, but they drive every technical choice downstream from replication topology to health-check placement.

Data is the long pole. Application tiers scale horizontally and redeploy anywhere in minutes; a two-hundred-gigabyte database does not. Asynchronous replication buys you availability at the price of a recovery-point gap, and knowing your actual tolerance for that gap - in minutes, in euros - changes which fancy technologies are even admissible.

Continue reading →

Networking

Structuring DNS for Reliability

September 22, 2026

Domain resolution outages produce an awkward paradox: server instances remain perfectly healthy while clients perceive complete downtime. Insulating against platform-level registrar outages demands secondary authoritative DNS providers operating over separate routing infrastructure.…

Engineering

A Field Guide to Graceful Degradation

July 4, 2026

Every system has a sequence in which its features should die. Recommendations fail before checkout; search suggestions fail before search; thumbnails fail before the image. Writing that order down - and enforcing it with dependency-aware timeouts and bulkheads - is what separates a partial outage fr…

Data

Cache Invalidation Patterns That Survive Traffic

May 13, 2026

While everyone chuckles at the difficulty of cache invalidation, systems routinely fall victim to stampedes when popular keys expire. Request collapsing represents the most effective safeguard: a single worker refreshes the expired entry while other concurrent threads consume slightly outdated value…

Operations

What Good Observability Actually Looks Like

August 11, 2026

Monitoring consoles sprawl uncontrollably while offering little insight during live incidents. True observability operates under inverted priorities: an on-call engineer gets paged, and telemetry systems must identify the root diff within sixty seconds.…

More reading

About us

Our contributors have spent years on-call for large platforms. This site collects the playbooks, postmortems and reference material we wish someone had handed us earlier.

More about the project →