Availability Multiplies Along a Synchronous Chain

Availability Multiplies Along a Synchronous Chain An architecture diagram generated by Archify. Request · Architecture component Request Service A · 99.9% · Architecture component Service A 99.9% Service B · 99.9% · Architecture component Service B 99.9% Service C · 99.9% · Architecture component Service C 99.9% Service D · 99.9% · Architecture component Service D 99.9% Service E · 99.9% · Architecture component Service E 99.9% Service F · 99.9% · Architecture component Service F 99.9% Served · 99.4%, about 4h/month · Architecture component · topology alone Served 99.4%, about 4h/month topology alone Legend Backend External

The arithmetic

  • • 0.999 to the sixth power is about 99.4%
  • • Roughly four hours of monthly failure from topology alone
  • • The fix is not making each service more reliable

Tails dominate

  • • Six hops at a 10ms median look fine
  • • A p99 of 200ms somewhere in the middle becomes your p50 under load
  • • Tail latency compounds along a chain in a way medians do not predict

The real measure

  • • The benefit is deploys not blocked on another team
  • • If that number is not improving, the split is not working
  • • Stop making the chain synchronous before adding reliability