• Irreversible Decisions
    • Constraints First
    • Coupling and Cohesion in Practice
    • Trade-off Sliders
    • The Boring Baseline
    • Monolith First, Split on Evidence
    • Service Boundaries That Survive Reorgs
    • Synchronous vs Asynchronous Integration
    • API Contracts and Versioning
    • Idempotency and Retries
    • Choosing a Datastore Without Regret
    • Schema Migrations Without Downtime
    • Consistency Models You Actually Need
    • Caching: The Four Questions
    • Event Sourcing and CDC: When It Pays
    • Trunk-Based Development and Branch Reality
    • Pipeline Design: Fast Feedback, Slow Gates
    • Progressive Delivery
    • Build Reproducibility and Artifact Promotion
    • Rollback Is a Feature
    • Kubernetes: What You Sign Up For
    • Infrastructure as Code That Doesn't Drift
    • Environments, Config, and Secrets
    • Multi-Tenancy and Cost Boundaries
    • The Internal Platform as a Product
    • SLOs, Error Budgets, and Saying No
    • Observability: What to Actually Wire
    • Capacity, Load Shedding, and Backpressure
    • Failure Modes: Timeouts, Breakers, Bulkheads
    • Incident Response and Blameless Postmortems
    • Threat Modeling in One Hour
    • Identity, Authentication, Authorisation
    • Secrets and Key Management
    • Supply Chain: Dependencies, SBOM, Signing
    • Least Privilege and Auditability
    • Architecture Decision Records That Get Read
    • Diagrams That Age Well
    • Runbooks and On-Call Docs
    • Design Reviews and RFCs
    • Building a Team Architecture Memory
    • GitHub
  • to navigate
  • to select
  • to close
    • Home
    • Reliability and Operations
    On this page
    monitor_heart

    Reliability and Operations

    Keeping it up, knowing what up means, and being useful at 3am: SLOs, observability, failure modes, incidents.

    speed

    SLOs, Error Budgets, and Saying No

    An SLO is not a reliability target. It is a negotiating instrument that makes the cost of reliability visible.

    visibility

    Observability: What to Actually Wire

    The goal is answering questions you did not anticipate. Most 'observability' spend buys dashboards for questions you already knew.

    compress

    Capacity, Load Shedding, and Backpressure

    A system without limits does not degrade under overload. It collapses, and then it stays down.

    warning

    Failure Modes: Timeouts, Breakers, Bulkheads

    Three patterns, routinely misconfigured in the same three ways. The defaults in most libraries are wrong for you.

    emergency

    Incident Response and Blameless Postmortems

    The goal during an incident is restoring service. The goal after is changing the system, not the person.


    © 2026 Architecture Field Notes. Built with Lotus Docs