Google Site Reliability Engineering Book
Free online, and the clearest account of what running large systems actually involves: error budgets, service level objectives, on call, postmortems and the economics of reliability. Reliability thinking is what makes a design answer sound like it came from someone who has operated something.
Book
System Design
1 to 2 months
Advanced
Free
What it is
The book that defined site reliability engineering as a discipline, published free online by the team that invented the role.
The ideas worth taking
Service level objectives and error budgets, which reframe reliability as a number you spend rather than a virtue you maximise. Blameless postmortems. The argument that toil is a measurable quantity you should be reducing. Load shedding and graceful degradation. These ideas transfer far beyond Google.
How to use it
The early chapters on service level objectives, error budgets and monitoring are the highest value and the most transferable. The later chapters are more specific to Google's internal scale and can be skimmed.
In an interview
Mentioning an error budget or a graceful degradation strategy in a design round signals operational experience more efficiently than almost anything else you can say.
Best for: backend, platform and infrastructure engineers, and senior design candidates.
Ready to start?
Opens on Google in a new tab.
Stuck on something specific?
Writing only gets you so far. If you want an answer to your situation rather than the general case, book a session and we will work through it together. Every session is free; a few slots open each week.
Follow along
New writing, resources and project ideas land here first.