Enhancing Digital Platforms with SRE
Site reliability engineering applied to a live platform — error budgets, monitoring, and deciding what uptime is worth.
Begin with measurement — you cannot set an objective for something you do not monitor.
Reliability as a product feature
In the age of digital services, reliability is not a background concern. Organisations like Amazon handle millions of transactions around the clock, where a momentary failure costs an amount that justifies almost any preventative spend.
Most organisations are not Amazon, but the principle scales down. If your service is how customers transact with you, its availability is part of the product, not part of the infrastructure budget.
In the age of digital services, reliability is not a background concern.
Where SRE came from
Site Reliability Engineering originated at Google, from a deceptively simple proposition: treat operations as a software problem.
Rather than scaling a service by adding people to run it, SRE applies engineering to the operational work itself — automating it, measuring it, and holding it to explicit reliability targets. That is the whole idea, and everything else follows from it.
What SRE teams actually do
- Define reliability numerically — service level objectives, and an error budget derived from them
- Automate toil, the manual repetitive work that scales with traffic
- Own monitoring and observability, so failures are seen before customers report them
- Run incident response, including blameless post-incident review
- Gatekeep on evidence: when the error budget is spent, the priority becomes stability rather than features
The error budget is the mechanism that makes SRE more than a job title. It converts "should we ship this or stabilise?" from an argument into a calculation.
What to actually monitor
Monitoring everything is the same as monitoring nothing — too much signal buries the alert that matters. The four measures worth building dashboards around are latency (how long a request takes), traffic (how much demand the service is under), errors (the rate of requests failing), and saturation (how close a resource is to its limit). Between them, they tend to catch a failing service well before a customer notices, which is the entire point of watching in the first place.
How teams are organised
There are two workable models. An embedded SRE sits within a product team and shares its priorities. A central SRE team builds the platform, tooling and standards that product teams use.
Larger organisations generally end up with both, and the failure mode is the same in either: an SRE team that becomes the operations department under a new name, absorbing all the toil instead of engineering it away.
Starting an SRE journey
Begin with measurement — you cannot set an objective for something you do not monitor. Then set one SLO on one important service and let the organisation experience what it means to hold to it.
From there, attack toil in order of how much time it consumes, and establish blameless post-incident reviews early. Both are cultural changes as much as technical ones, and both are much harder to introduce once the team is firefighting full time.
Get the full guide
Tell us who you are and the guide opens straight away, all 9 pages, to read online or download as a PDF.
Download the guide
What housing providers say
Named people at named organisations, in their own published words.
“We realised that we had a gap around the golden thread of data, in terms of the availability and accessibility of the data we held in seventeen different systems… That meant colleagues could immediately see all non-compliant properties.”
Jake Le Page Head of Building Safety Regulations Notting Hill Genesis
“Neo brought strong Dynamics 365 expertise, worked collaboratively with our internal teams and applied Microsoft best practice within a live operational environment… We would be pleased to recommend them as a Microsoft Dynamics 365 partner within the housing sector.”
Wayne Human Head of IT Change Sage Homes
“The successful deployment of this solution has significantly increased transparency and improved the operational efficiency of our contact centre, allowing us to deliver greater value to our customers.”
Philip Wragg Infrastructure Programme Manager
“The Neo Technology model allows us to scale our development capacity, accelerating our transformation programmes while future-proofing our business, while achieving substantial industry cost savings.”
Group CIO Notting Hill Genesis
Related proof
See how this works in practice
Thirty minutes, walked through by the people who deliver it. Bring your questions.