Your 3 AM outage may be your customer’s working morning.
Consider a service that fails while the local engineering team is asleep. SRE receives the page and starts investigating. The infrastructure looks healthy, but understanding the failure requires knowledge of the application that the responding engineer does not have. The people who built it have no agreed after-hours support responsibility.
Meanwhile, an international customer is trying to start their day. A feature they depend on is unavailable. Client services takes the calls and absorbs the frustration while SRE tries to work out what it can safely do.
Without a change in ownership, this can become a recurring pattern. The people taking the page feel unsupported, customer-facing teams bear the complaints, and engineering carries on with its next set of priorities. The same service eventually brings everyone back to the same conversation.
The ownership problem usually started well before that page.
Ownership Starts While the Service Is Being Built
Ownership becomes difficult when SRE is expected to support a service it had no input into building.
The development team chooses the framework, how the application will run, and where it will be hosted. Those decisions shape the work SRE will inherit. If SRE is brought in only at handoff, it loses the opportunity to question the operating assumptions while changing them is still relatively straightforward.
A weak handoff can also leave basic questions unanswered: expected API demand, memory requirements, CPU usage, or how the service behaves under load. SRE is left to figure those things out while being held responsible for keeping it available.
If a team will be expected to operate a service, it needs a voice in how that service is built. It needs enough evidence to judge whether the design can support the demands the business intends to place on it.
Google’s SRE engagement model describes involvement throughout the service lifecycle, including architecture and development. Production experience can improve a service before anyone picks up its pager.
Taking the First Page Does Not End Engineering’s Responsibility
SRE may be the first line of response. That arrangement can work when the engineering team remains available to help with the service knowledge and changes that recovery requires.
SRE cannot act on knowledge it was never given. A missing runbook or unfamiliar failure mode becomes much more serious when nobody with deeper application context is reachable.
The team that builds the service still has a responsibility to understand how it behaves in production. That responsibility includes helping during incidents and addressing the underlying problems afterward. A handoff to SRE does not make those obligations disappear.
The exact division of work will vary. What matters is whether the organization can reach the expertise it needs and get reliability work prioritized. Naming SRE as the owner is easy. Giving it a workable relationship with service engineering requires leadership to make those expectations explicit.
Growth Can Expose an Old Arrangement
A team may grow around one DevOps person who handles almost everything operational. Developers concentrate on building features because that is the arrangement the company established.
As the company grows, that person or team gets asked to support more services, more customers, and more complexity. Eventually, engineering is asked to take on production responsibilities it has never had before.
Engineers may question why the expectation has changed. They have never had to do this, it was never in their job description, or nobody told them they needed that knowledge. They may also be protective of their time, particularly when the new expectation extends into nights and weekends.
Leaders need to take those responses seriously. Writing application code and supporting it under production pressure require different experience. An engineer may be willing to help while being unsure how to do so safely.
Leaders need to understand what expectations they have set, what the team knows, and whether its workload allows it to take on production support. Simply announcing that everyone now owns production leaves those questions unanswered.
Small Teams Rely on Resourcefulness
Small teams often depend on people willing to step beyond their usual responsibilities. An engineer helps investigate an unfamiliar problem, a teammate shares what they know, and people find a way forward together. That resourcefulness and commitment can make a real difference while a startup is finding its footing.
Limited budgets and headcount can make formal support coverage difficult. Teams work with the people and experience they have. In that setting, it helps to be clear about who can respond and where they can turn for help. When SRE is expected to handle a service without engineering backup, it needs the knowledge and context to do so.
As the company grows, its support arrangements need room to grow too. More customers, services, and commitments create new demands on the team. Leaders can revisit how production work is shared and make space for more people to develop the experience it requires.
The goal is to preserve that willingness to help while giving people support they can depend on. Google’s guidance on being on-call highlights workload, engineering time, and compensation as parts of a sustainable arrangement. Resourcefulness remains valuable as the business builds the capacity to keep its customer commitments.
Explain What the Running Application Means to the Business
Bring the conversation back to the application itself. In development and in production, it remains part of the engineering team’s responsibility. It is how the business delivers value to customers.
Service commitments matter because customers build their own work around them. Keeping those commitments helps earn trust. Repeatedly leaving customers unable to use an important service puts that trust at risk and places more pressure on the people answering their calls.
Leaders need to explain that connection clearly. They also need to make their priorities reflect it. A team cannot consistently carry production responsibility if all of its time and attention are already committed to new features.
The support arrangement should fit the company’s stage, its customers, and the commitments it has made. Within that arrangement, SRE needs a voice, engineering needs to remain involved, and leadership needs to make the necessary decisions about capacity and priorities.
You cannot outsource reliability to the SRE team. The responsibility for delivering a working service extends across the people who design it, build it, operate it, and decide what the business promises its customers.
If production support is collecting around SRE while service ownership remains unclear, a Gone Rogue DevOps Assessment can help identify where responsibilities, capabilities, and business expectations have drifted apart.