Operations & support · reference function

Cloud Operations

Operate supported cloud services within clear reliability and support boundaries.

Reference mandate

Own or govern the run capability for accepted cloud platform services, including monitoring, incident, change, recovery and operational evidence.

Short

Operate supported cloud services within clear reliability and support boundaries.

Standard

Own or govern the run capability for accepted cloud platform services, including monitoring, incident, change, recovery and operational evidence.

Detailed

Own or govern the run capability for accepted cloud platform services, including monitoring, incident, change, recovery and operational evidence. The function maintains explicit decision boundaries, measurable outcomes, governed interfaces and a documented improvement loop for its scope.

Scope in

  • Monitoring and response
  • Incident and problem execution
  • Operational change
  • Recovery evidence

Scope out

  • Product roadmap ownership
  • Architecture policy
  • Accepting incomplete transition by default

Responsibilities

  • Monitoring and response
  • Incident and problem execution
  • Operational change
  • Recovery evidence

Services

  • Monitoring and response service
  • Incident and problem execution service
  • Operational change service
  • Recovery evidence service

Required capabilities

  • Cloud Operations Lead capability
  • Site Reliability Engineer capability
  • Major Incident Manager capability

Roles

  • Cloud Operations Lead
  • Site Reliability Engineer
  • Major Incident Manager

Decision rights

  • Declare major incident
  • Approve operational change
  • Accept operational risk within mandate

Key interfaces

  • Platform Engineering
  • Cloud Service Management
  • Vendor & Commercial Management

Inputs

  • Requirements from Platform Engineering
  • Requirements from Cloud Service Management

Outputs

  • Governed output to Cloud Service Management
  • Governed output to Vendor & Commercial Management

Governance forums

  • Operate design review
  • Cloud operating-model review

Measures

  • Availability
  • Mean time to restore
  • Recurring-incident reduction

Dependencies

  • Platform Engineering
  • Cloud Service Management
  • Vendor & Commercial Management

Sourcing options

  • Retained internal
  • Shared
  • MSP-supported

Organizational placements

  • Central cloud organization
  • Federated domain
  • Shared technology function

Decisions owned

  • Declare major incident
  • Approve operational change
  • Accept operational risk within mandate

Decisions contributed to

  • Contribute to decisions owned by Platform Engineering
  • Contribute to decisions owned by Cloud Service Management
  • Contribute to decisions owned by Vendor & Commercial Management
Design boundary

Common failure modes

  • Receives work without acceptance evidence
  • Provider accountability is unclear
  • Operational feedback never changes the platform
Sourcing principle

Run execution may be outsourced; internal service outcome and risk accountability remain.

Minimum retained capability

Run execution may be outsourced; internal service outcome and risk accountability remain.

Maturity guidance

Begin with named ownership and one measurable outcome; add delegation and automation only when evidence and capability are reliable.

Reference note

Vendor-neutral practitioner reference pattern; validate against organizational, regulatory and sourcing context.

Version 2026.3 · reviewed 8 September 2026