Clarify the task and decision
This guide turns sla and incident response into a reviewable operating workflow. It connects domain decisions, ownership, evidence, and acceptance so the result continues to work in production.
Define measurable availability, priority levels, response and restoration targets, communication duties, and remedies.
Practical workflow
- 1
Draw the data flow from collection through processing, logs, cache, support, backup, and deletion.
- 2
Classify each data category and connect it to purpose, legal role, location, recipient, and retention.
- 3
Request architecture, contractual, operational, and testing evidence for every material claim.
- 4
Score risk and record required controls, owners, acceptance evidence, and residual risk.
- 5
Approve only the documented configuration, then monitor subprocessors, incidents, changes, and deletion evidence.
Worked example or tool
An incident table aligns user impact, acknowledgement, update frequency, restoration, and post-incident review. In the tool, also record the baseline, owner, decision, evidence, open issue, and approval date. Use a real page or transaction so the team sees dependencies, exceptions, and the maintenance work that follows release.
| Decision point | Record | Acceptance criterion |
|---|---|---|
| Baseline | Observed current state | Source and date recorded |
| Decision | Selected option and rationale | Risk and audience considered |
| Evidence | Test, document, or measure | Reviewable and version-specific |
| Approval | Name, role, and date | All mandatory criteria met |
Define a measurable service boundary
State which production endpoints, user interfaces, batch jobs, and dependencies the availability target covers. Define successful service from the customer's perspective, including correct authentication and usable responses, rather than counting a server that returns errors as available.
Use one published formula: available minutes divided by scheduled service minutes, after only agreed exclusions. Set the measurement source, time zone, reporting interval, rounding method, maintenance rules, and treatment of partial degradation before comparing percentages.
Name every covered component and critical journey.
Define failure and degradation objectively.
Fix the measurement source and calculation method.
Limit exclusions and planned-maintenance windows.
Set priorities, clocks, and restoration targets
Define incident priority by user impact, data risk, reach, and workaround, not by the supplier's internal technical label. Distinguish acknowledgement, qualified response, mitigation, restoration, and permanent correction because each represents a different outcome.
Specify when clocks start, pause, and stop. Include nights, weekends, notification channels, customer dependencies, and escalation. For critical public services, pair supplier response targets with internal recovery objectives and a tested manual or alternative route.
- 1
Create examples for every priority level.
- 2
Set acknowledgement, update, mitigation, and restoration targets.
- 3
Name customer and supplier escalation roles.
- 4
Test the process in a timed exercise.
Communicate clearly and learn from incidents
An incident update should state confirmed impact, affected functions and regions, start time, current action, workaround, next update time, and contact. Separate facts from hypotheses. Keep a stable incident identifier across status page, email, support, and final report.
Require a post-incident report for severe events with timeline, contributing conditions, detection gap, containment, recovery, customer impact, corrective actions, owners, and due dates. Review whether actions reduce recurrence or only improve the wording of future reports.
Use a pre-approved update template.
Publish updates at the promised cadence.
Track corrective actions to verified completion.
Share relevant lessons with service owners and users.
Govern remedies and service evidence
Service credits should be automatic or easy to claim and should scale with impact, but they are not a substitute for resilience. Reserve stronger remedies for repeated failure, missed security duties, prolonged outage, or failure to complete corrective actions.
Review a monthly evidence pack containing raw availability, excluded minutes, incidents, target performance, recurring causes, support demand, changes, and capacity risks. Compare supplier data with customer monitoring. Use trends to trigger improvement plans, architecture changes, or exit preparation.
- 1
Reconcile supplier and customer measurements monthly.
- 2
Challenge every exclusion with evidence.
- 3
Apply credits and escalation consistently.
- 4
Trigger improvement or exit thresholds automatically.
Prepare the joint response before an incident occurs
Create a joint response matrix that maps incident types to supplier and customer duties. Cover availability failure, corrupted output, unauthorised access, suspected personal-data breach, lost audit records, model regression, excessive latency, quota exhaustion, and failed batch delivery. For each event, specify detection source, initial triage owner, evidence to preserve, authority to disable the service, regulatory assessment owner, communication approver, recovery route, and return-to-service criteria.
Synchronise operational and legal clocks. A supplier's critical support target does not replace statutory or contractual notification duties. The customer needs enough verified information to assess affected data, people, systems, time period, containment, likely consequences, and mitigation. Require the supplier to provide rolling facts as the investigation develops rather than waiting for a final report, while clearly marking uncertainty and subsequent corrections.
Exercise the matrix with realistic injects at least annually and after major architecture changes. Include unavailable contacts, incomplete logs, disagreement about priority, a public enquiry, and a failed first recovery attempt. Record decision times, missing information, manual workarounds, user impact, and improvement actions. The exercise should test leadership and communications as well as technical restoration. Close every action with evidence and update the SLA if the documented process cannot meet its own deadlines.
Map technical, privacy, security, service, and communication responsibilities.
Align supplier response targets with legal and organisational clocks.
Exercise incomplete information, escalation, workaround, and failed recovery.
Verify improvement actions and revise unachievable commitments.
Roles, evidence, and approval
Security and privacy reviews must describe the production configuration rather than a generic provider. Record the exact service, region, feature flags, optional telemetry, support access, subprocessors, encryption boundaries, retention settings, and customer responsibilities. Reassess after material architecture, contract, provider, or purpose changes and keep the decision linked to the evidence reviewed.
Operations and maintenance
The work does not end at publication. Link the language version or configuration to its source, monitor quality and service measures, and define concrete review triggers. Triggers include source changes, legal changes, new audience needs, recurring support questions, technical changes, and incidents. A named owner evaluates the trigger, opens a new revision when needed, and records renewed approval.
Release checklist
The complete processing path is documented.
Controller and processor roles are agreed.
Locations and subprocessors are evidenced.
Training and secondary use are explicitly addressed.
Access, encryption, logging, and incident controls are verified.
Retention and deletion are defined per data category.
International transfers and safeguards are documented.
Changes, audits, exit, and evidence ownership are assigned.