Practices: support and operations
Management practices are organizational resources and capabilities for getting work done, drawing on all four dimensions. This post covers the run-the-service five; the change-and-control set is in the next one.
- Incident management: minimize incident impact, restore service fast. An incident is an unplanned interruption or quality drop; swarming is a multi-discipline pile-on for complex incidents; a major incident is high impact demanding immediate coordinated action.
- Problem management: cut incident likelihood and impact by chasing causes. A problem is a cause, or potential cause, of incidents; a workaround is a temporary impact reducer; a known error is an analyzed problem without a permanent fix yet.
- Service desk: the entry point and single point of contact for users, capturing support demand. A moment of truth is any episode where a user forms an impression; service empathy is the ability to understand the user experience.
- Service request management: handle predefined user-initiated requests smoothly. A service request is pre-agreed normal delivery (password reset, software access); the request catalogue is the user-facing menu.
- Monitoring and event management: observe, analyze, and respond to state changes. An event is any significant change of state; an alert is a threshold reached or a failure notice.
The definitional splits are the exam meat: incident (unplanned interruption) vs problem (cause of incidents) vs service request (planned, pre-agreed). And workaround vs known error vs resolution.