Agentic AI Atlasby a5c.ai
OverviewWikiGraphFor AgentsEdgesSearchWorkspace
/
GitHubDocsDiscord
i.2Wiki
Agentic AI Atlas · Incident Management (Library)
library/incident-managementa5c.ai
Search the atlas/
Wiki · linked records

Article and nearby pages

I.Current articlepp. 1 - 1
accessibility (Library)Aerospace Engineering Specialization (Library)AI Agents and Conversational AI Specialization (Library)Algorithms and Optimization Specialization (Library)Arts and Culture Specialization (Library)ATDD/TDD Methodology (Library)
II.Documented nodesrefs · 1
specialization:incident-management
I.
Wiki article

library/incident-management

Reading · 8 min

Incident Management (Library) reference

Single owner for incident handling across the library. This specialization carries the flagship detection-to-postmortem lifecycle — severity classification, commander mobilization, parallel mitigation strands, severity-routed policy gates for every externally visible action, an adversarial postmortem-completeness gate, and kip-backed incident memory across runs.

Page nodewiki/library/incident-management.mdNearby pages · 135Documents · 1

Continue reading

Nearby pages in the same section.

accessibility (Library)Aerospace Engineering Specialization (Library)AI Agents and Conversational AI Specialization (Library)Algorithms and Optimization Specialization (Library)Arts and Culture Specialization (Library)ATDD/TDD Methodology (Library)authoring (Library)AutoMaker (Library)Automotive Engineering Specialization (Library)Backend Development (Library)BDD/Specification by Example (Library)Bioinformatics and Genomics Specialization (Library)Biomedical Engineering Specialization (Library)BMAD Method (Library)business/ (folded) (Library)Business Analysis and Consulting (Library)Business Strategy and Operations (Library)Business Strategy Specialization (Library)CC10X Methodology (Library)CCPM - Claude Code PM Methodology (Library)Chemical Engineering Specialization (Library)Civil Engineering Specialization (Library)ClaudeKit Methodology (Library)Cleanroom Software Engineering (Library)CLI and MCP Development Specialization (Library)Code Migration and Modernization Specialization (Library)COG Second Brain (Library)collaboration (Library)common-utilities (Library)Communication specialization (Library)Composition: Aerospace Flight Control (Waterfall + V-Model + Cleanroom + inline Formal Verification) (Library)Composition: Legacy Modernization (Event Storming + DDD + FDD + Strangler Fig + RUP) (Library)Composition: Open Source Data-Validation Framework (TDD + BDD + Kanban + XP + Continuous Deployment) (Library)Composition: Regulated Greenfield (V-Model + DDD + Cleanroom + Waterfall) (Library)Composition: SaaS Analytics Dashboard (JTBD + Impact Mapping + Spec-Kit + Kanban + XP) (Library)Composition: Smart Product Recommendations (DDD + Hypothesis-Driven Development + BDD + Kanban) (Library)Composition: Startup MVP (Shape Up + Example Mapping + TDD + Scrum) (Library)Computer Science Specialization (Library)Cryptography and Blockchain Development Specialization (Library)Customer Experience and Support Specialization (Library)customer-support (Library)Data Engineering, Analytics, and BI Specialization (Library)Data Privacy Compliance (Library)Data Science and Machine Learning Specialization (Library)Intelligence, Decision Support and Decision Making (Library)Desktop Product Development Specialization (Library)developer-relations (Library)DevOps, SRE, and Platform Engineering Specialization (Library)Digital Marketing and Content Strategy Specialization (Library)Domain-Driven Design (DDD) Methodology (Library)Double Diamond Methodology (Library)Education and Learning Specialization (Library)Electrical Engineering Specialization (Library)Embedded Systems Engineering Specialization (Library)Entrepreneurship and Startup Processes (Library)Environmental Engineering Specialization (Library)Event Storming (Library)Everything Claude Code Methodology (Library)Example Mapping Methodology (Library)Extreme Programming (XP) (Library)Feature-Driven Development (FDD) (Library)Finance, Accounting, and Economics Specialization (Library)FPGA Programming and Hardware Description Specialization (Library)Game Product Development Specialization (Library)Gas Town Methodology (Library)GPU Programming and Parallel Computing (Library)GSD-Adapted Workflows for Babysitter SDK (Library)Healthcare and Medical Management Specialization (Library)Human Resources and People Operations Specialization (Library)Humanities and Anthropology Specialization (Library)Hypothesis-Driven Development (Library)Impact Mapping Methodology (Library)Industrial Engineering Specialization (Library)internationalization (Library)Jobs to Be Done (JTBD) Methodology (Library)Kanban (Library)Knowledge Management (Library)Legal and Compliance Specialization (Library)Logistics and Operations Specialization (Library)Maestro App Factory (Library)Marketing and Brand Management Specialization (Library)Materials Science Specialization (Library)Mathematics Specialization (Library)Mechanical Engineering Specialization (Library)media (Library)Meta Specialization - Process, Skill, and Agent Creation (Library)Metaswarm Methodology (Library)MLOps (Library)Mobile Product Development Specialization (Library)Nanotechnology Specialization (Library)Network Programming and Protocols Specialization (Library)Observability specialization (Library)Enhanced Ontology-Driven Development (ODD) Methodology (Library)Operations Management Specialization (Library)Performance Optimization and Profiling Specialization (Library)Philosophy and Theology Specialization (Library)Physics Specialization (Library)Pilot Shell Methodology for Babysitter SDK (Library)Planning with Files (Library)Procurement — business domain specialization (Library)Product Management and Product Strategy Specialization (Library)Production contract (Library)Programming Languages and Compilers Development Specialization (Library)Project Management and Leadership Specialization (Library)Public Relations and Communications Specialization (Library)QA, Testing, and Test Automation (Library)Quantum Computing Specialization (Library)Release Engineering (Library)Research Specialization (Library)Robotics and Simulation Engineering Specialization (Library)RPIKit Methodology (Library)Ruflo Methodology (Library)RUP (Rational Unified Process) (Library)Sales and Business Development Specialization (Library)Scientific Discovery and Problem Solving Specialization (Library)Scrum (Library)SDK, Platform, and Systems Development (Library)Security, Compliance, and Risk Management Specialization (Library)Security Research and Vulnerability Analysis Specialization (Library)Shape Up (Library)Shared (Cross-Domain Assets) (Library)Social Sciences Specialization (Library)Software Architecture and Design Patterns Specialization (Library)sourcing/ (folded) (Library)Spec Kit Methodology (Library)Spiral Model (Library)Superpowers Extended Methodology (Library)Supply Chain Management Specialization (Library)Technical Documentation Specialization (Library)Travel (Curated-Dataset + SQL-Tool Pattern) (Library)UX/UI Design and User Experience Specialization (Library)V-Model Methodology (Library)Venture Capital and Investment Due Diligence Specialization (Library)Waterfall Methodology (Library)Web Product Development Specialization (Library)

Documented graph nodes

Records linked directly from this page’s Page node.

specialization:incident-management

Incident Management

Single owner for incident handling across the library. This specialization carries the flagship detection-to-postmortem lifecycle — severity classification, commander mobilization, parallel mitigation strands, severity-routed policy gates for every externally visible action, an adversarial postmortem-completeness gate, and kip-backed incident memory across runs.

Consolidation statement

The library census flagged three scattered near-misses that each owned a slice of incident handling. This specialization consolidates them under one flagship process:

  • **Supersedes** `specializations/devops-sre-platform/incident-response.js` — **deprecated, do not extend.** Harvested: commander/roles mobilization, parallel log/metrics/trace investigation (folded into the diagnosis strand), blast-radius fields, MTTR/MTTD metrics, Style-A task shape.
  • **Supersedes** `specializations/domains/business/customer-experience/itil-incident-management.js` — **deprecated, do not extend.** Harvested: categorization-informed severity rationale; the knowledge-base lookup is replaced by kipRecall, the post-incident-review loop by adversarialGate.
  • **Absorbs** `specializations/observability/incident-lifecycle.js` — structural seed: single-workflow lifecycle, severity matrix text, non-incident early exit, timeline accumulation, comms phase model with no-blame rules, iterative 3-pass diagnosis, SLO breach detection, follow-up issue creation.

The superseded files are **not deleted** by this consolidation — this README declares them deprecated pending a separate removal pass. New work goes here.

Migration notes

Old processOld inputsNew inputs mappingBehavioral deltas
devops-sre-platform/incident-response{ incidentType, severity, affectedServices, alertSource, description }signal: { source: alertSource, ref, firstSeenAt, symptomSummary: description, impactedSurfaces: affectedServices } + severityOverride: severitySeverity is classified by the process (the old hand-fed severity becomes severityOverride, which wins and is recorded in metadata). Production changes and all comms now sit behind severity-routed policy gates instead of running unguarded.
business/customer-experience/itil-incident-management{ incident, knowledgeBase }incident fields fold into signal{...}; knowledgeBase is replaced by the kip store (kipEnabled/kipDir/kipModel)KB lookup becomes kipRecall at detection; the post-incident-review loop becomes the adversarial postmortem-completeness gate with executed evidence.
observability/incident-lifecycle{ signal, onCall, commsChannels, slo }Same shape — onCall becomes oncall (adds commsLead, engineeringManager); adds severityOverride, customerFacing, postmortemRequired, kip knobsComms no longer publish directly: every status-page/customer message rides a policy gate. Postmortem gains the adversarial completeness gate and gated publication.

Module table — `incident-lifecycle.js` exports

ExportKindPurpose
process(inputs, ctx)orchestratorThe flagship lifecycle, phases P0–P10
SEVERITY_ROUTINGfrozen constPolicy-gate expert routing per severity (lookup via routingExpert)
COMMS_CADENCEfrozen constComms cadence rules per severity
REQUIRED_ROLESfrozen constRoles that must be staffed per severity
INCIDENT_SEVERITIESfrozen const['SEV1','SEV2','SEV3','SEV4']
routingExpert(actionId, severity)helperRouting lookup — **throws** on unknown action, unknown severity, or never-raised actions (no fallback expert)
commsPolicyFor(actionId, severity, customerFacing)helperCadence-table lookup — **throws** on unknown severity/action
detectClassifyTaskagent taskiml.detect-classify — severity classification
commanderAssignmentTaskagent taskiml.commander-assignment — roster mobilization
diagnoseTaskagent taskiml.diagnose — iterative root-cause passes
blastRadiusTaskagent taskiml.blast-radius — impact assessment
commsDraftTaskagent taskiml.comms-draft — drafts only, never publishes
mitigationPlanTaskagent taskiml.mitigation-plan — reversible-first plan
executeMitigationTaskagent taskiml.execute-mitigation — only after its gate approves
publishCommsTaskagent taskiml.publish-comms — only after its gate approves
verifyRecoveryTaskagent taskiml.verify-recovery — executed probe evidence
postmortemDraftTaskagent taskiml.postmortem-draft — blameless postmortem markdown
actionItemTrackingTaskagent taskiml.action-item-tracking — one issue per action item

All tasks are Style-A kind: 'agent' (zero kind: 'shell'), with per-effect io paths and labels, and every evidence-carrying output schema declares evidence { type: 'array', minItems: 1 }. Gate combinators (routedBreakpoint, adversarialGate, kipRecall, kipAssert) are imported from `../common-utilities/routed-gate-combinators.js`, not redefined.

Severity model

  • **SEV1** — user-facing outage on a critical path OR data loss/corruption risk.
  • **SEV2** — significant degradation with ongoing user impact (workaround may exist).
  • **SEV3** — partial degradation, limited user impact.
  • **SEV4** — internal-only, no user impact, cleanup-later.
  • **non-incident** — false alarm, duplicate, or expected behavior → early return before any commander is assigned or gate raised.

Routing table (`SEVERITY_ROUTING`, verbatim)

ActionSEV1SEV2SEV3SEV4
execute-prod-mitigationincident-commanderincident-commandertech-leadtech-lead
publish-status-pagecomms-leadcomms-leadcomms-lead (presentAlwaysApprove)not raised
send-customer-incident-commscomms-leadcomms-lead (only if customerFacing)not raisednot raised
publish-postmortemengineering-managerengineering-managertech-leadtech-lead (postmortem only when postmortemRequired===true)

Comms cadence table (`COMMS_CADENCE`, verbatim)

SeverityStatus pageCustomer comms
SEV1required, 30m updatesrequired
SEV2required, 60m updatesrequired iff customerFacing
SEV3discretionarynot required
SEV4not requirednot required

Lookups go through routingExpert(actionId, severity), which **throws** on an unknown severity, an unknown action, or an action that is never raised at that severity. There is deliberately no default expert — fallbacks are forbidden.

Policy-gated actions

Four actions are policy-gated. Convention: **breakpointId = actionId**, tags ['policy-gated', 'incident', '<sev>'] (severity tag interpolated per run), strategy single, expert from SEVERITY_ROUTING.

actionIdWhat it gatesRaised when
execute-prod-mitigationAny production change made to mitigate the incident (config change, rollback, failover, feature-flag kill)Always, before any production change. Never auto-approves (autoApproveAfterN is never set). Rejection → one re-plan pass; second rejection ends the run success:false with the executor never invoked.
publish-status-pagePublishing/updating the public status page entrySEV1/SEV2 required; SEV3 discretionary (presentAlwaysApprove:true); SEV4 never
send-customer-incident-commsDirect customer notifications about impact, workarounds, or resolutionSEV1 required; SEV2 required iff customerFacing; SEV3/SEV4 never
publish-postmortemPublishing the blameless postmortem to its audienceOnly after the adversarial completeness gate passes (or the owner accepts via escalation)

**Fail-closed posture:** a rejected or never-raised gate leaves its action unexecuted and unpublished — there is no alternate path around a gate. Every not-raised comms action is still recorded in commsLog as { required: false, raised: false }. Any harness-level auto-approval of a gate is surfaced in outputs.autoApprovals ({ breakpointId, phase, at }), which is **always present** in outputs, possibly empty.

The only non-policy breakpoint is iml.commander-assignment.roster-gap, raised solely when severity-required roles are unstaffed (sparse-breakpoint rule: an unstaffed incident is genuinely blocking).

Quality gate — postmortem completeness

adversarialGate(ctx, { gateId: 'iml.postmortem-completeness', ... }) fans out three independent critics over the postmortem draft (the postmortem author never reviews its own work):

CriticFocus
timeline-fidelity-criticEvery timeline entry in the postmortem matches the run journal and task timestamps — the comparison is EXECUTED (read both, diff them), with each verified/mismatched entry cited. The orchestrator accumulates the timeline itself, so the critic has ground truth to diff.
action-item-completeness-criticEvery action item has a named owner AND concrete due date AND category; every "what went poorly" finding maps to at least one action item — exact lines cited.
blameless-depth-criticContributing factors go beyond the proximate cause (systemic depth, 5-whys); language blameless; root cause stated with supporting evidence.

IRON-LAW rules (appended to every critic prompt): executed evidence only — run the actual cross-checks, a read-only skim is not evidence; cite file:line or executed-check output for every claim; passed:true with empty evidence is rejected by the combinator. Fix budget: maxFixAttempts (default 2) rounds of the default gateFixerTask between critic rounds; on exhaustion the combinator escalates to the owner via a routed breakpoint (iml.postmortem-completeness.gate-escalation).

kip incident memory

- { subject: 'incident:<incidentId>', predicate: 'has-signature', object: <signatureString> } - { subject: 'incident:<incidentId>', predicate: 'root-cause', object: <rootCause> } - { subject: 'incident:<incidentId>', predicate: 'mitigated-by', object: <planSummary>, props: { efficacy: 'effective'|'ineffective', reversible } } - one per action item: { subject: 'incident:<incidentId>', predicate: 'action-item', object: <title>, props: { owner, dueDate, category } }

  • **Recall at detection (P0)**: kipRecall(ctx, { kipDir, topic: 'incident signature: <symptomSummary> [<impactedSurfaces>]', kipModel, kind: 'incident-management' }) — prior incidents with matching signatures, known-good mitigations, and prior action-item outcomes are threaded into classification, diagnosis, and mitigation prompts. An empty store is initialized and reported as factCount: 0, never an error.
  • **Assert at close (P10)**, fact shapes:

Both touchpoints are wrapped in if (kipEnabled) (default true); the assert facts array is built unconditionally-non-empty when reached (the signature fact always exists).

Usage

bash
babysitter run:create \
  --process-file library/specializations/incident-management/incident-lifecycle.js \
  --inputs '{
    "signal": {
      "source": "alert",
      "ref": "pagerduty:P-4821",
      "firstSeenAt": "2026-07-23T09:14:00Z",
      "symptomSummary": "checkout API p99 latency 8x baseline, elevated 5xx on payment confirmation",
      "impactedSurfaces": ["checkout", "payments-api"]
    },
    "customerFacing": true,
    "oncall": {
      "primary": "alice",
      "secondary": "bob",
      "commsLead": "carol",
      "engineeringManager": "dana"
    },
    "commsChannels": [
      { "kind": "status-page", "target": "status.example.com" },
      { "kind": "slack", "target": "#incident-war-room" }
    ],
    "slo": { "mttdMinutes": 10, "mttrMinutes": 120 }
  }'

A signal like this classifies as SEV2: the status page gate routes to comms-lead (required, 60m cadence), customer comms is required because customerFacing is true, mitigation execution routes to incident-commander, and postmortem publication routes to engineering-manager after the completeness gate passes.

Non-interactive runs

Nothing policy-gated auto-approves **by design** — execute-prod-mitigation never sets autoApproveAfterN, and the comms/postmortem gates require an explicit approval. If a non-interactive harness auto-approves a breakpoint at its own level, that approval is recorded in outputs.autoApprovals as { breakpointId, phase, at } so the fail-closed posture stays auditable. autoApprovals is always present in outputs, even when empty.

Trail

Wiki

Library

Incident Management (Library)

Continue reading

accessibility (Library)
Aerospace Engineering Specialization (Library)
AI Agents and Conversational AI Specialization (Library)
Algorithms and Optimization Specialization (Library)
Arts and Culture Specialization (Library)
ATDD/TDD Methodology (Library)
authoring (Library)
AutoMaker (Library)

Page record

Open node ledger

wiki/library/incident-management.md

Documents

specialization:incident-management