Agentic AI Atlasby a5c.ai
OverviewWikiGraphFor AgentsEdgesSearchWorkspace
/
GitHubDocsDiscord
i.2Wiki
Agentic AI Atlas · MLOps (Library)
library/mlopsa5c.ai
Search the atlas/
Wiki · linked records

Article and nearby pages

I.Current articlepp. 1 - 1
accessibility (Library)Aerospace Engineering Specialization (Library)AI Agents and Conversational AI Specialization (Library)Algorithms and Optimization Specialization (Library)Arts and Culture Specialization (Library)ATDD/TDD Methodology (Library)
II.Documented nodesrefs · 1
specialization:mlops
I.
Wiki article

library/mlops

Reading · 10 min

MLOps (Library) reference

Flagship model-lifecycle for the library: dataset governance intake (parallel per-dataset lineage/consent/retention checks) -> eval-harness design -> executed training/eval runs -> an adversarial eval-review gate that RE-RUNS a sampled eval and diffs metrics -> a policy-gated model promotion with an executed serving smoke -> drift-monitoring setup with an executed drift-detection stub -> an adversarial drift-review gate -> a drift path with severity-routed escalation and a policy-gated rollback/retirement -> kip-backed model-registry memory. This is a brand-new specialization directory (verified: no prior mlops dir anywhere in library/).

Page nodewiki/library/mlops.mdNearby pages · 135Documents · 1

Continue reading

Nearby pages in the same section.

accessibility (Library)Aerospace Engineering Specialization (Library)AI Agents and Conversational AI Specialization (Library)Algorithms and Optimization Specialization (Library)Arts and Culture Specialization (Library)ATDD/TDD Methodology (Library)authoring (Library)AutoMaker (Library)Automotive Engineering Specialization (Library)Backend Development (Library)BDD/Specification by Example (Library)Bioinformatics and Genomics Specialization (Library)Biomedical Engineering Specialization (Library)BMAD Method (Library)business/ (folded) (Library)Business Analysis and Consulting (Library)Business Strategy and Operations (Library)Business Strategy Specialization (Library)CC10X Methodology (Library)CCPM - Claude Code PM Methodology (Library)Chemical Engineering Specialization (Library)Civil Engineering Specialization (Library)ClaudeKit Methodology (Library)Cleanroom Software Engineering (Library)CLI and MCP Development Specialization (Library)Code Migration and Modernization Specialization (Library)COG Second Brain (Library)collaboration (Library)common-utilities (Library)Communication specialization (Library)Composition: Aerospace Flight Control (Waterfall + V-Model + Cleanroom + inline Formal Verification) (Library)Composition: Legacy Modernization (Event Storming + DDD + FDD + Strangler Fig + RUP) (Library)Composition: Open Source Data-Validation Framework (TDD + BDD + Kanban + XP + Continuous Deployment) (Library)Composition: Regulated Greenfield (V-Model + DDD + Cleanroom + Waterfall) (Library)Composition: SaaS Analytics Dashboard (JTBD + Impact Mapping + Spec-Kit + Kanban + XP) (Library)Composition: Smart Product Recommendations (DDD + Hypothesis-Driven Development + BDD + Kanban) (Library)Composition: Startup MVP (Shape Up + Example Mapping + TDD + Scrum) (Library)Computer Science Specialization (Library)Cryptography and Blockchain Development Specialization (Library)Customer Experience and Support Specialization (Library)customer-support (Library)Data Engineering, Analytics, and BI Specialization (Library)Data Privacy Compliance (Library)Data Science and Machine Learning Specialization (Library)Intelligence, Decision Support and Decision Making (Library)Desktop Product Development Specialization (Library)developer-relations (Library)DevOps, SRE, and Platform Engineering Specialization (Library)Digital Marketing and Content Strategy Specialization (Library)Domain-Driven Design (DDD) Methodology (Library)Double Diamond Methodology (Library)Education and Learning Specialization (Library)Electrical Engineering Specialization (Library)Embedded Systems Engineering Specialization (Library)Entrepreneurship and Startup Processes (Library)Environmental Engineering Specialization (Library)Event Storming (Library)Everything Claude Code Methodology (Library)Example Mapping Methodology (Library)Extreme Programming (XP) (Library)Feature-Driven Development (FDD) (Library)Finance, Accounting, and Economics Specialization (Library)FPGA Programming and Hardware Description Specialization (Library)Game Product Development Specialization (Library)Gas Town Methodology (Library)GPU Programming and Parallel Computing (Library)GSD-Adapted Workflows for Babysitter SDK (Library)Healthcare and Medical Management Specialization (Library)Human Resources and People Operations Specialization (Library)Humanities and Anthropology Specialization (Library)Hypothesis-Driven Development (Library)Impact Mapping Methodology (Library)Incident Management (Library)Industrial Engineering Specialization (Library)internationalization (Library)Jobs to Be Done (JTBD) Methodology (Library)Kanban (Library)Knowledge Management (Library)Legal and Compliance Specialization (Library)Logistics and Operations Specialization (Library)Maestro App Factory (Library)Marketing and Brand Management Specialization (Library)Materials Science Specialization (Library)Mathematics Specialization (Library)Mechanical Engineering Specialization (Library)media (Library)Meta Specialization - Process, Skill, and Agent Creation (Library)Metaswarm Methodology (Library)Mobile Product Development Specialization (Library)Nanotechnology Specialization (Library)Network Programming and Protocols Specialization (Library)Observability specialization (Library)Enhanced Ontology-Driven Development (ODD) Methodology (Library)Operations Management Specialization (Library)Performance Optimization and Profiling Specialization (Library)Philosophy and Theology Specialization (Library)Physics Specialization (Library)Pilot Shell Methodology for Babysitter SDK (Library)Planning with Files (Library)Procurement — business domain specialization (Library)Product Management and Product Strategy Specialization (Library)Production contract (Library)Programming Languages and Compilers Development Specialization (Library)Project Management and Leadership Specialization (Library)Public Relations and Communications Specialization (Library)QA, Testing, and Test Automation (Library)Quantum Computing Specialization (Library)Release Engineering (Library)Research Specialization (Library)Robotics and Simulation Engineering Specialization (Library)RPIKit Methodology (Library)Ruflo Methodology (Library)RUP (Rational Unified Process) (Library)Sales and Business Development Specialization (Library)Scientific Discovery and Problem Solving Specialization (Library)Scrum (Library)SDK, Platform, and Systems Development (Library)Security, Compliance, and Risk Management Specialization (Library)Security Research and Vulnerability Analysis Specialization (Library)Shape Up (Library)Shared (Cross-Domain Assets) (Library)Social Sciences Specialization (Library)Software Architecture and Design Patterns Specialization (Library)sourcing/ (folded) (Library)Spec Kit Methodology (Library)Spiral Model (Library)Superpowers Extended Methodology (Library)Supply Chain Management Specialization (Library)Technical Documentation Specialization (Library)Travel (Curated-Dataset + SQL-Tool Pattern) (Library)UX/UI Design and User Experience Specialization (Library)V-Model Methodology (Library)Venture Capital and Investment Due Diligence Specialization (Library)Waterfall Methodology (Library)Web Product Development Specialization (Library)

Documented graph nodes

Records linked directly from this page’s Page node.

specialization:mlops

MLOps

Flagship model-lifecycle for the library: dataset governance intake (parallel per-dataset lineage/consent/retention checks) -> eval-harness design -> executed training/eval runs -> an adversarial eval-review gate that RE-RUNS a sampled eval and diffs metrics -> a policy-gated model promotion with an executed serving smoke -> drift-monitoring setup with an executed drift-detection stub -> an adversarial drift-review gate -> a drift path with severity-routed escalation and a policy-gated rollback/retirement -> kip-backed model-registry memory. This is a brand-new specialization directory (verified: no prior mlops dir anywhere in library/).

Composition map — callable upstream training stages (NOT superseded)

model-lifecycle.js consumes a **trained candidate** (modelVersion + artifactRef) and owns the governance / eval / promotion / drift / retirement lifecycle **around** it. Training itself is a pre-bar point task that can be delegated to either data-science-ml near-miss:

Upstream stageRole
`data-science-ml/model-training-pipeline.js`Hyperparameter tuning + experiment tracking producing the candidate model artifact that model-lifecycle P3 evaluates and promotes
`data-science-ml/automl-pipeline.js`Alternate: automated algorithm selection / ensembling producing a candidate; feeds the same P3 eval-harness inlet

These are mapped as callable upstream stages — **NOT superseded, NOT re-implemented**, nothing deprecated. mlo.training-run (P3, optional, gated on retrain) is the delegation seam.

Module table — `model-lifecycle.js` exports

ExportKindPurpose
process(inputs, ctx)orchestratorThe flagship lifecycle, phases P0–P8
MODEL_STAGESfrozen const['development','staging','production'] — ordered lifecycle stages
STAGE_PROMOTION_POLICYfrozen constPer-target-stage entry gate + accountable expert (lookup via stagePromotionPolicy)
DATASET_GOVERNANCE_CHECKSfrozen const['lineage','consent','retention'] — the three per-dataset checks (dsar-lifecycle shape)
DRIFT_SEVERITIESfrozen const['SEV1','SEV2','SEV3','SEV4'] — mirrored from release-lifecycle severity routing
DRIFT_ROUTINGfrozen constDrift escalation routing per severity (lookup via driftRouting)
stagePromotionPolicy(stage)helperPromotion-policy lookup — **throws** on unknown stage (no fallback policy)
driftRouting(severity, request?)helperRouting lookup — **throws** on unknown severity and on escalationExpert requests for immediate-rollback severities (no fallback route)
assertDriftSeverity(value, source)helperAccepts SEV1..SEV4 or 'none'; anything else **throws** naming the source
governanceCheckLabel(check)helperValidates a check name against DATASET_GOVERNANCE_CHECKS; **throws** on unknown check
datasetGovernanceCheckTaskagent taskmlo.dataset-governance-check — runs lineage/consent/retention for one dataset (fanned out per dataset)
retentionExecutionTaskagent taskmlo.retention-execution — executes the approved retention action; only inside dataset-retention-action approved
evalHarnessDesignTaskagent taskmlo.eval-harness-design — authors the benchmark harness + regression-threshold table
trainingRunTaskagent taskmlo.training-run — optionally (re)trains the candidate; delegable to the data-science-ml near-misses
evalRunTaskagent taskmlo.eval-run — actually runs one benchmark suite; one instance per evalSuite via ctx.parallel
promotionDeployTaskagent taskmlo.promotion-deploy — promotes exactly the evaluated version; only inside model-promotion-approval approved
promotionVerificationTaskagent taskmlo.promotion-verification — executes a serving smoke proving the promoted version serves
driftMonitorSetupTaskagent taskmlo.drift-monitor-setup — configures detectors + hooks, captures baseline, runs the drift-detection stub
driftTriageTaskagent taskmlo.drift-triage — SEV1..SEV4 classification grounded in the executed drift metrics
rollbackExecutionTaskagent taskmlo.rollback-execution — restores currentProductionRef / retires the superseded version; only inside model-rollback-approval approved
rollbackVerificationTaskagent taskmlo.rollback-verification — executed probes proving the restored version serves and the drift symptom is gone

Style note

All tasks are Style-A kind: 'agent' (zero kind: 'shell'), with per-effect io paths (tasks/<effectId>/input.json|result.json) and labels, and every gate / verification / executed-run output schema declares evidence { type: 'array', minItems: 1 }. Gate combinators (routedBreakpoint, adversarialGate, kipRecall, kipAssert) are imported from `../common-utilities/routed-gate-combinators.js`, not redefined. Timeline, breakpointsHit, and autoApprovals are accumulated in the orchestrator only — agents never write the timeline.

Stage model

Lifecycle stages (`MODEL_STAGES`, verbatim)

['development', 'staging', 'production'] — promotion advances one stage toward production; promotionTargetStage must be a member beyond development (default production). An unknown value throws (no fallback stage).

Promotion table (`STAGE_PROMOTION_POLICY`, verbatim)

Target stageEntry gateExpert
stagingmodel-promotion-approvalml-engineering-lead
productionmodel-promotion-approvalml-engineering-lead

Lookups go through stagePromotionPolicy(stage), which **throws** on an unknown stage — there is no fallback promotion policy.

Drift routing table (`DRIFT_ROUTING`, verbatim)

SeverityEscalation pathEscalation expert
SEV1immediate-rollback— (straight to the model-rollback-approval gate; expert lookup throws)
SEV2immediate-rollback— (straight to the model-rollback-approval gate; expert lookup throws)
SEV3remediation-choiceml-engineering-lead
SEV4remediation-choiceml-engineering-lead

All lookups (stagePromotionPolicy, driftRouting, assertDriftSeverity, governanceCheckLabel) **throw** naming the source on any unknown enum — there are no fallback rows anywhere.

Policy-gated actions

Three actions are policy-gated. Convention: **breakpointId = actionId**, strategy single.

actionIdExpertTagsRaised whenRejection behavior
dataset-retention-actiondata-governance-officer['policy-gated','mlops','dataset-governance']P1, only when a dataset governance check flags a deletion/retention-enforcement action; payload carries the dataset, proposedAction, check details, and priorKnowledge governance factsretentionExecutionTask never invoked; the action is recorded not-executed; if the un-actioned dataset is a required training/eval set the lifecycle **fails closed before eval**
model-promotion-approvalml-engineering-lead['policy-gated','mlops','promotion']P5, only after a passed mlo.eval-review gate; payload carries eval metrics, per-threshold pass/fail, eval-review evidence, governance clearance, target stageRun ends success:false, nothing promoted; promotionDeployTask never invoked (no alternate path)
model-rollback-approvalml-engineering-lead`['policy-gated','mlops','<sev>''retirement']`P7 drift path (SEV1/SEV2 immediately; SEV3/SEV4 after the remediation-choice picks rollback), and for retirement of a superseded production version

**Fail-closed posture:** there is no alternate execution path around a gate — the retention, promotion, and rollback executors are invoked **only** inside gate.approved === true branches, each with an explicit code comment that no other call site exists. **No gate in this process sets autoApproveAfterN**, and all three policy gates carry explicit code comments stating it must never be added. Any harness-level auto-approval is surfaced in outputs.autoApprovals ({ breakpointId, phase, at }), which is **always present** in outputs, possibly empty.

Quality gates

`mlo.eval-review` (P4) — RE-RUNS a sampled eval, diffs metrics

Runs over the eval report (per-suite metrics + deterministic threshold pass/fail), with the harness and evalSuites reachable in context. Failure (including an owner-rejected escalation) ends the run **before** model-promotion-approval is ever raised (fail closed).

CriticFocus
sampled-eval-reexecution-criticRE-RUNS a sampled subset of an evalSuite itself and DIFFS the fresh metrics against the reported metrics — raw re-execution outputs cited per sampled metric; a citation of the reported number without a fresh run is NOT evidence
regression-threshold-criticRecomputes every metric against regressionThresholds (min/max/baseline/maxRegression) and flags any breach the run under-reported — the recomputation is cited
eval-integrity-criticThe eval datasets are the governed holdout with no train/eval leakage — cross-checked against the P1 lineage clearance; datasets checked are cited

`mlo.drift-review` (P6) — RE-RUNS the drift-detection stub, fires the hooks

Runs over the drift-monitor spec + executed drift-detection run. A live breach in the executed run routes into the P7 drift path.

CriticFocus
drift-detection-reexecution-criticRE-RUNS the drift-detection stub itself over baselineWindow vs checkWindow and DIFFS its drift scores against the reported ones — raw executed outputs cited; a read of the reported drift score is not evidence
alerting-hook-criticProves each configured alerting hook ACTUALLY FIRES on a synthetic drift breach (executed) — the fired-alert output cited per hook

**IRON-LAW (executed evidence only):** the eval-review gate must re-run a sampled eval and diff; the drift-review gate must re-run the drift-detection stub and fire the hooks — reading a report or spec is not evidence, and passed:true with empty evidence is rejected by the combinator. Fix budget: maxFixAttempts (default 2) rounds of the built-in gateFixerTask; on exhaustion the combinator escalates to the owner via a routed breakpoint (mlo.eval-review.gate-escalation / mlo.drift-review.gate-escalation). A model **never promotes on a read-only review**.

Drift path

Entered only when the executed drift-detection run surfaces a live breach. driftTriageTask classifies SEV1..SEV4 grounded in the executed drift metrics (recommendation only). Then DRIFT_ROUTING:

  • **SEV1 / SEV2** -> straight to the model-rollback-approval gate.
  • **SEV3 / SEV4** -> one non-policy mlo.drift.remediation-choice breakpoint first (accept-drift/roll-forward vs rollback, expert ml-engineering-lead) — the **only non-policy breakpoint** in this process (sparse-breakpoint rule: the call is genuinely ambiguous at low severity). A roll-forward response ends the run success:false (drift accepted, no rollback gate raised); any other response proceeds to the gate.

On model-rollback-approval approved, rollbackExecutionTask restores exactly currentProductionRef and rollbackVerificationTask executes serving probes proving the restored version serves and the drift symptom is gone (bounded by driftPolicy.maxRollbackAttempts, default 1). **Retirement** of a superseded prior production version is the same gate surface (model-rollback-approval, tagged retirement): a clean promotion to production that supersedes a non-null currentProductionRef requests retirement of the old version through this gate.

kip model-registry memory

- { predicate: 'has-version', object: <modelVersion> } - { predicate: 'outcome', object: 'promoted'|'failed'|'rolled-back', props: { evalReviewPassed, promoted, driftDetected } } - one per eval suite: { predicate: 'eval-metric', object: <suite>, props: { thresholdsPassed } } - { predicate: 'promotion-decision', object: 'promoted-to-<stage>'|'not-promoted', props: { approved, verified } } - only when the drift path ran: { predicate: 'drift-incident', object: <triage.summary>, props: { severity, rollbackVerified } } and { predicate: 'rollback-lesson', object: '<severity>: <rationale>' }

  • **Recall (P0)**: kipRecall(ctx, { kipDir, topic: 'model registry: <modelName>@<modelVersion>', kipModel, kind: 'mlops-model-registry' }) — prior model performance, promotion decisions, and drift incidents threaded as priorKnowledge into every downstream agent. An empty store is initialized and reported as factCount: 0, never an error.
  • **Assert at close (P8)**, facts built deterministically in the orchestrator (never in an agent), subject model:<modelName>@<modelVersion>:

Both touchpoints are wrapped in if (kipEnabled) (default true). The assert facts are unconditionally non-empty when reached — the has-version and outcome facts always exist.

Inputs / outputs reference

Mirrors the JSDoc @inputs / @outputs in model-lifecycle.js. Required: model { modelName, modelVersion, artifactRef }, a non-empty datasets[] (each { name, uri, purpose }, purpose train|eval|holdout), a non-empty evalSuites[] (each { name, dataset, metrics[] }), and a regressionThresholds object mapping every referenced metric to a threshold (a referenced metric with no threshold **throws** before any run). Optional: promotionTargetStage (default production), retrain (default false), driftPolicy { maxRollbackAttempts, baselineWindow, checkWindow }, maxFixAttempts (default 2), kipEnabled/kipDir/kipModel, artifactsDir.

Usage

bash
babysitter run:create \
  --process-file library/specializations/mlops/model-lifecycle.js \
  --inputs '{
    "model": {
      "modelName": "fraud-scorer",
      "modelVersion": "2.4.0",
      "artifactRef": "s3://models/fraud-scorer/2.4.0/model.pt",
      "currentProductionRef": "s3://models/fraud-scorer/2.3.1/model.pt"
    },
    "datasets": [
      { "name": "txn-train", "uri": "s3://data/txn/train", "purpose": "train", "lineageRef": "dvc://txn@train" },
      { "name": "txn-holdout", "uri": "s3://data/txn/holdout", "purpose": "holdout", "consentBasis": "contract" }
    ],
    "evalSuites": [
      { "name": "accuracy-suite", "dataset": "txn-holdout", "metrics": ["auc", "precision"] }
    ],
    "regressionThresholds": {
      "auc": { "min": 0.9, "baseline": 0.94, "maxRegression": 0.01 },
      "precision": { "min": 0.85 }
    },
    "promotionTargetStage": "production"
  }'

For this run: the two datasets clear lineage/consent/retention in parallel (a flagged retention action would gate dataset-retention-action with the data-governance-officer), the eval harness is authored and every metric is checked for a threshold, accuracy-suite runs against the candidate and is scored deterministically, the mlo.eval-review gate re-runs a sampled eval and diffs it, model-promotion-approval (ml-engineering-lead) gates promotion to production, an executed serving smoke verifies it, mlo.drift-review re-runs the drift-detection stub, and — because a prior currentProductionRef exists — retirement of 2.3.1 is requested through the model-rollback-approval gate tagged retirement.

Non-interactive runs

Nothing policy-gated auto-approves **by design** — no gate in this process sets autoApproveAfterN, and the three policy gates must never gain it. If a non-interactive harness auto-approves a breakpoint at its own level, that approval is recorded in outputs.autoApprovals as { breakpointId, phase, at } with its phase provenance, so the fail-closed posture stays auditable. autoApprovals is always present in outputs, even when empty.

Trail

Wiki

Library

MLOps (Library)

Continue reading

accessibility (Library)
Aerospace Engineering Specialization (Library)
AI Agents and Conversational AI Specialization (Library)
Algorithms and Optimization Specialization (Library)
Arts and Culture Specialization (Library)
ATDD/TDD Methodology (Library)
authoring (Library)
AutoMaker (Library)

Page record

Open node ledger

wiki/library/mlops.md

Documents

specialization:mlops