agentic-chatops
AI agents that triage infrastructure alerts, investigate root causes, and propose fixes — while a solo operator sleeps.
For the complete technical reference, see README.extensive.md.

The Problem
One person. 310+ infrastructure objects across 6 sites. 3 firewalls, 12 Kubernetes nodes, self-hosted everything. When an alert fires at 3am, there's no team to call. There never is.
The Solution
Three agentic subsystems that handle the detective work — ChatOps (infrastructure), ChatSecOps (security), ChatDevOps (CI/CD) — built on n8n orchestration, Matrix as the human interface, and a tiered agent architecture (deterministic triage scripts → Claude Code → human). The human stays in the loop for every infrastructure change: the system never acts without a thumbs-up or poll vote, and since 2026-06-09 a remediation proposal cannot even reach the approval poll without a machine-computed consequence prediction attached (see Infragraph below).
What Makes This Different
Self-Improving Prompts — now with A/B trials (nobody else does this)
The system evaluates its own performance and auto-patches its prompts. Every session is scored by an LLM-as-a-Judge on 5 quality dimensions (gemma3:12b local-first since 2026-04-19; max-effort calibration via gw-mistral-large on the shared LiteLLM). When a dimension averages below threshold over 30 days, the preference-iterating patcher (IFRNLLEI01PRD-645, 2026-04-20) generates 3 candidate instruction variants (concise / detailed / examples) and assigns each future matching session to one arm via deterministic BLAKE2b hash — plus a no-patch control. A daily cron runs a one-sided Welch t-test once every arm reaches 15 samples; the winner is promoted only if it beats control by ≥ 0.05 points with p < 0.1. Otherwise the trial is aborted. Prompt-level policy iteration — no model weights are ever fine-tuned.
Session → LLM Judge (5 dims) → dimension trending below threshold
→ prompt-patch-trial.py generates 3 candidate variants + 1 control
→ future sessions hash-routed to arms → Welch t-test at 15+ samples/arm
→ winner promoted to config/prompt-patches.json (source: "trial:N:idx=I")
→ next eval cycle scores the new patch → loop continues
Infragraph — a Causal World Model with a Non-Bypassable Prediction Gate (2026-06-09)
The system maintains a causal dependency graph of the entire infrastructure (361 nodes / 468 edges in the causal layer; 721 entities / 661 relationships in the combined GraphRAG+infragraph knowledge graph) seeded daily from five truth layers — live Proxmox cluster API (0.95 confidence), LibreNMS dependency parents (0.90), NetBox devices + physical cables (0.85–0.90), operator-declared edges, and a statistical incident-co-occurrence miner deliberately capped at 0.75 — with per-edge dynamics (expected alert cascades, propagation delays, recovery times) learned from 159 chaos experiments and the full triage history. This is a genuine model-free → model-based shift enforced in control flow, not data:
- Prediction is computed outside the LLM — deterministic graph traversal (
infragraph-query.py), called by the n8n orchestrator, never at the model's discretion. - Prediction is mandatory — the Runner commits a plan-hash-keyed prediction artifact before any approval poll; a remediation proposal without one is rewritten to
[POLL-WITHHELD:NO-PREDICTION]and demoted to analysis-only. The kill-switch (INFRAGRAPH_DISABLED=1) fails the remediation lane closed. - Verification is mechanical — after execution, code (never the LLM that proposed the action) diffs observed alerts against the prediction and writes a
match / partial / deviationverdict; deviation = surprise = never auto-resolve.
The eval is falsifiable by design: a degree-preserving shuffled-graph negative control runs alongside every prediction. The 2026-05-11 cascade backtest passed the criterion (control ratio 0.367 ≤ 0.5×) only after four honest iteration rounds, each driven by what the misses revealed. Suppression authority is granted per rule by the operator — the system proposes (control YouTrack issue with evidence table), the human approves, and closing the control issue instantly revokes. Runbook: docs/runbooks/infragraph.md.
Autonomy-Forward Gate — Human as Circuit-Breaker, not Gatekeeper (2026-06-16)
Most "human-in-the-loop" systems assume the human is watching. Ours measured that the operator had voted on almost none of the approval polls in the prior two months — so the loop was a dead-end: reversible work stalled on a 30-min pause and genuinely-critical work paged no one. The fix is a 3-band risk gate (classify-session-risk.py): reversible, Infragraph-prediction-backed changes auto-resolve; a tightly-scoped critical set (HIGH-risk, P0-host blast, irreversible, model deviation) is the only thing that pages the operator by SMS. The safety floor is mechanical and non-configurable, and the whole gate flips on/off with a single touch/rm of a sentinel file — no workflow edit, instant kill-switch. Runbook: docs/runbooks/risk-based-auto-approval.md.
Self-Verifying Reliability Layer — the system watches itself (2026-06-21)
The defining failure mode here was never a crash — it was months of silent darkness (the auto-resolve pipeline dead across 5 layers, scanners dark 5 weeks, an apiserver crash-looping 27 days) where nothing alerted because standard alerting treats no data as no problem. So the autonomy loop is now a continuously-verified subsystem:
- Control-plane dead-man's-switch —
gateway-watchdog.shemits a heartbeat every 5 min via atrap … EXIT; a Prometheus alert with anabsent()clause pages by SMS if the heartbeat goes stale or vanishes (node_exporter/host down). It watches the thing that watches the pipeline. - Synthetic-incident canary — a daily probe drives the real classify→predict spine end-to-end against an isolated throwaway DB, so it proves the spine is alive (3 stages + plan-hash coherence) while structurally being unable to pollute production, collide a real fail-closed gate, or trigger remediation. A tier-1 SMS fires if it ever leaks a row into the live DB.
- False-auto-resolve governance — the system measures its own root-cause discipline: a pattern it auto-resolved that recurs within 24h is a false-auto-resolve, and a repeat offender (≥3×/30d) is auto-demoted so the gate escalates it instead of auto-closing it again — automatically, reversibly (30-day expiry), with no human review (human-as-circuit-breaker, not gatekeeper). Intentionally-suppressed flappy alerts are excluded so it never re-introduces suppressed noise.
- Bi-temporal knowledge — infragraph edges and compiled-wiki facts carry a contradiction/supersession axis with time-since-confirmation decay (reporting-only — it flags edges for re-ratification, never silently changes a prediction).
- Self-learning scheduled-reboot suppression (2026-06-29) — hosts with a discovered and promoted reboot schedule (observe-≥2-boots-before-live, strict DST-correct cron windows) get their on-schedule reboot alerts suppressed before any session spawns, with a two-phase verify that reopens + pages if the boot wasn't a clean
systemd-reboot. Safety floor: critical-never, allowlisted rules only, sentinel kill-switch, fail-open. Runbook:docs/runbooks/scheduled-reboot-suppression.md.
Orchestrator Control-Plane — the system governs itself (2026-06-26)
The agentic federation grew to ~10 subsystems and 363 components (320 at the 2026-06-26 landing) — Cronicle jobs, 57 n8n workflows, hooks, and the RAG / infragraph / teacher / chaos subsystems — coordinated only by convention, a shared SQLite, and the Prometheus textfile bus. Nothing owned their liveness as a set, and a 2026-06-25 audit proved the cost: MemPalace hooks, the OTel span sink, the tool-call log, and even the self-audit itself had run dark for weeks-to-months, each invisible because standard alerting reads no data as no problem. The fix is a thin governing layer (IFRNLLEI01PRD-1421) — three bricks built on the existing Prometheus + SQLite substrate, no platform rewrite (the research explicitly rejected adopting LangGraph / Temporal / Airflow / Dagster / Backstage):
- Component Registry (
scripts/registry-check.py) auto-discovers all 363 components (199 cronicle-job + 77 prom-writer + 57 n8n-workflow + 28 db-table + 2 cron, as of 2026-07-08), each with a declared liveness expectation — 15 critical, 0 critical-dark, ~10 known-dark-by-design. The dark-component failure class is now caught mechanically (RegistryCriticalDark, tier-1 SMS) instead of by a manual quarterly sweep. - Interaction Graph (
scripts/interaction-graph.py) static-analyzes 313 scripts into a read/write asset graph (Dagster's model in ~250 lines): currently 0 GAPs (the Session-End → reconcile orphan-consumer hole that silently darkened 4 analytics tables is closed), 0 cron-clashes, and 23 multi-writer conflicts surfaced for review. - Orchestration Benchmark (
scripts/orchestration-benchmark.py) replays a synthetic incident stream through the isolated classify→predict spine and scores 4 orchestration invariants — score 1.0, 4/4, including safety-composition: an irreversible incident is never auto-resolved, verified across the whole stream rather than case-by-case (OrchestrationSafetyFailure, tier-1).
All five rules are live in-cluster (infra MRs !347 + !348), and a fault-injection drill proved the alerts actually fire, not merely evaluate. The control-plane monitors its own three bricks — the who-watches-the-watcher gap is fully closed.
Plane-A self-healing platform controller — the actuator half (2026-06-26). The bricks detect; a Kubernetes-style self-healing operator (scripts/platform-controller.py, */5, armed) acts — closing the loop the absent human left open. It heals only idempotent platform operations: reactivate an inactive critical n8n workflow (it monitors all 57), re-run a failed safe-list Cronicle job, restart Cronicle, plus a consolidated watchdog heal-library. Heals are rate-limited by exponential heal-backoff → CrashLoopBackOff → SMS escalation, exactly as a k8s controller would. Crucially it draws the same Plane-A / Plane-B line k8s does between keeping pods alive and deciding app logic: Plane-A keeps the platform alive (crons, Cronicle, bricks, writers, n8n); Plane-B is the mission (resize a VM, reboot a host, resolve an incident) — the controller NEVER touches B. That stays the autonomy-forward lane's job. It consolidated the standalone watchdog into one operator, and carries its own dead-man.
Cronicle scheduler — every job has run-history now (2026-06-26). All cron jobs (180 at migration: 107 gateway + 72 agora-quant; 199 registered as of 2026-07-08) migrated off raw crontab to a native Cronicle scheduler: per-job run history, per-job-death alerting (the gap a flat crontab can't see — a crontab line that silently stopped firing looks identical to one that never existed), a REST API the registry seeds from, and auto-quarantine of a repeatedly-failing job.
A single realtime control-plane dashboard (grafana/orchestrator-control-plane.json, live at grafana.example.net/d/orchestrator-ctrl-plane) puts the whole thing on one pane of glass — the three bricks, the self-healing actuator, the scheduler, the decision plane, and the integrity / dead-man guarantees — 31 panels across 6 sections, refreshed every 30s.
The decision log itself is now tamper-evident: every governance decision (830 logged, 78% auto-approved, as of 2026-07-08) is chained by SHA-256 so any retroactive edit breaks the chain and pages by SMS (GovernanceChainBroken, tier-1). Observability is unified end-to-end — logging to self-hosted OpenObserve, Langfuse traces, and a fresh OTLP push — across ~1,700 metric series / 77 textfile writers / 74 in-repo alert rules, backed by the dead-man heartbeat, the synthetic-incident canary, and a deploy-drift guard.
Benchmarked against industry orchestration standards, the control-plane scores B+ (3.48 / 5) across 11 dimensions — strongest on the things almost nobody enforces: Plane-A / Plane-B separation enforced in code (not policy), reversibility-keyed human-in-the-loop, and independent mechanical verification of every outcome.
The whole layer governs roughly 10 subsystems · 363 components · 199 jobs · 57 n8n workflows · 53 DB tables · ~97K LOC across 433 scripts (2026-07-08) — and watches every one of them.
Benchmarked Against the Anthropic + OpenAI Agent Guides — 12/14 dimensions at A (2026-06-26)
The platform was scored as two separate, source-pure, adversarially-verified scorecards against Anthropic's Building Effective AI Agents (IFRNLLEI01PRD-1422) and OpenAI's A Practical Guide to Building Agents (-1423), then improved against what the misses revealed. 12 of 14 dimensions now sit at A (synthesis); the 2 remaining at B are deliberate operator decisions, not gaps — the rules blocklist is kept off the dispatched autonomous path, and the failure-threshold tripwire is a passive Matrix warning rather than an SMS page. Notable fixes shipped on the way: a model-router bug that counted markdown-table pipes instead of incident rows (pinning all 818 sessions to Opus — now low-risk alerts route to Sonnet behind a never-downgrade-risky floor), an OTLP trace export dead since ~March (a stale auth env shadowing the creds), a MemoryMax cgroup cap on dispatched sessions (the uncapped runaway class that wedged a host), and a concurrent-session tripwire that can now actually kill a runaway session on a token / cost / tool-call breach. LLM/agent traces now flow to a self-hosted Langfuse and dead-man "job never ran" liveness to a self-hosted Healthchecks.io — both composed alongside the bricks rather than replacing them.
Model Orchestration — centralized provider/model selection, the easy way (2026-06-28)
Which model runs on which component is centralized, not scattered across hardcoded IDs — and flippable with one command. Two planes (docs/model-provenance.md, MRs !116–!120):
- Claude Code (subscription, flat-rate): every
claudeinvocation — dispatched remediation,agent_as_tool,mr-review,parallel-dev, interactive — is routed by a single switch,scripts/claude-provider.sh{zai|anthropic|status}, which edits~/.claude/settings.json. Two providers: Z.ai (glm-5.2Opus-equivalent for--model opus,glm-4.7Sonnet-equivalent) and Anthropic Max (OAuth subscription).statusis authoritative for the live toggle — the operator flips it, so no document should claim a permanent default. Subscription auth can't proxy through a gateway, hence the direct route. - Eval layer (per-token API): the LLM judge, RAGAS, and the frontier cross-check route through the shared LiteLLM proxy to Mistral (
mistral-large-latest) + DeepSeek (deepseek-v4-pro), with local-Ollama fallback (never Anthropic). Per-component spend is tracked via LiteLLM tags. Per the operator directive, Mistral + DeepSeek are the only paid per-token APIs — zero Anthropic per-token spend. - Local ($0): judge / RAG synth-rewrite / embeddings / rerank / teacher on Ollama (
gemma3:12b,qwen2.5:7b,nomic-embed-text,bge-reranker-v2-m3).
The single source of truth is config/model-routing.json (resolved by scripts/lib/model_routing.py); the LiteLLM models+key are provisioned idempotently by scripts/litellm-gateway-setup.sh. To see "which model on which component now": python3 scripts/lib/model_routing.py --list for the intended-default catalog, plus bash scripts/claude-provider.sh status for the live Claude-Code provider (authoritative — the registry shows the intended default, status reflects the active settings.json toggle). This supersedes the old cc-cc/oc-* frontend/backend-pairing modes (OpenClaw retired).
Renovate MR Autonomy Lane — hands-off dependency updates with per-class gates (2026-07-06)
A self-hosted Renovate CE instance opens dependency-update MRs across the IaC estate; a dedicated n8n lane classifies each MR (classify-renovate-mr.py + a stateful-services manifest) and auto-merges + deploys + post-merge-verifies routine docker digest/patch bumps end-to-end — deterministic structural review, hard CI-green gate, snapshot-before-merge for stateful services, and a */15 reconciler. Anything consequential (Kubernetes, Helm, Terraform, OpenBao, Dockerfiles, majors) goes to a [POLL] + operator SMS instead — never auto-applied blind. Post-merge verification is 3-way: healthy / confirmed-bad → revert / inconclusive → escalate, never auto-revert. Armed via the ~/gateway.renovate_autonomy sentinel; first hands-off merges ran 2026-07-07. Runbook: docs/runbooks/renovate-mr-autonomy.md.
AI Planner Wired to Proven Ansible Playbooks
Before Claude Code investigates, a fast-tier planner (sonnet-tier, resolved by the centralized Model Orchestration layer) generates a 3-5 step investigation plan. The planner queries AWX for matching Ansible playbooks from 41 proven templates (maintenance, cert sync, K8s drain, PVE updates, DMZ deployments). Plans naturally include "Run AWX Template 64 with dry_run=true" as remediation steps — bridging AI reasoning with proven automation.
Predictive Alerting
Instead of only reacting after alerts fire, the system queries LibreNMS API daily for trending risk across both sites. Devices are scored on disk usage trends, alert frequency, and health signals. A daily top-10 risk report posts to Matrix before problems become incidents.
5-Signal RAG + GraphRAG + Staleness + Temporal Filter + mtime-Sort
Retrieval uses Reciprocal Rank Fusion across 5 signals (semantic + keyword + compiled wiki + MemPalace transcripts + chaos baselines), plus a GraphRAG + infragraph knowledge graph (721 entities, 661 relationships). Retrieval short-circuits via two intent detectors: temporal window ("last 48h", "72 hours ending YYYY-MM-DD") filters wiki on source_mtime, and mtime-sort intent ("name any three memory files created in the last 48h") bypasses semantic retrieval entirely and returns an mtime-ranked window. Results older than 7 days get age-proportional staleness warnings. A local qwen2.5:7b synth step composes cross-chunk answers when top rerank < threshold (rag-synth → Ollama under the centralized Model Orchestration layer). SYNTH_HAIKU_FORCE_FAIL env is retained for the failure-mode fallback path (429 / auth / timeout / network / empty).
Karpathy-Style Compiled Knowledge Base
Following Andrej Karpathy's LLM Knowledge Bases pattern: raw data from 7+ sources (575 memory files, 35 CLAUDE.md files, ~2,500 incidents, 107 docs, 22 skills, ~5,200 lab docs, as of 2026-07-08) is compiled into a browsable 88-article wiki with auto-maintained indexes, daily SHA-256 incremental recompilation, and contradiction detection. All articles embedded into RAG as the 3rd fusion signal.
Full Observability Stack with OTel
333K+ tool calls instrumented across 159 tool types with per-tool error rates and latency percentiles (2026-07-08). OTel spans exported to OpenObserve (OTLP; ~14K retained locally in SQLite). 13 Grafana dashboards (90+ panels, incl. the realtime orchestrator control-plane overview) covering ChatOps, ChatSecOps, ChatDevOps, and trace analysis. Infrastructure commands logged per-device in execution_log.
Formal Evaluation Pipeline
58 scenarios across 3 eval sets (22 regression + 20 discovery + 16 holdout) + 54 adversarial red-team tests. Prompt Scorecard grades 19 surfaces daily on 6 dimensions. Agent Trajectory scoring on 8 infra / 4 dev steps. A/B variant testing (react_v1 vs react_v2). CI eval gate blocks bad merges. Monthly eval flywheel cycle.
Structured Agentic Substrate — 9 adoptions from the OpenAI Agents SDK
The 2026-04-20 audit of openai/openai-agents-python flagged 11 gaps; 9 were implemented (issues IFRNLLEI01PRD-635..643). The system now has a versioned, typed, recoverable substrate the old string-based Matrix pipeline couldn't offer:
- Schema versioning on 9 session/audit tables + a central registry (
scripts/lib/schema_version.py) mirroring the SDK'sRunState.CURRENT_SCHEMA_VERSION/SCHEMA_VERSION_SUMMARIESpattern. Writers stampschema_version=CURRENT; readerscheck_row()fail-fast on future versions. - 13 typed events (
session_events.py) in a newevent_logtable —tool_started/ended,handoff_requested/completed/cycle_detected/compaction,reasoning_item_created,mcp_approval_*,agent_updated,message_output_created,tool_guardrail_rejection,agent_as_tool_call. Replaces free-form Matrix strings with Grafana-queryable structured telemetry. - Per-turn lifecycle hooks —
session-start.sh,post-tool-use.sh,user-prompt-submit.sh,session-end.sh(new — theon_final_outputequivalent) feeding asession_turnstable with per-turn cost, tokens, duration, tool count. - 3-behavior tool-guardrail taxonomy (
allow/reject_content/deny) inunified-guard.sh+audit-bash.sh+protect-files.sh.reject_contentsends Claude a retry hint instead of a wall;denyhard-halts. Every rejection is a typed event. HandoffInputDataenvelope (scripts/lib/handoff.py) — zlib-compressed base64 payload carryinginput_history,pre_handoff_items,new_items,run_context. 176 KB history → 752 B on the wire (0.43% ratio). Eliminates the "re-derive context via RAG" cost on escalation.- Transcript compaction (
scripts/compact-handoff-history.py) — opt-in per escalation. Localgemma3:12b(fast-tier fallback routed via the Claude-Code plane); circuit-breaker aware. - Agent-as-tool wrapper (
scripts/agent_as_tool.py) — wraps the 11 sub-agent definitions as callable tools so the orchestrator LLM can conditionally invoke them in the ambiguous-risk (0.4–0.6) band, complementing our deterministic routing. - Handoff depth counter + cycle detection (
scripts/lib/handoff_depth.py) —handoff_depth >= 5forces[POLL];>= 10hard-halts; any agent twice in the chain is refused and logged ashandoff_cycle_detected. - Immutable per-turn snapshots (
scripts/lib/snapshot.py) — a snapshot is captured BEFORE each mutating tool call (Bash,Edit,Write,Task; read-only tools skipped);rollback_to(id)restores any priorsessionsrow. 7-day retention.
Four new SQLite tables (event_log, handoff_log, session_state_snapshot, session_turns) bring the total to 35. Migrations 006–011 apply idempotently on both fresh and legacy DBs. Two follow-ups since then — the A/B prompt patcher (IFRNLLEI01PRD-645, prompt_patch_trial + session_trial_assignment) and the CLI-session RAG capture pipeline (-646/-647/-648, no new tables; chunks + tool calls + knowledge rows tagged issue_id='cli-<uuid>' on the existing schema) — the live total is now 53 tables / 31 schema-versioned (2026-07-08).
CLI-Session RAG Capture — interactive claude sessions flow into RAG too (2026-04-20)
Before this, only YT-backed Runner sessions had their transcripts/tool-calls/extracted knowledge written into the shared RAG tables. Interactive claude CLI sessions (human-in-the-loop dev work) were only captured by poll-claude-usage.sh for cost/tokens — their content was lost to retrieval.
A 3-tier pipeline (IFRNLLEI01PRD-646/-647/-648) closes the gap. A single cron line chains three idempotent steps over every CLI JSONL:
archive-session-transcript.pychunks exchange pairs →session_transcripts+nomic-embed-textembeddings + doc-chain refined summary atchunk_index=-1(sessions ≥ 5000 assistant chars).parse-tool-calls.pyextractstool_use/tool_resultpairs →tool_call_log(issue_id resolves tocli-<uuid>via patched path inference).extract-cli-knowledge.pyrunsgemma3:12bin strict-JSON mode over the summary rows →incident_knowledgewithproject='chatops-cli', embedded for retrieval.
Retrieval weights chatops-cli rows at CLI_INCIDENT_WEIGHT=0.75 by default so real infra incidents still win close ties. Byte-offset watermark skips unchanged files. Soak test (10 files): 12 chunks + 245 tool-call rows + 4 knowledge extractions — gemma correctly classified one sample as subsystem=sqlite-schema, tags=[schema, migration, versioning, data] at 0.95 confidence.
Skill Authoring Uplift — 6 dimensions closed vs google/agents-cli (2026-04-23)
A deep audit against google/agents-cli flagged 6 skill-authoring dimensions where we trailed (phase-gate choreography, discoverability, anti-guidance, inline behavioral anti-patterns, governance/versioning, skill index). An 11-commit uplift (IFRNLLEI01PRD-712 umbrella, Phases A→J) closed every gap. 0 reverts.
- Master phase-gate skill — new
.claude/skills/chatops-workflow/SKILL.mdcodifies the Phase 0→6 incident lifecycle (triage → drift-check → context → propose → approve → execute → post-incident). Force-injected into every Runner session's Build Prompt (marker-delimited for surgical removal; rollback anchor preserved at/tmp/runner-pre-IMMUTABLE.json). - Auto-generated skill index —
scripts/render-skill-index.pyemits a drift-gateddocs/skills-index.mdfrom all SKILL.md + agent frontmatter. Guarded bytest-656-skill-index-fresh.sh, refreshed as a pre-step of the daily 04:30 UTC wiki-compile cron. - Versioned + audited skills — every SKILL.md + agent frontmatter now carries
version: 1.x.0+requires: {bins, env}.scripts/audit-skill-requires.sh+ a Prometheus exporter feed two new alerts (SkillPrereqMissing,SkillMetricsExporterStale).scripts/audit-skill-versions.shwalks git history for body-changed-without-bump cases; semver convention atdocs/runbooks/skill-versioning.md. - Anti-guidance trailing clauses — every primary skill/agent description now ends with "Do NOT use for X (use /other-skill instead)". Measurably reduces over-routing to adjacent-sounding agents.
- Shortcuts-to-Resist tables inlined on 11 agents (46 rows drawn from
memory/feedback_*.mdwith source citations) — behavioral inoculation at the surface where the model is about to act. - Proving-Your-Work directive — new
check_evidence()inscripts/classify-session-risk.pyemits anevidence_missingrisk signal that forces[POLL]when CONFIDENCE ≥ 0.8 but the reply carries no tool output / code fence. Mirrored in the Runner's Prepare Result node to strip unearned[AUTO-RESOLVE]markers and prepend aGUARDRAIL EVIDENCE-MISSING:banner. - User-vocabulary map —
config/user-vocabulary.json(20 entries:"the firewall"→nl-fw01;gr-fw01,"xs4all"→"budget"post-2026-04-21 rename, etc.) scanned by the prompt-submit hook; every match emits a typedvocabularyevent toevent_log.
Scorecard delta: 3.94 → 4.94 average; 13/16 dimensions at 5/5 (was 9/16). Full memo: docs/scorecard-post-agents-cli-adoption.md. E2E hardened in the same batch via a J1–J5 pass: live vocabulary event captured by firing the real prompt-submit hook, promtool test rules executed inside the live Prometheus pod, force-injection proven by a real Runner session whose first tool call grepped for Phase 0 in the injected skill body.
NVIDIA DLI Cross-Audit + P0+P1 Implementation (2026-04-29)
The 19-transcript NVIDIA Deep Learning Institute Agentic AI Systems course (Vadim Kudlai) was the last major agentic-AI source not yet evaluated against this platform. The 12-dimension cross-audit on 2026-04-29 initially graded the system A (4.4/5.0) — the lowest of any of the 9 sources audited. A same-day implementation of all 7 P0+P1 items lifted it to A+ (4.83/5.0), putting the system at A+ across all 9 sources (aggregate A+ 4.79).
Shipped in 4 commits (G1–G4) under YouTrack umbrella IFRNLLEI01PRD-747 with children -748..-751. Six commits direct-pushed to main, zero reverts. 57/57 new QA tests pass.
- G1 — Long-horizon reasoning replay eval (
scripts/long-horizon-replay.py) replays the 30 longest historical sessions weekly (Mon 05:00 UTC), scoring trace_coherence, tool_efficiency, poll_correctness, cost_per_turn_z. Newlong_horizon_replay_resultstable;LongHorizonReplayStalealert. - G1 — Jailbreak corpus + Greek extension — 39 fixtures across the 5 NVIDIA-DLI-08 vectors (asterisk-obfuscation, persona-shift, retroactive-history-edit, context-injection, lost-in-middle-bait), including 8 Greek operator-language fixtures. Pure-regex
scripts/lib/jailbreak_detector.py; weekly regression cron (Wed 05:00 UTC);JailbreakBypassDetectedalert on any miss. - G2 — Intermediate semantic rail (DARK-FIRST) —
scripts/lib/intermediate_rail.py(heuristic + Ollama dual-backend) inserted as aCheck Intermediate RailCode node between Build Plan and Classify Risk in the Runner workflow (now 50 nodes). Emitsintermediate_rail_checkevent per session;IntermediateRailDriftHighalert at >20% out-of-dist over 24h. Observe-only — does NOT block; soft-gate evaluation deferred ≥7 days post-data. - G2 — Grammar-constrained decoding — JSON Schemas at
scripts/lib/grammars/passed to Ollama via theformatfield whenOLLAMA_USE_GRAMMAR=1(default on). Falls back toformat=jsonon schema rejection. Circuit-breaker semantics preserved. - G3 — Team-formation skill (
.claude/skills/team-formation/SKILL.mdv1.0.0) +scripts/lib/team_formation.pypropose a sub-agent roster per(alert_category, risk_level, hostname). Build Prompt injects a## Team Charter (advisory)section; same JSON emitted asteam_charterevent_log row. KNOWN_AGENTS inventory enforced against.claude/agents/*.md. - G3 — Inference-Time-Scaling explicit budget —
EXTENDED_THINKING_BUDGET_Senv var (+ optional per-category override) drives a## Reasoning BudgetBuild Prompt section;its_budget_consumedevent captures observed turns/thinking_chars at session end. - G4 — Server-side session-replay endpoint — new workflow
claude-gateway-session-replay.json(idlJEGboDYLmx25kBo) ACTIVE. POST/session-replayaccepts{session_id, prompt}, validates format, sqlite3-checks session existence inside the SSH command (the n8n task-runner sandbox blockschild_processin Code nodes), runsclaude -r, returns JSON. HTTP 404 on unknown session, HTTP 400 on malformed input.session_replay_invokedevent.
event_log schema bumped 1 → 4 (13 → 17 event types). 18 → 19 schema-versioned tables. 5 cron entries installed. 5 YouTrack issues all moved to Done via direct REST POST (the tonyzorin/youtrack-mcp:latest container's update_issue_state omits the $type: "StateBundleElement" discriminator — bug documented in memory/feedback_youtrack_mcp_state_bug.md).
Full state-of-the-platform reference: docs/agentic-platform-state-2026-04-29.md.
QA Suite — 834 known-passing tests, 85 suite files
scripts/qa/run-qa-suite.sh runs 85 suite files (78 suites + 7 e2e, ~7 min under full load; last full run 2026-07-08: 834 pass / 0 fail / 2 skip) with JSON scorecard + summary output, guarded by a per-suite QA_PER_SUITE_TIMEOUT wrapper (IFRNLLEI01PRD-724) that caps any slow/wedged suite at 120 s (a suite may declare a raise-only # QA_SUITE_TIMEOUT: <n> header for load headroom) and emits a synthetic FAIL record so the orchestrator never hangs silently:
- Per-issue suites — sanity + QA + integration for every adoption, plus 16 tests for the preference-iterating patcher (-645) and 12 tests for the CLI-session RAG pipeline (-646/-647/-648).
- Writer coverage — every script that
INSERTs into a versioned table is asserted to stampschema_version=1; same for all 5 n8n-workflow INSERT sites. - Pattern-by-pattern coverage — 53 deny-pattern tests + 32 reject-pattern tests.
- Payload shape — every one of the 13 event types round-trips through the CLI + Python paths.
- Concurrent-bump fuzz — 8 parallel
handoff_depth.bump()calls with no-lost-updates assertion. Surfaced and fixed a real race condition. - Mock HTTP server (
scripts/qa/lib/mock_http.py) — stdlib-only fake ollama/anthropic endpoints for testing successful compaction offline. - 6 e2e scenarios — happy path (all 9 adoptions in o
No comments yet
Be the first to share your take.