🐀 Fuzz your verifier before an RL agent does. Static + dynamic LLM security auditor to detect reward-hacking in RL post-training environmen...
#ai-safety
80 posts
Self-improving control kernel for AI coding agents (DAGx AGI Kernel): hard approval stops for irreversible actions, evidence-gated completio...
A small, vetted, self-evolving harness for Claude Code — curated, not dumped. Skills, safety guards, and a self-improvement loop.
Open skills, harnesses, hooks, and verifiers for giving probabilistic AI a control layer.
An AI coding agent guardrail — a CLI hook that blocks destructive git and filesystem commands and secret file access before they execute. Su...
An unreleased internal OpenAI model, very likely to be called GPT-6, was able to autonomously break out of its sandbox AND break into Huggin...
Make any coding agent work like a frontier model. Drop-in Agent Skills for disciplined planning, evidence-first debugging, and live-system s...
Subagent Verification for Claude AI Code Networks 2026
Automated Proof-of-Carrying Change Management for AIOps 2026
Runtime safety for AI coding agents with real-time enforcement, system-event monitoring, and long-horizon provenance. Supports Claude Code,...
Andrej Karpathy Joins Anthropic's DOOMSDAY Team, AI is a Math Genius, Live Callers! —Livestream 5/22
***ATTENTION: We’re looking for paid help at LessOnline from June 5-7. See link in the pinned comment!*** OpenAI pushes the frontier of mat...
Learn what AI researchers mean when they talk about sycophancy, when it's more likely to show up in conversations, and tactics you can use t...