MCP servers expose tools with no information about what they actually do at runtime. mcpsafetywarden sits between your agent and any MCP ser...
#ai-safety
80 posts
MCP & Claude Code security scanner — threat-models plugins, MCP servers, hooks, skills & connectors with an LLM before you trust them. Catch...
Glass Box Framework — runtime constitutional verification for AI answers. Trust Cards with claim-level reasoning chains, formal ECS scoring,...
Runtime artifact existence & freshness verification for AI agent completion claims — a lightweight, zero-LLM MCP gate that source-binds 'don...
AI Red Teaming / AI Safety に関する日本語リソースのキュレーションリスト
A safer MySQL CLI for AI coding agents: connection profiles, SSH tunnels, and automatic sensitive-data masking before query output reaches C...
🛡️ A curated list of resources on agent skills security: attacks, defenses, frameworks, and benchmarks for securing AI agent tool use and sk...
Open-source prompt injection detector — 5 layers, 91.7% F1, ~27ms, offline, Apache 2.0
Deterministic policy language for AI agents. Z3 + TLA+ dual-engine formal verification. Runtime enforcement <1ms.
Pre-execution policy engine for AI agents. Every tool call checked before execution.
All AIs are sycophants.
MCP EU AI Act Compliance Scanner - Open source tool to detect EU AI Act violations in codebases