Agentic security control plane for MCP and AI agent tool calls. MCP-native policy gateway with topology discovery and audit.
#ai-safety
74 posts
Dialectical reasoning architecture for LLMs (Thesis → Antithesis → Synthesis)
Security enforcement plugin for Claude Code. Blocks dangerous commands, audits every tool call, detects prompt injection.
LikenessGuard is an open-source reference project for pre-generation consent enforcement in AI image generation. V2 demonstrates the concept...
Your AI agent just burned $200. AgentGuard stops it at $5. Runtime cost guardrails for AI agents — budget enforcement, loop detection, kill...
Proof-backed AI agent for checking suspicious job posts, recruiter messages, and apply links, now live with case study and pilot intake.
MCP servers expose tools with no information about what they actually do at runtime. mcpsafetywarden sits between your agent and any MCP ser...
MCP & Claude Code security scanner — threat-models plugins, MCP servers, hooks, skills & connectors with an LLM before you trust them. Catch...
Glass Box Framework — runtime constitutional verification for AI answers. Trust Cards with claim-level reasoning chains, formal ECS scoring,...
Runtime artifact existence & freshness verification for AI agent completion claims — a lightweight, zero-LLM MCP gate that source-binds 'don...
AI Red Teaming / AI Safety に関する日本語リソースのキュレーションリスト
A safer MySQL CLI for AI coding agents: connection profiles, SSH tunnels, and automatic sensitive-data masking before query output reaches C...