HEWN 2.0 2026: AI Output Router for Precision Summaries & Polished Code
#benchmarking
9 posts
Library
150
HTML Updated 2d ago
kaderkck
0
0
Claude Skill
12
AI Agent plugin for Autoresearch with AI (Claude, OpenClaw, etc) to improve anything!
Shell MIT Updated 1w ago
proyecto26
0
0
Claude Skill
33
Use cultivar to test your Agent Skills, run them in sandboxes, and across different agents.
Python MIT Updated 1w ago
pinecone-io
0
0
MCP Server
168
Dialectical reasoning architecture for LLMs (Thesis → Antithesis → Synthesis)
Python MIT Updated 4mos ago
hmbown
0
0
MCP Server
146
Goku is an HTTP load testing application written in Rust
Rust MIT Updated 3w ago
jcaromiq
0
0
MCP Server
19
Auditable context capsules for LLM handoffs, coding agents, and OpenCode MCP workflows.
Python MIT Updated 1mo ago
raincherb
0
0
Claude Skill
14
A benchmarking harness for coding agents.
Python MIT Updated 2w ago
swival
0
0
MCP Server
40
Iterative agent harness improvement: run a coding agent on a hard task, generate the reusable tooling it was missing, qualify it, and replay...
Python Apache-2.0 Updated 2w ago
patrick-toulme
0
0
MCP Server
13
The open testing standard for voice AI agents. Deterministic + semantic + RAG augmented evaluation. Local first. Zero telemetry.
Python NOASSERTION Updated 4w ago