v0.3.0 · MIT
Make every model cheaper or better. Zero-config behavioral harness for Claude Code, with a reproducible 408-run benchmark
— · no license
DeepSeek V4 × J-Space capability realization report — benchmark evidence that J-Space reduces capability-realization loss on DeepSeek V4.
v0.54.0 · MIT
OMK — Observe. Measure. Know. Evidence-backed knowledge changes for AI applications.
v0.6.10 · MIT
Logic-first AI code review via semi-formal execution tracing (Premises → Trace → Divergence → Trigger → Remedy). Catches behavioral bugs, type-contract breaches & async hazards that linters miss. Six skills · Claude Code · Codex CLI · Gemini CLI.
— · Apache-2.0
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
— · MIT
Runtime evidence that helps agents trace, profile, and burn down hotspots in application and native code, GPU kernels, and inference stacks.
v0.3.3 · MIT
Autoharness — a self-learning skill layer for Claude Code — distills skills from your real sessions, updates them as you work, and prunes the ones that stop getting used. No daemon, no benchmark.
v0.32.0 · AGPL-3.0
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
v1.2.1 · MIT
Eco mode for Claude Code. /eco: -31% to -73% output tokens with critical findings intact; /eco-max: up to -75% with lowered effort. Measured hardest on Claude Fable 5 (fable5), deep-studied on Sonnet 5, works on Opus 4.8 too. We publish our negative results. 82 raw benchmark runs.
v2.0.0 · MIT
Multi-LLM entity enrichment: schemas, single/batch enrichment, fusion, model benchmarks.
— · Apache-2.0
The vendor-neutral protocol for AI agent safety & security — and the neutral benchmark that ranks the vendors.