— · MIT
Agent OS: the agent gets smarter on its own. We just hold the line: the grading command and expected result never make it into the success contract we hand it. Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 13 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.
v0.3.1 · Apache-2.0
Log Claude Code sessions to Opik, the open-source LLM observability and evaluation platform, built by Comet. Tracing, evaluation, and skills for observable AI applications.
— · Apache-2.0
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
— · MIT
Know if your agent skill actually works. A lightweight evaluation harness that tracks a success rate across Claude Code, Codex, Pi, and Hermes.
v0.15.0 · MIT
Open-source Agent Skills for Claude Code and Codex: ship-DAG orchestration, A-F code review, AI evals, CI gates, design, copy, SEO/AEO/GEO, Instagram growth, app shipping, creator rights, and consumer recovery.
— · MIT
Agentic RAG for local and self-hosted document search: hybrid retrieval, reranking and multimodal RAG on embedded LanceDB, with Docling parsing and an MCP server
v0.54.0 · MIT
OMK — Observe. Measure. Know. Evidence-backed knowledge changes for AI applications.
v0.7.2 · MIT
DeepSeek Harness plugin and bundled Skill for safely evolving Cordis Candidates with Harbor.
— · no license
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
— · MIT
One-stop handbook for building, deploying, and understanding LLM agents with 60+ skeletons, tutorials, ecosystem guides, and evaluation tools.
— · MIT
Provides ServiceNow ITSM tools (incidents, knowledge, CMDB, user lookups) to AI agents, with an evaluation harness that measures tool-selection accuracy, task completion, and latency.