Curated, ranked list of AI agent harnesses — the runtimes that close the loop between a stateless model and the outside world.
150+ harnesses · 12 categories · weekly-rescored · MCP-ready :
claude mcp add agent-harnesses -- uvx agent-harnesses-mcp| Project | Stars | Category | Tier | OSS |
|---|---|---|---|---|
| OpenClaw typescriptmulti-agent | 386k | Personal agent runtimes | complex | ✅ |
| superpowers memorycliide | 273k | Coding harness configs and SDKs | complex | ✅ |
| Hermes memorypythonprovider-agnostic | 231k | Personal agent runtimes | slightly complex | ✅ |
| n8n workflowlocaltypescript | 201k | Frameworks | complex | ⚠️ |
| opencode mcpprovider-agnosticclituitypescript | 198k | Coding agent products (IDEs, CLIs, full suites) | slightly complex | ✅ |
| AutoGPT memoryevalspython | 187k | Frameworks | complex | ⚠️ |
| Anthropic Skills | 170k | Coding harness configs and SDKs | mostly simple | ⚠️ |
| langflow low-codepython | 153k | Frameworks | complex | ✅ |
| Dify low-coderagpython | 153k | Frameworks | complex | ⚠️ |
| langchain python | 144k | Frameworks | complex | ✅ |
| GStack typescript | 128k | Coding harness configs and SDKs | slightly complex | ✅ |
| browser-use browserpython | 109k | Frameworks | slightly complex | ✅ |
| Gemini CLI mcpclitypescript | 107k | Coding agent products (IDEs, CLIs, full suites) | slightly complex | ✅ |
| Codex sandboxprovider-agnosticcli | 106k | Coding agent products (IDEs, CLIs, full suites) | slightly complex | ✅ |
| pi provider-agnostictuirust | 91.3k | Coding agent products (IDEs, CLIs, full suites) | slightly complex | ✅ |
| claude-mem memory | 90.9k | Memory and state | slightly complex | ✅ |
| MCP Servers mcpmemorytypescript | 89.6k | Plugins, MCPs, CLI tools | mostly simple | ✅ |
| addyosmani/agent-skills workflowide | 87.7k | Coding harness configs and SDKs | mostly simple | ✅ |
| OpenHands memorybrowsersandboxpython | 84.2k | Coding agent products (IDEs, CLIs, full suites) | complex | ⚠️ |
| DeerFlow memorymulti-agentsandboxpython | 80.1k | Research and task-specific harnesses | complex | ✅ |
| Daytona sandbox | 72k | Libraries and SDKs | slightly complex | ✅ |
| MetaGPT multi-agentpython | 69.9k | Multi-agent and orchestration | complex | ✅ |
| Open Interpreter clipython | 68k | Coding agent products (IDEs, CLIs, full suites) | mostly simple | ✅ |
| Headroom mcprag | 66.5k | Progressive disclosure harnesses | mostly simple | ✅ |
| Cline idetypescript | 66.3k | Coding agent products (IDEs, CLIs, full suites) | slightly complex | ✅ |
| Mem0 memorypython | 63.4k | Memory and state | slightly complex | ✅ |
| Context7 mcptrainingtypescript | 60.8k | Plugins, MCPs, CLI tools | super simple | ✅ |
| autogen multi-agentpython | 60.5k | Multi-agent and orchestration | complex | ✅ |
| OpenManus multi-agentpython | 58k | Multi-agent and orchestration | complex | ✅ |
| crewAI python | 57.2k | Multi-agent and orchestration | complex | ✅ |
| LiteLLM provider-agnosticpython | 56.5k | Libraries and SDKs | mostly simple | ✅ |
| goose mcprust | 52.9k | Coding agent products (IDEs, CLIs, full suites) | slightly complex | ✅ |
| awesome-claude-code | 52.4k | Coding harness configs and SDKs | super simple | ⚠️ |
| llama-index ragpython | 51.7k | Frameworks | complex | ✅ |
| chrome-devtools-mcp mcpbrowsertypescript | 49.3k | Plugins, MCPs, CLI tools | mostly simple | ✅ |
| aider mcpclipython | 48.3k | Plugins, MCPs, CLI tools | slightly complex | ✅ |
| nanobot mcpmemorylocalpython | 47.1k | Personal agent runtimes | mostly simple | ✅ |
| CowAgent memorypython | 46.5k | Personal agent runtimes | slightly complex | ✅ |
| agno memoryevalspython | 41.7k | Frameworks | complex | ✅ |
| awesome-cursorrules ide | 40.6k | Progressive disclosure harnesses | super simple | ✅ |
| langgraph workflowpython | 39.8k | Frameworks | slightly complex | ✅ |
| wshobson/agents multi-agentcliide | 38.9k | Coding harness configs and SDKs | super simple | ✅ |
| Khoj python | 36.5k | Personal agent runtimes | complex | ✅ |
| Playwright MCP mcpvisionbrowsertypescript | 36.2k | Plugins, MCPs, CLI tools | mostly simple | ✅ |
| continue idetypescript | 35.5k | Plugins, MCPs, CLI tools | complex | ✅ |
| DeepSeek-Reasonix memoryclituitypescript | 34.6k | Coding agent products (IDEs, CLIs, full suites) | slightly complex | ✅ |
| ChatDev python | 34k | Multi-agent and orchestration | slightly complex | ✅ |
| Langfuse evalstypescript | 33.2k | Observability and eval-ops | slightly complex | ✅ |
| github-mcp-server mcp | 32.3k | Plugins, MCPs, CLI tools | slightly complex | ✅ |
| cognee memoryragworkflowpython | 30.1k | Memory and state | slightly complex | ✅ |
| Graphiti (Zep) memoryragworkflowpython | 30k | Memory and state | slightly complex | ✅ |
| Composio sandboxtool-discoverypythontypescript | 29.7k | Libraries and SDKs | complex | ✅ |
| gpt-researcher multi-agentpython | 29k | Research and task-specific harnesses | complex | ✅ |
| smolagents sandboxpython | 28.8k | Libraries and SDKs | mostly simple | ✅ |
| openai-agents-python python | 28.7k | Multi-agent and orchestration | mostly simple | ✅ |
| semantic-kernel python | 28.5k | Frameworks | complex | ✅ |
| vibe-kanban | 27.8k | Coding agent products (IDEs, CLIs, full suites) | slightly complex | ✅ |
| deepagents multi-agentsandboxpythontypescript | 27.8k | Libraries and SDKs | slightly complex | ✅ |
| MLflow evalspython | 27.5k | Observability and eval-ops | complex | ✅ |
| crush memoryclitui | 27.4k | Coding agent products (IDEs, CLIs, full suites) | slightly complex | ⚠️ |
| mastra typedtypescript | 27.2k | Frameworks | slightly complex | ⚠️ |
| qwen-code sandboxclitypescript | 27.1k | Coding agent products (IDEs, CLIs, full suites) | slightly complex | ✅ |
| Kilo Code mcpcliidetypescript | 26.9k | Coding agent products (IDEs, CLIs, full suites) | slightly complex | ✅ |
| Symphony sandbox | 26.7k | Coding agent products (IDEs, CLIs, full suites) | complex | ✅ |
| beads memory | 26.4k | Memory and state | mostly simple | ✅ |
| Haystack memoryragpython | 26.2k | Frameworks | complex | ✅ |
| vercel/ai provider-agnostictypescript | 26.2k | Libraries and SDKs | slightly complex | ✅ |
| planning-with-files memory | 26.2k | Coding harness configs and SDKs | mostly simple | ✅ |
| oh-my-pi browserprovider-agnosticcliiderust | 25.2k | Coding agent products (IDEs, CLIs, full suites) | slightly complex | ✅ |
| Roo Code mcpworkflowidetypescript | 24.3k | Coding agent products (IDEs, CLIs, full suites) | slightly complex | ✅ |
| letta memorypython | 24.3k | Frameworks | mostly simple | ✅ |
| MCP Python SDK mcppython | 24k | Plugins, MCPs, CLI tools | mostly simple | ✅ |
| Stagehand browsertypescript | 24k | Frameworks | slightly complex | ✅ |
| agents.md idetypescript | 23.7k | Progressive disclosure harnesses | super simple | ✅ |
| Opik evalspython | 21.4k | Observability and eval-ops | slightly complex | ✅ |
| rasa voicepython | 21.3k | Frameworks | complex | ✅ |
| Google ADK evalssandboxpython | 21.1k | Frameworks | complex | ✅ |
| SWE-agent memoryevalspython | 20.1k | Coding harness configs and SDKs | slightly complex | ✅ |
| context-mode mcpmemorysandbox | 19.9k | Progressive disclosure harnesses | mostly simple | ⚠️ |
| pydantic-ai mcptypedprovider-agnosticpython | 19.3k | Libraries and SDKs | slightly complex | ✅ |
| Eliza memorymulti-agenttypescript | 19.1k | Personal agent runtimes | complex | ✅ |
| Agent Zero memorymulti-agentbrowsersandboxpython | 18.9k | Personal agent runtimes | slightly complex | ✅ |
| jcode mcpmemoryprovider-agnosticclirust | 17.7k | Coding agent products (IDEs, CLIs, full suites) | slightly complex | ✅ |
| Agent Lightning evalstrainingpython | 17.5k | Evaluation and benchmarking harnesses | complex | ✅ |
| OpenHarness (HKUDS) memorymulti-agent | 15.4k | Personal agent runtimes | complex | ✅ |
| eigent multi-agentlocal | 15k | Coding agent products (IDEs, CLIs, full suites) | complex | ✅ |
| botpress low-codetypescript | 14.9k | Frameworks | complex | ✅ |
| cc-haha memorymulti-agenttypescript | 14.1k | Coding agent products (IDEs, CLIs, full suites) | complex | ✅ |
| AutoResearchClaw multi-agent | 14k | Research and task-specific harnesses | complex | ✅ |
| E2B sandboxpython | 13.4k | Libraries and SDKs | slightly complex | ✅ |
| MCP TypeScript SDK mcptypescript | 13.2k | Plugins, MCPs, CLI tools | mostly simple | ✅ |
| Microsoft Agent Framework multi-agentworkflowpython | 12.8k | Multi-agent and orchestration | slightly complex | ✅ |
| Arize Phoenix evalspython | 11.1k | Observability and eval-ops | slightly complex | ⚠️ |
| hive multi-agentpython | 10.9k | Multi-agent and orchestration | complex | ✅ |
| MCP Inspector mcptypescript | 10.7k | Plugins, MCPs, CLI tools | super simple | ✅ |
| omnigent sandboxidepython | 8.9k | Multi-agent and orchestration | complex | ✅ |
| PraisonAI multi-agentpython | 8.9k | Multi-agent and orchestration | mostly simple | ✅ |
| MiroThinker evals | 8.4k | Research and task-specific harnesses | slightly complex | ✅ |
| get-shit-done clipython | 8.3k | Coding harness configs and SDKs | mostly simple | ✅ |
| R2R visionragworkflowpython | 8k | Frameworks | complex | ✅ |
| Claude Agent SDK mcpmemorypythontypescript | 7.9k | Coding harness configs and SDKs | complex | ✅ |
| agent-squad multi-agent | 7.7k | Frameworks | slightly complex | ✅ |
| Steel memorybrowserlocal | 7.5k | Libraries and SDKs | slightly complex | ✅ |
| MCP Registry mcp | 7.2k | Plugins, MCPs, CLI tools | slightly complex | ✅ |
| strands-agents mcpmulti-agenttypedpython | 6.9k | Libraries and SDKs | mostly simple | ✅ |
| Agent Governance Toolkit sandboxpython | 6k | Plugins, MCPs, CLI tools | slightly complex | ✅ |
| SWE-bench evalssandboxpython | 5.6k | Evaluation and benchmarking harnesses | slightly complex | ✅ |
| agents-cli evalscli | 5.6k | Coding harness configs and SDKs | mostly simple | ✅ |
| Cloudflare Agents memorytypescript | 5.4k | Libraries and SDKs | slightly complex | ✅ |
| AgentVerse multi-agentpython | 5.1k | Frameworks | complex | ✅ |
| skillhub localcliide | 4.9k | Coding harness configs and SDKs | mostly simple | ✅ |
| AG2 multi-agentpython | 4.9k | Multi-agent and orchestration | complex | ✅ |
| youtu-agent | 4.6k | Frameworks | mostly simple | ✅ |
| mcp-context-forge mcppython | 4.3k | Plugins, MCPs, CLI tools | complex | ✅ |
| AgentBench evalssandboxragworkflowpython | 3.7k | Evaluation and benchmarking harnesses | complex | ✅ |
| openai-agents-js multi-agentvoicetypescript | 3.6k | Libraries and SDKs | slightly complex | ✅ |
| Agent Sandbox memorysandboxlocal | 3.5k | Libraries and SDKs | slightly complex | ✅ |
| Bee Agent Framework mcpmulti-agentpythontypescript | 3.4k | Frameworks | complex | ✅ |
| cocoindex-code mcpcli | 2.6k | Plugins, MCPs, CLI tools | mostly simple | ✅ |
| inspect_ai evalssandboxpython | 2.6k | Evaluation and benchmarking harnesses | complex | ✅ |
| AgentStack | 2.2k | Frameworks | slightly complex | ✅ |
| agent-vault | 2.1k | Plugins, MCPs, CLI tools | mostly simple | ⚠️ |
| WebArena python | 1.6k | Evaluation and benchmarking harnesses | complex | ✅ |
| Docker MCP Gateway mcpsandboxcli | 1.5k | Plugins, MCPs, CLI tools | slightly complex | ✅ |
| Meta-Harness | 1.4k | Coding harness configs and SDKs | slightly complex | ✅ |
| AIlice sandboxpython | 1.4k | Personal agent runtimes | slightly complex | ✅ |
| WebVoyager evalsvision | 1.1k | Evaluation and benchmarking harnesses | slightly complex | ✅ |
| agent-qa mcpmemorysandboxclitypescript | 843 | Evaluation and benchmarking harnesses | slightly complex | ⚠️ |
| swe-smith trainingpython | 742 | Evaluation and benchmarking harnesses | slightly complex | ✅ |
| ARC-AGI-2 | 733 | Evaluation and benchmarking harnesses | super simple | ✅ |
| SWE-Gym evalstrainingpython | 721 | Evaluation and benchmarking harnesses | slightly complex | ✅ |
| inspect_evals evalssandbox | 627 | Evaluation and benchmarking harnesses | slightly complex | ✅ |
| open-harness mcpmulti-agenttypescript | 598 | Libraries and SDKs | slightly complex | ✅ |
| langgraph-bigtool tool-discoverypython | 554 | Progressive disclosure harnesses | slightly complex | ✅ |
| claw-code-agent mcprustpythontypescript | 543 | Coding agent products (IDEs, CLIs, full suites) | slightly complex | ❓ |
| RepoMaster workflowpython | 542 | Coding harness configs and SDKs | slightly complex | ❓ |
| MCP-Zero tool-discovery | 506 | Progressive disclosure harnesses | complex | ✅ |
| Terminal-Bench evalsclipython | 499 | Evaluation and benchmarking harnesses | slightly complex | ✅ |
| AgentSilex python | 454 | Frameworks | super simple | ✅ |
| openagents | 446 | Research and task-specific harnesses | complex | ✅ |
| AutoHarness memorymulti-agentprovider-agnosticpython | 367 | Coding harness configs and SDKs | super simple | ✅ |
| arc-agi-benchmarking evalsprovider-agnosticpython | 362 | Evaluation and benchmarking harnesses | mostly simple | ✅ |
| AgentBox sandboxtypescript | 352 | Coding agent products (IDEs, CLIs, full suites) | slightly complex | ✅ |
| AgentRL trainingpython | 339 | Multi-agent and orchestration | complex | ✅ |
| SuperAgentX multi-agentpython | 203 | Frameworks | mostly simple | ✅ |
| ToolGen tool-discoverypython | 184 | Progressive disclosure harnesses | complex | ✅ |
| Proliferate multi-agentsandboxidetypescript | 167 | Coding agent products (IDEs, CLIs, full suites) | complex | ✅ |
| VitaBench | 164 | Evaluation and benchmarking harnesses | complex | ✅ |
| LoopTroop typescript | 119 | Coding harness configs and SDKs | mostly simple | ✅ |
| AgencyBench evalssandboxpython | 93 | Evaluation and benchmarking harnesses | complex | ✅ |
| letta-evals memorypython | 83 | Evaluation and benchmarking harnesses | mostly simple | ✅ |
| Talon mcpmemoryclitypescript | 71 | Personal agent runtimes | slightly complex | ✅ |