Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Awesome List
Awesome list for AI agent harness engineering: tools, patterns, evals, memory, MCP, permissions, observability, and orchestration.
GitHub stars and default-branch commits for ai-boost/awesome-harness-engineering.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.
Model Context Protocol Servers
An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of tasks that could take minutes to hours.
The Memory Layer for AI Agents - Drop-in memory infrastructure for AI agents and apps. Context that persists. Built for production.
A programming framework for agentic AI
The best-benchmarked open-source AI memory system. And it's free.
Framework for orchestrating role-playing, autonomous AI agents. By fostering collaborative intelligence, CrewAI empowers agents to work together seamlessly, tackling complex tasks.
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
aider is AI pair programming in your terminal
Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps
Installable GitHub library of 1,800+ agentic skills for Claude Code, Cursor, Codex CLI, Gemini CLI, Antigravity, and more. Includes specialized plugins, installer CLI, bundles, workflows, and official/community skill collections.
Build resilient agents.
Multi-harness agentic plugin marketplace for Claude Code, Codex CLI, Cursor, OpenCode, GitHub Copilot, and Gemini CLI
Teams-first Multi-agent orchestration for Claude Code
Self-evolving Context Database for AI Agents. Unify Agent Memory, Knowledge RAG and Skills.
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers. Library, proxy, MCP server.
Cognee is the open-source AI memory platform for agents. Give your AI agents persistent long-term memory across sessions with a self-hosted knowledge graph engine.
Composio powers 1000+ toolkits, tool search, context management, authentication, and a sandboxed workbench to help you build AI agents that turn intent into action.
A lightweight, powerful framework for multi-agent workflows
The batteries-included agent harness.
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
Hindsight: Agent Memory That Learns
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
An open-source, code-first Python toolkit for building, evaluating, and deploying sophisticated AI agents with flexibility and control.
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
SWE-agent takes a GitHub issue and tries to automatically fix it, using your LM of choice. It can also be employed for offensive cybersecurity or competitive coding challenges. [NeurIPS 2024]
AI Agent Framework, the Pydantic way
The LLM Evaluation Framework
A collection of projects designed to help developers quickly get started with building deployable applications using the Claude API
SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.
AG-UI: the Agent-User Interaction Protocol. Bring Agents into Frontend Applications.
Structured Outputs
"OpenHarness: Open Agent Harness with a Built-in Personal Agent--Ohmo!"
Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞
Open Source framework for voice and multimodal conversational AI
Open-source, secure environment with real-world tools for enterprise-grade agents.
IronClaw is an Agent OS focused on privacy, security and extensibility
Multi-Agent Harness for Production AI
Build your own AI SRE agents. The open source toolkit for the AI era.
AI Observability & Evaluation
Secure, Fast, and Extensible Sandbox runtime for AI agents.
Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.
Build effective agents using Model Context Protocol and simple workflow patterns
OpenShell is the safe, private runtime for autonomous AI agents.
Open-source observability for your GenAI or LLM application, based on OpenTelemetry
Build an agent harness and control it end-to-end. Open-source SDK for production AI agents in Python & TypeScript - any model, any cloud.
NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.
The first "code-first" agent framework for seamlessly planning and executing data analytics tasks.
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓