Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Awesome List
Awesome list for AI agent harness engineering: tools, patterns, evals, memory, MCP, permissions, observability, and orchestration.
GitHub stars and default-branch commits for ai-boost/awesome-harness-engineering.
158 repos currently saved from this list.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
Hindsight: Agent Memory That Learns
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
An open-source, code-first Python toolkit for building, evaluating, and deploying sophisticated AI agents with flexibility and control.
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
SWE-agent takes a GitHub issue and tries to automatically fix it, using your LM of choice. It can also be employed for offensive cybersecurity or competitive coding challenges. [NeurIPS 2024]
Fast, efficient, battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipelines + LLM Agent, precise line-level comments, built-in multi-language ruleset (NPE, thread-safety, XSS, SQL injection), OpenAI & Anthropic compatible.
Context window optimization for AI coding agents. Sandboxes tool output (98% reduction), persists session memory, and enforces routing across 17 platforms via MCP + hooks.
AI Agent Framework, the Pydantic way
The LLM Evaluation Framework
A collection of projects designed to help developers quickly get started with building deployable applications using the Claude API
Browser Harness | Self-healing harness that enables LLMs to complete any task.
SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.
OpenWiki is a CLI that writes and maintains agent documentation for your codebase.
AG-UI: the Agent-User Interaction Protocol. Bring Agents into Frontend Applications.
Structured Outputs
"OpenHarness: Open Agent Harness with a Built-in Personal Agent--Ohmo!"
Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞
Open Source framework for voice and multimodal conversational AI
Open-source, secure environment with real-world tools for enterprise-grade agents.
The best agent harness.
IronClaw is an Agent OS focused on privacy, security and extensibility
Multiplayer agent harness for work
Multi-Agent Harness for Production AI
Build your own AI SRE agents. The open source toolkit for the AI era.
AI Observability & Evaluation
Secure, Fast, and Extensible Sandbox runtime for AI agents.
Instant, Concurrent, Secure & Lightweight Sandbox for AI Agents.
Visual testing tool for MCP servers
Free, local tool to track AI coding token usage and cost across 37 tools and agents (Claude Code, Cursor, Codex, Gemini and more), by model, project, and task. npx codeburn
Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.
A meta-skill that designs domain-specific agent teams, defines specialized agents, and generates the skills they use.
Build effective agents using Model Context Protocol and simple workflow patterns
The sandbox agent framework.
OpenShell is the safe, private runtime for autonomous AI agents.
Deepsec is a security harness for finding vulnerabilities in your codebase powered by coding agents
Open-source observability for your GenAI or LLM application, based on OpenTelemetry
Orchestrate sandboxed coding agents in TypeScript with sandcastle.run()
Build an agent harness and control it end-to-end. Open-source SDK for production AI agents in Python & TypeScript - any model, any cloud.
NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.
TypeScript AI agent orchestration framework with dynamic workflows. Describe the goal, not the graph: a coordinator plans the task DAG at runtime and runs it on any LLM (Claude, ChatGPT, Gemini, DeepSeek, or local models).
[EMNLP'23, ACL'24] To speed up LLMs' inference and enhance LLM's perceive of key information, compress the prompt and KV-Cache, which achieves up to 20x compression with minimal performance loss.
Persistent memory system for AI coding agents. Agent-agnostic Go binary with SQLite + FTS5, MCP server, HTTP API, CLI, and TUI.
The first "code-first" agent framework for seamlessly planning and executing data analytics tasks.
Self-hosted control plane for AI agents: dispatch tasks, review runs, track spend, and operate OpenClaw, Claude Code, Codex, and other runtimes.
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
Fast and Accurate Code Search for Agents. Uses 99% fewer tokens than grep+read
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
Nexent is a zero-code platform for auto-generating production-grade AI agents using Harness Engineering principles — unified tools, skills, memory, and orchestration with built-in constraints, feedback loops, and control planes.
AI Agent Governance Toolkit — Policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering for autonomous AI agents. Covers 10/10 OWASP Agentic Top 10.