Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Awesome List
An awesome list of Agent Harness engineering resources, including GitHub projects, tools, benchmarks, and practical guides.
GitHub stars and default-branch commits for Picrew/awesome-agent-harness.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
AI-Driven Life Cycle (AI-DLC) adaptive workflow steering rules for AI coding agents
Framework for evaluating and improving agents
Open-source security automation platform for teams and AI agents
Agent Skills as a Memory Layer
LangChain 🔌 MCP
Evaluation and Tracking for LLM Experiments and AI Agents
Devon: An open-source pair programmer
The platform for LLM evaluations and AI agent testing
agent-sandbox enables easy management of isolated, stateful, singleton workloads, ideal for use cases like AI agent runtimes.
Amazon Bedrock Agentcore accelerates AI agents into production with the scale, reliability, and security, critical to real-world deployment.
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
Security scanner for AI agents, MCP servers and agent skills.
The NVIDIA NeMo Agent toolkit is an open-source library for efficiently connecting and optimizing teams of AI agents.
A benchmark for LLMs on complicated tasks in the terminal
A production-oriented multi-agent orchestration framework.
Terminal coding agent powered by Kimchi's multi-model orchestration
How real engineers run Claude Code and Codex: spec-driven planning, enforced TDD, persistent memory, and quality enforcement on all levels. Make your agents production-ready.
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Code repo for "WebArena: A Realistic Web Environment for Building Autonomous Agents"
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
E2B Desktop Sandbox for LLMs. E2B Sandbox with desktop graphical environment that you can connect to any LLM for secure computer use.
Reference code for the Meta-Harness paper.
The Cloud Sandbox Built for AI Agents
OpenTelemetry Instrumentation for AI Observability
Evaluate and improve models and agents using environments
Minimal and readable coding agent harness implementation in Python to explain the core components of coding agents.
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers (aka The Norton for OpenClaw)
CUGA is an open-source generalist agent harness for the enterprise, supporting complex task execution on web and APIs, OpenAPI/MCP integrations, composable architecture, reasoning modes, and policy-aware features.
No description.
Your agent's favorite harness, built on Pydantic AI
A production-ready runtime framework for agent apps with secure tool sandboxing, Agent-as-a-Service APIs, scalable deployment, full-stack observability, and broad framework compatibility.
Security Governance for Agentic AI
Official AHE code — Agentic Harness Engineering: observability-driven automatic evolution of coding-agent harnesses (concurrent w/ meta-harness). NexAU-AHE reaches 84.7% ± 2.1 pass@1 on Terminal-Bench 2 (GPT-5.5). Lifts GPT-5.4 69.7→77.0% over 10 iters, beats Codex/ACE/Training-Free GRPO; frozen harness transfers to SWE-bench-Verified.
An agent benchmark with tasks in a simulated software company.
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
Collection of evals for Inspect AI
Sandboxed code execution for AI agents, locally or on the cloud. Massively parallel, easy to extend. Powering SWE-agent and more.
Open-source benchmark for browser AI agents on daily tasks.
The action firewall for AI agents. Enforce policy and human approval before risky tool calls, shell commands, workflows, and production changes, with auditable evidence.
A generative AI-powered framework for testing virtual agents.
The production-ready agent harness framework for Python
Secure runtime to sandbox AI agent tasks. Run untrusted code in isolated WebAssembly environments.
WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?
Open Python agent harness for production AI apps: tools, MCP, memory, workspace, telemetry, subagents, background tasks, and OmniServe APIs.
Tandem is the authority layer between AI agents and company tools: runtime policy for tools, memory, approvals, and audit trails.
Evaluation harness for OpenHands V1.
No description.
A verified version of the WebArena Benchmark
Benchmark for comparing agent harnesses on everyday online tasks — fixes the base model, varies the harness. Sister project of ClawBench, same scoring pipeline.