Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Awesome List
An awesome list of Agent Harness engineering resources, including GitHub projects, tools, benchmarks, and practical guides.
GitHub stars and default-branch commits for Picrew/awesome-agent-harness.
332 repos currently saved from this list.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
local multi-agent harness
Agent harness and tookit for building AI agents and agentic applications. CLI and SDKs included
Collection of evals for Inspect AI
Open-source cross-agent memory layer for coding agents via MCP. Compatible with Claude Code, Codex, Cursor, Windsurf, Gemini CLI, Antigravity, OpenClaw, Hermes Agent, Oh-my-Pi, Pi, Copilot, Kiro, OpenCode, and Trae.
🤖 MateClaw — Your second brain with Multi-Agent Orchestration, MCP Protocol, Skills & Memory, Dream, and Multi-Channel Support. Built on Spring AI Alibaba.
Sandboxed code execution for AI agents, locally or on the cloud. Massively parallel, easy to extend. Powering SWE-agent and more.
Open-source benchmark for browser AI agents on daily tasks.
Bring your own agent and build a self-improving agentic system. Automatically mine failures, optimize the agent harness, and gate against regressions.
The action firewall for AI agents. Enforce policy and human approval before risky tool calls, shell commands, workflows, and production changes, with auditable evidence.
An in-the-wild benchmark for AI agents in the OpenClaw Environment.
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Self-hosted Personal AI + agent runtime in .NET (NativeAOT-friendly)
Safe runtime for autonomous on-chain AI agents: isolated sandboxes, Library skills, encrypted secrets, and OKX read-only security checks.
Deterministic safety solutions for probabilistic AI agents
A collection of Model Context Protocol (MCP) servers, clients and developer tools by IBM.
A generative AI-powered framework for testing virtual agents.
The production-ready agent harness framework for Python
Secure runtime to sandbox AI agent tasks. Run untrusted code in isolated WebAssembly environments.
🛡️The governance runtime for AI agents. Intercept actions, enforce guard policies, require approvals, and produce audit-ready decision trails.
WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?
Open Python agent harness for production AI apps: tools, MCP, memory, workspace, telemetry, subagents, background tasks, and OmniServe APIs.
Secure local dev environment for collaboration with AI coding agents
The harness layer for Claude Code — a reference implementation of harness engineering with hook-enforced dual review, state-machine gates that survive context compaction, and fail-closed safety where it counts. Quality gates that AI can't skip.
Universally Triggered Agent Harness - An OpenClaw-like Inngest-powered personal agent
Runtime for long-horizon agents
HexAgent – An Agent harness that gives any LLM a computer to complete tasks the way humans do
Tandem is the authority layer between AI agents and company tools: runtime policy for tools, memory, approvals, and audit trails.
Evaluation harness for OpenHands V1.
This repository defines AGENT.md, a standardized format that lets your codebase speak directly to any agentic coding tool.
No description.
A verified version of the WebArena Benchmark
Benchmark for comparing agent harnesses on everyday online tasks — fixes the base model, varies the harness. Sister project of ClawBench, same scoring pipeline.