Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Awesome List
An awesome list of Agent Harness engineering resources, including GitHub projects, tools, benchmarks, and practical guides.
GitHub stars and default-branch commits for Picrew/awesome-agent-harness.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
Nexent is a zero-code platform for auto-generating production-grade AI agents using Harness Engineering principles — unified tools, skills, memory, and orchestration with built-in constraints, feedback loops, and control planes.
Open-source All in One AI agent workspace. Run any agent — Claude Code, Codex — across your tools (100+ integrations + MCP), apps, browser, and files, with shared memory. Built-in models or BYOK.
ByteRover CLI (brv) - The portable memory layer for autonomous coding agents (formerly Cipher)
TencentDB Agent Memory delivers fully local long-term memory for AI Agents via a 4-tier progressive pipeline, with zero external API dependencies.
Your agent in your terminal, equipped with local tools: writes code, uses the terminal, browses the web. Make your own persistent autonomous agent on top!
Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.
Open-source security automation platform for teams and AI agents
Agent Skills as a Memory Layer
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
The platform for LLM evaluations and AI agent testing
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
Official Microsoft Learn MCP Server and CLI tool – powering LLMs and AI agents with real-time, trusted Microsoft docs & code samples.
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
E2B Desktop Sandbox for LLMs. E2B Sandbox with desktop graphical environment that you can connect to any LLM for secure computer use.
LLM powered fuzzing via OSS-Fuzz.
Open-source AI agent harness in native Rust — GUI, CLI, headless, and webapp from one binary. Multi-provider, MCP, skills, plugins, agent teams.
A desktop multi-agent harness built with Rust, Tauri, and React, powered by langgraph-rust.
Evaluate and improve models and agents using environments
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
Tensorlake is a serverless runtime for sandboxes and deploying background agentic applications
An agent benchmark with tasks in a simulated software company.
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
Agent harness and tookit for building AI agents and agentic applications. CLI and SDKs included
Open-source benchmark for browser AI agents on daily tasks.
The action firewall for AI agents. Enforce policy and human approval before risky tool calls, shell commands, workflows, and production changes, with auditable evidence.
Self-hosted Personal AI + agent runtime in .NET (NativeAOT-friendly)
Safe runtime for autonomous on-chain AI agents: isolated sandboxes, Library skills, encrypted secrets, and OKX read-only security checks.
A collection of Model Context Protocol (MCP) servers, clients and developer tools by IBM.
The production-ready agent harness framework for Python
Secure runtime to sandbox AI agent tasks. Run untrusted code in isolated WebAssembly environments.
Open Python agent harness for production AI apps: tools, MCP, memory, workspace, telemetry, subagents, background tasks, and OmniServe APIs.
Benchmark for comparing agent harnesses on everyday online tasks — fixes the base model, varies the harness. Sister project of ClawBench, same scoring pipeline.