Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Awesome List
An awesome list of Agent Harness engineering resources, including GitHub projects, tools, benchmarks, and practical guides.
GitHub stars and default-branch commits for Picrew/awesome-agent-harness.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
AI Agent Governance Toolkit — Policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering for autonomous AI agents. Covers 10/10 OWASP Agentic Top 10.
Agenta is a workspace where you and your team build agents and automations.
Enterprise AI Platform with guardrails, MCP registry, gateway & orchestrator
An AI Gateway, registry, and proxy that sits in front of any MCP, A2A, or REST/gRPC APIs, exposing a unified endpoint with centralized discovery, guardrails and management. Optimizes Agent & Tool calling, and supports plugins.
Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.
Framework for evaluating and improving agents
Open-source security automation platform for teams and AI agents
Agent Skills as a Memory Layer
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
Evaluation and Tracking for LLM Experiments and AI Agents
Devon: An open-source pair programmer
The platform for LLM evaluations and AI agent testing
agent-sandbox enables easy management of isolated, stateful, singleton workloads, ideal for use cases like AI agent runtimes.
Amazon Bedrock Agentcore accelerates AI agents into production with the scale, reliability, and security, critical to real-world deployment.
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
The NVIDIA NeMo Agent toolkit is an open-source library for efficiently connecting and optimizing teams of AI agents.
A benchmark for LLMs on complicated tasks in the terminal
A production-oriented multi-agent orchestration framework.
Terminal coding agent powered by Kimchi's multi-model orchestration
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
LLM powered fuzzing via OSS-Fuzz.
1 place to call all your agents - OpenCode, Hermes, Claude Managed Agents, Cursor Agents API, DeepAgents.
The Cloud Sandbox Built for AI Agents
OpenTelemetry Instrumentation for AI Observability
Evaluate and improve models and agents using environments
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers (aka The Norton for OpenClaw)
CUGA is an open-source generalist agent harness for the enterprise, supporting complex task execution on web and APIs, OpenAPI/MCP integrations, composable architecture, reasoning modes, and policy-aware features.
No description.
A production-ready runtime framework for agent apps with secure tool sandboxing, Agent-as-a-Service APIs, scalable deployment, full-stack observability, and broad framework compatibility.
Security Governance for Agentic AI
An agent benchmark with tasks in a simulated software company.
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
Sandboxed code execution for AI agents, locally or on the cloud. Massively parallel, easy to extend. Powering SWE-agent and more.
Open-source benchmark for browser AI agents on daily tasks.
The action firewall for AI agents. Enforce policy and human approval before risky tool calls, shell commands, workflows, and production changes, with auditable evidence.
Deterministic safety solutions for probabilistic AI agents
The production-ready agent harness framework for Python
Open Python agent harness for production AI apps: tools, MCP, memory, workspace, telemetry, subagents, background tasks, and OmniServe APIs.
Evaluation harness for OpenHands V1.
No description.