Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Awesome List
An awesome list of Agent Harness engineering resources, including GitHub projects, tools, benchmarks, and practical guides.
GitHub stars and default-branch commits for Picrew/awesome-agent-harness.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Multi-Agent Harness for Production AI
AI Observability & Evaluation
Secure, Fast, and Extensible Sandbox runtime for AI agents.
🤗 ml-intern: an open-source ML engineer that reads papers, trains models, and ships ML models
Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!
An Open-Source Asynchronous Coding Agent
Multi-platform SDK for integrating GitHub Copilot Agent into apps and services
Cybersecurity AI (CAI), the framework for AI Security
Open source MCP Servers for AWS
Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.
PraisonAI 🦞 — Hire a 24/7 AI Workforce. Stop writing boilerplate and start shipping autonomous self-improving agents that research, plan, code, and execute tasks. Deployed in 5 lines of code with built-in memory, RAG, and support for 100+ LLMs.
Build effective agents using Model Context Protocol and simple workflow patterns
OpenShell is the safe, private runtime for autonomous AI agents.
Open-source observability for your GenAI or LLM application, based on OpenTelemetry
🧱 easy fast local-first microVM runtime and library
Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails so you stay focused on your agents core logic.
可私有部署的多租户知识智能体平台:统一 RAG、知识图谱、多智能体、MCP/Skills、沙盒与权限管理。Self-hosted knowledge agent platform for RAG, knowledge graphs and multi-agent workflows.
The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench verified!
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
Nexent is a zero-code platform for auto-generating production-grade AI agents using Harness Engineering principles — unified tools, skills, memory, and orchestration with built-in constraints, feedback loops, and control planes.
AI Agent Governance Toolkit — Policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering for autonomous AI agents. Covers 10/10 OWASP Agentic Top 10.
All-in-One Sandbox for AI Agents that combines Browser, Shell, File, MCP and VSCode Server in a single Docker container.
Agenta is a workspace where you and your team build agents and automations.
Our library for RL environments + evals
Enterprise AI Platform with guardrails, MCP registry, gateway & orchestrator
Next Generation Agentic Proxy for AI Agents and MCP servers
An AI Gateway, registry, and proxy that sits in front of any MCP, A2A, or REST/gRPC APIs, exposing a unified endpoint with centralized discovery, guardrails and management. Optimizes Agent & Tool calling, and supports plugins.
AI-Driven Life Cycle (AI-DLC) adaptive workflow steering rules for AI coding agents
Framework for evaluating and improving agents
Open-source security automation platform for teams and AI agents
Agent Skills as a Memory Layer
LangChain 🔌 MCP
Evaluation and Tracking for LLM Experiments and AI Agents
The platform for LLM evaluations and AI agent testing
agent-sandbox enables easy management of isolated, stateful, singleton workloads, ideal for use cases like AI agent runtimes.
Amazon Bedrock Agentcore accelerates AI agents into production with the scale, reliability, and security, critical to real-world deployment.
Security scanner for AI agents, MCP servers and agent skills.
The NVIDIA NeMo Agent toolkit is an open-source library for efficiently connecting and optimizing teams of AI agents.
A benchmark for LLMs on complicated tasks in the terminal
A production-oriented multi-agent orchestration framework.
Terminal coding agent powered by Kimchi's multi-model orchestration
How real engineers run Claude Code and Codex: spec-driven planning, enforced TDD, persistent memory, and quality enforcement on all levels. Make your agents production-ready.
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
Reference code for the Meta-Harness paper.
Evaluate and improve models and agents using environments
Minimal and readable coding agent harness implementation in Python to explain the core components of coding agents.
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
CUGA is an open-source generalist agent harness for the enterprise, supporting complex task execution on web and APIs, OpenAPI/MCP integrations, composable architecture, reasoning modes, and policy-aware features.
Your agent's favorite harness, built on Pydantic AI