Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Awesome List
An awesome list of Agent Harness engineering resources, including GitHub projects, tools, benchmarks, and practical guides.
GitHub stars and default-branch commits for Picrew/awesome-agent-harness.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
The LLM Evaluation Framework
A comprehensive collection of Agent Skills for context engineering, multi-agent architectures, and production agent systems. Use when building, optimizing, or debugging agent systems that require effective context management.
🐫 CAMEL: The first and the best multi-agent framework. Finding the Scaling Law of Agents. https://www.camel-ai.org
SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.
MCP Toolbox for Databases is an open source MCP server for databases.
Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hosted dashboard and advanced analytics with timeline and execution graph view
"DeepCode: Open Agentic Coding (Paper2Code & Text2Web & Text2Backend)"
"OpenHarness: Open Agent Harness with a Built-in Personal Agent--Ohmo!"
Supercharge Your LLM Application Evaluations 🚀
Security scanner for AI agent skills. Detect vulnerabilities, malicious patterns, security risks, prompt injection, data exfiltration, and supply-chain risks in Claude Code, Codex, and MCP skills before you install them.
Secure, Fast, and Extensible Sandbox runtime for AI agents.
Eigent: The Open Source Cowork Desktop - Local and Free Alternative to Claude Cowork and Codex
Personal memory across agents
Open-source, secure environment with real-world tools for enterprise-grade agents.
A framework for few-shot evaluation of language models.
One portable memory layer for every AI agent: local-first, Markdown-native, user-owned, and self-evolving across apps, tools, and workflows.
IronClaw is an Agent OS focused on privacy, security and extensibility
A framework for building, orchestrating and deploying AI agents and multi-agent workflows with support for Python and .NET.
TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.
Multi-Agent Harness for Production AI
AI Observability & Evaluation
Secure, Fast, and Extensible Sandbox runtime for AI agents.
🤗 ml-intern: an open-source ML engineer that reads papers, trains models, and ships ML models
Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!
The fullstack MCP framework to develop MCP Apps for ChatGPT / Claude & MCP Servers for AI Agents.
An Open-Source Asynchronous Coding Agent
Multi-platform SDK for integrating GitHub Copilot Agent into apps and services
Cybersecurity AI (CAI), the framework for AI Security
Open source MCP Servers for AWS
Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.
PraisonAI 🦞 — Hire a 24/7 AI Workforce. Stop writing boilerplate and start shipping autonomous self-improving agents that research, plan, code, and execute tasks. Deployed in 5 lines of code with built-in memory, RAG, and support for 100+ LLMs.
Build effective agents using Model Context Protocol and simple workflow patterns
OpenShell is the safe, private runtime for autonomous AI agents.
Open-source observability for your GenAI or LLM application, based on OpenTelemetry
Build an agent harness and control it end-to-end. Open-source SDK for production AI agents in Python & TypeScript - any model, any cloud.
🧱 easy fast local-first microVM runtime and library
Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails so you stay focused on your agents core logic.
可私有部署的多租户知识智能体平台:统一 RAG、知识图谱、多智能体、MCP/Skills、沙盒与权限管理。Self-hosted knowledge agent platform for RAG, knowledge graphs and multi-agent workflows.
The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench verified!
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
Nexent is a zero-code platform for auto-generating production-grade AI agents using Harness Engineering principles — unified tools, skills, memory, and orchestration with built-in constraints, feedback loops, and control planes.
AI Agent Governance Toolkit — Policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering for autonomous AI agents. Covers 10/10 OWASP Agentic Top 10.
SWE-bench: Can Language Models Resolve Real-world Github Issues?
Agenta is a workspace where you and your team build agents and automations.
Our library for RL environments + evals
Your agent in your terminal, equipped with local tools: writes code, uses the terminal, browses the web. Make your own persistent autonomous agent on top!
Enterprise AI Platform with guardrails, MCP registry, gateway & orchestrator
An AI Gateway, registry, and proxy that sits in front of any MCP, A2A, or REST/gRPC APIs, exposing a unified endpoint with centralized discovery, guardrails and management. Optimizes Agent & Tool calling, and supports plugins.
Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.