Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Awesome List
An awesome list of Agent Harness engineering resources, including GitHub projects, tools, benchmarks, and practical guides.
GitHub stars and default-branch commits for Picrew/awesome-agent-harness.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
"DeepCode: Open Agentic Coding (Paper2Code & Text2Web & Text2Backend)"
"OpenHarness: Open Agent Harness with a Built-in Personal Agent--Ohmo!"
Supercharge Your LLM Application Evaluations 🚀
Security scanner for AI agent skills. Detect vulnerabilities, malicious patterns, security risks, prompt injection, data exfiltration, and supply-chain risks in Claude Code, Codex, and MCP skills before you install them.
Personal memory across agents
Open-source, secure environment with real-world tools for enterprise-grade agents.
A framework for few-shot evaluation of language models.
One portable memory layer for every AI agent: local-first, Markdown-native, user-owned, and self-evolving across apps, tools, and workflows.
A framework for building, orchestrating and deploying AI agents and multi-agent workflows with support for Python and .NET.
Agent S: an open agentic framework that uses computers like a human
Multi-Agent Harness for Production AI
Give your agents the power of the Hugging Face ecosystem
AI Observability & Evaluation
Secure, Fast, and Extensible Sandbox runtime for AI agents.
🤗 ml-intern: an open-source ML engineer that reads papers, trains models, and ships ML models
Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!
An Open-Source Asynchronous Coding Agent
Cybersecurity AI (CAI), the framework for AI Security
Open source MCP Servers for AWS
"Context engineering is the delicate art and science of filling the context window with just the right information for the next step." — Andrej Karpathy. A frontier, first-principles handbook inspired by Karpathy and 3Blue1Brown for moving beyond prompt engineering to the wider discipline of context design, orchestration, and optimization.
Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.
PraisonAI 🦞 — Hire a 24/7 AI Workforce. Stop writing boilerplate and start shipping autonomous self-improving agents that research, plan, code, and execute tasks. Deployed in 5 lines of code with built-in memory, RAG, and support for 100+ LLMs.
Build effective agents using Model Context Protocol and simple workflow patterns
Open-source observability for your GenAI or LLM application, based on OpenTelemetry
Build an agent harness and control it end-to-end. Open-source SDK for production AI agents in Python & TypeScript - any model, any cloud.
可私有部署的多租户知识智能体平台:统一 RAG、知识图谱、多智能体、MCP/Skills、沙盒与权限管理。Self-hosted knowledge agent platform for RAG, knowledge graphs and multi-agent workflows.
Demo of a customer service use case implemented with the OpenAI Agents SDK
AIOS: AI Agent Operating System
The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench verified!
A simple SWE style browser agent framework that achieves SOTA results on long horizon web tasks.
A model-driven approach to building AI agents in just a few lines of code.
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
Nexent is a zero-code platform for auto-generating production-grade AI agents using Harness Engineering principles — unified tools, skills, memory, and orchestration with built-in constraints, feedback loops, and control planes.
AI Agent Governance Toolkit — Policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering for autonomous AI agents. Covers 10/10 OWASP Agentic Top 10.
SWE-bench: Can Language Models Resolve Real-world Github Issues?
All-in-One Sandbox for AI Agents that combines Browser, Shell, File, MCP and VSCode Server in a single Docker container.
The CLI and skills that turn any coding assistant into an expert at creating, evaluating, and deploying AI agents on Google Cloud.
Our library for RL environments + evals
Your agent in your terminal, equipped with local tools: writes code, uses the terminal, browses the web. Make your own persistent autonomous agent on top!
An AI Gateway, registry, and proxy that sits in front of any MCP, A2A, or REST/gRPC APIs, exposing a unified endpoint with centralized discovery, guardrails and management. Optimizes Agent & Tool calling, and supports plugins.
Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.
AI-Driven Life Cycle (AI-DLC) adaptive workflow steering rules for AI coding agents
Framework for evaluating and improving agents
Open-source security automation platform for teams and AI agents
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
LangChain 🔌 MCP
Evaluation and Tracking for LLM Experiments and AI Agents
Devon: An open-source pair programmer
Amazon Bedrock Agentcore accelerates AI agents into production with the scale, reliability, and security, critical to real-world deployment.
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.