UKGovernmentBEIS/inspect_evals
Collection of evals for Inspect AI
LoongFlow is an expert-grade Agent framework for Loop Engineering. Through a Plan-Execute-Summary loop and structured experiential memory, it enables AI to continuously think, execute, reflect, and evolve across complex software engineering, mathematical, and machine learning tasks.
Appears on
Quick read
Latest capture 2026-09-03 03:05
35 paths
Agent instructions and tool configuration found in this repository.
Agent instructions
Agent workspace 12
Claude Code 22
11 more paths detected.
7 observed captures since 2026-05-23. Observed captures are shown by default.
Stars from first capture +51
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Collection of evals for Inspect AI
🚀 EvoAgentX: Building a Self-Evolving Ecosystem of AI Agents
🚀 EvoAgentX: Building a Self-Evolving Ecosystem of AI Agents
Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.
Awesome resources for in-context learning and prompt engineering: Mastery of the LLMs such as ChatGPT, GPT-3, and FlanT5, with up-to-date and cutting-edge updates.
MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering