truera/trulens
Evaluation and Tracking for LLM Experiments and AI Agents
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
Appears on
Quick read
Latest capture 2026-09-11 03:03
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
5 observed captures since 2026-06-12. Observed captures are shown by default.
Stars from first capture +23
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Evaluation and Tracking for LLM Experiments and AI Agents
Open-source self-improving QA agent for software teams. A test harness with memory. Write tests in natural language for web and mobile. agent-qa learns from every run, adapts to UI changes, and catches regressions before you ship.
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
A generative AI-powered framework for testing virtual agents.
An agent benchmark with tasks in a simulated software company.
BenchClaw — Multi-dimensional AI agent evaluation with 17-judge AI Tribunal, 10 scoring dimensions, radar charts, and deception detection. Benchmark any LLM agent.