hidai25/eval-view
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
Repository profile
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
Repository updates
Get generated JudgmentLabs/judgeval development summaries by email, or follow the weekly and monthly RSS feeds.
Sign in to subscribe by email. RSS feeds are public.
Sign in to subscribeTracked growth, recent movement, and commit velocity from stored repository snapshots.
Latest capture 2026-07-31 03:09
4 captures since 2026-06-12
Stars from baseline +12
All tracked data
Frameworks, package managers, ecosystems, and dependency manifests found during catalog scans.
Scanned 2026-07-31 03:09
pyproject.toml
python ecosystem,
31 dependencies
uv.lock
python ecosystem,
0 dependencies
examples/basic-distributed-tracing/pyproject.toml
python ecosystem,
4 dependencies
examples/basic-evaluation/pyproject.toml
python ecosystem,
1 dependency
examples/basic-linked-trace/pyproject.toml
python ecosystem,
1 dependency
examples/basic-tracing/pyproject.toml
python ecosystem,
2 dependencies
examples/claude-agent-sdk/pyproject.toml
python ecosystem,
2 dependencies
examples/google-adk/pyproject.toml
python ecosystem,
2 dependencies
Searchable topics, generated tags, and stack labels that explain where this repository fits.
Agent instructions and tool configuration paths found in the repository tree.
Nearest indexed repositories by embedding similarity.
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.
The self-improving QA agent for software teams. A test harness with memory. Write tests in natural language for web and mobile. agent-qa learns from every run, adapts to UI changes, and catches regressions before you ship.
A generative AI-powered framework for testing virtual agents.
Evaluation and Tracking for LLM Experiments and AI Agents
An agent benchmark with tasks in a simulated software company.
BenchClaw — Multi-dimensional AI agent evaluation with 17-judge AI Tribunal, 10 scoring dimensions, radar charts, and deception detection. Benchmark any LLM agent.