confident-ai/deepeval
The LLM Evaluation Framework
Dingo: A Comprehensive AI Data, Model and Application Quality Evaluation Tool
Appears on
Quick read
Latest capture 2026-08-02 03:11
14 paths
Agent instructions and tool configuration found in this repository.
Agent instructions
Claude Code 8
Cursor 5
9 observed captures since 2026-05-22. Observed captures are shown by default.
Stars from first capture +29
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
The LLM Evaluation Framework
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
~95% on SimpleQA (e.g. Qwen3.6-27B on a 3090). Supports all local and cloud LLMs (llama.cpp, Ollama, Google, ...). 10+ search engines - arXiv, PubMed, your private documents. Everything Local & Encrypted.
Rigourous evaluation of LLM-synthesized code - NeurIPS 2023 & COLM 2024
Open source platform for AI Engineering: OpenTelemetry-native LLM Observability, GPU Monitoring, Guardrails, Evaluations, Prompt Management, Vault, Playground. 🚀💻 Integrates with 50+ LLM Providers, VectorDBs, Agent Frameworks and GPUs.