Evaluation
89 tools across 4 categories tagged with “evaluation”
dev tools
(63)RAG Failure Diagnostics Clinic
Diagnose and fix common RAG pipeline failure modes
Langfuse
Open-source LLM engineering platform — traces, evals, prompt management — 23K+ stars
promptfoo
Test and evaluate LLM prompts and agents — 11K+ stars
Arize Phoenix
Open-source LLM observability with tracing, evaluation, and datasets — 8K+ stars
Braintrust
AI evaluation and experiment tracking platform for production LLM apps
LangSmith
Observability, tracing, and evaluation platform for LLM applications
AgentBeats
Agent-based evaluation framework where AI judges assess other agents through standardized protocols
AgentCheck
Debug LLM agents with reproducible recordings and controlled interventions over MCP
AgentDS
Benchmark framework measuring AI agent performance vs human experts on data science tasks
Agora Consensus Agent
Multi-agent framework for autonomous bug detection in distributed consensus protocols
Agora Consensus
Multi-agent framework for autonomous bug detection in distributed consensus protocols
Alipay-PIBench
Benchmark for testing coding agents on real-world Alipay payment integration scenarios
ATLAS
LLM-powered automata learning for analyzing and explaining agent strategies
BackendForge
Benchmark framework for evaluating AI agents on end-to-end backend code generation
CapCode
Cheat-proof coding evaluation framework with performance-capped randomized tests
CausalFlow
Causal debugging framework that converts failed LLM agent traces into minimal counterfactual fixes
Change2Task
Transform merged PRs into verified coding agent tasks for training data generation
ClawTrap
MITM red-teaming framework for testing autonomous web agent security in real-world scenarios
CUA-Gym
Reinforcement learning benchmark suite for training GUI automation agents with verifiable rewards
CUA-Suite
Large-scale human-annotated video dataset for training computer-use agents
DeltaBox
Millisecond-level checkpoint/rollback system for stateful AI agent exploration
DiG-bench
70 game environments for benchmarking AI agent discovery and experimentation capabilities
EasyOPD
On-policy distillation framework for compressing LLMs with unified teacher-student training
Echo
Learning system that refines AI agents through user feedback on interaction logs
EpiBench
Verifiable benchmark for testing AI agents on epigenomics analysis workflows
EVMbench
Benchmark for evaluating AI agents on smart contract security tasks
EvoRepair
Self-evolving AI agent that learns from past fixes to automatically repair code vulnerabilities
FinToolBench
Specialized benchmark for evaluating LLM agents on real-world financial tool use and compliance
FluxBench
Benchmark framework for evaluating AI agents on complete chip design workflows
Front-End Checklist
Comprehensive web development checklist for HTML, CSS, and JavaScript best practices
Hallucination-Aware Layered Oversight
Zero-hallucination AI framework through architectural separation of generation and validation
HARP
Research platform for systematic HCI studies of human-AI interaction with live LLM agents
ICAE-Bench
Benchmark for evaluating coding agents on end-to-end project building from vague intent
IMBench
Benchmark for evaluating AI agents on intuitive robotic manipulation tasks
IssueTrojanBench
Security benchmark testing AI coding agents against malicious issue requests and backdoors
JetBrains AI Toolkit
In-IDE LLM tracing and evaluation for JetBrains developers
LEDGER
Claim-to-evidence trace graphs for auditing and verifying LLM agent reasoning
LH-Bench
Evaluation framework for measuring long-horizon agent workflows on enterprise tasks
MalSkillBench
Security benchmark for detecting malicious AI agent skills with runtime verification
MCP-in-SoS
Security risk assessment framework for evaluating open-source MCP server implementations
MCPEvol-Bench
Benchmark for testing LLM agent adaptability to evolving MCP server interfaces
OmniaBench
Comprehensive benchmark suite for evaluating general-purpose AI agents across diverse scenarios
PhysAssistBench
Benchmark for evaluating LLM agents in doctor-patient-EHR clinical assistance workflows
PostTrainBench
Benchmark for evaluating autonomous post-training of LLMs under compute constraints
ProofAgent Index
Governance readiness index for assessing AI agent production safety and compliance
QuoteBench
Benchmark isolating LLM code generation failures from command execution errors
Replica
Auto-evaluating framework for training AI agents to replicate scientific research papers
ScrambleToolBench
Benchmark testing agent reasoning in undocumented terminal environments
SlopCodeBench
Benchmark for measuring coding agent performance degradation across iterative tasks
SpecOps
Automated testing framework for GUI-based AI agents in real-world environments
SWE-Explore
Benchmark for evaluating coding agents' repository exploration and code understanding abilities
SWE-Touch
Benchmark for testing coding agents' handling of concurrent user edits in shared workspaces
TDAD
Pre-change impact analysis for AI coding agents using graph-based dependency mapping
TML-Bench
Benchmark for evaluating data science agents on Kaggle-style tabular ML tasks
ToolBench-X
Benchmark framework testing agent robustness with unreliable tools and error conditions
Translate-R1
RL-powered cost-aware translation optimizer for multilingual LLM agents
TRIM
Clean up bloated AI-generated code by removing speculative edits and abandoned search paths
Vera
Automated safety testing framework for LLM agents with risk discovery and verification
Veritas
LLM-powered binary analysis for detecting memory corruption vulnerabilities in stripped code
VRR-Stop
Principled stopping framework for LLM agent repair loops that prevents over-correction
VulnGym
Repository-level vulnerability detection benchmark for autonomous coding agents
WorkBuddy Bench
Multi-domain benchmark for evaluating coding agents across real-world software engineering tasks
ZEBRAARENA
Diagnostic simulation environment for testing LLM reasoning and tool use capabilities
frameworks
(14)Ares
Adaptive reasoning framework that optimizes LLM inference costs by dynamically scaling effort
Automat
Autonomous feature engineering for materials science via LLM-driven genetic programming
CAST
Case-driven framework that learns from execution trajectories to optimize LLM tool use
EvoTool
Evolutionary framework that self-optimizes tool-use policies for long-horizon LLM agent tasks
Glite ARF
Verifier-driven framework for parallel LLM coding agents with built-in audit capabilities
HarnessX
Composable runtime scaffolding for adaptive AI agents that learn from execution traces
Motif-3
MoE-based long-context agent framework with formal verification and multilingual support
OpenForgeRL
End-to-end RL framework for training stateful coding agents across diverse environments
ReflectFact
Self-reflective multi-agent system for verifying complex factual claims across multiple sources
ResidencyRL
RL framework for training medical AI agents through simulated clinical encounters
SPADE
Self-play RL framework where LLMs design environments and solve them for continuous improvement
TITAN Agent
Full-featured TypeScript agent framework with multi-agent orchestration and Mission Control UI
Vellum
Visual AI workflow builder with SDK and enterprise governance
VeriSkill
Self-evolving framework enabling LLM agents to learn and refine program verification skills
agents
(11)Agents-A1
Open-source multimodal agent model with image-text reasoning on Qwen 3.5 MoE architecture
AREX
Recursively self-improving research agent with constraint-wise verification
Code-Augur
Autonomous LLM agent for security vulnerability detection through specification inference
ImageEdit-R1
RL-powered multi-agent system for complex, instruction-based image editing
JetBrains Mellum2
12B parameter coding model with dedicated thinking mode for IDE integration
KARL
RL-trained enterprise search agents with multi-regime evaluation benchmark
MLEvolve
Self-evolving multi-agent system for automated ML algorithm discovery and optimization
QBugLM
Multi-agent framework for automated quantum software debugging with silent error detection
RepoLaunch
Automated repository testing agent for cross-language SWE dataset construction
ReViSQL
Text-to-SQL agent achieving human-level accuracy through step-by-step reasoning
Socratic-SWE
Self-improving coding agent that learns from its execution traces to evolve continuously