DeepYard

Evaluation

89 tools across 4 categories tagged with “evaluation

🔧

dev tools

(63)
R

RAG Failure Diagnostics Clinic

Diagnose and fix common RAG pipeline failure modes

OSSFree
Shubham Saboo
133.5Ktoday95
L

Langfuse

Open-source LLM engineering platform — traces, evals, prompt management — 23K+ stars

OSSfreemium
Langfuse
33.5Ktoday196
p

promptfoo

Test and evaluate LLM prompts and agents — 11K+ stars

OSSFree
promptfoo
24.4Ktoday318
A

Arize Phoenix

Open-source LLM observability with tracing, evaluation, and datasets — 8K+ stars

OSSfreemium
Arize AI
11.1K250.0K/wtoday207
B

Braintrust

AI evaluation and experiment tracking platform for production LLM apps

Commercialfreemium
Braintrust
1.2Ktoday93
L

LangSmith

Observability, tracing, and evaluation platform for LLM applications

Commercialfreemium
LangChain
1.0Ktoday103
A

AgentBeats

Agent-based evaluation framework where AI judges assess other agents through standardized protocols

OSSFree
Research
A

AgentCheck

Debug LLM agents with reproducible recordings and controlled interventions over MCP

OSSFree
Research
A

AgentDS

Benchmark framework measuring AI agent performance vs human experts on data science tasks

unknownFree
Research Team
A

Agora Consensus Agent

Multi-agent framework for autonomous bug detection in distributed consensus protocols

OSSFree
Research Team
A

Agora Consensus

Multi-agent framework for autonomous bug detection in distributed consensus protocols

OSSFree
Research Team
A

Alipay-PIBench

Benchmark for testing coding agents on real-world Alipay payment integration scenarios

OSSFree
Alipay/Research
A

ATLAS

LLM-powered automata learning for analyzing and explaining agent strategies

OSSFree
Research Collaboration
B

BackendForge

Benchmark framework for evaluating AI agents on end-to-end backend code generation

OSSFree
Research
C

CapCode

Cheat-proof coding evaluation framework with performance-capped randomized tests

OSSFree
Research Team
C

CausalFlow

Causal debugging framework that converts failed LLM agent traces into minimal counterfactual fixes

OSSFree
Research
C

Change2Task

Transform merged PRs into verified coding agent tasks for training data generation

OSSFree
Research
C

ClawTrap

MITM red-teaming framework for testing autonomous web agent security in real-world scenarios

OSSFree
Independent Researchers
C

CUA-Gym

Reinforcement learning benchmark suite for training GUI automation agents with verifiable rewards

OSSFree
Research
C

CUA-Suite

Large-scale human-annotated video dataset for training computer-use agents

OSSFree
Research Team
D

DeltaBox

Millisecond-level checkpoint/rollback system for stateful AI agent exploration

OSSFree
Research
D

DiG-bench

70 game environments for benchmarking AI agent discovery and experimentation capabilities

OSSFree
Research Team
E

EasyOPD

On-policy distillation framework for compressing LLMs with unified teacher-student training

OSSFree
Research
E

Echo

Learning system that refines AI agents through user feedback on interaction logs

OSSFree
Research
E

EpiBench

Verifiable benchmark for testing AI agents on epigenomics analysis workflows

OSSFree
Research
E

EVMbench

Benchmark for evaluating AI agents on smart contract security tasks

OSSFree
Research Team (Wang et al.)
E

EvoRepair

Self-evolving AI agent that learns from past fixes to automatically repair code vulnerabilities

OSSFree
Research Team
F

FinToolBench

Specialized benchmark for evaluating LLM agents on real-world financial tool use and compliance

OSSFree
Research Collaboration
F

FluxBench

Benchmark framework for evaluating AI agents on complete chip design workflows

OSSFree
Research Team (arXiv)
F

Front-End Checklist

Comprehensive web development checklist for HTML, CSS, and JavaScript best practices

OSSFree
David Dias
H

Hallucination-Aware Layered Oversight

Zero-hallucination AI framework through architectural separation of generation and validation

OSSFree
Research Team (arXiv)
H

HARP

Research platform for systematic HCI studies of human-AI interaction with live LLM agents

OSSFree
Research Community
I

ICAE-Bench

Benchmark for evaluating coding agents on end-to-end project building from vague intent

OSSFree
Research Community
I

IMBench

Benchmark for evaluating AI agents on intuitive robotic manipulation tasks

OSSFree
Research Team
I

IssueTrojanBench

Security benchmark testing AI coding agents against malicious issue requests and backdoors

OSSFree
Research Community
J

JetBrains AI Toolkit

In-IDE LLM tracing and evaluation for JetBrains developers

Commercialpaid
JetBrains
L

LEDGER

Claim-to-evidence trace graphs for auditing and verifying LLM agent reasoning

OSSFree
Research (Daehong Kim et al.)
L

LH-Bench

Evaluation framework for measuring long-horizon agent workflows on enterprise tasks

OSSFree
Research Project
M

MalSkillBench

Security benchmark for detecting malicious AI agent skills with runtime verification

OSSFree
Research Team
M

MCP-in-SoS

Security risk assessment framework for evaluating open-source MCP server implementations

OSSFree
Research
M

MCPEvol-Bench

Benchmark for testing LLM agent adaptability to evolving MCP server interfaces

OSSFree
Research
O

OmniaBench

Comprehensive benchmark suite for evaluating general-purpose AI agents across diverse scenarios

OSSFree
Research
P

PhysAssistBench

Benchmark for evaluating LLM agents in doctor-patient-EHR clinical assistance workflows

OSSFree
Research Team (Du et al.)
P

PostTrainBench

Benchmark for evaluating autonomous post-training of LLMs under compute constraints

OSSFree
Research Collaboration
P

ProofAgent Index

Governance readiness index for assessing AI agent production safety and compliance

OSSFree
Research
Q

QuoteBench

Benchmark isolating LLM code generation failures from command execution errors

OSSFree
Research
R

Replica

Auto-evaluating framework for training AI agents to replicate scientific research papers

OSSFree
Research Team
S

ScrambleToolBench

Benchmark testing agent reasoning in undocumented terminal environments

OSSFree
Research Team
S

SlopCodeBench

Benchmark for measuring coding agent performance degradation across iterative tasks

OSSFree
Research Collaboration
S

SpecOps

Automated testing framework for GUI-based AI agents in real-world environments

OSSFree
Research
S

SWE-Explore

Benchmark for evaluating coding agents' repository exploration and code understanding abilities

OSSFree
Research Team
S

SWE-Touch

Benchmark for testing coding agents' handling of concurrent user edits in shared workspaces

OSSFree
Research Team
T

TDAD

Pre-change impact analysis for AI coding agents using graph-based dependency mapping

OSSFree
Research Project
T

TML-Bench

Benchmark for evaluating data science agents on Kaggle-style tabular ML tasks

OSSFree
Independent
T

ToolBench-X

Benchmark framework testing agent robustness with unreliable tools and error conditions

OSSFree
Research (Yang Tian et al.)
T

Translate-R1

RL-powered cost-aware translation optimizer for multilingual LLM agents

OSSFree
Research Team
T

TRIM

Clean up bloated AI-generated code by removing speculative edits and abandoned search paths

OSSFree
Research Team (arXiv)
V

Vera

Automated safety testing framework for LLM agents with risk discovery and verification

OSSFree
Multiple Authors
V

Veritas

LLM-powered binary analysis for detecting memory corruption vulnerabilities in stripped code

OSSFree
Columbia University / King's College London
V

VRR-Stop

Principled stopping framework for LLM agent repair loops that prevents over-correction

OSSFree
Research Team (arXiv)
V

VulnGym

Repository-level vulnerability detection benchmark for autonomous coding agents

OSSFree
Research Team
W

WorkBuddy Bench

Multi-domain benchmark for evaluating coding agents across real-world software engineering tasks

OSSFree
Tencent
Z

ZEBRAARENA

Diagnostic simulation environment for testing LLM reasoning and tool use capabilities

OSSFree
Research Project
🧱

frameworks

(14)
A

Ares

Adaptive reasoning framework that optimizes LLM inference costs by dynamically scaling effort

OSSFree
Research Collaboration
A

Automat

Autonomous feature engineering for materials science via LLM-driven genetic programming

OSSFree
Trinity College Dublin
C

CAST

Case-driven framework that learns from execution trajectories to optimize LLM tool use

OSSFree
Research collaboration
E

EvoTool

Evolutionary framework that self-optimizes tool-use policies for long-horizon LLM agent tasks

OSSFree
Research Team (Yang et al.)
G

Glite ARF

Verifier-driven framework for parallel LLM coding agents with built-in audit capabilities

OSSFree
Glite
H

HarnessX

Composable runtime scaffolding for adaptive AI agents that learn from execution traces

OSSFree
Research Team (Tingyang Chen et al.)
M

Motif-3

MoE-based long-context agent framework with formal verification and multilingual support

OSSFree
Motif Technologies
O

OpenForgeRL

End-to-end RL framework for training stateful coding agents across diverse environments

OSSFree
Research Community
R

ReflectFact

Self-reflective multi-agent system for verifying complex factual claims across multiple sources

OSSFree
Research
R

ResidencyRL

RL framework for training medical AI agents through simulated clinical encounters

OSSFree
Research Collaboration
S

SPADE

Self-play RL framework where LLMs design environments and solve them for continuous improvement

OSSFree
Research (Bo Liu et al.)
T

TITAN Agent

Full-featured TypeScript agent framework with multi-agent orchestration and Mission Control UI

OSSFree
Djtony707
V

Vellum

Visual AI workflow builder with SDK and enterprise governance

Commercialfreemium
Vellum AI
V

VeriSkill

Self-evolving framework enabling LLM agents to learn and refine program verification skills

OSSFree
Research
🤖

agents

(11)
💬

prompts

(1)