DeepYard
Q

QuoteBench

Benchmark isolating LLM code generation failures from command execution errors

Open SourceFree

About

Research benchmark designed to measure where coding agents fail: at command generation or execution transport. Tests 56 one-shot tasks derived from 14 real incident families, using exact final-state validation to distinguish between LLM generation errors and system-level parsing/execution failures. Provides critical insights for improving agent reliability by identifying whether issues stem from model capabilities or infrastructure.

Details

Type
Integrations
Language

Tags

evaluationcoding-agentopen-sourceframework