Q
QuoteBench
Benchmark isolating LLM code generation failures from command execution errors
Open SourceFree
About
Research benchmark designed to measure where coding agents fail: at command generation or execution transport. Tests 56 one-shot tasks derived from 14 real incident families, using exact final-state validation to distinguish between LLM generation errors and system-level parsing/execution failures. Provides critical insights for improving agent reliability by identifying whether issues stem from model capabilities or infrastructure.
Details
| Type | |
| Integrations | |
| Language |
Tags
evaluationcoding-agentopen-sourceframework
Quick Info
- Organization
- Research
- Pricing
- open-source
- Free Tier
- Yes
- Updated
- Aug 16, 2026
Also in Dev Tools
C
Crawl4AI
Open-source web crawler optimized for LLMs and AI agents — 62K+ stars
OSSFree
unclecode
79.0K2d ago82
F
Firecrawl
Web scraping API built for LLMs — turn any website into LLM-ready data — 89K+ stars
OSSfreemium
Mendable
170.7Ktoday162
H
Headroom Context Optimization
Reduce LLM API costs by 50-90% through advanced context compression
OSSFree
Shubham Saboo
133.5Ktoday95