DeepYardDeepYard
I

ICAE-Bench

Benchmark for evaluating coding agents on end-to-end project building from vague intent

Open SourceFree

About

ICAE-Bench is a research benchmark designed to evaluate coding agents as interactive project builders rather than single-task completers. It tests agents on transforming incomplete product specifications into working software through multi-step workflows including requirement clarification, planning, tool use, debugging, and repository-level code construction. Addresses the emerging paradigm of 'vibe-coding' where users provide high-level intent instead of precise specifications, making it essential for evaluating next-generation AI coding assistants.

Details

Type
Integrations
Language

Tags

evaluationcoding-agentautonomousframeworkopen-sourcetool-use