DeepYardDeepYard
W

WorkBuddy Bench

Multi-domain benchmark for evaluating coding agents across real-world software engineering tasks

Open SourceFree

About

WorkBuddy Bench is Tencent's comprehensive evaluation suite designed to test coding agents across Code, Web, Office, and Security domains. It uses contamination-resistant tasks by reverse-engineering actual GitHub commits and pull requests, ensuring agents are tested on realistic software engineering challenges. The framework provides unified evaluation metrics and distribution-informed task construction, making it ideal for researchers and teams developing autonomous coding agents who need standardized, real-world benchmarks.

Details

Type
Integrations
Language

Tags

evaluationcoding-agentopen-sourceframeworkmulti-agent