WorkBuddy Bench
Multi-domain benchmark for evaluating coding agents across real-world software engineering tasks
About
WorkBuddy Bench is Tencent's comprehensive evaluation suite designed to test coding agents across Code, Web, Office, and Security domains. It uses contamination-resistant tasks by reverse-engineering actual GitHub commits and pull requests, ensuring agents are tested on realistic software engineering challenges. The framework provides unified evaluation metrics and distribution-informed task construction, making it ideal for researchers and teams developing autonomous coding agents who need standardized, real-world benchmarks.
Details
| Type | |
| Integrations | |
| Language |
Tags
Quick Info
- Organization
- Tencent
- Pricing
- open-source
- Free Tier
- Yes
- Updated
- Jul 24, 2026
Also in Dev Tools
Crawl4AI
Open-source web crawler optimized for LLMs and AI agents — 62K+ stars
Firecrawl
Web scraping API built for LLMs — turn any website into LLM-ready data — 89K+ stars
Headroom Context Optimization
Reduce LLM API costs by 50-90% through advanced context compression