DeepYardDeepYard
S

SWE-Touch

Benchmark for testing coding agents' handling of concurrent user edits in shared workspaces

Open SourceFree

About

Research benchmark framework that evaluates how coding agents respond when users modify code during active tasks. Introduces the Counter-Edits methodology to stress-test agent robustness in collaborative editing scenarios. Essential for developers building AI coding assistants that need to handle real-time workspace changes and maintain context during concurrent modifications.

Details

Type
Integrations
Language

Tags

coding-agentevaluationopen-sourceautonomous