DeepYardDeepYard
I

IMBench

Benchmark for evaluating AI agents on intuitive robotic manipulation tasks

Open SourceFree

About

IMBench is an open-source evaluation framework that tests AI agents' ability to combine reasoning with motor control in robotic manipulation scenarios. It measures how effectively agents convert high-level reasoning into physical actions and adapt to novel manipulation tasks. Addresses a critical gap in agent evaluation by focusing on the reasoning-to-action pipeline rather than just language or vision capabilities alone.

Details

Type
Integrations
Language

Tags

evaluationautonomousopen-sourceframework