DeepYard
D

DiG-bench

70 game environments for benchmarking AI agent discovery and experimentation capabilities

Open SourceFree

About

DiG-bench is a research benchmark consisting of 70 independent game environments designed to evaluate AI agents' ability to discover objectives and learn through experimentation. Unlike traditional benchmarks with known goals, DiG-bench tests agents in controlled environments where objectives are initially unknown, requiring autonomous exploration and hypothesis testing. Ideal for researchers evaluating discovery algorithms, reinforcement learning approaches, and autonomous agent behavior in open-ended scenarios.

Details

Type
Integrations
Language

Tags

autonomousevaluationopen-sourceframework