Spatial-Temporal World Understanding
Benchmarks and evaluation for precise world understanding in multimodal embodied settings.
This project direction focuses on evaluating whether modern multimodal models can reason about the world with enough spatial and temporal precision for embodied tasks.
Current public representative paper:
The broader goal is to expose the gap between generic multimodal understanding and the kind of grounded reasoning required by agents acting in real environments.