Spatial-Temporal World Understanding

Benchmarks and evaluation for precise world understanding in multimodal embodied settings.

This project direction focuses on evaluating whether modern multimodal models can reason about the world with enough spatial and temporal precision for embodied tasks.

Current public representative paper:

The broader goal is to expose the gap between generic multimodal understanding and the kind of grounded reasoning required by agents acting in real environments.

References

2025

  1. stibench.png
    Sti-bench: Are mllms ready for precise spatial-temporal world understanding?
    Y. Li, Y. Zhang, T Lin, and 4 more authors
    ICCV, 2025