VLA Models and Efficiency

Vision-language-action modeling with stronger spatial understanding and more efficient inference.

This line of work studies how to make vision-language-action models more spatially aware while keeping them practical to deploy.

Representative public papers:

The emphasis is on balancing semantic alignment, action capability, and computational efficiency.

References

2025

  1. evo0.png
    Evo-0: Vision-language-action model with implicit spatial understanding
    T Lin, Gen Li, Y. Zhong, and 5 more authors
    arXiv, 2025
  2. vla_pruner.png
    VLA-Pruner: Temporal-Aware Dual-Level Visual Token Pruning for Efficient Vision-Language-Action Inference
    Z. Liu, Y. Chen, H. Cai, T Lin, and 3 more authors
    arXiv, 2025