Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding

arXiv

Summary

Evo-0 investigates implicit spatial understanding in vision-language-action models for embodied manipulation.

My Contributions

  • Supported simulation data processing
  • Adapted the RLBench platform for experiments
  • Helped reproduce the model and test experimental settings
  • Participated in result analysis

Status

This work is available on arXiv.