Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding
Summary
Evo-0 investigates implicit spatial understanding in vision-language-action models for embodied manipulation.
My Contributions
- Supported simulation data processing
- Adapted the RLBench platform for experiments
- Helped reproduce the model and test experimental settings
- Participated in result analysis
Status
This work is available on arXiv.