Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model
Summary
Evo-Depth explores how explicit depth information can improve the spatial understanding of lightweight vision-language-action models for embodied manipulation.
My Contributions
- Led the core architecture design and implementation
- Advanced the training pipeline and experiment process
- Designed evaluation in both simulation and real-world settings
- Worked across RLBench, LIBERO, and MetaWorld
Status
This work is currently available on arXiv. I am listed as a co-first author.