Co-first author

Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model

arXiv

Summary

Evo-Depth explores how explicit depth information can improve the spatial understanding of lightweight vision-language-action models for embodied manipulation.

My Contributions

  • Led the core architecture design and implementation
  • Advanced the training pipeline and experiment process
  • Designed evaluation in both simulation and real-world settings
  • Worked across RLBench, LIBERO, and MetaWorld

Status

This work is currently available on arXiv. I am listed as a co-first author.