Co-first author Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model
A lightweight VLA model that incorporates depth cues to improve spatial grounding and embodied manipulation.
Selected work on spatially grounded VLA models, depth estimation, and embodied perception.
A lightweight VLA model that incorporates depth cues to improve spatial grounding and embodied manipulation.
A lightweight VLA model that preserves semantic alignment while improving efficiency for embodied tasks.
A region-aware monocular depth estimation framework that focuses depth modeling on target regions while preserving coherent global geometry.
A VLA model that improves embodied manipulation through implicit spatial understanding.