About
Embodied AI, 3D Vision, and Multimodal Foundation Models.
My research interests lie in spatially grounded vision-language-action (VLA) systems for robotic manipulation, particularly depth-aware perception, 3D scene understanding, and representation learning for embodied intelligence.
I previously studied embodied manipulation and spatial reasoning in VLA systems in the research group led by Associate Professor Bo Zhao at Shanghai Jiao Tong University. I am currently a Research Intern at The University of Hong Kong, learning about data retargeting and cross-embodiment alignment of action representations in the research group led by Professor Taku Komura.
Research Interests
- Spatially grounded vision-language-action models for robotic manipulation
- Depth-aware perception and 3D scene understanding in embodied systems
- Cross-embodiment action representation learning and retargeting
Selected Highlights
- Co-first author of Evo-Depth, focusing on spatial perception for embodied intelligence
I built this site as a record of my research journey, and plan to share publications, projects, and future work in embodied intelligence.