Focusable Monocular Depth Estimation
Summary
This paper introduces Focusable Monocular Depth Estimation (FDE), a task setting for region-aware depth prediction where the model must prioritize task-relevant target regions without losing global scene geometry.
Method
The proposed FocusDepth framework conditions monocular relative depth estimation on box or text prompts. Its Multi-Scale Spatial-Aligned Fusion module aligns prompt-aware features with depth features, enabling more precise foreground depth estimation and sharper target boundaries.
Why It Matters
The benchmark and method are especially relevant for embodied settings, where accurate depth around task-critical regions is often more important than uniform scene-level performance.
Status
Available on arXiv.