Focusable Monocular Depth Estimation

May 2026 arXiv

Summary

This paper introduces Focusable Monocular Depth Estimation (FDE), a task setting for region-aware depth prediction where the model must prioritize task-relevant target regions without losing global scene geometry.

Method

The proposed FocusDepth framework conditions monocular relative depth estimation on box or text prompts. Its Multi-Scale Spatial-Aligned Fusion module aligns prompt-aware features with depth features, enabling more precise foreground depth estimation and sharper target boundaries.

Why It Matters

The benchmark and method are especially relevant for embodied settings, where accurate depth around task-critical regions is often more important than uniform scene-level performance.

Status

Available on arXiv.