LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models

Published in CVPR, 2025

Recommended citation: Sun, F.-Y., Liu, W., Gu, S., Lim, D., Bhat, G., Tombari, F., Li, M., Haber, N., & Wu, J. (2025). "LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models." CVPR. https://arxiv.org/abs/2412.02193

LayoutVLM introduces a scene layout representation that lets vision-language models reason about numerical poses and spatial relations. It uses self-consistent decoding and differentiable optimization to improve instruction adherence and physical plausibility in generated 3D scenes.

Read the paper ยท View the code