Papers
arxiv:2609.26793

HARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis

Published on Sep 22
ยท Submitted by
Chen Wang
on Sep 23
Authors:
,
,
,
,

Abstract

Compositional 3D scene reconstruction has recently been explored from two directions: agentic reasoning that provides semantic understanding of spatial relationships but lacks precise alignment with input images; and visual geometry foundation models that predict dense point maps from input images but the reconstruction quality is limited. Therefore, recovering a complete 3D scene from a single monocular image with accurate inter-object relationships and high-fidelity reconstruction quality remains challenging. In this paper, we present HARMONY, a hierarchical chain-of-thought framework that leverages both agentic reasoning and visual geometry foundation. Given an image of an indoor scene, starting from an empty 3D floorplan, HARMONY first calibrates the camera against the reference image to establish a semantically-grounded spatial frame, then uses agentic VLM reasoning to recover the 3D room layout and an initial placement order. It then places the objects in a hierarchical order, from wall-mounted elements, free-standing furniture, to dependent decorations on top of furniture. We also use depth-first traversal for furniture so each placement conditions on previously resolved structure and a reflective feedback loop to avoid error accumulation. After each object placement by VLM, we use the point cloud estimations to perform geometry-based refinement so that the rendered image aligns better with the input. HARMONY can produce 3D scenes that are semantically consistent and perceptually aligned with the reference image, extending single-image compositional reconstruction to complex indoor scene images. Experiments on synthetic and real-world images demonstrate that HARMONY outperforms the evaluated reconstruction baselines, while qualitative comparisons with GPT-6 Astra suggest more faithful object arrangements and better preservation of scene details.

Community

Paper submitter

Excited to share our new ๐—ผ๐—ฝ๐—ฒ๐—ป-๐˜€๐—ผ๐˜‚๐—ฟ๐—ฐ๐—ฒ work, ๐—›๐—”๐—ฅ๐— ๐—ข๐—ก๐—ฌ, a framework for reconstructing compositional 3D scenes from single indoor images.

Much like a 3D artist, our VLM agent builds the scene in a ๐—ต๐—ถ๐—ฒ๐—ฟ๐—ฎ๐—ฟ๐—ฐ๐—ต๐—ถ๐—ฐ๐—ฎ๐—น ๐—ผ๐—ฟ๐—ฑ๐—ฒ๐—ฟ: first the room structure and camera, then wall-mounted objects, free-standing furniture, and finally dependent decorations. After each stage, it compares its render against the input image, reflects on what went wrong, and iteratively corrects the scene to prevent errors from accumulating.

Even with older VLM backbones, our framework remains ๐—ฐ๐—ผ๐—บ๐—ฝ๐—ฒ๐˜๐—ถ๐˜๐—ถ๐˜ƒ๐—ฒ ๐˜„๐—ถ๐˜๐—ต ๐—š๐—ฃ๐—ง-๐Ÿฒ ๐—”๐˜€๐˜๐—ฟ๐—ฎ ๐—ถ๐—ป ๐˜ƒ๐—ถ๐˜€๐˜‚๐—ฎ๐—น ๐—พ๐˜‚๐—ฎ๐—น๐—ถ๐˜๐˜†, ๐˜„๐—ต๐—ถ๐—น๐—ฒ ๐—ฎ๐—ฐ๐—ต๐—ถ๐—ฒ๐˜ƒ๐—ถ๐—ป๐—ด ๐—ฏ๐—ฒ๐˜๐˜๐—ฒ๐—ฟ ๐—ด๐—ฒ๐—ผ๐—บ๐—ฒ๐˜๐—ฟ๐—ถ๐—ฐ ๐—ฎ๐—น๐—ถ๐—ด๐—ป๐—บ๐—ฒ๐—ป๐˜ for single-image 3D scene reconstruction.

We also release ๐—›๐—”๐—ฅ๐— ๐—ข๐—ก๐—ฌ๐Ÿฏ๐Ÿฌ๐Ÿฌ, a benchmark for single-image 3D scene reconstruction with 300 scenes across three difficulty tiers: easy, medium, and complicated.

๐Ÿ”— Project: http://cwchenwang.github.io/harmony
๐Ÿ“Š Leaderboard: http://cwchenwang.github.io/harmony/harmony300
๐Ÿค— Data: http://huggingface.co/datasets/ShufanSun/harmony
๐Ÿ“„ Paper: https://arxiv.org/abs/2609.26793

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.26793
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.26793 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.26793 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.26793 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.