ArticuTable: Generating Instance-Level Interactive Rigid-Articulated 3D Tabletop Scenes from a Single Image

1 Wuhan University2 EPFL3 University of North Texas4 Australian National University

* Equal contribution   ·   † Corresponding author

arXiv:2610.05249 · 2026

Teaser video. A reconstructed tabletop scene with executable part-level motion.
EXPLORE THE RECONSTRUCTION

One image. An interactive world.

Rotate the tabletop, inspect each object, and control its recovered joints.

ArticuTable scene rendered in Isaac Sim

Click an object or choose it above, then drag its joint controls. Play follows the Isaac Sim video's joint trajectory. The poster and videos show the original simulator rendering; this WebGL view uses different lighting.

Abstract

Embodied agents benefit from 3D environments that combine visual fidelity to real-world observations with physical interactivity. Existing single-image tabletop reconstruction methods recover plausible scene geometry but typically represent objects as monolithic rigid bodies, limiting interaction to whole-object rigid motion and precluding executable part-level articulation. Meanwhile, recovering a scene layout consistent with the input view remains challenging because a single observation may admit multiple plausible pose-scale configurations.

We present ArticuTable, a single-image 3D tabletop reconstruction framework that recovers both executable part-level articulation and an input-view-consistent scene layout. For object modeling, generation-robust articulation modeling (GRAM) combines joint fitting guided by a multimodal large language model with semantic state reasoning to recover reliable joint parameters and valid motion ranges from imperfect proxy meshes. For scene layout, progressive semantic-geometric scene registration (PSGSR) progressively narrows the pose-scale search space using complementary metric, planar, and input-view constraints. We also contribute ArticuTable-100, a curated collection of 100 simulation-ready tabletop scenes. Extensive evaluation, including a user study, demonstrates strong performance across visual fidelity, scene consistency, articulation quality, physical plausibility, and simulation readiness.

Scenes in Motion

Switch between six additional simulation-ready scenes without opening a wall of video players.

Method

A single-image pipeline for reconstructing articulated assets and registering them in a consistent 3D scene.

ArticuTable pipeline: asset reconstruction, GRAM articulation modeling, PSGSR scene registration, and simulation-ready output
Figure 2. Overview of the ArticuTable framework.
01 · ARTICULATION

Generation-Robust Articulation Modeling

GRAM turns imperfect proxy meshes into executable assets by recovering joint structure, parameters, and motion ranges.

02 · REGISTRATION

Progressive Scene Registration

PSGSR aligns object position, orientation, and scale through metric, planar, and input-view constraints.

Reconstruction Results

Comparisons on generated and real-world tabletop images.

One RGB tabletop image reconstructed as an articulated scene under multiple joint configurations
Figure 1. Single-image reconstruction of articulated tabletop scenes under different joint configurations.
Qualitative comparisons between input images, baseline 3D reconstruction methods, and ArticuTable
Figure 3. Qualitative comparison with existing single-image tabletop reconstruction methods.

Interactive and simulation-ready

In-scene editing and physics validation with object removal, insertion, rearrangement, dragging, and articulated opening
Figure 6. In-scene editing and physics validation, including articulated opening.

Quantitative Evaluation

Scene reconstruction, articulated assets, component ablations, and human preference. Values below reproduce Tables 1–5 of the paper.

Scene-level reconstruction

ArticuTable leads on input-view consistency, perceptual assessment, and physical validity.

Table 1 · Scene-level results
MethodLPIPS ↓DINOv2 ↑CLIP ↑VF ↑IA ↑PP ↑Avg. ↑OR ↓ColO ↓ColS ↓
ACDC0.44470.51350.80583.9271.9203.8803.2424.2981.5302%32.00%
Gen3DSR0.27160.57040.83862.2204.6804.0133.6383.34911.3322%66.67%
MIDI0.43020.64690.86682.7872.9332.3602.6934.14237.9430%99.33%
TabletopGen0.39350.77880.89085.3804.8605.8935.3782.1621.4488%16.00%
ArticuTable0.23980.89910.92655.8676.2536.6006.2401.0490.3819%10.67%

VF: visual fidelity; IA: input alignment; PP: physical plausibility; OR: overall rank. ColO and ColS report collision rates.

Articulated asset quality

Lightwheel benchmark results, grouped by image- or mesh-conditioned input.

Table 2 · Lightwheel articulation benchmark
MethodInputTraining-freePrec. ↑Rec. ↑Rest mIoU ↑Art. gIoU ↑OC ↓AE ↓LE ↓
SINGAPOImage–37.6025.400.182−0.2630.03030.30°0.077
PActImage–24.6016.800.092−0.4340.05642.90°0.188
PhysX-AnythingImage–20.3019.200.093−0.4590.06438.80°0.123
URDF-Anything+Image–70.7023.900.260−0.1380.03350.40°0.128
Articulate AnyMeshMesh✓89.2042.500.4520.1580.01021.70°0.043
URDF-Anything+Mesh–77.5027.400.267−0.1100.02649.90°0.111
ParticulateMesh–89.9051.500.5760.3050.00920.90°0.040
GRAMMesh✓91.1546.540.5640.2830.01417.26°0.034

GRAM ablation

Table 3 · Component ablation
ComponentMetricEnabledDisabledReduction
MLLM-guided joint fittingAE ↓17.17°40.06°57.1%
LE ↓0.04170.052720.9%
Semantic state reasoningArt. PC ↓0.25020.26053.9%
OC ↓0.01520.017613.6%

PSGSR ablation

Table 4 · Registration stages
VariantTopInputSASPSLPIPS ↓DINOv2 ↑CLIP ↑
Point-cloud only–––0.34390.81020.9077
w/o Top-view×✓–0.25640.87020.9210
Top-view only✓××0.32350.85860.9156
w/o SASPS✓✓×0.25430.87620.9205
Full model✓✓✓0.23980.89910.9265

User study

59 participants; each method received 590 valid rankings.

Table 5 · Human preference rankings
MetricACDCGen3DSRMIDITabletopGenArticuTable
Mean rank ↓4.083.753.752.241.17
Median rank ↓44421
First-place rate ↑0.17%1.02%1.02%12.20%85.59%
Top-2 rate ↑6.61%11.86%8.81%74.58%98.14%

Citation

@misc{lv2026articutablegeneratinginstancelevelinteractive,
  title={ArticuTable: Generating Instance-Level Interactive Rigid-Articulated 3D Tabletop Scenes from a Single Image},
  author={Kai Lv and Yibo Yin and Lijun Guo and Heng Fan and Kaihao Zhang and Xingping Dong},
  year={2026},
  eprint={2610.05249},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2610.05249}
}