SciGroveBeta
Robotics

TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable Manipulation

Boyuan Wang, Yue Zhang, Xutao Xue, Xueyu Song, Yu Sun

Featured July 26, 2026

This analysis was generated by SciGrove. Upload your own PDFs or enter a DOI — and get the same AI breakdown on any paper.

Get started

AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.

Simply

This paper creates a huge dataset of realistic virtual tabletops by turning everyday photos into physics-ready 3D scenes, using a clever collision-fixing step to make sure objects don't overlap unnaturally, which helps robots learn to manipulate things better.

In depth
The paper introduces TableVerse, a Real2Sim pipeline that converts unstructured internet images into high-fidelity, physics-ready tabletop simulation environments. This framework deterministically reconstructs scenes with accurate metric scales and authentic object topologies, crucially addressing the physical implausibility and sparse clutter issues of prior generative methods. A key innovation is the Layout-Consistent Collision Rectification (LCCR) module, which geometrically disentangles intersecting meshes while preserving macroscopic layouts, followed by physics stabilization to create robust, interactive digital twins for robotic manipulation.

Key Takeaways

  • 1
    TableVerse is a novel Real2Sim pipeline that transforms unstructured internet images into high-fidelity, physics-ready tabletop simulation environments, overcoming limitations of synthetic scene generation.
  • 2
    The Layout-Consistent Collision Rectification (LCCR) module is introduced to geometrically disentangle intersecting meshes while preserving the original spatial layout, a critical step before physics simulation.
  • 3
    The TableVerse-100K Dataset is constructed, comprising 100,000 unique, physically consistent tabletop scenes with expert manipulation trajectories, providing an unprecedented scale for generalizable robotic policy learning.

Conceptual Flow

HIGH LEVEL
1
Methodology: From Real Photos to Virtual Scenes

The system takes a real photo, figures out what objects are there, builds them in 3D, fixes any overlaps, and then lets them settle naturally, creating a perfect virtual copy.

Real Photo
Convert & Refine
Virtual Scene
2
Results: Zero Collisions for Better Robot Training

The new method creates virtual scenes with zero object collisions, unlike older methods that often had many errors, making the scenes much more useful for training robots.

Old Methods: Many Collisions
New Method: No Collisions
Compare Scene Quality
Better Robot Training