GraphWrit3R: End-to-End 3D Scene Graph Writing
It aims to produce full 3D scene graphs directly from raw 3D inputs without intermediate object detection.
Current 3D scene graph generators rely on fragile, multi‑stage pipelines that require explicit intermediate steps and assume access to ground‑truth object annotations at inference time, limiting their applicability in real‑world settings and tying them to proprietary, slow models. The authors set out to build a single, open‑source system that ingests raw 3D data and outputs a structured scene graph in one pass.
The proposed system encodes point clouds with Sonata and Gaussian splats with Chorus, projects both modalities onto a shared voxel grid, and aligns them using a per‑voxel contrastive loss. A large language model then decodes the fused representation into a JSON script describing objects, their semantic attributes, and inter‑object relationships, all with a single weight set that supports open‑vocabulary queries.
Evaluated on the 3DSSG benchmark, the method achieves state‑of‑the‑art recall for object classes, predicates, and triplet predictions, outperforming prior approaches that still depend on ground‑truth detections. Qualitative examples and analyses of modality combinations, loss variants, and token fusion strategies further illustrate its robustness.
TakeawayGraphWrit3R shows that an end‑to‑end LLM‑driven pipeline can surpass detection‑dependent baselines on 3DSSG recall metrics.