SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes
Abstract
Lay Summary
Developing useful home robots requires exposing them to many realistic homes before deployment, but collecting real robot data is slow, expensive, and risky. Simulation can help, yet current simulated homes are often too empty and simple: they lack the clutter, movable cabinets, small objects, and physical details that make real homes hard for robots. We built SceneSmith, a system that turns a plain-language description into a complete indoor environment that a robot simulator can use directly. It plans the layout, adds furniture, fills shelves and tables with individual objects, creates or retrieves 3D models, and adds physical properties and collision shapes needed for simulation. This makes scenes interactive rather than just visually plausible: a robot can touch objects, pick them up, open cabinets, and be evaluated on tasks. In a user study, room-level SceneSmith scenes were preferred in 92% of realism judgments and 91% of prompt-faithfulness judgments on average against prior methods. SceneSmith also produced 3-6 times more objects while keeping collisions low and most objects physically stable. We further show that these scenes can be used to automatically evaluate robot behaviors from written task instructions. This could make robot training and evaluation more scalable, diverse, and realistic.