MaMi-HOI: Harmonizing Global Kinematics and Local Geometry for Human-Object Interaction Generation
Abstract
Lay Summary
Digital characters are increasingly used in games, virtual reality, robotics, and embodied AI, but making them interact naturally with objects is still difficult. A person should not only appear to understand an instruction such as “sit on the chair” or “pick up the box”; their body must also touch the object in the right places, at the right time, without looking stiff or unnatural. We found that current AI systems for generating human motion often focus on the overall meaning of an action while gradually losing awareness of the object’s exact shape and position. This can produce motions that look reasonable at first glance but fail at important physical details, such as a hand missing a handle or a body floating above a seat. We propose MaMi-HOI, a method that helps generated human motions balance two needs at once: smooth whole-body movement and accurate contact with nearby objects. It first restores fine object details when precise contact is needed, and then adjusts the full body so the motion remains natural rather than forced. This allows digital humans to move through 3D scenes, follow longer action plans, and interact with objects more reliably. Our work can help create more believable virtual characters and support future systems that need to understand and act in physical environments.