FlowHOI: Flow-based Semantics-Grounded Generation of Hand-Object Interactions for Dexterous Robot Manipulation

Jan 1, 2026·
Huajian Zeng
Huajian Zeng
,
Lingyun Chen
Jiaqi Yang
Jiaqi Yang
Yuantai Zhang
Yuantai Zhang
,
Fan Shi
,
Peidong Liu
Xingxing Zuo
Xingxing Zuo
· 0 min read
Abstract
We propose FlowHOI, a two-stage flow-matching framework that generates semantically grounded, temporally coherent hand-object interaction (HOI) sequences—comprising hand poses, object poses, and hand-object contact states—conditioned on an egocentric observation, a language instruction, and a 3D Gaussian splatting (3DGS) scene reconstruction. We decouple geometry-centric grasping from semantics-centric manipulation, conditioning the latter on compact 3D scene tokens and employing a motion-text alignment loss. Across the GRAB and HOT3D benchmarks, FlowHOI achieves the highest action recognition accuracy and a 1.7x higher physics simulation success rate than the strongest diffusion-based baseline, while delivering a 40x inference speedup. We further demonstrate real-robot execution on four dexterous manipulation tasks.
Type
Publication
arXiv preprint arXiv:2602.13444
Xingxing Zuo
Authors
Assistant Professor & Lab Director
Leading the Robotics Cognition and Learning (RCL) Group at MBZUAI, focusing on robot perception, 3D vision, and embodied AI.