HAND–OBJECT INTERACTION RECONSTRUCTION

4D-HOFHand-Object Flow Matching for
Feed-Forward 4D Interaction Reconstruction

EXPLORE IN 3D

Interaction, from every angle.

Input video RGB
4D-HOF Interactive reconstruction

Drag to rotate · Scroll or pinch to zoom

01 / OVERVIEW

Abstract

Existing methods for 4D hand-object reconstruction often rely on costly per-sequence optimization, while generative approaches typically synthesize interactions from random noise, which can lead to unstable interaction prediction. We introduce 4D-HOF, a feed-forward framework that reconstructs 4D hand-object interactions from coarse but informative estimates produced by vision foundation models. Concretely, we learn a conditional flow matching model that transports foundation-model-derived hand-object states toward an interaction manifold, allowing the model to correct errors in translation, rotation, and alignment in a feed-forward manner. A key advantage of our generative formulation is that it naturally enables test-time guidance within the transport process. Rather than applying a separate post-hoc optimization after reconstruction, we directly steer the evolving generative states using physical interaction constraints and observed 2D evidence, allowing the reconstruction to be refined as part of the generative process itself. By training the generative model on diverse datasets, 4D-HOF generalizes robustly to challenging in-the-wild scenarios. Experiments on out-of-domain benchmarks show that 4D-HOF achieves state-of-the-art performance, producing more stable and accurate 4D hand-object reconstructions.

02 / THE APPROACH

Method overview

View full-resolution figure
4D-HOF framework: contextual scene parsing, object and hand reconstruction, followed by flow matching refinement using image features and hand-object point clouds.
FIGURE 1
Overview of 4D-HOF. Given a monocular video, we construct initial hand-object states by leveraging foundation models, including (a) contextual scene parsing, (b) object reconstruction, (d) hand reconstruction, and (c) depth alignment for recovering metric geometry, producing coarse HOI initialization. We then formulate HOI refinement as a conditional generative bridge matching problem that takes the HOI initialization together with RGB and 3D cues as input and outputs the refined HOI reconstruction. At inference time, test-time guidance adjusts the evolving states to better satisfy physical and image-space constraints.
03 / SEE IT IN MOTION

Qualitative comparison

Same sequence. Different methods.
Explore geometry and image alignment side by side.

Choose a sequenceEXAMPLE 01 / 06

ABF14

Loading videos…GT = Ground truth

All available methods share playback controls. Reconstruction and overlay clips are shown at their original lengths.

04 / ADDITIONAL COMPARISON

GPT-6 Astra Ultra vs. 4D-HOF

For this experiment, we used GPT-6 Astra Ultra. A single run consumed approximately 35% of the weekly usage allowance on a $100 subscription.

ABF12

Overlay comparison
2 methods · Shared playbackGPT-6 Astra Ultra / 4D-HOF