Friday, 4 September 2026Support us

Aube.

News of progress
LabSingle source

Daxiao Robotics Open-Sources Egocentric 3D Hand Model

Languages for this article

Machine-translated from Chinese — read the original text. 6 languages available; yours is one click away.

Hands reach toward a workpiece and are briefly obscured. When they return to the camera’s view, the system still needs to pick up where the previous motion left off. This is precisely where egocentric operation videos are most likely to lose information. Daxiao Robotics, together with the Multimedia Lab at the Chinese University of Hong Kong, Nanyang Technological University and Shanghai Jiao Tong University, has open-sourced ACE-Ego-Hand. The model can analyze an entire video in one pass and directly output both hands’ pose, shape, visibility and 3D positions. The team has produced about 5,000 hours of training data.

Its approach differs from traditional pipelines. Earlier models typically detected hands frame by frame or performed temporal regression over video windows. Some also relied on multiple rounds of diffusion sampling to gradually generate intermediate video or pixel results before estimating hand pose. ACE-Ego-Hand converts a video diffusion model into a geometric encoder: the input first passes through a VAE for compression, the model performs a single forward pass on the latent variables, and spatiotemporal features are then read from the video diffusion Transformer module and passed to a bidirectional spatiotemporal decoder for output.

This bidirectional reasoning mechanism is designed specifically to handle segments in which hands disappear. The model does not use a causal mask or rely only on frames that have already occurred. Instead, it examines information from both before the hand leaves the field of view and after it reappears. The former provides the motion’s starting point, speed and direction; the latter provides an endpoint and pose constraints. Compared with one-way extrapolation after a hand leaves the frame, this approach is intended to reduce trajectory drift and preserve hand identity and motion continuity.

Camera changes are also incorporated into the design. During testing, ACE-Ego-Hand requires no explicit camera intrinsics—the camera’s own geometric parameters. For native fisheye ultra-wide-angle video, it requires neither advance calibration nor image undistortion. Instead, it directly predicts observation rays from the distorted image into 3D space and recovers the hands’ metric 3D positions in the camera coordinate system.

According to the team’s test results, the model reaches about 63.1 fps on a single A100 GPU (graphics processing unit), roughly 33x the speed of ViDiHand’s high-accuracy configuration. It can support production-line operation data collection, robot teaching, manual assembly process analysis, dexterous-hand motion retargeting and the creation of imitation-learning data.

about 63.1 fpsProcessing speed on a single A100 GPU

Sources — read the originals(Paris time)

第一电动网ZH
0000

Read next

Comments

Loading the thread…

Sign in to leave a comment. Sign in