Solutions

Structured 3D human motion for world models.

Convert visible human behavior into skeletal motion that can be evaluated alongside the source video or used as a separate structured dataset.

Video shows behavior, not skeletal state

Video captures appearance, context, and behavior. It does not directly expose the underlying skeleton, joint rotations, root trajectory, or the way those signals change from frame to frame.

A world model may learn useful structure from pixels alone. When an experiment requires explicit 3D human motion, however, raw video does not provide that representation by itself.

Extract a motion representation from video

Uthana converts a supported single-subject video into animated 3D skeletal motion. The self-serve workflow produces an editable motion that can be downloaded as FBX or GLB. Larger research or data programs can define a separate delivery specification around the intended experiment.

The motion does not replace the source video. It provides an additional representation of the observed human movement. Any required alignment fields, timestamps, labels, confidence signals, or tabular outputs must be defined and validated for the engagement rather than assumed to be universal.

Use captured and extracted motion for different jobs

Uthana’s current licensable studio dataset contains more than 150 hours of marker-based human motion. Studio capture provides controlled 3D reference data. Motion extracted from video can provide access to behavior that was not recorded in a capture volume.

The appropriate source depends on the objective. A team may need studio motion, motion derived from its own video, paired video and motion under a defined specification, or labels attached to an existing corpus. These are distinct data products and should be evaluated separately.

Commercial training and evaluation rights are available by agreement.

Start with a testable experiment

World-model work is a scoped part of Uthana’s current physical-AI offering. Begin with the observation format, target motion representation, alignment requirement, and evaluation method. Uthana can then identify what is available from the current products and what requires a data program or pilot.

Uthana does not claim that adding skeletal motion will improve a model without evaluation against the customer’s objective. The first engagement should produce a representative sample and a measurable acceptance test before scale is discussed.

Define the motion signal your experiment needs

Bring a representative video sample, the target representation, and the evaluation question. We will scope the motion extraction, data, and validation work around those inputs.