Models

Convert video to editable 3D motion.

Upload a single-person video recorded with an ordinary camera. Uthana estimates 3D skeletal motion over time and returns animation you can preview, retarget, and export.

From pixels to skeletal animation

A video records pixels, not an explicit 3D skeleton. Video-to-Motion estimates the subject's pose across frames and reconstructs the performance as temporally consistent skeletal animation.

The result is not motion baked into a rendered video. Preview the animation, apply it to a different bipedal character, refine it in an animation tool, or export it into the rest of the production pipeline.

Video-derived motion is an estimate. It is useful when motion needs to come from ordinary footage, but it should not be described as equivalent to marker-based studio capture or guaranteed physical ground truth.

Record footage that reconstructs well

The current self-serve workflow accepts MP4, MOV, and AVI files between 2 and 60 seconds, at 24–120 fps and 300–4096 px resolution.

For the strongest result:

  • Use one continuous shot with a single person in frame.
  • Keep the whole body visible, including the feet.
  • Hold the camera steady, or use a tripod.
  • Start from a standing position with both feet on the ground.
  • Wear fitted clothing, and shoot in even light against a contrasting background.
  • Avoid occlusion, heavy motion blur, and other people crossing the frame.

From reference performance to character animation

Use a phone or existing video as the creative reference for a gesture, athletic movement, dance, or acting choice. Generate the motion, apply it to the production character, and continue refining the curves in the DCC or engine where the rest of the work happens.

This is most useful when the performance itself matters and a text description would be too ambiguous. The video supplies the timing and movement reference; Uthana supplies editable 3D animation.

Structured motion from video collections

Video contains far more behavioral diversity than traditional motion-capture datasets, but raw footage exposes pixels rather than an explicit 3D skeleton. Uthana estimates temporally consistent skeletal motion from video so the movement can be processed as structured data alongside the source observations.

For teams working with owned or licensed video collections, Uthana can scope bulk extraction and enrichment as an enterprise engagement. The output is video-estimated motion, not marker-based ground truth. Its usefulness depends on source quality, the required representation, and the customer’s acceptance criteria.

The delivered skeleton, source-video alignment, labels, quality checks, file packaging, and permitted use are defined for the engagement. Use a sample to determine whether the estimated motion is suitable before scaling the collection.

Export the motion

Download self-serve motion as FBX, GLB, or BVH, with 24, 30, or 60 fps output options. Apply the motion to another bipedal character through Uthana retargeting.

Start with a video

Use the self-serve product for a single animation, or talk to Uthana about processing a larger video collection.