Turn video into editable 3D character animation.

Upload a video of one person. Uthana uses AI motion capture to reconstruct the performance as 3D skeletal animation you can preview, apply to a compatible character, and export.

AI motion capture from video

A video records pixels. It does not contain an explicit 3D skeleton. Uthana estimates the subject's pose across frames and reconstructs the performance as temporally consistent skeletal animation.

The result is editable character motion—not a new rendered scene. Preview it, apply it to another compatible bipedal character, refine it in an animation tool, or export it into the rest of the production workflow.

The self-serve workflow uses ordinary single-camera footage. It does not require a marker suit or multi-camera capture stage. Because the result is estimated from video, it should not be described as marker-based ground truth or guaranteed physical measurement.

Broadcast dunk footage beside the reconstructed 3D character motion↗
Dunk contest
VIDEO-TO-MOTION 2.1 · EDITABLE SKELETAL MOTION
Open interactive result ↗
EXAMPLES

From reference video to your character

01

Upload the performance

Start with a continuous video of one person performing the movement you want to capture. A phone video can work when the full body is visible and the footage follows the requirements below.

02

Generate and preview the motion

Uthana reconstructs the performance as full-body skeletal motion. Preview the result and inspect whether the timing, facing, major gestures, feet, and hands preserve the parts of the performance that matter.

03

Apply it to a compatible character

Use a built-in character or apply the motion to your own compatible bipedal character through Uthana retargeting. If the character does not have a usable skeleton, Auto-Rigging can add one before the motion is applied.

Apply motion to your character

Auto-rig an unrigged character

04

Export and keep editing

Download the animated result and continue working in the DCC or engine used by your production team. Review the curves, timing, contacts, and target-character fit as part of the normal animation workflow.

Record footage that reconstructs well

Current upload requirements

File typeMP4, MOV, or AVI
Duration2–60 seconds
Frame rate24–120 fps
SubjectOne person

Resolution limits and the full current list are in the API documentation.

For the strongest reconstruction

  • Use one continuous shot with one person in frame.
  • Keep the full body visible, including the feet.
  • Hold the camera steady or use a tripod.
  • Begin standing with both feet on the ground.
  • Use fitted clothing and clear contrast between the subject and background.
  • Shoot in even light without harsh shadows.
  • Avoid heavy motion blur, body-part occlusion, and other people crossing the frame.

These practices improve the information available to the model. They are not a guarantee that every video will reconstruct equally well.

Export editable character animation

Download self-serve results as FBX, GLB, or BVH, with 24, 30, or 60 fps output options. Choose a compatible character for the motion, then continue editing the result in the rest of the animation pipeline.

FORMATSFBXGLBBVH
FRAME RATES24 FPS30 FPS60 FPS
DOWNLOAD → INSPECT → KEEP EDITING
PRICING

Pricing based on output

Video-to-Motion uses pay-as-you-go credits. Pricing is based on output, with no required subscription or minimum commitment. The same account balance works across the web app and GraphQL API.

Download the motions you generate as many times as needed. Paid usage includes commercial use of generated outputs, subject to Uthana's Terms of Use.

Pay as you goNo required subscriptionShared web and API balance
See current Video-to-Motion pricing
API

Build with the Video-to-Motion API

Integrate motion capture from video through Uthana's GraphQL API. Submit a supported video, poll the asynchronous job, retrieve the completed motion result, and download it for a selected compatible character.

01Submit supported video
02Poll asynchronous job
03Retrieve completed motion
04Download for a compatible character

Video-to-Motion 2.0 is the current documented default and returns the base motion result. Video-to-Motion 2.1 returns refined motion results with model-native post-processing.

Read the Video-to-Motion API documentation

Choose video when the reference performance matters

Each model answers a different production question. Pick the input that carries the information you already have.

Video-to-Motion

You already have footage that shows the timing and body movement you want to reconstruct. The performance itself is the reference.

Try Video-to-Motion

Text-to-Motion

The performance is easier to describe than record.

Generate motion from text

Locomotion

Direction, speed, stride count, repeatability, and movement style should come from explicit controls.

Create controllable locomotion
VIDEO DATA

Structured motion from video collections

Video collections contain behavioral variety, but raw footage exposes pixels rather than an explicit 3D skeleton. Uthana can estimate temporally consistent skeletal motion from owned or licensed video so movement can be processed as structured data alongside the source observations.

Bulk extraction and enrichment are scoped engagements. The delivered skeleton, source-video alignment, labels, quality checks, packaging, and permitted use are defined for the project. Video-derived motion remains an estimate rather than marker-based ground truth, so begin with a representative sample and documented acceptance criteria.

Discuss video-motion data

Video-to-Motion FAQ

Can I use phone footage for AI motion capture?

Yes. A phone or ordinary camera can provide usable footage when the video meets the supported file requirements and keeps one person's full body clearly visible. Stable framing, even lighting, and limited occlusion generally provide a stronger source.

Can I apply the motion to my own 3D character?

Yes, if the character is compatible with Uthana's bipedal workflow. Use retargeting to apply the reconstructed motion. If the character is unrigged, Auto-Rigging can first add a compatible skeleton.

Does Video-to-Motion include finger movement?

Video-to-Motion includes finger motion, but the result depends on whether the hands are visible in the footage and whether the selected target character carries compatible finger joints. It should not be treated as a guarantee of exact finger contact.

What can I export?

Self-serve results can be downloaded as FBX, GLB, or BVH at 24, 30, or 60 fps. The result remains editable skeletal animation rather than a finished rendered scene.

How is Video-to-Motion priced?

Pricing is based on output, with no required subscription or minimum commitment. See the Pricing page for current rates.

When should I use video instead of a text prompt?

Use video when the timing and body movement of a specific reference performance matter. Use Text-to-Motion when the action is easier to describe, and use Locomotion when travel behavior should come from explicit controls.

Start from a performance

Upload a video, reconstruct the movement as editable 3D animation, and apply it to a compatible character.