Text-to-Motion 3.0

Generate 3D character animation from text.

Direct the action, sequence, pacing, posture, and style. Text-to-Motion 3.0 turns detailed prompts into editable full-body motion you can preview, apply to a compatible character, and export.

“A boxer crouches low with their weight on the back foot, raises both arms into a guard position, then shifts their weight forward onto the front foot and performs an uppercut”
Text-to-Motion 3.0 · Shown as generated

Direct more than an action

Text-to-Motion 3.0 responds to more than an action label. Describe the order of events, pace, posture, weight shifts, specific body movements, and overall style to give the model direction about how the performance should unfold.

Think of the prompt as actor direction. The result is still generative—not a locked physical constraint—so revise the prompt or generate another take when you want a different interpretation.

See what the prompt changes.

See how detailed prompts shape sequence, pacing, posture, and body direction.

All examples are Text-to-Motion 3.0 outputs shown as generated. Uthana generates the subject’s skeletal motion—not additional characters, props, or environments.

Sequence
“The character starts way back in the background, runs forward five steps as if panicked and scared, then realizes that they are being chased by a friend, not a monster, and starts to laugh with their hands on their knees”
The motion follows the prompt’s main beats: run forward, stop, then lean forward with hands toward the knees. The chasing friend is not generated.
Implied interaction
“A person bends down, slowly lifts a heavy box with both hands, staggers a step under the weight, then sets it down carefully on a table to their left”
The body performs the lift and set-down; the box and table are not generated.
Pacing and posture
Slow
“A person crosses the room with slow, hesitant steps, shoulders hunched, pausing halfway.”
Quick
“A character crosses the room with quick, confident strides, arms swinging freely.”

The same crossing action changes through pace, posture, step length, and arm swing. Same character; only the prompt changed.

Implied prop
“An archer creeps forward and is looking around. She sees a target and quickly fires an arrow, then immediately takes cover.”
The archer’s creep, draw, release, and take-cover are generated; the bow and arrow are implied, not rendered.

What you can direct with a prompt

Actions and sequences

Describe the main action and the order of multiple beats, using cues such as then, before, after, slowly, or suddenly.

Pacing and energy

Direct whether the movement should feel fast, slow, hesitant, urgent, tired, forceful, or relaxed.

Posture and body direction

Specify direction for the head, shoulders, arms, torso, feet, or shifts in weight.

Movement style

Combine action, pacing, and body language to shape how the performance reads.

These are language-based directions, not numeric or geometric constraints.

From prompt to editable character motion

01

Describe the performance

Write what the character should do and how the action should unfold. Add only the details that matter to the result.

02

Generate and preview

Generate full-body skeletal motion and preview it on a character. Revise the prompt or generate another take to explore a different interpretation.

03

Apply it to a character

Use a built-in character or upload a compatible bipedal character. Auto-Rigging can add a motion-ready skeleton; retargeting adapts the performance to its proportions.

04

Export and keep editing

Download FBX, GLB, or BVH at 24, 30, or 60 fps, then continue adjusting timing, poses, and curves in your DCC or engine.

Text-to-Motion 3.0 at a glance

Input

A natural-language description of a single full-body performance or a short sequence of related actions.

Output

Editable 3D skeletal motion for a character-animation workflow—not a rendered video, finished scene, or character mesh.

Access

Create in the Uthana web app or integrate through the GraphQL API. The V3 API uses an asynchronous generation job.

Length

For the V3 API, the documented target length is 4–10 seconds, with an 8-second default.

Pricing

Pay-as-you-go credits, charged per second of generated output.

Choose the right motion input

Text-to-Motion

Choose Text-to-Motion when an action or performance is easier to describe than to record or parameterize. V3 is the strongest fit when detailed direction matters.

Start creating

Video-to-Motion

Choose Video-to-Motion when you already have footage that shows the timing and movement you want to reconstruct as editable 3D motion.

Convert video to 3D motion

Locomotion

Choose Locomotion when direction, speed, stride count, repeatability, and movement style should come from explicit controls rather than a prompt.

Create controllable locomotion

What Text-to-Motion creates—and what it does not

Creates

Editable full-body skeletal motion—sequences, pacing, posture, and body direction, ready to retarget, edit, and export.

Does not create

Character meshes, props, environments, rendered scenes, or guaranteed contact against supplied scene geometry.

Language directs the performance qualitatively—it is not a deterministic numeric controller or an exact final-pose constraint. For explicit travel controls, use Locomotion.

Text-to-Motion FAQ

What is text-to-motion, and how is it different from text-to-video?

Text-to-motion converts a natural-language description into 3D skeletal motion for a character. Text-to-video creates rendered pixels and scenes. Uthana’s output remains editable, retargetable character animation for a 3D workflow.

Can I use my own 3D character?

Yes, if the character is compatible with Uthana’s bipedal workflow. An unrigged character can first go through Auto-Rigging, and motion retargeting then applies the generated performance to its skeleton and proportions.

Does Text-to-Motion create objects or environments?

No. Text-to-Motion can generate body movement that implies an interaction, but it does not generate the prop, environment, or additional character described by the prompt.

How do I access Text-to-Motion 3.0, and how is it charged?

Use Text-to-Motion in the Uthana web app or through the GraphQL API. Usage is pay as you go and charged per second of generated motion. See the pricing page for the current rate.

Direct your next character performance.

Describe the action and how it should unfold, then generate editable full-body motion with Text-to-Motion 3.0.