Direct the action, sequence, pacing, posture, and style. Text-to-Motion 3.0 turns detailed prompts into editable full-body motion you can preview, apply to a compatible character, and export.
Text-to-Motion 3.0 responds to more than an action label. Describe the order of events, pace, posture, weight shifts, specific body movements, and overall style to give the model direction about how the performance should unfold.
Think of the prompt as actor direction. The result is still generative—not a locked physical constraint—so revise the prompt or generate another take when you want a different interpretation.
See how detailed prompts shape sequence, pacing, posture, and body direction.
All examples are Text-to-Motion 3.0 outputs shown as generated. Uthana generates the subject’s skeletal motion—not additional characters, props, or environments.
Describe the main action and the order of multiple beats, using cues such as then, before, after, slowly, or suddenly.
Direct whether the movement should feel fast, slow, hesitant, urgent, tired, forceful, or relaxed.
Specify direction for the head, shoulders, arms, torso, feet, or shifts in weight.
Combine action, pacing, and body language to shape how the performance reads.
These are language-based directions, not numeric or geometric constraints.
Write what the character should do and how the action should unfold. Add only the details that matter to the result.
Generate full-body skeletal motion and preview it on a character. Revise the prompt or generate another take to explore a different interpretation.
Use a built-in character or upload a compatible bipedal character. Auto-Rigging can add a motion-ready skeleton; retargeting adapts the performance to its proportions.
Download FBX, GLB, or BVH at 24, 30, or 60 fps, then continue adjusting timing, poses, and curves in your DCC or engine.
A natural-language description of a single full-body performance or a short sequence of related actions.
Editable 3D skeletal motion for a character-animation workflow—not a rendered video, finished scene, or character mesh.
Create in the Uthana web app or integrate through the GraphQL API. The V3 API uses an asynchronous generation job.
For the V3 API, the documented target length is 4–10 seconds, with an 8-second default.
Pay-as-you-go credits, charged per second of generated output.
Choose Text-to-Motion when an action or performance is easier to describe than to record or parameterize. V3 is the strongest fit when detailed direction matters.
Start creatingChoose Video-to-Motion when you already have footage that shows the timing and movement you want to reconstruct as editable 3D motion.
Convert video to 3D motionChoose Locomotion when direction, speed, stride count, repeatability, and movement style should come from explicit controls rather than a prompt.
Create controllable locomotionEditable full-body skeletal motion—sequences, pacing, posture, and body direction, ready to retarget, edit, and export.
Character meshes, props, environments, rendered scenes, or guaranteed contact against supplied scene geometry.
Language directs the performance qualitatively—it is not a deterministic numeric controller or an exact final-pose constraint. For explicit travel controls, use Locomotion.
Text-to-motion converts a natural-language description into 3D skeletal motion for a character. Text-to-video creates rendered pixels and scenes. Uthana’s output remains editable, retargetable character animation for a 3D workflow.
Yes, if the character is compatible with Uthana’s bipedal workflow. An unrigged character can first go through Auto-Rigging, and motion retargeting then applies the generated performance to its skeleton and proportions.
No. Text-to-Motion can generate body movement that implies an interaction, but it does not generate the prop, environment, or additional character described by the prompt.
Use Text-to-Motion in the Uthana web app or through the GraphQL API. Usage is pay as you go and charged per second of generated motion. See the pricing page for the current rate.
Describe the action and how it should unfold, then generate editable full-body motion with Text-to-Motion 3.0.