Text-to-Motion 3.0 sharpens prompt response and detail. Get tips on writing prompts, picking models, and adding AI animation to your pipeline.

There is a difference between naming an action and directing a performance. Text-to-Motion 3.0 is designed to respond to details such as sequence, pacing, posture, body direction, and movement style—not only the action label.
Text-to-Motion 3.0 is Uthana's highest-quality text-to-motion model to date. It gives animators and developers more control through detailed natural-language direction while returning editable skeletal motion for a 3D workflow. This release note explains what changed, how to write useful prompts, when to use V3, and how to access it.
Text-to-Motion 3.0 is for users who need to describe not only the action, but how the performance should unfold. Prompts can combine ordered actions with pacing, posture, body direction, and style to create a more specific starting performance.
3.0 extends Uthana's text-to-motion product line. Previous models remain available and relevant for projects with different requirements.
The release introduces four key technical improvements, each answering a capability gap in prior solutions.
Prompt adherence measures how closely the output motion matches the input direction. If the prompt is "character walks cautiously, pauses, looks over their shoulder, then continues with nervous energy," prompt-adherent models will follow those precise actions and transitions, not flatten them into a generic walk cycle.
More specific terms for pacing, posture, body direction, and transitions give the model clearer performance information. The result can provide a more useful starting point for iteration, while still requiring review and, when needed, animation cleanup.
V3 is Uthana's highest-quality text-to-motion model to date. It produces editable full-body skeletal motion that can be retargeted and refined in Maya, Blender, Unity, or Unreal Engine. As with generated motion generally, review the result and make the production-specific adjustments the shot or character requires.
Text-to-Motion 3.0 accepts prompts in multiple languages. Results can vary by language and phrasing, so teams should test the language and vocabulary they plan to use.
V3 is designed for longer, more detailed direction than earlier models. A prompt can combine a main action with ordered beats, pacing, posture, body direction, and style. Keep the request within the scope of a short motion clip and describe the transitions between related actions.
The model returns the best results from precise, descriptive input. Here is an outline for maximizing output quality.
Write the primary movement requirement explicitly: walk, run, stumble, dodge, collapse, celebrate, search, argue, crouch, reach. Anchor prompts in the action needed. Start from the format: "[character] performs [action] with [pace or energy]." Expand as needed.
Cues such as cautious, tired, angry, confident, or relieved can shape posture, rhythm, and gesture. Pair the cue with observable body direction—for example, the size of the steps, the position of the shoulders, or the amount of energy in the arms—so the intended performance is physically clear.
Physical details provide descriptive guidance rather than locked technical constraints. Describe the head, shoulders, torso, arms, footwork, or weight shifts that matter to the performance. Use Locomotion instead when numeric travel controls are the priority.
Order words such as "then," "before," "after," "gradually," and "suddenly" help express the intended sequence. Keep the actions related and within the duration of a short generated clip. Review the result rather than assuming every requested beat will appear exactly.
Text-to-Motion 3.0 is priced above earlier Uthana text-to-motion models. Choose the model according to the amount of performance direction and output quality the job requires. Review the current rate on the pricing page.
Use V3 when detailed actor direction, body language, or a sequence of related actions matters—hero animations, cinematics, signature character actions, or shots that go through refinement and creative approval.
Earlier models may be a better fit for lower-cost exploration, simple utility actions, or generating multiple options at scale. Test the model on representative prompts before selecting it for a production workflow.
Animators can use 3.0 to start with a technically responsive raw performance and reduce overhead spent on blocking and initial performance iteration. Studios get flexibility in motion workflows without relying entirely on mocap for each variant. Game developers can generate bespoke motion for hero characters, NPCs, or cutscenes while retaining expressive control.
Developers can access 3.0 via the capabilities page and the GraphQL API reference. The async API runs with the create_text_to_motion_job mutation with model: "text-to-motion-3.0". The response returns a Job for polling completion status. Documentation covers parameterization and integration guidelines.
Detailed prompts give animators another way to communicate the intended performance before editing begins. The generated motion remains a starting asset: it does not replace animation judgment, character-specific staging, or final polish.
Prompts can now incorporate intent and attitude with the action directive. Example: change "character picks up an object" to "character picks up the object slowly, hesitates before gripping it, then holds it at arm's length as if unsure whether to keep it." The added notes give the model more information about the intended order, pacing, and body performance.
V3 provides an editable starting performance. Depending on the prompt, selected result, target character, and production standard, an animator may still adjust timing, poses, contacts, or character-specific staging.
To evaluate feature set and output quality, create an account on Uthana. Pricing is pay as you go — motion generation is charged per second of generated output, with no subscription or minimum commitment.
Animators and creators can sign up on Uthana, test prompts at varying specificity levels, and compare results across models. Compare an earlier model with 3.0 on the same prompts to see the difference in prompt responsiveness and detail retention. Select the model to match your production stage.
Technical teams and developers can access complete documentation, async usage patterns, and parameter guidance on the capabilities page. The API supports integration and automation; code examples are provided for Python, TypeScript, and cURL.
Text-to-Motion 3.0 is Uthana’s highest-quality model for turning detailed natural-language direction into editable 3D character motion. Use it when sequence, pacing, posture, body direction, or style matters to the performance; use an earlier model when a simpler, lower-cost generation is sufficient.
Explore the complete Text-to-Motion 3.0 workflow on the product page, start creating in Uthana, or review the API documentation for integration details.