Resources

3D human motion datasets compared

The best human motion dataset depends on the task. Dataset size is not enough — ask where the motion came from, whether it is measured or estimated, which representation it uses, which annotations were added, whether datasets overlap, and what the license permits.

Guide
Human motion data
Dataset facts checked August 28, 2026
Key point

These datasets are not seven independent piles of newly captured motion. Several reuse or annotate the same underlying sequences, so their hours cannot simply be added together.

AMASS is a foundational archive of optical motion-capture datasets in a common body representation. HumanML3D and KIT Motion-Language are established text-to-motion resources. BABEL adds dense action semantics to AMASS sequences. Motion-X expands toward expressive whole-body motion, hands, face, and mixed video-derived sources. MOSAIC packages multi-source human motion with Unitree G1-retargeted data for humanoid tracking and teleoperation research. Uthana provides separately captured, licensable studio motion for teams that need commercial training or evaluation rights.

The short selection guide

Which human motion dataset to start with by task, with the reason and what to check before use.
If your task is…Start withWhyCheck before use
General human-motion modellingAMASSBroad optical-mocap archive unified in a common body-model representationSource subsets, representation version, research-only licensing
Text-to-motion benchmark comparisonHumanML3DWidely used motion-caption benchmark with standardised 20 FPS preprocessingAMASS/HumanAct12 lineage, preprocessing choices, underlying licences
Smaller motion-language or robotics-oriented language workKIT Motion-LanguageNatural-language annotations with a unified MMM representation and robotics lineageSource terms, XML/MMM tooling, dataset version
Action recognition or temporal action localisationBABELSequence- and frame-level labels over AMASS motionIt is an annotation layer, not independent newly captured motion
Expressive whole-body generation with hands and faceMotion-XSMPL-X whole-body sequences with sequence and frame-level textMixed provenance, estimated video motion, subset-specific access and licences
Humanoid motion tracking or G1-reference workflowsMOSAICHuman AMASS-style files plus Unitree G1-retargeted NPZ dataUpstream tool/data licences, simulator/controller assumptions, composition
Commercial training or evaluation dataUthana Studio Motion DatasetMore than 150 hours of separately captured, full-body marker-based motion with finger data and commercial training and evaluation rights available by agreementSpecify the target skeleton, delivery format, and subset requirements

Dataset lineage matters

Many motion-dataset comparisons accidentally double-count derived data.

AMASS is not one capture session. It unifies multiple optical marker-based mocap datasets through the MoSh++ pipeline and a common body-model representation. Its original release reported more than 40 hours, more than 300 subjects, and more than 11,000 motions.

BABEL does not add a separate 43-hour capture collection. It adds sequence- and frame-level action labels to about 43 hours of AMASS motion.

HumanML3D is also derived. It preprocesses motion from AMASS and HumanAct12, standardises clips to 20 FPS, and pairs 14,616 motions with 44,970 text descriptions. Its 28.59 hours should not be added to AMASS as independent newly captured motion.

Motion-X combines eight existing datasets with motion estimated from online video and unifies the result in SMPL-X at 30 FPS. Its repository explicitly sends users back to original mocap sources and licences rather than redistributing all underlying motion.

MOSAIC combines optical, inertial, and generated human motion, stores human motion in an AMASS-style layout, and provides Unitree G1 motion retargeted through GMR and converted through BeyondMimic. It is both a motion collection and a task-specific processing pipeline.

Uthana's studio dataset sits outside that public-dataset lineage. The licensable motion was separately captured for production rather than aggregated from existing research datasets, and it is licensed directly rather than inherited through upstream research terms.

Master comparison table

Seven human motion datasets compared by provenance, scale, representation, annotations, robot-retargeted data, licence, and best use.
DatasetProvenanceReported scaleRepresentationAnnotationsRobot-retargetedLicence / accessBest suited for
AMASSMultiple optical marker-based mocap datasets unified with MoSh++Original release: >40 hours, >300 subjects, >11,000 motionsSMPL-family; repository documents SMPL+H and DMPL toolingInherits source motion metadata; not primarily a language datasetNo standard robot package in core AMASSFree for non-commercial scientific research, education or artistic projects; source terms also matterHuman-motion priors, representation learning, animation research, benchmark source data
KIT Motion-LanguageAggregated mocap databases in a unified Master Motor Map representation3,911 motions; 11.23 hours; 6,278 annotations (2016 release)MMM XML representationCrowdsourced free-form motion descriptionsRobotics lineage; no single standard robot packageDescribed as open; verify dataset and underlying source termsMotion-language research, retrieval, generation, human-to-robot semantics
HumanML3DDerived from AMASS and HumanAct12 with standardised preprocessing14,616 motions; 44,970 descriptions; 28.59 hoursStandardised joint-motion representation; 20 FPS3–4 sentence descriptions per motionNo core robot-retargeted releaseRepository access does not override AMASS or source-dataset termsText-to-motion training and evaluation, motion-text retrieval
BABELLanguage/action annotations over AMASS sequences~43 hours of AMASS; 28,055 sequence labels; 63,353 frame labels; 250+ categoriesAMASS motion representationSequence labels and temporally aligned frame-level action labels; overlaps allowedNo core robot-retargeted releaseSupported for academic research; AMASS access/licence required for motionAction recognition, temporal localisation, semantic motion segmentation
Motion-XEight existing datasets plus motion estimated from online video81.1K clips; 15.6M poses; 144.2 hours reported; 30 FPSSMPL-X whole-body motion81.1K sequence-level labels and 15.6M frame-level pose descriptionsNo single core robot packageAuthorisation is for non-commercial use; upstream datasets retain their own licencesWhole-body text-to-motion, hands/face, expressive motion, mesh recovery
MOSAICOptical Vicon mocap, inertial IO-AI motion, and GENMO-generated motionNo single independent total should be inferred; package is 45.9 GB as of this checkHuman motion in AMASS-style files; G1 motion in NPZGENMO prompts; source-specific metadata; adaptor data described separatelyYes — Unitree G1 motions retargeted with GMR, converted with BeyondMimicDataset card: CDLA-Permissive-2.0; external tools and source data retain their own termsHumanoid tracking, teleoperation, G1 reference motion, sim-to-real research
Uthana Studio Motion DatasetCaptured in an AAA production environment on a 120-camera marker-based system150+ hours currently licensableFull-body skeletal motion with fingers; target skeleton and delivery format can be customized, including UE5, MHR, and Unitree G1 workflowsAction labels, temporal segmentation, root trajectories, joint trajectories, foot contacts, and performer/session metadataMotion can be retargeted to the customer's target embodiment; Unitree G1 is a public output exampleCommercial training and evaluation rights available by agreementCommercial training or evaluation programs that require documented provenance, a target-specific representation, and support

Dataset facts checked August 28, 2026. Statistics describe the cited release or repository state, not a guarantee that every file is currently downloadable under one uniform licence.

How to choose a human motion dataset

Start with the learning task

A dataset built for text-to-motion is not automatically the best source for humanoid control. A dense action-label dataset is not automatically a clean generative-motion corpus. A large body-model dataset may not include hands, objects, or language.

Write the evaluation question first: are you training a motion generator, pose estimator, tracker, controller, or world model? Is the target a human body model, animation skeleton, or humanoid robot? Do you need whole-body kinematics, hands, face, objects, scenes, contacts, or robot actions? Is the data used for training, fine-tuning, evaluation, retrieval, or simulation? Must the output support commercial deployment?

Inspect provenance

Marker-based mocap records trajectories in a calibrated capture setup. Inertial capture measures orientation through wearable sensors. Video-derived datasets estimate 3D structure from pixels. Generated datasets synthesise motion from a model. Derived datasets may crop, resample, normalise, annotate, or retarget existing motion. These categories are all useful. They have different error models.

Match the representation

SMPL, SMPL+H, SMPL-X, MMM, raw joint arrays, BVH skeletons, and robot NPZ files are not interchangeable. Conversion can change joint counts, rotations, root motion, contacts, and body shape. Document hierarchy and joint names; global versus local positions; rotation representation; root and heading convention; coordinate axes and handedness; units and scale; frame rate and timestamps; body-shape parameters; contact and confidence fields; and missing-data and interpolation rules.

Evaluate annotations and temporal granularity

Sequence captions describe an entire clip. Frame-level labels identify when an action begins and ends. Pose descriptions describe individual frames. Contact states label physical relationships. These serve different objectives. For long, multi-action motion, a single caption can hide transitions and overlapping behaviour.

Check leakage and overlap

If train and test sets derive from the same underlying AMASS sequence, performer, or source dataset, evaluation can be misleading even when filenames differ. Track source identity through every derived dataset. For model comparisons, preserve the official split unless there is a documented reason to change it.

Treat licensing as a model constraint

Public download does not mean unrestricted commercial use. Some datasets permit non-commercial research only. Some require separate registration. Some derivative repositories distribute annotations but not the underlying motion. Mixed datasets can inherit multiple licence obligations. The safest operational practice is a licence manifest at the sequence or subset level.

Best datasets by task

Text-to-motion

Start with HumanML3D when comparability with established text-to-motion work matters. Use KIT-ML when its MMM representation, annotations, or robotics context fit the project. Add Motion-X when hands, face, expressive whole-body motion, or broader mixed-source coverage is required.

Action recognition and temporal localisation

BABEL is the clearest choice in this group because it provides sequence- and frame-level action labels over AMASS motion. It is designed to represent multiple and overlapping actions rather than only one caption per clip.

Human-motion priors

AMASS remains foundational because it aggregates many optical mocap datasets in a common representation. It offers broad human motion without imposing one text-generation preprocessing pipeline. Note that the AMASS licence grants non-commercial use only.

Humanoid tracking and teleoperation

MOSAIC is the most directly packaged option in this comparison. It includes human AMASS-style files, G1-retargeted motion, and a published downstream system. AMASS also remains an important upstream source, but requires a retargeting and control pipeline.

Commercial model training and evaluation

Public research datasets often do not provide the commercial rights or support needed for deployment. Uthana is the relevant option in this set when the requirement is a negotiated commercial licence, technical delivery specification, custom enrichment, or target-embodiment mapping. The buyer should still compare a Uthana sample with internal requirements — commercial availability is not a substitute for task fit.

Ground truth, estimated motion, and pseudo-ground truth

Marker-based mocap observes tracked markers through calibrated cameras. A processing pipeline converts those trajectories into a skeleton or body model. The result is not “raw truth” — fitting and cleanup still make assumptions — but its spatial provenance differs from a monocular estimate.

Video-derived motion estimates 3D body structure from pixels. It can reach behaviours, scenes, and scale that a studio program cannot easily reproduce. Its errors depend on visibility, camera motion, depth ambiguity, motion blur, clothing, and the pose model.

Generated motion samples from a learned model rather than reconstructing a specific performer. It can target rare or missing categories, but it also reproduces the model's training distribution and failure modes.

Use the labels precisely: captured (measured through a capture system); derived (transformed from an existing motion source); estimated or pseudo-ground-truth (inferred from video or another observation); generated (synthesised by a model); and retargeted (mapped from one body representation to another). A mixed dataset should preserve these fields at the sequence level.

What public motion datasets still miss

No dataset in this comparison resolves every gap. Common limitations include limited commercial rights; body-only motion without detailed fingers or face; weak coverage of object interaction and force; few paired video and precise 3D sequences; inconsistent contact annotations; narrow performer or environment diversity; short isolated clips without long transitions; derived datasets that obscure source overlap; no target-robot representation; and no synchronised robot actions, force, torque, or tactile signals.

These are acquisition questions, not just preprocessing questions. If the missing behaviour or signal is central to the model, a custom capture or labelling program may be more efficient than forcing an existing public benchmark to serve a different purpose.

Discuss a custom dataset · Explore labeling and enrichment

A dataset procurement checklist

Before training, request or document:

  1. Dataset version and release date.
  2. Sequence-level provenance and source dataset.
  3. Capture, estimation, generation, and retargeting method.
  4. Hours, clips, frames, performers, and duplicate policy.
  5. Skeleton, body model, joint count, and body coverage.
  6. Frame rate, coordinate system, units, axes, and root convention.
  7. Joint position and rotation representations.
  8. Labels, annotator process, segmentation, contacts, and confidence.
  9. Object, scene, video, audio, and other paired modalities.
  10. Train/validation/test split and source-level leakage controls.
  11. QA checks, rejection criteria, and known failure modes.
  12. Licence for training, evaluation, commercial deployment, derivatives, and redistribution.
  13. Delivery format, versioning, updates, and sample access.
  14. Target-embodiment mapping and downstream validation scope.

Frequently asked questions

What is the largest 3D human motion dataset?

“Largest” depends on the unit and provenance. Motion-X reports 81.1K clips, 15.6 million whole-body poses, and 144.2 hours, but combines existing datasets with motion estimated from online video. AMASS's original release reports more than 40 hours of optical mocap unified from multiple sources. Compare captured hours, derived clips, frames, and annotations separately.

Is HumanML3D part of AMASS?

HumanML3D derives motion from AMASS and HumanAct12, applies standardised preprocessing, and adds natural-language descriptions. It is a distinct benchmark package, but much of its underlying motion is not independent of AMASS.

Is BABEL a separate mocap dataset?

BABEL is primarily an annotation dataset over AMASS motion. It adds sequence- and frame-level action labels to about 43 hours of AMASS sequences rather than contributing a separate 43-hour capture collection.

Which dataset is best for text-to-motion?

HumanML3D is the most established benchmark in this comparison. KIT-ML is useful for a smaller motion-language dataset with robotics lineage. Motion-X is relevant when whole-body hands, face, and expressive SMPL-X motion matter.

Which dataset is best for humanoid robots?

MOSAIC is the most directly robot-oriented package compared here because it includes Unitree G1-retargeted motion and a published tracking/teleoperation pipeline. Human motion still requires target-specific feasibility and control validation.

Can AMASS be used commercially?

The standard AMASS licence grants use for non-commercial scientific research, non-commercial education, and non-commercial artistic projects. Commercial use requires separate rights and may also depend on the underlying source datasets.

Does a public dataset allow commercial model training?

Not automatically. Public access, open-source code, a research licence, and commercial model-training rights are different. Review the dataset licence, every upstream source licence, and any restrictions on derivatives or redistribution.

What is the difference between marker-based and video-derived motion data?

Marker-based motion is measured through a calibrated capture system and then fit to a skeleton or body model. Video-derived motion estimates 3D structure from pixels. Marker-based data offers stronger spatial provenance; video-derived data offers broader behavioural and environmental scale.

Motion data built for training

Evaluate a representative sample, the schema, provenance, and rights before choosing a dataset. Uthana provides licensed studio motion and can build custom capture, extraction, labeling, and embodiment-transfer programs around a defined training specification.

Uthana Studio Motion Dataset

What it is

Uthana’s current licensable studio dataset contains more than 150 hours of full-body, marker-based human motion with finger data, captured in an AAA production environment on a 120-camera system.

The public figure describes the currently licensable studio dataset. It should not be combined with Uthana’s separate motion-asset counts, or used to imply a public clip, frame, performer, prop, interaction, or face-coverage figure.

Representation and delivery

The target skeleton and delivery format can be customized for the buyer’s workflow. Examples include UE5, MHR, and Unitree G1 targets. The exact coordinate convention, frame rate, file structure, and delivery organization are specified for the licensed package rather than forced into one universal public format.

The licensable data includes action labels, temporal segmentation, root trajectories, joint trajectories, foot contacts, and performer/session metadata. Hand-object contact labels are not included as a standard annotation.

Request a representative sample and the corresponding specification before evaluating technical fit. Sample access is handled through the commercial team rather than a public download.

Best uses

  • Commercial model training or evaluation where provenance and support matter
  • Programs that need marker-based human motion under an agreed specification
  • Teams that require a specific character skeleton, robot embodiment, or delivery format

Commercial training and evaluation rights are available by agreement.

Limitations

Uthana does not publish performer, clip, or frame counts for the dataset. The public coverage claim is full-body motion with fingers. Face, prop, human-object interaction, multi-person, and hand-object contact-label coverage should not be inferred without a separate approved specification.

The data is human motion, not robot-native sensor, torque, force, or action data. Robot use requires a separate embodiment and validation step.

Related product pages

Studio Motion Dataset

Licensed marker-based studio motion. Commercial training and evaluation rights available by agreement.

Explore the studio dataset

Data specifications

Formats, skeletons, labels, metadata, and usage rights for Uthana motion data.

Review data specifications

Custom dataset programs

Capture, extract, label, and deliver motion to a defined training specification.

Discuss a custom dataset

Related resources

Resource

Human motion data for robotics

Read the guide
Resource

Motion retargeting explained

Read the guide
Resource

Best AI motion capture tools

Read the comparison