Open, well-documented training data remains one of the harder bottlenecks in fashion AI research. Benchmark datasets for garment recognition, body measurement estimation, fabric behaviour, and cloth simulation exist, but they are scattered across academic repositories, institutional pages, and supplementary materials from conference papers. This catalogue brings six of the most practically useful ones into one place, with notes on licence terms, approximate scale, and the model types each dataset suits. It is aimed at ML engineers and AI researchers working on fashion-specific tasks.
Key takeaways
- Several widely cited fashion datasets carry non-commercial research licences that restrict deployment in production systems without further negotiation.
- Cloth simulation benchmarks and fabric-property datasets are significantly underrepresented compared with garment-image datasets, creating a gap for researchers working on physical AI.
- Body-measurement datasets typically require careful demographic documentation before use in EU-regulated systems under the AI Act's requirements for high-risk applications.
- Dataset size alone does not predict utility: annotation quality, garment taxonomy consistency, and pose diversity matter more for downstream fine-tuning.
- Preprint repositories such as arXiv cs.GR are currently the fastest route to newly released simulation datasets before they reach formal publication.
What makes a fashion dataset research-ready?
Before cataloguing specific datasets, it is worth stating the evaluation criteria used here. A dataset is considered research-ready when it provides: a clear, machine-readable licence; a persistent download location (not a lab webpage that disappears when a PhD student graduates); documented annotation methodology; and enough metadata to assess demographic or garment-category coverage. Datasets that meet only some of these criteria are noted accordingly.
1. DeepFashion (MMLAB, CUHK)
DeepFashion is one of the most cited large-scale garment datasets in the academic literature. It contains roughly 800,000 images annotated with clothing category labels, attribute tags (collar type, sleeve length, pattern), landmark coordinates, and cross-modal retrieval pairs linking consumer photos to catalogue images. A subset, DeepFashion2, adds dense instance segmentation masks and clothing-item matching across domains.
What it is: A multi-task benchmark covering classification, retrieval, landmark detection, and segmentation across a broad taxonomy of garment categories.
Why it matters for researchers: The scale and annotation depth make it suitable for training and evaluating convolutional and transformer-based recognition models. The cross-domain pairs (street photo to shop photo) are particularly useful for virtual try-on research and domain-adaptation work.
What to watch: The licence is non-commercial research only. Any model trained on DeepFashion that you intend to deploy in a commercial product requires separate clearance. Annotation was performed at a single point in time, so trend-sensitive categories (silhouette shapes, colour palettes) reflect the period of collection rather than current fashion.
2. Fashion-MNIST (Zalando Research)
Fashion-MNIST was designed as a drop-in replacement for the original MNIST handwritten-digit dataset, using 70,000 greyscale images of ten clothing categories (T-shirt, trouser, pullover, dress, coat, sandal, shirt, sneaker, bag, ankle boot) at 28×28 pixels. It is released under the MIT licence.
What it is: A small-image classification benchmark intended for rapid prototyping, architecture comparison, and teaching. It is not a production-scale dataset.
Why it matters for researchers: The MIT licence is one of the most permissive available, which makes Fashion-MNIST useful for benchmarking experiments where you need a reproducible, legally uncomplicated baseline. Because the images are low-resolution and the category taxonomy is coarse, it is best used for ablation studies and sanity checks rather than as a primary training corpus.
What to watch: The resolution and category granularity are too limited for any task requiring fine-grained attribute recognition or realistic texture modelling. Researchers sometimes over-report performance on Fashion-MNIST as evidence of general fashion-AI capability, which it is not.
3. CLOTH3D (Inria / CVPR)
CLOTH3D is a large-scale synthetic dataset of 3D clothed human sequences, generated using physics-based simulation. It contains thousands of sequences of animated human bodies wearing a variety of garment types, with per-frame ground-truth 3D mesh geometry for both the body and the clothing layer.
What it is: A simulation-derived benchmark for 3D garment reconstruction, cloth deformation modelling, and body-shape estimation under clothing.
Why it matters for researchers: Acquiring ground-truth 3D geometry from real garments at scale is impractical, so synthetic datasets like CLOTH3D fill a gap that real-world capture cannot easily address. It is well suited to training and evaluating neural networks for garment mesh recovery, layered body modelling, and physics-informed cloth simulation — areas where NVIDIA Research and groups publishing on arXiv cs.GR have been active.
What to watch: Synthetic-to-real transfer remains an open problem. Models trained on CLOTH3D often require domain adaptation before they generalise to real camera footage. The garment taxonomy is also narrower than real-world wardrobes, and fabric mechanical properties in the simulation are approximated rather than measured from physical samples.
4. CAESAR (Civilian American and European Surface Anthropometry Resource)
CAESAR is a body-measurement dataset collected from thousands of civilian subjects in the United States, the Netherlands, and Italy, using 3D whole-body scanning. It records hundreds of body dimensions per subject alongside demographic metadata. Access is managed through the SAE International licensing process.
What it is: A structured anthropometric dataset for body-shape modelling, size-chart calibration, and fit-prediction research.
Why it matters for researchers: CAESAR is one of the few large-scale datasets that provides actual 3D body scans rather than self-reported measurements, which makes it substantially more reliable for training parametric body models (such as SMPL variants) and for validating size-recommendation systems. The multi-country sampling gives it broader demographic coverage than many alternatives.
What to watch: Licensing is not open in the conventional sense — access requires an institutional agreement with SAE International and carries a cost. The data was collected in the late 1990s and early 2000s, so it predates the body-composition shifts documented in more recent population surveys. For EU AI Act compliance purposes, the demographic documentation is relatively thorough, but the age of the data is a material limitation that should be disclosed in any system card or conformity assessment.
5. FabricNet / KTH-TIPS2 (texture and fabric properties)
KTH-TIPS2 (Textures under varying Illumination, Pose and Scale, second edition) is a material-texture dataset from KTH Stockholm covering 11 material classes including several textile categories (cotton, linen, wool, corduroy). Each class contains images captured under multiple illumination conditions, scales, and viewing angles.
What it is: A controlled texture-recognition benchmark for material classification under photometric variation.
Why it matters for researchers: Fabric texture recognition is a prerequisite for automated quality inspection, digital twin construction, and the kind of wear-comfort simulation described in recent Industry 4.0 research — including work on computational modelling of garment fabric texture that situates AI-driven simulation within fashion technology workflows. KTH-TIPS2 provides a reproducible baseline for evaluating texture-classification architectures before moving to domain-specific fabric datasets.
What to watch: The dataset covers material classes at a coarse level of granularity; it does not distinguish between weave structures (plain, twill, satin) or fibre compositions that matter for downstream simulation of mechanical behaviour. It is a starting point for texture modelling, not a complete fabric-property resource. Researchers needing mechanical parameters (tensile strength, bending rigidity, shear modulus) will need to supplement with physical measurement data or purpose-built datasets from textile engineering literature.
6. Human3.6M (with garment overlay use cases)
Human3.6M is a large-scale dataset of 3.6 million video frames of human subjects performing everyday actions, captured in a controlled laboratory environment with synchronised RGB cameras, depth sensors, and motion-capture markers. It is primarily a pose-estimation benchmark, but it has been widely used in clothed-human modelling research as a base for overlaying garment geometry.
What it is: A multi-modal motion and pose dataset used as a foundation for clothed-human reconstruction and animation research.
Why it matters for researchers: The combination of accurate 3D joint positions, video frames, and controlled capture conditions makes Human3.6M a reliable scaffold for research that needs to separate body pose from garment deformation — a core challenge in virtual try-on and cloth simulation. Meta AI Research (operating as Meta Superintelligence Labs) and other groups have built on datasets of this type for work on physical AI and human-body modelling.
What to watch: The licence restricts use to non-commercial research. Subject diversity is limited — the dataset was captured with a small number of subjects in a single laboratory — which constrains the generalisability of models trained on it. For EU AI Act purposes, the narrow subject pool is a documented limitation that affects any high-risk application relying on body-shape inference.
How do these datasets relate to EU AI Act obligations?
If your fashion AI system falls within a high-risk category under the EU AI Act — for example, a system used to make consequential decisions about individuals based on body measurements or biometric-adjacent data — you are required to document training data provenance, assess representativeness, and identify known biases. Several datasets listed here (CAESAR, Human3.6M) carry documented demographic limitations that must appear in your system's technical documentation and conformity assessment. Non-commercial research licences (DeepFashion, Human3.6M) also create a legal discontinuity if a research prototype transitions to a deployed product without renegotiating data rights.
Our earlier analysis of the gap between fashion AI ambitions and actual research readiness — covered in our piece on fashion AI ambitions versus readiness — is directly relevant here: the datasets available today were mostly designed for academic benchmarking, not for the documentation standards that regulated deployment now requires.
Where to find emerging datasets
For cloth simulation and physical AI specifically, the fastest route to newly released datasets is the arXiv cs.GR preprint feed, which carries graphics and simulation papers before formal publication. NVIDIA Research also releases open-source code libraries and datasets alongside its papers on generative AI and rendering; checking its publications page when a relevant paper appears is worthwhile, as supplementary data releases sometimes follow.
FAQ
Which of these datasets can be used in a commercial product without renegotiating the licence? Fashion-MNIST (MIT licence) is the clearest case for commercial use. The others — DeepFashion, CLOTH3D, Human3.6M — carry non-commercial research restrictions. CAESAR requires an institutional agreement. Always review the current licence terms directly with the dataset provider before deployment.
Are there open datasets specifically for fabric mechanical properties? Not in a single consolidated open repository. KTH-TIPS2 covers texture recognition. Mechanical parameters (bending rigidity, shear modulus) typically come from physical testing and are not yet standardised in open ML datasets. This is an active gap in the field.
How should I document dataset limitations for EU AI Act compliance? For each dataset used in training, record: the licence and its restrictions, the demographic and garment-category coverage, the data collection period, and any known biases documented by the dataset authors. This documentation forms part of the technical file required for high-risk AI systems.
Can synthetic datasets like CLOTH3D replace real-world data for cloth simulation? Synthetic datasets reduce the cost of ground-truth annotation, but synthetic-to-real transfer gaps remain a documented research problem. In practice, most robust simulation models combine synthetic pre-training with at least some real-world fine-tuning data.
Where do new fashion AI datasets typically appear first? Most are released alongside conference papers at CVPR, ICCV, ECCV, and SIGGRAPH, or as preprints on arXiv. The cs.GR and cs.CV sections of arXiv are the most relevant feeds for cloth simulation and garment vision research respectively.
Further reading
- Simulation Design for Wear Comfort of Garment Fabric Texture — ScienceDirect
- arXiv cs.GR — recent preprints in computer graphics and cloth simulation
- NVIDIA Research — publications and open-source releases
