Most fashion AI pilots do not fail because the model is wrong—they fail because no one agreed, before deployment, on what "working" would look like. If your success metric is an aggregate average across all SKUs and all seasons, a strong result in core basics can mask complete failure in the product categories the pilot was actually meant to improve. This guide walks programme managers and AI leads through a structured process for defining KPIs that are specific, scoped, and actionable before a single model goes live.
Key takeaways
- Scope your KPI to a named operation, a defined product category, and a fixed time window—not to the business overall.
- RAND research drawing on practitioner interviews found that in failed AI projects, the metrics measuring model success routinely diverge from the metrics that actually matter to the business.
- NIST's AI Risk Management Framework requires quantitative or qualitative measurement before deployment and continuously during operation—not only at the end of a pilot.
- A KPI without a baseline is a target without a reference point; collect baseline data before the pilot begins.
- Model cards—structured documentation of performance across subgroups—are a practical tool for exposing the disaggregated performance that aggregate averages conceal.
What you need before you start
- A defined pilot scope: which operation (demand forecasting, size recommendation, trend signal extraction, automated tech-pack generation, etc.), which product categories, and which markets or channels are in scope.
- Access to at least one full prior season of operational data for the in-scope categories, to establish a baseline.
- A named business lead and a named technical lead who will both sign off on the KPI definitions.
- A data governance check confirming that the data used to train and evaluate the model is permissible under your organisation's data protection obligations and, where relevant, the EU AI Act's requirements for the risk tier your system falls into.
- A version of the NIST AI Risk Management Framework MEASURE function bookmarked—it is the reference standard for what structured AI measurement looks like in practice.
Step 1 — Name the operation and resist the temptation to measure everything
Write one sentence that completes this template: "This pilot measures whether AI capability improves specific operational outcome for product category in channel or market over time window."
If you cannot complete that sentence without using the word "overall" or "general", the scope is too broad. A pilot that claims to improve "forecasting accuracy across the range" will produce an average that is technically true and operationally useless. Narrow to, for example, woven tops in the spring/summer transitional window, or size recommendation for extended sizing in the EU direct-to-consumer channel.
Expected result: A one-sentence scope statement agreed and signed by both the business lead and the technical lead.
Step 2 — Distinguish the model metric from the business metric
This is the single most common source of pilot failure. RAND research based on practitioner interviews found that in failed projects, business leaders and technical teams diverge on exactly this point: a business leader asks for a model that sets the right price, but the model is optimised for sales volume rather than profit margin—and nobody catches the mismatch until the pilot is over.
For each candidate KPI, write two columns:
| Model metric | Business metric it is supposed to serve |
|---|---|
| Mean Absolute Percentage Error (MAPE) on demand forecast | Reduction in end-of-season markdown rate for in-scope SKUs |
| Size recommendation acceptance rate | Return rate attributed to fit, for in-scope categories |
| Tech-pack generation time | Hours of technical design team time recovered per collection |
If you cannot draw a direct line from the model metric to a business metric that a commercial director would recognise, either reframe the model metric or reconsider whether this pilot addresses a real operational problem.
Expected result: A two-column table, reviewed by both leads, with no model metric left unmapped to a business outcome.
Note: The McKinsey State of Fashion annual report tracks AI adoption as a strategic priority across the industry. The consistent finding is that adoption is accelerating while measured return on investment remains uneven—precisely because pilots are evaluated on capability rather than on scoped business outcomes.
Step 3 — Set a baseline before the pilot begins
A KPI without a baseline is a wish, not a measurement. Before the model goes live, record the current-state value of every business metric in your two-column table, for the exact scope defined in Step 1.
Baseline collection rules:
- Use the same data extract, the same date range logic, and the same category filter you will use to measure the pilot outcome. Inconsistency in extraction method is a common source of spurious improvement.
- Record the baseline value, the extraction date, the data source, and the person responsible for the extract. Store this in a shared document that both leads can access.
- If historical data for the in-scope category is sparse—fewer than two full seasons—flag this as a measurement risk before the pilot begins, not after.
- Where your organisation uses a PLM system such as PTC FlexPLM to manage tech packs, bills of materials, and supplier timelines, that system is often the most reliable source of pre-pilot cycle-time data. Export and version the relevant records before any AI tooling touches the workflow.
Expected result: A baseline data file, dated and version-controlled, covering every business metric in scope.
Step 4 — Apply NIST's MEASURE function to structure your measurement plan
The NIST AI Risk Management Framework's MEASURE function requires that AI systems be tested before deployment and monitored continuously during operation—not evaluated only at the end of a fixed pilot window. Structure your measurement plan around three checkpoints:
- Pre-deployment: Validate model performance on a held-out test set drawn from the in-scope category and time period. Document accuracy, confidence intervals, and any known performance gaps by subgroup (size range, price tier, geography).
- In-pilot (weekly or bi-weekly): Track both the model metric and the business metric on a rolling basis. Set a threshold—agreed in advance—at which you will pause the pilot if the model metric degrades below an acceptable floor.
- End-of-pilot: Compare business metric outcomes against the baseline established in Step 3. Disaggregate by subgroup before reporting an aggregate figure.
The MEASURE function also requires documenting aspects of the system's trustworthiness. In a fashion context, this includes: data provenance (where did the training data come from, and is it representative of the in-scope categories?), known failure modes (does the model perform worse on new-season introductions than on replenishment lines?), and human oversight arrangements (who reviews model outputs before they affect a commercial decision?).
Expected result: A written measurement plan covering all three checkpoints, stored alongside the baseline file.
Step 5 — Disaggregate before you aggregate
An aggregate KPI that shows improvement may conceal that the model performs well on high-volume basics and poorly on the product categories that were the actual motivation for the pilot. The model cards framework proposed by Mitchell et al. in their paper on model reporting addresses this directly: it recommends documenting benchmarked evaluation across different conditions—demographic groups, product subgroups, or operational contexts—so that aggregate performance does not obscure subgroup failure.
Apply this logic to your pilot:
- Break your primary KPI by price tier (entry, mid, premium within scope).
- Break it by newness (new introductions vs. replenishment vs. carryover).
- Break it by geography or channel if the pilot spans more than one.
- Report each disaggregated result alongside the aggregate. If the aggregate improves but a critical subgroup worsens, that is a finding, not a footnote.
Expected result: A disaggregated results table, produced at the end-of-pilot checkpoint, that your steering committee reviews before any scale-up decision.
Step 6 — Define the decision rule before the pilot ends
A pilot without a pre-agreed decision rule produces a negotiation, not a decision. Before the pilot begins, document:
- The minimum improvement in the business metric required to recommend scale-up (e.g., a defined reduction in markdown rate, or a defined reduction in fit-related returns—expressed as a threshold, not as "meaningful improvement").
- The conditions under which the pilot will be extended rather than concluded (e.g., if the baseline data is found to be unreliable mid-pilot).
- The conditions under which the pilot will be stopped early (e.g., if the model metric falls below the floor set in Step 4).
- Who has authority to make each of these decisions.
This document should be reviewed by both leads and by any steering committee member who will receive the final report. It removes the incentive—common in practice—to reframe the evaluation criteria once results are visible.
Expected result: A one-page decision framework, signed before the pilot begins, referenced in the final report.
Troubleshooting common problems
The business lead and technical lead cannot agree on the business metric. This is a signal that the pilot is solving a technical problem that has not yet been translated into a commercial one. Pause scope definition and run a structured workshop to identify which operational decision the model output is intended to inform, and who currently makes that decision. The KPI follows from the decision, not from the model.
Baseline data does not exist for the in-scope category. Do not substitute a proxy from a different category or a different season. Either extend the pre-pilot period to collect a baseline, or narrow the scope to a category where baseline data exists. A pilot evaluated against a proxy baseline will produce findings that cannot be trusted.
The model performs well in testing but the business metric does not improve in-pilot. This is the divergence RAND's research identifies as a root cause of AI project failure. Return to the two-column table from Step 2 and examine whether the model metric genuinely drives the business metric, or whether other variables (promotional activity, supplier delays, channel mix shifts) are absorbing the effect. Document the confounders and adjust the evaluation accordingly.
Stakeholders want to report the aggregate result and move to scale-up. Require the disaggregated table from Step 5 to be presented alongside the aggregate before any scale-up decision. If a subgroup failure is present, it will surface at scale—better to find it now.
The pilot window is too short to observe the business metric. This is common in fashion, where the business outcome (markdown rate, return rate) is only observable at the end of a selling window. Either align the pilot window with the selling season, or use a leading indicator (e.g., sell-through rate at week four) as a proxy—but document it as a proxy and plan a follow-up measurement at season end.
What success looks like
A well-defined fashion AI pilot KPI framework produces three artefacts: a signed scope statement, a baseline data file, and a decision framework. At the end of the pilot, it produces a disaggregated results table and a recommendation that references the pre-agreed decision rule. The recommendation is credible because the criteria were set before anyone saw the results.
The Gartner research and advisory practice, which tracks enterprise AI adoption across sectors, consistently identifies measurement maturity—the ability to connect model outputs to business outcomes—as a differentiator between organisations that scale AI successfully and those that cycle through pilots without accumulating institutional capability. The fashion sector is not exempt from this pattern, as reporting from Just Style and others covering the industry's AI adoption trajectory makes clear.
FAQ
What is the right number of KPIs for a fashion AI pilot? One primary business KPI, supported by one or two model metrics that demonstrably drive it. More than three KPIs in a pilot typically signals that the scope is too broad or that stakeholders have not agreed on what the pilot is actually for.
How long should a fashion AI pilot run before evaluation? Long enough to observe the business metric in its natural cycle. For demand forecasting or markdown rate, that means at least one full selling window. For operational metrics like tech-pack generation time, four to six weeks of consistent use may be sufficient—but document the time window in advance.
What does NIST's MEASURE function require for a fashion AI pilot? Quantitative or qualitative measurement before deployment, continuously during operation, and at conclusion. It also requires documenting trustworthiness characteristics—data provenance, known failure modes, and human oversight arrangements—not only accuracy figures.
How do you handle a pilot where the model metric improves but the business metric does not? Treat it as a finding, not a failure to explain away. Document the confounders, examine whether the model metric genuinely drives the business outcome, and revise the KPI mapping before any scale-up decision. This is the most common and most instructive outcome of a well-measured pilot.
Should a fashion AI pilot KPI account for EU AI Act obligations? Yes, particularly if the system touches hiring, pricing, or consumer-facing recommendations at scale, which may attract higher-risk classifications. The measurement plan should include documentation of data quality, human oversight, and transparency measures that align with the risk tier assigned to the system under the Act.
Further reading
- RAND: The Root Causes of Failure for Artificial Intelligence Projects
- NIST AI RMF Core — MEASURE function
- Mitchell et al. — Model Cards for Model Reporting
