Case 03 · Taelor · AI product · Mar to Sep 2026

Teaching an AI stylist what "good" means

A "pairs with" recommender shipped to the stylist team in March. Within a week, almost half of its suggestions were rejected and few stylists used it. Six months later approval was 81%, six or more stylists used it every month, and shipments left in about ten minutes.

Problem45% of the first pairing reviews were negative
My contributionMeasures, pairing rules, UX specification and operating review
OutcomeApproval 52% → 81% · adoption broadened

Artifact and decisions

one issue · 5 linked fixes · numbers match the callouts

The issue

The recommender did not match how stylists actually pair clothes, so they ignored it

45% of the first 55 reviews were negative. Interviews said the colors and categories did not fit a stylist's taste; low usage came from slow loading and too little variety. The group below encodes the stylists' rules into the tool, gives them control over it, and measures whether it is earning its place.

Production-workflow recreationReady for a redacted product screenshot
Stylist tool · "pairs with" panel, aftersimplified recreation in English
Picked
Navy short-sleeve poloSize auto-selected · chest 40, slim
2Pairs with · 3 of 83Color: stonePrint: solid
Stone chinoSTRONG · sub-category rule
Olive shortsSTRONG · sleeve rule, color rule
Grey jeansCONDITIONAL · stylist pattern
Ranked by review and sell-through · avoids removedGoodBad
4Cart · 4 itemsRecommendations now appear here, next to what is already pickedfirst pick → ship: 9 min
1Monthly review · quality52% → 81%approval, May to Jul
Adoption6+ stylistsevery month, no single power user
5Speed9.6 minmedian first pick to ship

Items and sizes are illustrative. Rules, controls, feedback loop and the three measures are the real ones.

1Three measures, defined first
Targets · arguing about whether the tool helps

Quality as explicit good/bad reviews, adoption as stylists and selections per month, speed as first pick to ship with a 30-minute cut. Written down, then measured monthly.

2Pairing rules from real outfits
Targets · colors and categories that did not fit

Sub-category, sleeve-length, color and print matrices derived from customer-accepted outfits, each pairing marked Strong, Conditional or Avoid; business ranking by reviews and sell-through.

3Stylist controls and feedback
Targets · "it does not know my taste"

Stylists choose color and print for the suggestion set, see why each item was suggested, and rate it in one tap. Feedback volume grew 3.3× in three months and fed the next rules.

4Recommendations in the cart
Targets · low usage from clicks and scrolling

Suggestions surfaced where the stylist is already working, with sizes auto-selected, so the size panel rarely needs touching. Deduplication, diversity and fallback rules keep the set useful when inventory is thin.

5Monthly operating review
Targets · targets presented as results

One deck a month for the executive team: quality, adoption, speed, and a versioned roadmap that marks completions, delays and holds as they happen.

01 Discover45% negative in week one
02 DefineQuality, adoption, speed
03 PrioritizeRules before reach
04 BuildPairsWith v3.0 to v3.4
05 Measure+29 points, 6+ stylists, 9.6 min
06 LearnToward a shipment copilot
01Discover
PM strategy

The new styling tool had already cut time per shipment by 29% against the previous year. PairsWith was meant to take the next step, and its first week said otherwise: 55 reviews, 45% negative. I interviewed the stylists rather than reading the model logs. Two answers came back: the colors and categories did not fit how they pair, and the panel was slow and repetitive, so they stopped opening it.

Eng execution

Add-to-cart tracking for recommended items was still in development, so I made the explicit good/bad review the quality signal and wired it into a dashboard by stylist and month. For speed, I wrote the event definition: start at the first item pick per draft order, end at the next style preview or shipment end, gaps over 30 minutes excluded as inactive time.

start = first product.selection per draft_order end = next preview.start, else shipment.end cut = gaps > 30 min excluded
02Define
PM strategy

Three questions the tool had to answer every month: is it good, is it used, does it make the day faster. Three hypotheses about why it was failing: H1 quality is a rules problem, the model lacks the stylists' pairing logic; H2 usage is a friction problem, speed and variety; H3 shipment speed is driven by pick cadence, so better suggestions should show up as faster picks.

Evidence by September
H1 rules problem → yes: approval rose with each rule release H2 friction problem → yes: usage jumped once recs sat in the cart H3 pick cadence → yes: fast stylists differ on cadence, not review time
03Prioritize
PM strategy

Rules before reach. A recommender nobody trusts does not get more useful by being shown more often, so the choosing work (pairing rules, stylist controls) went ahead of the UI work (recommendations in the cart), and both went ahead of the larger bets: persona-based recommendations, styling notes, whole-shipment generation. Each bet was sequenced behind the evidence it needed, and put on hold when its constraint appeared: design capacity, engineering capacity, an expiring model credit.

Eng execution

The roadmap was kept as a versioned table: area, feature, purpose, engineering-complete date, testing-complete date, release date. Every month the dates were updated in place, so a slipped item read as slipped and a held item read as held, never as done.

04Build
PM strategy · the release train
v3.0 · AprColor and category matching tuned to each stylist's patternshipped v3.1 · MayStylists choose color and print for the suggestion setshipped v3.2 · MaySub-category pairing rules for varietyshipped v3.3 · JunLong and short sleeve pairing rulesshipped v3.4 · JunRecommendations shown directly in the cartshipped v4.0Persona and lookbook based recommendationson hold Notes · Apr–AugAI styling note V1.0 to V2.0: draft, prompt edits, assigned pairings, lookbookshipped Assistant · Aug–SepCustomer SMS and styling-event summaries between shipmentsshipped
Eng execution

Derived the color, print, pant and sub-category guardrails from customer-accepted outfits and packaged them as a rules file the recommender reads, with Strong, Conditional and Avoid states. Wrote the recommendation spec engineering built against: source markers on every suggestion, category switching, deduplication, diversity rules, fallback results when inventory is thin, and acceptance criteria for each. Built the click-through prototypes the team tested before a sprint was committed.

color_pairing_rules.json · {top: navy, bottom: stone} → STRONG subcategory_matrix · polo × chino → STRONG · polo × dress pant → CONDITIONAL
05Measure
Quality · share of pairings rated good
32%
52%
81%
MarMayJul 2026 · 253 reviews across the period

Approval rose with each rule release and held through August. Feedback volume grew 3.3× from March to May, so the signal got stronger as the number improved.

Adoption and speed
Stylists using AI pairing each month6+Selections in June, a full-month high58Concentrationno single power userMedian first pick to shipment9.6 minShipped within 30 min · within 60 min80% · 86%Stylists with an April baseline, faster by September41% to 45%Time per shipment vs the year before the new tool−29%
06Learn
PM strategy

A recommender earns adoption by matching the expert's rules, not by overriding them. Once the stylists trusted the pairs, the roadmap could move from rule-based assistance toward a human-reviewed shipment copilot: styling notes generated from assigned pairings, customer summaries between shipments, and whole-shipment drafts for a stylist to approve.

Eng execution

Define the metric before you argue about it. The shipment-time definition survived six months of reviews because it was written down first; every later debate was about the work, not the number.

What I would not claim

The approval gain is a team result across engineering, data and the stylists; I defined the measures, designed the rules and ran the reviews. The speed comparison uses the three stylists with an April baseline, so it is indicative rather than company-wide. Items marked on hold are targets, not results.

Rule this produced

Define the metric before you argue about it. Written down first, agreed, then measured.