Case 03 · Taelor · AI product · Mar to Sep 2026
Teaching an AI stylist what "good" means
A "pairs with" recommender shipped to the stylist team in March. Within a week, almost half of its suggestions were rejected and few stylists used it. Six months later approval was 81%, six or more stylists used it every month, and shipments left in about ten minutes.
Artifact and decisions
one issue · 5 linked fixes · numbers match the callouts
The issue
The recommender did not match how stylists actually pair clothes, so they ignored it
45% of the first 55 reviews were negative. Interviews said the colors and categories did not fit a stylist's taste; low usage came from slow loading and too little variety. The group below encodes the stylists' rules into the tool, gives them control over it, and measures whether it is earning its place.
Items and sizes are illustrative. Rules, controls, feedback loop and the three measures are the real ones.
Quality as explicit good/bad reviews, adoption as stylists and selections per month, speed as first pick to ship with a 30-minute cut. Written down, then measured monthly.
Sub-category, sleeve-length, color and print matrices derived from customer-accepted outfits, each pairing marked Strong, Conditional or Avoid; business ranking by reviews and sell-through.
Stylists choose color and print for the suggestion set, see why each item was suggested, and rate it in one tap. Feedback volume grew 3.3× in three months and fed the next rules.
Suggestions surfaced where the stylist is already working, with sizes auto-selected, so the size panel rarely needs touching. Deduplication, diversity and fallback rules keep the set useful when inventory is thin.
One deck a month for the executive team: quality, adoption, speed, and a versioned roadmap that marks completions, delays and holds as they happen.
The new styling tool had already cut time per shipment by 29% against the previous year. PairsWith was meant to take the next step, and its first week said otherwise: 55 reviews, 45% negative. I interviewed the stylists rather than reading the model logs. Two answers came back: the colors and categories did not fit how they pair, and the panel was slow and repetitive, so they stopped opening it.
Add-to-cart tracking for recommended items was still in development, so I made the explicit good/bad review the quality signal and wired it into a dashboard by stylist and month. For speed, I wrote the event definition: start at the first item pick per draft order, end at the next style preview or shipment end, gaps over 30 minutes excluded as inactive time.
Three questions the tool had to answer every month: is it good, is it used, does it make the day faster. Three hypotheses about why it was failing: H1 quality is a rules problem, the model lacks the stylists' pairing logic; H2 usage is a friction problem, speed and variety; H3 shipment speed is driven by pick cadence, so better suggestions should show up as faster picks.
Rules before reach. A recommender nobody trusts does not get more useful by being shown more often, so the choosing work (pairing rules, stylist controls) went ahead of the UI work (recommendations in the cart), and both went ahead of the larger bets: persona-based recommendations, styling notes, whole-shipment generation. Each bet was sequenced behind the evidence it needed, and put on hold when its constraint appeared: design capacity, engineering capacity, an expiring model credit.
The roadmap was kept as a versioned table: area, feature, purpose, engineering-complete date, testing-complete date, release date. Every month the dates were updated in place, so a slipped item read as slipped and a held item read as held, never as done.
Derived the color, print, pant and sub-category guardrails from customer-accepted outfits and packaged them as a rules file the recommender reads, with Strong, Conditional and Avoid states. Wrote the recommendation spec engineering built against: source markers on every suggestion, category switching, deduplication, diversity rules, fallback results when inventory is thin, and acceptance criteria for each. Built the click-through prototypes the team tested before a sprint was committed.
Approval rose with each rule release and held through August. Feedback volume grew 3.3× from March to May, so the signal got stronger as the number improved.
A recommender earns adoption by matching the expert's rules, not by overriding them. Once the stylists trusted the pairs, the roadmap could move from rule-based assistance toward a human-reviewed shipment copilot: styling notes generated from assigned pairings, customer summaries between shipments, and whole-shipment drafts for a stylist to approve.
Define the metric before you argue about it. The shipment-time definition survived six months of reviews because it was written down first; every later debate was about the work, not the number.
What I would not claim
The approval gain is a team result across engineering, data and the stylists; I defined the measures, designed the rules and ran the reviews. The speed comparison uses the three stylists with an April baseline, so it is indicative rather than company-wide. Items marked on hold are targets, not results.
Rule this produced
Define the metric before you argue about it. Written down first, agreed, then measured.