An e-commerce marketplace needs to classify every new product listing into a taxonomy of 12,000 leaf categories arranged in a 5-level hierarchy. Category assignment drives search filters, commission rates, tax treatment and compliance checks, so errors have real financial consequences.
Sellers currently pick their own category and get it wrong roughly 20% of the time: sometimes accidentally, sometimes deliberately to reach a lower commission tier.
You get 800,000 new listings per day. Each has a title, a description, 1-8 images, seller-provided attributes, and a price. You have 40 million historically labelled listings, but the labels come from the same unreliable seller-selected process, plus about 2 million listings audited by an internal operations team.
The taxonomy changes: roughly 200 categories are added, merged or split each quarter.
Build the architecture on a canvas: place the components, configure them, connect them into a data flow, and write a short reason for each one. The AI reviewer grades your design against a rubric written specifically for this problem.
You have 40M noisy labels and 2M clean ones. Exactly how do you use both?
Why is a flat 12,000-way softmax a bad idea here, and what do you do instead?
You can review 20,000 of 800,000 listings a day. Which 20,000, and why not simply the lowest-confidence ones?
Minimum 6 components · needs a wide desktop screen