How I Used AI to Classify Over 30,000 Products (Without Writing a Single Line of Code)

I designed an AI pipeline to clean and classify 30,000 e-commerce products using a public dataset, without writing the code myself. Here’s the design logic.

4 min read

If you're dealing with messy catalog data or thinking about where AI actually belongs in your ops workflows — this is the kind of problem I work on. Feel free to reach out.

Get in touch

Dataset: ASOS E-Commerce Dataset — 30,845 products by TrainingDataPro, via Kaggle.

The mechanical cleaning stage processed the dataset in seconds in the test environment. Classifying a 200-product sample with AI took under 10 minutes.

I did not review 30,000 products manually or write a custom script from scratch. Instead, I designed a pipeline that separated mechanical transformations from tasks requiring semantic judgment. AI-generated code handled the mechanical cleaning, while an AI model handled classification and enrichment. I did not write the code myself.

Here’s how the pipeline worked.

The dataset was a mess

I used a public ASOS product catalog — 30,845 rows, 9 columns — sourced from Kaggle. It looked clean at first glance. It wasn't.

  • 18 fully null rows

  • 873 duplicate SKUs

  • A category field that contained product names, not categories

  • Color values in three different cases — "Black", "BLACK", "black" — treated as separate values

  • Size fields that mixed size labels with stock status: "UK 14 - Out of stock"

  • Product descriptions stored as a stringified blob that needed parsing before anything could be extracted from them

This public example mirrors product attribute enrichment work I’ve done professionally. That work involved attributes such as subcategory, franchise, activity, division, color, and material. I’m using an open dataset here so I can explain the approach without exposing proprietary product data, taxonomy details, or internal implementation materials. The public dataset provides a way to explain the approach using open data, without exposing proprietary product information, taxonomy details, or internal implementation materials.

The most important decision: what AI should and shouldn't do

Most AI demos use AI for everything. That's the wrong move.

Normalizing "BLACK" to "Black" doesn't need a language model. Neither does deduplicating SKUs, parsing a price string, or splitting a size field. These are mechanical transformations. Code handles them faster and more predictably, without asking a language model to make decisions it does not need to make.

What code can't do is read this:

"Coats & Jackets by New Look. Low-key layering. Notch collar. Button placket. Tie waist. Regular fit."

...and reliably decide: is this Apparel > Jackets & Coats? Or Apparel > Suits? Or something else entirely?

That requires semantic understanding. That's where AI belongs. So I split the pipeline into two stages.

Stage 1 — Pure code. Drop nulls. Deduplicate. Normalize color casing. Parse prices (including "Now £22.50" sale formats). Split size strings into available sizes and a stock flag. Parse the description blob into five structured fields. No LLM API calls were needed for this stage. It ran locally and processed the 30,000 rows in seconds in the test environment.

Stage 2 — AI enrichment. Extract material and fit type from the description. Assign category and subcategory from a fixed taxonomy. Return a confidence score for every classification. The AI does the semantic work. The code validates the output — if the AI returns a category that isn't on the approved list, it gets flagged no matter how confident the model was.

The AI doesn't get the final say. The validation layer does.

Why I defined the taxonomy before anything else

This step matters more than the code.

Without a fixed taxonomy, the model might return “Casual Tops,” “Casual Top,” and “Top (Casual)” as three separate values for the same product type. These outputs may be understandable, but they are still unusable if they do not match the approved taxonomy. The model needs clear constraints before it can produce consistent catalog data.

I defined a fixed set of top-level categories and subcategories before the AI touched a single product.

One edge case that required real thought: the difference between Sports Lifestyle and Sports Performance. An adidas Originals track jacket is sports-inspired but built for everyday wear. A proper running top is functional gear. They look similar. They belong in different subcategories. The AI can often distinguish between them, but only when the taxonomy provides clear definitions and the prompt gives it the right context.

That's a human decision. It has to happen before the AI gets involved.

The confidence routing layer

Here's the part most pipelines skip.

For this demonstration, each classification returned a confidence score on a 0-to-1 scale. The pipeline routes each product into one of three buckets based on that score. I used three confidence bands:

  • Auto-approved: ≥ 0.80, provided the output matched the approved taxonomy

  • Human review: 0.50–0.79

  • Rejected: < 0.50

The human review queue is one of the most useful outputs in the pipeline.

It is not a failure state. It is the system being honest about where it is uncertain. A human reviewer can focus on the genuinely ambiguous cases instead of manually reviewing all 30,000 rows.

Why the design matters

The same design principles apply when enriching product data in a business setting. The challenge is not simply asking an AI model to fill in missing attributes. It is defining the allowed values, separating clear cases from ambiguous ones, validating the output, and routing uncertain classifications for review.

That is why the taxonomy, confidence thresholds, and validation layer matter. They turn an open-ended AI task into a controlled data workflow. Before deploying AI into an operational process, you need to decide what happens when the model is wrong.

If you haven’t answered that question, you haven’t finished designing the system.

What's next

A natural next step would be a lightweight web tool for spot-checking individual products before they go live. A catalog manager could paste a product name and description and receive a suggested category, subcategory, material, fit type, and confidence score without preparing a CSV or configuring a pipeline.

If I build it, I’ll share it here.

Jason Kester Hanani

Operations & Product Professional

contact@jasonkhanani.com