Skip to main content
Back to Blog
Catalog 8 min read By Priya Raman, CTO and Co-Founder

Product Attribution at Scale: Tagging Thousands of New Items

Product Attribution at Scale: Tagging Thousands of New Items

When a new product arrives in a retailer's catalog system, it typically comes with whatever attribute data the supplier chose to include. That data is almost never complete enough for the storefront. The supplier knows their product's dimensions and part number. Your storefront needs to know the material composition, the intended use case, the care instructions, the size chart mapping, and the category placement in a taxonomy that the supplier has never seen. The gap between what arrives and what you need is the attribution problem.

At small catalog volumes, a catalog team member can fill that gap manually. They read the product description, look at the images, consult supplier documentation, and fill in the fields. This works well for a few hundred new items per week. It stops working when you are onboarding thousands of new SKUs every week, or when your catalog spans dozens of product categories each with their own attribute requirements.

The Attribution Gap Is Predictable Before You Look at the Data

For any product type where the required attributes are not part of the supplier's standard data model, you will have attribution gaps. A sporting goods retailer bringing in apparel from a new vendor is going to receive weight, SKU, dimensions, and price. The fit type, activity category, material technology, and size chart will require work on the retailer's side. This is not a failure of the supplier relationship. It is a structural property of how product data flows in commerce supply chains.

The gap categories that appear most often in large catalog operations are: activity or use case classification (what is this product for, beyond its physical description), material and construction attributes (beyond what the supplier specifies, because search and filter require standardized vocabulary), size and fit attributes when the supplier uses non-standard size labeling, and category placement within the retailer's own taxonomy rather than the supplier's.

Each of these gaps has a different character. Activity classification is inferential: you need to look at the product holistically and decide what buyer intent it serves. Material attributes require translating supplier terminology into your standardized vocabulary. Size attributes require normalization against a reference chart. Category placement requires understanding your taxonomy's rules well enough to apply them to a new product type.

Why Manual Attribution Bottlenecks at Scale

The throughput limit on manual attribution is not the speed of individual catalogers. It is the context-switching cost of working across many product types, and the inconsistency that accumulates when many people make independent judgment calls on the same attribute fields.

Consider a catalog team of five people each handling incoming product attribution. Each person develops their own interpretation of ambiguous attribute values. One cataloger consistently tags products as "outdoor" where another tags the same product type as "all-terrain." One person uses "polypropylene" where another uses "synthetic." These are not mistakes. They are judgment calls made without a shared reference. But they create attribute inconsistency that degrades filter performance and makes search results less predictable.

The inconsistency compounds over time. A product tagged six months ago using the informal vocabulary from that period may not match the normalized vocabulary your search team introduced last quarter. The backlog of inconsistently tagged products grows faster than any manual cleanup effort can clear it.

What Automated Attribution Actually Does and Does Not Do Well

Automated attribute tagging using machine learning models works best on classification problems where there is a well-defined set of output values and sufficient training signal. Category placement from product title and description is a strong case. Material classification when supplier data includes some structured attributes is a strong case. Activity classification for established product types in a stable taxonomy is a strong case.

It works less well on novel product types the model has not seen before, on attributes that require reading physical samples or interpreting ambiguous supplier documentation, and on edge cases where the correct answer depends on nuances that do not appear in text data. The model that confidently classifies 90% of your incoming items correctly will produce confident wrong answers on some of the remaining 10%, and it usually cannot tell you which ones those are.

This is the boundary that matters most in practice. Automated attribution is not a replacement for human judgment. It is a way to handle the high-confidence, high-volume cases efficiently so that the human attention available for attribution goes to the cases that actually require it. A model that handles 85% of incoming attributions at acceptable quality, and routes the remaining 15% to human review with a clear indication of why the model was uncertain, is more valuable than a model that tries to handle 100% and gets the edge cases wrong without flagging them.

Confidence Thresholds and the Review Queue

The practical design of an automated attribution system requires a decision about what to do with low-confidence predictions. The two common approaches are: apply the prediction with a low-confidence flag for later audit, or route the item to a human review queue before it reaches the storefront.

Which approach is right depends on the cost of the error. For category placement, a wrong category assignment is immediately visible on the storefront and affects discoverability. For a non-critical attribute like secondary color descriptor, a wrong value is less immediately harmful. Your confidence threshold policy should be calibrated to the cost of the error type, not set uniformly across all attributes.

A well-designed system surfaces the low-confidence items to reviewers with the model's suggested value and the evidence the model used. The reviewer sees "Model suggests: Outdoor Apparel (71% confidence), based on title terms: waterproof, shell, wind-resistant" and can confirm or correct in a few seconds. This is faster than starting from a blank field, and it helps reviewers develop intuition for where the model tends to be systematically wrong.

Vocabulary Management as Infrastructure

One underinvested component of attribution at scale is the attribute vocabulary management system. For automated attribution to produce consistent values, the set of valid values for each attribute needs to be defined, maintained, and communicated to both the automation and the human reviewers.

In practice this is messier than it sounds. Attribute vocabularies expand as new product types arrive. A value that was not in the vocabulary for outdoor footwear six months ago becomes necessary when a new product category is added. The vocabulary management system needs a process for proposing new values, reviewing whether the new value is actually distinct from existing values or can be mapped to them, and rolling out approved new values to both the automation training data and the human reviewer reference guides.

Retailers who treat vocabulary management as a continuous process, rather than a one-time setup, have significantly better attribution consistency over time. Those who establish a vocabulary at system launch and then let it drift usually find that the automation starts producing more errors as the catalog grows into territory the original vocabulary did not anticipate.

The Connection to Search Performance

The reason attribution quality matters is that search and filter performance is directly downstream of it. A shopper filtering for "merino wool" returns results only for products with merino wool in the correct attribute field. If some merino wool products are tagged as "natural fiber" or "wool blend" rather than "merino wool," they do not appear in the filtered results. The shopper sees incomplete results and may conclude the retailer does not carry what they want.

This connection between attribution quality and search completeness is worth measuring explicitly. A category-level audit that identifies what fraction of products in each category have the critical attributes populated, and what fraction of filter queries in each category return zero results despite relevant products existing, can show you where the attribution gaps are costing you the most in actual shopper behavior. That measurement makes attribution quality a business metric, not just a data quality metric, and makes it easier to prioritize where to invest in improving the attribution workflow.

We want to be clear that improving attribution quality is one input to search performance among several. Product relevance, review quantity, and behavioral signals all contribute. Attribution improvements affect what products are retrievable, but they do not determine whether retrieved products are competitive on other dimensions. The goal of good attribution is to make sure the right products show up in the right searches, not to guarantee that they will be purchased once they appear.

More from the blog

Get started

See SuperCommerce in action

Request a live demo and we'll show you the platform running against a real catalog.