Skip to main content
Back to Blog
Merchandising 8 min read By Jordan Mills, CEO and Co-Founder

Why Merchandising Rules Break Down at Enterprise Catalog Sizes

Why Merchandising Rules Break Down at Enterprise Catalog Sizes

Merchandising rule systems in commerce platforms are built for human authors. Someone writes a rule: "products in the footwear category with review score above 4.2 and in-stock at the primary distribution center should be promoted to the top of search results in the footwear category." The rule is logical, auditable, and easy to explain. At 10,000 SKUs, a merchandising team can write rules, test them, and watch their effects unfold in a catalog they have some chance of actually knowing.

At 500,000 SKUs, the same rule system produces behavior that nobody fully understands. Not because the rules are wrong. Because the rules interact with a catalog that has grown past the point where any person can comprehend the totality of what the rules are doing.

The Interaction Problem

The core failure mode in enterprise-scale merchandising rules is rule interaction. A merchandising team writing rules over years will produce a system where rules conflict, override each other, and produce outcomes that none of the individual rules intended.

Consider a common sequence. A rule was written two years ago to boost products with high inventory levels in search rankings, written during a period when overstocked items were a problem and the team wanted to move inventory faster. A more recent rule was written to boost products with strong recent purchase velocity. A third rule suppresses products with review scores below 3.5. These rules seem reasonable in isolation. But they interact: a product with high inventory, strong recent velocity, and a 3.4 review score gets boosted by two rules and suppressed by a third. The net result depends on the weighting system, which was set by someone who is no longer on the team and has not been reviewed since.

Multiply this across a rule library with 150 rules written over three years by twelve different people, some of whom had conflicting business priorities, and you have a system that produces search results that the current merchandising team cannot fully explain or predict. When a major sales event produces unexpected search result quality, the root cause investigation requires archaeology through the rule history to find which combinations of rules produced the observed outcome.

Why Failures Cluster Around Sales Events

Merchandising rule failures tend to surface most visibly around major sales events: holiday promotions, clearance events, new product launches. There are several reasons for this clustering.

First, sales events introduce new rule additions. The team adds promotional boost rules for the event. These new rules enter a system of 150 existing rules that nobody has fully mapped. The new rules may interact with existing rules in ways that produce unexpected ranking behavior in categories nobody thought to check.

Second, inventory dynamics change sharply during sales events. Rules written for normal inventory distribution behave differently when large portions of the catalog are on promotion or when high-velocity sales are depleting certain categories faster than the rules anticipated. An inventory-boost rule that works well during normal periods can produce strange results when a high-velocity item depletes rapidly during a flash sale.

Third, traffic concentrations during sales events amplify the impact of any rule interaction problem. A misconfiguration that produces marginally wrong results during normal traffic becomes very expensive when ten times the normal traffic volume is running through the same search queries.

The result is that the merchandising team discovers problems in the worst possible moment: during a high-traffic, time-sensitive event when the cost of wrong search rankings is at its highest and the capacity to investigate and fix rule interactions is at its lowest.

The Attribution Problem in Rule Systems

A secondary failure mode specific to large catalogs is that rules written against attribute values depend on those attribute values being consistently populated across the catalog. A rule that says "boost outdoor apparel with moisture-wicking attribute" only functions correctly for products where the moisture-wicking attribute is populated and standardized. In a large catalog where attribute tagging has been done by many people over years, the same physical property may be tagged as "moisture-wicking," "moisture management," "sweat-wicking," or left blank.

The rule author assumed consistent attribute vocabulary. The catalog does not have it. The rule applies to a subset of the products it was intended to apply to, with the excluded products determined not by any business logic but by historical inconsistencies in how their attributes were recorded. The merchandising team cannot easily see this discrepancy because the rule system reports rule execution as successful. The products that should have been caught by the rule simply were not caught.

This is one of the reasons that catalog attribute quality and merchandising rule effectiveness are tightly coupled. Investing in attribute consistency upstream is not just about making search filters work better. It is also about making merchandising rules do what their authors intended.

The Auditing Gap

The most common operational response to merchandising rule complexity at scale is to add more auditing: manual reviews of search result quality in high-traffic categories before major events, sample-based checks on whether promotional boosts are applying correctly. This is better than nothing, but it has structural limits.

Manual auditing of search result quality can only check the categories and query patterns that someone thought to check. It cannot cover a 500,000-SKU catalog systematically. The rule interaction problems that cause the most damage are often in categories or query patterns that were not on the audit checklist, because nobody anticipated the problem would occur there.

Automated coverage testing, where every rule in the system is tested against a representative sample of catalog products to verify it applies to the intended population and not to unintended products, is more reliable than manual audit. It is also more expensive to build and maintain. The rule system needs to expose enough introspection capability to allow automated testing, and the test suite needs to be updated whenever rules change.

The teams that manage this well treat rule coverage testing as a required step in the rule publication workflow. A new rule does not go live until it has been run against a catalog sample and the results reviewed by at least one person other than the author. Rules that interact with existing rules in unexpected ways are flagged for review before publication. This slows down rule development slightly, and it is the right tradeoff when the cost of rule interaction failures is high.

Scale Changes the Risk Profile of Rule Systems

We are not arguing that rule-based merchandising systems are wrong or that they should be replaced. Rules are the right tool for expressing business logic that needs to be auditable and explainable. The problem is not rules. It is that the operational model for managing rules often does not scale with the catalog.

At 10,000 SKUs, the risk of an unexpected rule interaction is low because the catalog is small enough that the effects of any rule can be checked reasonably quickly. At 500,000 SKUs, the same rule system needs governance processes, automated coverage testing, and impact estimation tooling to be operated safely. These do not exist by default. They need to be built alongside the rule system as it scales.

The investment in rule governance tooling is not glamorous. It does not ship a new feature or add a new capability. It makes the existing capability more reliable. That is a hard case to make in a product roadmap discussion until the week before a major sales event when a rule interaction problem is causing visible damage to search result quality. The retailers who invest in this infrastructure before the crisis are the ones who do not experience the crisis.

To be honest about the limits of this framing: better rule governance reduces the frequency and severity of rule interaction problems, but it does not eliminate them entirely. Catalogs at enterprise scale are complex enough that some unexpected interactions will always occur. The goal of good rule governance is to make those occurrences detectable quickly and recoverable efficiently, not to guarantee they never happen.

More from the blog

Get started

See SuperCommerce in action

Request a live demo and we'll show you the platform running against a real catalog.