• From €1,250/month
  • Automate supplier data onboarding
  • Test it with your own product feed
Arnout Schutte

By Arnout Schutte, Co-founder / Managing Director

Why AI projects stall on poor product data.

Everyone is talking about AI. Copilots, agents, generative search engines and automated product copy are now on almost every roadmap.

Yet the difference between a successful AI application and a pilot that keeps stalling often does not sit in the model. It sits in the data AI has to work on.

And for retailers, brands and wholesalers, that data often starts in the same place: with the supplier.

Anyone serious about working with AI therefore has to start earlier in the product data flow: before PIM.

Everyone is talking about AI.

AI is now high on the agenda. Organisations experiment with agents, copilots, automatic content production, intelligent search functions and other applications that are meant to reduce manual work.

The promise is attractive: work faster, publish faster and do more with the same team.

But many discussions start with the wrong question. People look at which model, which tool or which AI supplier is best, while the quality of the underlying product data is often far more decisive for the final result.

The bottleneck usually does not sit in AI.

Modern AI models can work well with text, imagery and structured data. But they can only deliver reliable results if the underlying product data is sufficiently complete, consistent and understandable.

An agent that compares products needs reliable attributes. A copilot that generates product copy has to be able to trust specifications. A semantic search function works better when categories and features are logically built up.

AI can compensate a lot. But it cannot structurally solve the fact that the same property is named, measured or delivered differently by three suppliers.

Product data determines how reliably AI can work.

AI output is only as good as the product data it is based on. In practice, AI applications often get stuck on the same kind of problems:

  • missing attributes;
  • deviating units, such as centimetres at one supplier and millimetres at another;
  • inconsistent categorisation;
  • different values for the same feature;
  • missing or contradictory supplier information.

Each of those problems means someone has to check, correct or complete the output. For a handful of products, that is manageable. For thousands of products, multiple languages and multiple channels, it becomes a structural burden.

The technology can work, while the operational business case still disappears through the manual work around it.

The problem often arises before PIM.

Many organisations try to solve product data problems in PIM in the end. But by the time product data arrives there, many inconsistencies have already arisen.

Suppliers deliver Excel, CSV, XML, API data, media files and sometimes PDFs. Structures differ. Column names differ. Units differ. Categories differ.

If those differences are not resolved earlier in the flow, they are simply carried along into PIM.

You then do centralise the product data, but also the inconsistencies inside it.

At Azerty, more than 50 supplier feeds arrive, each with its own column names, value types and units. At Wehkamp, the challenge involves more than 2.5 million SKUs across multiple brands. At volumes like that, the structure at the front determines how much manual work remains later.

Learn more: Import & Onboarding and Supplier Data Onboarding.

Why supplier data is the foundation.

Supplier data is often the first point where differences become visible:

  • column names differ;
  • units differ;
  • mandatory attributes are missing;
  • product relations are recorded differently;
  • categories do not align with each other;
  • media arrives separately from the rest of the product data.

If you translate these differences at the source into one consistent structure, everything that follows benefits.

Supplier data → mapping → validation → data model → PIM → AI → channels

Your PIM receives better input. Workflows have to absorb fewer exceptions. AI receives consistent attributes to work with. And sales channels receive product data that was already checked earlier in the process.

That is not an extra layer. It is the foundation the rest leans on.

What this means for your product data architecture.

If AI becomes part of your product data process, the architecture has to ensure that AI works on controlled data. That asks for roughly three things.

Structure before PIM

Supplier data has to be able to be mapped, validated and normalised before it enters the rest of the product data flow.

The goal is not to make all suppliers deliver the same file. The goal is that different sources are ultimately interpreted in the same way.

One reliable data model

Attributes, categories, product relations and values have to receive the same meaning within the organisation.

A good data model creates the common structure on which PIM, workflows, AI and channels can build. You can read more about this at Data Model & Structure.

AI within rules and workflows

AI can help with classification, enrichment, translation, mapping and quality control. But automation has to fit within the rules of the organisation.

The value does not sit in generating as much as possible automatically, but in controlled automation where it is clear what AI may decide, what has to be checked, and when human review is necessary.

That is also how ConnectingTheDots is built: not as a separate AI layer on top of PIM, but as a product data platform in which onboarding, structure, PIM, workflows and AI are part of the same product data flow.

Learn more: AI Product Data Automation and Workflows & AI.

AI before and after PIM.

AI plays two different roles in the product data flow.

AI on product data

This concerns AI that helps to make product data itself better:

  • recognising attributes;
  • normalising values;
  • classifying products;
  • matching sources;
  • signalling missing data;
  • recognising inconsistencies.

AI with product data

This concerns applications that use reliable product data as input:

  • generating product copy;
  • translating;
  • semantic search;
  • product comparisons;
  • recommendations;
  • AI agents that use product information.

The better AI helps to structure product data before PIM, the more reliably AI can work with it after PIM.

What this means for the years ahead.

The direction is clear. AI will take a larger role in search, product advice, content production and automated commerce.

At the same time, marketplaces, data pools, customers and regulation ask for richer, better structured and better traceable product data. That increases the importance of a product data architecture and clear Product Data Governance that does not have to be rebuilt for every new application.

Organisations that structure product data closer to the source now do not have to solve the same data problem again for every new AI application.

Conclusion: do not start with the model.

The question about AI is not only which model you choose. The more important question is what that model will be able to rely on.

If product data is structured, validated and translated into one usable data model before PIM, AI does not become a separate experiment on top of polluted data. It becomes the next step in a product data flow that already works.

The AI strategy of tomorrow therefore does not start with the model. It starts with the product data of today.

Good product data starts before your PIM.

Frequently asked questions about AI and product data.

Curious what AI can do with your supplier data?

Show how one of your supplier feeds currently arrives. We will show where mapping, validation and automation can remove manual work.

Test my product feed

View Supplier Data Onboarding

Would you rather get acquainted first? Plan a demo