The filter stops grouping.
Faceted search matches strings. Three renderings of the same attribute across five hundred products means three facets where there should be one, and a category page that no longer narrows.
It has to happen every cycle, without a person coordinating it, and without the output quietly breaking the store.
What actually breaks
A blog post survives an inconsistent translation. A product listing does not — and none of the damage shows up where anyone is looking for it.
Faceted search matches strings. Three renderings of the same attribute across five hundred products means three facets where there should be one, and a category page that no longer narrows.
Shoppers type one of these words, not all three. Pages that used a different one do not come up lower down — they do not come up. And you will not spot it, because a search you never appeared in leaves nothing behind in any report.
No empty result, no complaint, just a shorter list of options and a session that ends. The purchase that did not happen leaves no event in any dashboard.
Someone shopping in a specialized category knows the vocabulary better than an AI model does. A store that calls the same thing three names reads as a store that does not stock it seriously — and that judgment is made in seconds, before price is even considered.
Two ways to build your language data
Terminology and style guides are usually treated as preparation — something to sort out before the real work starts. In practice there are two routes, and the difference is when the quality arrives rather than whether it does.
Prepare it up front
A terminologist works through your categories, your existing content and your markets before the first batch runs. It takes time and budget, and it buys you quality from the first cycle rather than the fifth.
Accumulate it as you go
Each batch produces a glossary, somebody settles it, and it merges into the repository. Nothing to approve and no delay to the start — but the first cycles are the weakest, and you are correcting in public while the record fills.
Either way the language data ends up in the same place. The only question is whether you would rather spend it in a budget line or in the quality of your first few refreshes; and that is not a translation decision.
How this one runs
One configuration, running in production. Parts of it are choices you can make differently.
Product content goes in as segments with your own identifiers, so nothing on your side has to track ours. Nothing is exported to a file and nothing is emailed.
Before anything is translated. Terms come from content in front of it, in context, rather than from a list written last year.
The glossary package goes to whoever knows that language and product. Once terms are agreed they merge into repository, making every later batch cheaper.
Automated translation with approved terminology enforced, right style guide applied, and memory of everything already confirmed in front of it.
Every segment is scored and result recorded. Routing failures into human revision is available; this catalogue puts human effort into sampling instead.
A share of output goes to professional reviewer whatever it scored. This keeps scoring honest rather than merely trusted.
Not one voice, several
Those are different problems and a catalogue has both. The material is either called the right thing or it is not, that is terminology. How the description reads is not binary at all, and it is not the same answer everywhere.
| Per product line | A safety-critical range and a seasonal accessory range do not want the same register, even in the same store. |
| Per market | What reads as confident in one language reads as overclaiming in another. Some markets have rules about what a product description may assert at all. |
| Per domain | Spec sheets, category pages and marketing copy come out of the same PIM and should not come out sounding the same. |
You can hold as many as you need and scope each one — by domain, by language, however your catalogue is actually divided. The right one is applied while the translation is produced, not chosen by whoever happens to be reviewing.
What the second cycle looks like
Approved terminology grows with every translation, and confirmed sentences are not paid for twice. So the second refresh is smaller than the first, and the tenth is smaller than the second because the record in front of it got bigger.
Preparing the language data up front does not change where that curve ends. It changes how steep the beginning is.
What sampling is for
Automatic evaluation decides what needs a person. Sampling decides whether the automatic evaluation can be believed.
A share of output goes to a professional reviewer whatever it scored. When a reviewer marks something the scoring passed, the threshold is set too low and you raise it; and you learn that from a colleague rather than from a customer. Without that step you are not measuring quality. You are measuring how confident a model is about itself.
Where else this shape fits
Nothing above depends on what this retailer sells. Three conditions produce the same problem in any category: more products than anyone can read, attributes that have to be exactly right, and each market using its own word for the same thing.
A part is found by its specification or not at all.
Fitment attributes that have to match across markets and marketplaces.
Tolerances, standards and ratings that survive no paraphrase.
The wrong word is a compliance question, not a wording one.
Dimensions, grades and certifications, in markets with different norms.
Materials and performance claims a customer filters on.
What these have in common is not the product. It is a buyer who can tell whether you know what you are selling, reading a catalogue nobody in your company has read end to end in any language.
Does this shape match yours?
Not the whole catalogue. One product category, one language you cannot check yourself, and a look at what comes back — including what the glossary extraction finds in content you have been publishing for years.
Questions
Practical answers for teams comparing catalogue translation workflows.
Not necessarily — terminology can be extracted from each batch and settled as the work goes. An existing termbase shortens both routes: TMX and TBX import.
In this configuration, an API call from the PIM. It could equally be a schedule, a webhook, or an automation platform — the process is the same whatever starts it.
In this setup, in two places: settling the glossary and reviewing the sample. Both are choices rather than requirements.
Translatable text usually concentrates in the parent record while variants multiply rows. What is in scope is decided by filters, set once per content shape rather than per run.
Yes. Whoever reviews and whoever translates is your choice, at your rates, in the same projects the API creates.
A model returns fluent output and forgets everything. The difference is not the translation step — it is everything around it: source comparison, terminology, review, and retained decisions.