The filter stops grouping.
Faceted search matches strings. Three renderings of one attribute across five hundred products produce three facets where there should be one. The category page no longer narrows.
The catalogue refresh has to happen every cycle, without a person coordinating it, and without the output quietly breaking the store.
What actually breaks
An inconsistent translation in a blog post costs you nothing. In a product listing it changes what the shopper can find.
Faceted search matches strings. Three renderings of one attribute across five hundred products produce three facets where there should be one. The category page no longer narrows.
Shoppers type one of these words, not all three. Pages that used a different word do not come up lower down. They do not come up. You will not spot the gap, because a search you never appeared in leaves nothing behind in any report.
There is no empty result and no complaint. The shopper sees a shorter list of options and closes the tab. The purchase that did not happen leaves no event in any dashboard.
Someone shopping in a specialized category knows the vocabulary better than an AI model does. A buyer who knows the category reads three names for one part as a sign that the store does not stock it seriously. That judgment takes seconds, before the price is read.
Two ways to build your language data
Most teams treat terminology and style guides as preparation, something to sort out before the real work starts. There are two routes. They differ in when the quality arrives, not in whether it does.
Prepare it up front
A terminologist works through your categories, your existing content and your markets before the first batch runs. It takes time and budget. The first cycle then comes out as good as the fifth would otherwise be.
Accumulate it as you go
Each batch produces a glossary, somebody settles it, and it merges into your language data. You approve nothing and the first batch starts immediately. The first cycles are the weakest, and the corrections happen after the pages are live.
Either way you end up with the same language data. The choice is whether you pay for it in a budget line or in the quality of your first few refreshes.
How this one runs
This is one way to configure it. Some of the steps are choices you can make differently.
Product content goes in as segments carrying your own identifiers, so nothing on your side has to track ours. No file is exported and nothing is emailed.
Before anything is translated. Terms come from content in front of it, in context, rather than from a list written last year.
The glossary package goes to whoever knows that language and product. Once terms are agreed they merge into repository, making every later batch cheaper.
Automated translation with approved terminology enforced, right style guide applied, and memory of everything already confirmed in front of it.
Every segment is scored and result recorded. Routing failures into human revision is available; this catalogue puts human effort into sampling instead.
A share of output goes to professional reviewer whatever it scored. This keeps scoring honest rather than merely trusted.
Not one voice, several
A catalogue has both problems. The material is either called the right thing or it is not, that is terminology. How the description reads is not binary at all, and it is not the same answer everywhere.
| Per product line | You would not write a safety-critical range and a seasonal accessory range in the same register, even in the same store. |
| Per market | What reads as confident in one language reads as overclaiming in another. Some markets set rules on what a product description may assert. |
| Per domain | Spec sheets, category pages and marketing copy may come out of the same PIM. They should not come out sounding the same. |
You can hold as many style guides as you need and scope each one by domain, by language, or however your catalogue is divided. The platform applies the scoped guide while the translation is produced, rather than leaving the choice to whoever happens to be reviewing.
What the second cycle looks like
Approved terminology grows with every translation, and you do not pay twice for the same confirmed sentence. The second refresh is smaller than the first. The tenth is smaller than the second, because the language data in front of it is larger.
Working up front on terminology does not change where the cost ends up. It changes how steep the first few cycles are.
What sampling is for
Automatic evaluation decides what needs a person. A reviewer reading a sample shows whether the automatic evaluation can be believed.
A share of every batch goes to a professional reviewer, whatever the score was. When the reviewer marks a segment that the scoring passed, the threshold is too low and you raise it. You learn that from a colleague rather than from a customer. Without that step you are not measuring quality. You are measuring how confident a model is about itself.
Where else this shape fits
Nothing above depends on what this retailer sells. Three conditions produce the same problem in any category: more products than anyone can read, attributes that have to be exact, and each market using its own word for the same thing.
A part is found by its specification or not at all.
Fitment attributes that have to match across markets and marketplaces.
Tolerances, standards and ratings that survive no paraphrase.
The wrong word is a compliance question, not a wording one.
Dimensions, grades and certifications, in markets with different norms.
Materials and performance claims a customer filters on.
What these have in common is not the product. Each one has a buyer who can tell whether you know what you are selling. Each one has a catalogue that nobody in your company has read end to end in any language.
Does this shape match yours?
Start with one product category and one language you cannot check yourself. You see what comes back, including the terms the glossary extraction finds in content you have been publishing for years.
Questions
Practical answers for teams comparing catalogue translation workflows.
Not necessarily — terminology can be extracted from each batch and settled as the work goes. An existing termbase shortens both routes: TMX and TBX import.
In this configuration, an API call from the PIM. It could equally be a schedule, a webhook, or an automation platform — the process is the same whatever starts it.
In this setup, in two places: settling the glossary and reviewing the sample. Both are choices rather than requirements.
Translatable text usually concentrates in the parent record while variants multiply rows. What is in scope is decided by filters, set once per content shape rather than per run.
Yes. Whoever reviews and whoever translates is your choice, at your rates, in the same projects the API creates.
A model returns fluent output and forgets everything. The difference is not the translation step — it is everything around it: source comparison, terminology, review, and retained decisions.