What we doPricing
Login
Use case · Product content for e-commerce

Three correct words for the same product, and the store stops working.

It has to happen every cycle, without a person coordinating it, and without the output quietly breaking the store.

ONE ATTRIBUTE · DE-DE · 500 PRODUCTSRostfreier StahlEdelstahlNichtrostender Stahl180 SKU240 SKU80 SKUALL THREECORRECTWHAT THE SHOPPER SEESFILTER · MATERIALSEARCH · "EDELSTAHL"Rostfreier StahlEdelstahlNichtrostender Stahl18024080Tick one and you see 180 of 500.The shopper cannot know there aretwo more boxes for the same thing.Edelstahlcompetitor.demarketplace.deyourstore.deOnly 240 of your 500 can show uphere. The rest used another word.

What actually breaks

Three correct words for the same thing, and the store stops working.

A blog post survives an inconsistent translation. A product listing does not — and none of the damage shows up where anyone is looking for it.

Navigation

The filter stops grouping.

Faceted search matches strings. Three renderings of the same attribute across five hundred products means three facets where there should be one, and a category page that no longer narrows.

SEO

You do not rank for the word the market types.

Shoppers type one of these words, not all three. Pages that used a different one do not come up lower down — they do not come up. And you will not spot it, because a search you never appeared in leaves nothing behind in any report.

Revenue

The customer does not report it.

No empty result, no complaint, just a shorter list of options and a session that ends. The purchase that did not happen leaves no event in any dashboard.

Trust

A specialist buyer notices before you do.

Someone shopping in a specialized category knows the vocabulary better than an AI model does. A store that calls the same thing three names reads as a store that does not stock it seriously — and that judgment is made in seconds, before price is even considered.

Terminology in a catalogue is not a quality preference. It is load-bearing.

Two ways to build your language data

Prepare it up front, or accumulate it as you go.

Terminology and style guides are usually treated as preparation — something to sort out before the real work starts. In practice there are two routes, and the difference is when the quality arrives rather than whether it does.

Prepare it up front

A terminologist works through your categories, your existing content and your markets before the first batch runs. It takes time and budget, and it buys you quality from the first cycle rather than the fifth.

Accumulate it as you go

Each batch produces a glossary, somebody settles it, and it merges into the repository. Nothing to approve and no delay to the start — but the first cycles are the weakest, and you are correcting in public while the record fills.

Either way the language data ends up in the same place. The only question is whether you would rather spend it in a budget line or in the quality of your first few refreshes; and that is not a translation decision.

How this one runs

An API call, six steps, and two places where a person decides.

One configuration, running in production. Parts of it are choices you can make differently.

Extract glossaryA person settles itTranslateEvaluatePublishedSampled reviewTERMINOLOGY · STYLE GUIDE · MEMORYWHATEVER IT SCORED
01

The PIM calls the API

Product content goes in as segments with your own identifiers, so nothing on your side has to track ours. Nothing is exported to a file and nothing is emailed.

02

A glossary is extracted from the batch

Before anything is translated. Terms come from content in front of it, in context, rather than from a list written last year.

03

Somebody settles it

A person decides

The glossary package goes to whoever knows that language and product. Once terms are agreed they merge into repository, making every later batch cheaper.

04

Translation runs

Automated translation with approved terminology enforced, right style guide applied, and memory of everything already confirmed in front of it.

05

Quality is evaluated against rules

Every segment is scored and result recorded. Routing failures into human revision is available; this catalogue puts human effort into sampling instead.

06

A reviewer samples it regardless

A person decides

A share of output goes to professional reviewer whatever it scored. This keeps scoring honest rather than merely trusted.

Not one voice, several

Terminology decides what a thing is called. A style guide decides how the sentence around it reads.

Those are different problems and a catalogue has both. The material is either called the right thing or it is not, that is terminology. How the description reads is not binary at all, and it is not the same answer everywhere.

Per product lineA safety-critical range and a seasonal accessory range do not want the same register, even in the same store.
Per market What reads as confident in one language reads as overclaiming in another. Some markets have rules about what a product description may assert at all.
Per domain Spec sheets, category pages and marketing copy come out of the same PIM and should not come out sounding the same.

You can hold as many as you need and scope each one — by domain, by language, however your catalogue is actually divided. The right one is applied while the translation is produced, not chosen by whoever happens to be reviewing.

What the second cycle looks like

The catalogue does not get translated twice.

Approved terminology grows with every translation, and confirmed sentences are not paid for twice. So the second refresh is smaller than the first, and the tenth is smaller than the second because the record in front of it got bigger.

Preparing the language data up front does not change where that curve ends. It changes how steep the beginning is.

What sampling is for

You cannot read a million words a month, and you should not have to trust a score nobody ever checked.

Automatic evaluation decides what needs a person. Sampling decides whether the automatic evaluation can be believed.

A share of output goes to a professional reviewer whatever it scored. When a reviewer marks something the scoring passed, the threshold is set too low and you raise it; and you learn that from a colleague rather than from a customer. Without that step you are not measuring quality. You are measuring how confident a model is about itself.

Where else this shape fits

Any catalogue where the words are part of the product.

Nothing above depends on what this retailer sells. Three conditions produce the same problem in any category: more products than anyone can read, attributes that have to be exactly right, and each market using its own word for the same thing.

Industrial supplies and MRO

A part is found by its specification or not at all.

Auto parts

Fitment attributes that have to match across markets and marketplaces.

Electronics and components

Tolerances, standards and ratings that survive no paraphrase.

Medical and laboratory supplies

The wrong word is a compliance question, not a wording one.

Building materials

Dimensions, grades and certifications, in markets with different norms.

Marine, outdoor and technical apparel

Materials and performance claims a customer filters on.

What these have in common is not the product. It is a buyer who can tell whether you know what you are selling, reading a catalogue nobody in your company has read end to end in any language.

Does this shape match yours?

Bring one product category and one market.

Not the whole catalogue. One product category, one language you cannot check yourself, and a look at what comes back — including what the glossary extraction finds in content you have been publishing for years.

Questions

Before the first batch

Practical answers for teams comparing catalogue translation workflows.

Do we need a termbase before we start?

Not necessarily — terminology can be extracted from each batch and settled as the work goes. An existing termbase shortens both routes: TMX and TBX import.

What triggers a run?

In this configuration, an API call from the PIM. It could equally be a schedule, a webhook, or an automation platform — the process is the same whatever starts it.

Does a person have to be involved in every cycle?

In this setup, in two places: settling the glossary and reviewing the sample. Both are choices rather than requirements.

What happens to variants?

Translatable text usually concentrates in the parent record while variants multiply rows. What is in scope is decided by filters, set once per content shape rather than per run.

Can we keep our own translators?

Yes. Whoever reviews and whoever translates is your choice, at your rates, in the same projects the API creates.

How is this different from putting our catalogue through a model?

A model returns fluent output and forgets everything. The difference is not the translation step — it is everything around it: source comparison, terminology, review, and retained decisions.