TextUnited
What we doPricing
Use case · Product content for e-commerce

Three correct words for the same product, and the store stops working.

The catalogue refresh has to happen every cycle, without a person coordinating it, and without the output quietly breaking the store.

ONE ATTRIBUTE · DE-DE · 500 PRODUCTSRostfreier StahlEdelstahlNichtrostender Stahl180 SKU240 SKU80 SKUALL THREECORRECTWHAT THE SHOPPER SEESFILTER · MATERIALSEARCH · "EDELSTAHL"Rostfreier StahlEdelstahlNichtrostender Stahl18024080Tick one and you see 180 of 500.The shopper cannot know there aretwo more boxes for the same thing.Edelstahlcompetitor.demarketplace.deyourstore.deOnly 240 of your 500 can show uphere. The rest used another word.

What actually breaks

Four things break, and none of them shows up in a report.

An inconsistent translation in a blog post costs you nothing. In a product listing it changes what the shopper can find.

Navigation

The filter stops grouping.

Faceted search matches strings. Three renderings of one attribute across five hundred products produce three facets where there should be one. The category page no longer narrows.

SEO

You do not rank for the word the market types.

Shoppers type one of these words, not all three. Pages that used a different word do not come up lower down. They do not come up. You will not spot the gap, because a search you never appeared in leaves nothing behind in any report.

Revenue

The customer does not report it.

There is no empty result and no complaint. The shopper sees a shorter list of options and closes the tab. The purchase that did not happen leaves no event in any dashboard.

Trust

A specialist buyer notices before you do.

Someone shopping in a specialized category knows the vocabulary better than an AI model does. A buyer who knows the category reads three names for one part as a sign that the store does not stock it seriously. That judgment takes seconds, before the price is read.

Terminology in a catalogue decides what the filter groups and what search returns.

Two ways to build your language data

Prepare it up front, or accumulate it as you go.

Most teams treat terminology and style guides as preparation, something to sort out before the real work starts. There are two routes. They differ in when the quality arrives, not in whether it does.

Prepare it up front

A terminologist works through your categories, your existing content and your markets before the first batch runs. It takes time and budget. The first cycle then comes out as good as the fifth would otherwise be.

Accumulate it as you go

Each batch produces a glossary, somebody settles it, and it merges into your language data. You approve nothing and the first batch starts immediately. The first cycles are the weakest, and the corrections happen after the pages are live.

Either way you end up with the same language data. The choice is whether you pay for it in a budget line or in the quality of your first few refreshes.

How this one runs

An API call, six steps, and two places where a person decides.

This is one way to configure it. Some of the steps are choices you can make differently.

Extract glossaryA person settles itTranslateEvaluatePublishedSampled reviewTERMINOLOGY · STYLE GUIDE · MEMORYWHATEVER IT SCORED
01

The PIM calls the API

Product content goes in as segments carrying your own identifiers, so nothing on your side has to track ours. No file is exported and nothing is emailed.

02

A glossary is extracted from the batch

Before anything is translated. Terms come from content in front of it, in context, rather than from a list written last year.

03

Somebody settles it

A person decides

The glossary package goes to whoever knows that language and product. Once terms are agreed they merge into repository, making every later batch cheaper.

04

Translation runs

Automated translation with approved terminology enforced, right style guide applied, and memory of everything already confirmed in front of it.

05

Quality is evaluated against rules

Every segment is scored and result recorded. Routing failures into human revision is available; this catalogue puts human effort into sampling instead.

06

A reviewer samples it regardless

A person decides

A share of output goes to professional reviewer whatever it scored. This keeps scoring honest rather than merely trusted.

Not one voice, several

Terminology decides what a thing is called. A style guide decides how the sentence around it reads.

A catalogue has both problems. The material is either called the right thing or it is not, that is terminology. How the description reads is not binary at all, and it is not the same answer everywhere.

Per product lineYou would not write a safety-critical range and a seasonal accessory range in the same register, even in the same store.
Per marketWhat reads as confident in one language reads as overclaiming in another. Some markets set rules on what a product description may assert.
Per domainSpec sheets, category pages and marketing copy may come out of the same PIM. They should not come out sounding the same.

You can hold as many style guides as you need and scope each one by domain, by language, or however your catalogue is divided. The platform applies the scoped guide while the translation is produced, rather than leaving the choice to whoever happens to be reviewing.

What the second cycle looks like

The catalogue does not get translated twice.

Approved terminology grows with every translation, and you do not pay twice for the same confirmed sentence. The second refresh is smaller than the first. The tenth is smaller than the second, because the language data in front of it is larger.

Working up front on terminology does not change where the cost ends up. It changes how steep the first few cycles are.

What sampling is for

You cannot read a million words a month, and you should not have to trust a score nobody ever checked.

Automatic evaluation decides what needs a person. A reviewer reading a sample shows whether the automatic evaluation can be believed.

A share of every batch goes to a professional reviewer, whatever the score was. When the reviewer marks a segment that the scoring passed, the threshold is too low and you raise it. You learn that from a colleague rather than from a customer. Without that step you are not measuring quality. You are measuring how confident a model is about itself.

Where else this shape fits

Any catalogue where the words are part of the product.

Nothing above depends on what this retailer sells. Three conditions produce the same problem in any category: more products than anyone can read, attributes that have to be exact, and each market using its own word for the same thing.

Industrial supplies and MRO

A part is found by its specification or not at all.

Auto parts

Fitment attributes that have to match across markets and marketplaces.

Electronics and components

Tolerances, standards and ratings that survive no paraphrase.

Medical and laboratory supplies

The wrong word is a compliance question, not a wording one.

Building materials

Dimensions, grades and certifications, in markets with different norms.

Marine, outdoor and technical apparel

Materials and performance claims a customer filters on.

What these have in common is not the product. Each one has a buyer who can tell whether you know what you are selling. Each one has a catalogue that nobody in your company has read end to end in any language.

Does this shape match yours?

Bring one product category and one market.

Start with one product category and one language you cannot check yourself. You see what comes back, including the terms the glossary extraction finds in content you have been publishing for years.

Questions

Before the first batch

Practical answers for teams comparing catalogue translation workflows.

Do we need a termbase before we start?

Not necessarily — terminology can be extracted from each batch and settled as the work goes. An existing termbase shortens both routes: TMX and TBX import.

What triggers a run?

In this configuration, an API call from the PIM. It could equally be a schedule, a webhook, or an automation platform — the process is the same whatever starts it.

Does a person have to be involved in every cycle?

In this setup, in two places: settling the glossary and reviewing the sample. Both are choices rather than requirements.

What happens to variants?

Translatable text usually concentrates in the parent record while variants multiply rows. What is in scope is decided by filters, set once per content shape rather than per run.

Can we keep our own translators?

Yes. Whoever reviews and whoever translates is your choice, at your rates, in the same projects the API creates.

How is this different from putting our catalogue through a model?

A model returns fluent output and forgets everything. The difference is not the translation step — it is everything around it: source comparison, terminology, review, and retained decisions.