A translation memory is one of the most important tools in professional translation, localisation and multilingual content operations, but it is also one of the easiest tools to misunderstand. People searching for what a translation memory is, how translation memory software works, CAT tools, translation units, exact matches, context matches, fuzzy matches, TMX files, pre-translation, translation-memory quality or how AI translation works with a translation memory are usually asking one deeper question: when is an old translation safe to reuse in a new context?
A good translation memory, often shortened to TM, is not simply a database of translated sentences and it is not a machine translator. It is a structured memory of earlier source segments and their target-language translations, usually stored as translation units and searched for identical or similar source text. The European Commission describes translation memories as databases of previously translated sentences and phrases that can suggest similar or identical past translations, while the Joint Research Centre describes translation units as sentence-sized or sub-sentence source-target pairs. In modern CAT tools, these memories can also carry context, project metadata, identifiers, document information and quality signals that help a translator decide whether a match deserves trust.
The hidden challenge is that similarity is not equivalence. A 100% source-text match can still be wrong for the new document if the context, speaker, product, terminology policy, legal status, gender, date, locale or intended meaning has changed. A fuzzy match can sometimes be more useful than an exact match because it exposes the change clearly. Translation memory therefore belongs inside a larger system of source analysis, terminology, context, review and quality assurance. The art is not merely remembering translations. It is remembering them without letting yesterday’s wording overrule today’s evidence.
What this article owns in the Master Art of Translation architecture
This node owns translation memory as a complete translation architecture.
It covers what a translation memory stores, how translation units are created, how exact and fuzzy matching works, why context changes trust, how translation memories differ from termbases and machine translation, how memories are built from new work or aligned legacy translations, how teams separate working, reference and master memories, how quality enters or leaves the system, how AI and machine translation interact with remembered human translations, how TMX exchange works at a practical level, and how to govern a memory across projects and years.
It does not replace the earlier source-analysis, equivalence, context, translation-unit or terminology nodes.
Those owners answer different questions.
Source analysis asks what the current source actually means.
Equivalence asks what should remain stable across languages when forms do not align one-to-one.
The context stack asks which surrounding evidence controls interpretation.
The translation-unit node asks what must move together.
The terminology system controls concepts and approved terms.
Translation memory begins after those decisions start producing reusable bilingual evidence.
Its job is to make good decisions reusable without making reuse automatic.
The hidden problem: useful memory can become inherited error
Translation memory promises speed.
Translate once.
Reuse later.
That promise is real.
The European Commission uses translation memories to support genuine reuse of translated material, and professional CAT systems retrieve identical and similar source segments so translators do not need to recreate every target segment from zero.
But reuse creates inheritance.
Every confirmed translation unit can become a suggestion for future work.
If the original target was excellent, the memory spreads excellence.
If it contained a subtle mistranslation, the memory can spread the mistranslation.
If it used an old product name, the old name can reappear.
If a legal term was changed by policy, old equivalents can survive in historical units.
If a sentence was valid only in one gendered context, the same target can be suggested when the new context changes.
If a translator normalised an intentional ambiguity, the normalisation can be repeated.
A translation memory therefore behaves less like a neutral archive and more like a production system.
What enters it matters.
What is trusted matters.
What is retired matters.
What context is stored matters.
Who is allowed to write to it matters.
A better mental model: the translation memory is a library of prior decisions
Do not imagine a TM as a bilingual dictionary.
A dictionary is organised primarily around lexical entries and senses.
A termbase is organised around concepts and terms.
A translation memory is organised around previously translated segments.
The most useful mental model is:
current source segment → search previous source segments → inspect candidate source-target pairs → compare current context with stored context → reuse, adapt or reject.
The target suggestion is not the answer.
It is evidence that another translator, reviewer or workflow once accepted a target expression for a similar source segment.
That history can be highly valuable.
It can also be misleading.
The translator must know what kind of evidence the memory contains.
What a translation unit is
The basic record in a translation memory is usually called a translation unit, or TU.
At minimum it contains:
a source segment,
a target segment,
and the language information connecting them.
In practical systems it may also include:
creation date,
modification date,
creator,
last editor,
client,
domain,
subdomain,
project,
document name,
file path,
segment key,
previous and next segment context,
quality status,
review status,
custom metadata,
and technical identifiers.
The translation unit is therefore more than two strings.
It can be a small packet of bilingual history.
The richer that history, the easier it becomes to answer the question that matters most:
Is this old target appropriate here?
Translation memory depends on segmentation
A TM does not remember an entire document as one undivided object.
Content is segmented.
Often the segment is a sentence.
It can also be a heading, list item, cell, title, short phrase or interface string.
The segmentation rules influence what gets stored and what can later match.
If two source files divide content differently, reuse may decrease even when the underlying wording is similar.
If a sentence is split incorrectly, the memory may contain fragments that lack enough context.
If multiple sentences are merged, a future match may become too specific to retrieve.
This is why the translation-unit architecture and the translation-memory architecture are separate but connected.
Segmentation decides the unit of stored evidence.
The memory decides how that evidence is retrieved and governed.
What happens when a segment is confirmed
In a typical CAT workflow, the translator enters or edits a target segment and confirms it.
The system writes the source and target pair into a translation memory that has write access.
Later, when another source segment appears, the CAT tool searches the assigned memories and returns candidate matches.
The exact implementation differs by tool.
The governing principle is stable:
confirmation is not merely a user-interface action.
It can be a database write.
That means confirmation should be treated as a quality event.
When a translator confirms a segment, the project may be teaching its future self.
Read memories and write memories
One of the most useful governance distinctions is between memories used for reference and memories allowed to receive new entries.
A read memory supplies suggestions.
A write memory receives confirmed translations.
These roles do not have to be identical.
A project might search:
a global corporate memory,
a product memory,
a client memory,
a historical archive,
and a project working memory.
But it may write only to the working memory during translation.
After review, approved segments can be promoted to a cleaner master memory.
This prevents provisional language from immediately contaminating the most authoritative resource.
The working-memory pattern
A working TM is the project’s active bilingual notebook.
Translators write to it.
Reviewers correct it.
QA catches errors before the work is promoted.
The working memory can tolerate revision because it is expected to change.
At project completion, approved units can be merged into a master TM.
This pattern is especially useful when:
many translators are involved,
content is high risk,
review changes are frequent,
terminology is still evolving,
or the project uses external providers.
The architectural principle is separation of production from authority.
The master-memory pattern
A master TM should represent approved reusable language.
It should not merely be the biggest memory.
It should be the most controlled.
A master resource benefits from:
clear ownership,
defined write permissions,
quality requirements,
terminology alignment,
metadata standards,
duplicate handling,
retirement procedures,
and periodic audit.
A smaller clean master TM can be more useful than a giant uncontrolled memory.
Scale is not the same as value.
The reference-memory pattern
Historical documents can be useful even when they are not authoritative enough to write directly into the master TM.
A reference TM lets translators search this material without giving it equal status.
Examples include:
legacy translations before a rebrand,
acquired-company content,
old legal templates,
third-party translations,
automatically aligned archives,
machine-translated corpora,
or projects translated under an outdated terminology policy.
Reference memories preserve institutional memory while keeping provenance visible.
Exact matches
An exact match generally means the current source segment is identical to a source segment stored in the TM under the tool’s comparison rules.
Many systems display this as a 100% match.
This sounds stronger than it is.
The score usually describes source similarity.
It does not independently prove target correctness.
Suppose the source string is:
“Open”
The stored translation was created for a button that opens a file.
The new source appears as a status label meaning a case remains unresolved.
The source text is exactly the same.
The target may need to be completely different.
Exact source identity does not guarantee functional identity.
Context matches
Professional tools often distinguish a context match from a basic exact match.
For running text, this can involve the preceding and following source segments.
For structured content, it can involve a segment key or identifier.
Phrase documentation, for example, distinguishes an in-context match from a simple 100% source match and can use preceding and following segments or keys to establish context. memoQ likewise supports context matches above the ordinary exact-match score.
The useful idea is not the vendor-specific number.
It is this:
same source + same relevant context = stronger reuse evidence.
Context raises confidence because the memory is not merely saying “I have seen these words.”
It is saying “I have seen these words in a matching neighbourhood or structural role.”
Why some tools show more than 100%
Scores such as 101% or 102% can surprise learners.
They do not mean the translation is mathematically more than completely identical.
They signal additional matching evidence.
For example, a tool may distinguish:
source identity,
source plus surrounding context,
or source plus multiple context conditions.
Treat these scores as tool-specific confidence categories.
Do not assume every CAT tool defines them identically.
The important question is always:
What matched besides the source characters?
Fuzzy matches
A fuzzy match occurs when the current source resembles a stored source but is not identical.
The tool calculates a similarity score.
A high fuzzy match may differ in one number, name, adjective, tag or short phrase.
A lower match may share only part of the wording.
Fuzzy matching is powerful because translation often contains recurring patterns.
Stored source:
“Click Save to close the window.”
Current source:
“Click Apply to close the window.”
The old target can provide useful structure.
But the translator must update the changed element and verify that the sentence still functions the same way.
A fuzzy match is a draft opportunity, not automatic truth.
Similarity is not semantic identity
memoQ’s own documentation warns that its comparison algorithm uses letters and words rather than meaning, so text that looks similar can have substantially different meaning.
This warning deserves architectural status.
Consider:
“The medicine must be taken with food.”
“The medicine must not be taken with food.”
The strings are highly similar.
The meanings conflict.
A fuzzy algorithm may correctly report strong textual similarity.
A safe workflow must still detect negation.
Other dangerous small changes include:
may → must,
before → after,
increase → decrease,
include → exclude,
minimum → maximum,
above → below,
except → including,
and positive → negative polarity.
High similarity can accompany high semantic risk.
The delta-first review method
When a TM produces a fuzzy match, do not read only the target.
First inspect the source difference.
Ask:
What changed between stored source and current source?
Then classify the change.
Name?
Number?
Negation?
Modality?
Terminology?
Actor?
Date?
Condition?
Product?
Gender?
Locale?
Register?
If the delta is low risk, adaptation may be quick.
If the delta affects logic or obligation, review slowly even when the fuzzy score is high.
The amount of visual difference is not the same as the amount of meaning difference.
Match percentage is an effort hint, not a quality score
A 95% TM match does not mean the target is 95% correct.
It normally means the current source is highly similar to a stored source according to the tool’s algorithm.
The stored target could be wrong.
The new context could differ.
A one-word change could reverse meaning.
A 75% match could contain a perfect reusable clause plus irrelevant material.
Match percentage is useful for:
ranking candidates,
estimating effort,
pre-translation rules,
and workflow planning.
It should not be treated as a semantic accuracy percentage.
The trust equation
A useful conceptual model is:
reuse trust = source similarity × context fit × provenance quality × terminology currency × target quality × project relevance.
This is not a literal formula.
It is a reminder that source similarity is only one variable.
A 100% match from an unreviewed legacy memory may deserve less trust than a 92% match from a current reviewed product memory.
A context match from the wrong legal jurisdiction may still be inappropriate.
A recent approved segment using current terminology may deserve more weight than an old exact match using deprecated language.
Trust is multidimensional.
Provenance: where did this match come from?
Every useful TM should make provenance visible where possible.
A translator benefits from knowing:
which memory supplied the match,
which client or product it belongs to,
which document created it,
who confirmed or reviewed it,
when it was created,
when it was last changed,
and whether it came from alignment, human translation, post-edited machine translation or another source.
Provenance turns a bilingual string into inspectable evidence.
Without provenance, a match can look authoritative merely because it appears in the CAT pane.
Metadata as retrieval control
Metadata can improve ranking when multiple matches exist.
Phrase, for example, can prioritise matches using fields such as client, domain, subdomain and filename.
That is valuable because identical source strings often occur in different semantic environments.
Consider:
“Charge”
in a battery manual,
a pricing interface,
a legal accusation,
and physics content.
A source-only TM may return several exact matches.
Domain metadata can help surface the relevant one.
Metadata does not replace human judgment.
It helps put better evidence closer to the top.
Segment keys and structured context
Software localisation frequently contains short strings.
Short strings have fewer lexical clues.
A segment key can be more informative than neighbouring text.
Examples:
menu.file.open
status.ticket.open
door.state.open
command.open_project
The visible English source “Open” may be identical in each case.
The keys reveal function.
Modern TM architectures should therefore treat structured identifiers as context, not technical debris.
For apps, games and software, key-aware translation memory can be dramatically safer than sentence-only memory.
Previous and next segment context
For prose, the previous and next source segments can help distinguish meaning.
Suppose the segment is:
“That is correct.”
In one context, “that” refers to a calculation.
In another, it refers to a legal interpretation.
The translation may differ in gender, demonstrative choice or register.
If the same segment appears with the same surrounding source, the context match becomes stronger evidence.
The memory is reconstructing a small discourse neighbourhood.
Context can expire
Stored context is historical.
The new document may look structurally similar while the underlying world has changed.
A company changed its name.
A law was amended.
A product feature was removed.
A title changed.
An organisation adopted inclusive language.
A locale style guide changed.
A scientific term was updated.
A context match still needs currency checks when the content domain evolves.
Memory is powerful precisely because it preserves the past.
That is also why it must not silently govern the future.
Translation memory versus termbase
A termbase answers:
What target term should represent this concept?
A translation memory answers:
How was a larger source segment translated before?
These resources reinforce each other.
A TM can show how approved terminology behaves inside full sentences.
A termbase can prevent old TM language from reintroducing deprecated terms.
When they disagree, do not choose automatically.
Investigate which resource is authoritative and current.
A strong workflow defines precedence.
For example:
current approved termbase > current style guide > current reviewed master TM > legacy reference TM.
The exact hierarchy varies by project.
The principle is explicit authority.
Translation memory versus bilingual glossary
A simple glossary may list source and target expressions.
A termbase can contain richer conceptual metadata.
A TM stores segment-level bilingual text.
The resources differ in granularity.
Glossary:
invoice → facture
Termbase:
concept record for invoice, definition, preferred French term, domain, forbidden variant, notes.
Translation memory:
“Please attach the invoice before submitting your claim.” → approved target sentence.
Do not expect one resource to do the job of all three.
Translation memory versus corpus
A parallel corpus is a collection of aligned source and target texts.
A TM is engineered for retrieval and reuse in translation workflows.
The boundary can blur.
A corpus can be converted into a TM through alignment.
A TM can be exported and studied as parallel language data.
But the governance expectations differ.
Corpus evidence may show how language has been translated historically.
A production TM is expected to help create current translations.
That demands stronger quality and provenance control.
Translation memory versus machine translation
Machine translation generates a target based on a model.
Translation memory retrieves stored target text associated with similar source text.
One generates.
One retrieves.
Modern CAT environments often show both.
This is useful because the translator can compare:
prior human-approved wording,
new machine-generated wording,
terminology suggestions,
and their own interpretation.
The presence of multiple suggestions should increase judgment, not decrease it.
Translation memory versus generative AI
Generative AI can translate, paraphrase, explain, compare and reason across large contexts.
A TM provides controlled institutional memory.
The two capabilities are complementary.
AI is flexible.
TM is traceable.
AI can generate a novel sentence when no close match exists.
TM can show exactly how a company translated a recurring regulated warning last month.
AI can help evaluate whether an old match fits new context.
TM can constrain AI with approved bilingual examples.
A mature architecture uses each for what it is good at.
Retrieval-augmented translation
One useful modern pattern is to retrieve relevant terminology, translation-memory matches, style guidance and reference examples before an AI system drafts the target.
This is a form of retrieval-augmented translation.
The retrieval layer reduces the need for the model to guess institutional language.
The generative layer can then adapt retrieved evidence to the current sentence and larger context.
But the architecture needs safeguards.
Retrieved language may be stale.
Multiple examples may conflict.
A model may overfit a near match.
Human or automated validation still needs to check meaning and policy.
Build a TM from new translation
The cleanest way to build a memory is through controlled current work.
Process:
analyse the source,
translate,
apply terminology,
review,
run QA,
confirm approved segments,
store them with context and metadata,
and promote them into the appropriate reusable resource.
This creates a memory whose provenance is known.
The challenge is ensuring that confirmation happens at the right stage.
If segments enter the master TM before review, the memory can preserve drafts.
Build a TM from aligned legacy documents
Organisations often have years of source documents and translations but no structured memory.
Alignment can recover value.
An alignment tool pairs corresponding source and target segments.
The resulting bilingual units can be imported into a TM.
This is powerful.
It is also risky.
Alignment errors can pair the wrong sentences.
Old translations may use obsolete terms.
Documents may not be true equivalents.
One side may contain edits absent from the other.
Automatically aligned material should therefore carry provenance and often an alignment penalty or reference status until reviewed.
Alignment quality checklist
Before trusting aligned legacy material, inspect:
document version equality,
language direction,
missing sections,
paragraph order,
tables,
lists,
footnotes,
captions,
headers and footers,
sentence splits,
merged segments,
numbers,
names,
and repeated boilerplate.
Random sampling is useful.
Risk-based sampling is better.
Inspect sections where:
the layout differs,
tables dominate,
OCR was used,
the source was edited after translation,
or the target contains known localisation changes.
TMX and exchange
TMX stands for Translation Memory eXchange.
It is a widely used XML-based format for exchanging translation-memory data between tools.
The European Commission’s public DGT and ECDC memories are distributed in TMX form, which illustrates its practical role as an interchange format.
A TMX file can contain translation units with language variants and metadata.
Export and import make translation assets portable.
But portability is not perfect.
Tools differ in how they represent:
context,
custom fields,
tags,
user metadata,
penalties,
and proprietary features.
Always test a sample export-import cycle before assuming a full migration preserves every useful signal.
TMX migration checklist
Before migrating:
record source tool and version,
export a backup,
document language codes,
count translation units,
list custom fields,
record context settings,
note tag behaviour,
identify forbidden or deprecated entries,
and save representative screenshots or reports.
After import:
compare TU counts,
test exact matches,
test context matches,
test fuzzy retrieval,
inspect tags,
inspect metadata,
verify language variants,
and search known records.
A migration is not complete because the file imported without an error message.
It is complete when retrieval behaviour remains trustworthy.
One enormous TM or several focused TMs?
A single giant memory seems simple.
Everything is searchable.
Nothing is lost.
But very large mixed-domain memories can create noise, maintenance difficulty and conflicting exact matches.
Phrase explicitly notes that very large TMs can slow searches and become harder to maintain, and recommends smaller well-managed memories over one uncontrolled giant resource.
Architecture should reflect meaningful boundaries.
Possible partitions include:
client,
brand,
product,
domain,
legal jurisdiction,
content type,
target locale,
or quality level.
Do not fragment so aggressively that useful reuse disappears.
Design the memory around decision boundaries.
Domain separation
A technical-support TM should not necessarily share equal authority with a marketing TM.
Both may contain the same word.
Their voice and terminology can differ.
Similarly:
legal,
medical,
academic,
consumer,
engineering,
and educational content
may need separate resources or metadata priorities.
A domain boundary reduces false reuse.
Locale separation
French for France and French for Canada are both French.
They are not identical localisation targets.
The same applies to many language families and regional standards.
Differences can include:
spelling,
terminology,
date format,
currency conventions,
legal language,
punctuation,
forms of address,
and institutional names.
A TM architecture must respect target locale.
Do not let broad language labels erase target-specific requirements.
Client separation
Agencies and large translation teams often work for multiple clients in the same domain.
Client A may prefer one term.
Client B may forbid it.
A shared undifferentiated TM can leak language from one client into another.
That can create:
brand inconsistency,
confidentiality risk,
contractual problems,
and terminology errors.
Use access control, separate memories or strong metadata boundaries.
Institutional memory should not become institutional leakage.
Confidentiality and translation memory
A TM may contain sensitive source and target text.
Contracts.
Medical information.
Internal product names.
Unreleased features.
Personal data.
Security details.
Private communications.
Do not treat the TM as harmless metadata.
It is content storage.
Govern:
who can access it,
where it is hosted,
how it is backed up,
how exports are controlled,
how long data is retained,
and how data is deleted when required.
Translation technology architecture is also information-governance architecture.
Ownership and intellectual property
Who owns the translation memory?
The client?
The agency?
The translator?
The platform?
The answer depends on contract, law and workflow.
Clarify ownership before years of valuable bilingual data accumulate.
Questions include:
Can the memory be reused for other clients?
Can freelancers retain a copy?
Can it be used to train machine-learning systems?
Can aligned public material be mixed with proprietary material?
Can the client request deletion?
Operational convenience should not substitute for explicit rights.
The contamination problem
A TM becomes contaminated when low-quality or inappropriate entries enter a resource that users trust.
Sources include:
unreviewed drafts,
machine translation saved as approved human translation,
bad alignments,
wrong-language imports,
wrong locale,
legacy terminology,
test data,
duplicate conflicts,
incomplete segments,
mismatched tags,
or reviewer experiments.
Contamination is dangerous because it is quiet.
The memory may continue functioning.
It simply begins recommending worse language.
Prevent contamination before it enters
Prevention is cheaper than cleanup.
Useful controls:
write only to designated working TMs,
run QA before confirmation where possible,
restrict master write access,
separate aligned or MT-derived resources,
require terminology checks,
store provenance,
lock approved context matches when appropriate,
and perform reviewer sign-off before promotion.
Phrase’s guidance on TM quality likewise emphasises QA before saving problematic content and maintaining feedback loops.
The general lesson is simple:
the memory should not learn faster than the organisation can review.
Detect contamination after it enters
Audit signals include:
sudden terminology drift,
multiple exact matches with conflicting targets,
unexpected client names,
unusual punctuation patterns,
machine-like phrasing,
wrong locale spelling,
very recent mass changes,
entries without provenance,
or matches from documents that should not exist in the resource.
Search by:
creator,
date,
document,
project,
term,
source pattern,
or target pattern.
TM maintenance should be queryable.
Duplicate source, different target
Sometimes multiple targets for the same source are legitimate.
“Home” can mean:
homepage,
physical residence,
navigation destination,
sports home team,
smart-device state.
Context distinguishes them.
Sometimes duplicates are accidental conflicts.
The architecture should preserve legitimate context-specific alternatives and resolve uncontrolled contradictions.
Do not deduplicate blindly.
First ask whether the duplicates encode different meanings.
Conflict resolution
When two TM entries compete, inspect:
context,
domain,
client,
locale,
date,
review status,
terminology policy,
and source provenance.
Choose the current authoritative version.
If the older target is genuinely wrong, correct or deprecate it.
If it remains valid in a different context, preserve it with clearer metadata.
The goal is not one target per source at all costs.
The goal is one appropriate target per meaning-and-context combination.
Deprecation
Translation memories need a way to stop old language from returning.
A termbase may mark a term as deprecated.
A TM often requires different mechanisms:
editing entries,
removing them,
applying penalties,
moving them to a legacy resource,
or preventing the old TM from participating in current projects.
Keep an audit trail where governance requires it.
Deletion solves retrieval.
Documentation preserves history.
Rebranding
Rebrands are classic TM hazards.
Old company names, product names, slogans and interface labels can exist in thousands of translation units.
A new style guide alone does not erase them.
After a rebrand:
update terminology,
search old names in TMs,
identify reusable units,
repair critical matches,
penalise or archive legacy resources,
and run targeted QA on new projects.
Memory makes old branding persistent.
Governance must make new branding stronger.
Regulatory change
Legal and medical content can become stale because rules, warnings, classifications or approved terminology change.
A translation memory does not know that a regulation changed unless the workflow tells it.
High-risk domains benefit from:
effective dates,
document versions,
jurisdiction metadata,
periodic audit,
and strict resource selection.
A historically correct translation can become operationally wrong.
Quality before quantity
Teams sometimes celebrate TM growth.
One million units.
Five million units.
Twenty languages.
But unit count is not a quality metric.
Ask instead:
What percentage is reviewed?
How current is the terminology?
How much context is stored?
How many duplicates conflict?
How much data is aligned legacy content?
Which domains are covered?
How often do translators accept matches unchanged?
How often are TM-origin errors found?
A memory’s value is the quality of reusable decisions, not database size.
Measuring reuse responsibly
Useful metrics include:
exact-match rate,
context-match rate,
fuzzy-match distribution,
new-word rate,
repetition rate,
acceptance rate,
edit distance after insertion,
TM-origin error rate,
time saved,
and reviewer changes.
But metrics can distort behaviour.
If translators are rewarded for high TM acceptance, they may accept poor matches.
If budgets assume every 100% match needs no review, errors can pass unchecked.
Use reuse metrics to understand work.
Do not let them replace judgment.
Pre-translation
Pre-translation automatically inserts target text before a translator manually opens each segment.
It can use:
context matches,
exact TM matches,
machine translation,
non-translatables,
or combinations depending on system settings.
The threshold matters.
A strict workflow might prefill only reviewed context matches.
A lower-risk workflow may insert high fuzzy matches for editing.
The key distinction is:
inserted ≠ approved.
Pre-translation accelerates drafting.
It should not silently convert suggestions into release-ready content.
Locking high-confidence matches
Some workflows lock high-confidence context matches.
This can reduce unnecessary editing and protect stable approved language.
It can also create risk if the source context appears identical while external facts changed.
Use locking only when the organisation knows what the match category guarantees.
A locked segment should be trusted because of evidence and policy, not because the software displayed a high number.
When exact matches still need review
Review exact matches when:
the project is high risk,
terminology changed,
style changed,
product names changed,
legal rules changed,
locale differs,
the memory has mixed quality,
context is absent,
the source is short,
or the segment contains dates, numbers, names or conditions.
“100%” is not a release exemption.
It is a retrieval result.
Fuzzy-match bands
Teams often define bands such as:
95–99,
85–94,
75–84,
below threshold.
The exact bands vary.
Do not attach universal semantic meanings to them.
A 99% change from “must” to “must not” is dangerous.
An 85% change in a repeated marketing sentence might be easy.
Risk-based review should inspect what changed, not only how much.
Numbers and tags
CAT tools can recognise numbers and inline tags and may adjust or penalise matches based on these differences.
This is useful for repetitive content.
It is also a reason to verify automation.
If the system automatically changes 2025 to 2026, did the surrounding target grammar need to change?
If a placeholder moved, does the target still read naturally?
If tags wrap gendered text, did the change affect agreement?
Automatic fix-up reduces mechanical effort.
It does not remove linguistic responsibility.
Translation memory and tags
Structured files contain inline formatting and functional tags.
A TM match may have the right words and wrong tag structure.
Strictness settings can affect whether the tool considers such content exact.
For technical localisation:
inspect tag position,
paired tag integrity,
placeholder identity,
and target grammar around variables.
A target string can be linguistically beautiful and technically unusable.
Translation memory and variables
Source:
“Welcome, {name}.”
A stored translation may fit most names.
But languages with grammatical gender, case or inflection may require context about the variable value.
Similarly:
“{count} file”
may need plural handling.
Do not assume a TM unit containing variables is fully context independent.
Internationalisation design affects translation-memory quality.
Translation memory and gender
A source such as “Project manager” may translate differently depending on the person’s gender in languages that mark it.
Phrase documents context matches as a way to distinguish target alternatives for identical source strings.
This is a powerful example.
The memory needs more than source text.
It needs enough context to select the correct target.
When the product data already knows gender or grammatical class, localisation architecture should consider whether that information can be exposed safely to translation.
Translation memory and politeness
“Send the file.”
An earlier target may use informal second person.
A new document may address customers formally.
The source is identical.
The social relationship changed.
A TM that stores only source-target strings cannot detect that.
Store or infer relevant project metadata.
Use separate style-controlled memories when necessary.
Translation memory and tone
Brand voice changes by content type.
A push notification may be concise.
A help article may be calm and explanatory.
A legal notice may be formal.
A campaign headline may be playful.
Identical source fragments can require different target choices.
Do not let cross-genre reuse flatten voice.
Translation memory and terminology updates
When a preferred term changes, two systems need attention.
Termbase:
mark the new preferred term.
TM:
find old segments that can reintroduce the deprecated form.
A terminology QA check can catch some old variants during new projects.
But critical high-frequency units may deserve proactive repair.
Terminology governance and TM maintenance should be linked.
Translation memory and style-guide updates
Style rules change.
Capitalisation.
Punctuation.
Inclusive language.
Date format.
Voice.
Contractions.
Spelling variant.
An old TM can silently preserve old style.
After a major style change:
audit frequent matches,
update high-value units,
use QA where possible,
and avoid treating old exact matches as immutable.
The memory should serve the current style guide, not compete with it.
Translation memory and vocabulary learning
Translation memory can teach vocabulary when used carefully.
A learner can search how a word appears across authentic sentence pairs.
This reveals:
sense variation,
collocation,
grammar,
register,
and phrase patterns.
But a TM is not automatically a balanced language-learning corpus.
It reflects the documents that entered it.
If the memory contains only software manuals, its vocabulary evidence is domain-biased.
Use it as contextual evidence, not a universal dictionary.
Translation memory and English learning
For English learners, TM examples can expose how the same source idea maps into different English structures.
This is useful for noticing:
articles,
tense,
word order,
collocation,
phrasal verbs,
modality,
and discourse patterns.
The protected How English Works system owns those English mechanisms in depth.
The translation-memory architecture simply shows how remembered bilingual examples can surface them during translation.
The repeatable TM decision method
When a match appears, use seven steps.
Step 1: identify match type.
Context, exact, fuzzy, subsegment or reference.
Step 2: inspect the source delta.
What changed?
Step 3: inspect context.
Same document role, speaker, product, domain and locale?
Step 4: inspect provenance.
Which memory, project, date and quality level?
Step 5: inspect terminology.
Still current?
Step 6: adapt and verify target.
Do not merely insert.
Step 7: decide whether the new result should be written back.
This converts TM use from reflex to controlled reasoning.
The TM trust ladder
Highest trust may belong to:
current reviewed context match from the correct product, locale and domain.
Then:
current reviewed exact match from the correct authoritative TM.
Then:
high fuzzy match from authoritative current resources.
Then:
approved reference material from nearby domains.
Then:
legacy or aligned material.
Then:
unreviewed or machine-derived memory.
The exact order varies.
The useful idea is to rank provenance and relevance, not just similarity percentage.
Failure mode: blind 100% acceptance
Symptom:
translators skip exact matches.
Cause:
match score mistaken for quality certification.
Risk:
stale, context-wrong or historically bad target survives.
Fix:
define when exact matches require review and store context.
Failure mode: one global TM for everything
Symptom:
too many conflicting matches.
Cause:
domains, clients and locales mixed without governance.
Risk:
wrong terminology, tone and confidential leakage.
Fix:
partition resources or prioritise with metadata.
Failure mode: master TM receives drafts
Symptom:
reviewers keep fixing the same error.
Cause:
translator confirmation writes directly into authoritative resource before review.
Risk:
error propagation.
Fix:
working TM → review → promotion.
Failure mode: terminology changes but TM does not
Symptom:
deprecated term keeps returning.
Cause:
termbase updated in isolation.
Risk:
inconsistent new content.
Fix:
targeted TM cleanup plus terminology QA.
Failure mode: aligned data treated as reviewed data
Symptom:
odd mismatches appear as strong suggestions.
Cause:
automated alignment imported without provenance or penalty.
Risk:
wrong source-target pairing.
Fix:
reference status, sampling, review and penalties.
Failure mode: wrong locale reuse
Symptom:
target is linguistically valid but locally wrong.
Cause:
language code too broad or resource selection too loose.
Risk:
brand, legal and usability problems.
Fix:
locale-aware TMs and explicit project settings.
Failure mode: match score drives billing and review policy blindly
Symptom:
high-score segments receive too little attention.
Cause:
effort model confused with risk model.
Risk:
critical small-delta errors escape.
Fix:
separate commercial weighting from quality requirements.
Failure mode: memory never gets cleaned
Symptom:
old products, old terms and duplicate conflicts accumulate.
Cause:
TM treated as archive with no lifecycle.
Risk:
retrieval noise and stale language.
Fix:
scheduled audit and ownership.
Failure mode: users do not know which TM they are writing to
Symptom:
project content lands in wrong resource.
Cause:
opaque CAT configuration.
Risk:
cross-client contamination.
Fix:
visible naming, permissions and project templates.
Failure mode: AI output enters TM with no distinction
Symptom:
memory looks human-reviewed but contains raw generated text.
Cause:
automatic confirmation or poor metadata.
Risk:
future translators trust unverified output.
Fix:
provenance fields and review gates.
Translation memory with machine translation
A mature CAT environment may use a decision order such as:
context TM match,
exact TM match,
high fuzzy TM,
machine translation,
manual translation.
But this order is not universal.
For creative content, a machine translation may be less useful than a lower TM match.
For new terminology, an old exact match may be worse than fresh MT plus termbase.
For regulated content, only approved TM segments may be pre-inserted.
Architecture should encode project risk.
Translation memory with generative AI
A useful AI-assisted workflow can provide:
current source segment,
surrounding context,
top TM matches,
term candidates,
style rules,
and known constraints
to the model.
Ask the AI to produce a target that explains when it deviates from the memory.
Then verify.
This makes the TM a controlled source of examples rather than a final answer.
AI can help maintain the TM
Potential uses include:
detecting duplicate target conflicts,
classifying domain,
flagging stale terminology,
finding suspicious source-target mismatches,
suggesting metadata,
clustering similar units,
and prioritising audit samples.
Do not let AI silently rewrite authoritative memories at scale.
Use it to find candidates for human or rules-based review.
TM as organisational memory
The most mature view of translation memory is organisational.
It captures:
how the organisation names things,
how recurring sentences are phrased,
how style evolved,
which translations were approved,
and how multilingual content connects across years.
That makes the TM strategically valuable.
It also makes governance strategic.
A translation memory is not merely a translator productivity file.
It is part of multilingual institutional knowledge.
A clean setup for a new translation programme
Create:
one working TM per controlled project or programme,
one reviewed master TM per meaningful domain or product,
reference TMs for historical or aligned material,
a termbase,
a style guide,
and a decision log.
Define:
who can read,
who can write,
who can promote,
who can delete,
who audits,
and when resources are retired.
Document the project template.
Good architecture reduces individual heroics.
A clean workflow for a recurring product
Before project:
attach correct master TM,
attach relevant reference TM,
attach current termbase,
set locale,
set context handling,
set pre-translation threshold,
and verify write target.
During project:
translate,
adapt matches,
confirm into working TM,
flag ambiguity,
use QA.
After project:
review,
repair terminology,
run final QA,
promote approved entries,
archive project memory if appropriate,
and record exceptions.
Repeat.
A clean workflow for a high-risk project
Use stricter controls.
No raw MT written into master resources.
Context matches reviewed according to risk.
Exact matches checked for changed external facts.
Terminology locked to approved policy.
Numbers and conditions verified separately.
Reviewer changes synchronised back into the TM.
Release only after in-context validation.
The higher the consequence of error, the stronger the memory governance.
TM audit: start with inventory
List every memory.
For each:
name,
language pair or multilingual structure,
owner,
purpose,
domain,
client,
locale,
unit count,
last updated,
quality status,
source of data,
write permissions,
and active projects.
Many organisations discover that they have resources nobody understands.
Inventory turns files into architecture.
TM audit: inspect quality
Sample:
recent entries,
old entries,
high-frequency matches,
duplicates,
aligned data,
MT-derived data,
and units containing critical terminology.
Check:
accuracy,
terminology,
locale,
style,
tags,
numbers,
and context.
Record defect patterns.
The purpose of audit is not to prove the memory is bad.
It is to find where governance should improve.
TM audit: inspect retrieval behaviour
A memory can contain good data and still retrieve badly.
Test representative source segments.
Do the right matches appear?
Are irrelevant memories outranking useful ones?
Does context work?
Are penalties applied?
Are deprecated translations still surfaced?
Does metadata prioritisation behave as expected?
Quality exists at retrieval time, not only storage time.
TM audit: inspect update flow
Ask:
When a reviewer changes a target, does the TM receive the correction?
When content is edited outside the CAT tool, is the memory updated?
When terminology changes, who triggers cleanup?
When a project closes, what becomes authoritative?
When a client requests deletion, what happens?
A memory can decay because the feedback loop is broken.
The feedback loop
Strong TM systems close the loop:
draft → review → correction → approved memory → future reuse → new review evidence.
Weak systems stop at:
draft → delivery.
Then reviewers fix content outside the system and the TM keeps the old error.
The next project repeats it.
Every repeated correction is a signal that feedback failed to reach memory.
A simple governance charter
Purpose:
what this TM is for.
Scope:
which content belongs.
Authority:
whether it is working, master or reference.
Write policy:
who can add units.
Review policy:
what must happen before promotion.
Metadata:
required fields.
Retention:
how long data is kept.
Deprecation:
how stale content is retired.
Security:
who can access exports.
Audit:
how often quality is sampled.
This one-page charter prevents years of ambiguity.
When not to reuse
Reject a TM match when:
meaning changed,
context changed materially,
terminology is deprecated,
the source was rewritten,
the target locale differs,
the old translation is wrong,
the genre changed,
the audience changed,
the legal or factual environment changed,
or the new project intentionally adopts a new voice.
Reuse is optional.
Fidelity is not.
When to trust reuse strongly
Trust rises when:
source is identical,
context is identical,
project domain matches,
locale matches,
resource is current,
entry was reviewed,
terminology is current,
no relevant external facts changed,
and the target still fits the current style guide.
This is why a context-rich approved TM can accelerate translation dramatically.
The architecture transforms past work into present evidence.
Translation memory and human expertise
TM does not replace the translator.
It changes where expertise is spent.
Less time retyping stable recurring language.
More time evaluating differences.
More time handling ambiguity.
More time checking context.
More time on genuinely new content.
This is the real productivity gain.
The tool reduces repetitive production while increasing the importance of judgment.
Translation memory and beginner errors
Beginners may:
accept high matches automatically,
edit target without checking source delta,
ignore metadata,
overwrite a correct contextual variant,
write drafts into master TMs,
or believe the TM is a dictionary.
Teaching should therefore begin with evidence hierarchy.
Every match should answer:
where did you come from?
why do you fit here?
what changed?
what must I verify?
Translation memory as a vocabulary laboratory
Search a source word across a TM.
Observe its translations by context.
Group them by sense.
Record collocations.
Compare register.
Notice when the target uses a phrase rather than one word.
Then check the Vocabulary Learning Hub for deeper lexical study.
This turns institutional bilingual data into a sense-disambiguation exercise.
Translation memory as an English laboratory
Search one English construction.
For example:
“may be required to”
“not only … but also”
“as long as”
“be expected to”
Compare how different source languages map into the English pattern.
Or reverse the direction.
The TM reveals recurring discourse and grammar patterns.
The protected How English Works system owns the underlying English mechanisms; this architecture uses TM evidence to surface them.
The principle of controlled reuse
Translation memory works when three things happen together.
Remember.
Compare.
Verify.
Remove comparison and you get blind reuse.
Remove verification and you get inherited error.
Remove memory and you waste repeated work.
The system is strongest when all three remain explicit.
The production rule
Every translation-memory suggestion should be treated as:
a candidate with provenance,
not a command.
That single rule protects the architecture.
The software retrieves.
The translator decides.
The reviewer calibrates.
The governance system preserves what deserves to become future evidence.
Translation-memory diagnostic laboratory
The fastest way to learn translation-memory judgment is to study cases where the source similarity looks reassuring but the reuse decision is not obvious. Each diagnostic below separates the visible match from the hidden decision. The point is not to memorise a list of exceptions. It is to build a habit: inspect the source delta, context, provenance, terminology and risk before accepting the target.
Diagnostic 1: negation flips inside a high fuzzy match
Stored source: “The device must remain connected during the update.”
Current source: “The device must not remain connected during the update.”
A similarity engine may consider the sentences extremely close because only one short word changed. Semantically, the operational instruction reverses.
The translator should mark the delta as critical, ignore the reassuring fuzzy score, rebuild the target from the current source, and verify the safety consequence independently. If a workflow automatically pre-translates high fuzzy matches, negation deserves a dedicated QA rule.
This case demonstrates the central principle: a small textual change can have a large semantic effect.
Diagnostic 2: modal strength changes
Stored source: “Users may reset the password.”
Current source: “Users must reset the password.”
“May” grants permission or possibility. “Must” imposes obligation.
The strings are nearly identical. The policy is not.
A translator should treat modal changes as high-risk deltas. If the stored target used a permissive construction, editing one word may not be enough because the target language may express obligation through different grammar, mood or sentence structure.
TM reuse is safe only after the new modal force is rebuilt naturally.
Diagnostic 3: a number changes but grammar changes too
Stored source: “1 file was deleted.”
Current source: “12 files were deleted.”
A CAT tool can recognise the number difference and may adjust it automatically.
But the target language may need:
plural changes,
agreement changes,
case changes,
or a different numeral construction.
Automatic number substitution is not the same as grammatical adaptation.
The translator should verify every target element controlled by the quantity, not only the digits.
Diagnostic 4: same short source, different function
Stored source: “Back”
Context: navigation button returning to the previous screen.
Current source: “Back”
Context: anatomical label on a medical diagram.
This is a perfect source-text match and a useless target reuse.
The correct signal is functional context, not lexical identity.
For interface strings, segment keys, screenshots and product metadata may be more important than neighbouring sentences.
This is why short-string TMs need structural context.
Diagnostic 5: same source, different gendered target
Stored source: “Project manager”
Context: a male employee.
Current source: “Project manager”
Context: a female employee.
In English the string is identical. In a target language that marks gender in job titles or modifiers, the target may differ.
A context-aware TM can store both target variants if the project exposes enough distinguishing context.
A source-only exact-match workflow can silently choose the wrong form.
Diagnostic 6: exact match after a rebrand
Stored source: “Contact Acme Cloud Support.”
Current source: “Contact Acme Cloud Support.”
The source text is identical because the source repository has not yet been updated, but the brand has officially renamed the service.
The translation memory faithfully returns the old approved target.
Operationally, both source and memory are stale.
This is a reminder that external reality can invalidate an exact match. A TM does not know that a rebrand happened unless governance connects product change to language assets.
Diagnostic 7: legal term changed by policy
Stored source: “The applicant may submit an appeal.”
The legal organisation later changes the official target-language term for “appeal.”
The new source is identical.
The master termbase contains the new approved term.
The TM contains the old target.
Which resource wins?
The project should define precedence. In many controlled environments, the current authoritative termbase or legal glossary should override older segment history. The TM should then be repaired so future exact matches stop reintroducing the deprecated term.
Diagnostic 8: same sentence, different jurisdiction
Stored source: “This agreement is governed by applicable law.”
Stored project: one jurisdiction.
Current project: another jurisdiction.
Even a generic sentence can belong to different legal drafting traditions. Terminology, capitalisation, formal register and established equivalents may differ.
Do not let client or jurisdiction boundaries disappear merely because the source text is identical.
Resource selection should happen before translation begins.
Diagnostic 9: old translation was fluent but wrong
Stored source: “The treatment delayed disease progression.”
Stored target accidentally means “prevented disease progression.”
The memory returns a 100% match.
Nothing in the match score reveals the earlier semantic error.
Only bilingual review catches it.
A translation memory amplifies past quality. This case shows why master resources need quality provenance and why repeated reviewer corrections should trigger root-cause analysis in the TM rather than isolated fixes in final files.
Diagnostic 10: aligned legacy mismatch
Two old documents were automatically aligned.
Source segment A was paired with target segment B by mistake.
The resulting TU is imported into a reference TM.
Months later it appears as a fuzzy match.
The target seems unrelated, but a hurried translator assumes the old project used unusual wording.
Alignment provenance would have warned that the entry deserves lower trust.
Automatically aligned data should be visibly distinguished from reviewed production memory.
Diagnostic 11: the source changed its audience
Stored source: “Please provide your date of birth.”
Stored context: internal staff form.
Current context: public-facing child-registration form.
The sentence is identical, but the target may need different politeness, explanatory language or form conventions depending on locale and audience.
Audience is not stored in the characters.
Project metadata and style guidance must carry it.
Diagnostic 12: technical term becomes ordinary language
Stored source: “The object is created at runtime.”
Context: software development.
Current source: “The object is displayed in the museum.”
The word “object” has a technical sense in one document and an ordinary sense in another.
A fuzzy or subsegment match can tempt reuse of the wrong lexical choice.
Domain metadata and current context should dominate surface repetition.
Diagnostic 13: same terminology, different collocation
Stored source: “Apply pressure to the wound.”
Current source: “Apply the setting to all users.”
The English verb “apply” repeats, but the target language may need completely different verbs for physical pressure and configuration assignment.
TM search is most useful when it retrieves larger contextualised units rather than encouraging word-level substitution.
The termbase controls concepts; the TM supplies phrase evidence.
Diagnostic 14: punctuation difference hides meaning
Stored source: “Let’s eat, Grandma.”
Current source: “Let’s eat Grandma.”
The strings differ only by punctuation.
The meanings are famously different.
A tool configured to be permissive about punctuation could rank them as highly similar.
The translator should understand what the tool’s exact-match rules ignore. Punctuation can be semantic, not merely typographic.
Diagnostic 15: a tag difference changes emphasis
Stored source contains emphasis tags around “not.”
Current source contains the same words but emphasis tags around “today.”
A permissive tag-matching setting may still return a strong match.
If the visual emphasis carries contrastive meaning, the target structure may need reworking.
Technical tags and linguistic focus can interact.
Diagnostic 16: same text, different screen space
Stored source: “Continue with setup”
Context: desktop dialog.
Current source: “Continue with setup”
Context: narrow mobile button.
The target may be too long for the new interface.
A pure TM sees equality.
Localisation sees a constraint change.
The translator may need an approved shorter target while preserving function.
Store interface context where possible and test in product.
Diagnostic 17: source string moved from title to sentence
Stored source: “Account Security”
Context: page heading.
Current source: “Account security is important.”
A subsegment match can suggest a capitalised title-form target.
The sentence may require different case, article, inflection or word order.
Subsegment reuse must respect grammatical environment.
Diagnostic 18: terminology is current but tone is stale
A company keeps the same product vocabulary but changes brand voice from formal to conversational.
Old TMs contain correct terms inside highly formal sentence patterns.
Exact matches remain linguistically valid yet stylistically off-brand.
This is why style-guide change can require TM audit even when terminology does not change.
Diagnostic 19: locale shift inside one language
Stored target is Portuguese for Portugal.
Current project is Portuguese for Brazil.
The source matches exactly.
Target terminology, spelling, pronouns, punctuation and product phrasing may differ.
Do not rely on broad language labels.
The target locale belongs in resource selection and metadata.
Diagnostic 20: named entity changed
Stored source: “The report was submitted to the Ministry of Education.”
A government department has since been renamed.
The source template remains old.
The TM correctly reproduces the historical name.
The new document needs the current official name.
Named entities require external verification when they can change over time.
Translation memory preserves historical language; it does not independently validate current institutions.
Diagnostic 21: date-sensitive meaning
Stored source: “The current rate is 5%.”
The segment was translated three years ago.
Current source is identical.
The figure may or may not still be current.
A translator should not infer factual validity from TM recency unless the source owner confirms it.
The role of translation is not to silently fact-check or rewrite, but high-risk factual staleness should be flagged.
Diagnostic 22: one source, two legitimate translations
Source: “Home”
In a website header, the approved target means homepage.
In a real-estate application, it means residence.
A TM that permits context-specific alternatives should preserve both.
A cleanup process that deduplicates them into one target would destroy useful distinctions.
Duplicate detection must be semantic, not merely string-based.
Diagnostic 23: reviewer fix never reached the TM
The translator used a wrong target.
A reviewer corrected the final Word document after export.
Nobody updated the translation memory.
Next month the same source returns the old wrong target.
The team blames the translator for repeating the error.
The actual failure is the feedback loop.
Final-file edits must flow back into reusable assets when they affect future translation.
Diagnostic 24: raw machine translation enters master memory
An automated workflow pre-translates with MT and confirms segments automatically.
Those targets are saved to the same TM used for approved human work.
Later users cannot tell which entries were reviewed.
The resource’s authority becomes ambiguous.
Machine-derived entries need clear provenance or separation, especially before they can become master evidence.
Diagnostic 25: a translator improves style but corrupts consistency
The TM contains an approved recurring warning.
A translator dislikes the repetition and rewrites each occurrence with synonyms.
The final document sounds varied but no longer uses controlled safety language.
Not every stylistic improvement belongs in a TM.
In regulated or procedural content, stable repetition can be intentional.
Diagnostic 26: context match from wrong product version
The surrounding source segments are identical to an old manual.
The product model, however, changed internally.
The context match is technically strong because the text neighbourhood matches.
Operational meaning may still differ if the hardware or procedure changed.
Document version and product metadata can matter beyond textual context.
Diagnostic 27: target contains a deprecated inclusive-language pattern
The organisation updates its inclusive-language policy.
The source remains unchanged.
Exact matches preserve the old target constructions.
This is a governance event.
Update the style guide, then identify high-frequency or high-risk TM units affected by the policy. Do not expect new translators to manually override every exact match forever.
Diagnostic 28: fuzzy match with safe structural reuse
Stored source: “Select the file and click Upload.”
Current source: “Select the folder and click Upload.”
The difference is local and low risk if the target language uses the same grammatical frame.
A high fuzzy match can save real time.
The translator updates the changed noun, checks agreement, verifies the UI label “Upload,” and confirms the result.
This is the positive case: controlled reuse reduces repetitive work without weakening judgment.
Diagnostic 29: exact context match with high confidence
A recurring regulated disclaimer appears in the same product, same document type, same locale and same approved master TM. The previous and next segments match, terminology is unchanged and the unit was recently reviewed.
This is a strong reuse candidate.
The translator still remains accountable, but the evidence supports minimal intervention.
The architecture earns speed by improving provenance and context, not by asking people to care less.
Diagnostic 30: low fuzzy match with one valuable phrase
A 62% match is too different to reuse as a whole sentence, but it contains an approved technical phrase that is difficult to translate consistently.
The translator extracts the stable phrase, checks the termbase and rewrites the rest.
Low overall match does not mean zero value.
Subsegment evidence can still help when used selectively.
What these diagnostics teach
Across all thirty cases, the same pattern appears.
Translation-memory quality is not a property of the percentage alone.
It emerges from the relationship between:
current source,
stored source,
current context,
stored context,
resource provenance,
target-language correctness,
terminology policy,
project metadata,
and present-day requirements.
The stronger the architecture around those variables, the more confidently an organisation can reuse previous work.
The weaker the architecture, the more a high match score becomes false reassurance.
The architecture of trust: separate similarity, authority and risk
Translation memory becomes much easier to govern when three questions are separated.
Question one: how similar is the current source to a stored source?
That is the retrieval question.
Question two: how authoritative is the stored translation?
That is the provenance question.
Question three: how costly would an error be in the current context?
That is the risk question.
A CAT-tool score mostly helps with the first question.
Metadata, workflow status and resource design help with the second.
Project classification helps with the third.
Do not collapse all three into one percentage.
The similarity axis
Source similarity can be thought of in broad categories:
context match,
exact source match,
high fuzzy match,
medium fuzzy match,
low fuzzy match,
subsegment match,
and no useful match.
These categories estimate textual proximity.
They do not say whether the old target is approved, current or relevant.
The authority axis
Authority can be ranked independently.
For example:
Level A: reviewed master translation under current policy.
Level B: reviewed project translation from a trusted recent source.
Level C: approved historical translation with known provenance.
Level D: aligned legacy material.
Level E: post-edited MT with uncertain review history.
Level F: raw machine or AI output.
Level G: unknown-origin bilingual data.
A 100% match from Level G may deserve less trust than an 88% match from Level A.
The risk axis
Risk can also be classified.
Low risk:
internal understanding,
draft notes,
low-impact informational text.
Moderate risk:
public marketing,
general help content,
education,
customer support.
High risk:
legal,
medical,
financial,
safety,
eligibility,
compliance,
security,
regulated instructions.
The same match category should be reviewed differently across these layers.
A practical reuse matrix
Context match + high authority + low risk:
strong candidate for pre-translation and light review.
Context match + high authority + high risk:
strong candidate, but still bilingual review.
Exact match + medium authority + moderate risk:
review context and terminology.
High fuzzy + high authority + high risk:
inspect the source delta carefully before adapting.
Low fuzzy + high authority + low risk:
reuse only useful fragments.
Exact match + unknown authority:
do not treat as approved merely because the score is high.
This matrix makes project policy explainable.
Build naming conventions that reveal purpose
A TM name should tell the user what it is.
Bad:
TM1
Global
Old
Main
Client
Good:
ACME_ProductX_UI_fr-FR_MASTER_2026
ACME_ProductX_UI_fr-FR_WORKING_Q3
ACME_LegacyDocs_fr-FR_REFERENCE
EDUCATION_Science_en-zh_REVIEWED
Names should encode enough context that a translator can see whether the resource belongs in the project.
Avoid overlong cryptic codes nobody understands.
The naming system should reduce mistakes.
Define language direction clearly
Some systems support multilingual memories.
Others work with one source language and one or more targets.
Some can reverse language direction under certain settings.
Do not assume a memory built English→French is equally safe for French→English.
The target side of an old translation was produced under different constraints than a source authored originally in that language.
Reverse use can be useful for search.
It should not automatically imply equivalent authority.
Source-language quality affects TM quality
A TM preserves the source too.
If the source contains:
typos,
bad punctuation,
inconsistent terminology,
broken placeholders,
ambiguous abbreviations,
or OCR errors,
future matching becomes noisier.
Controlled authoring improves translation memory before translation starts.
Stable source text produces stable retrieval patterns.
This is another reason the translation architecture begins upstream.
Use source normalisation carefully
Some systems ignore or normalise differences such as:
capitalisation,
whitespace,
numbers,
punctuation,
or tags.
This can improve recall.
It can also hide meaningful distinctions.
Before enabling aggressive normalisation, ask whether those features carry meaning in the content type.
“US” and “us” are not always equivalent.
“Polish” and “polish” are not always equivalent.
A decimal comma and decimal point can change a number.
Normalisation policy should follow domain risk.
Case sensitivity
Case may be irrelevant in some languages and critical in others.
It may distinguish:
proper nouns,
acronyms,
interface labels,
variables,
legal defined terms,
or sentence position.
A case-insensitive match can be useful as a candidate.
Do not assume the target casing is automatically reusable.
Whitespace
Extra spaces are often harmless.
In code, templates or fixed-format content they may matter.
Line breaks can affect subtitles, poetry, UI layout and legal forms.
A translation memory system should know when formatting differences are linguistic noise and when they are content.
Punctuation policy
Punctuation is not universal decoration.
A question mark changes speech act.
A colon can introduce a condition.
A semicolon can separate legal obligations.
Quotation marks can signal irony or defined terms.
Ellipses can mark omission or hesitation.
Do not configure punctuation tolerance without understanding the genre.
Segmentation rules as infrastructure
Segmentation rules decide where one translation unit ends and another begins.
Abbreviations can cause false sentence breaks.
Decimal numbers can split incorrectly.
Headings may be merged with body text.
Bullets may fragment into useless pieces.
Custom segmentation rules can improve reuse, but every change affects future compatibility.
Document the rules used to build important memories.
Segment joining and splitting
Translators sometimes join two source segments to produce a natural target or split a long segment for clarity.
Tools differ in how these operations interact with TM storage.
Before using aggressive joining or splitting, know what will be stored.
A future project may not segment the same way.
Useful bilingual decisions can become hard to retrieve if structural choices are inconsistent.
Store context deliberately
Context is not free.
It increases specificity.
That is good when identical source strings need different targets.
It can reduce matches when surrounding segments change even though the current target remains reusable.
Choose context settings based on content.
Running prose benefits from neighbouring segments.
Software strings often benefit from keys.
Tables may need cell identifiers.
Forms may need field names.
A single context strategy does not fit every file type.
Context identifier strategy
For structured content, stable IDs can be powerful.
A translation unit for:
settings.notifications.email.subject
is more informative than one for the visible string:
Subject
Stable keys help match content even when file order changes.
But keys also change during refactoring.
If developers rename identifiers casually, context-match rates can fall.
Localisation architecture therefore benefits from collaboration between developers and language teams.
Document metadata strategy
Useful document-level metadata may include:
product,
version,
content type,
authoring team,
jurisdiction,
campaign,
release,
platform,
audience,
and confidentiality class.
Do not store every possible field.
Store fields that improve retrieval, audit or governance.
Metadata should earn its maintenance cost.
User metadata
Creator and last-modified-by fields can help investigations.
They can reveal:
which vendor produced an entry,
which reviewer changed it,
when a problematic batch entered,
or which workflow step created it.
But user metadata can also create privacy and access concerns.
Store what the organisation genuinely needs.
Time metadata
Creation and modification dates are valuable for currency checks.
A recent entry is not automatically correct.
An old entry is not automatically obsolete.
But date helps answer:
Did this unit predate the rebrand?
Did it exist before the regulation changed?
Was it created during the bad migration?
Did it come from the project with known QA problems?
Time is evidence.
Version metadata
Version is especially important for products and regulated documents.
A target approved for Product 3.0 may not fit Product 4.0.
A warning from an old medical device version may no longer apply.
If version is meaningful to translation, make it retrievable.
Build a promotion pipeline
A robust translation-memory system often uses stages.
Stage 1: draft working memory.
Stage 2: reviewed working memory.
Stage 3: approved project memory.
Stage 4: master reusable memory.
Promotion can be manual or automated.
What matters is that authority increases only after evidence increases.
Do not let a first draft become master data merely because a translator pressed confirm.
Define promotion criteria
Possible criteria:
bilingual review complete,
terminology QA passed,
automated QA passed,
client approval received,
in-context review complete,
no unresolved comments,
correct locale verified,
legal or subject-matter sign-off complete where needed.
Not every project needs every gate.
High-risk work often does.
Define demotion criteria
A master unit may need to lose authority when:
a term is deprecated,
a product is retired,
a regulation changes,
a translation error is found,
a client changes policy,
or a source-target pair is discovered to be misaligned.
Demotion can mean:
edit,
penalty,
archive,
remove,
or move to legacy reference.
Lifecycle includes both promotion and retirement.
Penalties as a trust signal
Some CAT tools allow penalties to reduce the apparent score of matches from less trusted resources.
This is useful when:
aligned content is helpful but not reviewed,
legacy terminology may be stale,
or vendor quality is mixed.
A penalty does not fix the data.
It changes ranking.
Use penalties to encode caution, not to avoid cleanup forever.
Multiple memories and priority order
Projects may search several TMs.
Priority can reflect:
product specificity,
client authority,
domain relevance,
locale,
quality level,
or recency.
A common pattern is:
product master first,
client general second,
domain reference third,
legacy reference last.
But if match scores always outrank metadata or priority, understand the tool’s retrieval logic.
Test actual behaviour rather than assuming configuration intent equals runtime result.
Search manually, not only through automatic matches
Concordance or TM search lets a translator search words or phrases directly in the memory.
This is useful when the automatic segment match is weak but a recurring phrase is important.
Search can reveal:
how a term appears in context,
which collocations are common,
which target variants exist,
and which clients use them.
Manual TM search turns memory into a research resource.
Concordance search and lexical evidence
Suppose the translator needs a target for “raise a concern.”
The current full sentence has no useful match.
A concordance search may show ten approved examples.
The translator can inspect which verb-noun collocation is stable in the target.
This is where TM supports vocabulary depth without replacing the termbase.
Phrase-level reuse
Subsegment matching and concordance are especially useful for:
legal boilerplate,
technical phrases,
recurring instructions,
product names,
formulaic correspondence,
and academic expressions.
But phrase-level reuse has grammar risk.
A phrase copied from one sentence may need inflection, agreement or word-order changes in another.
Reuse the concept and collocation, not necessarily the exact surface form.
Boilerplate
Some organisations have highly repetitive content.
Terms and conditions.
Safety warnings.
Regulatory statements.
Form instructions.
Boilerplate is ideal for controlled reuse if:
the source is stable,
the target is approved,
and version governance is strong.
A context-rich TM can dramatically reduce work.
But boilerplate is often high risk, so approval and change control matter more than creativity.
Marketing content
Marketing has lower lexical repetition and higher sensitivity to tone.
Translation memory can still help with:
brand phrases,
product names,
repeated feature descriptions,
legal disclaimers,
and campaign continuity.
Do not let a high TM match prevent transcreation where the brief calls for adaptation.
Reuse is subordinate to communicative function.
Literary content
Literary translation may gain less from sentence-level TM reuse because context, voice and rhythm dominate.
However, a memory can still support:
names,
invented terms,
recurring motifs,
titles,
place names,
formulaic dialogue,
and series consistency.
The architecture should not force technical-workflow assumptions onto literature.
Academic content
Academic translation benefits from reuse of:
methodological phrases,
institution names,
discipline terminology,
standard headings,
ethics statements,
and recurring project descriptions.
But claims, numbers and evidence require fresh scrutiny.
A near-identical sentence can differ in one hedge that changes epistemic strength.
Legal content
Legal TMs can be extremely valuable and extremely dangerous.
Useful reuse includes:
defined terms,
recurring clauses,
institution names,
standard notices,
and established formulations.
Risks include:
jurisdiction,
version,
amendment history,
party names,
dates,
modal force,
and defined-term scope.
Legal memory should have strong provenance and access controls.
Medical content
Medical translation memory can preserve controlled wording across:
instructions for use,
patient information,
warnings,
clinical documentation,
and product materials.
High-risk deltas include:
dose,
frequency,
contraindication,
body site,
age group,
measurement,
and risk category.
Numbers and negation deserve special QA.
Software localisation
Software is a classic TM environment.
It offers high repetition and strong structural context.
Useful signals include:
segment keys,
screenshots,
platform,
character limits,
variables,
and string comments.
The biggest source of error is treating short strings as context-free.
Game localisation
Games combine software strings with narrative.
The same source may occur in:
menus,
dialogue,
quests,
items,
tutorials,
combat messages.
A TM needs context to avoid flattening character voice and world lore.
Speaker metadata can matter as much as string keys.
Customer support
Support content repeats:
procedures,
product names,
troubleshooting steps,
policy language,
and empathy phrases.
TMs improve consistency.
But product versions and current policy change quickly.
Support memories need freshness.
Education
Educational content can reuse:
curriculum terminology,
question stems,
instruction verbs,
rubric language,
and recurring explanations.
But assessment items require special care.
A reused translation can alter difficulty or clue an answer if the context changes.
Government and public services
Public-service TMs can support consistency across:
forms,
eligibility rules,
rights information,
instructions,
and official notices.
Governance must protect:
legal meaning,
access conditions,
deadline language,
appeal rights,
and institutional names.
Security content
Cybersecurity text changes rapidly.
Old TMs may contain obsolete product names, threat terms or procedures.
Do not assume historical reuse is current best practice.
Metadata and effective dates matter.
Financial content
Financial translation memory can help with recurring reports and disclosures.
High-risk fields include:
percentages,
currencies,
dates,
accounting terms,
legal entity names,
and forward-looking statements.
Exact-match review should include numerical validation.
Build an audit sample
A useful audit sample can combine:
random units,
high-frequency units,
recent units,
very old units,
duplicate-source units,
units with critical terms,
aligned units,
MT-derived units,
and entries from low-confidence vendors.
This gives a better picture than a purely random sample.
Audit severity
Classify findings.
Critical:
could cause harm, legal exposure, unsafe action or major misinformation.
Major:
changes meaning, terminology, instruction or user action materially.
Minor:
style, punctuation or low-impact fluency.
Prefer severity based on consequence, not embarrassment.
A tiny missing “not” is critical.
A clumsy but accurate sentence may be minor.
Audit root cause
For every repeated problem, ask why it entered.
Was it:
source ambiguity,
translator error,
review failure,
term update failure,
alignment error,
migration loss,
wrong TM assignment,
permission problem,
automation rule,
or missing context?
Fixing individual entries without fixing the source process guarantees recurrence.
TM cleanup campaign
A focused cleanup can proceed in waves.
Wave 1:
critical safety and legal errors.
Wave 2:
deprecated terminology.
Wave 3:
wrong names and brands.
Wave 4:
duplicate conflicts.
Wave 5:
style and locale inconsistencies.
Wave 6:
low-value noise and obsolete fragments.
Prioritise by future retrieval impact.
Bulk editing
Bulk TM updates can save time.
They can also create large-scale damage.
Before a global replace:
define exact scope,
export backup,
test sample,
check inflection,
check false positives,
and verify after change.
Replacing a term inside sentence-level targets is not always grammatically safe.
Backups
A translation memory is a valuable database.
Back it up before:
migration,
bulk editing,
merging,
deduplication,
major terminology updates,
or tool upgrades.
A backup should be restorable, not merely exported.
Test recovery periodically for critical resources.
Retention
Not every TM should live forever.
Reasons to retire include:
client contract ending,
data retention rules,
product end-of-life,
confidentiality requirements,
or complete replacement by a clean new resource.
Retention policy should distinguish operational memory from historical archive.
Deletion requests
If translation memories contain personal or client-sensitive data, deletion may need to propagate through:
working memories,
master memories,
exports,
backups,
vendor copies,
and local freelancer resources.
Architecture should know where data went.
Data lineage matters.
Merger and acquisition scenario
Two organisations merge.
Both have:
their own TMs,
product terminology,
style guides,
and client histories.
Do not simply merge all TMs.
First map:
overlapping products,
conflicting terms,
locale policies,
quality levels,
ownership rights,
and sensitive data.
Create reference layers before master consolidation.
Organisational merger does not imply semantic equivalence.
Re-platforming scenario
The organisation changes CAT/TMS provider.
Plan:
inventory resources,
export representative samples,
map fields,
test TMX or supported exchange formats,
verify context,
verify tags,
verify metadata,
test match scoring,
test user permissions,
run parallel projects,
then migrate at scale.
Tool migration is a production change, not a file-copy exercise.
Vendor handoff scenario
A new language vendor takes over.
Provide:
approved master TM,
reference TMs labelled clearly,
termbase,
style guide,
decision log,
project template,
and write policy.
Do not hand over an unexplained folder of TMX files.
The receiving team should know which resource is authoritative.
Multi-vendor scenario
If several vendors translate the same product, central governance is essential.
Otherwise:
each vendor builds its own memory,
terminology diverges,
review fixes remain local,
and the organisation pays repeatedly for the same problems.
Use shared master resources with controlled write paths and transparent provenance.
Freelancer scenario
Independent translators also benefit from TM governance.
Even one person can separate:
client-specific memories,
personal general reference,
working projects,
and subject-domain resources.
Do not reuse confidential client content across unrelated customers.
Professional ethics applies at small scale too.
Small-team scenario
A small educational or publishing team may not need enterprise complexity.
Minimum viable setup:
one clean reviewed memory per major content family,
one working memory for current work,
one term list,
one style sheet,
periodic backup,
and a simple rule for promotion.
Governance should be proportional, not absent.
When TM ROI is high
TM provides strong value when content has:
repetition,
version updates,
stable terminology,
recurring templates,
long product lifecycles,
many target languages,
or multiple translators.
Examples:
software,
manuals,
policies,
support libraries,
regulated documentation.
When TM ROI is lower
TM may provide less sentence-level leverage for:
one-off literary works,
highly creative campaigns,
transcreation,
short bespoke speeches,
or content rewritten radically in every version.
Even then, terminology and names may still benefit from structured memory.
Cost estimation
Translation projects often use match categories in pricing.
A context or exact match may be priced differently from fuzzy or new words.
This can reflect effort.
But contractual discount bands should not dictate quality policy.
A highly discounted exact match may still need review.
Commercial models and safety models are different systems.
Weighted word counts
Some tools apply different weights to match bands to estimate effort.
This is useful for planning.
But a weighted count is not an objective measure of cognitive work.
A short legal fuzzy match can require more thought than a long new marketing sentence.
Use statistics as estimates, not laws of human effort.
Productivity measurement
Better metrics include:
time per reviewed segment,
edit effort by match type,
error rate by match type,
acceptance without change,
and reviewer correction rate.
This shows whether the TM actually helps.
If high fuzzy matches take longer than new translation, the resource or thresholds may be poorly configured.
Quality measurement
Track errors by origin.
Was the error:
introduced in new translation,
copied from TM,
introduced by MT,
introduced during post-editing,
or caused by terminology conflict?
If many errors originate from TM reuse, fix the memory rather than only retraining translators.
Match acceptance rate
A high acceptance rate can mean:
excellent memory quality.
It can also mean:
translators are not reviewing.
Interpret metrics with context.
Compare acceptance with downstream error rate.
Edit distance
Edit distance from inserted TM target to final approved target can estimate reuse usefulness.
Low edit distance across many context matches suggests strong memory quality.
High edit distance on exact matches signals a governance problem.
But style-heavy languages can produce large formal edits even when semantic reuse is useful.
Use the measure comparatively, not absolutely.
Time-to-repair
When a TM error is discovered, measure how quickly it is:
identified,
corrected in the current file,
corrected in reusable memory,
and prevented in future work.
Fast containment is a quality capability.
The memory health dashboard
A useful dashboard can include:
total active TUs,
reviewed percentage,
units by age,
units by provenance,
duplicate conflicts,
deprecated-term hits,
error reports,
top contributing projects,
and audit status.
Do not overload the dashboard.
Show measures that trigger decisions.
Translation memory and continuous localisation
Continuous localisation means source content changes frequently and translations flow continuously.
TM is central because small updates create many near matches.
Key needs:
stable string IDs,
branch/version awareness,
rapid terminology updates,
automated QA,
clear review thresholds,
and feedback from production.
A stale TM can spread mistakes at continuous-delivery speed.
Translation memory and content management systems
When translation is connected directly to a CMS, final edits may occur in the CMS after CAT review.
Those edits need a return path.
Otherwise the CMS becomes more current than the TM.
Define which system is authoritative and how corrections synchronise.
Translation memory and source control
Software projects use version control.
Localization assets benefit from similar thinking.
Track:
which source release generated the unit,
which target release approved it,
which changes were made,
and when.
TM systems are not Git repositories, but version discipline improves reproducibility.
Translation memory and branching
Product branches can diverge.
A long-term-support branch may keep old terminology while the main branch adopts new wording.
If both write into one undifferentiated TM, cross-branch contamination can occur.
Use metadata or separate resources where the language genuinely diverges.
Translation memory and release trains
For regular release cycles, define a rhythm.
Before release:
freeze terminology where necessary.
During translation:
write to working resource.
During review:
repair errors.
At release:
promote approved entries.
After release:
audit recurring changes and archive the cycle.
Memory governance becomes part of release operations.
Translation memory and rollback
If a bad translation batch is promoted, can the organisation roll back?
Useful controls:
timestamped backups,
batch identifiers,
project metadata,
creator fields,
and export snapshots.
A rollback plan matters when one automation can modify thousands of entries.
Translation memory and observability
Operational systems need observability.
Questions:
Which TM supplied this target?
Why did it rank first?
Which metadata matched?
Was a penalty applied?
When was the unit created?
Who changed it?
Without traceability, users cannot debug retrieval decisions.
Good tooling should make match provenance inspectable.
Translation memory and explainable AI
When AI receives TM examples, record which examples were provided where feasible.
If a generated target is challenged, reviewers can inspect whether the model was influenced by stale or conflicting memory.
Retrieval provenance makes AI-assisted translation more explainable.
Translation memory and prompt design
A structured AI prompt can say:
Use the following TM matches as reference, not authority.
Prefer current termbase terminology.
Preserve current source meaning even when it conflicts with old examples.
Flag contradictions between retrieved examples.
Do not invent missing context.
This prompt encodes the architecture directly.
Translation memory and automated validation
Automation can compare AI or MT output against TM and termbase signals.
Possible checks:
approved recurring warning changed unexpectedly,
forbidden term reappeared,
named entity differs from master translation,
number differs,
critical segment has no trusted match,
or high-confidence match was heavily altered.
These checks flag review targets.
They should not automatically reject every legitimate variation.
Translation memory and semantic search
Traditional TM fuzzy matching is often surface-oriented.
Newer systems can use semantic similarity.
This can retrieve paraphrases that share meaning but not wording.
That is potentially powerful.
It also increases the need for provenance and context because semantically related sentences can still differ in modality, condition or factual detail.
Semantic retrieval expands recall.
It does not eliminate verification.
Surface similarity and semantic similarity together
A mature retrieval system can combine:
string similarity,
semantic similarity,
terminology overlap,
context,
metadata,
and quality status.
The best candidate is not necessarily the one with the highest score on any single dimension.
This moves translation memory from simple lookup toward evidence ranking.
Human override
No matter how sophisticated ranking becomes, the translator must be able to reject a match.
Systems should not punish appropriate non-reuse.
Automation exists to reduce unnecessary work, not to force historical wording into unsuitable contexts.
Reviewer override
Reviewers also need visibility into match provenance.
A reviewer seeing an odd target should be able to determine:
was it new translation?
TM reuse?
MT?
AI?
This helps identify whether the issue is individual or systemic.
Governance council for large programmes
Large multilingual organisations may benefit from a small governance group covering:
localisation operations,
language leads,
terminology,
product,
legal or compliance where needed,
security,
and technology.
The group does not review every sentence.
It defines resource policy.
Governance for education and smaller organisations
Smaller organisations can use a lightweight version.
One owner can maintain:
naming,
backups,
term updates,
reviewed status,
and periodic audit.
The key is accountability.
A resource nobody owns will eventually become stale.
The authority map
Create a simple table.
Resource:
Master TM.
Authority:
high.
Purpose:
approved reuse.
Writes:
restricted.
Resource:
Working TM.
Authority:
provisional.
Purpose:
current project.
Writes:
translator and reviewer.
Resource:
Legacy TM.
Authority:
reference.
Purpose:
historical search.
Writes:
none.
Resource:
Raw MT cache.
Authority:
low.
Purpose:
draft assistance.
Writes:
automated.
This map prevents users from treating all bilingual data equally.
The escalation map
If a match conflicts with the termbase:
follow terminology authority or escalate.
If two master matches conflict:
inspect context and governance history.
If a high-risk clause has only legacy matches:
translate fresh and review.
If current source appears factually stale:
query source owner.
If the TM contains confidential material from another client:
stop and report resource contamination.
Escalation rules reduce improvisation.
The minimum viable TM policy
Even a small team should answer:
What is a TM?
Which TM is authoritative?
Who can write to it?
When is a unit considered reviewed?
How are terminology changes propagated?
How are backups made?
How are client boundaries protected?
What happens when an error is found?
Eight answers prevent many future problems.
The mature TM policy
A large programme may additionally define:
metadata schema,
locale policy,
context strategy,
pre-translation thresholds,
penalties,
alignment policy,
MT/AI provenance,
TMX migration standards,
retention,
privacy,
security,
vendor access,
audit frequency,
promotion workflow,
rollback,
metrics,
and deprecation rules.
Maturity means explicit decisions.
The central design principle
A translation memory should remember enough to be useful and forget enough to avoid becoming a museum of obsolete decisions.
That balance requires curation.
Memory without curation becomes accumulation.
Curation without memory becomes repeated work.
The architecture is the loop that keeps reuse current.
Build a translation memory from zero: a complete playbook
A new translation programme should not begin by importing every bilingual file it can find. Begin by defining the memory’s job.
Step 1: define scope.
Which product, domain, client, content type and locale does this memory serve?
Step 2: define authority.
Will it be working, master or reference?
Step 3: define source quality.
Will entries come from reviewed human translation, aligned legacy documents, post-edited machine translation or mixed sources?
Step 4: define metadata.
Which fields will materially help retrieval, audit or governance?
Step 5: define context.
Neighbouring segments, keys, IDs, document metadata or none?
Step 6: define write policy.
Who can confirm into it?
Step 7: define review gate.
When does a unit become reusable authority?
Step 8: define backup and export.
How will the resource be protected and migrated?
Step 9: define terminology precedence.
What happens when the TM and termbase disagree?
Step 10: define audit cadence.
Who samples the memory and when?
A TM built this way has a purpose before it has volume.
Clean an existing translation memory: a complete playbook
Start with a backup.
Then inventory the resource.
Identify language direction, target locales, unit count, age distribution, creators, projects, metadata fields, context settings and known problem periods.
Next, search for high-impact defects: deprecated brand names, deprecated terms, known mistranslations, wrong locale variants, broken tags, suspicious numbers, duplicate source entries with conflicting targets, and entries created during known migration or vendor incidents.
Separate problems into classes.
Critical meaning errors.
Terminology drift.
Context conflicts.
Legacy-but-valid language.
Stylistic inconsistency.
Noise.
Fix critical meaning first.
Then reduce future recurrence.
A clean TM is not one with no history.
It is one whose history is correctly labelled.
Migrate a translation memory: a complete playbook
Migration should preserve usefulness, not merely record count.
Before migration, freeze major TM edits if possible, export backups, document unit counts, document language codes, document custom fields, document context strategy, capture sample matches and record known edge cases.
Create a pilot export.
Import it into the new system.
Test exact source match, context match, high fuzzy match, lower fuzzy match, concordance search, tags, numbers, segment keys, metadata, duplicate behaviour and language variants.
Compare results.
Only then migrate the full resource.
After migration, run the same representative searches again.
If the new platform scores differently, document that change.
A successful migration reproduces reliable decisions, not identical user-interface numbers.
Merge two translation memories: a complete playbook
Never merge blindly.
First compare scope, authority, locale, domain, client, terminology date, style policy and quality history.
Then classify overlap.
Same source, same target.
Same source, different legitimate contextual target.
Same source, conflicting target.
Near-duplicate source.
Obsolete target.
Wrong-locale target.
Unknown-provenance entry.
Create resolution rules.
Keep context-specific alternatives.
Choose current authoritative targets for true conflicts.
Move uncertain material to reference status.
Merge only after the conflict policy exists.
Split a translation memory: a complete playbook
A giant global TM may need separation.
Choose split dimensions that reflect meaningful language decisions.
Product.
Brand.
Client.
Domain.
Locale.
Regulatory jurisdiction.
Quality level.
Do not split merely because the file is large.
After split, test whether translators still retrieve common reusable language.
Over-fragmentation can destroy recall.
The ideal partition reduces noise without hiding relevant evidence.
Repair a translation memory after a terminology change
Start with the termbase.
Record the old term, new preferred term, effective date, domain, locale and exceptions.
Search the TM for the old target.
Do not global-replace immediately.
Inspect grammar.
The old term may have inflected forms, compound forms, proper-name uses, historical quotations or contexts where it remains valid.
Prioritise high-frequency units and exact-match candidates.
Then add QA that flags the old term in new projects.
This creates both proactive and reactive control.
Repair after a brand rename
Search old brand name, old product name, old slogan, old URL text, old support name and old company abbreviations.
Identify whether source text also needs updating.
Update master resources first.
Archive historical translations if they remain legally or historically relevant.
Add forbidden-term QA for new content.
Notify vendors.
A rename is an ecosystem change.
Repair after a discovered mistranslation
Suppose one target phrase is known to be wrong.
Find the exact target string, morphological variants, source segments that generated it, projects that reused it and any derived memories.
Correct the current master.
Then determine whether released content needs remediation.
The TM is only one propagation surface.
A serious error may already exist in websites, apps, PDFs and downstream vendor memories.
Repair after a bad vendor batch
Identify the batch boundary: project ID, date range, creator, vendor account, files and import source.
Export affected units.
Sample them.
If quality is consistently low, remove or quarantine the batch.
If mixed, review high-risk and high-reuse units first.
Do not punish the entire TM because one period was bad.
Use provenance to isolate damage.
Repair after a broken alignment
Alignment errors often show as targets that are fluent but unrelated.
Find entries sharing alignment creator, import batch, document pair, date or custom tag.
Move them to quarantine.
Re-align representative files.
Compare quality.
If systematic, rebuild the whole aligned batch.
Bad alignment is structural; sentence-by-sentence patching may be inefficient.
Repair after wrong-locale contamination
Search for locale markers: spelling, currency, date format, government terms and common vocabulary differences.
Then use project metadata to identify the source batch.
Move valid content to the correct locale TM if rights and policy allow.
Do not simply delete useful bilingual work if it has a legitimate home.
TM governance roles
A mature programme can separate responsibilities.
Translator: creates and adapts target content.
Reviewer: checks fidelity and target quality.
Language lead: owns language policy and high-level consistency.
Terminologist: owns concepts and approved terms.
Localization engineer: owns file handling, tags, segmentation and technical context.
Project manager: owns resource assignment, deadlines and vendor flow.
TM steward: owns memory health, permissions, promotion and cleanup.
Not every organisation needs separate people.
But every responsibility should have an owner.
The TM steward
The TM steward asks questions nobody else has time to ask.
Why are there five exact targets for this source?
Why did a legacy memory outrank the master?
Why are old brand terms returning?
Why did unit count jump overnight?
Why are reviewer corrections not reaching the memory?
Stewardship turns translation memory from passive storage into maintained infrastructure.
Quarterly TM review
A practical quarterly review can include backup verification, new resource inventory, unit growth, deprecated terminology scan, duplicate-conflict scan, high-risk audit sample, locale contamination check, permission review, vendor access review and restoration test.
The review can be small.
Consistency matters more than ceremony.
Annual architecture review
Once a year, ask whether the TM structure still matches the organisation.
New products?
New markets?
New vendors?
New laws?
New style guide?
New CAT platform?
New AI workflow?
New privacy rules?
A structure designed three years ago may no longer fit.
Translation memory in an AI-first workflow
An AI-first workflow should still distinguish retrieval from generation.
Layer 1: authoritative terminology.
Layer 2: reviewed TM matches.
Layer 3: style and context.
Layer 4: AI draft.
Layer 5: automated checks.
Layer 6: human review according to risk.
If the AI sees old matches, tell it which resources are authoritative.
If examples conflict, require it to flag the conflict.
If no reliable evidence exists, allow fresh translation instead of forcing imitation.
Retrieval pack for AI
A useful retrieval pack may contain current source, previous and next source segment, document purpose, target locale, top three TM matches, match provenance, approved terms, forbidden terms, style rules and risk notes.
The pack should be compact enough to remain relevant.
More context is not always better if it introduces contradictory examples.
AI output provenance
If AI contributes to a target that enters a TM, record this where practical.
Useful labels include human translated, human reviewed, MT post-edited, AI assisted, AI generated reviewed, aligned legacy and unknown.
The goal is not to stigmatise technology.
It is to preserve evidence about how the unit was produced.
Confidence without provenance is dangerous
An AI system can generate an elegant target.
A TM can return a 100% match.
Neither surface proves authority.
The translation architecture should ask:
Where did this come from?
What controls applied?
What context supports it?
What changed since then?
That is evidence-based translation.
Advanced diagnostics: governance, reuse and production edge cases
Diagnostic 31: exact match from the wrong client
A translator opens Client B’s project, but Client A’s memory is attached by mistake.
The source sentence is common: “Thank you for contacting us.”
The retrieved target is grammatically excellent, yet Client A uses informal address while Client B requires formal address.
The error is invisible if review looks only at source-target meaning.
Resource assignment is part of translation quality.
The fix is not translator vigilance alone. Use project templates, clear TM names and restricted client boundaries.
Diagnostic 32: same company, different brand voice
A parent company owns two brands using the same English source copy.
Brand One is playful and casual.
Brand Two is formal and restrained.
One corporate TM contains both targets.
The source match is exact.
The style match is not.
Brand metadata or separate authoritative resources may be necessary when voice differs systematically.
Diagnostic 33: same product, different audience
A software company publishes an administrator guide and a consumer help guide.
The source sentence “Configure the authentication provider” appears in both.
The technical translation is appropriate for administrators.
It may be opaque for general users.
If the source itself is reused unchanged across audiences, the translator must rely on document purpose and audience metadata.
Diagnostic 34: punctuation-only change that changes function
Stored source: “Warning: Do not restart.”
Current source: “Warning — do not restart?”
The second may appear in a support conversation as a question, not a command.
A permissive matching rule can rank them highly.
Punctuation is sometimes semantic evidence.
Do not configure universal punctuation tolerance without genre awareness.
Diagnostic 35: a year update changes legal effect
Stored source: “Valid through 2025.”
Current source: “Valid through 2026.”
The edit seems mechanical.
Yet the target may use an inflected date phrase, different calendar conventions or surrounding agreement.
Verify the entire date expression.
Numbers can control grammar.
Diagnostic 36: currency identity changes
Stored source: “Price: USD 50.”
Current source: “Price: SGD 50.”
The amount is unchanged.
The monetary meaning is not.
A target locale may use currency symbols differently.
A visually small change can alter the economic fact.
Diagnostic 37: source typo fixed, target remains reusable
Stored source: “Your acccount is ready.”
Current source: “Your account is ready.”
The old target is correct because the translator understood the typo.
This is a safe fuzzy-reuse case.
After reuse, consider cleaning the old malformed source entry if it adds retrieval noise.
Diagnostic 38: source typo had changed meaning
A source typo accidentally formed another valid word.
The translator previously translated the wrong apparent meaning.
The source is now corrected.
A fuzzy match returns the old target.
The correct action is fresh interpretation, not cosmetic editing.
Diagnostic 39: quotation context changes authority
Stored source: “We must act now.”
Context: current company policy.
Current context: quotation from a historical speech.
A new translation may require an established published quotation, historical register or citation fidelity.
Exact wording does not imply identical translation function.
Diagnostic 40: canonical quotation outranks internal history
A famous passage has a recognised target-language translation.
The TM contains an earlier ad hoc version.
The current publication requires the canonical form.
External authority can outrank internal memory.
Institutional reuse should not isolate translators from established public-language conventions.
Diagnostic 41: button becomes voice command
Stored source: “Play”
Context: media-player button.
Current source: “Play”
Context: voice-command grammar.
The target may need a different form.
Functional context matters even when the source word is identical.
Diagnostic 42: placeholder expands semantic class
Stored source: “Delete {item}?”
Originally, {item} represented files.
A new product release allows it to represent people, folders and projects.
The old target may use grammar that assumes an inanimate object.
A source string can stay stable while the variable domain changes.
Internationalisation metadata must evolve with product semantics.
Diagnostic 43: confidential cross-client match
A freelancer’s private global TM contains previous client content.
A distinctive sentence appears as a strong match in another client’s project.
Even if the target is linguistically perfect, reuse may violate confidentiality or contract.
A translation-memory decision can be ethically unavailable despite linguistic relevance.
Diagnostic 44: public corpus mixed with approved internal memory
A team imports public parallel data directly into its production master TM.
Retrieval improves, but users can no longer distinguish internally approved translations from external examples.
The resource becomes epistemically ambiguous.
Keep public corpora as reference evidence unless reviewed and promoted.
Diagnostic 45: short source with several exact targets
Source: “Close”
Stored targets correspond to:
close a window,
near in distance,
close a bank account,
intimate relationship.
The TM is not wrong for containing alternatives.
The source is under-specified.
Use context keys, surrounding text or interface information.
Diagnostic 46: inclusive-language policy update
An organisation changes preferred wording for disability, gender or identity-related language.
Old exact matches remain structurally correct but policy-stale.
A style-guide update should trigger targeted TM maintenance.
Do not make translators manually fight the same exact match forever.
Diagnostic 47: content moves from internal to public
Internal documentation becomes a public help article.
The same source sentences are retained.
The stored target uses internal abbreviations and specialised shorthand.
The exact matches are accurate for the old audience and inaccessible for the new one.
Audience shift can invalidate perfect source matches.
Diagnostic 48: product term becomes ordinary word
A string that was once a product name is later used generically.
The TM preserves branding capitalisation and proprietary treatment.
The new context may require a common noun.
Entity identity must be checked.
Diagnostic 49: date string is ambiguous across source locales
Stored source: “03/04/2026.”
The old project interpreted it as 3 April.
The new source locale means 4 March.
The source characters are identical.
Locale metadata determines meaning.
Diagnostic 50: measurement policy changes
A manual previously preserved imperial units only.
New policy requires metric plus imperial.
The source wording remains largely unchanged.
Old exact matches no longer satisfy publication requirements.
Style and regulatory policy can invalidate structurally accurate memory.
Diagnostic 51: reviewer change never reaches the TM
A reviewer improves a recurring warning in the final PDF.
The CAT memory keeps the earlier wording.
The next release restores the old version.
The organisation appears to forget its own correction.
The real defect is feedback synchronization.
Diagnostic 52: two reviewers create conflict
Reviewer A approves one target.
Reviewer B later prefers another and writes it into another project memory.
Now identical sources produce competing exact matches.
This is not a database problem.
It is an unresolved language-policy decision.
Escalate, decide, document and clean the memory.
Diagnostic 53: source deletion is missed
Stored source has an extra condition clause.
The new source removes it.
A fuzzy target is inserted.
The translator edits the visible changed phrase but forgets to remove the old target condition.
Deletions can be harder to notice than additions.
Delta review must check both directions.
Diagnostic 54: new exception is overlooked
Stored source: “All users must register.”
Current source: “All users except guests must register.”
The added exception is short.
Its policy impact is large.
Small fuzzy deltas deserve risk classification based on meaning, not character count.
Diagnostic 55: invisible-character pollution
Imported memories may contain non-breaking spaces, zero-width characters or inconsistent Unicode forms.
Targets can look identical while search, export or QA behaves strangely.
Normalize carefully.
Some invisible characters are noise.
Others are linguistically or typographically necessary.
Training exercise: dangerous delta ranking
Give learners ten source pairs and hide their match percentages.
Ask them to rank semantic risk.
Examples should include:
one number change,
one negation,
one name,
one punctuation mark,
one modal,
one harmless typo.
Then reveal the fuzzy scores.
The lesson is that textual distance and semantic risk are separate.
Training exercise: trust ladder
Give five TM matches with different provenance.
For each include:
match score,
TM name,
date,
review status,
domain,
locale.
Students choose which candidate deserves first inspection and explain why.
This trains evidence ranking.
Training exercise: termbase conflict
Provide an exact TM match containing a deprecated target term.
Provide a current termbase showing a new preferred term.
Ask students to decide:
use TM unchanged,
follow termbase,
query owner,
or repair both.
The answer depends on resource authority, but they must recognise the conflict.
Training exercise: contextual duplicates
Give identical source strings with different targets.
Students classify them as:
legitimate context variants,
wrong locale,
style variants,
historical variants,
or errors.
This teaches that deduplication requires interpretation.
Training exercise: alignment audit
Provide a small legacy source-target set with several shifted alignments.
Students identify:
correct pairs,
missing segments,
merged segments,
and mispaired translations.
They then decide whether the resource belongs in master or reference status.
Training exercise: cleanup priority
Present a fictional TM containing:
one safety mistranslation,
twenty deprecated product names,
hundreds of punctuation inconsistencies,
a wrong-locale import batch,
and raw MT.
Ask students what to fix first.
Risk and future propagation should determine order.
Translation memory for individual learners
A learner can build a personal TM from carefully checked exercises.
The memory can reveal:
recurring grammar patterns,
vocabulary choices,
collocations,
and repeated mistakes.
Do not store unchecked answers as if they were authoritative.
A learning memory should document improvement, not fossilise errors.
Translation memory for teachers
Teachers can use bilingual sentence banks to demonstrate:
sense selection,
collocation,
register,
grammar transformation,
and acceptable variation.
Label pedagogical examples clearly.
A deliberately literal teaching example is not automatically a model professional translation.
Translation memory for source writers
Writers can use TM statistics and repeated translator queries as feedback on source quality.
If one source sentence repeatedly produces divergent translations, it may be ambiguous.
If small source rewrites destroy match rates, authoring may be unstable.
Translation-memory evidence can therefore improve upstream writing.
Translation memory for terminology teams
Terminologists can search TMs for:
real usage,
emerging variants,
collocations,
and inconsistent equivalents.
The memory provides context examples.
It does not automatically define preferred terms.
Frequency and authority are different.
Translation memory for reviewers
Reviewers can use concordance search to see whether a phrase is:
an approved standard,
a one-off variation,
or a recurring inconsistency.
Review should still allow legitimate context-driven variation.
Consistency is not sameness at all costs.
Translation memory for project managers
Project managers use TM analysis to estimate:
reuse,
effort,
time,
and cost.
They should understand that match bands are workflow indicators, not quality certification.
A project with many exact matches can still require rigorous review if the memory is old or high risk.
Translation memory for localization engineers
Localization engineers control:
segmentation,
file filters,
tags,
keys,
placeholders,
context extraction,
imports,
exports,
and automation.
Their technical choices determine what linguistic evidence the TM can store and retrieve.
Engineering is part of translation quality.
Translation memory for leadership
At organisational level, a TM is reusable multilingual intellectual capital.
Benefits include:
faster releases,
less repeated translation,
stronger terminology consistency,
institutional continuity,
and lower avoidable cost.
Risks include:
data leakage,
stale language,
error propagation,
vendor lock-in,
poor migration,
and uncontrolled AI reuse.
The investment is not merely a software licence.
It is governed memory.
Frequently asked question: what is a translation memory?
A translation memory is a database of previously translated source segments and their target translations, stored as translation units so identical or similar source text can retrieve earlier targets for reuse.
Frequently asked question: what is a translation unit?
A translation unit is a stored source-target pair, often a sentence or smaller segment, plus language information and sometimes context or metadata.
Frequently asked question: what is an exact match?
An exact match usually means the current source segment is identical to one stored in the TM under the tool’s comparison rules.
It does not prove that the target is correct for today’s context.
Frequently asked question: what is a context match?
A context match adds evidence beyond source identity.
Depending on the system, context may include previous and next segments, segment keys or identifiers.
Some tools label these matches above 100%.
Those numbers are tool conventions.
Frequently asked question: what is a fuzzy match?
A fuzzy match is a similar but non-identical stored source segment.
The target can be adapted as a draft if the differences are understood and verified.
Frequently asked question: does 95% mean 95% correct?
No.
It usually describes source similarity.
One changed word can reverse meaning.
Frequently asked question: what is TMX?
TMX means Translation Memory eXchange.
It is a widely used XML-based format for exchanging translation-memory data between tools.
Migration should still be tested because tool-specific metadata and context can behave differently.
Frequently asked question: does translation memory replace a translator?
No.
It reduces repetitive work and retrieves prior evidence.
The translator still decides whether the old target fits the current meaning, context and purpose.
Frequently asked question: does a TM replace a termbase?
No.
A termbase controls concepts and preferred terminology.
A TM stores larger translated segments.
They are complementary.
Frequently asked question: is translation memory the same as machine translation?
No.
TM retrieves stored translations.
Machine translation generates a new target using a model.
Modern workflows often use both.
Frequently asked question: can generative AI use a TM?
Yes.
TM matches can be supplied as retrieved examples or constraints.
The AI should be instructed to follow current source meaning and current terminology even when an old example conflicts.
Frequently asked question: should raw AI output be stored in the master TM?
Not as if it were reviewed approved content.
If AI output enters a TM, preserve provenance and require the appropriate review gate before promotion.
Frequently asked question: should every 100% match be locked?
No universal rule exists.
Locking may be efficient for stable reviewed low-risk content.
High-risk or rapidly changing content may still require review.
Frequently asked question: how many TMs should a project use?
Enough to provide relevant authoritative and reference evidence without creating avoidable noise.
More memories are not automatically better.
Frequently asked question: should every company use one global TM?
Not when products, brands, clients, domains, locales or quality levels diverge substantially.
Focused resources or strong metadata can protect relevance.
Frequently asked question: how often should a TM be cleaned?
Audit periodically and after major events such as:
rebrand,
terminology change,
vendor change,
migration,
regulatory change,
or discovery of a systemic error.
Frequently asked question: can a TM contain confidential information?
Yes.
A TM can contain segmented versions of contracts, medical documents, internal products and personal data.
Treat it as a content repository with access, retention and security rules.
Frequently asked question: what makes a good TM?
A high-quality TM has accurate targets, current terminology, clear provenance, useful context, correct locale, controlled write access, explicit authority and active maintenance.
Frequently asked question: what is the biggest TM mistake?
Treating retrieval as approval.
A match is evidence.
The current translation must still satisfy the current source.
Translation-memory project checklist
Before work begins:
correct memory?
correct locale?
correct write target?
correct termbase?
correct context settings?
correct metadata?
correct pre-translation threshold?
correct permissions?
During translation:
source delta checked?
context checked?
provenance checked?
terminology checked?
numbers checked?
tags checked?
uncertainty flagged?
working memory updated?
Before release:
review complete?
QA complete?
review corrections synchronized?
deprecated terms absent?
critical matches verified?
locale consistent?
in-context review complete?
After release:
approved units promoted?
working TM archived deliberately?
new terms captured?
systemic errors repaired?
backup created?
decision log updated?
The master rule for translation memory
Translation memory should reduce repeated thinking, not replace necessary thinking.
Reuse what remains true.
Rebuild what changed.
Retire what is obsolete.
Label what is uncertain.
Preserve context.
Preserve provenance.
Keep terminology current.
Keep authority explicit.
That is how a memory becomes an asset rather than a trap.
Where this node leads next
This Translation Memory System sits after the Translation Unit and Terminology System because reusable memory needs stable units and controlled vocabulary.
From here, the Master Art of Translation architecture can continue into human translation workflow, machine translation, generative AI translation, post-editing, translation quality assurance, revision, localisation, transcreation, interpreting, literary translation, technical translation, legal translation, medical translation, academic translation, business translation, software localisation, SEO translation, subtitle translation and language-pair applications.
The same rule follows every branch:
past bilingual evidence is useful only while current meaning, context and purpose still support it.
Professional reference layer
The European Commission Directorate-General for Translation explains that translation memories store sentences and phrases already translated and can suggest identical or similar previous translations, supporting speed and consistency.
The European Commission Joint Research Centre distributes multilingual translation-memory resources such as DGT-TM and ECDC-TM and describes translation memories as collections of source-target translation units.
memoQ documentation distinguishes exact, context and fuzzy matches and explicitly notes that similarity scoring works on text rather than meaning, which is why semantically different content can still look similar to the retrieval algorithm.
Phrase documentation describes context matches, previous and next segment context, segment keys, metadata prioritisation, pre-translation and TM maintenance.
These references support the professional mechanics used throughout this article.
No vendor score should be mistaken for a universal quality percentage.
Final conclusion
Translation memory is simple to define and difficult to govern.
At its simplest, it stores source segments and target translations.
At its best, it becomes a trusted multilingual memory of approved decisions.
The difference is architecture.
A strong TM tells users where a target came from.
It stores enough context to distinguish identical strings with different functions.
It separates provisional work from approved memory.
It lets current terminology override stale history.
It preserves legacy evidence without pretending legacy equals authority.
It works with machine translation and generative AI without losing provenance.
It is backed up, audited, migrated carefully and cleaned when language changes.
Most importantly, it never confuses similarity with truth.
The current source remains the primary evidence.
The current context remains the interpretive frame.
The current brief remains the purpose.
The memory helps the translator remember what worked before.
The art is knowing when it still works now.
Continue through the eduKateSG translation architecture
Master Translation root: Master Art of Translation | The Complete System for Moving Meaning Between Languages
Source analysis: How to Read a Source Text Before You Translate It
Equivalence: How Equivalence Works When Languages Do Not Match One-to-One
Context: The Context Stack
Translation Unit: The Translation Unit
Terminology System: The Terminology System
Vocabulary: Vocabulary Learning Hub
English: How English Works
External professional references
European Commission Directorate-General for Translation
European Commission JRC — DGT Translation Memory
memoQ — Translation Memory and Term Base
Phrase — Translation Memories Overview
Phrase — Translation Memory Match Context
Translation-memory lifecycle: from creation to retirement
A translation memory should have a lifecycle just as a product, policy or dataset does.
Creation is only the beginning.
A healthy lifecycle includes:
creation,
active use,
review,
promotion,
maintenance,
audit,
migration,
archive,
and retirement.
The resource should not remain permanently authoritative merely because it exists.
Its authority should depend on current relevance and known quality.
Creation stage
At creation, record why the TM exists.
A useful creation record includes:
owner,
language or locale,
product or domain,
source of data,
quality status,
context strategy,
write permissions,
and intended lifetime.
This prevents future teams from discovering a large bilingual resource with no explanation.
Active-use stage
During active use, monitor:
which projects write to the memory,
which projects read from it,
how often matches are accepted,
what reviewers change,
and whether terminology remains aligned.
The most dangerous time for a TM is often when it is growing quickly.
Growth can hide contamination.
Review stage
Review should not focus only on random linguistic quality.
Also inspect architecture.
Are correct projects attached?
Are users writing to the intended resource?
Are new entries carrying context?
Are vendor-specific memories leaking into master resources?
Are current terms winning over historical ones?
A technically clean TM can still be badly governed.
Promotion stage
Promotion turns provisional language into reusable authority.
Promotion should be explicit.
A translated segment can be grammatically good yet still be waiting for:
legal approval,
client approval,
product review,
terminology decision,
or in-context testing.
Do not equate translator confirmation with organisational approval unless the workflow deliberately defines it that way.
Maintenance stage
Maintenance includes:
updating terms,
repairing errors,
removing noise,
resolving conflicts,
refreshing metadata,
archiving obsolete content,
and checking access.
The best maintenance is event-driven.
A rebrand should trigger brand-term checks.
A regulatory change should trigger affected-domain checks.
A new locale should trigger resource separation.
A major tool migration should trigger retrieval tests.
Archive stage
Archive memories when they remain historically useful but should no longer drive current production.
Archived memories can support:
research,
historical comparison,
old-product maintenance,
legal record keeping,
and forensic investigation.
But they should be clearly labelled and usually penalised or removed from normal current projects.
An archive is evidence.
It is not automatically current authority.
Retirement stage
Retire a TM when its continued operational presence creates more risk than value.
Reasons include:
product end-of-life,
client contract termination,
data-retention requirements,
complete replacement by a new reviewed resource,
or irreversible contamination.
Retirement should include a record of what happened to the data and whether backups remain.
Security model for translation memory
Security should match content sensitivity.
A TM can contain highly specific fragments of confidential documents.
The fact that the document is segmented does not make the information harmless.
Control:
authentication,
role permissions,
download rights,
export rights,
API access,
vendor access,
local copies,
backup access,
and deletion procedures.
If contractors can export a master TM, understand where those exports can persist after the project ends.
Least-privilege access
Not every translator needs access to every memory.
Grant only the resources necessary for the project.
Least privilege reduces:
accidental cross-client reuse,
confidentiality breaches,
wrong-domain suggestions,
and clutter.
Access control improves both security and linguistic relevance.
Export control
TMX and other exports are portable.
That portability is useful for ownership and migration.
It also makes data easy to copy.
Define:
who may export,
where exports may be stored,
how they are transferred,
whether encryption is required,
and when copies must be deleted.
Backup security
Backups protect against loss.
They can also preserve data past its intended retention date.
A deletion policy that ignores backup copies may be incomplete.
Define how backups are rotated, restored and eventually expired.
Multilingual translation-memory architecture
Large organisations may manage dozens of languages.
Do not assume one architecture fits all.
Some language pairs have high repetition.
Others involve heavy morphology that reduces surface-match scores.
Some target languages require gender or case information absent from the source.
Some scripts need special normalisation.
Some locales share a language but differ substantially in terminology.
The architecture should remain conceptually consistent while allowing language-specific configuration.
Morphologically rich languages
In highly inflected languages, small grammatical changes can alter many target words.
A high source match may still require broad target editing.
Conversely, a low string-similarity score can hide strong semantic reuse.
Do not compare translator productivity across languages using match percentages alone.
Languages with omitted subjects
When the target or source language often omits subjects, context can become more important than segment-level wording.
An isolated TU may not carry enough information to recover person, number, gender or social relationship.
Store or expose broader context where possible.
Languages with grammatical gender
Identical English source strings may require different targets depending on gender.
This creates legitimate duplicate targets.
Use context rather than trying to force a universal one-source-one-target rule.
Languages with honorific systems
Social relationship can control verb forms, pronouns, titles and sentence endings.
A source language such as English may leave these distinctions under-specified.
Project metadata and speaker context can be essential.
Script and normalization issues
Unicode allows visually similar text to be represented in different ways.
Inconsistent normalization can reduce retrieval or create duplicate entries.
Technical teams should understand the scripts they support and test normalization carefully.
Do not apply blanket transformations that damage meaningful distinctions.
Transliteration policy
Names and non-Latin scripts may be transliterated under different systems.
A TM can preserve old romanisation conventions long after policy changes.
Keep transliteration policy linked to term and name governance.
Translation memory across related locales
Related locales can share large amounts of language.
This creates opportunities for reuse.
It also creates false confidence.
A French-France target may be understandable in French-Canada while still violating local terminology or legal conventions.
Cross-locale reuse should be deliberate and visibly lower authority unless approved.
Machine translation fallback policy
When no adequate TM match exists, a workflow may call machine translation.
Define the threshold.
For example:
context and exact TM first,
high fuzzy TM next,
MT after that,
manual translation when neither is useful.
But high-risk projects may choose a different order.
A current MT result plus current termbase can sometimes outperform an old weak TM match.
The policy should be evidence-driven.
Generative AI fallback policy
Generative AI can consider more context than a sentence-level TM and can rewrite around retrieved examples.
That is useful when:
source structure changes,
multiple TM examples conflict,
or target style needs adaptation.
But AI should not silently convert low-authority TM examples into authoritative-sounding prose.
Prompt it to distinguish:
approved examples,
reference examples,
and current instructions.
Translation-memory QA layer
A robust QA layer can check:
numbers,
units,
dates,
names,
tags,
placeholders,
forbidden terms,
missing translations,
target-source length anomalies,
duplicate inconsistencies,
and punctuation.
Some checks are deterministic.
Others require linguistic judgment.
Use automation for what it can reliably detect.
Reserve human attention for ambiguity and meaning.
QA before TM write
The ideal moment to stop a defect is before it enters authoritative memory.
Where the tool supports it, run critical checks before confirming or promoting segments.
Examples:
forbidden terminology,
missing numbers,
broken tags,
empty target,
wrong locale spelling,
and known safety phrases.
Prevention reduces cleanup cost.
QA after TM write
Not every problem can be caught immediately.
Post-write audits can search for:
deprecated terms,
suspicious target-source mismatches,
duplicate conflicts,
unusually short or long targets,
and units created by risky automation.
Use batch identifiers to make investigation easier.
Quality feedback from production
Real users reveal translation problems reviewers may miss.
Support tickets.
Search queries.
Product analytics.
User complaints.
Legal corrections.
Clinical feedback.
When production evidence reveals a translation defect, repair both the released content and the reusable memory.
Otherwise the system will recreate the defect.
The correction propagation rule
A correction is complete only when every authoritative layer that can reproduce the error has been addressed.
Possible layers:
source content,
termbase,
style guide,
translation memory,
machine-translation customisation,
AI retrieval store,
CMS,
app resource file,
PDF,
vendor memory,
and released product.
Translation memory is part of a propagation graph.
The multilingual knowledge graph idea
At scale, terminology, TMs, style guides, documents and product metadata form a connected knowledge system.
A term points to a concept.
A TM unit shows the concept in sentence context.
A document supplies genre.
A product supplies version.
A locale supplies conventions.
A reviewer supplies quality status.
Future multilingual systems will increasingly retrieve across these layers rather than treating each as an isolated file.
The architectural lesson remains stable:
preserve provenance and relationships.
The case for keeping human decisions legible
Automation becomes more powerful when past decisions are explicit.
Why was this term chosen?
Why was this exact match rejected?
Why was this resource penalised?
Why did this locale use another form?
A decision log can preserve reasoning that a TM alone cannot store cleanly.
Memory stores the result.
Governance should store the reason when the reason matters.
A final ten-question TM audit
1. Which memory is authoritative for this project?
2. Who can write to it?
3. What quality gate exists before promotion?
4. What context is stored?
5. How are terminology changes propagated?
6. How are aligned or machine-derived entries labelled?
7. How are wrong-locale or wrong-client entries prevented?
8. How are corrections from review returned?
9. How are backups, exports and confidential data controlled?
10. How does the organisation retire obsolete language?
If these ten questions have clear answers, the TM is likely functioning as managed infrastructure rather than accidental history.
The final synthesis
Translation memory solves a real problem: humans should not be forced to retranslate stable language repeatedly.
But the solution creates a second problem: remembered language can outlive the conditions that made it correct.
The complete system therefore needs two capacities at once.
Capacity one: remember reliably.
Capacity two: challenge memory when current evidence changes.
That is why the best translation-memory workflow is not maximum reuse.
It is maximum justified reuse.
A translator should reuse with confidence when the source, context, terminology, provenance and purpose align.
The translator should rebuild when any of those conditions breaks.
A reviewer should correct the memory when repeated evidence shows a problem.
A language lead should update authority when terminology or style changes.
An engineer should preserve context and technical integrity.
A project manager should attach the right resources.
A steward should keep the resource healthy.
The memory is shared.
So is responsibility.
Release policy: when a remembered translation is ready to ship
A translation memory match becomes production-ready only when three independent conditions are satisfied.
First, the current source must be understood.
Second, the retrieved target must be appropriate for the present context.
Third, the target must meet the current project’s release requirements.
This distinction matters because a TM can help with the second condition without proving the first or third.
An exact match may come from a good memory, yet the source owner may have reused stale wording.
A context match may fit linguistically, yet the project may require a new inclusive-language policy.
A reviewed target may be correct, yet the final interface may truncate it.
Release readiness is therefore broader than memory confidence.
Decision matrix: low-risk exact match
Scenario:
A recently reviewed help-centre sentence appears unchanged in the same product, same locale, same content type and same release family.
The termbase is unchanged.
The style guide is unchanged.
The surrounding context matches.
Recommended action:
reuse the target,
perform a quick bilingual check,
confirm technical integrity,
and allow normal QA to complete the release.
Why:
all major evidence layers agree.
The purpose of a well-governed TM is to make this kind of routine reuse fast.
Decision matrix: high-risk exact match
Scenario:
A medical warning appears as an exact match from the correct reviewed master memory.
The product and locale are the same.
Recommended action:
reuse as a candidate,
verify source meaning,
verify numbers, negation and conditions,
check whether regulatory wording has changed,
and include the segment in the required high-risk review.
Why:
high authority reduces linguistic uncertainty, but consequence of error remains high.
Exact match does not cancel risk policy.
Decision matrix: current fuzzy match with one critical change
Scenario:
Stored source:
“Do not use above 30°C.”
Current source:
“Do not use below 30°C.”
The match is textually very high.
Recommended action:
treat as a fresh high-risk translation decision,
not a cosmetic edit.
Verify the condition independently.
Why:
the changed word controls the operational threshold.
The semantic delta outweighs the match score.
Decision matrix: old exact match, new terminology policy
Scenario:
The source sentence is identical.
The target uses a term deprecated last month.
The current termbase contains the replacement.
Recommended action:
apply the current preferred term,
check grammatical consequences,
update the target,
and repair the TM so the old exact match stops returning.
Why:
current terminology authority outranks stale segment history.
Decision matrix: reference TM versus master TM
Scenario:
A 100% match appears from a legacy reference TM.
A 92% match appears from a current reviewed master TM.
The two targets use different terminology.
Recommended action:
inspect both,
prefer current terminology and project relevance,
and do not assume the 100% source match is superior merely because the number is higher.
Why:
match score and resource authority measure different things.
Decision matrix: two exact matches with different targets
Scenario:
The current source is identical to two stored source units.
Both are 100%.
Targets differ.
Recommended action:
inspect context, locale, domain, date, source document and review status.
Determine whether both are legitimate context variants or whether one is obsolete.
Why:
multiple exact matches are not automatically a database defect.
They may encode real contextual distinction.
Decision matrix: machine translation versus weak TM
Scenario:
The best TM match is 58% from old unrelated content.
Current MT output follows current terminology and reads naturally.
Recommended action:
use whichever candidate provides the better starting point after bilingual verification.
Do not force TM reuse for ideological reasons.
Why:
retrieval and generation are tools.
Current meaning is the authority.
Decision matrix: AI draft with strong TM evidence
Scenario:
A generative AI system receives three reviewed TM matches, current termbase entries and the full paragraph.
It produces a target that differs from the closest exact historical phrasing.
Recommended action:
compare the AI deviation with current source and style.
Accept the new wording if it better fits present context and remains compliant.
Do not treat historical wording as untouchable.
Why:
memory supports judgment.
It does not prohibit better current translation.
Decision matrix: source author changes intent without changing many words
Scenario:
A marketing sentence is reused in a new campaign, but the call to action now targets existing customers instead of prospects.
Most source words remain the same.
Recommended action:
review communicative purpose before reusing the old target.
Why:
audience and intent can change while lexical similarity remains high.
Decision matrix: same source across web and app
Scenario:
A sentence exists in a desktop web page and a mobile app.
The source is identical.
The mobile interface has strict character limits.
Recommended action:
preserve function and terminology while adapting target length if project policy allows.
Store context-specific variants with keys or metadata.
Why:
layout constraint is part of translation context.
Decision matrix: reviewer rejects an accepted TM match
Scenario:
The translator inserts a reviewed master-TM target unchanged.
The reviewer rejects it because current context differs.
Recommended action:
treat the reviewer decision as evidence about context, not as disobedience to the TM.
If the old target remains valid elsewhere, keep it with clearer context.
If it is wrong generally, repair it.
Why:
human review exists partly to detect when historical memory no longer fits.
Decision matrix: target edited outside the translation system
Scenario:
After translation, a product manager changes the wording directly in the CMS.
Users prefer the new wording.
The TM still contains the old target.
Recommended action:
decide whether the CMS edit is approved language.
If yes, feed it back into the reusable memory and style resources.
Why:
otherwise the next release will overwrite the product manager’s accepted improvement.
Decision matrix: new locale derived from an existing locale
Scenario:
A team launches a new regional locale using an existing target-language TM as reference.
Recommended action:
copy only when rights and governance permit,
mark the inherited data as reference,
identify locale differences,
review high-frequency units,
and promote only locally approved translations.
Why:
understandability is not the same as locale correctness.
Decision matrix: emergency translation
Scenario:
A time-critical public notice must be translated immediately.
The TM contains a strong reviewed match from an earlier event.
Recommended action:
reuse high-confidence stable language,
but verify dates, locations, authorities, instructions and emergency details independently.
Why:
emergency speed increases the value of controlled reuse while also increasing the cost of carrying over one stale fact.
Decision matrix: educational assessment reuse
Scenario:
A question instruction appears as an exact match in a test.
Recommended action:
verify that the target preserves the same cognitive demand, difficulty and answerability.
Do not vary command words casually.
Why:
assessment translation measures learning.
A small wording change can change what the item tests.
Decision matrix: historical text
Scenario:
A phrase appears in an old archive and matches a modern translation unit.
Recommended action:
check period meaning and publication purpose before importing modern wording.
Why:
historical translation may need to preserve period register, obsolete terminology or source ambiguity that modern production memory normalises away.
A release-gate checklist for TM-origin segments
Before a TM-origin target reaches publication, ask:
Does the current source mean what the stored source meant?
Is the stored context comparable?
Is the memory authoritative for this project?
Is the target locale correct?
Is terminology current?
Is the target still grammatically appropriate after all source changes?
Are numbers, units, dates and names correct?
Are conditions, exceptions, modality and negation preserved?
Are tags and placeholders intact?
Does the target fit the current interface or document?
Has the required risk-level review occurred?
If the answer to any critical question is no, the match is not release-ready.
Why this final gate matters
Translation memory creates speed by collapsing repeated effort.
Release governance creates safety by restoring the checks that repetition might otherwise bypass.
The two are not enemies.
The best system uses trusted memory to make routine work fast and uses risk-aware checks to slow down only where meaning can drift.
That is the mature balance:
fast where evidence is strong,
careful where consequences are high,
and always accountable to the current source.
