A translation memory becomes more valuable when it remembers good decisions and more dangerous when it remembers bad ones. The same feature that makes a TM efficient—automatic reuse of past source-target pairs—can spread outdated terminology, old product names, inconsistent tone, segmentation noise, mistaken numbers, superseded legal language and low-quality machine drafts into future work. Translation memory maintenance is therefore not cosmetic housekeeping. It is quality control for one of the most reusable language assets in a localization programme.
Searches for translation memory cleanup, translation memory maintenance, TM hygiene, translation memory management, TM deduplication, clean translation memory, translation memory audit, translation memory quality, remove old translations from TM and translation memory best practices all point to one operational problem: reuse is only as trustworthy as the data being reused. A high match percentage is not an achievement when the memory is confidently suggesting yesterday’s error.
This guide explains how to maintain translation memory as a governed production asset. It covers master versus project memories, approval state, duplicate units, contradictory targets, obsolete terminology, source-version drift, segmentation, metadata, context, domain separation, machine-translation contamination, reviewer corrections, global replacements, archiving, sampling, concordance testing, performance, portable exports, backup-before-cleanup, change logs and safe revalidation. The aim is not to make the database look tidy. The aim is to make future reuse more trustworthy.
This article extends eduKateSG’s Translation Memory System owner. That article explains what TM is and how reuse, context and governance fit together. This page owns the maintenance job: how to find decay inside an existing memory, correct it without destroying valuable history, and stop bad entries from returning to production. For recovery and backup strategy, use Back Up Translation Memories, Termbases and Project Assets.
Quick answer: translation-memory maintenance is controlled forgetting
A healthy translation memory does not keep every historical segment equally available forever. It preserves trusted reusable language, marks or removes content that should no longer be proposed, separates domains that should not contaminate one another, and records enough provenance to explain why an entry is trusted. Maintenance is therefore a controlled forgetting process. The organisation remembers what remains useful and deprecates what has become misleading.
- Protect first: create a recoverable backup before any bulk maintenance.
- Classify trust: distinguish approved human translation from draft, machine, legacy and imported material.
- Find conflicts: identify duplicate sources with contradictory targets and outdated terminology.
- Use metadata: preserve project, client, domain, date, locale and status where they improve decisions.
- Archive instead of blindly deleting: remove unsafe entries from active reuse while preserving history when needed.
- Feed corrections back: reviewer fixes should repair the reusable asset, not only the current document.
- Measure reuse quality: a good memory reduces correction effort rather than merely increasing match percentages.
1. Decide which memory is allowed to influence production
Many organisations accumulate translation memories organically. One comes from an agency, another from an internal team, another from a migrated CAT tool, another from a product acquisition, and another from years of machine-assisted work. If every memory is attached to every project, the translator sees a large pile of history with little information about which history deserves trust.
Start by mapping memories by role. A production master memory should contain the strongest reusable material. A project memory may collect current work before it is promoted. A legacy archive may remain searchable without being allowed to auto-propagate. A reference memory from a partner may offer suggestions at lower priority. A machine-output memory, if retained at all, should never masquerade as human-approved production language.
This hierarchy matters because maintenance cannot succeed if everything is treated as equivalent. The first cleaning decision is architectural: which resources can write into the future? Once that boundary is explicit, the team can spend its effort on the assets that actually affect new translations.
2. Back up before bulk cleanup
Translation-memory maintenance can be destructive. Global find-and-replace, deduplication, merge operations, field changes and bulk deletions can remove or alter thousands of units in seconds. The correct first step is therefore a dated, recoverable backup of the memory plus enough metadata to identify its state.
Keep the backup outside the immediate working copy. Where practical, retain both a native backup for rich same-platform restoration and a portable TMX export for interoperability. Record memory name, language pair, export date, tool version and a short reason for the maintenance event. If the cleaning process later reveals a mistaken assumption, restoration should be ordinary, not heroic.
Do not use backup as an excuse to perform reckless transformations. A backup is a safety net, not a substitute for controlled change. Large maintenance batches should still be staged, reviewed and sampled before they replace the active production memory.
3. Define the trust classes inside the memory
A TM entry has more meaning when the team knows where it came from. An approved target accepted by a subject-matter reviewer is not equivalent to an imported unknown segment from ten years ago. A post-edited machine translation is not automatically equivalent to a human translation that passed legal review. Treating them as identical throws away useful evidence.
Define a small set of trust classes that the tools can represent: for example approved, reviewed, working, legacy, machine-origin, deprecated and blocked. The exact labels matter less than consistency. The purpose is to control how aggressively each class can be reused, who can promote it and what additional review is required.
When metadata is missing, do not invent confidence. Unknown provenance should remain unknown until reviewed. It can be useful reference without entering the highest-trust production layer.
4. Find exact duplicate sources with contradictory targets
One of the clearest signs of TM decay is the same source segment appearing multiple times with different target translations and no useful context explaining why. Some variation is legitimate: different domains, clients, eras, audiences or locales can require different wording. The problem is unlabelled contradiction inside a resource that presents all alternatives as equally reusable.
Group exact source duplicates and inspect their targets. If the differences are domain-specific, preserve them with metadata or separate memories. If one target reflects a deprecated term, mark or archive it. If two targets are stylistic alternatives and the current style guide prefers one, promote the preferred form and reduce the other form’s reuse priority.
Do not deduplicate by simply keeping the newest segment. Recency is evidence, not proof of correctness. The newest entry may be an accidental edit. Use approval status, project context, terminology and current product policy to decide.
5. Look for near-duplicates that hide terminology drift
Exact duplicates are easy. Near-duplicates are often more important. A product feature may be renamed from Workspace to Project Space, while the old term remains embedded in thousands of otherwise useful sentences. The TM then suggests high fuzzy matches containing terminology that should never return.
Search by deprecated source and target terms, product names, legal names and campaign phrases. Use concordance to see where they appear across the memory. Build a controlled replacement plan only when the change is truly systematic. A term that maps neatly in one sentence may require grammatical adjustment in another, so global replacement must be reviewed in language context.
After a terminology migration, create a regression query: deprecated forms should either disappear from the active memory or remain clearly marked as historical. This prevents the next translator from having to rediscover that the organisation changed vocabulary months ago.
6. Separate obsolete product content from reusable language
A retired feature does not make every sentence that mentioned it linguistically useless. Some generic constructions may remain valuable; other segments are dangerous because they reference functions, prices, policies or names that no longer exist. Maintenance should distinguish the language pattern from the obsolete fact.
Use project and product metadata to identify segments tied to retired content. Archive them from active production reuse when they could create factual confusion. If the CAT tool supports penalties or reference-only memories, use those controls rather than deleting valuable history indiscriminately.
Consider legal and policy content especially carefully. A beautifully translated old warranty clause can be harmful if reused after obligations change. Reuse quality depends on semantic currency, not linguistic elegance alone.
7. Repair segments that contain wrong numbers, names or identifiers
A TM can preserve errors that reviewers corrected only in the final target file. Numbers, product codes, dates and names deserve special attention because a single old mismatch can propagate repeatedly through exact or fuzzy reuse.
Run automated checks for source-target number differences where numbers should be preserved, placeholder mismatches, suspicious identifier changes and protected names. Do not treat every difference as an error; locale formatting and translated measurements may legitimately change representation. The check should produce candidates for review, not silently rewrite the memory.
When a factual mismatch is confirmed, correct it at the reusable source and document the change. Otherwise the project file may be repaired while the old bad unit continues to suggest itself in every future job.
8. Audit segmentation before blaming translation quality
Translation memory depends on segmentation. Poor segmentation can make a memory look inconsistent even when individual translations are good. Broken OCR, hard line breaks, list markers, abbreviations, HTML extraction and sentence-boundary rules can create fragmented or merged units that match poorly later.
Sample low-performing or noisy portions of the memory and inspect the source segmentation. Are headings joined to body text? Are two sentences fused? Are bullet numbers stored as part of the segment? Are line breaks creating dozens of tiny fragments? Some problems can be fixed by realignment or by improving future extraction rather than editing target language.
Be cautious when changing segmentation globally because existing bilingual alignment may depend on the current rules. A maintenance project can clean historic data, while the engineering or CAT configuration must separately prevent the same segmentation defect from returning.
9. Remove accidental source-target misalignment
Legacy memories built by alignment tools or bulk imports can contain source and target sentences that are not actually translations of one another. These entries are particularly dangerous because they may still look fluent independently. An exact source match can then insert a completely unrelated target sentence with high confidence.
Use bilingual similarity, length anomalies, terminology mismatch and random human sampling to identify suspicious alignment. High-value recurring source segments deserve manual verification. When importing aligned historical documents, keep the imported memory at a lower trust level until representative quality has been tested.
Do not rely on fluency as proof of alignment. A fluent wrong sentence can be more dangerous than obvious gibberish because reviewers may skim past it.
10. Keep machine translation from contaminating the trusted memory
AI and machine translation can accelerate drafting, but raw output should not automatically become permanent high-trust memory. If every generated segment is saved before review, the TM becomes an amplifier for whatever errors the system produced that day.
Define the promotion rule. Machine-origin output can stay in a project layer until a qualified reviewer accepts it. Only approved target text should enter the master production memory, or the origin and review state should remain explicit if the workflow keeps mixed material together.
This matters even more when TM data later becomes reference material for AI translation or model adaptation. Poor memory hygiene does not stay inside CAT tools; it can influence other automation layers. A clean TM is therefore part of AI quality governance as well as translator productivity.
11. Feed final reviewer corrections back into the memory
One of the most wasteful localisation patterns is correcting the same error in every project because the correction never reaches the reusable asset. If a reviewer fixes terminology, punctuation, tone or factual wording in the final document, the corresponding TM unit should be updated when that correction is generally reusable.
Build the feedback path explicitly. Decide which review stage can update the master memory, how conflicts are resolved and whether corrections require language-owner approval. Track repeated overwrites. If translators consistently replace the same TM suggestion, the memory is telling you it has become a liability.
Do not promote every stylistic preference blindly. Some corrections are document-specific. The maintainer’s job is to decide whether the change belongs to this sentence only, this domain, this locale, or the reusable language system.
12. Use metadata to prevent wrong-domain reuse
A sentence can have different valid translations in engineering, finance, marketing, education or law. Without domain metadata, the TM may present a technically correct but contextually wrong target as a perfect match.
Store the metadata that actually changes reuse decisions: product, domain, client, locale, content type, project date, approval status and perhaps confidentiality class. Avoid adding fields that nobody maintains. Metadata earns its place when it helps filter, rank, audit or retire entries.
If the tool supports context matches, make sure resource keys, preceding segments or document structure are captured consistently. A 100-percent text match with matching context can deserve more trust than the same text appearing in an unrelated location.
13. Split memories when one master becomes semantically incoherent
Centralization can reduce duplication, but one enormous universal memory is not automatically better. If unrelated business units, clients, regulated domains and brand voices share one resource, contradictory valid translations may constantly compete.
Use a layered design. A broad corporate memory can contain genuinely universal language. Product or domain memories can override it. Highly sensitive client or legal memories may remain isolated. The translator receives a ranked stack rather than one undifferentiated database.
The reverse problem also occurs: dozens of tiny project memories prevent good reuse. Maintenance can consolidate resources whose content and governance are compatible while keeping genuine boundaries visible.
14. Archive deprecated entries instead of erasing evidence
Deletion is appropriate for some corrupted or unauthorized data, but archiving is often safer for obsolete yet historically meaningful translations. An archive preserves audit history, helps investigate old documents and provides evidence of how terminology evolved without letting the old target appear as an active high-confidence suggestion.
Use status fields, separate archive memories or reduced priority according to the tool. The key operational condition is that deprecated entries cannot silently return to new production content. Test this after maintenance by querying known old terms and confirming their active reuse path.
Retention rules still apply. Confidential or personal information should not be preserved indefinitely merely because it exists in a TM. Archive policy must align with security, contractual and privacy obligations.
15. Clean tags, placeholders and formatting artefacts carefully
Memories imported from multiple tools can contain inconsistent inline tags, encoded entities, escaped markup and placeholder conventions. These artefacts can reduce matching or generate technically invalid reuse. Yet stripping tags blindly can damage the relationship between text and structure.
Identify the active toolchain’s expected representation and normalize only with a tested migration. Check placeholder identity, order flexibility and tag pairing. If old memories contain literal markup because extraction failed, consider separating or realigning them rather than forcing a global regex replacement.
Run structural QA after cleanup. A memory that looks cleaner in a text editor may still produce broken files if protected syntax was altered.
16. Inspect capitalization and punctuation inconsistency without over-normalizing
Repeated punctuation and capitalization variation can clutter TM suggestions, especially for UI labels and headings. Some differences are noise; others are meaningful because target-language capitalization conventions depend on sentence position, title style or product grammar.
Use style-guide rules to identify genuine inconsistencies. Normalize only when the function is the same. A sentence-start version and a menu-label version may legitimately differ even if the source text looks similar. Likewise, punctuation can carry meaning or follow locale typography.
The objective is not mechanical uniformity. It is predictable reuse for equivalent contexts. Maintenance that flattens all variation can make the memory less linguistically intelligent.
17. Sample high-frequency matches before low-frequency history
A large memory may contain millions of units. Cleaning everything with equal effort is rarely efficient. Prioritize segments that are reused frequently, appear in high-risk content, trigger repeated reviewer edits, contain important terminology or originate from uncertain imports.
Use match logs, translation volume and overwrite patterns where available. A bad segment reused two thousand times deserves attention before an obscure archived sentence that never appears. This is risk-based maintenance rather than random tidying.
Still include random sampling. Frequency-based approaches can miss rare but severe defects, especially in regulated or safety-related content. Combine targeted and random review.
18. Use concordance as a diagnostic tool
Concordance search shows how a term or phrase has been translated across the memory. It is one of the simplest ways to see inconsistency, domain variation and terminology migration. Search key product names, core verbs, legal phrases, UI nouns and frequently corrected words.
Look for patterns rather than demanding one target everywhere. If three translations correspond to three distinct meanings, the memory may be correct. If the same meaning appears under five variants because different vendors worked independently, terminology governance is missing. Use examples to decide which variation is legitimate and which is drift.
After cleanup, rerun the same concordance queries. Maintenance should create observable improvement, not just a modified file date.
19. Audit old entries against current terminology and style
Language assets age even when they contain no obvious errors. Brand voice changes. Inclusive-language guidance changes. Product terminology changes. Regulatory vocabulary changes. A target sentence accepted in 2019 may no longer fit the organisation’s 2026 communication policy.
Use age as a review signal, especially for high-frequency segments. Compare older entries against the current termbase and style guide. Do not refresh merely to make the prose sound fashionable; update when present-day guidance or product reality genuinely makes the old form less suitable.
Preserve publication history when needed. The goal is future reuse quality, not rewriting the historical record. An old translation may remain correct for its original document while being inappropriate as a reusable suggestion today.
20. Tune fuzzy-match behaviour after cleanup
A clean memory can still create poor productivity if match thresholds and penalties are misconfigured. Very low fuzzy matches may distract translators; exact matches from weak resources may deserve penalties; cross-domain memories may need lower priority.
Observe actual editing behaviour. If reviewers heavily rewrite high-percentage matches, the threshold is giving a false sense of safety. If useful near-matches never surface, it may be too strict. Context matches and approved status can justify higher trust than text similarity alone.
Tuning should follow maintenance, not replace it. Reducing the visibility of dirty data is weaker than repairing or deprecating that data.
21. Measure time-to-edit, overwrite patterns and defect recurrence
Match rate is an incomplete TM metric. A memory that produces many matches but requires heavy correction can increase cognitive load and propagate mistakes. Better indicators include how often suggestions are accepted, how much they are edited, whether terminology defects recur, and whether reviewers repeatedly overwrite the same units.
Use these signals to prioritize maintenance. If one domain shows rising correction effort, inspect its memory layer. If translators frequently ignore a supposedly trusted match, ask why. Quality analytics should help identify bad reusable knowledge rather than reward volume for its own sake.
Do not turn translator speed into a simplistic individual-performance score. Editing time is affected by content difficulty, context, tool friction and quality expectations. Use it as an asset diagnostic with context, not as a blunt ranking system.
22. Maintain a change log for major TM operations
Large maintenance events should be explainable later. Record what changed, why, who approved it, which backup precedes the change, which filters or scripts were used and what validation followed. This is especially important for global replacements, merges, domain splits and terminology migrations.
A change log makes rollback rational. It also prevents future teams from undoing intentional decisions because they cannot tell whether a pattern is historical accident or governed policy.
Keep the log readable. The objective is operational traceability, not bureaucracy. A concise record linked to the backup and maintenance report is usually more useful than thousands of undocumented manual edits.
A safe translation-memory cleanup sequence
- Inventory active, project, legacy, machine-origin and archive memories.
- Export a dated native backup and portable copy before bulk change.
- Define trust classes and approval states.
- Profile duplicate sources, contradictory targets and deprecated terminology.
- Inspect segmentation, alignment, placeholders and tag artefacts.
- Prioritize high-frequency, high-risk and frequently overwritten entries.
- Correct or deprecate factual errors and obsolete product language.
- Split domains or clients when one memory has become semantically incoherent.
- Archive historical material that should remain searchable but not reusable.
- Rebuild trust metadata and context fields where evidence exists.
- Test concordance, fuzzy matching and representative projects after cleanup.
- Promote the cleaned memory only after human sampling and rollback readiness.
Worked scenarios
Scenario 1: a product rename keeps returning in old translations
A feature changed from Team Space to Workspace. Current translators use Workspace correctly, but fuzzy matches repeatedly insert the old target because thousands of older units contain Team Space. The maintenance task begins with concordance across source and target, identifies the affected units, distinguishes historical documents from reusable UI content, and marks the old terminology deprecated in active memories.
A blind replace would be risky because some old release notes legitimately describe the retired feature. The cleanup therefore updates reusable current-domain segments, archives historical ones and adds a QA rule for the deprecated term. Future projects stop inheriting the old name without erasing history.
Scenario 2: an agency migration delivers several overlapping TMs
The organisation receives five TMX files named master, final, client-final, reviewed and latest. They overlap heavily and contain contradictory targets. Instead of merging everything into production, the team profiles overlap, samples each resource, identifies the one with reliable approval metadata and imports the others at lower trust for reference.
Only after conflicts are classified does consolidation occur. The resulting memory keeps useful history while avoiding the common migration mistake of converting uncertainty into one giant apparently authoritative database.
Scenario 3: raw machine translations entered the master TM
An automated workflow wrote every AI draft into the same memory used for approved translations. Reviewers later notice that exact matches sometimes contain unreviewed wording. The team identifies entries by creation user, timestamp and project metadata, moves uncertain material to a separate low-trust layer and rebuilds the promotion rule so only reviewed segments enter the production memory.
The incident also becomes a governance test: every future integration must prove what state is written to which memory.
Scenario 4: numbers are repeatedly corrected in financial documents
Reviewers keep changing values that entered target segments incorrectly after a historic alignment project. Automated candidate detection finds source-target number mismatches in that imported TM. High-risk entries are manually checked; confirmed misalignments are removed or repaired; the aligned memory remains lower-trust until a larger sample passes.
The main lesson is that fluent text was not enough. The reusable asset needed factual integrity.
Scenario 5: two business units need different valid terminology
A shared corporate memory produces constant terminology fights because the same source term has one approved meaning in healthcare software and another in logistics. The maintenance answer is not to choose one universal winner. The team creates a small shared general memory and separate domain memories with ranked priority and metadata.
Reuse improves because each valid translation appears in the context where it belongs.
Scenario 6: reviewers rewrite perfect matches more than fuzzy matches
Analytics show that old 100-percent matches receive heavy edits while newer fuzzy matches are often accepted. This counterintuitive pattern suggests memory decay rather than translator inconsistency. Sampling reveals that the older exact matches reflect a superseded style guide.
The team uses age, domain and overwrite frequency to prioritize a targeted refresh. Match percentage remains useful, but no longer substitutes for trust.
Twenty translation-memory maintenance checks
- Search for duplicate source segments with more than one target.
- Search deprecated product names in both source and target.
- Sample entries imported from unknown vendors or legacy systems.
- Check whether machine-origin segments can enter the trusted master without review.
- Compare reviewer corrections against the TM version used during translation.
- Inspect high-frequency exact matches with high overwrite rates.
- Find source-target number differences in domains where values should be preserved.
- Check placeholder and tag parity in frequently reused segments.
- Inspect very short or very long segments for segmentation noise.
- Search for OCR artefacts and broken line-break fragments.
- Compare active TM terminology with the current termbase.
- Identify entries tied to retired products or expired policies.
- Check whether one memory mixes incompatible domains or clients.
- Confirm archived material cannot surface as high-trust auto-propagation.
- Review old capitalization and punctuation patterns against current style.
- Test concordance for core terms before and after cleanup.
- Confirm context metadata is populated consistently where the tool uses it.
- Check fuzzy thresholds and penalties after the cleaned memory is deployed.
- Restore the pre-cleanup backup in a test environment before deleting it.
- Record the maintenance event, approval and validation result.
Release checklist for a cleaned TM
- A recoverable pre-cleanup backup exists.
- High-trust production memory is clearly separated from unknown and machine-origin data.
- Contradictory exact matches have been resolved or contextualized.
- Deprecated terminology cannot silently re-enter new work.
- Obsolete factual and policy content has been archived or blocked from reuse.
- Segmentation and alignment problems were checked before linguistic rewriting.
- Numbers, identifiers, placeholders and tags passed structural sampling.
- Reviewer corrections have a route back into reusable assets.
- Domain boundaries and memory priorities are explicit.
- Concordance and representative project tests show improvement.
- Fuzzy matching and penalties reflect trust, not only text similarity.
- The maintenance event is documented and reversible.
Frequently asked questions
How often should a translation memory be cleaned?
There is no universal calendar. Review frequency should follow change rate, reuse volume and risk. Major product renames, vendor migrations, style-guide changes, repeated reviewer corrections or rising edit effort are strong triggers for maintenance even if the last scheduled audit was recent.
Should duplicate translation units always be deleted?
No. Different targets may be valid in different domains, clients or contexts. Remove meaningless redundancy and corruption, but preserve legitimate variation with metadata or separate memories.
Should old translations be deleted?
Often they should be archived or deprecated rather than erased. Historical translations can remain useful for old documents and audits while being excluded from active high-trust reuse.
Can AI clean translation memory automatically?
Automation can identify duplicates, anomalies, terminology conflicts and suspicious entries, but bulk changes still need governed validation. A fluent model can misclassify legitimate domain variation as inconsistency, so high-impact decisions require human review.
What is TM hygiene?
It is the ongoing practice of keeping translation memory accurate, current, well-scoped and reusable by correcting or deprecating bad units, maintaining metadata and terminology alignment, and preventing unreviewed content from contaminating trusted resources.
Is a 100-percent TM match safe to use automatically?
Not by percentage alone. Trust also depends on context, provenance, approval status, domain, currency and whether the source meaning is genuinely identical in the new use.
What should happen to machine-translated segments?
Keep their origin and review status visible. Raw machine output should not silently enter the highest-trust production memory. Promote it only after the workflow’s required human or automated approval conditions are met.
How do we know cleanup worked?
Run before-and-after concordance, representative projects and editing-effort samples. The cleaned memory should surface fewer deprecated or contradictory suggestions, require less unnecessary correction and preserve the correct reusable language for its domains.
Selected references and next routes
- RWS Studio API documentation: Maintaining Translation Memories — examples of browsing, global changes, field operations, duplicate cleanup and filtering as TM maintenance tasks.
- Master Art of Translation | The Translation Memory System.
- Master Art of Translation | The Terminology System.
- Back Up Translation Memories, Termbases and Project Assets So Localization Work Can Be Recovered.
Conclusion: a translation memory should deserve its confidence
The purpose of translation memory is not to preserve every sentence forever. Its purpose is to make trustworthy prior work available when that prior work genuinely helps. Maintenance protects that purpose by separating approved knowledge from noise, current terminology from historical terminology, valid domain variation from contradiction, and reusable language from obsolete facts.
A well-maintained TM feels quieter. Translators spend less time arguing with old suggestions. Reviewers correct fewer recurring errors. Product renames stop resurfacing. High-confidence matches begin to deserve their confidence again. That is the real measure of TM hygiene: not a smaller database, but a more reliable memory of how the organisation chooses to say things now.