A translation error taxonomy is useful only if it tells you what failed, not merely that a reviewer disliked a sentence. “Wrong,” “awkward,” “bad style” and “not natural” may express dissatisfaction, but they do not create a repeatable quality system. Professional review becomes far more useful when errors are classified by the kind of damage they cause: meaning, completeness, terminology, reference, logic, register, grammar, locale, formatting, compliance or process.
Searches for translation error taxonomy, translation error categories, translation quality errors, translation review categories, translation QA errors, translation mistake analysis, translation quality assessment, translation error severity and how to evaluate translation quality all circle the same professional problem: if reviewers cannot name the failure consistently, teams cannot measure it, learn from it or prevent it from recurring.
This guide builds a practical taxonomy from first principles. It separates semantic errors from stylistic preferences, distinguishes terminology inconsistency from ordinary lexical variation, treats names and numbers as factual integrity, gives logic and reference their own categories, and shows how error classification should connect to severity, root cause and corrective action. The aim is not to turn language into a spreadsheet. The aim is to make review evidence precise enough to improve the next translation rather than merely correct the current one.
Quick Read
A translation error taxonomy is a shared classification system for describing what went wrong. It lets translators, reviewers, project managers and subject experts distinguish different failure types, compare recurring patterns, set priorities and decide what needs a process fix rather than a one-sentence edit.
The one-sentence answer
Classify errors by the function they damage—meaning, completeness, terminology, facts, logic, reference, language, register, locale, format or compliance—then record severity separately from category.
Why categories and severity must be separated
A category answers, “What kind of problem is this?” Severity answers, “How much does this problem matter here?” Those are different questions. A punctuation error in a casual article may be minor. A punctuation error in a numerical expression or legal clause can change interpretation. A terminology variation may be harmless in an essay but critical in a regulated document.
If a taxonomy mixes category and severity, reviewers start inventing labels such as “serious terminology” or “small grammar,” and the system becomes inconsistent. Better architecture uses a stable set of categories and a separate severity scale based on consequence.
This also reduces emotional review. A translator can disagree with a severity decision while still agreeing that the issue belongs to, for example, reference or terminology. The conversation becomes about evidence and impact rather than personal taste.
1. Meaning transfer error
A meaning transfer error occurs when the target communicates a different proposition from the source. The error may reverse an action, change who did what, alter cause and effect, change time, strengthen or weaken certainty, or substitute one concept for another.
This is the central semantic category. It should not be used for every sentence a reviewer would phrase differently. The question is whether the target reader receives a materially different meaning.
Example: source meaning, “The medicine may reduce symptoms.” Target meaning, “The medicine reduces symptoms.” The grammar may be perfect, but uncertainty has been removed. That is a meaning transfer error, not a style issue.
2. Omission
An omission occurs when meaning-bearing source content disappears from the target without authorisation. Omitted content can be a whole sentence, a qualifier, a list item, a negative, a condition, an example or a small word that carries important scope.
Not every missing surface word is an omission. Languages package meaning differently. A source article, pronoun or grammatical marker may have no independent target equivalent because its meaning is expressed elsewhere. Judge meaning, not word count.
A useful review note identifies what information was lost and why that loss matters. “Missing ‘unless approved’ changes the condition” is better than “omission.”
3. Addition
An addition occurs when the target introduces information, certainty, evaluation or instruction that the source does not support and the brief does not authorise. Translators sometimes add material to improve clarity, especially when the source relies on shared cultural or grammatical context. Some explicitation is legitimate. The error begins when clarification becomes new content.
Example: a source says “the proposal was criticised.” The target says “the proposal was widely criticised for being ineffective.” If neither “widely” nor “ineffective” is supported, the translation has acquired new claims.
Addition is particularly important in educational, legal, medical, journalistic and historical translation, where the target should not quietly improve the evidence.
4. Terminology error
A terminology error occurs when a domain concept is represented by the wrong target term or by a prohibited variant in a context where controlled terminology matters. Terminology is not simply “difficult vocabulary.” It is the language by which a domain keeps concepts distinct.
A reviewer should be able to point to a definition, termbase, authoritative usage or domain convention. Personal preference is not enough. If two target words are both ordinary synonyms and no controlled distinction exists, the issue may be style rather than terminology.
Terminology errors can be severe even when the sentence remains understandable because they can damage legal identity, technical precision, searchability or cross-document consistency.
5. Terminology inconsistency
Wrong terminology and inconsistent terminology are related but not identical. A translator may use two individually acceptable terms for one concept. In creative prose that may be good style. In technical documentation it may make readers think two concepts exist.
This category is useful because the corrective action differs. Wrong terminology requires choosing the correct concept label. Inconsistency requires deciding whether variation is permitted and, if not, standardising the approved form.
Do not penalise natural grammatical variation. Singular, plural, inflected or syntactically transformed forms can still represent one controlled term correctly.
6. Name and entity error
This category covers people, organisations, products, programmes, laws, institutions, places and other named entities. Errors include misspelling, wrong transliteration, unofficial translation, outdated organisation name or confusion between similar entities.
Names deserve their own category because the corrective process is often research rather than linguistic editing. A spellchecker cannot know whether an institution has an official target-language form.
Entity errors can also break search, indexing and records. A small orthographic difference may look stylistic while actually producing a different identity.
7. Number, date, unit or factual-token error
Numbers, percentages, dates, currencies, measurements, model numbers, addresses and identifiers form a factual integrity category. They are often copied rather than translated, which makes them easy to overlook during linguistic review.
Errors include transposed digits, changed decimal marks, wrong date order, incorrect conversion, missing unit, changed currency, altered version number and corrupted reference code.
This category should also include a value that is numerically copied correctly but presented under the wrong locale convention when that convention changes interpretation.
8. Negation and polarity error
Negation is important enough to track separately when teams work with high-risk content. A missing or misplaced negative can reverse a rule while leaving the target sentence fluent.
Examples include “not all” becoming “none,” “must not” becoming “does not have to,” or a negative condition attaching to the wrong clause. Double negatives and negative prefixes can also cause polarity drift.
Separating this category makes root-cause analysis easier because teams can build targeted QA checks for negative forms.
9. Modality and obligation error
This category covers must, should, may, might, can, required, permitted, recommended and comparable systems in other languages. The failure occurs when the target changes obligation, permission, possibility, probability or recommendation.
Modality errors are common because languages distribute force across verbs, adverbs, particles and context. Editors may also strengthen sentences unintentionally in the pursuit of confident prose.
A review note should state the source force and the target force: “Source permits; target requires.” That is much more actionable than “modal wrong.”
10. Condition and scope error
A condition error occurs when if, unless, except, only, provided that, before, after or another logical constraint attaches to the wrong action or loses part of its scope.
These errors frequently arise when long source sentences are split or reordered. The target may read better sentence by sentence while changing which rule applies to which case.
Reviewers should test the logic by describing the true and false cases. If the target produces different outcomes, the translation has changed the rule.
11. Reference and antecedent error
Reference errors occur when pronouns, demonstratives, labels or repeated noun phrases point to the wrong entity. The target can be grammatical and still assign responsibility to the wrong person or object.
Languages differ in pronoun omission, gender, number and repetition. A translator sometimes has to make explicit what the source leaves recoverable from context. That makes reference a genuine reasoning task.
This category is especially valuable in long documents because cross-sentence reference problems are easy to miss in segment-level review.
12. Logical connector error
Connectors encode relationships such as cause, result, contrast, concession, sequence and addition. Because, therefore, however, although, meanwhile and their equivalents are not interchangeable.
A target that changes “although” to simple “and” may remove the concession. A target that turns “after” into “because” changes sequence into causation.
Classifying these problems separately helps teams recognise that discourse logic is more than grammar.
13. Temporal relationship error
This category covers tense, aspect, chronology and duration when the target changes the timeline. The issue may involve whether an action is completed, ongoing, repeated, future, prior to another event or still true.
Not every tense mismatch is an error because languages organise time differently. Review the resulting temporal meaning, not the surface morphology.
A useful note is specific: “Target makes the approval precede the inspection; source says inspection happens first.”
14. Register error
Register concerns the social situation: formal versus informal, specialist versus general, distant versus familiar, deferential versus direct. A register error occurs when the target language places the message in the wrong social relationship for the brief.
A formal legal notice rendered as casual conversation is a register error. A friendly help message rendered in ceremonial prose is also a register error even if every proposition is accurate.
Reviewers need reference examples or style-guide rules. Otherwise register becomes personal taste disguised as quality control.
15. Tone or stance error
Tone overlaps register but can be tracked separately when emotional stance matters. A source may be reassuring, cautious, neutral, urgent, apologetic, celebratory or restrained.
An editor who makes a cautious sentence enthusiastic has changed more than style. In customer communication, crisis messaging, medical information or public notices, emotional stance can affect trust and behaviour.
Again, use evidence. “Target sounds more accusatory because it replaces the source’s conditional wording with direct blame” is actionable.
16. Grammar error
A grammar error violates target-language grammatical rules or produces a construction unacceptable for the intended standard. This includes agreement, case, syntax, morphology and sentence formation.
Do not use grammar as a catch-all label for any sentence you dislike. An unusual sentence can be grammatical. A stylistically weak sentence can be grammatical. The category should describe a real target-language structural problem.
Grammar errors often have low severity in casual content and high severity when they create ambiguity or damage professional credibility.
17. Collocation and idiomaticity error
A target can be grammatical yet sound unlike normal target-language usage because the words do not conventionally combine. Collocation problems include unnatural verb-noun, adjective-noun or preposition patterns.
This category is useful for separating “not how people say it” from grammar. Evidence can come from corpus patterns, authoritative domain texts or native editorial judgment.
Be careful with rare but legitimate phrasing. Frequency is evidence, not an absolute rule.
18. Clarity or readability error
Clarity errors arise when the target is unnecessarily difficult to understand because of sentence structure, information order, unclear reference or excessive compression, even though the core meaning is technically present.
This category should be tied to audience. A specialist paper can legitimately be dense. A safety instruction cannot rely on readers reconstructing a six-clause sentence under pressure.
Readability is not a licence to simplify technical content beyond accuracy. The job is to make the intended meaning accessible to the intended reader.
19. Locale and convention error
Locale errors include spelling variety, date format, decimal convention, punctuation style, quotation marks, address format, currency display and other region-specific norms defined by the project.
The sentence may be perfectly correct in another target locale and still be wrong for the specified one. That is why the taxonomy should treat locale as a project requirement rather than a universal grammar rule.
Locale categories are especially useful in global English, Spanish, Portuguese, French, Chinese and other languages with multiple regional standards.
20. Formatting and structural error
This category covers missing headings, list structure, emphasis, paragraph breaks, table alignment, links, footnotes and other document-level structures that affect meaning or usability.
A translated list may preserve all words while attaching a qualifier to the wrong bullet. A heading may be demoted into body text. A link may point to the source-language page. These are not merely cosmetic.
Formatting severity depends on function: a missing bold word may be negligible, while a broken table column can change which number belongs to which label.
21. Tag, placeholder or variable error
Digital translation often contains markup and dynamic variables. Missing braces, altered tag order, translated code tokens or misplaced placeholders can break rendering or change grammar around inserted content.
This category is highly automatable. QA tools can compare tag sets, placeholder counts and protected tokens, but human review is still required to confirm that surviving variables sit in grammatically correct positions.
Technical errors deserve visibility because a linguistically flawless translation that crashes a screen is still a failed deliverable.
22. Compliance or instruction error
A translation can violate explicit project requirements even when the language is otherwise good. Examples include translating text marked “do not translate,” failing to use an approved disclaimer, exceeding a character limit, omitting mandatory terminology or ignoring a required citation form.
This category keeps project instructions separate from universal language quality. A reviewer can say, “The translation is linguistically acceptable but does not comply with the brief.”
That distinction is fairer and more useful than labelling every instruction violation a linguistic error.
23. Consistency error
Consistency covers repeated content that should behave the same way across a document or product: headings, recurring phrases, UI actions, terminology, punctuation patterns and parallel procedures.
Do not demand consistency where context changes meaning. The same source word may require different translations in different senses. The same target phrase may need inflection. Consistency means stable decisions under stable conditions.
Good review notes identify the earlier approved pattern or style rule being violated.
24. Unsupported preference
This is an unusual but important category for review governance: a proposed change that does not correct a demonstrable error and is not required by the brief, termbase, style guide or target-language norm.
Recording unsupported preference does not mean reviewers cannot improve prose. It means the team distinguishes optional editing from objective defect correction. That distinction is essential when measuring quality or evaluating translators.
Without it, every reviewer can make a different stylistic version and report all changes as “errors,” making quality scores meaningless.
Severity: how much does the error matter?
After category, assign severity based on consequence. A useful scale can be simple: critical, major, minor and preference—or another set agreed by the team. The labels matter less than the decision rules.
A critical issue can cause serious harm, legal exposure, unsafe action, major financial loss, wrong identity, system failure or a fundamentally false message. A major issue materially changes meaning, task success, professional usability or a controlled requirement. A minor issue damages quality without materially changing the core message. A preference is a defensible alternative rather than a defect.
Severity should consider audience, content type and context. A missing comma in an essay may be minor; the same punctuation inside a numerical format may be major. A terminology variant in literary prose may be preference; in a dosage instruction it may be critical.
How to write a useful error note
A strong review comment contains four things: the location, the observed target problem, the source or rule it conflicts with, and the correction or desired outcome. For example: “Paragraph 6: target uses ‘recommended,’ but source states a mandatory requirement. Change to the approved obligation form in the style guide.”
This format reduces argument because the reviewer reveals the reasoning. It also creates training data for the team: future translators can learn the pattern rather than memorise one corrected sentence.
Root cause is not the same as error category
Two terminology errors can have different causes. One may come from an outdated glossary. Another may come from a translator who did not search the glossary. A number error may come from a source update that arrived after translation. A reference error may come from missing screenshot context.
Taxonomy describes the visible failure. Root-cause analysis asks why the system allowed it. Keep those layers separate. Otherwise teams blame translators for failures caused by poor source content, broken tools or contradictory instructions.
Common root causes include ambiguous source, missing context, inadequate domain knowledge, outdated resources, rushed schedule, inconsistent review, tool limitations, source change, copy-paste error and unclear ownership.
A practical error-review workflow
- Read enough context to understand the source and target purpose.
- Identify the concrete defect before choosing a category.
- Classify the error by the function it damages.
- Assign severity from consequence, not annoyance.
- Cite the source, terminology, style rule or target-language evidence.
- Propose a correction or outcome.
- Record recurring patterns separately for root-cause analysis.
- Do not count reviewer preferences as translator errors.
- Calibrate difficult examples with multiple reviewers.
- Periodically simplify categories that nobody can apply consistently.
Calibration drills
Drill 1: “Could” becomes “must”
Category: modality and obligation. Severity depends on context. In a casual suggestion it may be major; in policy, legal or safety text it may be critical because permission or possibility has become requirement. Root cause might be an editor strengthening prose rather than a translator misunderstanding vocabulary.
Drill 2: an official institution name is translated literally
Category: name and entity. The sentence may remain understandable, but the target no longer uses the institution’s official identity. Severity rises where legal recognition, search, credentials or public trust matter.
Drill 3: “20% increase” becomes “increase to 20%”
Category: number/factual relationship or meaning transfer. The digits are preserved but the mathematical relation changes. This is why number checking cannot be reduced to matching numerals.
Drill 4: two acceptable synonyms are used for a marketing adjective
Probably not an error unless the style guide requires one form. If the change does not create a conceptual distinction, it may be acceptable variation or preference. A taxonomy should protect translators from false precision as well as identify real defects.
Drill 5: a pronoun points to the wrong agency
Category: reference and antecedent. If responsibility is reassigned, severity may be major or critical. The root cause could be ambiguous source grammar, especially if two agencies were possible antecedents.
Drill 6: the target is accurate but far too formal for a children’s app
Category: register, possibly clarity. The problem is not semantic transfer but audience fit. The reviewer should refer to the brief or style guide rather than claiming universal linguistic incorrectness.
Drill 7: a placeholder is deleted
Category: tag, placeholder or variable. If the product crashes or shows blank information, severity can be critical. This is a strong candidate for automated QA because the token can often be checked mechanically.
Drill 8: a translator adds a helpful explanation not present in a historical quotation
Category: addition. Even if the explanation is factually correct, it changes the documentary record unless the project explicitly permits explanatory annotation.
Drill 9: a sentence is grammatical but uses an improbable collocation
Category: collocation and idiomaticity. Severity is usually minor unless the odd phrasing causes misunderstanding or damages a high-visibility brand message.
Drill 10: a date 03/04/2026 is copied unchanged into a locale that reads the order differently
Category: locale and factual-token integrity. The digits match but reader interpretation may not. A safer target may spell the month or follow the declared locale convention.
Drill 11: “although” becomes “therefore”
Category: logical connector. The relationship moves from concession to result. This is a semantic failure even if both sentences are individually plausible.
Drill 12: reviewer changes “begin” to “start” with no rule or contextual reason
Category: unsupported preference, not translator error. The change can still be accepted if it improves flow, but it should not distort quality metrics.
How error data should be used
Error counts are valuable only when they lead to better decisions. If terminology errors recur, update the termbase or training. If reference errors cluster in one source type, improve source preparation. If placeholder failures recur, add automated checks. If reviewers disagree constantly about register, strengthen the style guide and calibration examples.
Do not use raw counts without exposure. A 100,000-word project will naturally have more opportunities for error than a 1,000-word project. Do not compare translators across radically different domains without context. Do not treat every minor punctuation issue as equivalent to a changed medical instruction.
The purpose of taxonomy is learning and control, not the production of impressive-looking numbers.
Frequently asked questions
How many categories should a taxonomy have?
Enough to separate recurring failure types that require different corrective actions, but not so many that reviewers spend more time choosing labels than explaining problems. Start compact and split categories only when the distinction changes decisions.
Should grammar and style be separate?
Usually yes. Grammar concerns target-language structural correctness. Style concerns appropriateness, clarity, voice and conventions within the brief. Mixing them encourages subjective edits to be reported as hard errors.
Should terminology inconsistency be its own category?
It can be useful when consistency is operationally important. Wrong term and inconsistent term have different causes and remedies, so separating them often improves analysis.
Can one issue belong to two categories?
Yes, but choose a primary category when possible to avoid double-counting. A wrong date format can be both locale and factual integrity; select the category that best explains the failure and record the secondary observation in the note.
How should reviewer preferences be handled?
Keep them visible but separate from defect counts. Optional edits can improve prose without implying that the original translator made an error.
Can AI classify translation errors?
AI can help suggest categories and find mechanical patterns, but reliable classification still requires source understanding, project rules and consequence. Automated confidence should never replace evidence in high-risk decisions.
Should error taxonomy be shown to translators?
Yes. A hidden scoring system encourages surprise and defensiveness. Shared categories let translators understand expectations, self-review against the same framework and challenge misclassification constructively.
What is the most important rule?
Name the observed failure before judging the translator. A precise category should clarify the problem, not provide a new label for blame.
Concluding idea
A translation error taxonomy is a map of failure modes. Its value is not in the labels themselves but in the distinctions they force a team to make: wrong meaning is not awkward style; terminology is not ordinary vocabulary; factual integrity is not punctuation; register is not personal taste; severity is not category; root cause is not blame.
When those distinctions are clear, review becomes teachable. Translators can improve specific habits. Reviewers can justify changes. Project managers can target process fixes. Content owners can see when the source caused the problem. Quality stops being a vague impression and becomes a conversation about evidence, consequence and better systems.
Continue through the wider system at Master Art of Translation | The Complete System for Moving Meaning Between Languages.
