VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Master Art of Translation | The Translation Quality System — How Specifications, Error Severity, Evaluation, Revision and Release Confidence Work Together

EDKSG-TRANS-MASTER-WORLD-090

Translation quality is one of the most searched and most easily oversimplified questions in professional language work. People search for translation quality, translation quality assurance, translation evaluation, translation accuracy, translation error types, MQM, translation revision, bilingual review, proofreading, quality scores, translation quality metrics and how to know whether a translation is good because they want a reliable way to separate an acceptable target text from one that merely sounds fluent. The real answer is not a single score. Translation quality is a system of specifications, evidence, error detection, severity, risk, human judgment, automated checks and release decisions.

A good translation-quality system begins before translation. It defines what the translation is for, who will read it, what risks follow from error, which terminology is authoritative, what style and locale conventions apply, which source details must remain invariant, and what level of review is required. ISO 5060:2024 provides current international guidance for evaluating human translation, post-edited machine translation and unedited machine translation output through an analytic approach based on error types, penalty points, error scores and quality ratings. The European Commission’s current translation-quality framework similarly treats quality as accurate, clear and fit for purpose, with quality controls selected according to document type, risk profile, intended use and reader expectations.

The hidden problem is that quality cannot be inferred from fluency alone. A polished target can still mistranslate a legal obligation, omit a medical condition, change a percentage, choose the wrong person for a pronoun, flatten uncertainty, misuse a technical term or break an interface placeholder. Equally, a stiff sentence can be accurate but fail its target reader. The translation-quality system therefore has to answer two questions at once: did the target preserve what mattered in the source, and does the target function correctly for its intended reader and use?

What this article owns in the Master Art of Translation architecture

This node owns the complete architecture of translation quality.

It covers:

quality specifications,

fit for purpose,

quality risk,

translation error taxonomies,

error severity,

MQM-style analytic evaluation,

ISO 5060 evaluation principles,

accuracy,

completeness,

terminology,

target-language quality,

locale,

format,

technical integrity,

revision,

review,

sampling,

automated QA,

quality scoring,

quality gates,

release confidence,

correction,

incident learning,

and continuous improvement.

It does not replace the Human Translation System.

That node owns the people and professional workflow that produce and revise target text.

It does not replace the Machine Translation System.

That node owns non-human generation, post-editing and automation controls.

It does not replace the Terminology System, Translation Memory System, Context Stack or Source Analysis.

Those systems supply evidence.

The Translation Quality System evaluates whether the output built from that evidence is fit to use.

Quality is not one thing

Translation quality has several dimensions.

A target can be:

accurate but unnatural,

natural but inaccurate,

terminologically correct but incomplete,

complete but tonally wrong,

linguistically excellent but technically broken,

or mechanically correct but unfit for its reader.

This is why one global label such as “good translation” is too coarse.

A useful quality architecture separates dimensions.

Fit for purpose

“Fit for purpose” is one of the strongest practical ideas in translation quality.

The same text can require different quality priorities depending on use.

Internal gist translation:

meaning enough for understanding.

Public help article:

accuracy plus clarity and usability.

Legal clause:

accuracy, scope, defined terms and legal force.

Medical instruction:

accuracy, dosage, conditions and safety.

Marketing slogan:

brand effect plus factual constraint.

Literary passage:

meaning, voice, imagery, rhythm and ambiguity.

Quality is always quality for a purpose.

Specifications come before evaluation

You cannot evaluate a translation fairly unless you know what it was supposed to do.

Specifications may include:

source language,

target language,

target locale,

audience,

purpose,

domain,

genre,

terminology,

style guide,

format,

risk class,

delivery channel,

review level,

and acceptance threshold.

Without specifications, reviewer disagreement increases.

Why vague quality requirements fail

“Make it accurate.”

“Make it natural.”

“Use professional language.”

These are useful aspirations.

They are not operational specifications.

What does accurate mean for:

a poem?

a safety label?

a user interface?

a financial report?

The quality system needs enough detail to judge the output.

Risk-based quality

The consequences of error vary.

Low-risk error:

awkward phrase in an internal note.

Medium-risk error:

confusing instruction on a website.

High-risk error:

wrong eligibility condition.

Critical error:

wrong medical dose.

Quality resources should follow risk.

This is why the European Commission’s current framework explicitly links quality control to document type, risk profile, use and reader needs.

Risk is not difficulty

A simple sentence can be high risk.

“Take 5 mL.”

Few words.

Potentially serious consequence.

A complex literary paragraph may be difficult but lower operational risk.

Difficulty and risk should be tracked separately.

Quality risk model

A conceptual model:

quality risk = likelihood of error × consequence of error × exposure.

Likelihood:

How hard is the text?

How reliable is the workflow?

Consequence:

What happens if wrong?

Exposure:

How many people will see or act on it?

This is not a literal universal formula.

It is a decision framework.

Quality assurance versus quality evaluation

These terms are related but not identical.

Quality assurance is the system of preventive and checking processes used to reduce defects.

Quality evaluation assesses translation output against defined criteria.

ISO 5060:2024 focuses on evaluation of translation output, not the broader process of quality assurance or corrective action.

That distinction is useful.

QA asks:

How do we prevent and catch errors?

Evaluation asks:

How good is this output according to a defined method?

Preventive quality

Preventive quality happens before or during translation.

Examples:

clear source,

strong brief,

qualified translator,

termbase,

style guide,

approved TM,

context,

training,

and correct file handling.

The cheapest error is one never created.

Detective quality

Detective quality finds errors after they appear.

Examples:

self-revision,

bilingual revision,

monolingual review,

subject expert review,

automated QA,

and in-context testing.

A mature system uses both prevention and detection.

Corrective quality

Corrective quality responds after a defect is found.

Fix the target.

Fix reusable resources.

Fix process.

Train people.

Update tests.

Quality management becomes a feedback loop.

The quality loop

Specify.

Produce.

Check.

Evaluate.

Release.

Observe.

Correct.

Learn.

Then specify better next time.

Translation quality is cyclical.

Accuracy

Accuracy asks whether the target preserves source meaning.

Check:

actors,

actions,

objects,

time,

conditions,

negation,

modality,

quantity,

attribution,

certainty,

and reference.

Accuracy is not literal word matching.

Completeness

Completeness asks whether all required source content appears.

Check:

sentences,

paragraphs,

headings,

lists,

captions,

footnotes,

tables,

labels,

and notes.

Fluent omission remains omission.

Addition

A quality system also checks unsupported addition.

Added cause.

Added certainty.

Added gender.

Added explanation.

Added example.

Added legal specificity.

Some explicitation may be legitimate.

It should be justified.

Terminology

Terminology quality asks whether domain concepts use approved and consistent target terms.

Check:

preferred terms,

defined terms,

forbidden variants,

abbreviations,

capitalisation,

and concept consistency.

A translation can be accurate in general language and wrong professionally because terminology is off.

Target-language quality

Target-language quality asks whether the target is competent language.

Grammar.

Syntax.

Collocation.

Punctuation.

Spelling.

Cohesion.

Register.

Naturalness.

The target should not require the reader to reconstruct source grammar mentally.

Locale quality

A translation can be valid in the language and wrong for the locale.

Check:

spelling,

date format,

number format,

currency,

quotation marks,

address format,

official institutions,

and local terminology.

Style

Style quality asks whether the target matches project voice.

Formal.

Conversational.

Technical.

Academic.

Plain language.

Brand voice.

Style errors can be important even when semantic meaning survives.

Register

Register includes social relation and situational language level.

A target can become:

too casual,

too formal,

too intimate,

too bureaucratic,

or too technical.

Quality includes appropriateness.

Tone

Tone can be:

reassuring,

urgent,

cautious,

critical,

empathetic,

humorous.

A target that preserves facts but loses tone may fail communication.

Technical integrity

Structured translation must preserve:

tags,

variables,

placeholders,

IDs,

code,

links,

format markers,

and functional strings.

A linguistically perfect target can be technically unusable.

Formatting

Check:

headings,

lists,

tables,

captions,

footnotes,

page order,

and cross-references.

Format can encode relationships.

Usability

A translation can be accurate but hard to use.

Can the reader:

complete the form?

follow the procedure?

find the setting?

understand the warning?

Quality can be functional.

Accessibility

Check:

clear labels,

screen-reader strings,

captions,

alt text,

reading order,

plain-language needs,

and accessible formatting.

Quality includes access.

Translation error taxonomy

An error taxonomy gives names to recurring defect types.

Useful top-level categories can include:

accuracy,

terminology,

linguistic conventions,

style,

locale,

format,

and technical integrity.

MQM is a widely used framework for analytic translation-quality evaluation and can be configured with error types and scoring models.

The exact taxonomy should fit the project.

Why categories matter

If reviewers only say:

“bad translation”

the organisation cannot learn.

If they classify:

mistranslation,

omission,

wrong term,

number error,

register error,

tag error

then root causes become visible.

Category is not severity

A terminology error can be minor.

Or critical.

A punctuation error can be minor.

Or change a number.

Keep type and consequence separate.

Severity levels

A practical system may use:

critical,

major,

minor.

Some frameworks use additional levels.

The important part is defining them.

Critical error

A critical error can cause:

harm,

wrong legal action,

clinical risk,

security failure,

financial loss,

denial of rights,

or severe misinformation.

Critical errors can fail a translation regardless of overall fluency.

Major error

A major error materially affects:

meaning,

instruction,

term,

reader action,

or important content.

It may not create catastrophic consequence.

Minor error

A minor error does not materially change core meaning or function but reduces quality.

Examples:

punctuation,

small fluency issue,

non-critical inconsistency.

Neutral or preference

Some changes are not errors.

Two translations can both be acceptable.

Review systems should avoid converting preference into defect count.

This is essential for fair evaluation.

MQM thinking

MQM encourages analytic quality evaluation using defined error categories and severity.

The strength is not that one universal score solves translation.

The strength is that evaluation becomes explicit.

What kind of error?

How serious?

Where?

Why?

A metric can then aggregate according to project rules.

Scoring

A scoring model may assign penalty points.

For example:

critical = fail,

major = larger penalty,

minor = smaller penalty.

The exact weights vary.

Do not treat one score formula as universal truth.

Error score versus quality rating

A raw error score can be converted into a rating.

Pass.

Fail.

Excellent.

Acceptable.

Needs revision.

The rating should map to project acceptance criteria.

Normalisation

Long documents naturally contain more opportunities for errors than short ones.

Scores can be normalised by:

word count,

segments,

or other units.

Choose a method consistently.

Sampling

ISO 5060 explicitly discusses sampling.

Sampling is useful when full evaluation is expensive.

A sample can estimate quality.

But sample design matters.

Random sampling

Random sampling gives broad representation.

It can miss rare critical sections.

Risk-based sampling

Target:

legal clauses,

warnings,

numbers,

new terminology,

new translator sections,

or low-confidence machine output.

This increases sensitivity to consequence.

Hybrid sampling

Combine:

random sample,

risk sample,

and known-problem sample.

This is often more useful than one method alone.

Sample size

The right sample depends on:

document size,

risk,

quality history,

and decision consequence.

Do not use one fixed percentage blindly.

Sampling bias

If reviewers always sample easy introduction text, quality estimates become optimistic.

Sample across:

beginning,

middle,

end,

tables,

lists,

and difficult sections.

Evaluation unit

Decide what is being evaluated.

Word?

Segment?

Sentence?

Document?

Project?

Different units support different questions.

Segment-level evaluation

Useful for precise error tagging.

Risk:

losing paragraph context.

Document-level evaluation

Useful for:

cohesion,

style,

voice,

and functionality.

Harder to score consistently.

A mature system uses both where needed.

Reference translation evaluation

Automated metrics often compare output against reference translation.

Human evaluation does not need one single reference.

Several target solutions may be valid.

Human evaluator competence

ISO 5060 also addresses evaluator competence.

Evaluators need enough:

source understanding,

target-language skill,

translation knowledge,

domain knowledge,

and rubric understanding

to judge fairly.

The evaluator is part of measurement.

Reviewer calibration

Two evaluators can label the same issue differently.

Calibration reduces variance.

Use shared examples.

Discuss category.

Discuss severity.

Document rulings.

Inter-rater agreement

When formal evaluation matters, compare reviewer agreement.

Large disagreement means:

rubric unclear,

training weak,

or categories too ambiguous.

Evaluation quality itself needs QA.

Error-location evidence

A useful evaluation record includes:

source excerpt,

target excerpt,

error category,

severity,

comment,

and proposed correction.

This makes the finding auditable.

Reviewer rationale

“Wrong” is weak feedback.

“Changes possibility to certainty” is useful.

“Violates approved termbase entry X” is useful.

“Official target institution name is Y” is useful.

Quality evidence should teach.

Translation quality and specifications

A target cannot be judged solely against general language norms.

If the brief requires:

formal address,

then informal address is a defect.

If the brief requires:

plain language,

then unnecessary jargon is a defect.

Specifications operationalise quality.

Specification hierarchy

Possible authority order:

law/regulation,

client requirement,

project brief,

termbase,

style guide,

approved reference,

general language preference.

Conflicts should be escalated.

Quality gate architecture

Gate 1:

source accepted.

Gate 2:

brief complete.

Gate 3:

resources current.

Gate 4:

translation complete.

Gate 5:

self-revision complete.

Gate 6:

required revision/review complete.

Gate 7:

formal QA passed.

Gate 8:

evaluation threshold met.

Gate 9:

in-context check complete.

Gate 10:

release approved.

Quality gates convert policy into workflow.

Gate 1: source quality

Is the source complete?

Readable?

Correct version?

If not, translation evaluation later becomes unfair.

Gate 2: specifications

Audience?

Purpose?

Locale?

Risk?

Acceptance threshold?

Without them, evaluation is subjective.

Gate 3: resources

Current:

termbase,

style guide,

TM,

reference.

Stale resources can create systematic errors.

Gate 4: completeness

No missing translation.

No hidden content.

No skipped tables.

Gate 5: self-revision

Translator confirms first quality pass.

Gate 6: independent quality control

Revision or review chosen according to risk.

Gate 7: automated QA

Check mechanical defects.

Gate 8: evaluation

If formal sampling or score required, perform it.

Gate 9: context

Test target where used.

Gate 10: release confidence

All required evidence available?

Then approve.

Release confidence

Release confidence is not certainty that no error exists.

It is justified confidence that:

required process completed,

known risks were controlled,

evaluation threshold met,

and critical issues were resolved.

This is a more realistic quality goal than “perfect.”

Zero-defect myth

No human system can guarantee zero errors in all translation.

A quality system should aim for:

appropriate error prevention,

detection,

and consequence control.

Perfection claims are less useful than transparent process.

High-quality does not mean maximum review

If a low-risk internal email receives five reviewers, resources are wasted.

Quality is efficient fit for purpose.

Under-review

If a critical medical warning receives only spell-check, process is inadequate.

Over-review

If every stylistic alternative is debated for hours in a temporary internal note, process is inefficient.

Right-sized review

Match quality control to:

risk,

lifespan,

audience,

and publication impact.

Quality assurance automation

Automated QA can reliably detect many formal defects.

Numbers.

Tags.

Placeholders.

Missing punctuation.

Repeated spaces.

Terminology.

Forbidden terms.

Use automation for repeatable rules.

Automated QA limits

Automation can flag:

“must” missing

only if configured.

It may not understand whether target modal force is appropriate.

Human judgment remains necessary for semantics.

Number QA

Compare source and target numbers.

Flag:

missing,

added,

changed,

or formatting differences.

Then human checks whether transformation was intentional.

Date QA

Dates may legitimately reformat.

QA should flag difference, not automatically call it error.

Terminology QA

Flag missing preferred terms or forbidden variants.

But grammar can require inflection.

Rules must fit target language.

Tag QA

Check:

paired tags,

order,

missing tags,

extra tags.

Technical errors are machine-detectable.

Placeholder QA

Source and target should preserve required variables.

Check set equality where appropriate.

Length QA

Large length anomalies can indicate:

omission,

addition,

or wrong segmentation.

Length is a clue, not proof.

Repeated-target QA

Different source segments sharing identical target may indicate overgeneralisation.

Sometimes legitimate.

Use as diagnostic.

Inconsistency QA

Same source repeated with different targets.

May be error.

May be contextual variation.

Human review decides.

Spell-check

Necessary.

Not sufficient.

A correctly spelled wrong word remains wrong.

Grammar-check

Helpful for target language.

Can introduce incorrect “fixes” in specialist or creative text.

Review suggestions.

AI-assisted QA

AI can help find:

possible omissions,

tone mismatch,

and inconsistent reasoning.

It can also hallucinate defects.

Use as assistant, not sole authority.

Human revision

Bilingual revision remains central for source fidelity.

The European Commission defines revision as a bilingual second-pair-of-eyes comparison of translation and source for accuracy and completeness.

Human review

The Commission distinguishes review as monolingual attention to clarity, tone and audience suitability.

This division is useful.

Subject review

Domain expert verifies concept and procedure.

Not every project needs it.

High-risk specialist work often benefits.

Proofreading

Final surface and production check.

Do not rely on proofreading to detect mistranslation.

In-context linguistic review

Evaluate final rendered product.

UI.

PDF.

Website.

Subtitle.

Form.

Context can reveal errors invisible in bilingual segments.

Quality incident

A quality incident occurs when translation defect reaches production and matters enough to require response.

Examples:

wrong legal condition,

unsafe warning,

wrong name,

incorrect amount,

broken UI.

Incident response

Contain.

Correct.

Verify.

Propagate correction.

Analyse root cause.

Update system.

Root causes

Source problem.

Translator error.

Machine error.

Termbase gap.

TM contamination.

Reviewer miss.

QA configuration.

File-processing error.

Deployment error.

Quality management should fix root causes.

Corrective action

If wrong term:

fix termbase.

If repeated number errors:

improve number QA.

If reviewer inconsistency:

calibrate.

If source ambiguity:

improve authoring.

If wrong TM:

fix resource assignment.

Preventive action

Use incident evidence to prevent recurrence.

Quality matures through failures that are studied rather than hidden.

Quality metrics

Possible metrics:

critical error rate,

major errors per thousand words,

terminology adherence,

QA warnings,

review edit rate,

rework rate,

incident rate,

reader task success.

No metric should dominate blindly.

Speed metric

Words per hour can matter operationally.

Do not infer quality from speed alone.

Acceptance rate

High reviewer acceptance can indicate good translation.

Or superficial review.

Interpret with error outcomes.

Edit distance

Low edit distance can indicate strong initial quality.

Or reviewer restraint.

Use context.

Query rate

Useful queries can improve quality.

Do not punish translators for identifying ambiguity.

Incident rate

Low incident rate is good.

But under-reporting can make it look artificially low.

Create safe reporting culture.

Customer complaints

Complaints are useful evidence.

Not every complaint is a translation error.

Classify.

Reader comprehension

For public communication, comprehension testing can be more meaningful than error count alone.

Task success

Can users complete:

form,

navigation,

procedure?

Translation quality can be observed in behaviour.

Search success

For multilingual websites, target readers should be able to find content.

SEO localisation affects access but should not distort meaning.

Quality and cost

Higher review levels cost more.

The question is whether cost matches risk.

Do not compare workflows only by price per word.

Cost of error

Quality investment should consider downstream error cost.

One critical translation error can cost more than extensive review.

Cost of over-quality

Excessive review of low-risk content also has cost:

delay,

staff time,

and reduced capacity for high-risk work.

Optimal quality

Optimal quality means the right quality for purpose with efficient risk control.

Translation quality for machine output

ISO 5060 applies evaluation guidance to:

human translation,

post-edited MT,

and unedited MT.

The same core question remains:

Does the output satisfy defined requirements?

Provenance can affect expected workflow.

Output evaluation still looks at the target.

Raw MT quality

Raw MT may pass for low-risk gist.

It may fail for publication.

Use case matters.

Post-edited MT quality

Full post-editing can target professional publication quality.

Evaluation should assess final output, not punish it merely because MT contributed.

Human translation quality

Human origin does not guarantee quality.

Evaluate against the same task requirements.

Quality equality principle

Do not assume:

human = good

machine = bad.

Evaluate output and process evidence.

Risk and provenance inform review.

Error severity and automation

A system may automatically fail any translation containing:

critical error.

Major-error threshold may trigger fail.

Minor-error threshold may influence rating.

Rules should be documented.

Penalty weights

If major = 5 and minor = 1, that is a project scoring choice.

Do not present arbitrary weights as universal truth.

Pass/fail threshold

Set threshold before evaluation where possible.

Changing threshold after seeing results undermines measurement.

Gold standard

There is rarely one perfect target sentence.

Quality evaluation should allow legitimate variation.

Error tolerance

Different projects tolerate different minor defects.

Critical errors may have zero tolerance.

Define policy.

Translation evaluation report

A useful report contains:

scope,

sample method,

rubric,

categories,

severity,

score,

critical findings,

examples,

and decision.

Executive summary

For managers:

pass/fail,

critical risks,

systemic issues,

recommended action.

Do not drown decision-makers in every comma.

Linguist report

For translators:

detailed examples,

rationale,

patterns,

and resource updates.

Quality dashboard

Useful signals:

critical errors,

major error rate,

top categories,

language pairs,

vendors,

domains,

trend.

Avoid vanity metrics.

Trend analysis

One month of data may be noise.

Look for recurring patterns.

Vendor comparison

Compare only similar:

content,

risk,

language pair,

and review conditions.

Otherwise ranking is unfair.

Translator comparison

Quality metrics can support development.

Do not reduce people to one score.

Calibration data

Use anonymised examples to train reviewers and translators.

Quality culture

A strong quality culture rewards:

useful queries,

early risk detection,

and correction reporting.

It does not reward hiding errors.

Blame-free root cause

Blame-free does not mean no accountability.

It means investigate system conditions before simplistic punishment.

Accountability

People still own decisions.

Systems should make good decisions easier.

Quality and specifications: the deepest connection

Most quality disputes are specification disputes in disguise.

One reviewer expects formal language.

Another expects conversational language.

One client wants literal titles.

Another wants SEO adaptation.

If expectations are not explicit, evaluation becomes argument.

The specification-first rule

Before translation starts, decide:

what counts as success.

Then quality evaluation becomes much fairer.

The quality contract

A project-quality contract can include:

content type,

audience,

risk,

required accuracy,

terminology source,

style source,

review steps,

acceptance threshold,

and correction process.

This can be formal or lightweight.

Quality architecture for a small team

Minimum:

brief,

competent translator,

self-revision,

QA,

spot review according to risk,

correction loop.

Quality architecture for a large enterprise

Specifications,

risk classification,

qualified linguists,

termbases,

TMs,

automated QA,

revision/review rules,

formal evaluation,

language leads,

incident response,

dashboards,

continuous improvement.

Quality architecture for education

Teach learners to separate:

meaning accuracy,

target naturalness,

and mechanics.

Do not give one opaque grade.

Quality architecture for translation training

Students need error categories and severity.

They should justify why an error matters.

Quality architecture for AI-era translation

AI increases output volume.

Quality systems must scale evaluation intelligently.

Risk routing.

Sampling.

Automated checks.

Human expert attention.

Quality architecture becomes more important as generation becomes cheaper.

Translation-quality diagnostic laboratory

The fastest way to understand quality is to examine outputs that look acceptable at first glance but fail under a more disciplined evaluation. Each case below separates error type from severity and shows why one global “good/bad” judgment is too crude.

Diagnostic 1: missing negation

Source:

“Do not disconnect the device.”

Target:

“Disconnect the device.”

Category:

accuracy — mistranslation.

Severity:

critical in a safety context.

Why:

the target reverses required action.

A scorecard should not allow twenty minor style points to hide one critical reversal.

Diagnostic 2: modal change

Source:

“Users may cancel at any time.”

Target:

“Users must cancel at any time.”

Category:

accuracy — modality.

Severity:

major or critical depending on legal consequence.

The sentence remains grammatical.

Meaning does not.

Diagnostic 3: hedge strengthened

Source:

“The results may indicate…”

Target:

“The results prove…”

Category:

accuracy — overtranslation / epistemic force.

Severity:

major in research.

The target invents certainty.

Diagnostic 4: omission of exception

Source:

“All employees except contractors must complete training.”

Target omits “except contractors.”

Category:

accuracy — omission.

Severity:

major.

Scope changes.

Diagnostic 5: wrong actor

Source:

“The regulator instructed the company to revise its filing.”

Target assigns revision to the regulator.

Category:

accuracy — relationship.

Severity:

major.

Role structure changed.

Diagnostic 6: pronoun ambiguity resolved incorrectly

Source:

“Amira told Sofia that she had passed.”

Target identifies Amira as the person who passed.

Context later shows Sofia.

Category:

accuracy — reference.

Severity:

major in a narrative or record.

Diagnostic 7: number changed

Source:

“0.5 mg”

Target:

“5 mg”

Category:

accuracy — number.

Severity:

critical in medicine.

This is why hard-detail errors need independent handling.

Diagnostic 8: unit changed

Source:

“5 mg”

Target:

“5 g”

Category:

accuracy — unit.

Severity:

critical.

Digit identity alone is not quantity fidelity.

Diagnostic 9: percentage point changed to percent

Source:

“rose by 5 percentage points”

Target:

“rose by 5%”

Category:

accuracy — quantitative relation.

Severity:

major in finance or statistics.

Diagnostic 10: date ambiguous

Source:

“03/04/2026”

Target chooses March 4.

Source locale intended 3 April.

Category:

accuracy / locale.

Severity:

major if deadline-related.

Diagnostic 11: currency symbol localised incorrectly

Source:

“SGD 100”

Target:

“$100”

Target locale interprets $ as USD.

Category:

accuracy / locale.

Severity:

major financial.

Diagnostic 12: defined term varied

Source legal text defines “Supplier.”

Target alternates among three synonyms.

Category:

terminology.

Severity:

major if conceptual distinction is created.

Diagnostic 13: preferred term ignored

Current termbase requires “data breach.”

Target uses deprecated equivalent.

Category:

terminology.

Severity:

minor, major or critical depending on policy.

Diagnostic 14: term correct, grammar wrong

Approved technical noun inserted in wrong case.

Category:

linguistic convention plus terminology integration.

Severity:

minor or major depending on comprehensibility.

Diagnostic 15: proper name translated

Brand or institution translated literally despite official target form.

Category:

accuracy — entity / terminology.

Severity:

major if identity is altered.

Diagnostic 16: transliteration variant not authorised

Person’s name receives a new romanisation.

Category:

entity accuracy.

Severity:

major in official records.

Diagnostic 17: quotation altered

Target changes wording inside quotation marks.

Category:

accuracy — quotation.

Severity:

major in journalism, evidence or scholarship.

Diagnostic 18: indirect speech turned into direct quotation

Target adds quotation marks to a paraphrase.

Category:

accuracy — attribution.

Severity:

major.

Diagnostic 19: source attribution removed

Source:

“Officials said the measure may…”

Target states the proposition as fact.

Category:

accuracy — attribution and certainty.

Severity:

major.

Diagnostic 20: causation invented

Source:

“Complaints increased after the change.”

Target:

“Complaints increased because of the change.”

Category:

accuracy — logical relation.

Severity:

major.

Diagnostic 21: word sense wrong but sentence fluent

“Charge” in electricity translated as fee.

Category:

accuracy — lexical sense.

Severity:

major in technical content.

Diagnostic 22: false friend

Related-language cognate selected incorrectly.

Category:

accuracy — lexical.

Severity:

depends on consequence.

Diagnostic 23: collocation unnatural

Every word individually correct.

Target phrase not used naturally.

Category:

linguistic convention — collocation.

Severity:

minor unless it impairs understanding.

Diagnostic 24: source syntax copied

Target reads like source grammar.

Category:

linguistic convention — grammar / syntax.

Severity:

minor to major.

Diagnostic 25: register too casual

Formal public notice uses intimate slang.

Category:

style / register.

Severity:

major if it damages institutional suitability.

Diagnostic 26: register too formal

Children’s instructions become bureaucratic.

Category:

style / audience fit.

Severity:

major if readers cannot understand.

Diagnostic 27: tone flattened

Empathetic support message becomes cold.

Category:

style / tone.

Severity:

minor or major depending on service context.

Diagnostic 28: sarcasm lost

Source ironic praise becomes sincere praise.

Category:

accuracy / pragmatics.

Severity:

major in dialogue.

Diagnostic 29: idiom translated literally

Target phrase meaningless.

Category:

accuracy / idiom.

Severity:

major.

Diagnostic 30: metaphor flattened unnecessarily

Literary image becomes plain explanation.

Category:

style / accuracy of effect.

Severity:

major in literary translation.

Diagnostic 31: glossary variation creates false distinction

One concept translated with two near-synonyms.

Reader assumes two concepts.

Category:

terminology consistency.

Severity:

major in technical text.

Diagnostic 32: repeated source translated inconsistently

Same UI string appears with two targets in same function.

Category:

consistency.

Severity:

minor to major.

Diagnostic 33: different source translated identically

Two button actions become same target, causing ambiguity.

Category:

accuracy / functional distinction.

Severity:

major.

Diagnostic 34: locale spelling wrong

en-GB project uses en-US spellings.

Category:

locale.

Severity:

minor unless brand/legal policy says otherwise.

Diagnostic 35: decimal separator wrong

1.5 displayed as 1,5 under wrong locale policy.

Category:

locale / number.

Severity:

can be critical.

Diagnostic 36: date format wrong

Target language correct, locale convention wrong.

Category:

locale.

Severity:

minor or major depending on ambiguity.

Diagnostic 37: quotation marks wrong

Language-specific typography ignored.

Category:

locale / punctuation.

Severity:

usually minor.

Diagnostic 38: punctuation changes meaning

Missing comma creates ambiguity or changes legal scope.

Category:

accuracy or linguistic convention.

Severity:

based on consequence.

Diagnostic 39: missing list item

Everything else perfect.

Category:

completeness — omission.

Severity:

major if instruction lost.

Diagnostic 40: missing footnote

Category:

completeness / format.

Severity:

major if footnote qualifies claim.

Diagnostic 41: table row misaligned

Target labels shifted one row.

Category:

format / accuracy.

Severity:

major or critical.

Diagnostic 42: broken placeholder

{username} becomes translated text.

Category:

technical integrity.

Severity:

major because software can fail.

Diagnostic 43: tag missing

HTML close tag omitted.

Category:

technical integrity.

Severity:

major.

Diagnostic 44: link text correct, URL wrong

Category:

technical / functional.

Severity:

major.

Diagnostic 45: subtitle too long

Meaning accurate but unreadable in available time.

Category:

functional quality.

Severity:

major for audiovisual use.

Diagnostic 46: UI text truncates

Button label clipped.

Category:

functional / layout.

Severity:

major if action unclear.

Diagnostic 47: alt text omitted

Visual content inaccessible.

Category:

accessibility / completeness.

Severity:

major for required accessibility.

Diagnostic 48: form label too vague

Source field distinction disappears.

Category:

functional accuracy.

Severity:

major.

Diagnostic 49: SEO keyword forced unnaturally

Target search phrase stuffed into prose.

Category:

style / functional quality.

Severity:

minor to major.

Diagnostic 50: SEO keyword literal but wrong intent

Page ranks for irrelevant target query.

Category:

functional localisation.

Severity:

major for discoverability.

Error type versus source of error

An evaluator should distinguish what went wrong from why it went wrong.

Error type:

wrong number.

Possible root cause:

translator typo,

OCR corruption,

machine output,

source ambiguity,

or reviewer edit.

Evaluation describes output.

Root-cause analysis improves process.

Why this separation matters

If every wrong number is blamed on translators, an OCR pipeline bug remains.

If every terminology issue is blamed on MT, a stale termbase remains.

If every style issue is blamed on reviewer preference, the brief remains vague.

Quality management needs both layers.

The evaluation record

A strong evaluation entry can include:

segment ID,

source excerpt,

target excerpt,

category,

subcategory,

severity,

comment,

recommended correction,

specification violated,

and evaluator.

This makes findings inspectable.

Error category design

Do not build a taxonomy with hundreds of categories nobody can apply consistently.

Start with a manageable core.

Expand only where decisions require it.

Minimal error taxonomy

Accuracy.

Terminology.

Target language.

Style.

Locale.

Format / technical.

This can be enough for many teams.

Expanded accuracy taxonomy

Mistranslation.

Omission.

Addition.

Untranslated content.

Wrong relationship.

Wrong reference.

Wrong logic.

Wrong number.

Wrong name.

Wrong unit.

Wrong date.

Expanded terminology taxonomy

Wrong term.

Inconsistent term.

Deprecated term.

Forbidden term.

Defined-term violation.

Abbreviation error.

Expanded target-language taxonomy

Grammar.

Syntax.

Agreement.

Collocation.

Spelling.

Punctuation.

Cohesion.

Word form.

Expanded style taxonomy

Register.

Tone.

Voice.

Readability.

Brand style.

Unnecessary verbosity.

Unnecessary literalness.

Expanded locale taxonomy

Number format.

Date format.

Currency.

Address.

Quotation marks.

Spelling variant.

Institutional naming.

Expanded technical taxonomy

Tag.

Placeholder.

Markup.

Link.

Code.

Layout.

Truncation.

Reading order.

Severity calibration examples

Minor:

awkward but understandable collocation.

Major:

wrong term for a key product feature.

Critical:

wrong dosage.

Minor:

one inconsistent comma.

Major:

punctuation that changes scope.

Critical:

missing negative in safety instruction.

Severity should follow effect.

Criticality matrix

Ask:

Can this error cause physical harm?

Can it change a legal right?

Can it cause material financial loss?

Can it expose security?

Can it change access to service?

Can it misidentify a person?

If yes, critical classification may be justified.

Major-error matrix

Ask:

Does it materially distort meaning?

Could it mislead action?

Does it damage document purpose?

Does it violate key terminology?

If yes, major.

Minor-error matrix

Ask:

Does meaning remain intact?

Is function preserved?

Is issue mainly polish or convention?

If yes, minor.

Preference matrix

Ask:

Is the existing target acceptable under specifications?

Would reviewer version simply be another valid choice?

If yes, do not count as error.

The importance of not over-penalising variation

Translation has multiple valid solutions.

If evaluators penalise every difference from their preferred wording, metrics become reviewer-style scores rather than quality scores.

This damages trust.

Evaluator training

Train evaluators on:

taxonomy,

severity,

specifications,

acceptable variation,

domain,

and evidence standards.

Evaluation is a skill.

Calibration pack

Create ten to twenty examples with agreed decisions.

Use them when onboarding new evaluators.

Borderline cases

Keep examples of difficult distinctions.

Minor vs major.

Style vs accuracy.

Term vs general lexical error.

This improves consistency.

Adjudication

When evaluators disagree:

compare source,

specification,

taxonomy,

and consequence.

Use a senior adjudicator for high-stakes formal evaluation.

Record principle.

Quality score design

One possible scoring model:

critical errors cause automatic fail.

Major errors carry larger penalties.

Minor errors smaller penalties.

Normalise by words.

But this is only one design.

Weighted errors

Weights should reflect project priorities.

Technical manual may weight:

accuracy and terminology heavily.

Marketing may also weight style and brand.

Public-service form may weight usability and access.

Category weights

Be cautious.

If terminology is weighted too strongly, reviewers may ignore serious general accuracy errors.

Core semantic accuracy usually deserves strong protection.

Quality thresholds

Define:

pass,

conditional pass,

fail

before evaluation where possible.

Conditional pass may require correction.

Acceptance after correction

A failed sample can be corrected.

Do not assume first evaluation is final.

Quality system should support remediation.

Re-evaluation

After correction, check:

affected segments,

systemic pattern,

and sample if needed.

Do not blindly accept “fixed.”

Sampling design for formal evaluation

Step 1:

define population.

Step 2:

identify high-risk strata.

Step 3:

choose random component.

Step 4:

choose risk component.

Step 5:

define sample size.

Step 6:

document method.

Repeatability matters.

Stratified sampling

Sample across:

sections,

file types,

translators,

language variants,

or content classes.

This prevents one easy area dominating.

Cluster risk

If one translator handled one chapter, sample that chapter enough to detect pattern.

New-vendor sampling

Sample more heavily during onboarding.

Reduce only after stable evidence.

New-engine sampling

Machine-translation changes deserve increased sampling until performance is established.

High-risk full review

Sampling may be inappropriate when every unit matters.

Examples:

medication dose,

legal rights,

emergency instructions.

Translation quality and process evidence

A clean evaluation score does not prove process was good.

A strong process does not prove output is error-free.

Use both output evidence and process evidence.

Process evidence

Qualified translator.

Revision completed.

QA completed.

Termbase used.

Source version controlled.

Queries resolved.

Output evidence

Sample score.

Critical-error count.

Reader test.

In-context pass.

Together they support release confidence.

Release confidence levels

Low confidence:

unknown provenance,

no review,

no QA.

Moderate:

qualified translator,

self-review,

basic QA.

High:

independent revision,

current terminology,

QA,

context check.

Very high:

high-risk specialist review,

formal evaluation,

sign-off.

These are conceptual, not universal labels.

Confidence should be explainable

If someone asks:

Why are we releasing this translation?

The team should be able to answer with evidence.

Quality debt

Quality debt is unresolved translation risk carried forward.

Examples:

stale TM,

unresolved term conflicts,

unreviewed legacy content,

missing context,

or temporary machine output left live.

Like technical debt, it accumulates.

Managing quality debt

Inventory.

Prioritise by risk and traffic.

Fix high-impact items first.

Do not rewrite everything blindly.

Legacy content audit

Classify:

trusted current,

needs spot check,

needs full review,

deprecated,

archive.

This is more efficient than assuming old = bad.

Quality regression

A new release can reduce translation quality.

Causes:

new term,

new translator,

new engine,

new style guide,

new UI layout.

Regression testing applies to language too.

Linguistic regression test

Maintain representative approved examples.

After system changes, compare.

Terminology regression

Search for forbidden old term.

UI regression

Check overflow and truncation after target update.

Functional regression

Confirm forms, links and variables still work.

Quality dashboard design

Show decisions, not decorative statistics.

Useful panels:

critical incidents,

top error categories,

trend by locale,

review backlog,

terminology defects,

release blocks.

Dashboard anti-pattern

One red/green “quality score” with no explanation.

Managers cannot act.

Quality governance roles

Quality manager.

Language lead.

Translator.

Reviser.

Reviewer.

Subject expert.

Localization engineer.

Project manager.

Release owner.

Each owns different risk.

Quality manager

Owns system:

rubric,

sampling,

calibration,

metrics,

incidents,

improvement.

Language lead

Owns language-community quality:

terminology,

style,

review decisions,

calibration.

Workflow manager

Owns risk classification and review assignment.

This mirrors the European Commission’s quality architecture at a general conceptual level.

Quality assistants / formal checks

Mechanical QA can be delegated to people or systems focused on formal integrity.

This preserves expert linguistic attention.

Continuous improvement

Quarterly ask:

Which errors repeat?

Which specifications are unclear?

Which termbase entries cause conflict?

Which reviewers disagree most?

Which file types fail?

Which language pairs have incidents?

Then change system.

Quality maturity model

Level 1:

spell-check and hope.

Level 2:

translator self-review.

Level 3:

shared terminology and independent review.

Level 4:

risk-based QA and formal taxonomy.

Level 5:

analytic evaluation, sampling and incident learning.

Level 6:

continuous multilingual quality governance across human and machine output.

Maturity means repeatability and evidence.

Evaluation design: from specifications to a scorecard

A translation-quality scorecard should be the last step of a specification process, not the first.

Begin with:

what the content is,

what the reader needs,

what errors matter,

how serious those errors can become,

what review resources exist,

and what release decision the evaluation must support.

Then build the scorecard.

If a team copies another organisation’s scorecard without its context, the numbers can look precise while answering the wrong question.

Step 1: define the evaluation purpose

Possible purposes include:

accept or reject delivered translation,

compare vendors,

compare machine-translation engines,

monitor one translator,

audit legacy content,

calibrate reviewers,

evaluate training progress,

or decide whether content is safe to publish.

One scorecard can sometimes support several purposes.

Do not assume it always should.

Step 2: define evaluation scope

What is included?

Source fidelity?

Target grammar?

Terminology?

Style?

Formatting?

Technical integrity?

Usability?

SEO?

Accessibility?

The scope should match the decision.

Step 3: define error categories

Start with a core taxonomy.

Expand only if more detail changes action.

A medical programme may need:

dose,

frequency,

route,

contraindication

as high-visibility subtypes.

A software programme may need:

placeholder,

tag,

string function,

truncation.

Taxonomy follows risk.

Step 4: define severity

Use descriptions tied to consequence.

Critical:

can cause serious harm or render content unusable.

Major:

materially changes meaning or function.

Minor:

degrades quality without materially changing core function.

Preference:

acceptable alternative; no penalty.

Step 5: define weights

If using numerical penalties, document them.

Example:

major = 5.

minor = 1.

critical = automatic fail.

This is one possible project rule.

The exact numbers are not universal.

Step 6: define denominator

Per 1,000 words?

Per 100 segments?

Per document?

Per task?

The denominator affects interpretation.

Choose one that supports comparison.

Step 7: define threshold

Pass if score below threshold?

Fail if critical error exists?

Conditional pass if correctable?

Set rules before evaluation.

Step 8: define evaluator qualifications

Who may score?

Language specialist?

Subject expert?

Trained reviewer?

Do they need both source and target languages?

Formal evaluation is only as consistent as the evaluators.

Step 9: define sample

Full document or sample?

Random or risk-weighted?

How large?

Document method.

Step 10: define reporting

Who receives:

detailed error table?

summary score?

release decision?

recommendations?

Evaluation exists to support action.

Scorecard example: general public information

Categories:

accuracy,

terminology,

target language,

style,

locale,

format.

Severity:

critical,

major,

minor.

Critical rule:

no critical error allowed.

Major threshold:

low.

Style tolerance:

moderate.

Why:

public text needs meaning and reader clarity.

Scorecard example: medical device

Categories:

accuracy,

clinical terminology,

numbers/units,

warnings,

target language,

technical integrity.

Critical:

dosage,

wrong action,

missing contraindication,

wrong device setting.

Style errors may matter less than safety.

Scorecard example: legal

Categories:

accuracy,

defined terms,

obligation/permission,

conditions/exceptions,

entities,

numbers/dates,

target legal drafting.

Critical:

changed right or obligation.

Scorecard example: marketing

Categories:

factual accuracy,

brand terminology,

claim limits,

target style,

tone,

CTA function,

SEO where required.

A slightly literal sentence may be a quality defect even if factually accurate because purpose is persuasion.

Scorecard example: software UI

Categories:

function,

terminology,

locale,

length,

placeholder integrity,

grammar,

consistency.

One mistranslated button can block user action.

Scorecard example: subtitles

Categories:

meaning,

timing,

reading speed,

speaker identity,

tone,

line breaks,

cultural references.

Quality includes audiovisual constraints.

Scorecard example: education

Categories:

meaning,

command words,

difficulty,

answer scope,

terminology,

target readability,

format.

A translation that helps too much can be wrong for assessment.

Scoring example 1

Document:

1,000 words.

Errors:

2 minor grammar errors.

1 major terminology error.

Weights:

minor 1.

major 5.

Raw penalty = 7.

If threshold is 10 and no critical error exists, document passes.

But the team may still require correction before release.

Passing evaluation does not mean leave known errors unfixed.

Scoring example 2

Document:

10,000 words.

One critical safety error.

Many otherwise excellent segments.

Automatic fail.

This demonstrates why severity can override average quality.

Scoring example 3

Document:

200-word legal clause.

One major condition error.

Penalty normalisation may look large because text is short.

That is appropriate if the error changes clause meaning.

Scoring example 4

Marketing page contains no accuracy errors but severe tone mismatch.

If brand purpose is central, evaluation can fail based on style.

Fit for purpose matters.

Scoring example 5

Internal gist MT has several minor grammar errors but accurate core information.

If purpose is internal comprehension, it may pass.

The same output would fail publication standard.

Evaluation without scoring

Not every quality review needs numbers.

A qualitative decision can be:

accept,

accept with corrections,

revise and resubmit,

reject.

Qualitative evaluation may be better for creative work.

When scoring is useful

Vendor management.

Large-scale sampling.

Engine comparison.

Trend analysis.

Training.

Formal acceptance.

When scoring can mislead

Literary translation.

Transcreation.

Small samples.

Uncalibrated reviewers.

Ambiguous specifications.

Do not quantify uncertainty you have not resolved.

Translation quality and MQM-Core

A core taxonomy provides common ground across organisations.

Teams can then extend for domain needs.

This improves comparability without forcing every project into one identical model.

Translation quality and MQM-Full

A more detailed taxonomy can support:

research,

large enterprise programmes,

or specialised diagnosis.

More categories create more training burden.

Use only the detail you can apply consistently.

Error granularity

Should a wrong date be:

accuracy > number?

locale > date?

format > convention?

Choose one policy.

Consistency matters more than philosophical perfection.

Primary error principle

When one defect could fit several categories, choose the category closest to root linguistic effect.

Then use comment for detail.

Avoid double-penalising one underlying defect unless it causes separate independent problems.

Error span

Mark enough target text to show defect.

Do not tag entire paragraph for one word unless entire paragraph is wrong.

Precise spans improve analysis.

Repeated errors

If one wrong term repeats twenty times, how many errors?

Possible policies:

count every occurrence,

count first plus consistency penalty,

or count capped occurrences.

Define before scoring.

Repetition policy trade-off

Counting every repeat reflects user exposure.

Capping prevents one systemic term from overwhelming all other data.

The right method depends on evaluation goal.

Systemic errors

A repeated term error is often more important as root cause than as count.

Report both:

frequency,

systemic cause.

Critical error override

Some systems use automatic fail for one critical error.

This is appropriate when consequence is unacceptable.

Define critical examples clearly.

Near-critical errors

Borderline cases should go to adjudication.

Do not let one evaluator improvise catastrophic labels casually.

Severity calibration workshop

Give reviewers 20 examples.

Each independently rates.

Compare.

Discuss consequence.

Create shared anchors.

Repeat periodically.

Anchor examples

Critical anchor:

“0.5 mg” → “5 mg.”

Major anchor:

“may” → “must” in policy.

Minor anchor:

awkward but understandable collocation.

Preference anchor:

two equally natural synonyms.

Anchors improve consistency.

Evaluator drift

Over time, reviewers can become:

stricter,

more lenient,

or idiosyncratic.

Monitor calibration.

Vendor-specific leniency

Do not unconsciously grade preferred vendors more gently.

Blind samples can reduce bias.

Translator-specific expectations

A famous translator can still make an error.

A junior translator can produce excellent work.

Evaluate output.

Machine bias

Do not assume machine output deserves harsher or softer scoring.

If evaluating final output, apply task criteria consistently.

Provenance can still influence required review process.

Reference bias

If evaluator has one reference translation, they may penalise valid alternatives.

Train for equivalence, not reference imitation.

Source bias

Evaluators can assume source wording must be mirrored.

Target-language quality may require restructuring.

Target bias

Monolingual elegance can distract from source mismatch.

Bilingual evaluation protects fidelity.

Evaluation fatigue

Long scoring sessions reduce consistency.

Use:

manageable batches,

breaks,

and sampling.

Double evaluation

High-stakes formal programmes can have two independent evaluators on selected samples.

Use adjudication for disagreements.

Expensive but useful for calibration.

Evaluation audit

Periodically audit evaluators.

Check whether their labels match rubric.

Quality evaluation needs its own quality control.

Sampling architecture in depth

A sample should represent both ordinary content and important risk.

One useful formula conceptually combines:

random coverage + risk targeting + novelty targeting.

Random:

general quality.

Risk:

critical sections.

Novelty:

new translator, new engine, new terminology.

Random sample design

Use actual random selection where possible.

Do not let project manager handpick “representative” easy pages.

Systematic sample

Every nth segment can be practical.

Risk:

periodic structure may bias.

Stratified sample

Divide by:

section,

content type,

translator,

or risk.

Then sample each.

This improves coverage.

Cluster sampling

Select whole pages or sections.

Useful for document coherence.

Less precise statistically.

Risk-based sample

Always include:

warnings,

legal conditions,

numbers,

new terms,

names,

and changed content.

Change-based sample

For updated documents, sample changed segments more heavily.

Old approved unchanged content may need less review.

New-person sample

New translator or reviewer receives higher sample rate during onboarding.

Reduce after evidence.

New-tool sample

New MT engine, prompt or CAT migration deserves extra quality sampling.

High-traffic sample

For web content, pages with highest user exposure deserve more review.

Exposure is part of risk.

Sample escalation

If sample finds major pattern, expand evaluation.

Sampling should have escalation rules.

Example:

one critical error → full review.

three major same-category errors → expand sample.

Sample termination

If sample remains clean beyond threshold, release can proceed according to policy.

Document decision.

Confidence intervals and formal statistics

Large programmes may use statistical confidence intervals.

Small translation teams may not need them.

Do not use complex mathematics without meaningful data assumptions.

Practical sampling over pseudo-science

A transparent 10% risk-weighted sample can be more useful than a mathematically impressive model built on inconsistent reviewer labels.

Measurement quality begins with reliable categories.

Quality report anatomy

Section 1:

project and scope.

Section 2:

specifications.

Section 3:

method and sample.

Section 4:

score or decision.

Section 5:

critical findings.

Section 6:

systemic patterns.

Section 7:

recommended corrections.

Section 8:

process improvements.

Translation-quality executive report

Keep concise:

release status,

critical risk,

top three issues,

action owner,

deadline.

Managers need decisions.

Translator feedback report

Include:

examples,

rationale,

patterns,

and priorities.

Feedback should improve future work.

Vendor corrective action request

When supplier quality fails:

show evidence,

specification violated,

severity,

required corrective action,

and re-evaluation plan.

Avoid vague accusations.

Quality incident report

Include:

what happened,

impact,

languages affected,

content affected,

containment,

correction,

root cause,

preventive action.

Quality debt register

List unresolved known risks.

Item.

Locale.

Traffic.

Severity.

Owner.

Planned action.

Do not let “temporary” bad translation disappear into memory.

Quality backlog prioritisation

Priority score can consider:

severity,

user exposure,

age,

legal risk,

and repair cost.

Fix high consequence first.

Legacy translation clean-up

Do not rewrite all legacy content automatically.

Audit first.

Some old translations are excellent.

Some are obsolete.

Some only need term updates.

Quality inheritance

New content can inherit old defects through:

TM,

templates,

copy-paste,

CMS clones,

or AI retrieval.

Quality system must govern reuse.

Quality provenance

Know whether content originated from:

human translation,

MT,

post-editing,

legacy import,

or unknown source.

This helps set review.

Quality at source

Bad source increases translation error risk.

Measure source readability and consistency where useful.

ISO 5060 notes its guidance can also support evaluation of source texts intended for translation.

That is an important upstream connection.

Source-quality categories

Ambiguity.

Inconsistency.

Missing information.

Broken syntax.

Wrong number.

Undefined acronym.

Poor formatting.

Flagging source defects prevents unfair translator blame.

Translation quality and controlled authoring

Clear source terms and references improve all languages.

Quality can be engineered upstream.

Translation quality and terminology governance

Termbase quality affects translation quality.

Audit termbase:

definitions,

preferred terms,

status,

locale,

and currency.

Translation quality and TM governance

A contaminated TM can systematically lower output.

Audit:

wrong targets,

stale terms,

wrong locale,

and provenance.

Translation quality and model governance

MT or AI models can change.

Quality benchmark after change.

Translation quality and prompt governance

Prompt changes can alter output.

Version and test.

Translation quality and file engineering

Broken extraction creates translation defects.

Quality includes:

input pipeline,

export,

and round-trip testing.

Translation quality and localisation

Localization quality includes function, not only language.

Date pickers.

Sorting.

Search.

Layout.

Plural rules.

Quality system should connect linguistic and product QA.

Translation quality and accessibility

Accessibility defects can be quality defects.

Missing alt text.

Broken reading order.

Unclear labels.

Translation quality and security

Wrong translation of security instruction can create risk.

Confidentiality failures in translation workflow are also quality-management failures at service level, even if target text is linguistically correct.

Translation quality and privacy

Names and personal data must be handled correctly.

Do not overexpose content in evaluation systems.

Translation quality and ethics

A quality system should not optimise for a score at expense of reader rights, source fidelity or translator fairness.

Goodhart’s law in translation quality

When a metric becomes a target, people may game it.

Examples:

avoid reporting queries to look fast.

make fewer reviewer changes to improve acceptance rate.

choose easy samples.

Quality governance must guard against metric gaming.

Balanced scorecard

Combine:

output quality,

process quality,

user outcome,

and improvement.

No one metric owns truth.

Quality trend

Track whether:

critical errors fall,

term consistency improves,

reviewer agreement rises,

incidents decline.

Trend can be more informative than isolated score.

Quality by language pair

Do not compare raw rates without context.

Different languages and content differ in difficulty.

Quality by domain

Technical and marketing defects look different.

Domain-specific dashboards help.

Quality by source quality

A translator working on broken source should not be compared directly to one receiving polished source without context.

Quality by reviewer

Monitor reviewer variance.

A severe reviewer can make one vendor look worse.

Calibration matters.

Quality by release

Track whether quality improves after new termbase or workflow.

This tests interventions.

Quality experimentation

When changing process:

pilot.

Measure before/after.

Do not assume change helps.

Example intervention: new terminology QA

Before:

many wrong-term errors.

After:

fewer term errors, slightly more false warnings.

Decision:

tune rules.

Example intervention: mandatory second review

Before:

critical errors in low-risk documents rare.

After:

quality slightly better, delivery much slower.

Decision:

apply only to high-risk documents.

Risk-based design beats universal rules.

Example intervention: AI-assisted QA

Before:

reviewers miss omissions.

After:

AI flags many real omissions plus false positives.

Decision:

use AI as triage with human confirmation.

Translation-quality programme charter

Purpose:

ensure target output meets specifications with proportionate risk control.

Scope:

languages, content types, providers.

Methods:

revision, review, QA, evaluation, sampling.

Roles:

owners and escalation.

Metrics:

limited actionable set.

Incident process:

defined.

Continuous improvement:

scheduled.

Quarterly quality review

Review:

incidents,

error trends,

reviewer calibration,

termbase issues,

TM quality,

new tools,

and legacy debt.

Annual quality architecture review

Ask:

Do categories still fit?

Have new domains emerged?

Are risks changing?

Are standards updated?

Do machine outputs need different sampling?

Is reviewer training current?

Quality systems must evolve.

Genre-specific quality architecture

Translation quality changes shape across genres. The quality system should preserve a common core while changing the weight of different checks.

Legal translation quality

Primary risks:

rights,

obligations,

defined terms,

conditions,

exceptions,

jurisdiction,

dates,

parties,

and cross-references.

Quality priorities:

semantic fidelity,

defined-term consistency,

modal precision,

scope,

and legal readability.

Typical critical defects:

must/may reversal,

missing exception,

wrong party,

wrong date,

wrong amount,

wrong defined term.

Review model:

qualified translator,

bilingual revision,

legal subject review where necessary,

formal QA.

Medical translation quality

Primary risks:

dose,

frequency,

route,

contraindication,

device setting,

risk warning,

patient eligibility,

and clinical uncertainty.

Quality priorities:

clinical meaning,

plain-language usability where patient-facing,

terminology,

numbers,

and safety.

Typical critical defects:

decimal error,

unit error,

dose-frequency change,

omitted warning.

Review model:

specialist translator,

bilingual revision,

medical expert,

hard-detail QA.

Financial translation quality

Primary risks:

currency,

rates,

percentages,

percentage points,

accounting concepts,

forward-looking claims,

and legal disclosures.

Quality priorities:

quantitative accuracy,

terminology,

claim strength,

table/narrative consistency.

Typical critical defects:

wrong currency,

wrong sign,

decimal error,

material disclosure omission.

Technical translation quality

Primary risks:

procedure,

sequence,

component names,

units,

warnings,

and drawings.

Quality priorities:

operational accuracy,

terminology,

controlled repetition,

diagram alignment.

Typical critical defects:

reversed sequence,

wrong part,

missing safety instruction.

Software localisation quality

Primary risks:

function,

short-string ambiguity,

placeholder integrity,

layout,

locale,

and consistency.

Quality priorities:

correct string function,

product terminology,

variables,

UI fit,

and in-context usability.

Typical critical defects:

wrong button action,

broken variable,

mistranslated permission,

unusable workflow.

Marketing translation quality

Primary risks:

claim distortion,

brand voice,

CTA failure,

cultural mismatch,

and search-intent mismatch.

Quality priorities:

factual fidelity,

persuasive function,

tone,

target-market naturalness,

and SEO where relevant.

Typical major defects:

“up to” removed,

unsupported superlative,

CTA weakened,

brand tone lost.

Literary translation quality

Primary risks:

voice,

rhythm,

ambiguity,

imagery,

character distinction,

and historical register.

Quality priorities:

interpretive fidelity,

stylistic effect,

coherence,

and literary quality.

Formal error scoring may be less useful than expert comparative review.

Academic translation quality

Primary risks:

claim strength,

method,

evidence,

citations,

hedging,

and terminology.

Quality priorities:

epistemic accuracy,

discipline style,

citation integrity,

and clear argument.

Typical major defects:

correlation turned into causation,

suggestion turned into proof,

citation scope changed.

Public-service translation quality

Primary risks:

eligibility,

rights,

deadline,

required documents,

appeal,

and navigation.

Quality priorities:

access,

clarity,

legal meaning,

and reader action.

Typical critical defects:

wrong threshold,

missing exception,

wrong deadline.

Education translation quality

Primary risks:

learning objective,

command word,

difficulty,

assessment construct,

and age appropriateness.

Quality priorities:

task equivalence,

curriculum terminology,

reading level,

and instructional clarity.

Subtitle quality

Primary risks:

timing,

reading speed,

speaker attribution,

compression,

and scene context.

Quality priorities:

meaning,

brevity,

natural dialogue,

and audiovisual fit.

Game localisation quality

Primary risks:

branching logic,

UI function,

lore terminology,

character voice,

and variables.

Quality priorities vary by string type.

One scorecard may need subprofiles.

Quality-profile architecture

Instead of one universal QA setting, create profiles.

General prose.

Legal.

Medical.

Software.

Marketing.

Education.

Each profile can define:

error categories,

severity rules,

automated checks,

sampling,

and release threshold.

Profiles turn risk policy into reusable configuration.

General-prose profile

Automated checks:

numbers,

tags,

terminology,

spelling.

Human checks:

meaning,

naturalness,

style.

Legal profile

Automated checks:

numbers,

defined terms,

cross-references,

dates,

names.

Human checks:

modal force,

scope,

conditions,

legal concepts.

Medical profile

Automated checks:

numbers,

units,

drug/device names,

forbidden terms.

Human checks:

clinical meaning,

instructions,

warnings,

patient readability.

Software profile

Automated checks:

variables,

tags,

length,

duplicates,

terminology.

Human checks:

function,

context,

UI naturalness.

Marketing profile

Automated checks:

names,

claims,

required terms,

length.

Human checks:

voice,

persuasion,

culture,

SEO intent.

Quality specification card

A small project can use a one-page card.

Content:

Public help article.

Risk:

Medium.

Target:

fr-FR.

Audience:

General users.

Accuracy priority:

High.

Style:

Friendly professional.

Terminology:

Product glossary v4.

Review:

Bilingual revision.

QA:

Numbers, links, terms, tags.

Release:

CMS preview.

This card reduces ambiguity.

Quality acceptance matrix

Green:

No critical errors.

Major errors below threshold.

All known errors corrected.

QA passed.

Yellow:

No critical errors.

One or more correctable major issues.

Release blocked until repair.

Red:

Critical error.

Systemic misunderstanding.

Wrong locale.

Incomplete translation.

Requires rework and re-evaluation.

Conditional release

Sometimes non-critical content must ship with a known minor issue.

Document exception.

Owner.

Reason.

Planned correction.

Do not let exceptions disappear.

Release waiver

A waiver should never hide critical safety or legal risk.

Use only where policy permits.

The translation-quality release gate

Before publication, verify:

specifications met,

required review complete,

critical errors absent,

known major errors corrected,

QA complete,

format validated,

approval present.

The gate should be binary enough to stop bad release.

Gate evidence

A release gate can store:

QA report,

review status,

evaluation score,

query closure,

version ID,

approval.

This creates auditability.

Quality ownership model

Requester owns:

clear requirements.

Source owner owns:

source accuracy.

Translator owns:

translation decisions.

Reviser owns:

independent comparison.

Reviewer owns:

target-reader suitability.

Subject expert owns:

domain validation.

Quality manager owns:

system and metrics.

Release owner owns:

final gate.

Shared responsibility is not diluted responsibility.

Requester responsibility

Quality failures can begin with vague requests.

“Translate this” is not enough for high-stakes work.

Requester should define purpose and audience.

Source-owner responsibility

If source contains contradictory facts, translation cannot create truth.

Source owner resolves.

Translator responsibility

Do not guess critical ambiguity.

Use resources.

Self-revise.

Flag risk.

Reviser responsibility

Check evidence.

Do not overedit preferences.

Escalate systemic issues.

Reviewer responsibility

Represent target reader.

Do not rewrite source meaning.

Subject-expert responsibility

Validate concept.

Do not assume language authority unless qualified.

Quality-manager responsibility

Create fair rubrics.

Calibrate.

Analyse trends.

Improve process.

Release-owner responsibility

Confirm gates.

Do not waive silently.

The specification chain

Source specification.

Translation specification.

Review specification.

Evaluation specification.

Release specification.

Each layer should connect.

Source specification example

US English source.

Technical support.

Product X v3.2.

All warnings final.

No changes after freeze.

Translation specification example

Japanese.

Formal customer support.

Approved product termbase.

Retain English API names.

Review specification example

Full bilingual revision.

Focus on product actions and terminology.

Evaluation specification example

10% random sample + all warnings.

Critical error automatic fail.

Release specification example

In-context review passed.

No open severity-1 issue.

This chain makes quality traceable.

Formal quality evaluation versus production QA

Production QA aims to make the translation better.

Formal evaluation aims to judge output.

The evaluator may not edit everything.

Separate judging from fixing when measurement objectivity matters.

Why evaluation after correction can hide initial quality

If evaluating vendor performance, record initial output before internal corrections.

Otherwise the score reflects reviewer work.

Why release quality should reflect final output

For users, final target is what matters.

Maintain both metrics if needed:

delivered quality,

released quality.

Quality at handoff

A translation can degrade during:

copy-paste,

DTP,

CMS import,

software integration.

Perform post-handoff QA.

DTP quality

Check:

fonts,

line breaks,

hyphenation,

tables,

captions,

reading order,

and missing text.

CMS quality

Check:

correct locale,

links,

metadata,

heading structure,

and no fallback language.

Software-build quality

Check:

string loading,

variables,

plural rules,

RTL,

truncation,

and encoding.

Subtitle-build quality

Check:

timecodes,

line breaks,

overlaps,

speaker labels.

Print quality

Check final proof, not only source file.

Quality transfer across formats

When Word becomes PDF, or XLIFF becomes app strings, each transformation creates risk.

Quality system follows the content.

Correction governance

A correction should have:

issue ID,

severity,

affected locale,

affected content,

owner,

status,

root cause,

and propagation list.

This may be lightweight.

Correcting one published page

Fix target.

Check reusable sources.

Republish.

Verify cache and search index where necessary.

Correcting many pages

Use search/replace only with controlled scope.

Test.

Backup.

Verify grammar.

Correcting a term globally

A target-language term may inflect.

Blind string replacement can break grammar.

Use linguistic review.

Correcting a number

Search for same wrong number in all outputs.

Check whether error came from source, TM or template.

Correcting a name

Identity errors may appear in:

body,

metadata,

alt text,

PDF,

search snippets.

Propagation matters.

Correction priority

Critical live error:

immediate.

Major high-traffic:

urgent.

Minor low-traffic:

scheduled.

Corrigenda

The European Commission publicly describes formal correction procedures for significant translation errors in legal acts.

The broader lesson is that important translation systems need formal correction paths rather than silent ad hoc edits.

Quality transparency

Internal transparency:

who changed what and why.

External transparency:

where regulation or product policy requires it.

Quality documentation burden

Do not document every comma.

Document decisions proportional to risk.

Quality and speed

Quality systems should accelerate safe work by preventing rework.

Clear terminology saves time.

Calibrated reviewers save debate.

Automated QA saves manual counting.

Quality is not always slower.

Quality bottlenecks

Common bottlenecks:

late queries,

unclear terms,

reviewer conflicts,

broken files,

source changes,

and last-minute DTP.

Track bottlenecks.

Quality lead time

Measure from source freeze to approved target.

Break down:

translation,

review,

query wait,

QA,

engineering.

This reveals process delay.

Queue risk

High-risk content waiting behind low-risk volume is bad prioritisation.

Use risk-based queues.

Quality capacity planning

Allocate senior reviewers to:

new translators,

new domains,

high-risk content,

and incidents.

Do not spend scarce expertise evenly.

Quality and fatigue

Reviewer fatigue increases missed errors.

Limit batch size.

Use automation.

Rotate tasks.

Quality and cognitive diversity

Different reviewers notice different defects.

One strong at source fidelity.

One at target style.

Role differentiation can improve coverage.

Quality and language community

Each language has specific:

grammar,

typography,

register,

and locale issues.

Global rubric plus language-specific guidance is powerful.

This mirrors the layered structure of the European Commission’s quality system: overarching framework plus operational and language-specific guidance.

Language-specific guidelines

Examples:

French punctuation.

German compounds.

Arabic RTL.

Japanese honorifics.

Chinese names and scripts.

Spanish regional variation.

One global checklist cannot encode every language detail.

Quality and terminology community

Terminologists can monitor cross-language concept consistency.

Quality is not only per-document.

Quality and institutional voice

Language leads maintain stable institutional voice.

Translation quality includes cross-document coherence.

Quality and continuous learning

Review data can become:

training examples,

rubric anchors,

style rules,

and termbase updates.

Quality feedback meeting

Monthly 30 minutes.

Top repeated errors.

One calibration example.

One process issue.

One resource update.

Small regular learning beats annual blame.

Quality and psychological safety

Translators should feel safe reporting uncertainty and mistakes.

Hidden uncertainty is more dangerous than visible uncertainty.

Quality and professional pride

Quality culture should value:

evidence,

reader impact,

and learning

rather than perfection theatre.

Quality and automation era

As AI produces more text, quality bottleneck moves from generation to evaluation.

The scarce resource becomes expert attention.

Risk routing is therefore central.

Expert attention budget

Spend humans on:

critical segments,

ambiguous content,

low-confidence output,

new terminology,

and reader-sensitive text.

Let automation handle deterministic checks.

Quality triage

Green segments:

trusted exact matches, low risk.

Amber:

fuzzy, new terms, complex syntax.

Red:

safety, law, dose, money, eligibility.

Review intensity follows colour.

Quality triage for human translation

Even human-drafted content can be red.

Origin does not lower consequence.

Quality triage for MT

Even high-confidence MT can be red.

Confidence does not lower consequence.

Quality triage for legacy content

Old approved content can become red after policy change.

Freshness matters.

Final principle of risk-based quality

Quality control should follow potential harm, not prestige of tool, translator or source.

That is the fairest and safest architecture.

Source quality and translation quality

A translation-quality system can fail if it evaluates the target without acknowledging the source.

A source can contain:

contradictions,

missing information,

bad OCR,

inconsistent terminology,

wrong numbers,

ambiguous pronouns,

broken tables,

unfinished sentences,

or impossible dates.

The translator can respond professionally.

They cannot create source truth from nothing.

Source defect versus translation defect

Source:

“Submit by 31 February.”

Target preserves the date exactly.

The target is faithful to a defective source.

Quality action:

flag source defect.

Do not score translator as inaccurate for preserving evidence unless specification required correction.

Source defect silently repaired

Translator changes 31 February to 28 February without confirmation.

Now translation differs from source.

Even if likely intended, it is unsupported.

Category:

accuracy — addition/alteration.

Quality system should reward query behaviour.

Source ambiguity preserved

Source deliberately ambiguous.

Target preserves same ambiguity.

This may be high quality.

A reviewer who “clarifies” it can introduce error.

Source ambiguity unresolved but target forces choice

Target language requires gender.

Source does not.

Translator chooses randomly.

Quality defect:

unsupported addition.

Better process:

neutral reformulation or query.

Source readability and quality risk

Poor source readability increases translator uncertainty.

Project manager should increase review or seek source revision.

Risk begins upstream.

Source evaluation before translation

For high-risk projects, evaluate source for:

clarity,

consistency,

terminology,

completeness,

and format.

This reduces downstream quality cost.

Controlled authoring quality

Controlled source language can improve translation consistency by:

one term per concept,

clear references,

stable sentence patterns,

explicit units,

and limited ambiguity.

Not every genre should be controlled.

Literature depends on style.

The source-quality feedback loop

Translation queries are data.

If translators repeatedly ask:

What does “it” refer to?

Which unit?

Which product?

Then source authoring needs improvement.

Quality systems should send feedback upstream.

Target quality and translationese

Translationese is target language influenced unnaturally by source patterns.

Signs:

source-like word order,

over-explicit pronouns,

unnatural collocations,

repeated nominalisation,

foreign punctuation patterns.

Target quality review catches it.

Accuracy versus naturalness

These are not opposites.

A high-quality translation aims for both.

But when naturalness edits change meaning, accuracy wins.

When literal structure creates wrong meaning or unreadability, restructuring is necessary.

Fluency trap

Evaluators may give high scores to fluent output.

Use bilingual evidence.

A polished false statement remains inaccurate.

Literalness trap

Evaluators may penalise natural target variation because it differs from source wording.

Evaluate meaning and function.

Quality and ambiguity

Quality does not always mean resolving ambiguity.

If source ambiguity is purposeful or unresolved, target may need to preserve it.

Quality and explicitation

Target languages sometimes require explicit information implicit in source.

This can be legitimate.

The information must be inferable from source/context.

Do not invent.

Quality and implicitation

Target may omit information grammatically recoverable or redundant in target language.

Check whether meaning remains.

Not every surface omission is semantic omission.

Quality and compensation

A stylistic effect may be moved elsewhere in target.

Literary evaluation should allow compensation.

Rigid segment-level scoring can miss it.

Quality and transcreation

Transcreation may intentionally move away from literal source.

Specifications should define what must remain invariant:

claim,

brand promise,

CTA,

identity.

Evaluation follows those invariants.

Quality and adaptation

Units, dates, examples and cultural references may be adapted.

If specified, adaptation is not error.

If unapproved, it can be.

Quality and localisation engineering

A string can be linguistically correct and fail because product internationalisation is poor.

Example:

one variable design cannot support target grammar.

Do not blame translator for architecture limitation.

Record engineering defect.

Defect ownership taxonomy

Translation defect.

Source defect.

Terminology defect.

Engineering defect.

Process defect.

Specification defect.

Evaluation defect.

This helps organisations fix right layer.

Quality and service quality

Translation-service quality includes more than text.

On-time delivery.

Security.

Correct files.

Communication.

Query management.

ISO 17100 frames translation services as process plus resources.

Output evaluation alone is narrower.

Quality and turnaround

Late perfect translation can fail business purpose.

Speed can be part of specification.

But impossible deadline can force quality risk.

Trade-offs should be explicit.

Quality and confidentiality

A linguistically perfect target produced through unauthorised data exposure is not a high-quality professional service.

Security belongs in service quality.

Quality and accessibility

If translation makes a service inaccessible to disabled users, target function fails.

Accessibility requirements can be specifications.

Quality and inclusivity

Language should follow current approved inclusive policy when project requires.

Do not infer identity.

Do not erase source meaning.

Quality and political neutrality

For political or contested content, preserve source framing and attribution.

Do not insert translator opinion.

Quality includes neutrality relative to source.

Quality and religious terminology

Established doctrinal terms may have community-specific authority.

Generic synonym can be inaccurate.

Use appropriate source authority.

Quality and indigenous language governance

Community standards may define quality.

External metrics should not override community linguistic authority.

Quality and low-resource languages

Evaluation resources may be weaker.

Need stronger community expertise and qualitative review.

Do not assume major-language metrics transfer.

Quality and sign languages

Signed-language translation quality may involve:

spatial grammar,

facial markers,

timing,

and video production.

Text-centric rubrics are insufficient.

Quality and speech translation

Interpreting and speech translation have different constraints.

This article focuses on translation output; interpreting needs its own architecture.

Quality and multimodal translation

Comics.

Maps.

Charts.

Video.

Interfaces.

Meaning can be distributed across visual context.

Evaluation must inspect final multimodal product.

Quality and image text

OCR and image replacement can create errors.

Check rendered image, not only extracted text.

Quality and PDF

Text can disappear during export.

DTP can break reading order.

Final PDF proof is necessary.

Quality and spreadsheets

Formulas, IDs and hidden cells can be corrupted.

Linguistic QA plus data integrity.

Quality and slides

Text expansion can obscure charts or overflow.

Final deck review.

Quality and web pages

Check:

metadata,

links,

navigation,

forms,

and fallback language.

Quality and apps

Check:

states,

buttons,

errors,

plural forms,

and variables.

Quality and APIs

Machine-translated API responses need structural validation.

Language quality and schema integrity both matter.

Translation-quality training programme

Week 1:

specifications.

Students receive same source with different briefs.

They see quality priorities change.

Week 2:

accuracy errors.

Mistranslation, omission, addition, logic.

Week 3:

terminology and names.

Week 4:

target-language quality and style.

Week 5:

severity.

Students rate consequences.

Week 6:

revision and review.

Bilingual versus monolingual tasks.

Week 7:

QA tools.

Numbers, tags, terms, placeholders.

Week 8:

MQM-style taxonomy and scorecards.

Week 9:

sampling and evaluator calibration.

Week 10:

incident analysis and corrective action.

Quality becomes a discipline.

Training exercise: one source, three purposes

Source:

“Please remain seated.”

Purpose A:

aircraft safety instruction.

Purpose B:

school assembly note.

Purpose C:

humorous poster.

Students define different quality priorities.

Training exercise: error versus preference

Give ten target alternatives.

Students decide:

error,

acceptable variation,

or preference.

This reduces reviewer overreach.

Training exercise: severity without category

Give only consequence.

Students assign critical/major/minor.

Then show error type.

This teaches separation.

Training exercise: category without severity

Give same error type in three contexts.

Wrong number in:

casual blog,

invoice,

medication label.

Severity changes.

Training exercise: sampling plan

Students receive 100-page manual.

They design:

random sample,

risk sample,

and escalation rule.

Training exercise: evaluator calibration

Five reviewers score same passage.

Compare disagreement.

Revise rubric.

Training exercise: QA false positive

Tool flags number difference because date was localised correctly.

Students decide warning versus error.

Automation flags.

Humans judge.

Training exercise: source defect

Source contradicts table.

Students decide:

translation error?

source query?

subject expert?

Training exercise: release gate

Students receive project status with:

one open major query,

QA passed,

review complete.

Can it release?

No.

Quality gate prevents process bypass.

Training exercise: incident root cause

Published wrong term.

Students trace:

termbase,

TM,

review,

source,

and deployment.

Find actual cause.

Teaching quality to Primary learners

Use simple questions.

Did anything disappear?

Did anything new appear?

Are names and numbers same?

Does sentence make sense?

This builds verification habit.

Teaching quality to Secondary learners

Add:

tone,

register,

connectives,

pronouns,

and word sense.

Students compare translations and justify.

Teaching quality to advanced learners

Add:

rubrics,

severity,

sampling,

terminology,

and revision.

Students act as evaluators.

Translation quality as critical thinking

Evaluation asks:

What claim is source making?

What evidence does target preserve?

What changed?

What consequence follows?

This is structured critical reasoning.

Translation quality as vocabulary precision

Quality review exposes near-synonym differences.

Vocabulary Learning Hub owns deeper lexical learning.

Quality System measures whether the choice works here.

Translation quality as English precision

How English Works owns grammar and discourse.

Quality System checks whether target English uses them correctly.

Quality and the Translation Memory System

Bad evaluation decisions can pollute TM.

Only approved corrected targets should become authoritative reuse.

Quality and the Terminology System

Error trends can reveal missing or bad termbase entries.

Quality feeds terminology.

Quality and the Machine Translation System

MT evaluation data can guide:

engine routing,

prompt changes,

glossary updates,

and review thresholds.

Quality and the Human Translation System

Revision data can guide:

translator training,

reviewer calibration,

and assignment.

Quality and the Context Stack

Many errors are context errors.

Quality findings should identify missing context sources.

Quality and the Translation Unit

Segment-level evaluation can miss cross-sentence issues.

Use document-level review for cohesion.

Quality and Equivalence

Evaluation asks whether relevant invariants survived despite form change.

Equivalence architecture defines those invariants.

Quality and Source Analysis

If source was misread, target quality fails upstream.

Quality feedback can improve source-analysis training.

Frequently asked question: what is translation quality?

Translation quality is the degree to which a target text satisfies defined requirements for meaning, language, terminology, function, audience and technical integrity.

Frequently asked question: what is translation quality assurance?

Translation QA is the system of preventive and checking processes used to reduce translation defects before release.

Frequently asked question: what is translation quality evaluation?

Translation evaluation assesses output against criteria, often using error categories, severity and a score or rating.

Frequently asked question: what is ISO 5060?

ISO 5060:2024 is the current international standard titled “Translation services — Evaluation of translation output — General guidance.”

It covers evaluation of human translation, post-edited MT and unedited MT output, evaluator competence and sampling.

Frequently asked question: what is MQM?

MQM means Multidimensional Quality Metrics.

It is a community framework for analytic translation-quality evaluation using structured error typologies and scoring models.

Frequently asked question: is MQM an ISO standard?

MQM is a separate community framework.

ISO 5060 provides international-standard guidance on translation output evaluation.

They can inform one another but should not be presented as the same thing.

Frequently asked question: what is a critical translation error?

A critical error is one whose consequence can be severe enough to make the translation unsafe or unusable, such as wrong medical dose or reversed safety instruction.

Frequently asked question: what is a major error?

A major error materially changes meaning, function or important terminology but may not create catastrophic harm.

Frequently asked question: what is a minor error?

A minor error reduces language or style quality without materially changing the core meaning or function.

Frequently asked question: is a style preference an error?

Not necessarily.

If the existing target satisfies specifications and is natural, an alternative preferred wording may be merely preference.

Frequently asked question: can translation quality be measured with one score?

It can be summarised with a score for some purposes, but one score hides dimensions.

Always preserve underlying error data.

Frequently asked question: can machine metrics replace human evaluation?

No universal automated metric captures all translation meaning, context and function.

Metrics can support system comparison.

Human evaluation remains important.

Frequently asked question: how much should be sampled?

It depends on risk, document size, history and decision.

High-risk content may need full review.

Frequently asked question: should exact TM matches be evaluated?

If risk or context warrants it, yes.

Historical approval is evidence, not eternal guarantee.

Frequently asked question: is human translation always high quality?

No.

Human output can contain serious errors.

Quality depends on competence and process.

Frequently asked question: is machine translation always lower quality?

No.

For some tasks it can be excellent.

Evaluate output against requirements.

Frequently asked question: why do reviewers disagree?

Translation permits acceptable variation, and severity judgments can differ.

Clear specifications and calibration reduce disagreement.

Frequently asked question: what is bilingual revision?

A source-target comparison by a second linguist to verify accuracy and completeness.

Frequently asked question: what is monolingual review?

A target-language review focused on clarity, tone and reader suitability.

Frequently asked question: what is proofreading?

Late-stage checking of spelling, punctuation, typography, formatting and production details.

Frequently asked question: what is in-context review?

Checking translation inside final website, app, PDF, subtitle or other product.

Frequently asked question: what is a quality gate?

A required checkpoint that blocks release until defined conditions are met.

Frequently asked question: what is release confidence?

Justified confidence that required specifications, review and checks have been satisfied.

It is not a guarantee of zero error.

Frequently asked question: how do you improve quality?

Use:

better source,

better specifications,

better assignment,

terminology,

revision,

QA,

evaluation,

and feedback loops.

Frequently asked question: what if a translation fails evaluation?

Correct defects.

Investigate systemic causes.

Re-evaluate according to policy.

Frequently asked question: what if the reviewer is wrong?

Review findings should be evidence-based.

Use adjudication for disagreement.

Reviewers are also accountable.

Frequently asked question: how do you evaluate creative translation?

Use criteria aligned to purpose:

meaning,

effect,

voice,

brand,

or literary function.

Rigid literal scoring may be inappropriate.

Frequently asked question: how do you evaluate legal translation?

Prioritise:

rights,

obligations,

scope,

defined terms,

conditions,

entities,

and legal drafting conventions.

Frequently asked question: how do you evaluate medical translation?

Prioritise:

clinical meaning,

dose,

frequency,

units,

warnings,

terminology,

and patient usability.

Frequently asked question: how do you evaluate software localisation?

Prioritise:

string function,

terminology,

variables,

UI fit,

locale,

and product behaviour.

Frequently asked question: can a translation pass with minor errors?

Depending on project threshold, yes.

Known errors should normally be corrected before release when practical.

Frequently asked question: should critical errors be tolerated?

High-risk systems often use zero tolerance for critical errors.

Frequently asked question: why is source quality part of translation quality?

Poor source can create ambiguity and inconsistency in every language.

Upstream quality reduces downstream defects.

Frequently asked question: what is the best quality model?

There is no one universal model.

The best model is explicit, risk-based, repeatable and fit for the content.

The core quality doctrine

Specify what matters.

Prevent avoidable error.

Detect remaining error.

Classify clearly.

Measure consistently.

Release proportionately.

Correct transparently.

Learn continuously.

That is the Translation Quality System.

Release confidence architecture

A translation can pass individual checks and still lack release confidence if the process evidence is incomplete.

Release confidence should answer:

Do we know the source version?

Do we know the target locale?

Do we know who produced or reviewed the translation?

Were the required quality gates completed?

Did the output meet the defined acceptance threshold?

Are critical risks closed?

Has the target been checked in its final environment?

Confidence is built from evidence.

Confidence layer 1: source confidence

Questions:

Is the source final?

Complete?

Correct version?

Readable?

Are known source defects documented?

If source confidence is low, translation confidence cannot be high.

Confidence layer 2: specification confidence

Is the brief clear?

Audience?

Purpose?

Risk?

Terminology?

Style?

Locale?

If specifications are unclear, quality decisions are unstable.

Confidence layer 3: production confidence

Was the translator or engine appropriate?

Were resources current?

Were queries resolved?

Was draft complete?

Confidence layer 4: review confidence

Did required:

self-revision,

bilingual revision,

monolingual review,

subject review

occur?

Were reviewers qualified?

Confidence layer 5: QA confidence

Were:

numbers,

terms,

tags,

placeholders,

links,

and formatting

checked?

Confidence layer 6: evaluation confidence

If formal evaluation was required:

Was sampling documented?

Were evaluators calibrated?

Did score meet threshold?

Any critical error?

Confidence layer 7: context confidence

Was target checked in final:

web,

app,

PDF,

video,

form,

or other environment?

Confidence layer 8: approval confidence

Who approved release?

Is the decision recorded?

Confidence layer 9: correction confidence

If error appears, can it be:

found,

fixed,

propagated,

and prevented?

A system without correction path has fragile confidence.

Release-confidence statement

For high-risk work, a concise internal statement can read:

Source v4.2 translated to de-DE using Termbase 7.

Full bilingual revision completed.

Legal SME review completed.

Automated QA passed.

20% formal ISO 5060-style sample passed with no critical errors.

In-context PDF review passed.

Approved for release by Legal Localization Lead.

This is actionable confidence.

Confidence is not perfection

No statement should claim zero possibility of error.

The goal is justified release.

Confidence degradation

Confidence can degrade after release.

New law.

Product change.

Term update.

Source correction.

Style change.

A previously trusted translation may need revalidation.

Freshness metadata

For long-lived content, record:

last review date,

source version,

term policy version,

and owner.

Revalidation triggers

Trigger re-review when:

source changes,

regulation changes,

product changes,

rebrand,

new locale policy,

critical incident,

or major tool migration.

Translation quality incident management

Quality incidents should be handled like production incidents.

Detection.

Triage.

Containment.

Correction.

Verification.

Root cause.

Preventive action.

Closure.

Detection sources

User report.

Reviewer.

Monitoring.

Legal team.

Customer support.

Translator.

Automated QA.

Search audit.

Incident triage

Ask:

Is anyone at risk?

Are rights affected?

Is money affected?

Is security affected?

How many users?

How many languages?

How long live?

Incident levels

Severity 1:

critical immediate harm or major legal/security consequence.

Severity 2:

significant functional or reputational impact.

Severity 3:

moderate user misunderstanding.

Severity 4:

minor language defect.

These operational incident levels can differ from translation-error severity but should map clearly.

Containment

Options:

unpublish page,

disable locale,

roll back translation,

replace warning,

block affected feature,

or publish corrected file.

Contain first when harm is possible.

Correction

Use independent verification.

Do not let the same unchecked process that created the defect approve the correction alone.

Propagation audit

Search:

same term,

same string,

same TM unit,

same source template,

same model prompt,

same vendor batch,

other locales.

One visible defect may be systemic.

Root-cause analysis example 1: wrong term

Finding:

wrong term in 20 pages.

Root cause:

deprecated term remained in master TM.

Corrective action:

repair TM,

add forbidden-term QA,

update pages.

Root-cause analysis example 2: wrong number

Finding:

5 mL instead of 0.5 mL.

Root cause:

OCR extraction removed decimal.

Corrective action:

fix OCR pipeline,

add numeric comparison,

recheck affected scans.

Root-cause analysis example 3: wrong locale

Finding:

fr-FR target published on fr-CA page.

Root cause:

deployment mapping error.

Corrective action:

fix locale routing,

audit release process.

Not a translator error.

Root-cause analysis example 4: omission

Finding:

one bullet missing in target.

Root cause:

source file filter skipped text box.

Corrective action:

fix extraction and re-run all files.

Root-cause analysis example 5: reviewer overcorrection

Finding:

legal meaning changed during “style” review.

Root cause:

monolingual reviewer edited controlled legal language without source.

Corrective action:

role boundary and permissions.

Corrective action versus preventive action

Corrective:

fix existing defect.

Preventive:

change system to stop recurrence.

Both matter.

CAPA thinking

Regulated industries often use corrective and preventive action concepts.

Translation programmes can borrow the mindset without unnecessary bureaucracy.

Ask:

what happened,

why,

what fixes now,

what prevents later.

Incident closure

Close only when:

live target fixed,

reusable assets fixed,

root cause addressed,

affected scope checked,

and owner accepts closure.

Post-incident learning

Share:

anonymised example,

quality rule,

and updated process.

Do not repeat same mistake privately across teams.

Quality and change management

Every translation system changes.

New model.

New vendor.

New termbase.

New style guide.

New CMS.

New reviewer.

Changes can improve or degrade quality.

Change-risk assessment

Before change:

what can break?

which content affected?

what benchmark?

what rollback?

Then pilot.

New vendor quality plan

Week 1:

high sampling.

Week 2:

feedback.

Month 1:

calibration.

After stability:

reduce sample according to evidence.

New translator quality plan

Assign representative but controlled work.

Provide feedback quickly.

Do not wait for large failure.

New reviewer quality plan

Use calibration set.

Compare with senior evaluator.

Monitor preference edits.

New MT engine quality plan

Benchmark.

Pilot.

Increased sampling.

Monitor incidents.

New terminology quality plan

Test high-frequency contexts.

Search old term.

Update TM.

Run QA.

New style guide quality plan

Identify rules that affect existing exact matches.

Update reviewers.

New CMS quality plan

Round-trip sample.

Check missing content, links, metadata and locale.

New DTP process quality plan

Pilot export.

Check fonts, RTL, tables, page breaks.

Quality regression suite

Maintain fixed examples of:

critical terms,

numbers,

conditions,

UI strings,

and known historic errors.

Re-run after change.

Golden segments

A set of approved high-risk translations can function as regression anchors.

If new workflow changes them unexpectedly, investigate.

Quality canary

Release small amount to limited audience first where possible.

Monitor.

Then scale.

Translation quality and vendor contracts

Contracts can define:

specifications,

review obligations,

quality thresholds,

correction timelines,

and confidentiality.

Avoid vague “industry standard quality.”

Service acceptance

Client should know how to accept or reject delivery.

Evidence-based process reduces disputes.

Dispute resolution

If vendor and client disagree:

return to source,

brief,

taxonomy,

severity,

and authority.

Not reputation.

Quality penalties

Financial penalties can create perverse incentives if poorly designed.

Focus first on correction and prevention.

Quality bonuses

Reward consistent quality, useful queries and improvement.

Do not reward hiding errors.

Quality and procurement scoring

Procurement can evaluate:

price,

domain expertise,

security,

process,

technology,

and quality evidence.

Lowest price alone can create downstream cost.

Quality and outsourcing

Outsourced translation still needs internal quality ownership.

Vendor is part of system.

Quality and in-house teams

In-house teams need calibration too.

Familiarity can hide assumptions.

Quality and freelancers

Freelancers can deliver high quality with clear briefs, references and review.

Organisation size does not define quality.

Quality and crowdsourcing

Crowdsourcing can work for some low-risk community content.

Needs:

voting,

moderation,

terminology,

and review.

Not suitable for every risk.

Quality and community translation

Community expertise may be authoritative for certain languages and cultural terms.

Quality governance should include community voice.

Quality and literary editors

Literary quality may require editor-translator dialogue rather than error counting.

Use appropriate model.

Quality and transcreation reviewers

Review against:

brand promise,

claim,

tone,

and effect.

Not literal correspondence alone.

Quality and interpreting

Interpreting requires different evaluation architecture:

accuracy,

delivery,

real-time constraints.

Keep separate future node.

Translation-quality research

Research can compare:

human,

MT,

post-edited MT,

different prompts,

different reviewers.

Use transparent methodology.

Evaluation dataset design

Dataset should represent actual content.

Avoid cherry-picked easy segments.

Annotation guideline

Formal research needs precise category definitions.

Adjudication in research

Resolve evaluator disagreement.

Report agreement.

Benchmark leakage

If model trained on reference translations, benchmark may overstate generalisation.

Consider data provenance.

Quality and generative AI evaluation

LLMs can produce several valid targets.

Reference-overlap metrics become even less complete.

Human analytic evaluation remains important.

LLM-as-judge caution

An LLM can help evaluate translation.

Risks:

bias,

inconsistency,

hallucinated errors,

shared model assumptions.

Use with human validation.

AI-assisted evaluator

Potential workflow:

AI flags likely errors.

Human confirms category and severity.

This can scale attention.

AI evaluation prompt

Provide:

source,

target,

specification,

termbase,

error taxonomy.

Ask for candidates, not final authority.

Quality model ensembles

One deterministic QA.

One AI detector.

One human reviewer.

Different systems catch different errors.

Escalation based on agreement

If tools agree clean:

lower priority.

If tools disagree:

human inspect.

This is one possible triage strategy.

Quality economics in AI era

Generation cost falls.

Evaluation cost becomes larger share.

Invest in:

sampling,

automation,

and expert routing.

Quality bottleneck shift

Old bottleneck:

producing target.

New bottleneck:

verifying target at scale.

This is why Translation Quality System deserves its own architecture node.

Quality and trust

Users trust translated content because they cannot compare source.

Quality system protects that asymmetry.

Quality and accountability

Every important release should have an accountable owner.

Tools cannot own accountability.

Quality and transparency inside teams

People should know:

why a segment failed,

what severity means,

how score calculated.

Opaque quality scoring creates distrust.

Quality and fairness to translators

Do not penalise:

source defects,

preference edits,

or wrong reference data

as translator errors.

Measurement fairness improves behaviour.

Quality and fairness to reviewers

Reviewers need:

time,

source access,

and clear rubrics.

Unrealistic review speed lowers evaluation quality.

Quality and fairness to vendors

Compare equivalent projects.

Adjust for:

source quality,

domain,

and risk.

Quality and fairness to users

Do not lower quality because target language has fewer users if consequences remain high.

Language access is a quality principle.

The quality-policy document

A concise policy can state:

quality objectives,

risk classes,

review levels,

critical error rules,

rubric,

roles,

sampling,

incidents,

and improvement.

Example risk class A

Critical:

medical, legal rights, safety.

Review:

full bilingual + subject expert.

Critical error tolerance:

zero.

Risk class B

High:

public operational content, financial, security.

Review:

full bilingual revision.

Risk class C

Medium:

public information, marketing, education.

Review:

revision or monolingual review based on task.

Risk class D

Low:

internal gist, ephemeral communication.

Review:

self-check or sampling.

Risk classes are project choices

Do not universalise letters.

Use whatever system people understand.

Quality exception process

If deadline prevents normal review:

document exception,

risk owner approves,

define compensating control,

schedule post-release review.

Never hide skipped gate.

Compensating control

Example:

urgent alert cannot receive full revision.

Use approved template,

second rapid checker,

and post-release audit.

Risk is managed, not ignored.

Quality escalation tree

Critical error found:

stop release.

Major pattern:

expand review.

Specification conflict:

escalate to requester.

Term conflict:

terminologist.

Domain uncertainty:

SME.

Reviewer disagreement:

language lead.

Technical failure:

engineer.

Quality ownership RACI

Responsible:

does task.

Accountable:

owns outcome.

Consulted:

provides expertise.

Informed:

needs status.

Large programmes may use RACI to clarify quality responsibility.

Quality and handoffs

Each handoff should include:

version,

status,

open issues,

and next owner.

Handoff loss creates defects.

Quality status labels

Draft.

Translated.

Self-reviewed.

Revised.

Reviewed.

QA passed.

Approved.

Released.

Corrected.

Labels make workflow visible.

Quality and traceability

For high-risk content, maintain history.

Who changed critical clause?

Why?

When?

Traceability supports investigation.

Quality and immutable audit records

Some regulated systems need controlled records.

Follow applicable policy.

Do not overengineer ordinary content.

Quality and documentation proportion

Document more when:

risk high,

regulation,

or complex team.

Document less when:

low-risk simple content.

Proportion is a quality principle.

Final confidence matrix

High source clarity + strong process + clean evaluation = high confidence.

Low source clarity + strong process = still query risk.

High source clarity + weak process = uncertain output.

Clean sample + poor sampling design = false confidence.

Strong score + unresolved critical query = do not release.

Confidence integrates evidence.

The quality manager’s final question

Not:

What is the score?

But:

What decision does this evidence justify?

That is the purpose of evaluation.

Worked evaluation case 1: a medical instruction

Source:

“Take one tablet every six hours as needed. Do not exceed four tablets in 24 hours.”

Target A:

Natural language, but says “take four tablets daily.”

Evaluation:

Major or critical accuracy error.

Why:

maximum daily limit became the routine schedule.

Target B:

Preserves one tablet every six hours and maximum four in 24 hours, but contains one awkward collocation.

Evaluation:

Meaning accurate.

One minor target-language error.

Release:

Correct collocation, then approve.

Lesson:

severity matters more than surface polish.

Worked evaluation case 2: a legal clause

Source:

“The tenant may terminate this agreement if rent remains unpaid for 30 days.”

Target A:

Uses equivalent of “must terminate.”

Category:

accuracy — modality.

Severity:

critical or major depending on legal effect.

Target B:

Uses correct permission but varies defined “agreement” term later.

Category:

terminology consistency.

Severity:

major.

Lesson:

legal quality is controlled logic plus terminology.

Worked evaluation case 3: research abstract

Source:

“The findings suggest a possible association.”

Target:

“The findings prove a relationship.”

Categories:

accuracy — epistemic force.

Severity:

major.

Target prose may be elegant.

Quality fails because evidence status changed.

Worked evaluation case 4: software button

Source:

“Archive”

Context:

verb action moving item to archive.

Target is noun meaning archive.

Category:

functional accuracy.

Severity:

major because button action unclear.

In-context review catches faster than sentence evaluation.

Worked evaluation case 5: marketing headline

Source:

“Save up to 30%.”

Target:

“Save 30%.”

Category:

accuracy — quantifier / claim.

Severity:

major because advertising claim strengthened.

Worked evaluation case 6: public-service form

Source:

“Applicants must be at least 18.”

Target:

“Applicants must be 18.”

Category:

accuracy — threshold.

Severity:

major.

Eligibility changed.

Worked evaluation case 7: technical warning

Source:

“Disconnect power before removing cover.”

Target reverses order.

Category:

accuracy — sequence.

Severity:

critical.

Worked evaluation case 8: museum label

Source:

“Probably made in the late 17th century.”

Target:

“Made in 1690.”

Category:

accuracy — uncertainty and invented precision.

Severity:

major scholarly.

Worked evaluation case 9: child-facing text

Target uses advanced academic vocabulary while source is simple.

Meaning correct.

Category:

style / audience suitability.

Severity:

major if children cannot understand.

Fit for purpose fails.

Worked evaluation case 10: internal gist translation

Target has clumsy grammar but preserves core information accurately.

Purpose:

internal understanding.

Evaluation:

May pass low-risk gist threshold.

Same target would fail publication threshold.

Worked evaluation case 11: subtitle

Target faithfully includes every source word but exceeds reading speed severely.

Category:

functional / audiovisual.

Severity:

major.

Literal completeness can reduce subtitle quality.

Worked evaluation case 12: poetry

Translator changes word order and chooses nonliteral phrase to preserve rhyme and image.

Segment-level literal checker flags mismatch.

Expert evaluation finds artistic equivalence.

Lesson:

quality rubric must fit genre.

Worked evaluation case 13: website metadata

Body translation excellent.

Meta title remains source language.

Category:

completeness / localisation.

Severity:

major for discoverability.

Worked evaluation case 14: table

Every cell translated correctly but one target shifted into adjacent row.

Category:

format / data relation.

Severity:

critical in financial table.

Worked evaluation case 15: form placeholder

Target contains translated “{email}” variable.

Category:

technical integrity.

Severity:

major.

Worked evaluation case 16: personal name

Target changes official spelling of person name.

Category:

entity accuracy.

Severity:

major.

Worked evaluation case 17: casual conversation

Source joke translated literally.

Reader misses humour.

Category:

pragmatics / style.

Severity:

depends on purpose.

In entertainment, major.

In rough gist, minor.

Worked evaluation case 18: academic citation

Target translates article title in bibliography against project citation policy.

Category:

style / citation integrity.

Severity:

minor or major depending on traceability.

Worked evaluation case 19: emergency alert

Target is grammatically perfect but omits location.

Category:

omission.

Severity:

critical because action cannot be directed.

Worked evaluation case 20: cybersecurity advisory

Source:

“may allow remote code execution under specific conditions.”

Target:

“allows remote code execution.”

Category:

accuracy — uncertainty + condition.

Severity:

major.

Worked evaluation case 21: insurance exclusion

Source:

“Coverage does not apply to…”

Target:

“Coverage may not apply to…”

Firm exclusion becomes uncertainty.

Category:

accuracy — modality.

Severity:

major.

Worked evaluation case 22: e-commerce specification

Source:

“Compatible with X200 and X220 only.”

Target omits “only.”

Category:

accuracy — exclusivity.

Severity:

major because customers may buy incompatible item.

Worked evaluation case 23: education assessment

Source command:

“Compare.”

Target command means “describe.”

Category:

functional accuracy.

Severity:

major because task construct changes.

Worked evaluation case 24: survey

Source response scale:

“Strongly disagree / disagree / neutral / agree / strongly agree.”

Target middle item becomes “uncertain.”

Category:

measurement equivalence.

Severity:

major because response construct changes.

Worked evaluation case 25: historical quotation

Target modernises offensive historical term to neutral modern language.

If project purpose is historical evidence, category:

accuracy / register.

Severity:

major.

If project specification requires modernised adaptation, it may be correct.

Specifications decide.

Worked evaluation case 26: support article

Source tells user to reset router only after three checks.

Target moves reset instruction to opening.

Category:

sequence / functional.

Severity:

major.

Worked evaluation case 27: accessibility

Visible UI translated.

Screen-reader ARIA label remains source language.

Category:

accessibility / completeness.

Severity:

major.

Worked evaluation case 28: bilingual government publication

One language version omits an entire annex.

Category:

completeness.

Severity:

critical if annex carries legal effect.

Worked evaluation case 29: term variation

Source uses same defined technical term 20 times.

Target has 4 synonyms.

Category:

terminology.

Severity:

major if concept identity matters.

Worked evaluation case 30: source variation

Source uses one ordinary verb in several senses.

Target uses different appropriate verbs.

Automated consistency tool flags variation.

Human evaluator rejects warning as false positive.

Lesson:

consistency is semantic, not mechanical sameness.

Worked evaluation case 31: MT hallucination

Target adds explanatory sentence absent from source.

Category:

accuracy — addition.

Severity:

major.

Worked evaluation case 32: human embellishment

Human translator adds rhetorical flourish.

Same category as MT hallucination if unsupported.

Origin does not change output defect.

Worked evaluation case 33: reviewer damage

Initial target correct.

Reviewer changes “may” to “must” for style.

Released target wrong.

Root cause:

reviewer error.

Quality system evaluates final output and process.

Worked evaluation case 34: termbase defect

All translators consistently use wrong target because approved termbase is wrong.

Output errors are terminology defects.

Root cause is central resource.

Quality management should repair upstream.

Worked evaluation case 35: source defect

Source says 2025 in paragraph and 2026 in table.

Target follows paragraph.

Evaluator should flag source inconsistency, not automatically penalise translation without policy.

Worked evaluation case 36: DTP defect

Translation correct in source file.

Final PDF drops last line.

Output has omission.

Root cause is production.

Release QA must catch it.

Worked evaluation case 37: locale mismatch

Spanish translation correct but uses Spain terminology for Mexican audience.

Category:

locale.

Severity:

depends on functional effect.

Worked evaluation case 38: cultural mismatch

Target example is understandable but culturally inappropriate for target audience.

Category:

style/localisation.

Severity:

based on purpose.

Worked evaluation case 39: wrong language variant fallback

fr-CA page serves fr-FR translation.

Category:

localisation deployment.

Not translator defect.

Worked evaluation case 40: empty string

One interface string untranslated.

Category:

completeness / technical.

Severity:

major if UI action blocked.

The evaluation handbook for reviewers

Before evaluation:

read specifications.

Know target locale.

Know error taxonomy.

Know severity rules.

Know whether you are evaluating final output or draft.

Know whether correction is allowed during evaluation.

Then begin.

Reviewer pass 1: source-target meaning

Ignore style preferences.

Find:

mistranslation,

omission,

addition,

logic,

reference,

numbers,

names.

Reviewer pass 2: terminology

Use authoritative termbase.

Check concept, not only string.

Reviewer pass 3: target-language conventions

Grammar.

Syntax.

Collocation.

Punctuation.

Spelling.

Reviewer pass 4: style and audience

Tone.

Register.

Readability.

Genre.

Reviewer pass 5: locale and format

Dates.

Currency.

Addresses.

Tags.

Placeholders.

Layout.

Reviewer pass 6: severity

After identifying type, ask consequence.

Do not decide severity before understanding error.

Reviewer pass 7: report systemic patterns

If same issue repeats, note root-cause candidate.

Evaluator self-check

Before finalising score:

Did I penalise preferences?

Did I apply severity consistently?

Did I double-count one defect?

Did I ignore source defect?

Did I follow project specification?

This reduces evaluator noise.

Adjudicator checklist

When two evaluators disagree:

What is exact source meaning?

What does brief require?

Is target acceptable language?

Which taxonomy definition applies?

What consequence?

Then decide.

Building an evaluator calibration corpus

Collect real anonymised examples.

Tag with agreed:

category,

severity,

rationale.

Include borderline cases.

Update after policy changes.

This corpus becomes institutional quality memory.

The quality of the rubric itself

A rubric can be defective.

Signs:

reviewers disagree constantly,

categories overlap,

severity unclear,

results do not correlate with user problems,

teams game scores.

Audit rubric.

Rubric simplification

If 50 categories produce poor agreement, reduce.

More detail is not automatically better.

Rubric expansion

If one category “accuracy” hides important patterns such as numbers and names, add useful subcategories.

Error taxonomy lifecycle

Draft.

Pilot.

Calibrate.

Use.

Review.

Revise.

Rubrics evolve.

Quality and machine learning data

Quality annotations can train models.

But inconsistent annotations create noisy training data.

Calibration matters beyond human management.

Quality labels for AI

If using evaluation data for model training, preserve:

source,

target,

category,

severity,

context,

and adjudication.

Privacy in quality data

Evaluation examples may contain sensitive content.

Anonymise where possible.

Control access.

Quality and vendor confidentiality

Do not publish vendor defect examples publicly without permission.

Use anonymised internal training.

Quality and translator dignity

Feedback should address output, not character.

“This clause changes obligation” is professional.

“You don’t understand legal translation” is not useful.

Quality and learning culture

People improve when feedback is specific, fair and consistent.

Translation-quality coaching

For repeated errors, assign focused practice.

Pronoun reference.

Modal verbs.

Numbers.

Terminology.

Do not prescribe generic “be more careful.”

Targeted coaching example

Problem:

repeated loss of hedging.

Exercise:

translate 20 academic sentences containing:

may,

might,

suggest,

appears,

likely.

Review epistemic scale.

Targeted coaching: numbers

Exercise:

dates,

ranges,

percentages,

units,

currency.

Use hard-detail checklist.

Targeted coaching: terminology

Build concept definitions.

Use termbase.

Compare near-synonyms.

Targeted coaching: target language

Read monolingual target corpus.

Rewrite translationese.

Quality and professional development

Review data can guide individual learning plans.

Use patterns over time.

Quality and certification exams

Formal translation exams often evaluate accuracy and target language under time constraints.

The same principles apply, though professional production allows more resources.

Quality and classroom grading

Teachers should explain:

why points lost,

error category,

severity.

Opaque marks teach less.

Quality and peer review

Students can review each other with simplified rubric.

This develops metalinguistic awareness.

Quality and self-assessment

Translators can score sample of own work.

Compare with reviser.

This calibrates self-review.

Quality and reflection

After project:

What error surprised me?

What term did I learn?

What process would I change?

Reflection converts work into expertise.

Frequently asked question: what is the difference between QA and QC?

Usage varies.

A practical distinction is:

quality assurance = planned system to prevent and detect issues.

quality control = specific checking activities.

Many organisations use terms differently.

Define your local usage.

Frequently asked question: what is the difference between revision and evaluation?

Revision aims to improve target through source-target comparison.

Evaluation aims to assess output against criteria, sometimes without correcting it.

Frequently asked question: what is the difference between review and proofreading?

Review addresses target communication.

Proofreading focuses on final surface/production defects.

Frequently asked question: can one person do all quality steps?

For low-risk work, one person can do translation plus self-revision and QA.

Independent checks add value as risk increases.

Frequently asked question: how many reviewers are enough?

Enough to control project risk.

More reviewers can create conflict.

Use roles, not arbitrary headcount.

Frequently asked question: should reviewers rewrite style?

Only when target violates specification, language norms or reader needs.

Do not rewrite valid language merely for personal taste.

Frequently asked question: how do you score repeated errors?

Choose a documented policy.

Count all, cap, or treat systemically.

Consistency matters.

Frequently asked question: how do you compare vendors fairly?

Use comparable content, same specifications, same sample method and calibrated evaluators.

Frequently asked question: how do you compare MT engines fairly?

Use same source, context, glossary, evaluation rubric and blinded output where possible.

Frequently asked question: can a translation be accurate but fail?

Yes.

It can be unreadable, wrong locale, technically broken or unfit for audience.

Frequently asked question: can a translation be natural but fail?

Yes.

It may add, omit or distort source meaning.

Frequently asked question: what is a quality threshold?

A predefined acceptance boundary, such as maximum weighted errors per 1,000 words and zero critical errors.

Frequently asked question: what is sampling in translation evaluation?

Evaluating selected portions of a larger translation to estimate or monitor overall quality.

Frequently asked question: when should you review everything?

When risk, regulation or content consequence justifies full review.

Frequently asked question: what is risk-based sampling?

Selecting additional evaluation items because they have higher potential consequence, such as warnings, legal conditions or numbers.

Frequently asked question: what is evaluator calibration?

Training reviewers to apply categories and severity consistently using shared examples.

Frequently asked question: what is inter-rater agreement?

The degree to which different evaluators make consistent judgments.

Low agreement signals measurement problems.

Frequently asked question: what is quality debt?

Known unresolved translation risk carried forward, such as stale content or unreviewed legacy strings.

Frequently asked question: what is a translation quality incident?

A translation defect in production significant enough to require organised response and corrective action.

Frequently asked question: should every error be corrected?

Known production errors generally should be corrected when practical, prioritised by severity and impact.

Frequently asked question: what does “fit for purpose” mean?

The translation performs the job required for its intended audience and use while meeting relevant accuracy and quality requirements.

Frequently asked question: what is the best translation-quality framework?

The best framework is one that is explicit, risk-based, trained, reproducible and aligned to the project.

ISO 5060 and MQM provide strong current reference points.

Frequently asked question: can quality be automated?

Parts can.

Numbers, tags, terminology and placeholders can be checked automatically.

Meaning and context still require human judgment in many cases.

Frequently asked question: can AI evaluate translation quality?

AI can flag potential issues and assist triage.

Human confirmation remains important, especially for high-risk decisions.

Frequently asked question: why do quality systems need terminology?

Consistent approved terminology reduces conceptual drift and makes evaluation objective.

Frequently asked question: why do quality systems need translation memory governance?

Old errors can be reused automatically.

Quality must control the memory feeding future work.

Frequently asked question: why do quality systems need source feedback?

Repeated translation problems often originate in source ambiguity or inconsistency.

Frequently asked question: why does locale matter to quality?

A translation can be correct language but wrong regional convention, institution or term.

Frequently asked question: why does final layout matter?

Users consume rendered output, not CAT segments.

Layout can hide, truncate or misalign meaning.

The professional reference frame

ISO 5060:2024 is currently published as “Translation services — Evaluation of translation output — General guidance.”

Its public abstract states that it covers human translation, post-edited machine translation and unedited machine translation output; evaluator qualifications and competences; sampling; and an analytic approach based on error types, penalty points, error scores and quality ratings.

The European Commission’s current translation-quality framework states that its translations should be accurate, clear and fit for purpose, and it uses a risk-based approach tied to document type, intended use and reader expectations.

The Commission also distinguishes bilingual revision from monolingual review and describes quality as a shared responsibility across quality managers, language communities, workflow managers, translators and formal quality checks.

MQM provides a current community framework for multidimensional analytic translation evaluation, with configurable error typologies and scoring models.

These sources support the architecture here without requiring every organisation to copy one exact scoring formula.

Continue through the Master Art of Translation architecture

Root:

Master Art of Translation | The Complete System for Moving Meaning Between Languages.

Then:

Source Analysis.

Equivalence.

Context Stack.

Translation Unit.

Terminology System.

Translation Memory System.

Machine Translation System.

Human Translation System.

This Translation Quality System is the layer that determines whether output from any of those workflows is safe, fit and ready to use.

Vocabulary route

When quality defect comes from:

wrong sense,

collocation,

near-synonym,

register,

or lexical precision,

route to the Vocabulary Learning Hub for deeper lexical learning.

How English Works route

When quality defect comes from:

English syntax,

tense,

modality,

reference,

cohesion,

or information structure,

route to How English Works.

Terminology route

When quality defect is concept-to-term control, route to the Terminology System.

Translation Memory route

When defect comes from reused historical language, route to the Translation Memory System.

Machine Translation route

When defect comes from automated generation, prompt, model or post-editing process, route to the Machine Translation System.

Human Translation route

When defect comes from translator/reviser workflow, assignment, query handling or professional review, route to the Human Translation System.

The final quality equation

Conceptually:

translation quality =
clear specifications
× source understanding
× competent production
× authoritative terminology
× context
× appropriate review
× technical integrity
× risk-aware evaluation
× correction learning.

This is not arithmetic.

It is a reminder that quality can collapse when one essential layer fails.

Final synthesis

Translation quality is not a feeling.

It is not fluency alone.

It is not a perfect score.

It is not a reviewer making many edits.

It is not an expensive vendor.

It is not a human-versus-machine label.

Quality is the degree to which a target text, produced through an appropriate process, satisfies explicit requirements for meaning, reader, function and risk.

That requires a system.

The system begins with specifications.

It identifies what can go wrong.

It separates error type from severity.

It uses qualified people.

It automates deterministic checks.

It evaluates output with a transparent rubric.

It samples intelligently when full evaluation is impractical.

It blocks critical defects.

It checks the final product.

It records release evidence.

It corrects incidents.

It repairs upstream resources.

It learns.

The most mature quality question is therefore not:

“Is this translation good?”

It is:

“What evidence shows this translation is fit for this purpose, for this reader, at this risk level, and ready to release?”

When the organisation can answer that clearly, translation quality stops being an argument and becomes an architecture.


Architecture links

Master Translation root: Master Art of Translation | The Complete System for Moving Meaning Between Languages

Source Analysis: How to Read a Source Text Before You Translate It

Equivalence: How Equivalence Works When Languages Do Not Match One-to-One

Context Stack: The Context Stack

Translation Unit: The Translation Unit

Terminology System: The Terminology System

Translation Memory System: The Translation Memory System

Machine Translation System: The Machine Translation System

Human Translation System: The Human Translation System

Vocabulary: Vocabulary Learning Hub

English: How English Works

External professional references

ISO 5060:2024 — Translation services — Evaluation of translation output — General guidance

European Commission — Translation quality

Multidimensional Quality Metrics — MQM Council

Final release-decision handbook

A translation-quality system ultimately exists to support decisions.

Release now.

Correct first.

Expand review.

Escalate.

Reject.

Archive.

Re-translate.

Those decisions should follow evidence rather than mood.

Decision: release now

Appropriate when:

no critical errors,

known major errors corrected,

required review complete,

QA passed,

context check complete,

specifications satisfied.

Minor stylistic alternatives should not block release indefinitely.

Decision: correct first

Appropriate when:

one or more major errors are known but localised and fixable.

Correct.

Verify.

Then release.

Decision: expand review

Appropriate when sample reveals:

repeated error pattern,

unknown scope,

or one suspicious critical area.

Do not assume sampled defect is isolated.

Decision: full re-review

Appropriate when:

critical error suggests systemic misunderstanding,

wrong source version,

wrong locale,

or bad terminology resource.

The cost of broader review is justified.

Decision: reject and retranslate

Appropriate when:

target is fundamentally misaligned,

post-editing would cost more than fresh translation,

or trust in source-target relation is broken.

Decision: quarantine

Useful for:

legacy TM,

unknown-provenance translation,

or vendor batch under investigation.

Do not let uncertain content feed new work.

Decision: archive

Appropriate when content remains historically valid but should not drive current production.

Decision: retire

Appropriate when:

content obsolete,

risk high,

and no operational reason to preserve it in active system.

Release decision and residual risk

Every release contains residual risk.

The question is whether residual risk is acceptable under specifications.

Quality management makes that decision explicit.

Quality debt in detail

Quality debt is any known weakness postponed rather than resolved.

Examples:

unreviewed old pages,

temporary MT still live,

deprecated term in archive,

missing locale review,

weak context strings,

or inconsistent glossary.

Quality debt is not automatically failure.

Unmanaged quality debt is.

Quality-debt register

Fields:

item,

locale,

content type,

risk,

user exposure,

owner,

date identified,

planned action,

target date.

This prevents known problems from disappearing.

Quality-debt priority

Priority can combine:

severity,

traffic,

legal/safety consequence,

and likelihood of reuse.

A low-traffic typo can wait.

A stale safety instruction cannot.

Debt interest

Quality debt can grow.

Old bad target enters TM.

TM reuses it.

New pages inherit it.

One issue becomes twenty.

This is quality interest.

Debt prevention

Do not promote temporary output to authoritative memory.

Label provisional content.

Use expiry dates.

Audit legacy reuse.

Expiry policy

Some translations should have expiry or review date.

Examples:

temporary policy,

time-sensitive travel advice,

legal notice,

product beta.

Freshness is quality.

Quality release calendar

Long-lived multilingual sites can schedule:

monthly critical-page checks,

quarterly term audits,

annual legacy review.

The cadence depends on change rate.

Evergreen content

Stable educational concept may need less frequent re-review.

Still revisit when terminology or standards change.

High-change content

Software help.

Policies.

Prices.

Regulations.

Needs stronger freshness triggers.

Quality owner per content family

Assign owner for:

legal pages,

product UI,

support,

marketing,

education.

Ownership accelerates corrections.

Evaluator governance

Evaluators should have a code of practice.

Apply rubric.

Separate preference.

Document evidence.

Respect source.

Respect target norms.

Escalate uncertainty.

Avoid personalisation of feedback.

Evaluator conflict of interest

A person evaluating their own translation may be less independent.

Self-review is useful.

Formal evaluation may require another person.

Evaluator domain conflict

A linguist may lack subject expertise.

Use SME input for technical disputes.

Evaluator language direction

A target-language expert may judge style well.

Source-language competence needed for fidelity.

Team evaluation can divide roles.

Evaluator workload

Too many segments reduce consistency.

Sampling exists partly because evaluator attention is finite.

Evaluator tool support

Use:

annotation interface,

termbase,

search,

QA reports,

and source context.

Do not force evaluators to work blind.

Evaluator comments

Keep concise and diagnostic.

Bad:

“Awkward.”

Better:

“Target collocation is non-idiomatic; standard target usage is X.”

Bad:

“Wrong.”

Better:

“Changes possibility to obligation.”

Evaluator confidence

Allow:

certain,

probable,

needs adjudication

where useful.

Not every quality call is equally clear.

Adjudication governance

Adjudicator should not merely pick a favourite.

Review:

source,

brief,

target norms,

taxonomy,

and consequence.

Record ruling if reusable.

Quality arbitration with client

Sometimes client insists on wording that linguists consider poor.

Clarify whether issue is:

terminology preference,

legal requirement,

brand voice,

or factual error.

If client wording changes meaning, document risk.

Client-approved error

Approval does not transform an objective meaning error into correct translation.

It may become an accepted client risk.

Document separately.

Quality exception log

For each exception:

what standard rule skipped?

why?

who approved?

what compensating control?

when expires?

Compensating control examples

No bilingual reviser available before emergency release.

Compensation:

use approved template,

second monolingual check,

full bilingual review within two hours.

Exception remains visible.

Quality escalation after release

If user reports issue:

do not dismiss because translation passed QA.

Real-world evidence overrides procedural pride.

Investigate.

User feedback taxonomy

Confusing.

Wrong.

Offensive.

Too formal.

Wrong term.

Broken link.

Wrong number.

Untranslated.

Collect patterns.

Search-log quality signal

Users searching wrong target term may indicate terminology mismatch.

Search analytics can reveal vocabulary problems.

Support-ticket signal

Repeated confusion about translated instruction may indicate usability defect.

Conversion signal

Marketing performance can indicate target-market fit.

Do not equate poor conversion automatically with translation error.

Safety signal

Any report of unsafe action gets immediate linguistic review.

Quality and accessibility feedback

Screen-reader users may reveal labels that visual reviewers miss.

Include diverse user feedback.

Quality and multilingual parity

One language should not systematically receive lower review merely because traffic is lower when consequence is same.

Quality policy should be fair.

Quality and resource constraints

When resources limited, prioritise risk.

Translate fewer critical pages well rather than everything badly.

Quality and staged localisation

Stage 1:

critical journeys.

Stage 2:

high-value content.

Stage 3:

long-tail pages.

Review level can vary.

Quality and fallback language

If translation unavailable, is fallback better than poor MT?

Depends on audience.

Define product policy.

Quality and machine disclosure

Unreviewed MT may need label.

Reviewed professional target may not.

Disclosure is product and policy decision.

Quality provenance metadata

Internal metadata can record:

human,

MT,

post-edited,

reviewed,

sampled,

approved.

This enables risk-aware future reuse.

Quality and translation memory promotion

Only promote target to master TM when required quality state reached.

This prevents quality debt multiplying.

Quality and terminology promotion

New term should have:

concept definition,

target form,

domain,

status,

and authority.

A reviewer correction alone is not always enough to create term policy.

Quality and style promotion

Repeated style correction can become documented rule.

This reduces future evaluator disagreement.

Quality and source-author guidance

Repeated source issue can become authoring rule.

Example:

avoid ambiguous “this” in safety procedures.

Quality feedback improves source.

Quality as organisational memory

A mature quality system remembers:

what failed,

what worked,

what reviewers agreed,

what terms changed,

and why.

It does not merely remember final target strings.

Final 30-point quality release checklist

1. Correct source version.
2. Correct target locale.
3. Purpose defined.
4. Audience defined.
5. Risk defined.
6. Terminology current.
7. Style specification current.
8. Translation complete.
9. Self-revision complete.
10. Required independent revision complete.
11. Required review complete.
12. Required SME check complete.
13. Numbers verified.
14. Names verified.
15. Dates verified.
16. Units verified.
17. Conditions verified.
18. Negation verified.
19. Modality verified.
20. Terminology consistent.
21. Locale conventions correct.
22. Tags/placeholders intact.
23. Format complete.
24. QA passed.
25. Formal evaluation passed if required.
26. Critical errors zero.
27. Open major issues zero.
28. In-context review complete.
29. Approval recorded.
30. Correction path known.

This checklist does not replace professional judgment.

It makes judgment systematic.

The quality-system stop rule

Stop release if:

critical error exists,

source version uncertain,

required review incomplete,

major unresolved query,

wrong locale,

or technical integrity broken.

A stop rule protects teams from deadline pressure.

The quality-system restart rule

Resume when:

problem corrected,

scope checked,

required reviewer confirms,

and gate evidence updated.

Final architecture position

The Translation Quality System is not downstream decoration.

It touches every earlier node.

Source Analysis creates interpretive quality.

Equivalence creates cross-language meaning quality.

Context Stack creates evidence quality.

Translation Unit creates segmentation quality.

Terminology creates concept quality.

Translation Memory creates reuse quality.

Machine Translation creates generation quality.

Human Translation creates accountable production quality.

This node integrates them into release quality.

Closing rule

A translation is ready when the evidence for release is stronger than the residual risk of error.

That evidence comes from:

clear specifications,

competent production,

appropriate review,

transparent evaluation,

technical validation,

and correction readiness.

Quality is therefore not the absence of uncertainty.

It is disciplined control of uncertainty.

Acceptance thresholds and re-evaluation

A quality threshold should answer a practical release question.

Not:

Is this translation flawless?

But:

Has the translation met the required quality level for this use?

That means acceptance criteria should be defined before final evaluation.

Possible rules:

zero critical errors,

maximum weighted error score,

all major errors corrected,

mandatory terminology compliance,

successful in-context test,

or subject-expert approval.

Different projects can combine these.

Threshold design for low-risk content

Low-risk internal content may tolerate:

minor grammar,

awkward style,

or inconsistent punctuation

if meaning remains clear.

Threshold can prioritise semantic accuracy and basic usability.

Threshold design for medium-risk public content

Require:

no critical errors,

very low major-error rate,

clear target language,

terminology consistency,

and final web/product check.

Threshold design for high-risk content

Require:

zero critical errors,

no unresolved major meaning error,

qualified independent revision,

hard-detail checks,

and formal sign-off.

One quality profile should not be copied across all risk classes.

Threshold drift

Teams sometimes lower standards informally because deadlines tighten.

If threshold must change, make it explicit.

Who approved?

Why?

What compensating control?

Hidden threshold drift undermines quality governance.

Re-evaluation after correction

When errors are corrected, do not automatically assume the output now passes.

Ask:

Was the correction local?

Could the same error exist elsewhere?

Did correction introduce new grammar or formatting defects?

If systemic, expand review.

Correction verification

A second person may verify high-risk corrections.

Example:

wrong medical dose fixed.

Independent check confirms source, target, unit and final PDF.

Re-evaluation sampling

After systemic correction, sample affected pattern.

Example:

old term replaced across 500 pages.

Check several grammatical environments.

Do not trust bulk replacement blindly.

Release after rework

A failed vendor delivery can become acceptable after rework.

Keep initial quality metrics separate from final release quality.

This supports fair vendor management and user protection.

Initial quality versus released quality

Initial quality answers:

What did producer deliver?

Released quality answers:

What did users receive after correction?

Both matter.

Quality of corrections

Late corrections can introduce new errors.

Treat correction as translation work.

Review proportional to risk.

Emergency correction

When critical error is live:

speed matters.

Use a small verified correction team.

Contain first.

Then complete deeper investigation.

Post-release verification

After publication, confirm:

right locale live,

right version live,

no cache serving old target,

links correct,

layout intact,

metadata correct.

Release process can fail after linguistic approval.

Search-engine verification

For web content, ensure:

translated title,

meta description,

canonical relationships,

and indexability

match project intent.

Quality includes discoverability when specified.

Analytics verification

If target users abandon a translated form unusually, investigate.

Not every metric drop is translation error.

But analytics can be a quality signal.

Support feedback

Create path for support teams to flag language issues.

Frontline staff see real misunderstandings.

Reader correction form

For public educational or reference content, a simple feedback channel can surface:

typos,

wrong terms,

and outdated facts.

Moderate and verify submissions.

Quality after source change

When source updates:

identify target impact.

Do not leave translations pointing to old version silently.

Version relationships should be trackable.

Stale translation detection

Signals:

source modified date newer,

old product name,

old term,

old legal reference,

or old screenshot.

Freshness checks belong to long-lived quality.

Quality maintenance lifecycle

Create.

Review.

Release.

Monitor.

Refresh.

Correct.

Archive.

Quality continues after publication.

The final acceptance statement

For significant projects, a concise internal statement can document:

quality profile used,

revision/review completed,

evaluation result,

open exceptions,

release approval.

Example:

“Translation evaluated under Public Information profile. Full bilingual revision completed. Automated QA passed. Risk-based sample contained no critical errors and all major errors were corrected. In-context web review passed. Approved for publication.”

This is stronger than:

“Looks good.”

Why quality statements matter

They make release reasoning visible.

They help future teams understand why content was trusted.

They support audits and corrections.

They reduce institutional amnesia.

The final quality principle

Never confuse absence of reported errors with evidence of quality.

Evidence comes from:

specification,

appropriate production,

appropriate checking,

and a release decision tied to risk.

A translation that nobody examined may be perfect.

The organisation does not know that.

A translation that passed a transparent, calibrated, risk-appropriate system has justified confidence.

That is what professional quality architecture should deliver.

Final verification margin: quality after approval

Approval should be the beginning of controlled release, not the end of attention.

Before the translation is considered fully closed, verify one last time that the approved target is the same target users actually receive.

Check the live version.

Check the locale.

Check the filename or page ID.

Check that no later edit bypassed review.

Check that the target did not revert to an older translation memory match.

Check that the page builder, CMS, app build, subtitle export or DTP process did not remove text.

The final quality object is the delivered experience.

Quality after deployment

A translation can pass every linguistic gate and still fail because:

the wrong locale is mapped,

the wrong file is uploaded,

a cache serves the old version,

a variable breaks,

a table shifts,

or a font cannot render required characters.

Post-release verification connects linguistic quality with operational reality.

The final verification owner

Assign one person or role to confirm:

approved version equals deployed version.

This can be:

project manager,

localization engineer,

language lead,

or release owner.

The role matters more than the title.

Close only when evidence is complete

A project can be marked closed when:

required output exists,

required quality checks passed,

approved version is live or delivered,

open critical issues are zero,

and correction paths remain known.

This final closure prevents the common gap between “translation finished” and “translation safely delivered.”

The final standard

The strongest quality systems do not ask translators, reviewers or tools to be infallible.

They assume humans and machines can make errors.

They design the workflow so important errors are unlikely to survive every layer.

That is the real meaning of translation quality assurance:

not perfection by promise,

but reliability by architecture.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading