Localization quality assurance becomes useful when reviewers classify problems consistently enough that the organization can learn from them. A redlined document full of personal preferences may improve one file, but it does not tell you whether recurring defects come from source ambiguity, terminology drift, machine translation, translator training, missing context, product layout or an unstable review process.
Searches for LQA, localization quality assurance, translation error taxonomy, MQM, translation error severity, LQA score, translation quality evaluation, translation quality metrics and localization QA framework describe the need for repeatable analytic evaluation. ISO 5060:2024 addresses human evaluation of translation output using error types and penalty-based approaches, while MQM provides a structured error typology and terminology for analytic quality assessment.
This guide shows how to turn those ideas into an operational LQA system without letting scores become the purpose of translation. It covers error types, severity, neutral preferences, penalties, sampling, critical errors, reviewer calibration, defect evidence, root cause, normalization, thresholds, trend analysis, supplier comparison, corrective action and the important distinction between measuring translation output and improving the process that produced it.
This article belongs to eduKateSG’s Master Art of Translation architecture. It extends the professional workflow layer without replacing the existing owners for terminology, file preparation, release control, regression testing or general translation quality.
Quick answer
A useful LQA model has clear specifications, a manageable error taxonomy, severity definitions tied to user impact, calibrated evaluators, sampling rules, transparent scoring and a corrective-action loop. It should distinguish true errors from preferential edits, keep critical-risk gates visible, and never let one aggregate score replace professional judgment about meaning, safety or fitness for purpose.
- Specify: define what the translation is supposed to achieve.
- Classify: use stable error types with clear examples.
- Grade: tie severity to impact, not reviewer irritation.
- Sample: choose content that represents risk and use.
- Calibrate: train reviewers on shared cases.
- Score carefully: normalize where needed without hiding critical defects.
- Improve: connect error patterns to root cause and corrective action.
1. Begin with specifications, not error labels
An error is a violation of a requirement, not merely something a reviewer dislikes. MQM terminology frames translation errors relative to rules of good writing or translation based on specifications.
Professional method. Document audience, purpose, terminology, style, locale, risk and required fidelity before evaluation begins. The aim is to make the decision repeatable, because localization problems become expensive when the correct fix exists only in one reviewer’s memory.
Failure mode. Reviewers mark personal rewrites as errors because there is no shared target specification. Formal versus conversational wording cannot be judged fairly if the project’s register was never defined.
Verification. For each proposed error, identify the violated requirement. If the result still depends on guesswork, return to the source context, locale requirement, product state or quality specification before approving the translation.
2. Keep the taxonomy small enough to use
A taxonomy should create useful distinctions without becoming an encyclopedia. Too many categories reduce reviewer agreement and make trend data sparse.
Professional method. Start with broad dimensions such as accuracy, terminology, linguistic conventions, style, locale conventions and design/functional issues, then add subtypes only where decisions benefit. The aim is to make the decision repeatable, because localization problems become expensive when the correct fix exists only in one reviewer’s memory.
Failure mode. Reviewers choose among dozens of nearly identical labels. A simple ‘mistranslation’ subtype can be more reliable than forcing a fine theoretical distinction nobody can reproduce.
Verification. Measure how often reviewers disagree between neighbouring categories. If the result still depends on guesswork, return to the source context, locale requirement, product state or quality specification before approving the translation.
3. Separate accuracy from style
A sentence can preserve meaning yet violate style, or sound elegant while changing meaning. These defects need different fixes and different risk treatment.
Professional method. Classify meaning changes under accuracy and target-language presentation under style or linguistic-convention dimensions. The aim is to make the decision repeatable, because localization problems become expensive when the correct fix exists only in one reviewer’s memory.
Failure mode. A rewritten sentence is marked ‘accuracy’ merely because the reviewer prefers another expression. Changing ‘may’ to ‘will’ is accuracy; replacing an acceptable phrase with a smoother synonym may be style or preference.
Verification. Ask whether the source claim itself changed. If the result still depends on guesswork, return to the source context, locale requirement, product state or quality specification before approving the translation.
4. Treat terminology as concept control
Terminology errors are more than vocabulary taste. Approved terms often represent product, legal or technical concepts that must remain stable.
Professional method. Link terminology findings to a governed termbase and distinguish mandatory terms from allowed variants. The aim is to make the decision repeatable, because localization problems become expensive when the correct fix exists only in one reviewer’s memory.
Failure mode. Every synonym becomes a terminology error even though no controlled term exists. Using a deprecated feature name can be a terminology error; using one natural synonym in general prose may not be.
Verification. Point to the concept record or style requirement. If the result still depends on guesswork, return to the source context, locale requirement, product state or quality specification before approving the translation.
5. Use a neutral preference state
Review comments can be useful even when the original is not wrong. MQM terminology recognizes that issues can be preferential changes rather than penalized errors.
Professional method. Allow reviewers to suggest alternatives without assigning error severity when the target already meets specification. The aim is to make the decision repeatable, because localization problems become expensive when the correct fix exists only in one reviewer’s memory.
Failure mode. Preference inflation damages trust in LQA and punishes legitimate variation. Two natural sentences may both be acceptable; one can still be preferred for brand voice.
Verification. Ask whether the existing text violates a requirement before scoring it. If the result still depends on guesswork, return to the source context, locale requirement, product state or quality specification before approving the translation.
6. Define severity through user consequence
Severity should express impact, not how obvious an error looks. A tiny word can reverse a legal condition while a conspicuous typo may be harmless.
Professional method. Define levels such as critical, major, minor and neutral using consequences for meaning, task completion, safety, legal obligation, trust and visibility. The aim is to make the decision repeatable, because localization problems become expensive when the correct fix exists only in one reviewer’s memory.
Failure mode. Long awkward sentences get major penalties while a subtle reversed number is rated minor. A dosage, payment or consent error can be critical even if it is only one character.
Verification. Write the user or business consequence in the severity rationale. If the result still depends on guesswork, return to the source context, locale requirement, product state or quality specification before approving the translation.
7. Keep critical errors outside the average
Some defects should block acceptance regardless of aggregate score. A weighted average can hide one catastrophic change among thousands of correct words.
Professional method. Define explicit critical-error gates for high-risk content and release decisions. The aim is to make the decision repeatable, because localization problems become expensive when the correct fix exists only in one reviewer’s memory.
Failure mode. A translation passes because its normalized score is high despite one safety-changing mistranslation. A reversed eligibility condition should not disappear statistically inside a large manual.
Verification. Check critical gates before calculating overall acceptance. If the result still depends on guesswork, return to the source context, locale requirement, product state or quality specification before approving the translation.
8. Choose penalty weights for decisions, not tradition
Severity points are a model of impact. Weights influence pass/fail outcomes and supplier behaviour.
Professional method. Set weights that reflect organizational risk and test them on historical cases before using them contractually. The aim is to make the decision repeatable, because localization problems become expensive when the correct fix exists only in one reviewer’s memory.
Failure mode. The team copies a 1/5/25 scheme without checking whether it matches actual consequences. A high penalty for critical errors may be appropriate but should be validated against real acceptance decisions.
Verification. Run old projects through the proposed weights and compare with expert judgment. If the result still depends on guesswork, return to the source context, locale requirement, product state or quality specification before approving the translation.
9. Normalize scores transparently
Raw error counts cannot compare samples of very different size. LQA often needs errors or penalty points relative to words, segments or another denominator.
Professional method. Choose a denominator meaningful to the content and publish the formula and limitations. The aim is to make the decision repeatable, because localization problems become expensive when the correct fix exists only in one reviewer’s memory.
Failure mode. One 500-word sample and one 5,000-word sample are compared by total points alone. Penalty points per thousand words can aid comparison while still keeping critical errors separate.
Verification. Make sure stakeholders can reproduce the score from the annotated sample. If the result still depends on guesswork, return to the source context, locale requirement, product state or quality specification before approving the translation.
10. Sample by risk as well as randomness
A random sample can miss the parts that matter most. ISO 5060 notes the role of sampling, but operational programs also need representative high-risk content.
Professional method. Combine representative sampling with targeted review of critical journeys, new terminology, machine-translated content or changed components. The aim is to make the decision repeatable, because localization problems become expensive when the correct fix exists only in one reviewer’s memory.
Failure mode. A random 1,000 words contains only easy descriptive text while payment instructions go unreviewed. A release sample can deliberately include checkout, legal and new-feature copy alongside randomly selected general text.
Verification. Map sample coverage to user journeys and content risk. If the result still depends on guesswork, return to the source context, locale requirement, product state or quality specification before approving the translation.
11. Calibrate reviewers before production scoring
Evaluator consistency is part of measurement quality. Two reviewers can assign different type and severity to the same issue.
Professional method. Use shared examples, independent annotation, reconciliation and periodic re-calibration. The aim is to make the decision repeatable, because localization problems become expensive when the correct fix exists only in one reviewer’s memory.
Failure mode. Supplier comparisons actually reflect reviewer personality. A calibration set can contain borderline preference/minor and minor/major cases.
Verification. Track agreement and investigate categories with persistent disagreement. If the result still depends on guesswork, return to the source context, locale requirement, product state or quality specification before approving the translation.
12. Require evidence in annotations
An LQA record should explain what is wrong and why. Vague comments cannot support root-cause analysis or fair appeals.
Professional method. Capture source, target, error span, type, severity, rationale and corrected version where appropriate. The aim is to make the decision repeatable, because localization problems become expensive when the correct fix exists only in one reviewer’s memory.
Failure mode. Tickets say ‘awkward’ or ‘bad translation’ with no requirement or impact. A terminology issue can cite the approved termbase entry.
Verification. Ask whether another reviewer could evaluate the same finding without speaking to the original annotator. If the result still depends on guesswork, return to the source context, locale requirement, product state or quality specification before approving the translation.
13. Separate output evaluation from process QA
Evaluating a translation sample is not the same as assuring the entire localization process. ISO 5060 focuses on translation-output evaluation and explicitly does not cover all quality-assurance processes or corrective actions.
Professional method. Use LQA findings as one evidence source alongside workflow, tooling, source quality and production QA. The aim is to make the decision repeatable, because localization problems become expensive when the correct fix exists only in one reviewer’s memory.
Failure mode. Management believes a score alone proves the whole localization system is healthy. A strong sample cannot reveal strings that were never extracted or screens nobody reviewed.
Verification. Maintain separate controls for coverage, structural QA and in-context testing. If the result still depends on guesswork, return to the source context, locale requirement, product state or quality specification before approving the translation.
14. Track root cause separately from error type
What went wrong in the output and why it happened are different fields. The same terminology defect can come from missing termbase entries, poor context, translator error or stale automation.
Professional method. After classification, assign or investigate root-cause codes at the process level. The aim is to make the decision repeatable, because localization problems become expensive when the correct fix exists only in one reviewer’s memory.
Failure mode. The dashboard reports 200 terminology errors but never changes the system that creates them. If most term errors originate from missing source-side concept approval, retraining translators alone will not solve them.
Verification. Each recurring defect family should map to a plausible prevention owner. If the result still depends on guesswork, return to the source context, locale requirement, product state or quality specification before approving the translation.
15. Trend patterns over time
LQA data becomes powerful when it shows change. One sample can be noisy, while repeated evaluation can reveal whether interventions work.
Professional method. Track comparable metrics, defect types, severity and reviewer conditions across releases. The aim is to make the decision repeatable, because localization problems become expensive when the correct fix exists only in one reviewer’s memory.
Failure mode. Every month uses a different taxonomy and sample strategy, making charts meaningless. A decline in placeholder or terminology defects after tooling changes can validate the intervention.
Verification. Confirm the comparison periods used equivalent definitions and denominators. If the result still depends on guesswork, return to the source context, locale requirement, product state or quality specification before approving the translation.
16. Compare suppliers carefully
Quality scores can support supplier management only when evaluation conditions are comparable. Different content difficulty, languages and reviewers can distort rankings.
Professional method. Control sample type, locale, reviewer calibration and risk mix; combine score data with delivery and rework context. The aim is to make the decision repeatable, because localization problems become expensive when the correct fix exists only in one reviewer’s memory.
Failure mode. A supplier handling legal content appears worse than one handling repetitive UI strings. A normalized score is still not a complete measure of service quality.
Verification. Document what variables were controlled before drawing conclusions. If the result still depends on guesswork, return to the source context, locale requirement, product state or quality specification before approving the translation.
17. Close the corrective-action loop
Quality measurement should change future output. Without action, LQA becomes expensive annotation.
Professional method. Connect repeated findings to source repair, termbase updates, style guidance, training, automation or regression tests. The aim is to make the decision repeatable, because localization problems become expensive when the correct fix exists only in one reviewer’s memory.
Failure mode. The same error categories remain high for months. Frequent mistranslation of one ambiguous source phrase can be prevented by rewriting the source and adding context.
Verification. For the top recurring defects, name the implemented prevention and later evidence. If the result still depends on guesswork, return to the source context, locale requirement, product state or quality specification before approving the translation.
A repeatable operating sequence
A mature LQA program moves from specification to evaluation to improvement while keeping measurement limits visible.
- Define audience, purpose, terminology, style and risk requirements.
- Choose a compact taxonomy and write examples.
- Define neutral, minor, major and critical severity rules.
- Calibrate evaluators on shared cases.
- Choose representative plus risk-targeted samples.
- Annotate errors with reproducible evidence.
- Apply transparent penalty and normalization rules.
- Check critical-error gates separately.
- Analyze type, severity and root-cause patterns.
- Assign corrective actions to process owners.
- Re-evaluate comparable samples after changes.
- Revise the model when categories or weights fail to produce useful decisions.
Treat this sequence as a loop. A late defect can reveal an earlier design assumption, missing context or weak specification. Repair the upstream cause where possible so the same class of problem becomes less likely in the next language, screen, release or evaluation sample.
Worked scenarios
1. Preference-heavy reviewer
One reviewer rewrites fluent target text extensively and produces far more errors than peers. The hidden risk is measurement capturing taste instead of specification violations.
Recalibrate against shared cases, use a neutral preference state and require a violated requirement for scored errors. The useful question is not merely whether the sentence sounds good in isolation, but whether the translated experience still performs the intended job under the real conditions in which a user, reviewer or system encounters it.
2. Critical mistranslation inside a strong score
A large sample contains one reversed safety instruction but very few other problems. The hidden risk is aggregate score hiding non-compensatory risk.
Apply the critical-error gate before the normalized score and block acceptance until the defect and its root cause are resolved. The useful question is not merely whether the sentence sounds good in isolation, but whether the translated experience still performs the intended job under the real conditions in which a user, reviewer or system encounters it.
3. Terminology defects spike after a product launch
Reviewers find many inconsistent feature names. The hidden risk is blaming translators for an upstream concept-governance gap.
Check whether the approved terminology existed before translation, update the termbase and measure the next release under the same taxonomy. The useful question is not merely whether the sentence sounds good in isolation, but whether the translated experience still performs the intended job under the real conditions in which a user, reviewer or system encounters it.
4. Two vendors evaluated on different content
One handles repetitive UI, the other new policy documents. The hidden risk is quality ranking confounded by difficulty and risk mix.
Use comparable samples or stratify by content type and avoid an overall winner claim unsupported by equivalent conditions. The useful question is not merely whether the sentence sounds good in isolation, but whether the translated experience still performs the intended job under the real conditions in which a user, reviewer or system encounters it.
5. Minor layout issue repeated 400 times
Each instance has low severity but appears in every screen. The hidden risk is individual severity hiding systemic cost.
Keep instance severity minor while flagging the root cause and frequency as a high-priority process defect. The useful question is not merely whether the sentence sounds good in isolation, but whether the translated experience still performs the intended job under the real conditions in which a user, reviewer or system encounters it.
6. LQA dashboard improves but user complaints rise
Scores look better while production localization problems increase. The hidden risk is output-sample measurement missing coverage and context defects.
Add in-context, coverage and production evidence; do not treat the LQA score as the complete quality system. The useful question is not merely whether the sentence sounds good in isolation, but whether the translated experience still performs the intended job under the real conditions in which a user, reviewer or system encounters it.
LQA taxonomy and severity: twenty professional practice cases
Use these cases to practise diagnosis rather than memorising slogans. For each case, identify the owning layer, the evidence required, the safest first action and the final release test.
1. The translation is correct in the CAT tool but wrong on screen
Inspect the rendered component, surrounding labels, dynamic values and available space. Decide whether the defect belongs to translation, layout, source context or product logic before changing words. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
2. A field rejects a perfectly legitimate user value
Separate business rules from culturally narrow validation assumptions. Ask whether the system needs the restriction or merely inherited it from one market. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
3. Two reviewers classify the same problem differently
Return to the error definition and severity criteria. If the framework cannot produce repeatable judgments, refine the rubric rather than arguing from preference. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
4. A webhook event arrives twice
Treat duplicate delivery as normal distributed-system behaviour. Use a stable event identifier or equivalent deduplication record before applying the same translation-state change again. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
5. A screenshot shows text clipped by one word
Check whether the string can be improved naturally, but also test whether the component is undersized for the target language. Do not force translators to compensate permanently for a layout bug. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
6. A user has only one name
Do not invent a family name field value. Store the name faithfully and change the form or downstream assumptions that demanded two Western-style parts. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
7. An error is frequent but low impact
Track frequency and severity separately. Many minor issues can indicate a systemic process problem without pretending each instance has the impact of a critical mistranslation. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
8. A synchronization call times out after the server may have processed it
Retry only under an idempotent or deduplicated design so uncertainty about the first response does not create duplicate projects, jobs or translations. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
9. A translated button is ambiguous only in one workflow state
Review the string in the exact state where the ambiguity appears. Context can change the action a short label appears to name. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
10. An address has no postal code
Allow for locales where postal codes are absent instead of generating fake data merely to satisfy a globally required field. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
11. A reviewer wants to mark every stylistic preference as an error
Use a neutral or preferential-change category where the quality model permits it, and reserve error penalties for violations of specifications or genuine quality requirements. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
12. An API consumer applies events out of order
Compare event version or current resource state rather than assuming arrival order is authoritative. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
13. A translated dialog looks fine until a user name becomes very long
Test real and extreme dynamic values in context. The correct localization unit includes the runtime content envelope, not just the static source string. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
14. A phone number parses but is not actually reachable
Distinguish structural possibility or numbering-plan validity from evidence that a number is assigned to a real user and currently reachable. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
15. One severe error is hidden inside a strong average quality score
Keep severity and critical-risk gates visible. Aggregate scores should not allow one safety- or obligation-changing error to disappear inside many correct segments. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
16. A webhook signature is valid but the event is stale
Authenticate the sender and separately check event relevance, version, object state and whether newer changes supersede it. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
17. The product uses the correct translation but the surrounding icon changes the meaning
Treat in-context review as a complete communication check. Text, icon, hierarchy and interaction state can jointly create the user’s interpretation. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
18. A form forces title choices such as Mr or Mrs
Ask whether the data is genuinely required. If not, make it optional or remove it instead of forcing users to reveal irrelevant personal information. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
19. An LQA program produces hundreds of defects but no fixes
Add root-cause and corrective-action loops. Measurement that never changes source content, tooling, training or review is reporting, not quality improvement. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
20. A localization webhook is missed during downtime
Use provider delivery history, replay/redelivery capabilities or reconciliation queries so the system can recover state rather than assuming every event arrives exactly once. Then state what would make you reverse the decision. This last step matters because a professional rule is stronger when its boundary conditions are visible.
Finally, test the decision outside the original example: another locale, another screen size, another user profile, another reviewer or another retry path. Robust localization survives changed conditions.
Release checklist
- Specifications exist before evaluation.
- Taxonomy categories are distinct enough for reliable use.
- Preferences are not automatically scored as errors.
- Severity is tied to user and business impact.
- Critical defects have non-compensatory gates.
- Penalty weights have been tested on real decisions.
- Score formulas and denominators are transparent.
- Sampling covers representative and high-risk content.
- Evaluators are calibrated and periodically rechecked.
- Annotations include evidence and rationale.
- Error type and root cause are tracked separately.
- Findings lead to corrective action and later verification.
Frequently asked questions
What is LQA?
Localization quality assurance commonly refers to structured review of localized output, often using error types, severity and scoring to evaluate whether the translation meets specifications. A professional workflow should record the rule, the evidence behind it and the condition that would make the team revisit the decision.
What is MQM?
Multidimensional Quality Metrics is a structured framework for analytic translation-quality evaluation with a defined error typology and related terminology. A professional workflow should record the rule, the evidence behind it and the condition that would make the team revisit the decision.
What does ISO 5060:2024 cover?
It provides general guidance on human evaluation of translation output, including analytic error-based approaches, evaluator competence and sampling. A professional workflow should record the rule, the evidence behind it and the condition that would make the team revisit the decision.
Should every edit count as an error?
No. A reviewer can suggest a preferable wording even when the existing target meets specifications; quality systems should distinguish preferences from real violations. A professional workflow should record the rule, the evidence behind it and the condition that would make the team revisit the decision.
How should severity be assigned?
By likely impact on meaning, user task, trust, safety, legal obligation or other project requirements rather than by how noticeable the issue is. A professional workflow should record the rule, the evidence behind it and the condition that would make the team revisit the decision.
Can an overall score hide a critical error?
Yes, which is why high-risk programs often use critical-error gates that cannot be compensated for by many correct words. A professional workflow should record the rule, the evidence behind it and the condition that would make the team revisit the decision.
How large should an LQA sample be?
There is no universal number. Sample design depends on risk, content volume, variability, decision purpose and the confidence needed. A professional workflow should record the rule, the evidence behind it and the condition that would make the team revisit the decision.
How do we improve reviewer agreement?
Use explicit definitions, shared examples, independent calibration exercises and reconciliation of borderline cases. A professional workflow should record the rule, the evidence behind it and the condition that would make the team revisit the decision.
Selected references and next routes
- ISO 5060:2024 — Translation services — Evaluation of translation output
- Multidimensional Quality Metrics (MQM)
- eduKateSG: Measure Translation Quality Without Letting Metrics Distort the Work
Conclusion
An LQA framework is valuable when it turns subjective review into evidence the organization can act on. That requires more than a score: clear specifications, stable error definitions, meaningful severity, calibrated evaluators and transparent sampling.
The deeper goal is prevention. When quality data reveals where errors come from and the team changes source content, context, terminology, training or tooling accordingly, LQA becomes part of a learning system rather than a monthly grading ritual.