VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How People Translate Quickly | Project Analysis and Match Bands: Estimate Real Translation Effort Before You Start

People searching CAT tool project analysis, translation memory analysis report, fuzzy match bands, weighted word count translation, translation repetition count, or how to estimate translation effort from TM matches are trying to answer a question that raw word count cannot answer: how much human work is actually inside this project? A 20,000-word file with mostly context matches and repetitions is not the same workload as a 20,000-word file of entirely new technical prose.

A fast translation workflow uses project analysis and match bands before serious drafting begins. Modern CAT tools commonly report categories such as context matches, 100% exact matches, high fuzzy matches, lower fuzzy matches, repetitions, and new words. Some systems also calculate a weighted word count by assigning different effort weights to different match categories. The percentages are not universal truth, but they make the hidden structure of the workload visible.

This article owns one reader job: use a CAT analysis report to estimate attention, sequence work, and detect project risk before translating sentence one. It is not a pricing article, not a general translation-speed benchmark, and not a tutorial on fuzzy-match editing. The focus is operational planning: convert raw file size into an evidence-based map of how much work is new, reusable, repetitive, or uncertain.

Quick answer

Run project analysis before translating.

Inspect at least:

  • new/no-match words;
  • context matches;
  • 100% matches;
  • high fuzzy bands;
  • lower fuzzy bands;
  • repetitions;
  • locked or excluded content;
  • weighted word count if available.

Then ask:

Where is the real human effort?

Use the answer to:

  • schedule difficult work;
  • decide whether resources are attached correctly;
  • spot suspiciously low leverage;
  • estimate review burden;
  • sequence files;
  • identify whether TM cleanup or terminology setup is worth doing before drafting.

Analysis is not a guarantee.

It is a workload model.

Raw word count is not effort

Suppose Project A has 10,000 source words.

Project B also has 10,000.

Project A contains:

  • 6,000 context/exact match words;
  • 2,000 repetitions;
  • 1,000 high fuzzy;
  • 1,000 new.

Project B contains:

  • 500 exact;
  • 500 high fuzzy;
  • 9,000 new.

Same word count.

Completely different translation workload.

This is why experienced localization workflows separate volume from effort.

The CAT analysis report is a bridge between the two.

What match bands represent

Translation-memory analysis groups source segments by how similar they are to stored material.

Common categories include:

  • context match / 101% or equivalent;
  • exact 100%;
  • 95–99%;
  • 85–94%;
  • 75–84%;
  • lower fuzzy bands;
  • no match;
  • repetition.

The exact ranges differ by platform.

The principle is stable.

Higher similarity often means more reusable target material.

But similarity is not the same as effort.

A one-word change can be crucial.

Context matches: low expected effort, not zero thought

A context match usually means the same source text appears in the same or highly similar context.

This can be a strong reuse signal.

In a clean approved TM, context matches may require only light verification.

But check:

  • terminology currency;
  • source metadata;
  • locale;
  • current style;
  • high-risk numbers or conditions.

Weighted models often assign low effort to these matches.

That makes sense statistically.

It does not make them infallible.

100% matches: exact text, context still matters

A 100% match tells you the source text is identical to stored source text.

That can save substantial work.

But short strings and context-sensitive sentences may need inspection.

Analysis therefore gives you a planning assumption:

likely low effort.

It does not give you a semantic guarantee.

High fuzzy matches: often the sweet spot

A 95–99% match frequently differs by:

  • one number;
  • one name;
  • one product code;
  • punctuation;
  • a short modifier;
  • negation.

These segments can be extremely fast to edit.

They can also be extremely dangerous if the changed element controls meaning.

From a planning perspective, high fuzzy matches usually require less time than new translation but more deliberate comparison than exact matches.

This is why separate bands are useful.

Mid fuzzy matches: reusable structure, more rewriting

An 80–94% match may preserve:

  • terminology;
  • sentence skeleton;
  • recurring phraseology.

But clause structure or argument relationships may have changed.

The translator can still benefit from the match.

Editing effort becomes more variable.

Some 85% matches are nearly ready.

Some are slower than translating from zero.

A project analysis report averages this variability.

Low fuzzy matches: evidence, not necessarily leverage

Low matches can contain useful terms or phrases.

But if the target requires major restructuring, pre-inserting them can create editing debt.

When the analysis shows a large lower-fuzzy bucket, ask:

  • Are the attached TMs relevant?
  • Is this domain mismatch?
  • Are source templates similar only superficially?
  • Would a higher working threshold reduce noise?

The analysis can diagnose resource quality, not only volume.

Repetitions: high leverage if context is stable

A repeated source segment can be translated once and reused.

This can dramatically reduce work in:

  • manuals;
  • forms;
  • software;
  • catalogs;
  • policies.

But repetition count alone can be misleading.

A short string such as “Open” may repeat in multiple functions.

Context matters.

Treat repetition as a leverage category, then confirm whether auto-propagation is safe.

Internal repetitions versus TM matches

A repetition exists inside the current project.

A TM match comes from previous bilingual history.

Both reduce new formulation.

They differ in provenance.

A repetition may become reusable only after the first current occurrence is translated.

This matters for scheduling.

Translating the first occurrence early can unlock later repetitions.

Worked example 1: two 25,000-word manuals

Manual A:

  • 50% context/exact;
  • 20% repetitions;
  • 15% high fuzzy;
  • 15% new.

Manual B:

  • 10% exact;
  • 5% repetitions;
  • 10% high fuzzy;
  • 75% new.

A deadline based only on 25,000 raw words treats them equally.

An analysis-based plan does not.

Manual B needs substantially more fresh linguistic work.

Manual A may need more consistency verification because of reuse.

Different profile.

Different plan.

Weighted word counts

A weighted word count converts match categories into estimated effort.

Conceptually:

raw words × category weight = weighted effort

Example:

  • new words weight 100%;
  • high fuzzy maybe 40–60%;
  • exact matches maybe 20–30%;
  • repetitions lower;
  • context matches lower still.

Vendors use different scales.

Do not treat one set of weights as scientifically universal.

The value is comparative.

Weighted counts help you model workload more realistically than raw counts.

Why default weights may be wrong for you

A translator working in legal content may spend significant time reviewing exact matches.

A translator localizing highly repetitive UI may review context matches extremely quickly.

A language pair with heavy morphology may require more editing in fuzzy matches.

A project with unreliable TM may turn exact matches into expensive verification.

Therefore:

weighted count = model, not measurement.

Calibrate weights with your own observed effort when planning matters.

Use analysis before pricing

Even if this article is not about pricing, pricing exposes why analysis matters.

If a quote is based on raw word count only, the client may overpay for heavily reused material or underfund a mostly-new technical project.

A match-band report provides a more transparent basis for estimating labor.

The same logic helps internal planning.

You need to know whether the workload is new creation, reuse review, or mixed editing.

Use analysis before assigning translators

Suppose one file contains mostly repetitions and exact matches.

Another contains dense new regulatory text.

Do not assign both based only on raw word count.

A balanced team assignment can use:

  • weighted words;
  • risk;
  • domain expertise;
  • deadline;
  • review burden.

This is especially valuable in multi-file programs.

Use analysis to order files

Sometimes one file should be translated first because it contains the first occurrence of many repetitions.

Or because it establishes terminology used in later files.

Analysis can reveal which files contain:

  • most new text;
  • most repeated text;
  • highest TM leverage;
  • unusual low-match clusters.

Sequence strategically.

The first file can create leverage for the next.

Use analysis to detect a missing TM

You expect 70% reuse.

The analysis shows 5%.

Something is wrong.

Possibilities:

  • wrong TM not attached;
  • target locale mismatch;
  • TM disabled;
  • metadata filtering too strict;
  • source was heavily revised;
  • segmentation changed.

This is one of the fastest diagnostic uses of analysis.

Do not begin translating 50,000 words before asking why expected leverage disappeared.

Use analysis to detect a dirty TM

You expect good reuse.

The analysis shows thousands of low fuzzy matches and very few high matches.

That can indicate:

  • old style;
  • mixed domain;
  • poor segmentation;
  • noisy aligned data.

Again, the report is a diagnostic signal.

It tells you that the resource environment may deserve attention before drafting.

Worked example 2: missing context matches

A software update should largely repeat last month.

Analysis shows many 100% matches but almost no context matches.

Possible causes:

  • string IDs changed;
  • file order changed;
  • context metadata was not imported;
  • source pipeline changed.

This matters because ordinary exact matches may require more review than stable in-context matches.

The report reveals a context-loss problem.

Worked example 3: sudden repetition spike

A file has 60% repetitions.

Great?

Maybe.

Inspect the source.

The repetitions might be:

  • genuine repeated UI labels;
  • blank-like boilerplate;
  • malformed OCR;
  • duplicated source blocks;
  • accidentally imported hidden content.

Analysis can expose source-quality problems.

High leverage is not automatically good news.

The analysis sanity check

After generating the report, ask:

  • Does this profile make sense for the project?
  • Is the new-word share believable?
  • Are exact matches coming from the expected TM?
  • Are repetitions plausible?
  • Are locked segments excluded?
  • Are numbers and tags counted as expected?
  • Is the source word count roughly consistent with the original file?

A wildly surprising report deserves investigation.

Match percentage is not correctness percentage

A 99% fuzzy match is not “99% correct.”

A 100% match is not “100% safe.”

A match score describes source similarity under the tool’s algorithm.

It does not directly measure:

  • target quality;
  • terminology currency;
  • context;
  • legal validity;
  • factual correctness.

Use match bands to estimate work.

Use review to determine correctness.

Weighted counts and high-risk text

High-risk content can break the effort model.

A short exact clause may still require legal review.

A 99% match may differ by “not.”

A single changed dosage can require careful verification.

Add risk on top of match analysis.

A good workload model has at least two axes:

  • expected editing effort;
  • consequence of error.

Create a risk-adjusted view

You do not need a complex formula.

Mark files or sections:

  • low risk;
  • normal;
  • high risk.

Then read the match-band profile through that lens.

High exact-match coverage in a safety manual does not justify eliminating human review.

It may justify moving review from retranslation to verification.

That is still a major speed gain.

Use analysis with pre-translation

Analysis tells you what resources can contribute.

Pre-translation applies those resources.

A good sequence is:

  1. run analysis;
  2. inspect match distribution;
  3. set pre-translation thresholds;
  4. pre-translate;
  5. review according to match type.

Do not configure automation blindly.

The report should inform it.

Use analysis with TM penalties

A penalized TM may lower effective match scores.

This changes match-band distribution.

If you add a penalty to legacy memory, some 100% source matches may now fall into a lower effective category.

The report then better reflects the extra review effort.

This is useful.

The analysis should mirror trust, not only text overlap.

Use analysis with project templates

If you handle the same client repeatedly, project templates can include:

  • TM set;
  • QA profile;
  • pre-translation rules;
  • analysis settings.

Now each new project produces comparable reports.

Trend data becomes more meaningful.

You can see whether leverage is rising or falling across releases.

Track analysis over time

For recurring work, record:

  • total raw words;
  • weighted words;
  • new-word share;
  • exact/context share;
  • repetition share.

Patterns can reveal:

  • better TM accumulation;
  • source reuse decline;
  • terminology churn;
  • structural changes in the product.

Translation memory is a learning asset.

Analysis can show whether it is actually creating leverage.

Failure mode 1: treating weighted count as objective truth

The model says 6,000 weighted words.

The translator finishes in half the expected time.

Or double.

That does not mean the tool failed.

The weights were assumptions.

Calibrate them.

Do not confuse a planning model with a stopwatch.

Failure mode 2: ignoring review effort

A project has enormous exact-match coverage.

The quote or schedule assumes almost no work.

But client policy requires full bilingual review of all reused segments.

The analysis was read too narrowly.

Match leverage reduces drafting effort.

Review policy can reintroduce effort.

Failure mode 3: counting repetitions as universally cheap

Repeated sentence:

Apply

appears in multiple contexts.

One translation is propagated everywhere.

Several are wrong.

The repetition discount was applied before context was understood.

Leverage requires functional sameness.

Failure mode 4: trusting raw analysis after resource changes

You add a new TM.

Do not assume the old report still represents the project.

Rerun analysis.

The workload model changed.

This is particularly important before assignment or deadline planning.

Failure mode 5: analyzing after translation begins

By then, new confirmed segments may have entered the working TM.

The report can show more leverage than existed at the start.

For baseline planning, analyze before substantial translation.

If you run later, label the report accordingly.

Failure mode 6: using analysis as a productivity scorecard for individuals

A translator receives a project with 3,000 weighted words.

Another receives 3,000 weighted words.

Their tasks may still differ in risk, domain, language direction, source quality, and review policy.

Do not use weighted counts as a simplistic human performance ranking.

They estimate text leverage.

They do not measure professional difficulty completely.

A five-minute analysis routine

Before starting:

  1. run project statistics;
  2. note raw and weighted words;
  3. scan match bands;
  4. inspect repetition share;
  5. verify the expected TM is active;
  6. inspect one sample from each major band;
  7. identify the highest-risk new-text area;
  8. decide work order.

Five minutes can prevent hours of poorly sequenced work.

Sample the bands

Open:

  • one context match;
  • one 100%;
  • one 95–99%;
  • one middle fuzzy;
  • one repetition;
  • one new segment.

Ask:

Do these categories behave the way I expect?

If not, adjust your plan.

This connects statistics with actual text.

Build your own effort ratios

After several projects, compare actual time by band.

Perhaps you discover:

  • context match review = 10% of new translation time;
  • exact review = 20%;
  • high fuzzy = 45%;
  • mid fuzzy = 75%;
  • new = 100%.

These ratios can improve internal planning.

Do not copy somebody else’s weights without testing.

Workload by file, not just project total

A project-level average can hide one difficult file.

Inspect file-level analysis when:

  • work is assigned across people;
  • one file is safety-critical;
  • deadlines are staggered;
  • source authors differ.

You may find one file is 90% reused and another 90% new.

Average coverage is operationally misleading.

Segment count also matters

Word count can hide UI-heavy projects with thousands of tiny segments.

Each segment may require:

  • context check;
  • confirmation;
  • tag handling;
  • screen lookup.

A 5,000-word UI project with 3,000 segments can feel different from a 5,000-word article with 100 paragraphs.

Use segment count alongside words.

Character count matters in some languages

Languages without whitespace word boundaries may be analyzed using characters rather than words.

Do not compare raw metrics across language systems blindly.

The metric should match the target workflow.

Analysis for subtitles

Subtitle projects may need additional workload factors:

  • timing;
  • character limits;
  • reading speed;
  • audiovisual context.

TM match bands remain useful but do not capture full effort.

Again, analysis is one layer.

Analysis for spreadsheets

Spreadsheet projects may have many repeated labels and IDs.

Analysis can show high repetition.

But structural validation is extra work.

Do not let a low weighted count erase file-integrity checks.

Analysis for legal work

Legal templates often create strong TM leverage.

Review policy may require exact clause verification anyway.

Use analysis to distinguish new drafting from clause validation.

Do not treat legal approval as a generic match discount.

Analysis for technical manuals

This is a strong environment for match-band planning because:

  • terminology repeats;
  • procedures recur;
  • previous versions exist;
  • small updates create high fuzzy matches.

Project analysis can often predict effort reasonably well if the TM is clean.

A practical project-analysis worksheet

Record:

Raw words: total visible source volume. Segments: interaction count. Context/exact: likely light review. High fuzzy: focused change review. Mid fuzzy: moderate editing. Low/no match: fresh translation. Repetitions: current-project reuse. Weighted words: modeled effort. Risk notes: material whose review requirement exceeds its match band. Resource notes: missing or weak TM, terminology gaps, segmentation anomalies.

This creates a concise planning snapshot.

Transfer: study planning

Students can use the same logic.

A textbook chapter contains:

  • known concepts;
  • partially familiar concepts;
  • repeated exercises;
  • entirely new material.

Raw page count is not study effort.

Classifying familiarity produces a better plan.

Transfer: software estimation

Developers estimate tasks by reuse and novelty.

A feature built from known components differs from one requiring new architecture.

Translation analysis makes the same distinction explicit.

The deeper principle: plan from expected human decisions, not raw file size

A project is not heavy because it contains many words.

It is heavy because it contains many decisions that humans still need to make.

Context matches reduce decisions.

Repetitions reduce decisions.

Fuzzy matches partially reduce decisions.

New text creates new decisions.

A good CAT analysis report approximates this structure.

Use it as a map.

Then add risk, domain knowledge, and review policy.

That is a far better basis for speed than staring at raw word count.

Advanced practice: build an effort curve instead of one weighted number

A single weighted total is convenient.

It can also hide important differences.

Two projects may both show 8,000 weighted words.

Project A:

  • mostly 95–99%;
  • small new bucket.

Project B:

  • large new bucket;
  • many context matches.

The editing rhythm differs.

Build a quick effort curve:

  • low-touch reuse;
  • focused edits;
  • heavy edits;
  • new translation.

Now you know not only total expected effort but how the effort is distributed.

This helps with scheduling.

A day of mostly focused fuzzy edits feels different from a day of entirely new technical translation.

Match-band volatility

Some bands are predictable.

Context matches from a clean TM may be consistently easy.

Mid fuzzy matches can be volatile.

One 85% segment needs two word changes.

Another requires full restructuring.

When planning, assign more uncertainty to volatile bands.

This prevents false precision.

Confidence intervals for workload

You do not need statistical software.

Use a range.

Example:

Estimated weighted effort: 6,000–7,500 new-word equivalents.

The range acknowledges variation in:

  • source quality;
  • language pair;
  • review depth;
  • TM trust.

A range is often more honest than “6,482 weighted words.”

Worked example 4: same weighted count, different risk

Project A is internal product documentation.

Project B is a patient-facing medical leaflet.

Both show 5,000 weighted words.

Project B needs:

  • more careful exact-match review;
  • stronger terminology verification;
  • higher consequence checks.

Match analysis estimates reuse.

Risk modifies the human process.

Do not collapse both into one number.

Worked example 5: TM penalty reshapes analysis

A legacy TM contains many exact source matches.

Before penalty:

  • 40% exact.

After a 10% resource penalty:

  • many matches fall into a lower effective band.

This is useful.

The report now reflects that humans must spend more attention because the historical targets are less trusted.

The source did not change.

The workload model did.

Worked example 6: project template hides stale resources

A template automatically attaches last year’s TM.

Analysis shows high leverage.

But the client has changed terminology.

The high exact-match count looks attractive and is operationally dangerous.

Always interpret analysis with resource currency.

A strong match from stale knowledge is not cheap work.

Analysis and source novelty

High new-word share can mean:

  • genuinely new content;
  • source rewrite;
  • new product area;
  • failed TM attachment;
  • segmentation change.

Do not jump directly to “more translation.”

Ask why novelty increased.

The reason can change your next action.

Analysis and content churn

For recurring releases, compare the new-word share over time.

If it rises from 10% to 70%, something changed in the content pipeline.

Possible explanations:

  • templates were rewritten;
  • source authors changed;
  • sentence standardization declined;
  • string IDs changed;
  • source generation changed.

Translation analysis can become a content-engineering signal.

Repetition leverage across files

Some tools analyze repetitions across the whole project.

Others calculate within files or according to settings.

Know the analysis scope.

If the same sentence appears in ten files, cross-file repetition can create major leverage.

If analysis is file-local, that leverage may be hidden.

Scope affects the report.

First-occurrence cost

Repetitions are cheap only after the first instance is solved.

For a project with 1,000 repetitions of 20 source phrases, the real linguistic work may concentrate in 20 first occurrences.

Translate those first.

Then propagation or reuse can handle the rest.

Analysis can inform sequence.

Exact-match review policy

Define policy before estimating.

Possible policies:

Light

Context matches sampled, exact matches glanced.

Standard

Every reused segment read against source.

High-risk

Every reused segment fully verified with terminology and data checks.

The same analysis report maps to different human effort depending on policy.

Machine translation coverage versus TM analysis

Some platforms include MT or quality-estimation categories separately.

Do not treat MT-filled new segments as TM leverage.

A new source segment remains new source text even if a machine generated a draft.

The human task becomes post-editing rather than blank-page translation.

Analysis vocabulary should reflect this distinction.

Analysis and quality estimation

TM match percentage and MT quality estimation are different signals.

TM match:

similarity to previous source text.

QE:

predicted quality of generated target.

A project can contain both.

Use them separately.

Do not blend percentages mentally.

Segment length affects edit cost

A 95% match on a four-word segment may mean one word changed.

That is 25% of the sentence.

A 95% match on a 40-word segment may contain a smaller proportional change.

Band labels summarize similarity.

Segment length still matters.

Fuzzy match location matters

Changed text at the beginning of a segment can affect:

  • subject;
  • tense;
  • framing.

Changed text at the end may affect only one modifier.

The same numerical match can require different cognitive effort.

This is why fuzzy-match diffing remains necessary.

Build a calibration sample

Before committing to the weighted estimate:

  1. sample 10 segments per major band;
  2. time or roughly rate editing effort;
  3. compare to your assumed weights.

This takes little time on a large project.

It can improve the whole estimate.

Calibrate by language pair

Language pairs differ in:

  • word order;
  • morphology;
  • article systems;
  • grammatical gender;
  • script.

A fuzzy match that is easy in one pair may require more structural editing in another.

Maintain separate effort assumptions for major pairs if planning precision matters.

Calibrate by domain

Legal fuzzy matches may be slower than marketing fuzzy matches at the same score because wording precision is stricter.

Technical exact matches may be faster than literary exact matches because controlled terminology is stable.

Domain changes the conversion from match band to labor.

Calibrate by resource source

Matches from:

  • client-approved TM;
  • aligned legacy corpus;
  • vendor TM;
  • machine-origin TM;

should not share identical effort assumptions if trust differs.

Analysis becomes more realistic when match score and resource provenance are both considered.

Analysis before and after cleanup

If you clean a TM or fix segmentation rules, rerun analysis.

Compare:

  • exact/context share;
  • fuzzy distribution;
  • weighted total.

This quantifies the value of the cleanup.

Sometimes a one-hour resource fix reduces thousands of weighted words.

That is a powerful optimization.

Analysis before and after term extraction

Term extraction may not change match bands directly.

It can still reduce the effort inside new words.

A project with many new words but a prepared glossary may be faster than the same raw analysis suggests.

Again, analysis is one layer.

Analysis and project deadlines

Turn the effort model into a schedule:

  • new-heavy file early;
  • repetition-heavy file after first occurrences are solved;
  • high-risk file when reviewers are available;
  • low-touch exact-match pass in lower-energy periods.

This is workload shaping.

Analysis and parallel work

If several translators work simultaneously, repetitions can lose leverage if the same first occurrence is translated differently in multiple files.

Coordinate:

  • shared TM;
  • first-occurrence ownership;
  • terminology.

The project analysis can reveal where parallelization helps and where it creates duplicated decisions.

Analysis and reviewer workload

Reviewers may receive:

  • new text;
  • fuzzy edits;
  • reused exact text.

Decide whether reviewer depth changes by band.

A review plan can be more efficient than one universal pass.

For example:

  • full review of new/high fuzzy;
  • targeted terminology/data review of exact matches;
  • sample context matches.

Use risk policy.

A reviewer-facing analysis report

Share a simplified view:

  • new;
  • fuzzy;
  • reused;
  • repetitions;
  • high-risk exceptions.

The reviewer does not need every CAT statistic.

They need to know where attention is likely to pay off.

Analysis and external review packages

If a bilingual review package will be sent externally, use analysis to decide scope.

Do you need to send:

  • all segments;
  • only new/fuzzy;
  • only high-risk exact matches;
  • only reviewer comments?

A smaller package can produce better review attention.

Analysis and budgeting attention

Attention is finite.

The report helps budget it.

Spend strongest attention on:

  • new content;
  • low-confidence fuzzy matches;
  • high-risk changed details.

Spend lighter but nonzero attention on:

  • stable context matches;
  • verified repetitions.

This is the human version of weighted counts.

Beware false precision from decimals

A report may calculate exact percentages to two decimal places.

Human effort is not that precise.

Do not debate whether a project is 6,421 versus 6,503 weighted words.

Use the numbers to improve decisions, not create numerical theater.

Use analysis as an early-warning system

Unexpected patterns can warn you about:

  • missing resources;
  • source duplication;
  • segmentation changes;
  • TM contamination;
  • content churn.

This diagnostic value may be more important than the effort estimate itself.

A monthly leverage dashboard for recurring programs

For continuous localization, track:

  • raw words;
  • weighted words;
  • new share;
  • context/exact share;
  • repetition share;
  • average turnaround;
  • reviewer corrections by band.

Over time, you can see whether the translation system is learning.

If TM leverage rises while review errors stay low, the infrastructure is compounding.

If leverage rises but errors rise too, reuse quality may be deteriorating.

When not to rely heavily on analysis

Be cautious when:

  • source is creative;
  • TM is mixed quality;
  • segmentation is unstable;
  • review rules require full re-reading;
  • file structure adds large nonlinguistic effort;
  • MT post-editing dominates.

The report still helps.

It simply explains less of total effort.

A mature interpretation

A beginner sees:

12,000 words.

A more experienced translator sees:

12,000 raw words, 4,500 weighted, 65% high-trust reuse, 800 new technical terms, two high-risk sections, one file with broken segmentation.

That second description is closer to the actual job.

Project analysis trains this richer view.

Final analysis check: compare estimate with reality

After the project is complete, compare the original analysis with what actually happened.

Record roughly:

  • actual translation time;
  • review time;
  • which match band consumed more effort than expected;
  • which band was easier than expected;
  • whether repetitions behaved safely;
  • whether exact matches needed terminology repair;
  • whether file structure added unmodeled work.

Then adjust your future planning assumptions.

This turns CAT analysis into a learning system rather than a static report. If high fuzzy matches consistently take 70% of new-translation effort for your language pair, a 40% weight is too optimistic for your workflow. If context matches are nearly frictionless, your estimate can reflect that.

The most useful project-analysis model is therefore not the vendor default. It is the model that gradually learns your real combination of language pair, domain, resource quality, and review policy.

Use the first completed file to recalibrate the rest

In a multi-file project, the first completed file gives you better evidence than the original analysis alone. Compare its predicted match profile with the real editing experience before committing the same schedule to every remaining file.

If high fuzzy matches took much longer than expected because terminology changed, raise the effort assumption for the rest. If repetitions were genuinely safe and fast, preserve that leverage. If exact matches needed heavy review because the TM was stale, adjust the remaining workload immediately.

This mid-project recalibration is especially valuable when deadlines are tight. Planning should improve as evidence improves. The analysis report starts the estimate; completed work refines it.

Summary

Project analysis and match bands help people translate quickly by revealing how much of the source is new, repeated, or reusable from translation memory before work begins.

The reliable workflow is:

run analysis → inspect context/exact/fuzzy/repetition bands → sanity-check resources → review weighted counts as estimates → add risk and review policy → sequence work accordingly

Do not interpret match percentages as correctness probabilities.

Do not treat weighted counts as universal truth.

Do use analysis to detect missing resources, surprising repetition, source problems, and workload imbalance.

The point is simple:

estimate human effort from the structure of the decisions, not from raw word count alone.

Frequently asked questions

What is a CAT project analysis?

It is a report that counts source volume by translation-memory match category, repetitions, and other project states so teams can estimate workload.

What are fuzzy match bands?

They are ranges such as 95–99% or 85–94% that group segments by similarity to previous TM source text.

What is a weighted word count?

It is an effort model that multiplies raw words in each match band by an assumed workload percentage.

Is weighted word count the same as actual work time?

No. It is a planning estimate. Actual effort depends on language pair, risk, TM quality, source quality, and review requirements.

Why analyze before translating?

Confirmed work can enter the working TM and change later match results. Pre-translation analysis provides a cleaner baseline.

What do repetitions mean?

They are repeated source segments within the current project or analysis scope. They may be reusable after the first occurrence is translated.

Are 100% matches free work?

No. They often require less effort, but context, terminology, and review policy can still require attention.

Why can a 99% match be dangerous?

The changed 1% might be a number, negation, deadline, name, or condition that controls the sentence meaning.

How can analysis reveal a project setup problem?

Unexpectedly low reuse can indicate a missing TM, wrong locale, changed segmentation, context loss, or resource misconfiguration.

Should I use default match weights?

They are a starting model. Calibrate them with observed effort for recurring workflows.

Internal-link opportunities

This article can naturally connect to:

  • How People Translate Quickly | Pre-Translation — for applying resources after analysis.
  • How People Translate Quickly | Fuzzy Match Diffing — for editing high fuzzy matches.
  • How People Translate Quickly | Context Matches — for interpreting high-trust reuse.
  • How People Translate Quickly | Repetition Auto-Propagation — for current-project repetition leverage.
  • How People Translate Quickly | TM Penalties and Match Thresholds — for resource trust that affects analysis categories.
  • How People Translate Quickly | Translation Speed Benchmarks — for human throughput rather than CAT leverage analysis.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading