People searching translation project analysis, CAT tool analysis, weighted word count translation, TM analysis report, fuzzy match bands, repetitions in translation pricing, or translation memory leverage report are trying to answer a question that should be solved before full-speed translation begins: how much of this file is genuinely new work?
A raw word count cannot answer that by itself. A 10,000-word project with thousands of approved context matches, exact translation-memory matches, and internal repetitions is not the same workload as a 10,000-word project containing entirely new text. Current CAT and TMS systems therefore analyze files against translation memories, repetition patterns, and other project resources, then break the volume into match bands. Some workflows also convert those bands into a weighted word count or effective volume for scheduling, quoting, and assignment.
This article has one dominant job: read a translation-project analysis before you commit to the workload, so you can distinguish raw volume from actual translation effort. It does not set a universal price grid, tell agencies what to pay translators, or replace a detailed risk assessment. It explains how people translate quickly by knowing where the leverage is before they begin.
Quick answer
Before promising a deadline or dividing a project, analyze the source against the resources that will actually be used.
Look for:
- total words or characters;
- context or in-context matches;
- 100% exact matches;
- high fuzzy matches;
- medium and low fuzzy matches;
- no-match or new words;
- internal repetitions;
- cross-file repetitions;
- machine-translation categories if the platform reports them;
- locked or non-translatable content where relevant.
Then ask:
- Are the attached translation memories current and trustworthy?
- Are repetitions actually reusable?
- Are exact matches contextually safe?
- Are fuzzy matches concentrated in easy or risky material?
- Does the analysis include technical complexity such as tags, tables, or file engineering?
- Does the weighted model reflect the real review burden?
- What part of the project still needs fresh human decision-making?
The central rule is:
raw words measure size; analysis measures leverage; neither alone measures risk.
Why project analysis makes translation faster
Translation speed is strongly affected by what the translator must decide from scratch.
Consider two files.
File A
10,000 words.
- 4,000 context or exact matches;
- 2,000 repetitions;
- 2,000 high fuzzy matches;
- 2,000 new words.
File B
10,000 words.
- 200 exact matches;
- 100 repetitions;
- 400 fuzzy matches;
- 9,300 new words.
The raw volume is identical.
The likely work pattern is not.
File A may contain substantial reusable material.
File B is close to a new translation.
If you promise the same turnaround merely because both files say “10,000 words,” you ignore the structure of the work.
What a CAT analysis actually does
A translation-project analysis usually compares source segments against one or more translation memories and scans the project for internal repetition.
The tool then groups words or segments according to similarity or match status.
Common categories include:
- context match;
- 100% match;
- 95–99% fuzzy match;
- 85–94% fuzzy match;
- 75–84% fuzzy match;
- lower fuzzy bands;
- no match;
- repetition.
Names and thresholds vary between systems.
The important concept is stable:
The analysis estimates how much prior reusable evidence exists for the current source.
It is a map of expected leverage.
Analysis is only as good as the resources attached
Suppose you analyze a file against a huge general translation memory.
The report shows many matches.
But the memory contains:
- another client’s language;
- outdated terminology;
- wrong locale;
- old product names;
- inconsistent quality.
The analysis looks efficient.
The project may not be.
This is why project analysis should use the same resources that will govern the real translation job.
A misleading resource produces misleading leverage.
Step 1: verify the resource set before trusting the numbers
Before running or interpreting an analysis, check:
- client;
- product;
- domain;
- target locale;
- date;
- memory priority;
- whether legacy resources are read-only or active;
- whether current terminology changed.
Do not calculate workload against a memory that the translator should not actually trust.
The number of matches is less important than the quality of matches.
Step 2: distinguish match percentage from effort
A 95% match can be easy.
It can also be dangerous.
Old source:
The device may be used below 40°C.
New source:
The device may not be used below 40°C.
The segment is highly similar.
The meaning is opposite.
A weighted model might discount this heavily because only one token changed.
Human review cannot.
Match percentage measures textual similarity.
It does not measure semantic consequence.
The risk token problem
Certain changes deserve disproportionate attention:
- not;
- except;
- must;
- may;
- before;
- after;
- above;
- below;
- less;
- more;
- dates;
- money;
- dosage;
- units;
- thresholds.
A high fuzzy match containing one changed risk token may deserve more effort than a completely new but simple sentence.
So the analysis is a workload signal, not an automatic priority order.
Step 3: understand context matches
Some tools provide a category stronger than ordinary 100% identity.
Names differ:
- context match;
- 101% match;
- in-context exact match;
- ICE match;
- double-context match.
The basic idea is that the current source segment is identical and additional context—such as surrounding segments or identifiers—also agrees.
These matches can be strong reuse candidates when:
- memory provenance is good;
- product context is stable;
- source structure is trustworthy.
But even strong matches can become unsafe after:
- product rename;
- legal update;
- locale change;
- terminology migration;
- changed interface function.
The analysis tells you confidence is higher.
It does not make verification unnecessary in every project.
Step 4: treat 100% matches as exact text, not universal truth
An ordinary 100% match says the source segment is textually identical to a stored source segment under the tool’s matching rules.
That is useful.
But ask:
- same client?
- same product?
- same audience?
- same date?
- same screen?
- same defined term?
- same jurisdiction?
- same target locale?
Exact text can perform different jobs.
A short string such as “Open” demonstrates the problem.
It may be:
- command;
- adjective;
- status;
- menu label.
Project analysis counts the match.
A translator interprets the function.
Step 5: read fuzzy bands as repair zones
A fuzzy match is best understood as:
previous target text that may be cheaper to repair than to rebuild.
The closer the source match, the more likely that is.
But not always.
High fuzzy match
Usually worth inspecting first.
Possible changes:
- one noun;
- one date;
- small clause;
- formatting;
- number.
Medium fuzzy match
May provide:
- useful syntax;
- terminology;
- partial phrase reuse.
But can require substantial rewriting.
Low fuzzy match
May create more interference than help.
A translator might spend longer deleting and repairing a weak match than writing naturally from scratch.
Good workflows learn when to abandon the match.
A fuzzy-match stop rule
Ask:
Is the previous target reducing my decision load?
If yes, repair.
If no, clear it and translate fresh.
The existence of a match is not a command to use it.
Step 6: understand repetitions
Repetition is often one of the most valuable categories.
If the same source segment appears twenty times, the translator should not solve it twenty times.
But the project analysis may assume reuse that depends on actual workflow behavior.
Current platforms can count internal matches differently depending on whether auto-propagation or another reuse mechanism is applied.
So ask:
- are repetitions exact?
- are they context-sensitive?
- will the CAT tool propagate them?
- can legitimate exceptions be preserved?
- are repeated segments spread across files?
Repetition statistics describe potential leverage.
Real leverage requires a reuse process.
Internal repetition versus TM match
These are different.
Translation memory match
The source resembles historical material.
Internal repetition
The source repeats inside the current project.
An internal repetition can be very strong evidence because the current project context may be consistent.
But short strings and context-dependent fragments still need care.
Step 7: understand weighted word count
A weighted word count converts match categories into one effective number.
Simple conceptual formula:
weighted words = sum(words in each band × assigned effort weight)
Example:
- 2,000 new words × 100%;
- 1,000 medium fuzzy × 70%;
- 1,000 high fuzzy × 40%;
- 1,000 exact × 20%.
Weighted total:
2,000 + 700 + 400 + 200 = 3,300 weighted words.
The raw file contains 5,000 words.
The model estimates 3,300 full-word equivalents of effort under that particular grid.
This can help with:
- quoting;
- planning;
- assignment;
- productivity comparison.
But the weights are assumptions.
There is no universal grid that is correct for every job.
Why weighted words are not physical units
A kilogram means the same mass regardless of who measures it correctly.
A weighted word does not have that kind of universal meaning.
Its value depends on:
- discount grid;
- match thresholds;
- resource quality;
- language pair;
- content type;
- review policy;
- project risk.
Two companies can analyze the same file and produce different weighted totals because their effort models differ.
Treat weighted words as a model.
Not as natural law.
Step 8: distinguish pricing model from effort model
A commercial rate grid and a cognitive effort model are related but not identical.
An agency may pay a fixed percentage for 100% matches.
That does not prove every 100% match requires exactly that percentage of work.
Pricing can reflect:
- contracts;
- market norms;
- negotiated rates;
- business policy.
Effort depends on the actual segment.
Do not use a pricing grid as a substitute for project-risk reasoning.
Worked example 1: technical manual
Raw volume: 18,000 words.
Analysis:
- 5,500 repetitions;
- 3,000 context matches;
- 2,500 exact matches;
- 2,000 high fuzzy;
- 1,500 medium fuzzy;
- 3,500 new.
At first glance, the job is highly leveraged.
Then the project manager notices:
- product name changed;
- units differ by market;
- many fuzzy matches contain changed measurements;
- repeated warnings require safety review.
The analysis still saves planning time.
But the weighted estimate needs adjustment.
The right conclusion is not:
Analysis failed.
The right conclusion is:
Analysis measured textual leverage; project risk adds another dimension.
Worked example 2: marketing brochure
Raw volume: 4,000 words.
Analysis:
- 1,500 exact or high fuzzy;
- 2,500 new.
Looks reusable.
But the client requests:
- new tone;
- major brand repositioning;
- transcreation.
Historical matches may be poor fit.
A smaller raw file can require more creative effort per word than a much larger technical manual.
This is why content type matters.
Worked example 3: recurring monthly report
Raw volume: 12,000 words.
Analysis:
- 8,000 exact/context matches;
- 2,000 repetitions;
- 1,500 fuzzy;
- 500 new.
The previous translation was reviewed and approved.
Terminology is stable.
This is a high-leverage project.
The translator can concentrate on:
- new figures;
- changed narrative;
- changed headings;
- new conclusions.
Project analysis has correctly narrowed the work.
Step 9: use analysis to order the translation
An analysis can help decide sequence.
Possible strategy:
- new terminology-rich sections first;
- high-frequency repeated decisions;
- high fuzzy matches containing risk tokens;
- ordinary new content;
- stable exact matches;
- low-risk repetitions.
The right order depends on project structure.
The principle is to settle high-leverage decisions early.
Step 10: use analysis to split work across translators
A 20,000-word project should not necessarily be split into four equal 5,000-word chunks.
One chunk may contain:
- 4,000 repetitions;
- little new work.
Another may contain:
- 5,000 new words;
- dense terminology.
Raw division creates unequal workloads.
Use analysis to understand effective effort before assignment.
Still preserve coherent sections where possible.
Fragmenting a document purely for arithmetic can harm consistency.
Segment count versus word count
Word count tells one story.
Segment count tells another.
Two files may both contain 5,000 words.
File A:
- 150 long segments.
File B:
- 1,200 short UI strings.
The second can be slower because each string creates context switching and individual decisions.
Analysis should therefore consider:
- words;
- segments;
- context type;
- repetition;
- short-string ambiguity.
Volume metrics need interpretation.
Character counts
For some languages or markets, character count is more meaningful than word count.
CAT platforms may report:
- words;
- characters with spaces;
- characters without spaces;
- pages.
Do not compare metrics blindly across language pairs.
The planning unit should match the project’s established method.
Analysis after source changes
If the source files change, rerun the analysis.
Why?
Because revisions can alter:
- repetition counts;
- match bands;
- total volume;
- new terms;
- cross-file leverage.
A quote based on yesterday’s source is not necessarily valid for today’s source.
Version identity matters.
Analysis after TM changes
If you attach or update a major translation memory, rerun the analysis where practical.
The match distribution can change dramatically.
This is especially relevant when:
- client supplies new TM;
- master memory is cleaned;
- aligned legacy material is added;
- old resources are removed.
Workload depends on resource state.
Analysis and project templates
A reliable recurring project template can ensure the analysis uses the right:
- memories;
- locales;
- settings;
- thresholds.
Otherwise two project managers may analyze similar jobs against different resources and obtain incomparable results.
Standardized setup improves planning quality.
Analysis and terminology
A project can show excellent TM leverage but still contain many new terms.
Term extraction and project analysis solve different questions.
Project analysis:
How much historical sentence-level reuse exists?
Term extraction:
Which lexical decisions will recur?
Use both on terminology-heavy projects.
Analysis and machine translation
Some platforms report MT categories or estimated MT leverage.
Be careful.
Machine-generated output is not equivalent to approved TM.
The effort to post-edit depends on:
- domain;
- engine;
- language pair;
- source quality;
- terminology;
- quality threshold.
Do not assign a universal low effort weight to MT merely because it is prefilled.
Measure actual post-editing performance.
Analysis and quality estimation
Quality estimation can further triage machine-generated segments by predicted confidence.
That is another layer.
Project analysis answers:
How much content falls into each reuse class?
Quality estimation asks:
Which machine-generated outputs may deserve less or more human attention?
Keep these tools conceptually separate.
Failure mode 1: quoting from raw words only
Result:
- high-leverage jobs are overestimated;
- new-content jobs are underestimated;
- assignment becomes uneven.
Repair:
- run project analysis before committing.
Failure mode 2: trusting weighted words blindly
A safety-critical manual looks cheap because many fuzzy matches are high.
Repair:
- add risk review;
- inspect changed tokens;
- adjust schedule.
Failure mode 3: analyzing against the wrong TM
Report looks excellent.
Translator cannot safely use the matches.
Repair:
- verify resources before analysis.
Failure mode 4: treating all 100% matches as zero work
Exact matches still may require:
- context check;
- terminology check;
- number check;
- locale check.
Repair:
- define approval policy.
Failure mode 5: treating every fuzzy match as helpful
A weak match can create editing interference.
Repair:
- use a stop rule and translate fresh when necessary.
Failure mode 6: ignoring repetitions
The translator repeatedly solves identical material.
Repair:
- use auto-propagation or controlled repetition reuse.
Failure mode 7: ignoring technical complexity
A tag-heavy XML file and a clean Word document may have identical word counts.
The XML job may require more technical verification.
Repair:
- add file-engineering complexity to planning.
Failure mode 8: using one grid for every content type
Legal, marketing, UI, and technical work differ.
Repair:
- calibrate effort models using actual historical data.
A simple workload matrix
Think in two dimensions.
Leverage
How much can be reused?
- high;
- medium;
- low.
Risk
How costly is a wrong decision?
- high;
- medium;
- low.
Then classify.
High leverage, low risk
Fastest candidate.
High leverage, high risk
Reuse, but verify carefully.
Low leverage, low risk
Mostly fresh drafting.
Low leverage, high risk
Slowest and most attention-intensive.
This matrix is more useful than word count alone.
Add a complexity modifier
A mature estimate may consider:
effective linguistic volume × complexity modifier
Complexity factors:
- poor source writing;
- heavy terminology;
- difficult file type;
- many tags;
- ambiguous UI strings;
- external research;
- legal or medical consequence;
- extensive review requirements.
You do not need a perfect formula.
You need to stop pretending every word is equal.
Calibrate with completed jobs
After delivery, compare:
- predicted weighted volume;
- actual translator hours;
- actual review hours;
- QA time;
- rework.
Over several projects, patterns emerge.
Perhaps 95–99% matches in your domain take more time than the default grid assumes.
Perhaps repetitions are nearly free.
Perhaps UI strings have high overhead.
Use evidence.
A personal translator benchmark
Freelancers can do the same.
For ten projects, record:
- raw words;
- weighted words;
- project type;
- hours;
- effective words per hour;
- review time.
Then compare.
This helps answer:
Which kinds of leverage actually make me faster?
The answer may differ from market assumptions.
Analysis for learners
Students can practise project analysis without professional software.
Take a 1,000-word text and mark:
- sentences already translated before;
- repeated sentences;
- partially similar sentences;
- completely new sentences.
Then estimate effort.
This teaches an important principle:
translation speed depends on reuse structure.
Do not confuse analysis with progress
A pre-project analysis predicts potential reuse.
A project-progress report describes completed work.
They are not the same.
A file may have many exact matches but still be unreviewed.
A file may be 90% translated even if analysis originally predicted little leverage.
Keep forecast and status separate.
Progress metrics can change during work
Internal matches may appear only after earlier segments are translated.
Some tools count repetitions differently as content becomes confirmed.
This means pre-project statistics and live editor statistics may not match perfectly.
Understand the platform’s counting logic before interpreting discrepancies.
Why “zero weighted words” needs caution
A system may assign zero weighted effort to fully automated or highly trusted categories.
That can make sense under a specific workflow.
But zero in a model does not mean zero operational responsibility.
Someone may still need to ensure:
- correct source version;
- correct target locale;
- correct resource;
- correct export.
Use zero only when governance supports it.
A five-minute analysis routine
Before accepting a job:
Minute 1
Check source volume and file types.
Minute 2
Check resources and match distribution.
Minute 3
Inspect high fuzzy samples and exact-match context.
Minute 4
Inspect repetitions, tags, tables, and risky content.
Minute 5
Adjust the time estimate and identify open questions.
Five minutes can prevent a day of scheduling trouble.
A project-manager handoff note
When assigning work, give the translator more than raw word count.
Useful handoff:
- raw words;
- weighted words;
- match bands;
- TM status;
- repetition strategy;
- risk notes;
- deadline;
- review expectations.
The translator can then plan intelligently.
The ethical use of weighted counts
Weighted counts should make work more transparent, not hide effort.
A fair model acknowledges:
- fuzzy matches can vary;
- review still requires attention;
- difficult languages differ;
- technical tasks exist outside word count.
Use weighted metrics as shared planning evidence.
Do not turn them into unquestionable truth.
What analysis cannot see well
A CAT analysis may not understand:
- conceptual difficulty;
- ambiguous source writing;
- cultural adaptation;
- political sensitivity;
- legal consequence;
- rhetorical creativity;
- broken author intent;
- need for external research.
Human preflight must add these factors.
The relationship to document preflight
Document preflight asks:
What kind of text is this, what is wrong with it, and what hidden difficulty exists?
Project analysis asks:
