A professional translation programme should know what content exists before it decides what content to translate. A multilingual content inventory is the working map that shows pages, files, product strings, documents, campaigns, help articles, training materials and other source assets that may need translation. It records what each asset is, who owns it, whether it is current, which languages already exist, what risk it carries, and whether it deserves to enter the localization workflow at all.
Searches for multilingual content inventory, localization content audit, translation content audit, multilingual content management, what content should be translated, translation scope, localization inventory and website translation audit all point to the same operational problem. Organisations often begin with a language request before they have a reliable picture of the source estate. That makes budgets unstable, deadlines unreliable and quality control unnecessarily difficult.
This guide shows how to build a translation-ready content inventory without turning the work into an endless spreadsheet exercise. It explains what to count, what not to count, how to connect source assets to owners and audiences, how to identify duplicates and stale content, how to mark risk and update frequency, how to record language coverage, and how to turn the inventory into a decision system for translation rather than a passive catalogue.
This article belongs to eduKateSG’s Master Art of Translation architecture. It complements the existing workflow articles on preparing source content, estimating translation effort, managing translation projects and retiring stale translations, but it owns a different question: what exactly is in the multilingual content estate before work is commissioned?
Quick Read: what a multilingual content inventory actually does
A multilingual content inventory is a structured record of source content and its translation state. At minimum, each row or record should identify the asset, its location, its owner, its purpose, its audience, its current source version, the languages in which a valid translation exists, the date of the last meaningful change, and any constraints that affect translation. More mature inventories also track traffic, legal importance, revenue importance, support demand, update frequency, localization status, terminology dependencies and whether the content is planned for retirement.
The inventory is not the translation memory, termbase or style guide. Those are linguistic assets. It is not the project tracker either, although project status may link to it. The inventory is closer to a map of the things that might generate translation work. It makes visible the relationship between source content, multilingual copies and the people responsible for keeping both aligned.
The most useful principle is simple: do not translate an unknown estate. If nobody knows whether a page is current, duplicated, owned, heavily used or likely to disappear next month, translation is being purchased before the organisation has answered the more basic question of whether that source asset deserves to survive.
1. Begin with the content system, not the word count
A common first instinct is to ask how many words need translation. Word count matters for effort, but it is a poor starting map. Ten thousand words spread across two stable manuals behave very differently from ten thousand words spread across five hundred interface strings, product descriptions, archived pages and constantly changing support articles. The number of words says little about ownership, change frequency, risk or technical complexity.
Start by identifying the content systems that produce source material. That may include a website CMS, product repository, documentation platform, learning-management system, design tool, source-code repository, customer-support knowledge base, document management system, marketing automation platform and shared drives. Each system creates a different kind of translation problem. The inventory should reflect those boundaries because ownership, update signals and technical extraction methods often differ by system.
Once the systems are visible, word count becomes useful inside a larger context. You can then say that a certain knowledge-base collection contains forty thousand words, changes weekly, drives a large share of support traffic and currently exists in three of the six required languages. That is operational information. “Forty thousand words” alone is not.
2. Define the unit you are inventorying
An inventory becomes confusing when one row sometimes means a website, sometimes a page, sometimes a paragraph and sometimes an entire product. Decide what the working unit is. For websites, the unit may be a URL or CMS entry. For software, it may be a resource file, screen, feature or release bundle. For documents, it may be a controlled document rather than each file copy. For training, it may be a course module rather than each slide.
The right level is the smallest unit that can sensibly be owned, updated, approved and translated as a coherent object. If the unit is too large, meaningful differences disappear. One product may contain high-risk legal terms and low-risk marketing copy that should not share one translation priority. If the unit is too small, the inventory becomes maintenance work rather than decision support.
Use hierarchy where necessary. A record can belong to a product, section, campaign or document family. That lets teams view the estate at several scales without mixing the identity of individual assets. The goal is not to create the perfect database model. The goal is to make each translation decision point traceable to a real source object.
3. Record a canonical source location
Every inventory item should point to the authoritative source. This sounds obvious until teams discover the same brochure in email attachments, shared drives, old campaign folders and a CMS. Translators then receive whichever copy is easiest to attach, not necessarily the current version.
A canonical source location solves part of that problem. It may be a URL, document ID, repository path, CMS record ID or controlled file location. What matters is that the inventory distinguishes the source of truth from copies. If the source lives in a system with version history, record the stable identifier rather than only a human-readable filename that may change.
This becomes especially important when multilingual content must be updated later. The translation should remain connected to the canonical source object even if the title, folder or public URL changes. Stable identity prevents the content estate from fragmenting every time the source moves.
4. Add a human owner and an operational owner
Content ownership is often more complicated than a single name. One person may own the meaning, another the publishing system, and another the multilingual programme. The inventory should make those responsibilities visible rather than forcing every future translation question to begin with detective work.
A useful pattern is to record at least two roles. The content owner is accountable for whether the source is current and correct. The operational owner knows how the asset is stored, released or changed. In some environments a third role, the language owner, is useful for target-language approval or terminology governance.
Do not make the field depend on one individual forever. People leave. Prefer teams, functions or durable ownership groups where possible, with the current contact attached. Ownership should answer a practical question: if the translator finds an ambiguity or the localized page is two versions behind, who can make a decision?
5. Record purpose, audience and user action
The same source content may deserve very different localization treatment depending on what it does. A legal notice, emergency instruction, sales landing page, internal training memo and archived blog post can contain similar word counts but have different consequences when misunderstood.
For each inventory item, record a short purpose label such as sell, explain, instruct, warn, document, support, certify or onboard. Then record the main audience and, where useful, the action the user is expected to take. These fields later help with prioritization because they expose why the content exists.
Keep the categories small enough to use consistently. The inventory is not a place for a paragraph-length brief for every page. It should capture enough context to distinguish materially different functions. Detailed translation briefs can be created after an asset is approved for translation.
6. Capture language coverage as a matrix, not a yes/no field
“Translated” is too vague for a multilingual estate. A source asset may exist in French and Japanese but not Arabic. The Japanese version may match the current source while the French version is six months behind. One locale may be fully human reviewed while another exists only as a machine draft. A single translated/not-translated column hides all of this.
Use a language coverage matrix or linked records. For each target locale, record a status such as not started, in progress, current, stale, under review, machine-only, retired or intentionally not localized. The exact vocabulary can be simpler, but the model should distinguish presence from validity.
This is also where locale specificity matters. Spanish for Mexico and Spanish for Spain may be separate operational targets. English for one regulated market may require a different approved text from global English. An inventory built around broad language names can accidentally merge content that is legally or commercially distinct.
7. Track source freshness before translation freshness
A translation can only be current relative to a known source version. Record when the source was last materially changed and, where systems support it, a version ID, revision number or content hash. This gives multilingual teams a way to detect drift instead of relying on memory.
Not every edit should trigger translation. A spelling correction, analytics tag or formatting change may not alter meaning. If possible, distinguish meaningful source change from purely technical or cosmetic change. That lets the inventory drive smarter update workflows rather than generating noise.
For frequently changing content, update frequency is itself a planning signal. A stable annual policy and a daily market commentary should not share the same localization operating model. Freshness data helps teams choose automation, staffing and review depth appropriate to the content’s real behaviour.
8. Mark stale, duplicate and near-duplicate source assets
One of the most valuable outcomes of a localization inventory is discovering content that should not be translated. Websites often contain old campaign pages, duplicated help articles, obsolete PDFs, alternative URLs carrying the same text and documents retained only because nobody has decided to remove them.
Mark exact duplicates and near-duplicates before translation. Exact duplicates should normally point to one source owner. Near-duplicates require judgment: they may represent legitimate market, product or audience differences, or they may be accidental copies that have drifted.
This step is not a mass-deletion exercise. The inventory should surface candidates and ownership questions. Retirement remains a governed decision. The point is to avoid paying for four translations of four source pages when the organisation later discovers that three were redundant copies of the same information.
9. Distinguish current, legacy and planned content
An estate is not static. Some content is live, some is legacy but still legally required, some is scheduled to disappear, and some has not yet launched. Add a lifecycle field that reflects those states.
Current content may be a translation candidate. Legacy content may need preservation but not expansion into new languages. Planned content may deserve translation preparation before publication. Retiring content may need coordinated multilingual withdrawal rather than new translation. These categories prevent teams from treating every discoverable source asset as equally active.
Lifecycle labels also improve budget conversations. It becomes easier to explain why a large archive is excluded from a new-language launch or why a small set of future product pages deserves early investment despite having no current traffic.
10. Add risk as a property of the content, not the language
Risk is often discussed at project level, but the inventory should record it at the asset level whenever consequences differ. High-risk content includes material where mistranslation could affect health, safety, legal rights, financial decisions, eligibility, security, technical operation or significant reputation.
Use a small number of clear risk bands and define what puts content in each band. Avoid using “high risk” as a synonym for “important to the marketing team.” The label should connect to review requirements. For example, high-risk material may require specialist translation, independent revision and explicit sign-off, while low-risk internal content may use a lighter process.
Risk fields later help both prioritization and quality planning. They also make it possible to calculate how much of the untranslated estate carries real consequence, rather than assuming the largest volume is automatically the most urgent.
11. Record update frequency and volatility
A one-time translation decision can create a permanent maintenance burden. Content that changes daily, weekly or with every software release requires a continuous multilingual process. Content that changes once every two years behaves like a controlled document.
Add a volatility field based on observed behaviour rather than optimistic promises. High-volatility content may benefit from connected CMS workflows, translation memory, automation and clear source freeze rules. Low-volatility content may justify more intensive one-time review because the translation will be reused for a long period.
Volatility also changes prioritization. A highly valuable page that changes every day may be expensive to keep synchronized, while a stable page with slightly less traffic may deliver more reliable multilingual value for the same operational effort.
12. Add usage signals without letting analytics become the whole decision
Traffic, downloads, search impressions, support views, conversion contribution and product usage can help distinguish heavily used content from content that almost nobody sees. These signals are useful, especially when the estate is too large to localize at once.
However, usage should not become the sole priority score. A rarely visited emergency procedure can be more consequential than a popular lifestyle article. A low-traffic legal page may be mandatory. A new-market launch page may have no historical traffic because the market does not yet exist.
Treat analytics as evidence about use, not as a complete definition of value. The inventory should let usage sit beside risk, strategic importance, audience need and maintenance cost.
13. Record revenue and service importance carefully
Commercial organisations may want to know which assets influence revenue, activation, retention or support cost. Public and educational organisations may use different value signals such as service access, completion rates or learning outcomes. The field should reflect the organisation’s real mission.
Avoid pretending that every page can be assigned an exact monetary value. Simple bands are often better: direct transaction, supports transaction, reduces support burden, required for trust, informational only. The purpose is to compare content classes, not to manufacture precision.
These value labels become powerful when paired with language demand. A product page important to conversion may deserve priority in a market where users repeatedly abandon the source-language version. The inventory provides the evidence needed to see that relationship.
14. Capture technical constraints before translation begins
Some content is easy to extract and republish. Other content is embedded in images, PDFs, source code, design files, database fields or third-party platforms. Technical accessibility affects scope and cost, so the inventory should record it.
Useful fields include content format, extraction method, character limits, markup type, placeholder presence, right-to-left implications, image text, subtitle timing, supported locale architecture and whether the asset can be synchronized automatically. Do not turn the inventory into a full engineering specification; simply expose constraints that materially change the translation workflow.
This can prevent a common failure: approving a language based on visible text and discovering later that half the user journey sits inside unextractable graphics or a vendor platform that does not support the locale.
15. Connect assets to terminology and product domains
Translation quality improves when content is linked to the domain knowledge it needs. An inventory can include a product, subject or terminology-domain field that tells downstream workflows which termbase, style guide or specialist pool should apply.
This is especially useful in large organisations where “translation” covers unrelated material such as medical content, software interfaces, marketing campaigns and legal policies. The same target language may require different specialists and different approved vocabularies.
The field can remain simple: product family, department, document type or controlled domain label. Its value comes from routing, not description. A clear domain label helps the correct linguistic assets and reviewers attach automatically later.
16. Identify content relationships and dependencies
Pages and documents rarely stand alone. A help article may link to a policy. An interface string may depend on a screenshot. A course lesson may reference a worksheet. A legal notice may be reused in several product flows. Translation planning improves when these dependencies are visible.
You do not need to model every hyperlink. Focus on relationships that affect release or meaning. If one source update requires three related translations to change together, that relationship belongs in the inventory or the system connected to it.
Dependencies help prevent partial localization. A translated landing page that sends the reader to untranslated checkout instructions may create a worse experience than delaying the page until the whole critical path is ready.
17. Separate source inventory from translation asset inventory
The source estate and the linguistic asset estate overlap but are not the same. Source inventory answers what content exists. Translation assets include translation memories, termbases, style guides, glossaries, reference translations, language-specific legal wording and reviewer instructions.
Keep the concepts separate even if the same platform stores both. This makes governance clearer. Source owners decide whether a page remains valid; language owners decide whether an approved term remains valid. One cannot automatically substitute for the other.
The distinction also makes audits easier. A team can ask whether all active source content is represented in the multilingual workflow and separately ask whether all active languages have current linguistic resources.
18. Do not treat crawler output as the finished inventory
Website crawlers are useful discovery tools. They can enumerate URLs, status codes, titles, word counts and duplicate patterns quickly. But a crawl is evidence, not governance. It does not know which page is authoritative, whether a URL is legally required, whether a campaign is over or who owns the content.
Import crawler output into a review process rather than publishing it as truth. Combine technical discovery with CMS metadata, analytics, ownership information and human review. The same principle applies to file-system scans and repository scripts.
An automated inventory that nobody trusts becomes another data set to reconcile. A smaller, verified inventory that reflects operational reality is more useful than a perfect-looking export with no accountable owner.
19. Build the minimum viable inventory first
Large organisations can spend months designing an ideal content model. That delays the decisions the inventory was meant to support. Start with the smallest fields needed to answer the next real localization question.
A practical minimum may include asset ID, canonical source, title, owner, content type, purpose, audience, status, last meaningful update, current language coverage, risk band, update frequency and next action. Add analytics, value, dependencies and technical detail when they materially improve decisions.
The inventory should earn its complexity. Every field creates maintenance work. If nobody uses a field to route, prioritize, budget, review or retire content, consider whether it belongs in the core view.
20. Decide how the inventory stays current
A one-time audit becomes obsolete as soon as the source changes. The inventory needs an update mechanism. The best approach depends on the systems involved: automatic sync from CMS metadata, release-event updates, scheduled audits, source-owner review or a combination.
Assign responsibility for data quality. Translators should not become the default owners of source metadata simply because they notice problems. Source teams should own source truth; localization teams should own language-state truth; technical teams may own automated connections.
Use timestamps and confidence. If a field has not been verified for a year, the inventory should make that visible. Unknown is a valid state. False certainty is more dangerous because it causes teams to act without checking.
21. Turn inventory fields into translation decisions
The inventory becomes valuable when it drives action. A content record should be able to move toward one of several decisions: translate now, translate later, translate only for selected locales, update an existing translation, retire the translation, retain source only, investigate ownership, or block pending source cleanup.
Those decisions should be explainable from fields rather than intuition alone. A high-risk, high-use, stable asset serving a target market with strong language demand is an obvious candidate. A stale duplicate with no owner and planned retirement is an obvious non-candidate.
The middle cases are where the inventory earns its keep. It lets teams discuss trade-offs with shared evidence instead of arguing from whichever page or stakeholder is loudest that week.
22. Use the inventory to prevent orphan translations
An orphan translation is a target-language asset that no longer has a clear source owner, source version or maintenance path. These are expensive because nobody knows whether they are safe to update, safe to publish or safe to remove.
Prevent orphaning by requiring every multilingual asset to point back to a canonical source record. When source content is retired, the relationship triggers a multilingual retirement decision. When ownership changes, the new owner inherits responsibility for the translated versions as well.
This connection is especially important during migrations. A new CMS often focuses attention on source content, while translated pages are copied later as an afterthought. Inventory relationships make multilingual migration part of the same system instead of a separate rescue project.
23. Use the inventory during market and language expansion
When a new target language is proposed, the inventory can generate a realistic scope instead of a vague “translate the website” request. Teams can filter by content status, user journey, risk, market relevance, product availability and technical support.
This leads to staged language launches. Critical navigation, onboarding, transactional pages, help content and high-risk policies may launch first. Deep archive material may remain source-language only. The choice becomes deliberate and documented.
A staged launch is not automatically lower quality. It can be higher quality because the translated surface is complete where it matters. The inventory helps teams avoid the appearance of multilingual coverage without the actual ability to complete essential tasks in the target language.
24. Use the inventory to detect translation debt
Translation debt is the gap between what multilingual users need and what the content system can reliably maintain. It appears as stale pages, missing locales, disconnected translations, unresolved updates, outdated terminology and target content that no longer maps cleanly to source.
An inventory can make this debt measurable. Count current translations, stale translations, missing high-priority locales, assets with no owner and items whose source changed after the target was last reviewed. The exact metric matters less than making the maintenance gap visible.
Do not turn debt into a shame score. It is a planning signal. Some debt is rational because resources are limited. The professional question is whether the organisation knows where the debt sits and whether the most consequential gaps are being reduced.
25. Use the inventory during provider handoff
When translation providers change, the content inventory gives the new team a map of the operational estate. Without it, handoff often begins with folders, memories and partial lists assembled by whoever happens to be available.
Provide the inventory together with translation memories, termbases, style guides, decision logs and current project status. The inventory tells the new provider what content those language assets are meant to support.
This also makes provider boundaries clearer. The organisation retains ownership of its source map even when translation production is outsourced. A vendor may help maintain fields, but the inventory should not become inaccessible when a contract ends.
A practical multilingual content inventory schema
A useful inventory does not need hundreds of columns. The following field set is broad enough for serious work while remaining understandable:
- Asset ID: stable identifier independent of title or URL.
- Canonical source location: authoritative CMS record, file or repository object.
- Title or label: human-readable name.
- Content type: webpage, UI, document, video, course, support article, policy or other class.
- Owner: accountable source team and operational contact.
- Purpose: explain, transact, instruct, warn, support, sell, certify or another defined function.
- Audience: principal user group.
- Lifecycle state: planned, current, legacy, retiring or archived.
- Source version: revision ID, date or meaningful-change marker.
- Language coverage: status for every required locale.
- Risk: consequence band tied to review requirements.
- Volatility: typical update frequency.
- Usage: traffic, product use, downloads or another relevant signal.
- Value: transaction, service, support, strategic or compliance importance.
- Technical constraints: format, extraction, placeholders, character limits, media dependencies.
- Domain: product or subject area for terminology routing.
- Dependencies: related assets that must remain coordinated.
- Next action: translate, update, investigate, hold, retire or no localization planned.
Worked example: a multilingual help centre
Imagine a help centre with 1,200 articles. The naive scope is 1,200 articles multiplied by six target languages. The inventory reveals a different picture. Two hundred articles belong to an old product scheduled for retirement. One hundred and fifty are near-duplicates created during previous migrations. Three hundred have almost no traffic but include a handful of legally important instructions. Four hundred serve the current product, and sixty of those account for most support visits.
The inventory also shows that German coverage is mostly current, French is uneven, Japanese exists only for the current product, and two newly requested languages have no translations. Some high-traffic articles change every week because the interface is evolving, while account-security articles are stable but high risk.
Now the translation problem is visible. The launch can prioritize the current product, critical journeys, security content and high-use support pages. Duplicates can be resolved before translation. Retiring product content can be excluded from new-language expansion. Frequently changing pages can be connected to a continuous localization process rather than treated as one-time jobs.
The inventory did not reduce the value of translation. It made the work more serious by connecting language effort to the actual content system.
Worked example: an educational content library
An educational publisher may have lessons, worksheets, teacher guides, videos, quizzes, answer keys and parent notices. If these are inventoried only as files, the translation team can miss instructional dependencies. A student worksheet translated into a new language may still point to an untranslated video or use terminology different from the teacher guide.
Organize the inventory around learning units and dependency relationships. Mark the learner age, subject, curriculum role, assessment consequence and whether the content can be adapted or must remain aligned to a source standard. Record which assets form a complete lesson path.
Translation priority can then protect educational completeness. Instead of translating the most visible worksheets first, the organisation can localize coherent learning sequences that a student can actually complete.
Worked example: a software product
Software inventories often need two connected layers: product strings and supporting content. The application may contain UI labels, validation messages, onboarding flows and notifications, while the wider product estate contains release notes, documentation, marketing pages and help articles.
Do not assume the string repository alone represents the user experience. A user can complete an interface in their language and then hit an untranslated support page at the first difficult moment. The inventory should connect product features to the surrounding content needed to use them successfully.
Release cadence matters. If the product changes weekly, the inventory should receive version signals automatically where possible. Source changes that affect localized strings, screenshots and documentation can then be grouped rather than discovered separately after release.
Common failure: inventorying everything before deciding why
Some teams respond to uncertainty by collecting every possible field across every possible system. Six months later they have a giant data project and still cannot answer which pages should be translated next.
Reverse the logic. Start with the decisions you need to make. If the immediate problem is launching two new languages, inventory the content needed to define launch scope. If the problem is stale translations, capture source-version and language-state data. If the problem is cost, add volume and volatility. Expand the schema only when a new decision requires new evidence.
The inventory is infrastructure for action, not an archive of everything knowable about content.
Common failure: assuming every existing translation is valid
Discovery tools often show that a translated file or page exists. Existence is not the same as validity. A page may be incomplete, machine generated, based on an old source, translated for another locale, never approved or disconnected from the current design.
Use status labels that express confidence. “Exists” can be a discovery fact. “Current and approved” should require stronger evidence. This distinction prevents inherited multilingual content from quietly entering production as if its quality and provenance were known.
Where provenance is missing, mark it as unknown and audit it before reuse. That is more professional than pretending uncertainty has disappeared because a file opens successfully.
Common failure: making the localization team responsible for source hygiene
Localization often exposes duplicate, broken or ownerless source content because translators must understand what they are translating. That does not mean localization should permanently own every content-governance problem.
Use the inventory to route problems back to source owners. Localization can flag a contradiction, stale page or duplicate. The responsible source team decides whether to correct, merge, retire or preserve it. Clear boundaries prevent the translation programme from becoming an unofficial repair department for the entire organisation.
At the same time, localization should not simply accept poor source hygiene. Translation multiplies ambiguity. The inventory creates a formal point at which weak source content can be stopped before it spreads across languages.
Common failure: inventorying by language instead of by source identity
If each language team builds its own content list, the organisation quickly creates several incompatible maps. French may identify pages by URL, Japanese by spreadsheet row, Arabic by filename, and nobody can reliably tell which target items correspond to one source asset.
Anchor the inventory in source identity and attach language states to that identity. Language-specific details can live in linked records, but the relationship back to the same canonical source must remain intact.
This one design choice dramatically simplifies updates, retirement, audits and coverage reporting because all languages can be compared against the same source estate.
A 30-day implementation path
Week 1: define the decision and unit. Choose the content system and business question. Decide what one inventory record represents. Create a minimal schema and test it on twenty representative assets.
Week 2: discover and verify. Export or crawl candidate assets, then have source owners verify canonical locations, lifecycle states and ownership. Mark duplicates and unknowns instead of forcing uncertain answers.
Week 3: attach language state. Add target-language coverage, source-version relationships and confidence status. Connect major technical constraints and risk categories.
Week 4: turn the map into decisions. Filter the estate into translate now, translate later, update, investigate, retire and intentionally source-only. Use the result to define the next translation programme scope.
After thirty days, the inventory does not need to be complete for every content system. It needs to be trustworthy enough to support a real decision and structured enough to expand without losing identity.
Quality checks for the inventory itself
- Can each active multilingual asset be traced to one canonical source?
- Can a source owner be identified without asking the localization team?
- Does language status distinguish current from merely existing?
- Can the team tell which source version each important translation represents?
- Are high-risk assets visible even when traffic is low?
- Are duplicates and retirement candidates marked rather than silently translated?
- Can a new target language be scoped from the inventory without recrawling the entire estate?
- Can stale translations be identified by comparing source and target update state?
- Can the system distinguish unknown information from confirmed information?
- Does every field support a real decision, routing action or quality check?
Frequently asked questions
Is a multilingual content inventory the same as a translation management system?
No. A translation management system manages translation workflows and linguistic assets. A content inventory maps the source assets that may create those workflows. The two can be connected, and in mature environments they often should be, but they answer different questions.
Should we inventory every page before translating anything?
Not necessarily. Start with the content systems and user journeys relevant to the next decision. A complete enterprise inventory may take time. A verified inventory of the launch-critical estate can still support responsible action.
Can a website crawler build the inventory automatically?
It can discover much of the visible web estate and provide useful metadata, but it cannot reliably determine ownership, strategic value, legal necessity, lifecycle intent or whether two similar pages are legitimate variants. Human governance remains necessary.
What if nobody knows who owns a page?
Record the owner as unknown and treat that as a risk signal. Do not invent ownership. Ownerless content should usually be investigated before expensive translation because nobody may be available to answer source questions or approve later updates.
How often should the inventory be updated?
The answer depends on volatility. Frequently changing product content may need event-driven or automated updates. Stable policy libraries may be reviewed periodically. The key is to match update frequency to the rate at which the source estate changes.
Should machine-translated pages count as translated?
Record them as a distinct state. A machine-generated target can be operationally useful, but it should not be indistinguishable from a reviewed and approved translation if the quality process is different.
What is the most important field?
Canonical source identity is the foundation because language states, ownership, versions and retirement decisions all need something stable to attach to. Without source identity, the inventory becomes a list of disconnected copies.
How does the inventory help with translation cost?
It prevents unnecessary work, exposes duplicates, separates stable from volatile content, identifies reusable domains, and creates realistic scope. It also helps budget review depth according to risk rather than applying one expensive process to everything.
Can the inventory include non-text assets?
Yes. Images with text, audio, video, subtitles, downloadable files and interactive elements can all create translation requirements. If a user must understand the asset to complete a journey, it belongs in the content model even if the translated words live elsewhere.
What happens after the inventory is built?
The next step is prioritization. The inventory tells you what exists and its relevant properties; a prioritization model uses those properties to decide what should be translated first, what should wait, and what should not be translated at all.
The larger lesson: translation begins with content governance
A multilingual programme is often described as a language operation, but its first dependency is content governance. Translators cannot keep an estate synchronized if the organisation does not know what the estate contains, which version is authoritative or who owns the source meaning.
A strong multilingual content inventory turns the invisible content system into something inspectable. It prevents teams from translating stale copies, exposes incomplete language journeys, supports realistic budgeting, gives market launches a defensible scope and makes multilingual maintenance possible after the launch excitement is over.
The professional question is therefore not “How many words do we have?” It is “Which source assets exist, which ones matter, who owns them, how do they change, and what is the multilingual state of each?” Once that map exists, translation can become a managed system rather than a sequence of requests.
Continue with Retire Stale Translations Without Leaving Multilingual Content Behind for the other end of the content lifecycle, or return to the Master Art of Translation for the wider architecture.
