Quality Can No Longer Be Evaluated Only Through Linguistic Metrics

For a long time, quality in translation and localization was assessed through a relatively simple lens: grammar, syntax, terminology, spelling, and fluency. Those criteria still matter. But they are no longer enough.

In environments where content moves across product, marketing, support, legal, audiovisual, gaming, documentation, and interfaces, a more important question has emerged: does this content actually work in its context of use?

In other words, a translation can be linguistically correct and still fail where it matters most.

It can:

  • slow the user down,
  • create ambiguity in an interface,
  • weaken brand consistency,
  • break immersion,
  • drive support tickets,
  • increase regulatory risk,
  • or undermine an international launch.

That is why quality can no longer be treated as a simple total of linguistic errors. It must be evaluated as a matter of usability, perception, and continuity.

Linguistic metrics are still necessary, but they are incomplete

Traditional LQA approaches still provide real value. They help teams:

  • identify objective errors,
  • categorize issues,
  • prioritize fixes,
  • compare vendors or workflows,
  • and define tolerance thresholds by content type.

The problem is not that these metrics exist. The problem starts when teams expect them to answer questions they were never designed to cover.

For example:

  • Is the text clear in a mobile interface?
  • Is a subtitle readable at the right pace?
  • Does a campaign preserve the same brand promise across markets?
  • Is a legally acceptable translation also safe to publish?
  • Does a series of content remain consistent across episodes, versions, seasons, or updates?

None of those questions can be resolved by counting linguistic errors alone.

Quality depends on the context of use

This is the central point: linguistic quality is not absolute. It depends on the type of content, its intended use, cultural expectations, level of risk, and the business outcome it is meant to support.

An error message, a product page, a subtitle, narrative dialogue, a CRM email, a knowledge base article, and an acquisition banner should not be judged in the same way.

The book Localisation linguistique : du chaos à la stratégie makes this point clearly: quality should be assessed in proportion to the content, its use, and its impact. Good evaluation is not about abstract perfection. It is about fitness for real purpose.

That changes how quality should be designed and measured.

A correct translation can still be a poor user experience

This is one of the most common blind spots.

A piece of content can be:

  • grammatically flawless,
  • terminologically compliant,
  • smooth to read,

and still be ineffective in practice.

Here are a few common cases.

In product

A UI label may be linguistically correct but too long for the screen, unclear in the user journey, or inconsistent with other steps in the flow.

In marketing

A call to action may be well translated and still lose persuasive force, clarity, or cultural fit.

In support

An article may be correct but fail to help users solve their problem quickly enough, increasing support contact volume.

In narrative content

A line may be faithful to the source and still break a voice, tone, rhythm, or sense of immersion built over time.

In all of these cases, the quality issue is not primarily linguistic. It is a usage issue.

Quality is also about consistency over time

One of the most important contributions of current quality approaches is the reminder that content almost never exists in isolation.

It sits within multiple forms of continuity:

  • terminological continuity,
  • tonal continuity,
  • narrative continuity,
  • UX continuity,
  • brand continuity,
  • continuity across formats,
  • and continuity across versions and updates.

This matters especially in high-volume environments or recurring publication cycles.

Serial content and consistency

In content published across episodes, seasons, versions, or successive releases, quality is not limited to each segment in isolation. It depends on the ability to maintain stable reference points from one piece of content to the next.

A translation can be excellent sentence by sentence and still create a degraded overall experience if:

  • terminology starts to drift,
  • characters or personas change voice,
  • writing conventions vary without reason,
  • visual, textual, and contextual references stop aligning.

At that point, quality becomes a discipline of continuity, not just of correctness.

Quality perception is multimodal

In many formats, meaning does not depend on words alone.

Images, lettering, visual hierarchy, timing, audio, layout, paratext, screen constraints, and accessibility conventions all shape perceived quality.

Audiovisual

In subtitling, linguistic accuracy is not enough. Readability, segmentation, display pace, text density, and harmonization across episodes matter just as much.

Comics, visual content, and social media

Text interacts with the image. A wording choice that is theoretically accurate can still feel wrong if the lettering overflows, the visual joke disappears, or the text clashes with what the image already suggests.

Product and interfaces

A localized string is not “good” simply because it is correct. It also needs to be understandable in its actual placement, compatible with technical constraints, and consistent with the rest of the journey.

In other words, quality is often situated. It cannot be evaluated outside its display and interaction context.

Fluency is not a reliable synonym for quality

Another common trap is to overvalue fluency. Text that “sounds good” is not automatically high-quality text.

It may be:

  • too far from the source intent,
  • overly smoothed out and missing useful nuance,
  • imprecise from a product standpoint,
  • misleading in a regulated context,
  • or inconsistent with approved terminology or brand voice.

Fluency can be one component of quality. It is not the same thing as quality.

A serious evaluation model needs to distinguish multiple dimensions instead of collapsing judgment into a general reading impression.

What metrics alone do not show: publication risk

Another limit of purely linguistic approaches is that they do not answer the most critical business question well: is this content safe to publish?

A text may show few visible linguistic errors and still be problematic because it creates risk elsewhere:

  • liability,
  • compliance,
  • brand,
  • customer impact,
  • operations,
  • time to fix,
  • strategic impact.

This is an important shift. The question is no longer just about the quality of the text. It becomes a question of decision quality.

From that perspective, a linguistically correct translation may still be unacceptable if it:

  • creates room for contractual misunderstanding,
  • weakens a regulated message,
  • damages a brand promise,
  • causes usage errors,
  • or becomes too costly to correct after publication.

Quality assessment therefore needs to account for actual publication risk, not just the linguistic surface of the content.

Quality must be connected to user experience

The more organizations localize product, self-service, and transactional content, the clearer one fact becomes: users do not evaluate localization like linguists. They experience it.

What they perceive is not an LQA score. It is:

  • how easy something is to understand,
  • how much confidence they have in the message,
  • how consistent the journey feels,
  • how clear the expected action is,
  • how much cognitive effort is required,
  • whether the experience feels intentionally designed for them.

That is why stronger quality practices now connect assessment to signals closer to real-world use, such as:

  • journey completion rates,
  • onboarding success,
  • lower support ticket volume,
  • content engagement,
  • retention,
  • conversion,
  • internal reviewer satisfaction,
  • stakeholder confidence in published content.

These indicators do not replace linguistic analysis. They complement it where linguistic analysis alone no longer goes far enough.

Toward a three-level model: language, use, and confidence

To move beyond a narrow view of quality, it is useful to assess it across three levels.

1. Linguistic quality

This is the foundation: accuracy, grammar, terminology, style, omissions, meaning errors, and local conventions.

2. Usability quality

Is the content clear, navigable, relevant, culturally appropriate, and effective in its real-world context?

3. Confidence quality

Do product, marketing, legal, support, and in-market teams trust what is being published? Are validation signals strong enough to support a release decision?

This third dimension is often underestimated. Yet when reviewer confidence, market confidence, or business-team confidence is low, friction increases: late feedback, over-validation, launch delays, bottlenecks, and heavier reliance on post-publication fixes.

Not all content needs the same quality model

A mature quality policy does not apply the same standard or the same thresholds to every content type.

It is more useful to think in content classes.

High-risk content

Examples: legal, regulated, medical, financial, product safety, critical documentation.

Here, evaluation needs to be strict, with strong attention to accuracy, compliance, and publication risk.

Conversion- or brand-driven content

Examples: campaigns, web pages, brand messaging, CRM, acquisition.

Quality here needs to account for persuasion, tonal consistency, cultural adaptation, readability, and business performance.

Interface and support content

Examples: UI, help centers, transactional emails, onboarding.

The key criterion becomes functional clarity: helping users understand quickly, act correctly, reduce effort, and avoid mistakes.

Narrative and immersive content

Examples: games, audiovisual content, content series, brand storytelling.

Quality depends heavily on continuity of voice, tone, rhythm, references, and overall experience.

This kind of segmentation makes it possible to define error categories, tolerance thresholds, and validation methods that are actually useful.

How to evolve your evaluation model

Moving from purely linguistic quality to situated quality does not mean lowering standards. It means broadening the framework.

Here is a practical approach.

1. Define quality by content type

For each content family, explicitly define:

  • the intended use,
  • the level of risk,
  • cultural expectations,
  • the stakeholders involved,
  • the KPIs that may be affected.

Without this step, evaluations stay too generic.

2. Keep a structured LQA layer

Maintain error categories, severity levels, and thresholds. They still matter for creating a shared baseline and making deviations visible.

But adapt them to the content instead of applying them uniformly.

3. Add a contextual evaluation layer

Include criteria such as:

  • clarity in real context,
  • consistency across screens or assets,
  • readability,
  • cultural appropriateness,
  • brand consistency,
  • continuity over time,
  • ease of action for the user.

4. Assess publication risk

Before release, ask not only “is this correct?” but also:

  • is this publishable without disproportionate risk?
  • what happens if we publish it as is?
  • what will a late fix cost?
  • what is the likely impact on customer experience or compliance?

5. Connect quality to downstream metrics

Pair linguistic evaluation with operational and business indicators such as:

  • support contacts,
  • resolution time,
  • journey completion,
  • conversion,
  • engagement,
  • retention,
  • publishing speed,
  • number of internal review rounds,
  • time spent in review.

This link is essential if quality is to move beyond theory.

6. Measure consistency over time

Create mechanisms to monitor quality across publication cycles:

  • living glossaries,
  • voice and tone guides,
  • libraries of approved examples,
  • cross-functional review practices,
  • controls for variation across versions.

Consistency is not a fixed state. It is a governance practice.

7. Involve non-linguistic teams

Product, marketing, support, legal, in-market teams, and operations all need to contribute to the definition of quality.

Why? Because a strong quality decision rarely depends on linguistic judgment alone. It depends on balancing accuracy, usage, risk, brand, and speed.

The real shift: from text quality to experience quality

This is the deeper change.

For years, many organizations treated quality as a property of an isolated text. That is no longer enough. Today, quality is increasingly a property of the full localized experience.

That experience needs to be:

  • correct,
  • understandable,
  • consistent,
  • culturally appropriate,
  • workable in its real context,
  • safe to publish,
  • sustainable over time.

This is especially true in an era of high-volume content flows, automation, simultaneous launches, and hybrid workflows that combine AI with human expertise. The faster the system moves, the more expensive a narrow definition of quality becomes.

Conclusion

Linguistic metrics are not going away. They remain a necessary foundation for any serious quality practice.

But they can no longer carry the full definition of quality on their own.

A mature organization now evaluates quality across several dimensions: linguistic quality, usability quality, consistency over time, user perception, stakeholder confidence, and business risk.

That shift is what allows teams to answer the real question: not just “Is this translation correct?” but “Is it appropriate, consistent, reliable, and ready to deliver the intended outcome in the real world?”


Photo by Patricia Serna from Unsplash