Inside the CulturalFit Score: grading cultural accuracy, not just grammar. | Qorrvio
BlogProduct & Technology

Inside the CulturalFit Score: grading cultural accuracy, not just grammar.

Linh Dao Sørensen·Co-Founder & CTO
26 May 2026
8 min read

When we started building Qorrvio, we spent the first six months talking to localization buyers — heads of international, e-commerce directors, senior brand managers responsible for cross-border rollouts. The question we kept asking was simple: how do you know when localization is good?

Almost universally, the answer came back as some variation of: "We ask a native speaker." And when we pushed on how that review was structured, it was usually informal, inconsistent, and unable to produce an answer that could be compared across linguists, markets, or time.

That's the problem the CulturalFit Score is designed to solve. Not whether copy is grammatically correct — there are mature tools for that — but whether it is culturally accurate: whether it sounds like it was written in that language, for that market, by someone who lives there.

What the score actually measures

The CulturalFit Score is a composite score from 0 to 100, generated by combining AI analysis with structured human review. It evaluates localized content on four dimensions, each of which we identified through our initial research as independently capable of making copy feel foreign even when it is technically correct.

Dimension 1: Tone match

Does the energy of the source content survive in the target language? A brand that communicates with confidence and warmth in English should communicate with confidence and warmth in Korean. Tone match measures whether the emotional register of the localized content is consistent with both the source and with what that register implies to a native reader of the target language. A brand that uses casual, conversational English copy often needs substantially different structural choices in German to achieve the same feeling of approachability — the words change, but the experience of reading them should not.

Dimension 2: Idiomatic naturalness

Is the copy natural or does it read like a translation? This is the dimension that machine translation has historically struggled with most. Idioms, collocations, fixed expressions, and the specific combinations of words that native speakers reach for intuitively are almost impossible to generate from statistical pattern-matching alone. Our linguists flag unnatural phrasings — grammatically legal constructions that no native speaker would actually produce — and the AI component learns from these flags over time, building market-specific naturalness models that improve with every project.

Dimension 3: Brand-glossary adherence

Every brand using Qorrvio can upload a glossary: product names, proprietary terms, preferred translations for brand-specific concepts, and terminology that should never be translated. Brand-glossary adherence measures consistency — whether the localized content uses these terms correctly and consistently throughout. This dimension is purely deterministic: either the glossary is followed or it isn't. We surface adherence at the term level so brands can see exactly which terms caused deductions.

Dimension 4: Cultural appropriateness

Does the content respect the cultural context of the target market? This is the broadest and hardest dimension to evaluate systematically. It encompasses formality level, references that may carry different weight in different markets, visual culture signals embedded in copy (references to seasons, celebrations, family structures), and the specific trust signals that cause a shopper in a given market to feel safe. Our linguists are domain specialists — they work in fashion, consumer electronics, health and beauty, or home goods, not across all categories simultaneously — so their cultural appropriateness judgements are informed by category-specific market knowledge, not just general language fluency.

"Grammar checkers have been solved. The hard problem is whether your copy feels right. That's what we built a scoring system to answer."

How the AI and human layers interact

The CulturalFit Score is not a purely AI score, and it is not a purely human score. Both matter. The AI component provides speed, consistency, and scale: it can evaluate thousands of segments simultaneously, flag potential issues for human review, and apply patterns learned from the entire corpus of projects on the platform. The human component provides depth, context, and the lived cultural knowledge that machine systems cannot yet replicate.

In practice, the pipeline works like this: a completed localization project is first run through our AI quality layer, which produces a preliminary CulturalFit Score and a list of specific segments that fall below threshold on one or more dimensions. These flagged segments go to a senior reviewer — a different linguist from the one who produced the translation — who evaluates each flag, accepts or overrides it with a documented reason, and can add additional flags the AI missed. The final score is the reconciled output of both layers.

This architecture means the score is auditable. Every deduction has a source: either an AI flag that the human reviewer confirmed, or a human flag that the reviewer added. Brands can see exactly what caused any score below their quality floor, and request remediation on specific segments rather than entire projects.

Calibration and market specificity

One of the early engineering challenges was that "naturalness" and "cultural appropriateness" are not universal concepts — a score of 85 on idiomatic naturalness means something different for Swiss German than for Brazilian Portuguese, because the reference corpuses, the stylistic norms, and the expectations of native readers differ. We addressed this through market-specific calibration: our scoring models are trained and validated separately for each language-market pair, not at the language level alone.

This matters because French for France and French for Québec have substantially different idiomatic norms. Spanish for Spain and Spanish for Mexico diverge in ways that go well beyond vocabulary differences. Brazilian Portuguese and European Portuguese have enough stylistic distance that content optimised for Lisbon can read as foreign in São Paulo. The CulturalFit Score treats these as distinct markets, because brands operating in them experience them as distinct markets.

What the data shows

Across the 3,200+ brands that have run projects through Qorrvio, we have accumulated a substantial dataset linking CulturalFit Scores to downstream commercial outcomes. The correlation between score and conversion performance is strong enough that we now surface it proactively: brands targeting premium markets — where customer expectations are highest and competitive localization quality is most differentiated — we recommend a quality floor of 90. For volume-tier markets where speed is prioritised over polish, an 82–85 floor typically produces commercially adequate results.

The score also functions as a signal for linguist matching. Our platform tracks each linguist's average CulturalFit Score by domain and market over time. Brands with premium quality requirements are matched with linguists whose track record puts them consistently above the required floor — not just linguists who claim expertise in a language pair.

"Every deduction has a source. Brands can see exactly what caused any score below their quality floor."

What we're building next

The current CulturalFit Score evaluates copy after localization is complete. The next version — in private beta with a subset of enterprise brands — provides real-time scoring during the linguist's workflow, flagging potential issues before the project is submitted. This closes the feedback loop faster and reduces remediation cycles. We expect it to lift average CulturalFit Scores by 4–6 points across the platform when it reaches general availability.

We are also expanding the score's scope to cover visual-copy interaction: the relationship between images, design, and text in a localized storefront. A product image that works well in one market can undermine culturally accurate copy in another if the visual signals conflict. That's a harder problem, but one worth solving.

See how CulturalFit Scores factor into Qorrvio's plan tiers and pricing.

View Pricing →

Ready to go further?

See what Qorrvio can do for your brand.

Start free