Most SEO writing about BERT and MUM treats them as a capability upgrade — “Google understands language better now, so write naturally.” Accurate in the loosest sense. Useless in practice. What actually changed is where in the ranking pipeline relevance gets evaluated, and that shift has specific, exploitable consequences for how you build and structure content.
The Pre-BERT Relevance Model (And Why It Still Matters)
Before BERT rolled out in 2019, Google’s relevance scoring at query time leaned heavily on term frequency and positional signals. A page about “best practices for Python logging” ranked partly because those exact tokens appeared in the title, H1, and body at the right density. The system was essentially asking: does this document contain the query tokens in high-signal positions?
That model had a well-known failure mode: it couldn’t distinguish between pages that mentioned a concept and pages that explained it. It also couldn’t handle prepositional nuance. The example Google used in the BERT announcement was “2019 brazil traveler to usa need a visa” — the word “to” flips the entire meaning, and pre-BERT models routinely dropped it.
The old model gave you a clear optimization lever. Put the right tokens in the right places, relevance score goes up. BERT broke that lever without replacing it with anything as simple.
What BERT Actually Does at Query Time
BERT (Bidirectional Encoder Representations from Transformers) runs on the query side of the relevance calculation, not just the document side. This is the part most SEO content skips. Google uses BERT to build a richer representation of what the query is actually asking — context, implied intent, the relationship between words — before it even looks at your page.
The comparison isn’t “your page keywords vs. the query keywords.” It’s “your page’s semantic content vs. a BERT-encoded understanding of what the user needs.” That gap matters most on long-tail and conversational queries, where the surface-level tokens almost never tell the whole story.
Concretely: a query like “can I take ibuprofen after drinking” doesn’t rank pages that contain “ibuprofen” + “drinking” + “after.” It ranks pages whose content maps onto the actual underlying question — drug metabolism, timing, safety thresholds. Answer the surface query but not the implied one, and BERT’s query representation and your document’s representation won’t align well. You lose relevance score even if you technically match the keywords.
Where MUM Is Different — and More Disruptive
MUM (Multitask Unified Model), which Google started deploying in 2021, operates at a different level. BERT is better at understanding a single query in context. MUM can hold multiple subtasks in memory simultaneously — processing a question that requires synthesizing information across formats, languages, and topics in a single inference pass.
The practical consequence: MUM shifts the relevance question from “does this page answer the query?” to “does this page, or the set of pages on this site, satisfy the full informational need behind the query?”
That’s not a minor distinction. A single page that answers the explicit question but ignores the contextual subtasks — related decisions, prerequisites, downstream implications — looks weaker to MUM than a page that addresses the full need, even if both technically answer the query tokens equally well.
This is why deeply structured, multi-subtopic pages keep outperforming tight, focused pages on complex queries. It’s not that word count proxies for quality. It’s that richer pages satisfy more of the subtasks MUM is evaluating, so they score better on the relevance dimension MUM is weighting.
What This Actually Breaks in Standard Keyword Targeting
The standard workflow — pick a focus keyword, build a page around it, optimize title/H1/meta — still works for simple navigational and transactional queries. BERT’s impact there is minimal because those queries are low-ambiguity: “Nike Air Max 270 buy” doesn’t have much contextual nuance to decode.
Where it breaks down is informational queries with meaningful prepositional or contextual structure. MUM compounds this by making topical coverage a relevance factor, not just a quality factor.
The failure modes we see most often:
- Targeting the query token, missing the intent vector. A page optimized for “remote work productivity” that covers tactics but ignores the underlying tension — distraction, isolation, async coordination — will match the keyword but score poorly against MUM’s subtask model for that query cluster.
- Over-relying on keyword density as a relevance signal. BERT doesn’t care if your target phrase appears eight times or two times. It’s reading semantic structure, not counting tokens. Dense repetition can actually hurt if it crowds out the contextual content that builds relevance in BERT’s model.
- Building isolated pages when clusters are required. For MUM-relevant queries — complex, multi-part, research-oriented — a single page rarely holds rankings long-term. The pages that do tend to be deeply interconnected with supporting content that handles the subtasks MUM is evaluating across the broader topic space.
How to Adjust Your Content Planning Accordingly
Not going to tell you to “write for humans, not search engines” — that framing has always been incomplete. Write for the actual informational need, which means doing the work to map what BERT and MUM are evaluating for your target query class.
For BERT: stress-test your content against prepositional and contextual variants of your target query. If your page about “project management for remote teams” answers “what tools should remote teams use” but not “how remote teams should adapt processes they already have,” you’re answering a different BERT-encoded query than the one driving most of the traffic in that cluster.
For MUM: map the subtasks. What decisions, prerequisites, and downstream questions does someone carry when they search this query? A page targeting “how to set up a home network” should probably address “what equipment do I actually need,” “how do I handle multiple floors,” and “what can go wrong during setup” — not because those are keywords you want to rank for, but because MUM is evaluating whether your page satisfies the full informational context of that query type.
The practical tool here is the “People Also Ask” expansion on your target queries, gone two levels deep. The PAA boxes surface exactly the subtasks Google’s models have identified as related to the primary query. If your page doesn’t address them, you’re leaving MUM-relevance on the table.
One Genuine Caveat Worth Sitting With
We don’t have direct visibility into how Google weights BERT and MUM outputs relative to other ranking signals for any given query. Google has confirmed both are in production, but how much they dominate the relevance score versus traditional signals — PageRank, anchor text, domain-level quality signals — varies by query type and vertical. YMYL queries almost certainly run a different signal weighting stack than hobbyist-niche informational content.
So the move isn’t “optimize for BERT and MUM specifically.” It’s “stop optimizing exclusively for pre-BERT relevance signals and start building content that satisfies the full informational need behind your target queries.” The models explain why that matters. They’re not a separate thing to game.
Next time a page drops on a complex informational query despite solid keyword targeting, pull the PAA box expansion and check how many of those subtasks your page actually answers. That gap is almost always the MUM-relevance gap.

