MUVERA - The Technology That Will Speed Up Google Search

MUVERA - The Technology That Will Speed Up Google Search

When you run a search on Google, it finds and presents the most relevant results from billions of pages within seconds. Understanding how it does this brings us to the concept of "embedding": a method for translating the meaning of words, numbers, and images into numerical values that computers can process. Artificial intelligence assigns each piece of content a coordinate in a virtual space, positioning semantically similar content close together. The search engine's algorithm then finds the most accurate results by measuring how close these coordinates are to one another.

Until recently, each piece of content was represented by a single vector. This was fast, but accuracy suffered because a single vector couldn't capture the depth of the content. More advanced multi-vector models then began to emerge. These models didn't reduce each word to a single vector; instead, they represented the whole piece of content as the sum of many vectors. This produced more accurate results, but at a slower speed. MUVERA combines the strengths of both approaches, addresses their drawbacks, and delivers the most accurate results as fast as possible.

Embedding Technology - Speed or Accuracy?

For years, search engines have relied on single-vector embedding technology. Think of it as summarizing an entire book with a single word: search for that word, and you can easily find the book among billions of others, but you miss everything else the book describes.

To address this shortcoming, multi-vector embedding technologies such as ColBERT were developed. These models summarize each chapter of a book separately, then bring the summaries together, so a search touching one chapter can draw on the book as a whole. This approach finds what you're looking for more accurately, but it takes longer: creating hundreds of vectors instead of one increases the amount of data to process, and as the similarity calculation grows more sophisticated, the processing involved becomes more complex too. As a result, search time and cost both increase.

This is where MUVERA comes in, combining the strengths of both approaches. It does this by reducing the multi-vector set to a single vector, making the process more efficient. This method is called Fixed Dimensional Encoding, or FDE for short. MUVERA takes the complex but semantically rich vector set that multi-vector embedding produces for a piece of content and compresses it into a single FDE vector.

This compression is the first step in MUVERA's two-step process. The search doesn't rely on the compressed FDE vector alone; it also uses an elimination method to find the best result.

First, it runs a very fast search on the FDEs to narrow the field down to a small group of the most likely candidates. Then it moves to a slower analysis, this time examining the original multi-vector sets. This stage, called "re-ranking," lets MUVERA combine the efficiency of fast systems with the accuracy that comes from richer data.

Part of what makes the FDE method so efficient and smart is that it treats queries and documents differently. MUVERA uses asymmetric encoding: it sums the vectors in a user's search query while averaging the vectors in the documents being searched. This lets it determine whether what a query is looking for actually exists in a document. Because FDEs don't depend on a specific dataset, this structure can also adapt to data that keeps changing and expanding.

Tired of reading?

You can also listen to this article as a podcast we made with Google NotebookLM, available on Spotify.

MUVERA in Statistics

  • Compared to the previous most advanced system, MUVERA answers search queries 90% faster on average, while improving accuracy by an average of 10%.

  • Compared to traditional embedding methods, MUVERA scans 5 to 20 times fewer candidate documents to reach the same accuracy rate.

  • The key to MUVERA's performance gain is its ability to compress by a factor of 32 without a significant drop in search quality, which lets it increase its query capacity per second by 20 times.

  • With single-vector embedding, reaching an 80% accuracy rate required 300 unique document candidates. MUVERA reaches the same rate with 60 candidates, meaning that even in the most efficient scenario, it processes 5 times fewer candidates.

MUVERA is more than a technological advance in the digital world. It's an innovation that will fundamentally change how comfortably we access information, and it deserves a place in digital marketers' strategies. From now on, instead of asking, "What searches is my target audience performing?" professionals should ask, "What problem is my target audience trying to solve, and what information do they want to access?" In this new era, strategy should shift from a keyword focus to a context focus. To succeed, put user intent at the center of your strategy, and cover each topic in full detail and from multiple perspectives.

Related service

Bring this question into your SEO plan

If this article points to a problem on your site, our SEO consultancy can help you decide what to check first and who should own the next step.
Explore SEO services
Ezgi Gülsen Yaylı
Ezgi Gülsen Yaylı