Semantic & Hybrid Search Development
Hybrid search should earn each ranking signal query by query, with exact terms, semantic matches, filters, and reranking traced to reviewer judgments.
Some queries depend on an exact product code. Others use language the source never does. We combine lexical and semantic retrieval with filters and reranking, then judge the result query by query.
You receive a working retrieval service plus the query inventory, calibration findings, and accepted trade-off note needed for later relevance reviews.


Some of the 500+ brands we've worked with
See all referencesSteps, gates, and who decides
How we work
We start with the queries and judgments your reviewers can stand behind. Those cases guide the retrieval design, the instrumentation, and every calibration decision.
Collect information needs
Reviewers bring real queries, expected sources, user context, filters, and examples of current failure. Together they form the evaluation set for the search decision.
- AI assist
- The material is sorted into candidate queries and filter patterns for review.
- Human gate
- Does the query set cover the important vocabulary, filters, and failure slices? Your search owner confirms the query set covers the priority needs.


Design the retrieval mix
We compare the retrieval options on the agreed query set and make the order of signals explicit. Filters and exceptions are treated as ranking rules with observable effects.
- AI assist
- Scores candidate retrieval mixes against the collected query set for a first ranking.
- Human gate
- Is the selected design the least complex option that covers the agreed needs? Your search owner picks the retrieval mix worth building.


Build and instrument search
The service records how candidates were found, filtered, scored, and reranked. A reviewer can follow a missing or surprising result back to the stage that changed it.
- AI assist
- Flags queries where the trace does not explain why a result ranked as it did.
- Human gate
- Can reviewers trace a result through every retrieval and ranking stage? Before calibration, your engineering lead signs off on the retained trace fields.


Calibrate relevance and handoff
Ordinary and difficult queries go through the same review. We tune only where the evidence supports a change, then document weak query slices, accepted trade-offs, and who keeps the evaluation set current.
- AI assist
- Generates adversarial and edge-case queries to widen the calibration set.
- Human gate
- Do critical query slices meet the agreed relevance and exception criteria? The person who owns relevance here accepts the trade-offs before handoff.


Named artifacts you keep
What you get
The service is handed over with the query judgments and traces behind its ranking choices, giving the search owner a practical basis for later recalibration.


Dashboard
Tested hybrid retrieval service and relevance report
The working lexical, semantic, filtered, and reranked retrieval path with results for representative information needs.


Risk register
Query assumptions and source-dependency inventory
The query set, relevance assumptions, source and metadata dependencies, known limitations, and unresolved exceptions.


Test evidence
Retrieval calibration and exception report
Relevance findings for ordinary and difficult queries, including lexical, semantic, filter, and reranking failure slices.


Decision record
Accepted ranking trade-offs and review note
The accepted retrieval design, relevance trade-offs, open conditions, owners, and next review point.
Scope and honest limits
When to bring us in
This work fits when exact matches, semantic similarity, metadata filters, and ranking priorities solve different parts of the same search problem.
A good fit when
- Your search team has real queries and expected sources, but the evidence is scattered across reviewers, filters, and source systems.
- Your exact product codes are easy to find while semantically similar language disappears, and another important query slice shows the opposite problem.
- Relevance reviewers disagree on the trade-offs, so nobody can say which weak query slices are acceptable or who keeps the evaluation set current.
- Lexical and semantic retrieval solve different queries, but filters and reranking have not been tested against the same information needs.
- Candidate retrieval designs look plausible, yet their dependencies, failure cases, and critical query slices have not been compared side by side.
- Reviewers find surprising results, but no record names the exception decision, its owner, or what the handoff carries.
- A retrieval design has been selected, but its decision gates and query evidence do not show why that option won or when it needs review.
Better handled as other work when
- You need the search system or its content cleared by legal, regulatory, audit, or certification reviewers. That decision stays with those reviewers.
- One ranking design must serve every query, user, and future corpus change. This build is calibrated to the tested query set and current content.
- You need the corpus repaired, the application redesigned, or the search service run in production. Those tasks sit beyond the agreed retrieval boundary.
If one of these is closer to your situation, start here instead: View the parent service
Engineers who ship production AI
This is the part of Zeo that writes and ships code. Our senior engineers build agents, chatbots, and RAG pipelines, along with the automation and data work around them, and they keep operating those systems once they're live. We've worked with more than 500 brands since 2011.
Tools we use
Tools behind this work
Coherethe reranker applied after lexical and semantic candidates are already combined
Qdrantthe single engine combining lexical and semantic retrieval with native score fusion
Voyage AIthe dense embedding covering a query phrased in language the source never used
Sentence Transformersthe self-hosted encoder keeping dense-side query latency off an external API
Ragasthe query-by-query relevance score matching this page's own calibration method
Next step
Calibrate search with the queries people actually use


Before you decide
























