← Corpus / lossless-monorepo / agent-skill
lossless-monorepo/agent-skills/chroma-agent-skills/src/chroma-cloud/templates/search-api
An expressive and flexible API for doing dense and sparse vector search on collections, as well as hybrid search
- Path
- agent-skills/chroma-agent-skills/src/chroma-cloud/templates/search-api.md
Search() API
The Search API provides a fluent, composable interface for building complex queries. It’s more expressive than the basic query method and supports advanced features like hybrid search with rank fusion.
Note: The Search API is only available on Chroma Cloud and is designed to work with Collection Schemas.
The Search() API acts as both query() and get() because the search expression that is being passed in ultimately decides what type of query to issue.
When to use Search() vs query()
Use query() when:
- You need simple semantic search
- You want the most straightforward API
Use search() when:
- You need hybrid search combining dense and sparse indexes
- You want fine-grained control over ranking and filtering
- You’re building complex queries with multiple conditions
- You need to select specific fields to return
Note that the Search() class uses a builder pattern, so if you call a method on it, it does not mutate that instance, it returns a copy with that mutation, so it needs re-assignging to the variable that is referencing it.
The search() method on a collection is able to take a single Search class instance or an arry of them, so the return value of the search() method on a collection is a SearchResult class, which has a rows() method, which will give you an array of array of results. So index 0 of the return value of rows() will be the array of the first Search class instance results.
Ranking and Scoring
Ranking expressions score and order results. Lower scores = better matches (distance-based). When rank is omitted, results are returned in index/insertion order.
Document inclusion rules when combining multiple Knn expressions:
- A document must appear in at least one
Knn’s top-limitresults to be scored - It must also appear in every
Knnwheredefault=None(the default); otherwise it’s excluded - Set a
defaultvalue on aKnnto assign a fallback score for documents missing from its results, keeping them in the pool. This is usually what you want when combining multipleKnnexpressions, otherwise the intersection rule often filters out too many candidates.
Expressions support arithmetic (+, -, *, /, unary -) and math functions (exp, log, abs, min, max) for combining and transforming scores. Numbers are auto-converted to constants, or use Val(x) explicitly. A common pattern is a weighted sum like Knn(query=q) * 0.7 + Knn(query=q, key="sparse_embedding") * 0.3 — but note this mixes two raw distance spaces, which may have very different scales. For robust hybrid ranking across dense + sparse, prefer RRF (below).
Two limits to keep straight:
Knn(limit=N)— how many candidates thatKnnscores (default 16). Raise it for better recall at the cost of latency.Search.limit(N)— how many results are returned after ranking.
For rank fusion (RRF), pass return_rank=True (Python) / returnRank: true (TypeScript) on each Knn so it emits rank positions (0, 1, 2…) instead of raw distances — see the Hybrid Search section below.
Setup
{{CODE:imports}}
Filtering with Key (K)
The Key class (aliased as K for brevity) provides a fluent interface for building filter expressions. Think of it like a query builder for metadata, document content, and IDs.
{{CODE:k}}
Ranking with Knn
Knn (k-nearest neighbors) is how you specify which embeddings to search and how to score results. Each Knn finds the nearest neighbors for a given query in a specific index.
The limit parameter controls how many candidates each Knn considers. A higher limit means more candidates are scored, which can improve recall but increases latency.
{{CODE:knn}}
Basic search example
Here’s a complete example showing the typical flow: create a collection, add documents, and search.
{{CODE:base-example}}
Hybrid search with Reciprocal Rank Fusion (RRF)
Hybrid search combines results from multiple indexes (typically dense + sparse) to get better results than either alone. RRF is a rank fusion algorithm that merges ranked lists without needing score normalization.
How RRF works:
- Each
Knnproduces a ranked list of candidates - Documents are scored based on their rank position in each list:
1 / (k + rank) - Scores are weighted and summed across all lists
- Final results are sorted by combined score
The k parameter (default 60) controls how much weight top-ranked documents get relative to lower-ranked ones. Higher k values make rankings more uniform.
{{CODE:rrf}}
Building effective hybrid search
For best results with hybrid search:
- Use comparable limits for each
Knnso both indexes contribute meaningfully - Weight based on your data: keyword-heavy content might favor sparse; conceptual content might favor dense
- Start with 0.7/0.3 weighting (dense/sparse) and adjust based on evaluation
- Use
returnRank: truewhen combining with RRF, as RRF operates on ranks, not distances
Note that return ranks from RRF are netagive and the value furthest from 0 is the closest to the original query.