Skip to content
AI News HubLIVE
Public articles 18Collected articles 29Trust 82Refresh 120 min
Health HealthySource type OfficialFull-text rights Official full textLast ingested 2026-09-17ID qdrant-blogStatus Enabled

Official vector database and AI infrastructure feed; confirm reuse terms before full body display.

Latest public articles

Hyperbolic Embeddings in Qdrant

We choose embedding models, dimensions, and indexes. The geometry usually comes with the package. But why use a flat space, and what else could we choose? What Is a Manifold, and Where Do Our Vectors Live? A manifold is the space our embeddings live in. For embeddings, we care about the geometry we give that space. It determines how we measure distance, what the shortest path looks like, and how much room there is as we move outward.

Qdrant BlogIn-site articleHyperbolic Embeddings in Qdrant

Hybrid Search in Qdrant

A search result can look plausible and still be wrong. Dense retrieval can return a document on the right topic but miss an exact identifier copied into the query. Sparse retrieval can miss a relevant document when the query describes it with terms the corpus doesn’t use. Either way, your logs record a successful query. Hybrid search runs dense and sparse retrieval over the same query, then merges their result lists. Dense retrieval adds semantic similarity, so paraphrases can rank together. Sparse retrieval adds weighted term matching for exact words and identifiers.

Qdrant BlogIn-site articleHybrid Search in Qdrant

When Your Collection Outgrows RAM

Once a collection no longer fits in RAM, the kernel evicts vector pages, and the next query waits on a disk read to get them back. Quantization buys that memory back. Qdrant keeps a compressed copy of each dense vector in RAM and moves the full-precision originals to disk. TurboQuant is the method measured here. It rotates each vector before compressing it, which spreads the error evenly across coordinates, and its bits parameter sets the depth from bits4 down to bits1. Start at bits4, a good default for many workloads at eight times compression.

Qdrant BlogIn-site articleWhen Your Collection Outgrows RAM

When Is a Reranker Worth It?

Before you tune a reranker, use the pre-tuning checks to verify index state and set a labeled baseline. Your candidate list can already contain documents your ranking never shows. Score those candidates as if they were perfectly ordered, then compare that with the score your pipeline returns today. The gap between the two is everything a better ranking stage could recover, so measure it before you reach for a model. Use nDCG@10, which grades the top 10 results and gives more credit to relevant documents near the top.

Qdrant BlogIn-site articleWhen Is a Reranker Worth It?

How to Tune Hybrid Search in Qdrant

Before you tune fusion, use the pre-tuning checks to verify index state and set a labeled baseline. Hybrid search retrieves dense and sparse candidate lists, then fuses them into one ranking. The dense prefetch finds similar meaning; the sparse prefetch finds matching keywords. Fusion reorders the candidates the prefetches return, so a document missing from both lists cannot appear in the result. Confirm Fusion Beats Either Prefetch Before tuning, compare dense retrieval, sparse retrieval, and default Reciprocal Rank Fusion (RRF) at k=2 and equal weights. Score all three with nDCG@10, which grades the top 10 results and gives more credit to relevant documents near the top.

Qdrant BlogIn-site articleHow to Tune Hybrid Search in Qdrant

Candidate Depth: How Much Retrieval Is Enough?

Before you tune candidate depth, use the pre-tuning checks to verify index state and set a labeled baseline. Everything below measures against that baseline. Candidate depth is the number of candidates a retrieval stage passes to a later ranking stage. It matters only when a later stage can use the extra candidates. In hybrid search, every prefetch carries its own limit, and a multi-stage query that nests one prefetch inside another sets a depth at each level. In dense-only or sparse-only search, it is the number of candidates you pass to a reranker or other downstream stage.

Qdrant BlogIn-site articleCandidate Depth: How Much Retrieval Is Enough?

What to Check Before Tuning a Qdrant Collection

Before you change a setting, decide what better retrieval means for your workload. The right document at rank one, more candidates for a reranker, lower latency, and a smaller memory footprint each favor different settings, so pick your goal first. If your labeled queries can’t detect the improvement you’re chasing, you won’t be able to tell whether a change helped. Some settings are there to verify correctness, not to tune performance. If a vector is unindexed, a sparse vector is missing the IDF modifier, or the BM25 average length is wrong, the results are invalid. Any benchmark or comparison you run after that will reflect a broken setup. This article shows you how to check each setting and what the correct state looks like.

Qdrant BlogIn-site articleWhat to Check Before Tuning a Qdrant Collection

Filtered Vector Search: What ACORN Fixes, and What Fixes ACORN

Filtered vector search breaks when metadata filters turn a healthy nearest-neighbor graph into scattered islands. HNSW’s m parameter controls how many links each point gets. At Qdrant’s default m=16, the one-million-point collection benchmarked below averaged about 21 links per node on layer 0. Filter out 96% of the points and fewer than one link per node survives on average, so traversal can get stranded before it reaches the true nearest matches.

Qdrant BlogIn-site articleFiltered Vector Search: What ACORN Fixes, and What Fixes ACORN

Predicting Weak Retrieval Without an LLM

Most retrieval systems use a single pipeline for all queries, which either under-serves hard queries or wastes compute on easy ones. This article presents cheap signals—like score spread and retriever agreement—to detect weak retrieval without an LLM, enabling selective escalation only when needed.

Qdrant BlogIn-site articlePredicting Weak Retrieval Without an LLM

TurboQuant in Qdrant

Qdrant 1.18 ships TurboQuant, a new rotation-based vector quantization method from Google Research, with extensions for production embeddings. It offers 4-bit, 2-bit, 1.5-bit, and 1-bit options, outperforming or matching Scalar Quantization (SQ) and Binary Quantization (BQ) in compression and recall. The article explains the algorithm, Qdrant's enhancements (length renormalization and per-coordinate calibration), and benchmark results.

Qdrant BlogIn-site articleTurboQuant in Qdrant

Fine-Tuning Sparse Embeddings for E-Commerce Search | Part 2: Training SPLADE on Modal

This is Part 2 of a 5-part series on fine-tuning sparse embeddings for e-commerce search. It covers training a SPLADE model on Modal's serverless GPUs using the Amazon ESCI dataset, including data loading, product text formatting, Modal setup, model creation, training function, SpladeLoss, YAML configuration, parallel hyperparameter sweeps, and a pitfall to avoid.

Qdrant BlogIn-site articleFine-Tuning Sparse Embeddings for E-Commerce Search | Part 2: Training SPLADE on Modal

Fine-Tuning Sparse Embeddings for E-Commerce Search | Part 1: Why Sparse Embeddings Beat BM25

This article is Part 1 of a series on fine-tuning sparse embeddings for e-commerce search. It explains why dense embeddings fail in product search due to blurred exact matches, and how sparse embeddings preserve critical details. It introduces the SPLADE model, query expansion, and Qdrant's native sparse vector support. The fine-tuned system achieves a 29% improvement over BM25 on the Amazon ESCI dataset.

Qdrant BlogIn-site articleFine-Tuning Sparse Embeddings for E-Commerce Search | Part 1: Why Sparse Embeddings Beat BM25

Distance-based data exploration

This article explores how Qdrant's Distance Matrix API enables efficient data exploration through dimensionality reduction, clustering, and graph-based visualization, helping to uncover hidden structures in large unstructured datasets.

Qdrant BlogIn-site articleDistance-based data exploration

Qdrant Summer of Code 2024 - ONNX Cross Encoders in Python

In this article, Huong (Celine) Hoang shares her experience integrating ONNX cross-encoders into the FastEmbed library during Qdrant's Summer of Code 2024. The project enables re-ranking search results using relevance scores, enhancing context-aware search applications. Key challenges included building a new input-output scheme, tokenization, model loading, and testing. The functionality is available in FastEmbed 0.4.0.

Qdrant BlogIn-site articleQdrant Summer of Code 2024 - ONNX Cross Encoders in Python

What is a Vector Database?

An introduction to vector databases, covering how they represent unstructured data as vectors, their architecture (collections, distance metrics, storage), core operations (indexing, searching, updating, deleting), advanced features (dense/sparse vectors, hybrid search, quantization, sharding, replication, multitenancy, security), and practical use cases.

Qdrant BlogIn-site articleWhat is a Vector Database?

What is Vector Quantization?

Vector quantization compresses high-dimensional vectors to reduce memory usage and speed up search operations in large datasets. This article covers scalar, binary, and product quantization methods, along with techniques like oversampling, rescoring, and io_uring to balance accuracy and performance.

Qdrant BlogIn-site articleWhat is Vector Quantization?

Qdrant 1.7.0 has just landed!

Qdrant 1.7.0 introduces native support for sparse vectors, enabling keyword-based search and hybrid search; a new Discovery API for more precise vector search, including discovery search and context search; user-defined sharding for flexible data distribution; snapshot-based shard transfer for efficient cluster scaling; and various performance improvements.

Qdrant BlogIn-site articleQdrant 1.7.0 has just landed!

On Unstructured Data, Vector Databases, New AI Age, and Our Seed Round.

Qdrant announces $7.5M seed funding led by Unusual Ventures. The article discusses the importance of vector databases in the AI age, the explosive growth of unstructured data, and Qdrant's progress as an open-source vector similarity search solution.

Qdrant BlogIn-site articleOn Unstructured Data, Vector Databases, New AI Age, and Our Seed Round.

All sources