Qdrant published version 1.19.2 on October 5, 2026, with fixes for sparse-vector ingestion, filtered search and storage recovery. For teams running retrieval-augmented generation (RAG), the release warrants a staged upgrade evaluation—especially where writes, tenant filters or schema changes exercise the documented failure paths.
For RAG teams, filtered retrieval and restart recovery need separate acceptance checks. Prioritize correct results and recoverable data before throughput gains. The checklist below is a proposed validation plan, not a report of tests performed by RohitAI. Availability in every Qdrant Cloud deployment and across all container architectures has not been established.
Before changing the server version
Check the starting version and clients. Qdrant’s upgrade guidance requires the latest patch of each intermediate minor version, including for single-node deployments, and recommends updating compatible SDKs first. For deployments entering 1.19, migrate legacy search, recommend and discover endpoints to the unified Query API; that removal belongs to the 1.19 minor release, not this patch.
Preserve a usable recovery source. Collection snapshots are per-node and exclude aliases; full-storage snapshots are for single-node recovery. Rehearse restoring into an isolated compatible target and verify point IDs, payloads and representative queries. The snapshot documentation supports restoration within the same minor at an equal-or-newer patch, or into the next minor—not an assumed downgrade to an older binary. Retain the pre-upgrade backup and any inputs needed to replay subsequent writes.
Hold other variables steady. Record the server artifact, configuration, collection schema, replication, optimizer settings and effective feature flags. Keep embeddings and queries unchanged for the database comparison. Changing encoders is a separate decision; the Embed 5 re-indexing guide covers that evaluation.
Sparse ingestion: compare the same workload
The sparse-upsert PR reduces repeated posting-list work. Its contributor reports these REST upload times for one million SPLADE vectors with 126 nonzeros, one 16-core node, 1,000-point batches and wait=true. These are upstream measurements; the baseline is described as dev, not a reproduced 1.19.1 comparison.
Collection settings | Before | After | Calculated time ratio |
|---|---|---|---|
Defaults | 72.2 seconds | 23.7 seconds | 3.0× |
indexing_threshold: 0 | 839.8 seconds | 29.5 seconds | 28.5× |
Ratios are before-time divided by after-time. After the fix, disabling indexing was still about 24% slower in this setup: 29.5 / 23.7 − 1. The larger improvement does not make that setting the preferred configuration. The PR’s separate roughly 457× result is an in-memory index microbenchmark, not REST or RAG throughput.
Proposed check: Replay append, overwrite, shuffled-insert and delete workloads on equivalent deployments. Fix hardware, batches, concurrency and optimizer settings. Record points per second, write p95, query p95 during ingestion and peak memory. Include common-term posting lists; do not adopt an upstream multiplier as your acceptance threshold.
Filtered search: check known answers, not just latency
The filtered-HNSW fix supplies matching seed points from the payload index when no graph entry point passes the filter. The documented case uses m=0, payload_m=16 and a tenant index, then narrows the query with a product filter. Matching records could exist despite an empty response; this is not a claim that every filtered query was affected.
Proposed check: Use tenant-only, tenant-plus-attribute, rare-match and true no-answer fixtures. Compare returned IDs and Recall@k with an exact or known reference set, and verify that every result satisfies its filter. Keep thresholds and quantization fixed. Record latency separately: reaching valid candidates may cost more than incorrectly returning nothing. Investigate unexpected empty results before changing the generator prompt or lowering similarity thresholds.
Multivectors: exercise an existing collection
The named-multivector fix addresses empty placeholders when a vector field is added to an already indexed collection. A fresh collection with every multivector populated misses that scenario.
Proposed check: In a disposable copy, add the intended named multivector, leave some older points without it, populate a subset and allow optimizer merges. Exercise both the original vector search and the new named path, then restart. Check optimizer errors and results using the production datatype; upstream regression coverage names f32, u8 and f16.
Recovery: distinguish configuration, durability and readiness
Proxy persistence lets Qdrant save pending segment changes during flushes and replay them after restart, allowing the write-ahead log (WAL) to advance while proxies exist. But the implementation is feature-gated: persist_proxy_segments is false in the v1.19.2 defaults. Deployment modes can enable it. Inspect effective configuration rather than assuming the version alone delivers shorter recovery; do not enable every experimental flag to obtain one feature.
Other fixes are separate: flush accounting stops claiming an operation is persisted before it is fully applied, while proxy change propagation follows operation-version order. Recovery validation therefore needs to cover record contents and schema/index changes, not just process startup time.
The Gridstore mapping fix logs mapping changes with length and CRC32 protection before updating the mapping file, then replays valid entries after restart. A payload-index recovery fix allows malformed derived indexes to be regenerated. Neither establishes that upgrading reconstructs previously lost source payloads or protects against every storage failure.
Proposed check: With representative writes and background optimization, measure time to readiness and time to correct queries separately. Compare completed writes’ IDs and payloads against the operation log, inspect retained WAL size, and exercise payload-index and named-vector changes. Record whether proxy persistence is active. Any abrupt-stop or storage-fault exercise belongs in a disposable test environment; a normal restart does not validate torn-write protection.
Distributed deployments: verify peer access and convergence
When service.enforce_internal_auth is enabled, the internal gRPC validator accepts the read-write api_key and alt_api_key, not read-only keys or JWTs. This change leaves public API authentication unchanged. Enforcement remains opt-in; installing the patch does not enable it automatically.
Check credential types and peer connectivity without logging key values. If rotation uses an alternate-key-only stage, verify the outgoing-request fallback to alt_api_key. Match auth sequencing to the starting peer versions; do not disable existing enforcement as a generic upgrade step.
The new cluster.p2p.host setting can bind internal gRPC separately from the public service. Leaving it unset preserves the prior binding. If using it, select a reachable interface IP, verify advertised peer addresses and restrict the internal port to intended peers. A new bind option or startup warning is not evidence of network isolation.
For recovery, test peer convergence and unrelated collection availability. The release separately fixes stale snapshot transfers resumed by recovering peers and holding the collections lock during Raft snapshot reconciliation. These are distributed-state operations, not ordinary collection backup archives.
Snapshot storage: test the real restore path
Collection snapshots gain Google Cloud Storage and Azure Blob Storage backends, alongside local and S3 storage. Optional object-key prefixes leave the old layout unchanged when omitted; a prefix is not an access-control policy and does not move existing backups.
If adopting a backend or prefix, verify creation, listing, download and restore using the deployment’s actual identity permissions. Inspect object placement and access to older snapshots. The backend PR reports emulator coverage and explicitly notes that real-bucket GCS uploads were not exercised there, so support alone is not a verified production recovery result.
The rollout decision
Define acceptable result correctness, error rates, query tail latency and recovery time before the canary. Expand only after the relevant checks pass. Missing records, unexplained empty answers, authentication failures or stalled replica convergence should stop expansion and trigger the prepared recovery plan—not an improvised binary downgrade.
For rolling restarts, Qdrant’s availability guidance requires every collection to have replication factor at least two. Allow replication and transfers to settle between nodes. Single-node or replication-factor-one deployments need a restart interruption. For a dense-only, static workload, sparse throughput is not the deciding metric; the applicable storage, retrieval and recovery checks still are.
Methodology: AI-assisted reporting and analysis based on Qdrant’s release notes, merged PRs, version-pinned source and official documentation checked on October 5, 2026. Benchmark numbers are contributor-reported; ratios above are arithmetic from those figures. RohitAI did not run Qdrant, reproduce the benchmarks or perform the proposed restore and failure tests.
