Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions _includes/configuration/rq-compression-parameters.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -2,5 +2,7 @@
| :---------------------- | :------ | :------ | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `rq`: `bits` | integer | `8` | The number of bits used to quantize each data point. Value can be `8`, `4` or `1`, but not every index type accepts all three. The `hnsw` index type accepts `8`, `4` and `1`. The `flat` index type accepts only `8` and `1`. The `hfresh` index type accepts only `1`. <br/><br/>This parameter is fixed once RQ is enabled and cannot be changed afterwards. <br/> <br/>Learn more about [8-bit](/weaviate/concepts/vector-quantization#8-bit-rq), [4-bit](/weaviate/concepts/vector-quantization#4-bit-rq) and [1-bit](/weaviate/concepts/vector-quantization#1-bit-rq) RQ. |
| `rq`: `rescoreLimit` | integer | `20` (`hnsw`, 8-bit and 4-bit)<br/>`512` (`hnsw`, 1-bit)<br/>`-1` (`flat`) | The minimum number of candidates to fetch before rescoring. Mutable at any time. <br/><br/>The default depends on the vector index type, and under `hnsw` also on `bits`: `20` for 8-bit and 4-bit RQ, and `512` for 1-bit RQ. Under the `flat` index type the default is `-1`, which lets Weaviate pick the limit. <br/><br/>The Java client sends this parameter under a field name that Weaviate does not read, so values set from that client are ignored and the server default applies. <br/><br/>These defaults apply to the `hnsw` and `flat` index types. For the HFresh index, see [HFresh index parameters](/weaviate/config-refs/indexing/vector-index#hfresh-index-parameters). |
| `rq`: `centering` | boolean | `false` | Centers the vectors on their mean before 4-bit RQ compression. Only valid on the `hnsw` index with `bits` set to `4`. Not supported for multi-vector embeddings without MUVERA. <br/><br/>Weaviate computes the mean from a sample of up to `trainingLimit` vectors. Compression starts once a shard holds more than `trainingLimit` vectors. <br/><br/>Cannot be changed once RQ is enabled. <br/><br/>Added in `v1.39.3` |
| `rq`: `trainingLimit` | integer | `10000` | The maximum number of vectors per shard sampled to compute the mean for `centering`. Must be greater than `0` when `centering` is `true`. Has no effect without `centering`. <br/><br/>Added in `v1.39.3` |
| `rq` : `cache` | boolean | `false` | Whether to cache the vectors in memory.<br/> (only when using the `flat` vector index type) |
| `vectorCacheMaxObjects` | integer | `1e12` | Maximum number of objects in the memory cache. By default, this limit is set to one trillion (`1e12`) objects when a new collection is created. For sizing recommendations, see [Vector cache considerations](/weaviate/concepts/vector-index#vector-cache-considerations). |
2 changes: 2 additions & 0 deletions _includes/feature-notes/hfresh_multivector.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
:::info Added in `v1.40`
:::
2 changes: 2 additions & 0 deletions _includes/feature-notes/meta.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
:::info Added in `v1.39.3`
:::
5 changes: 1 addition & 4 deletions _includes/feature-notes/rq-4bit.mdx
Original file line number Diff line number Diff line change
@@ -1,5 +1,2 @@
:::caution Preview — added in `v1.39.0`

**4-bit Rotational quantization (RQ)** for the **HNSW vector index** was added in **`v1.39.0`** as a preview feature. The API may change in future releases.

:::info Added in `v1.40`
:::
2 changes: 1 addition & 1 deletion _includes/multi-vector-compress.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
:::info Added in `v1.30`
:::

Multi-vector embeddings (implemented through models like ColBERT, ColPali, or ColQwen) represent each object or query using multiple vectors instead of a single vector. Just like with single vectors, multi-vectors support [PQ](/weaviate/configuration/compression/pq-compression), [BQ](/weaviate/configuration/compression/bq-compression), [RQ](/weaviate/configuration/compression/rq-compression), [SQ](/weaviate/configuration/compression/sq-compression), or no compression.
Multi-vector embeddings (implemented through models like ColBERT, ColPali, or ColQwen) represent each object or query using multiple vectors instead of a single vector. On an `hnsw` index, multi-vectors support [PQ](/weaviate/configuration/compression/pq-compression), [BQ](/weaviate/configuration/compression/bq-compression), [RQ](/weaviate/configuration/compression/rq-compression), [SQ](/weaviate/configuration/compression/sq-compression), or no compression. On an `hfresh` index (from `v1.40`), the MUVERA encoding is always compressed with 1-bit RQ.

During the initial search phase, compressed vectors are used for efficiency. However, when computing the `MaxSim` operation, uncompressed vectors are utilized to ensure more precise similarity calculations. This approach balances the benefits of compression for search efficiency with the accuracy of uncompressed vectors during final scoring.
Loading
Loading