Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions _includes/feature-notes/hfresh_multivector.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
:::info Added in `v1.40`
:::
11 changes: 11 additions & 0 deletions docs/weaviate/concepts/indexing/vector-index.md
Original file line number Diff line number Diff line change
Expand Up @@ -272,6 +272,16 @@ HFresh only supports `cosine` and `l2-squared` distance metrics. Dot product is

For configuration details, see the [HFresh index parameters](../../config-refs/indexing/vector-index.mdx#hfresh-index-parameters).

### Multi-vector embeddings on HFresh

import HFreshMultiVector from '/_includes/feature-notes/hfresh_multivector.mdx';

<HFreshMultiVector />

HFresh can index [multi-vector embeddings](../../configuration/compression/multi-vectors.md), such as ColBERT or ColPali representations, but only with MUVERA encoding. On HNSW, MUVERA is optional but on HFresh it is what makes multi-vector search work, so you have to turn it on when you create the collection. For how the encoding itself works, and how candidates are scored exactly with MaxSim, see [MUVERA encoding](../../configuration/compression/multi-vectors.md#muvera-encoding).

For the configuration parameters, see [Multi-vector embeddings on HFresh](../../config-refs/indexing/vector-index.mdx#hfresh-multi-vector).

## Vector cache considerations

For optimal search and import performance, previously imported vectors need to be in memory. A disk lookup for a vector is orders of magnitudes slower than memory lookup, so the disk cache should be used sparingly. However, Weaviate can limit the number of vectors in memory. By default, this limit is set to one trillion (`1e12`) objects when a new collection is created.
Expand Down Expand Up @@ -333,6 +343,7 @@ Here's a quick guide to choosing the right index:

| Feature | Flat | HNSW | HFresh |
| ----------------------------- | -------------------------------- | --------------------------- | --------------------------------------------------- |
| Multi-vector support | No | Yes, MUVERA optional | Yes, MUVERA required |
| Memory usage | Very low | High | Low |
| Search speed (small datasets) | Fast | Very fast | Moderate |
| Search speed (large datasets) | Slow | Very fast | Fast |
Expand Down
23 changes: 23 additions & 0 deletions docs/weaviate/config-refs/indexing/vector-index.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -191,11 +191,34 @@ HFresh only supports `cosine` and `l2-squared` distance metrics. Dot product is
| `replicas` | integer | `4` | No | Number of posting lists in which a vector is added. Min: `1`, Max: `10`. |
| `searchProbe` | integer | `256` | Yes | Number of posting lists to search during a query. The default is `256` in `v1.36.20`, `v1.37.10`, `v1.38.2` and later. Earlier releases on each of those lines default to `64`. |
| `rq` | object | -- | Partial | Rotational quantization (RQ) compression configuration. RQ is mandatory for HFresh and cannot be turned off. Its `bits` value is fixed at `1`; a request that sets a wider width is rejected. Its `rescoreLimit` (default `350`), the number of candidates rescored against uncompressed vectors, is mutable at runtime. |
| `multivector` | object | -- | No | Multi-vector configuration. On HFresh, multi-vector embeddings require MUVERA encoding. See [Multi-vector embeddings on HFresh](#hfresh-multi-vector).<br/><br/>Added in `v1.40` |

:::tip Tuning HFresh recall
Start with the defaults. If recall is too low, increase `searchProbe` (search more posting lists per query) or the RQ `rescoreLimit` (rescore more candidates with full-precision vectors). Both are mutable at runtime and take effect **without reindexing**.
:::

### Multi-vector embeddings on HFresh {#hfresh-multi-vector}

import HFreshMultiVector from "/_includes/feature-notes/hfresh_multivector.mdx";

<HFreshMultiVector />

HFresh supports multi-vector embeddings only with [MUVERA encoding](../../configuration/compression/multi-vectors.md#muvera-encoding), so `multivector.muvera.enabled` is required.

:::caution On HFresh, `multivector.enabled` alone has no effect
The collection is created, but it behaves as a single-vector collection and rejects multi-vector imports and queries. Set `multivector.muvera.enabled`. Multi-vector search on HFresh also needs a named vector with `"vectorizer": "none"`.
:::

| Parameter | Type | Default | Mutable | Details |
| :--------------------------- | :------ | :------ | :------ | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `multivector.muvera.enabled` | boolean | `false` | No | Enables MUVERA encoding, which is required for multi-vector search on HFresh. Setting it to `true` also turns on `multivector.enabled`. Cannot be changed after the collection is created. |

For what `ksim`, `dprojections` and `repetitions` mean and what they default to, see [MUVERA encoding](../../configuration/compression/multi-vectors.md#muvera-encoding).

On HFresh these three have upper limits, checked whenever you create or update a collection: `ksim` at most `10`, `dprojections` at most `1024`, `repetitions` at most `256`, and at most 1,048,576 float32 values (4 MiB per vector) for the encoding they produce together so the three maximums cannot be combined. Set the values when you create the collection, changing them later is accepted but has no effect.

On a multi-vector index, `searchProbe` and `rq.rescoreLimit` set the [routing and rescore budgets](../../concepts/indexing/vector-index.md#multi-vector-embeddings-on-hfresh), `max(limit, searchProbe)` and `max(limit, rq.rescoreLimit)`. Because both include the query limit, neither parameter can reduce the number of results returned.

## Quantization parameters

### RQ parameters
Expand Down
6 changes: 5 additions & 1 deletion docs/weaviate/configuration/compression/multi-vectors.md
Original file line number Diff line number Diff line change
Expand Up @@ -101,8 +101,12 @@ These parameters can be used to fine-tune MUVERA:
the dimensionality of the final encoding but can lead to better approximation
of the original multi-vector similarity.

:::note MUVERA is required on the HFresh index
MUVERA is optional on the `hnsw` index. The `hfresh` index supports multi-vector embeddings only with MUVERA encoding, so there you always have to set `multivector.muvera.enabled`. See [Multi-vector embeddings on HFresh](../../config-refs/indexing/vector-index.mdx#hfresh-multi-vector). Added in `v1.40`.
:::

:::note Quantization
Quantization is also available as a compression technique for multi-vector embeddings. It reduces the memory footprint of individual vectors by approximating their values with less precision. Just like with single vectors, multi-vectors support [PQ](./pq-compression.md), [BQ](./bq-compression.md), [RQ](./rq-compression.md) and [SQ](./sq-compression.md) quantization.
Quantization is also available as a compression technique for multi-vector embeddings. It reduces the memory footprint of individual vectors by approximating their values with less precision. Just like with single vectors, multi-vectors on an `hnsw` index support [PQ](./pq-compression.md), [BQ](./bq-compression.md), [RQ](./rq-compression.md) and [SQ](./sq-compression.md) quantization.
:::

## Further resources
Expand Down
2 changes: 1 addition & 1 deletion docs/weaviate/manage-collections/vector-config.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -236,7 +236,7 @@ Adding a new vector to the collection definition [won't trigger vectorization fo

<MultiVector/>

Multi-vector embeddings, also known as multi-vectors, represent a single object with multiple vectors, i.e. a 2-dimensional matrix. Multi-vectors are currently only available for HNSW indexes for named vectors. To use multi-vectors, enable it for the appropriate named vector.
Multi-vector embeddings, also known as multi-vectors, represent a single object with multiple vectors, i.e. a 2-dimensional matrix. Multi-vectors are available for `hnsw` and, since `v1.40`, `hfresh` indexes ([with MUVERA encoding only](../config-refs/indexing/vector-index.mdx#hfresh-multi-vector)), for named vectors. To use multi-vectors, enable it for the appropriate named vector.

<Tabs className="code" groupId="languages">
<TabItem value="py" label="Python">
Expand Down
Loading