diff --git a/_includes/feature-notes/hfresh_multivector.mdx b/_includes/feature-notes/hfresh_multivector.mdx
new file mode 100644
index 000000000..ac311cde0
--- /dev/null
+++ b/_includes/feature-notes/hfresh_multivector.mdx
@@ -0,0 +1,2 @@
+:::info Added in `v1.40`
+:::
diff --git a/docs/weaviate/concepts/indexing/vector-index.md b/docs/weaviate/concepts/indexing/vector-index.md
index 7b78bec8d..02474a8b2 100644
--- a/docs/weaviate/concepts/indexing/vector-index.md
+++ b/docs/weaviate/concepts/indexing/vector-index.md
@@ -272,6 +272,16 @@ HFresh only supports `cosine` and `l2-squared` distance metrics. Dot product is
For configuration details, see the [HFresh index parameters](../../config-refs/indexing/vector-index.mdx#hfresh-index-parameters).
+### Multi-vector embeddings on HFresh
+
+import HFreshMultiVector from '/_includes/feature-notes/hfresh_multivector.mdx';
+
+
+
+HFresh can index [multi-vector embeddings](../../configuration/compression/multi-vectors.md), such as ColBERT or ColPali representations, but only with MUVERA encoding. On HNSW, MUVERA is optional but on HFresh it is what makes multi-vector search work, so you have to turn it on when you create the collection. For how the encoding itself works, and how candidates are scored exactly with MaxSim, see [MUVERA encoding](../../configuration/compression/multi-vectors.md#muvera-encoding).
+
+For the configuration parameters, see [Multi-vector embeddings on HFresh](../../config-refs/indexing/vector-index.mdx#hfresh-multi-vector).
+
## Vector cache considerations
For optimal search and import performance, previously imported vectors need to be in memory. A disk lookup for a vector is orders of magnitudes slower than memory lookup, so the disk cache should be used sparingly. However, Weaviate can limit the number of vectors in memory. By default, this limit is set to one trillion (`1e12`) objects when a new collection is created.
@@ -333,6 +343,7 @@ Here's a quick guide to choosing the right index:
| Feature | Flat | HNSW | HFresh |
| ----------------------------- | -------------------------------- | --------------------------- | --------------------------------------------------- |
+| Multi-vector support | No | Yes, MUVERA optional | Yes, MUVERA required |
| Memory usage | Very low | High | Low |
| Search speed (small datasets) | Fast | Very fast | Moderate |
| Search speed (large datasets) | Slow | Very fast | Fast |
diff --git a/docs/weaviate/config-refs/indexing/vector-index.mdx b/docs/weaviate/config-refs/indexing/vector-index.mdx
index c7a13cec8..fd2dfb3bb 100644
--- a/docs/weaviate/config-refs/indexing/vector-index.mdx
+++ b/docs/weaviate/config-refs/indexing/vector-index.mdx
@@ -191,11 +191,34 @@ HFresh only supports `cosine` and `l2-squared` distance metrics. Dot product is
| `replicas` | integer | `4` | No | Number of posting lists in which a vector is added. Min: `1`, Max: `10`. |
| `searchProbe` | integer | `256` | Yes | Number of posting lists to search during a query. The default is `256` in `v1.36.20`, `v1.37.10`, `v1.38.2` and later. Earlier releases on each of those lines default to `64`. |
| `rq` | object | -- | Partial | Rotational quantization (RQ) compression configuration. RQ is mandatory for HFresh and cannot be turned off. Its `bits` value is fixed at `1`; a request that sets a wider width is rejected. Its `rescoreLimit` (default `350`), the number of candidates rescored against uncompressed vectors, is mutable at runtime. |
+| `multivector` | object | -- | No | Multi-vector configuration. On HFresh, multi-vector embeddings require MUVERA encoding. See [Multi-vector embeddings on HFresh](#hfresh-multi-vector).
Added in `v1.40` |
:::tip Tuning HFresh recall
Start with the defaults. If recall is too low, increase `searchProbe` (search more posting lists per query) or the RQ `rescoreLimit` (rescore more candidates with full-precision vectors). Both are mutable at runtime and take effect **without reindexing**.
:::
+### Multi-vector embeddings on HFresh {#hfresh-multi-vector}
+
+import HFreshMultiVector from "/_includes/feature-notes/hfresh_multivector.mdx";
+
+
+
+HFresh supports multi-vector embeddings only with [MUVERA encoding](../../configuration/compression/multi-vectors.md#muvera-encoding), so `multivector.muvera.enabled` is required.
+
+:::caution On HFresh, `multivector.enabled` alone has no effect
+The collection is created, but it behaves as a single-vector collection and rejects multi-vector imports and queries. Set `multivector.muvera.enabled`. Multi-vector search on HFresh also needs a named vector with `"vectorizer": "none"`.
+:::
+
+| Parameter | Type | Default | Mutable | Details |
+| :--------------------------- | :------ | :------ | :------ | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
+| `multivector.muvera.enabled` | boolean | `false` | No | Enables MUVERA encoding, which is required for multi-vector search on HFresh. Setting it to `true` also turns on `multivector.enabled`. Cannot be changed after the collection is created. |
+
+For what `ksim`, `dprojections` and `repetitions` mean and what they default to, see [MUVERA encoding](../../configuration/compression/multi-vectors.md#muvera-encoding).
+
+On HFresh these three have upper limits, checked whenever you create or update a collection: `ksim` at most `10`, `dprojections` at most `1024`, `repetitions` at most `256`, and at most 1,048,576 float32 values (4 MiB per vector) for the encoding they produce together so the three maximums cannot be combined. Set the values when you create the collection, changing them later is accepted but has no effect.
+
+On a multi-vector index, `searchProbe` and `rq.rescoreLimit` set the [routing and rescore budgets](../../concepts/indexing/vector-index.md#multi-vector-embeddings-on-hfresh), `max(limit, searchProbe)` and `max(limit, rq.rescoreLimit)`. Because both include the query limit, neither parameter can reduce the number of results returned.
+
## Quantization parameters
### RQ parameters
diff --git a/docs/weaviate/configuration/compression/multi-vectors.md b/docs/weaviate/configuration/compression/multi-vectors.md
index 549cf4940..97fda93ed 100644
--- a/docs/weaviate/configuration/compression/multi-vectors.md
+++ b/docs/weaviate/configuration/compression/multi-vectors.md
@@ -101,8 +101,12 @@ These parameters can be used to fine-tune MUVERA:
the dimensionality of the final encoding but can lead to better approximation
of the original multi-vector similarity.
+:::note MUVERA is required on the HFresh index
+MUVERA is optional on the `hnsw` index. The `hfresh` index supports multi-vector embeddings only with MUVERA encoding, so there you always have to set `multivector.muvera.enabled`. See [Multi-vector embeddings on HFresh](../../config-refs/indexing/vector-index.mdx#hfresh-multi-vector). Added in `v1.40`.
+:::
+
:::note Quantization
-Quantization is also available as a compression technique for multi-vector embeddings. It reduces the memory footprint of individual vectors by approximating their values with less precision. Just like with single vectors, multi-vectors support [PQ](./pq-compression.md), [BQ](./bq-compression.md), [RQ](./rq-compression.md) and [SQ](./sq-compression.md) quantization.
+Quantization is also available as a compression technique for multi-vector embeddings. It reduces the memory footprint of individual vectors by approximating their values with less precision. Just like with single vectors, multi-vectors on an `hnsw` index support [PQ](./pq-compression.md), [BQ](./bq-compression.md), [RQ](./rq-compression.md) and [SQ](./sq-compression.md) quantization.
:::
## Further resources
diff --git a/docs/weaviate/manage-collections/vector-config.mdx b/docs/weaviate/manage-collections/vector-config.mdx
index b27bf712b..eeee8fb22 100644
--- a/docs/weaviate/manage-collections/vector-config.mdx
+++ b/docs/weaviate/manage-collections/vector-config.mdx
@@ -236,7 +236,7 @@ Adding a new vector to the collection definition [won't trigger vectorization fo
-Multi-vector embeddings, also known as multi-vectors, represent a single object with multiple vectors, i.e. a 2-dimensional matrix. Multi-vectors are currently only available for HNSW indexes for named vectors. To use multi-vectors, enable it for the appropriate named vector.
+Multi-vector embeddings, also known as multi-vectors, represent a single object with multiple vectors, i.e. a 2-dimensional matrix. Multi-vectors are available for `hnsw` and, since `v1.40`, `hfresh` indexes ([with MUVERA encoding only](../config-refs/indexing/vector-index.mdx#hfresh-multi-vector)), for named vectors. To use multi-vectors, enable it for the appropriate named vector.