Repository navigation
Add OTLP Metrics ingest support - #2127
gareth-ellis wants to merge 31 commits into
Conversation
| * ``body``: Raw binary protobuf payload (bytes). | ||
|
|
||
| Optional parameters: | ||
| * ``endpoint``: OTLP endpoint path. Defaults to ``/_otlp/v1/metrics``. |
There was a problem hiding this comment.
Could a signal-type attribute be considered with possible values of metrics, logs, or traces instead of endpoint?
There was a problem hiding this comment.
🟡 Changes recommended
Retry parameters are dropped, prepared variants can resolve incorrectly, and corpus validation can accept corrupt or mismatched files.
12 open findings
Validate offset indexes against the protobuf corpus · New Reject truncated records during corpus scanning · New Reject negative retry values · New Keep prepared OTLP format variants in the same directory · New Filter corpora to the selected OTLP document set · New Forward OTLP retry options from track parameters · New Bound OTLP batches by cumulative bytes · New Correct documented prepared filenames to include .json · New Correct the documented batch count to 360 · New Document installation of the OTLP conversion extra · New Exclude retry waits from measured service time · New Restore the missing test method boundary · New
What changed in this PR
Adds end-to-end OTLP metrics benchmarking through Elasticsearch’s native OTLP endpoint.
Changes:
- Adds OTLP corpus conversion, preparation, partitioning, and gzip support.
- Introduces the
otlp-ingestoperation with retries and metrics. - Adds dependencies, documentation, and comprehensive tests.
| File | Description |
|---|---|
uv.lock |
Locks OTLP protobuf dependencies. |
pyproject.toml |
Defines OTLP and development dependencies. |
create-notice.sh |
Adds dependency license notices. |
esrally/utils/io.py |
Converts and reads protobuf corpus files. |
esrally/track/track.py |
Registers the OTLP format and operation. |
esrally/track/params.py |
Adds OTLP corpus partitioning and parameters. |
esrally/track/loader.py |
Adds format-aware corpus preparation. |
esrally/driver/runner.py |
Implements OTLP HTTP ingestion and retries. |
tests/utils/io_test.py |
Tests conversion and record indexing. |
tests/track/params_test.py |
Tests OTLP parameter sourcing. |
tests/track/loader_test.py |
Tests corpus preparation and parsing. |
tests/driver/runner_test.py |
Tests OTLP request and retry behavior. |
docs/track.rst |
Documents the new format and operation. |
docs/otlp_ingest.rst |
Adds the OTLP usage guide. |
docs/index.rst |
Includes the guide in documentation navigation. |
docs/advanced.rst |
Updates preparation architecture documentation. |
🧠 Review effort: Balanced
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| # like bulk: trust an existing offset file and only verify when we have to build it | ||
| if os.path.exists(pb_file.pb_path + ".offset"): | ||
| return True |
There was a problem hiding this comment.
This is intentional. The same behavior is present in case of bulk corpus. The assumption is if someone uploaded offset file into the corpus store they want to reduce the corpus preparation time aggressively. We can change this assumption for both OTLP and bulk in unison but not in this PR.
There was a problem hiding this comment.
🟡 Changes recommended
Offline preparation and prepared-variant path resolution can prevent valid local OTLP corpora from running.
6 open findings
Resolve prepared files using the operation's required format · New Validate offset indexes against the protobuf corpus Treat offline download errors as missing optional prebuilt files · New Document atomic fallback across all prepared format variants · New Correct prepared filenames to include the complete source path · New Align partitioning documentation with sparse-index implementation · New
11 resolved since last review
Reject truncated records during corpus scanning Bound OTLP batches by cumulative bytes Forward OTLP retry options from track parameters Filter corpora to the selected OTLP document set Keep prepared OTLP format variants in the same directory Reject negative retry values Restore the missing test method boundary Exclude retry waits from measured service time Document installation of the OTLP conversion extra Correct the documented batch count to 360 Correct documented prepared filenames to include .json
🧠 Review effort: Balanced
| suffixes = DOCUMENT_SET_FORMATS[document_set.source_format].prepared_file_suffixes() | ||
| resolved = first_existing_with_any_suffix(data_root, document_set.document_file, suffixes) |
| pb_archive_path = pb_path + archive_ext | ||
| try: | ||
| preparator.downloader.download(document_set.base_url, pb_archive_path) | ||
| except exceptions.DataError: |
| if len(data_root) == 1: | ||
| preparator.prepare_document_set(document_set, data_root[0], fmt) | ||
| # attempt to prepare everything in the current directory and fallback to the corpus directory | ||
| elif not preparator.prepare_bundled_document_set(document_set, data_root[0], fmt): | ||
| preparator.prepare_document_set(document_set, data_root[1], fmt) |
| When a challenge runs ``otlp-ingest`` with multiple clients (``"clients": N``), Rally splits the corpus across clients so each client reads a distinct, non-overlapping slice of the records. Partitioning uses the ``.pb.offset`` index for O(1) seek to each client's starting record. | ||
|
|
||
| Each client's slice size is ``floor(total_records / N)``; the final client gets any remainder. This guarantees each record is sent exactly once per pass across all clients. |
… include the complete source path' Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>



This commit adds support for indexing to the _otlp/v1/metrics endpoint of elasticsearch.
There's quite a bit more than just a new runner here, as follows: