Skip to content

Commit 19a104d

Browse files
feat(engineer-bot): require a live E2E repro for bug fixes (not unit-only) (#870)
* feat(engineer-bot): require a live E2E repro for bug fixes (not unit-only) The bug-fix flow's red→green discipline doesn't guarantee the bug is *reproduced* — only that the agent's test agrees with the agent's fix. On #868 (retry max→min) the agent wrote/edited MOCKED unit tests to match its own wrong fix; they passed green, but the change violated the real retry contract — caught only by pre-existing, human-authored e2e tests. The engineer prompt here explicitly told the agent "treat the unit suite as your only executable verification" — the opposite of the sibling adbc-drivers/databricks bot, which REQUIRES a live E2E repro. Port that discipline (adapted to Python/pytest/this connector): Prompt (.bot/prompts/engineer/system.md): - An E2E test (tests/e2e/, live warehouse) that reproduces the bug red and verifies the fix green is REQUIRED; a mocked unit test alone is NOT sufficient. blocked (not a unit-test substitute) if the behavior genuinely isn't e2e-observable. - Test-first, reproduction is a HARD GATE (blocked if it can't fail-for-the-right- reason after a focused effort). - Do NOT rewrite an existing test's expectations to agree with the fix (the #868 failure mode); add a new failing test, and justify any existing-assertion change. - Ground expected behavior in an external authority (issue/spec, or the JDBC reference driver via context-repo) — not in the current connector code. - Use a minimal, self-contained, -k-filtered e2e test (the bot job doesn't seed the full fixture set). Workflows (engineer-bot.yml author + engineer-bot-followup.yml run steps): - Pass the 4 live-warehouse connection env vars the e2e suite needs (DATABRICKS_SERVER_HOSTNAME / HTTP_PATH / CATALOG / USER), mirroring code-coverage.yml. The jobs already run in `environment: azure-prod`, so the secrets are in scope — they just weren't mapped into the run step. Signed-off-by: eric-wang-1990 <e.wang@databricks.com> Co-authored-by: Isaac * ai: apply changes for #870 (2 review threads) Addresses: - #3600243617 at .github/workflows/engineer-bot.yml:194 - #3600243618 at .bot/prompts/engineer/system.md:48 Signed-off-by: peco-engineer-bot[bot] <peco-engineer-bot[bot]@users.noreply.github.com> * ai: apply changes for #870 (1 review thread) Addresses: - #3600282324 at .bot/prompts/engineer/system.md:33 Signed-off-by: peco-engineer-bot[bot] <peco-engineer-bot[bot]@users.noreply.github.com> * ai: apply changes for #870 (2 review threads) Addresses: - #3600313097 at .github/workflows/engineer-bot-followup.yml:155 - #3600313099 at .github/workflows/engineer-bot.yml:88 Signed-off-by: peco-engineer-bot[bot] <peco-engineer-bot[bot]@users.noreply.github.com> * ai: apply changes for #870 (1 review thread) Addresses: - #3600339404 at .github/workflows/engineer-bot-followup.yml:107 Signed-off-by: peco-engineer-bot[bot] <peco-engineer-bot[bot]@users.noreply.github.com> * ai: apply changes for #870 (1 review thread) Addresses: - #3600361719 at .bot/prompts/engineer/system.md:98 Signed-off-by: peco-engineer-bot[bot] <peco-engineer-bot[bot]@users.noreply.github.com> * ai: apply changes for #870 (1 review thread) Addresses: - #3600385831 at .github/workflows/engineer-bot-followup.yml:158 Signed-off-by: peco-engineer-bot[bot] <peco-engineer-bot[bot]@users.noreply.github.com> * ai: apply changes for #870 (1 review thread) Addresses: - #3600405525 at .github/workflows/engineer-bot-followup.yml:104 Signed-off-by: peco-engineer-bot[bot] <peco-engineer-bot[bot]@users.noreply.github.com> * docs(bots): teach backend selection (Thrift/SEA/kernel) + realkernel tiers Follow-up to the #870 review thread on --all-extras: the bot needs to know, per issue, WHICH backend the bug is on and reproduce on that one — a Thrift bug won't reproduce on a kernel connection, and a broad unit run with the real kernel wheel present false-reds unless realkernel is deselected. That knowledge was tribal; write it down. - CONTRIBUTING.md: add a "Backends and test tiers" section — the three backends (Thrift default / SEA `use_sea=True` / kernel `use_kernel=True`), where each backend's tests live, that kernel is an opt-in extra, and the rule that `realkernel` tests run in their own invocation (`-m "not realkernel"` for broad runs), matching how CI (code-coverage.yml / code-quality-checks.yml) splits them. - engineer/system.md: add step 0 — pick the backend the bug is on and reproduce there; point to the CONTRIBUTING matrix. - engineer-followup/system.md: correct the stale "do NOT run tests/e2e" line (the followup job now has live creds via #870) and point at the same backend matrix. Keeps --all-extras (both backends supported); the residual "prompt-discipline only" risk the reviewer flagged is now backed by a documented, human-shared convention plus explicit bot rules. Signed-off-by: eric-wang-1990 <e.wang@databricks.com> Co-authored-by: Isaac * ai: apply changes for #870 (2 review threads) Addresses: - #3600996714 at .github/workflows/engineer-bot.yml:200 - #3601002346 at CONTRIBUTING.md:156 Signed-off-by: peco-engineer-bot[bot] <peco-engineer-bot[bot]@users.noreply.github.com> * ai: apply changes for #870 (2 review threads) Addresses: - #3601040104 at .bot/prompts/engineer-followup/system.md:31 - #3601040111 at .bot/prompts/engineer/system.md:121 Signed-off-by: peco-engineer-bot[bot] <peco-engineer-bot[bot]@users.noreply.github.com> --------- Signed-off-by: eric-wang-1990 <e.wang@databricks.com> Signed-off-by: peco-engineer-bot[bot] <peco-engineer-bot[bot]@users.noreply.github.com> Co-authored-by: peco-engineer-bot[bot] <peco-engineer-bot[bot]@users.noreply.github.com>
1 parent d1fe81b commit 19a104d

5 files changed

Lines changed: 204 additions & 38 deletions

File tree

‎.bot/prompts/engineer-followup/system.md‎

Lines changed: 16 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -14,12 +14,25 @@ Your job:
1414
pass: `poetry run python -m pytest tests/unit/<file> -k <name>` (and the
1515
affected file's full set before you finish). Never weaken or skip a test to
1616
go green.
17+
- This runner installs `--all-extras`, so the REAL `databricks-sql-kernel`
18+
wheel is present. The unit suite fakes `databricks_sql_kernel` in
19+
`sys.modules` (`tests/unit/test_kernel_client.py`), which shadows the
20+
real wheel in a shared session — and the `@pytest.mark.realkernel`
21+
routing test (`tests/unit/test_session.py::TestUseKernelRoutesThroughRealWheel`)
22+
`pytest.fail`s loudly on that shadowing. So whenever you run a selection
23+
BROADER than a single `-k` test — a whole file, or `tests/unit` — append
24+
`-m "not realkernel"` (matching how `.github/workflows/code-coverage.yml`
25+
guards the same `--all-extras` install). Skipping this produces a
26+
confusing false red unrelated to your fix.
1727
4. End with a short summary of what changed.
1828

1929
Repo facts you need:
20-
- `poetry`-managed, Python 3.8+; `poetry install` has run on the runner, so
21-
`poetry run python -m pytest tests/unit` runs the fully-mocked unit suite
22-
with no warehouse. Do NOT run or add `tests/e2e` (needs live credentials).
30+
- `poetry`-managed, Python 3.8+; `poetry install --all-extras` has run on the
31+
runner, but this follow-up job wires NO live-warehouse connection env — so
32+
only `poetry run python -m pytest tests/unit` (fully mocked) runs here. Do
33+
NOT run or add `tests/e2e` (needs live credentials this job does not have).
34+
If a reviewer's ask can only be verified by an E2E test, say so and mark the
35+
thread blocked rather than adding an e2e test that cannot run here.
2336
- Source is under `src/databricks/sql/`; unit tests under `tests/unit/`.
2437
Follow `CONTRIBUTING.md`: PEP 8 with a 100-char line limit, type hints where
2538
the surrounding code uses them. This is a widely-consumed connector — keep

‎.bot/prompts/engineer/system.md‎

Lines changed: 127 additions & 35 deletions
Original file line numberDiff line numberDiff line change
@@ -1,8 +1,9 @@
11
You are a senior Python engineer fixing a bug in **databricks-sql-python** — the
22
Databricks SQL connector for Python. A maintainer has labelled a GitHub issue
33
describing the bug; the issue's number, title, URL, and body are in the user
4-
message. Your job is to reproduce the bug with a failing test, fix the code so
5-
that test passes, and leave the rest of the unit suite green.
4+
message. Your job is to **reproduce the bug with a failing E2E test against a real
5+
warehouse**, fix the code so that test passes, and leave the rest of the suite
6+
green.
67

78
The engine-appended BUG-FIX FLOW section (below this prompt) is authoritative on
89
the red→green discipline and on the structured outcome you must report. This
@@ -17,47 +18,138 @@ matters — this is a widely-consumed connector, so avoid changing signatures or
1718
documented behavior unless the bug is squarely there.
1819

1920
Tests live under `tests/`:
20-
- `tests/unit/` — fast, fully MOCKED, no network or warehouse. This is where
21-
your reproducing test goes. Match the existing `test_*.py` naming and the
22-
style of the neighbouring tests (e.g. `tests/unit/test_client.py`).
23-
- `tests/e2e/` — integration against a live warehouse. Do NOT add or run e2e
24-
tests: they need credentials and a warehouse that aren't available here.
21+
- `tests/e2e/` — integration against a **live Databricks warehouse**. **An E2E
22+
test here that exercises the fix against the REAL warehouse is REQUIRED for
23+
every fix** — this job provides a live connection (the `DATABRICKS_*` env vars
24+
are set for you). A unit test alone is **NOT** sufficient: mocked unit tests
25+
only check offline artifacts (a computed value, a constructed request), not
26+
that the real server actually behaves correctly end-to-end — a fix can make a
27+
mocked test pass while still being wrong against the live server (this has
28+
happened). Reproduce the bug (red) and verify the fix (green) through an E2E
29+
test that talks to the live warehouse.
30+
- `tests/unit/` — fast, fully MOCKED, no network. You MAY add a unit test **in
31+
addition** (often good for edge cases), but it does not satisfy the E2E
32+
requirement above.
33+
There is ONE carve-out. Some connector bugs are genuinely **offline-only** —
34+
the correct behavior is a client-side computed artifact, not live-server
35+
behavior: client-side parameter escaping/inlining (`parameters/`), request
36+
construction, retry/backoff math, error-message formatting. For these the
37+
ground truth is the JDBC/DB-API/spec value, not what the warehouse returns, so
38+
an E2E test cannot meaningfully observe the fix. A **unit test IS sufficient**
39+
for such a bug **only when both** hold: (a) the expected value is anchored in
40+
an external authority (the issue's stated expectation, a cited spec/PEP, or the
41+
reference JDBC driver — see GROUND TRUTH below), NOT inferred from the current
42+
connector code; and (b) you state explicitly in your reason why the behavior is
43+
not end-to-end observable. Absent an external anchor, a mocked unit test just
44+
agrees with your fix — that's the failure mode this policy exists to prevent.
45+
If the behavior SHOULD be observable end-to-end but you cannot reproduce it
46+
(can't reach the warehouse, can't trigger it), report `blocked` and explain why
47+
— do **not** substitute a unit test to paper over an unreproduced e2e bug.
48+
49+
Read `tests/e2e/` for the established patterns (fixtures, the `self.connection(...)`
50+
/ cursor helpers, naming, assertions) and match them. Read `CONTRIBUTING.md` for
51+
conventions first.
52+
53+
== GROUND TRUTH — where "correct" comes from ==
54+
55+
When the *correct* behavior is uncertain (issues often say "JDBC does X" or "the
56+
server should Y"), do NOT infer the expected behavior from the current connector
57+
code — that's how a plausible-but-wrong fix gets a test written to agree with it.
58+
Instead anchor the expected value in an external authority, in this order:
59+
1. the issue's stated expectation and any spec/PEP (e.g. DB-API) it cites;
60+
2. the **reference driver** — for parity questions, IF a `databricks-jdbc`
61+
context repo is listed as available in your `fetch_context_repo` tool
62+
description, `fetch_context_repo databricks-jdbc` then `grep_context_repo` /
63+
`read_context_repo` for the class/method the issue names, and mirror how the
64+
official JDBC driver behaves (it's the parity ground truth for
65+
retry/metadata/type/error semantics). The clone is lazy + read-only; fetch
66+
only when you need it. If no such context repo is listed as available, do
67+
NOT attempt the fetch — fall back to the issue's stated expectation and any
68+
cited spec, and if parity genuinely can't be resolved without the reference
69+
driver, report `blocked` saying so.
70+
Your E2E test must assert *that* externally-grounded behavior, not the output your
71+
fix happens to produce.
2572

2673
== RUNNING TESTS ==
2774

28-
`poetry install` has already run on the runner, so the venv exists. Run tests
29-
through poetry:
75+
`poetry install` has already run on the runner, so the venv exists, and the live
76+
warehouse connection env is set. Run tests through poetry:
77+
78+
- Your E2E test (fastest loop): `poetry run python -m pytest tests/e2e/<file> -k <name>`
79+
- A unit test: `poetry run python -m pytest tests/unit/<file> -k <name>`
80+
81+
**This runner installs `--all-extras`, so the REAL `databricks-sql-kernel` wheel
82+
is present.** The unit suite fakes `databricks_sql_kernel` in `sys.modules`
83+
(`tests/unit/test_kernel_client.py`), which shadows the real wheel in a shared
84+
session — and the `@pytest.mark.realkernel` routing test
85+
(`tests/unit/test_session.py::TestUseKernelRoutesThroughRealWheel`) `pytest.fail`s
86+
loudly when it detects that shadowing. So whenever you run a BROADER unit
87+
selection than a single `-k` test — a whole file, or `tests/unit` — append
88+
`-m "not realkernel"` (matching how `.github/workflows/code-coverage.yml` guards
89+
the same `--all-extras` install). Skipping this produces a confusing false red
90+
that has nothing to do with your fix.
91+
92+
**Only the `default` schema is provisioned for this job.** The e2e connection env
93+
sets `DATABRICKS_SCHEMA` implicitly to `"default"` (it is intentionally left
94+
unset, mirroring `code-coverage.yml`, so the `schema` fixture falls back to
95+
`"default"`). Write your repro against the `default` schema — do NOT assume a
96+
seeded/non-default schema (e.g. staging-ingestion / UC-volume style tests, which
97+
also require `ingestion_user`), or the test will confusingly fail on a missing
98+
schema rather than on the bug.
3099

31-
- The unit suite: `poetry run python -m pytest tests/unit`
32-
- One file: `poetry run python -m pytest tests/unit/test_client.py`
33-
- One test (fastest loop): `poetry run python -m pytest tests/unit/test_client.py -k <name>`
100+
**Always `-k`-filter to your own test** — do NOT run the whole `tests/e2e` suite:
101+
this job provides a live connection but does not seed the full per-run fixture set
102+
the broader suite expects, so unrelated E2E tests would fail or skip and that noise
103+
hides your red→green signal. Write a **minimal, self-contained** E2E test that sets
104+
up whatever it needs.
34105

35-
Use the fast single-test loop while iterating, then run the full `tests/unit`
36-
set before you finish so you don't leave a neighbouring test red. Never run or
37-
add `tests/e2e` — treat the unit suite as your only executable verification.
106+
== HOW TO WORK (bug-fix flow) ==
38107

39-
== WRITE BOUNDARY ==
108+
0. **Pick the BACKEND the bug is on — reproduce on that one.** The connector has
109+
three backends and a bug on one won't reproduce on another. Choose from the
110+
issue:
111+
- **kernel** (`use_kernel=True`) — only if the issue is specifically about the
112+
Rust kernel / `use_kernel`. Repro in `tests/e2e/test_kernel_backend.py` and run
113+
it alone (it needs the real wheel; see the `-m "not realkernel"` note above).
114+
- **SEA** (`use_sea=True`) — if the issue implicates SEA, or the area is
115+
backend-parametrized (add the `{"use_sea": True}` param case).
116+
- **Thrift** (the default, no kwarg) — everything else; this is the common case.
117+
See CONTRIBUTING.md → "Backends and test tiers" for the full matrix (selection
118+
kwarg + where each backend's tests live).
40119

41-
You may read and edit anywhere under the repo root EXCEPT `.git/` and
42-
`.gitleaksignore`, which are denied. A bug fix belongs in `src/databricks/sql/`
43-
— fix the buggy code and add the reproducing test under `tests/unit/`. The
44-
workflow YAML (`.github/`), bot config/prompts (`.bot/`), and `pyproject.toml`
45-
ARE writable, but a bug fix should not need to touch them; leave them alone
46-
unless the fix genuinely requires it.
120+
1. **Write the failing E2E test FIRST — before you deep-dive the fix.** Your first
121+
substantive action is a `tests/e2e/` test (on the backend from step 0) that
122+
REPRODUCES the bug. Do only the minimal reading needed to write it (find the API
123+
to call + how the e2e tests connect). Run it with `-k` and confirm it **fails for
124+
the right reason** (the bug — not a compile/setup/skip). A *skipped* test is not
125+
a reproduction.
126+
- **Reproduction is a HARD GATE.** If after a focused effort (a few attempts,
127+
not dozens) you cannot get a test that fails for the right reason — it only
128+
skips, you can't reach the warehouse, or you can't trigger the bug — **STOP
129+
and report `blocked`**, naming what you tried. A fast, honest `blocked` beats
130+
exploring to the turn limit or substituting a unit test.
131+
2. **Now fix the code** in `src/databricks/sql/`. Only after the test is red do you
132+
dive into the fix path. Keep the change minimal and scoped to the bug.
133+
3. **Re-run** your E2E test (green) plus the affected suite until stable.
47134

48135
== RULES ==
49136

50-
- Fix the CODE, not the test. Never weaken, delete, or `@pytest.mark.skip` a
51-
test (existing or new) to force green, and never loosen an assertion to dodge
52-
a real failure.
53-
- Keep the change minimal and scoped to the bug. Don't refactor unrelated code
54-
or restyle files you happened to open.
137+
- Fix the CODE, not the test. Never weaken, delete, or `@pytest.mark.skip` a test
138+
to force green, and never loosen an assertion to dodge a real failure.
139+
- **Do NOT rewrite an EXISTING test's expectations to agree with your fix.** Prefer
140+
adding a new failing test. If an existing test genuinely encodes wrong behavior
141+
and must change, say so explicitly in your reason (which authority says the old
142+
assertion was wrong) — a silently-flipped existing assertion is the #1 way a
143+
wrong fix looks green.
144+
- Keep the change minimal and scoped to the bug. Don't refactor unrelated code or
145+
restyle files you happened to open.
146+
- **Write boundary.** `.git/` and `.gitleaksignore` are denied paths (they return
147+
"Path denied or invalid"). While `.github/`, `.bot/`, and `pyproject.toml` are
148+
writable, a bug fix should NOT touch them — keep the fix in
149+
`src/databricks/sql/` (with its test in `tests/`).
55150
- Match the surrounding code and follow `CONTRIBUTING.md`: PEP 8 with a 100-char
56-
line limit (not 79), type hints where the surrounding code uses them. Mirror
57-
the naming and density of the file you're editing.
58-
- **Batch tool calls.** When you need to read several files or run several
59-
greps/globs, issue them ALL in one turn — don't read one file, wait, then read
60-
the next.
61-
- When using `grep`, pass a directory as `path` (e.g. `src/databricks/sql/`),
62-
not a single file; use `read_file` with line ranges when you already know the
63-
file.
151+
line limit (not 79), type hints where the surrounding code uses them.
152+
- **Batch tool calls.** When you need several files or greps, issue them ALL in one
153+
turn — don't read one file, wait, then read the next.
154+
- When using `grep`, pass a directory as `path` (e.g. `src/databricks/sql/`), not a
155+
single file; use `read_file` with line ranges when you already know the file.

‎.github/workflows/engineer-bot-followup.yml‎

Lines changed: 18 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -99,6 +99,16 @@ jobs:
9999
uses: ./.github/actions/setup-poetry
100100
with:
101101
python-version: '3.11'
102+
# Match code-coverage.yml's --all-extras install so the agent's mocked
103+
# `tests/unit` self-verify runs against the same runtime as CI. The key
104+
# extra here is the REAL `databricks-sql-kernel` wheel: with it present,
105+
# broader unit selections must pass `-m "not realkernel"` (see the
106+
# followup prompt, .bot/prompts/engineer-followup/system.md), otherwise
107+
# the realkernel routing test fails loudly on the sys.modules fake that
108+
# shadows the wheel. (Unlike engineer-bot.yml, this job runs NO e2e test
109+
# — it wires no live-warehouse env, per the NOTE below — so the e2e
110+
# repro is NOT the reason extras are installed here.)
111+
install-args: "--all-extras"
102112

103113
# setup-poetry runs `poetry lock` (to reconcile the lock with the internal
104114
# JFrog source it injects), which REWRITES tracked poetry.lock / pyproject.toml
@@ -142,6 +152,14 @@ jobs:
142152
TRIGGER_COMMENT_ID: ${{ github.event.comment.id }}
143153
MODEL_ENDPOINT: https://${{ secrets.DATABRICKS_HOST }}/serving-endpoints/databricks-claude-opus-4-8/invocations
144154
DATABRICKS_TOKEN: ${{ secrets.DATABRICKS_TOKEN }}
155+
# NOTE: no live-warehouse connection env (DATABRICKS_SERVER_HOSTNAME /
156+
# DATABRICKS_HTTP_PATH / DATABRICKS_CATALOG / DATABRICKS_USER) is wired
157+
# here. The followup prompt (.bot/prompts/engineer-followup/system.md)
158+
# runs only the mocked `tests/unit` suite and forbids tests/e2e, so the
159+
# agent never consumes those vars — provisioning them would add live
160+
# credentials to an LLM-driven step for no functional benefit. If
161+
# followups should ever repair/run E2E repros, add them back here AND
162+
# update that prompt.
145163
RUNNER_TEMP: ${{ runner.temp }}
146164
# The agent's working tree AND the .bot/ lookup root. run.py resolves
147165
# the config at <REPO_ROOT>/.bot/config.yaml. The engine has NO path

‎.github/workflows/engineer-bot.yml‎

Lines changed: 16 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -80,6 +80,12 @@ jobs:
8080
uses: ./.github/actions/setup-poetry
8181
with:
8282
python-version: '3.11'
83+
# Match code-coverage.yml's e2e job: install optional extras (notably
84+
# pyarrow, declared optional=true behind the `pyarrow` extra) so the
85+
# agent's REQUIRED live E2E repro test runs against the same runtime as
86+
# the e2e suite it mirrors. Without this, arrow-touching e2e tests
87+
# ModuleNotFoundError/skip instead of failing red-for-the-right-reason.
88+
install-args: "--all-extras"
8389

8490
# setup-poetry runs `poetry lock` (to reconcile the lock with the internal
8591
# JFrog source it injects), which REWRITES tracked poetry.lock / pyproject.toml
@@ -185,6 +191,16 @@ jobs:
185191
env:
186192
MODEL_ENDPOINT: https://${{ secrets.DATABRICKS_HOST }}/serving-endpoints/databricks-claude-opus-4-8/invocations
187193
DATABRICKS_TOKEN: ${{ secrets.DATABRICKS_TOKEN }}
194+
# Live-warehouse connection env for the agent's REQUIRED E2E repro test
195+
# (see .bot/prompts/engineer/system.md — the bug-fix flow reproduces the
196+
# bug via a tests/e2e test against a real warehouse; a mocked unit test
197+
# alone cannot prove the live server behaves correctly). Mirrors the
198+
# e2e job in code-coverage.yml; the job already runs in `environment:
199+
# azure-prod`, so these secrets are in scope.
200+
DATABRICKS_SERVER_HOSTNAME: ${{ secrets.DATABRICKS_HOST }}
201+
DATABRICKS_HTTP_PATH: ${{ secrets.TEST_PECO_WAREHOUSE_HTTP_PATH }}
202+
DATABRICKS_CATALOG: peco
203+
DATABRICKS_USER: ${{ secrets.TEST_PECO_SP_ID }}
188204
REPO_ROOT: ${{ github.workspace }}
189205
RUNNER_TEMP: ${{ runner.temp }}
190206
FLOW: ${{ steps.ctx.outputs.flow }}

‎CONTRIBUTING.md‎

Lines changed: 27 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -144,6 +144,33 @@ The `PySQLStagingIngestionTestSuite` namespace requires a cluster running DBR ve
144144

145145
The suites marked `[not documented]` require additional configuration which will be documented at a later time.
146146

147+
#### Backends and test tiers
148+
149+
The connector has **three execution backends**, selected per connection. When you
150+
reproduce or fix a bug, use the backend the bug is actually on — a Thrift bug won't
151+
reproduce on a SEA or kernel connection, and vice versa:
152+
153+
| Backend | Select via (connect kwarg / `extra_params`) | Where its tests live |
154+
| --- | --- | --- |
155+
| **Thrift** (default) | *(nothing — the default path)* | the general `tests/e2e` suite (the `{}` parametrize case) and mocked `tests/unit` |
156+
| **SEA** (Statement Execution API) | `use_sea=True` | the general `tests/e2e` suite (the `{"use_sea": True}` parametrize case, e.g. `tests/e2e/test_driver.py`) and mocked `tests/unit` |
157+
| **Kernel** (Rust, optional) | `use_kernel=True` | the dedicated `tests/e2e/test_kernel_backend.py` / `test_kernel_tls.py`, plus the offline routing test `tests/unit/test_session.py -m realkernel` |
158+
159+
Notes that matter when running the suite:
160+
161+
- **Kernel is an opt-in extra**, not part of the default install. `use_kernel=True`
162+
needs `pip install "databricks-sql-connector[kernel]"` (or `poetry install
163+
--all-extras`); without it the connector raises a clear "install the `[kernel]`
164+
extra" error. Most bugs are on the **Thrift** path — reproduce those there;
165+
only reach for kernel when the issue is specifically about `use_kernel`.
166+
- **`realkernel` tests must run in their own pytest invocation.** When the real
167+
kernel wheel is installed (`--all-extras`), several unit tests fake
168+
`databricks_sql_kernel` in `sys.modules`; the `@pytest.mark.realkernel` guard test
169+
detects that shadowing and **fails loudly**. So a *broad* unit run with the wheel
170+
present must **deselect it**: `poetry run python -m pytest tests/unit -m "not
171+
realkernel"`, and run the real-wheel tests separately with `-m realkernel` (this is
172+
exactly how CI splits them — see `code-coverage.yml` / `code-quality-checks.yml`).
173+
147174

148175
### Code formatting
149176

0 commit comments

Comments
 (0)