diff --git a/AGENTS.md b/AGENTS.md index d7e8ff4f..4311816c 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -320,7 +320,7 @@ The 2.0 rewrite lands as underscore-private modules alongside the v1 code. These - **Method organization**, fixed section order in every class: fields + `__post_init__` validation → alternative constructors → dunders (construction/equality → protocol → operators) → properties → public methods by concern (access → editing → comparison → rendering delegates) → private helpers last, except a helper serving exactly one section may sit at that section's head. Sanctioned deviation, facade layer only: `HumanName` and the shim `Constants` organize by v1 concern groups (`# -- render defaults --`, `# -- config / parsing --`, `# -- fields --`, ..., dunders and pickle last) — the classes mirror v1's own surface and die in 3.0; the canonical order still binds every core type. - **Validation is eager and fail-loud**: every `raise` states the offending value, the expected form, and the fix. Exception taxonomy: wrong type — including wrong element type inside a collection, bare `str` where an iterable of strings is expected, or a `Mapping` where a plain iterable is expected — raises `TypeError`; well-typed but unacceptable values raise `ValueError`; failed enum lookups stay `ValueError` for any input (stdlib `EnumType` precedent). **When the message hands the reader code to paste, that code has to survive a type checker** — nameparser ships `py.typed`. #337's segmenterless warning offered `Policy(segment_scripts=())`, an `arg-type` error, because these fields are annotated with what they STORE rather than everything the constructor accepts. Prefer the `frozenset()` / `()` spellings in messages and docstrings, and pin the offered spelling in a test — the warning tests matched on `ja_segmenter` and never checked the actionable half of the message. **A warning emitted in `Parser.__post_init__` needs `parser_for` to re-emit it from its own frame** (the `catch_warnings(record=True)` block at its return): `__post_init__`'s `stacklevel` is sized for direct `Parser(...)` construction, and through `parser_for`'s extra frame the default one-line rendering attributes the warning to the library's own `return Parser(...)` — the exact call the message tells the user to change becomes invisible. No single stacklevel serves both entry points; a new construction warning gets the re-emission for free, but a new CONSTRUCTION SITE for `Parser` inside this package needs its own re-emission or its callers get library-attributed warnings (#337 review). - **Guard, hint, and emit for the WHOLE family, and parametrize the test over it**: a check added to one member of a set belongs on all of it, and the test must sweep the family, not one example. This session shipped `_reject_str_and_mapping` on `Policy` but not `PolicyPatch`, the bytes decode hint on three of five config entry points, and a regex-sync roster missing four of its copies — each a separate follow-up bug that a `{class} × {field} × {bad-value}` parametrization would have caught and a per-example test hid. When you find you're guarding member N, grep for the other members first. -- **Ambiguities are emitted at the DECISION site**: an `Ambiguity` records a fork the parse had to call, not a token that sits in an ambiguous vocabulary. Emit where the branch is taken — the trailing-suffix peel (read once in `_group` for a main segment since #614, its picks reported by `_assign` once the roles they took are known, except a pick the particle chain takes into the name, which the chain reports and the read then gives up), the delimiter escape's follow-up in `classify` — never by scanning for a `vocab:*-ambiguous` tag. The same tagged token is a genuine fork in one position and unremarkable in another (`do` mid-name in "Joao da Silva do Amaral de Souza" chooses nothing). **A branch that runs but changes nothing is not a decision either** -- the prefix chain's `merge_pieces(pieces, ptags, k, j)` executes even when `j == k + 1`, folding a piece into itself, and keying on "the code got here" reported a fork for all ambiguous particles on "Do Van Jr." (`Dr.` when that was written, before #367 made a plain title transparent and put the shape out of the loop's reach entirely), where the particle stayed a lone leading name piece — the GIVEN name under the default order, the family name under `FAMILY_FIRST` — and `_assign` reported the same token again. Check that the branch actually claimed something (`j > k + 1`) before recording. Structure often settles the question before it arises, which is why `PARTICLE_OR_GIVEN` is not emitted on the `FAMILY_COMMA` path's WHOLLY-FAMILY read -- the comma fixed which piece is the family -- and `SUFFIX_OR_NAME` is not emitted for "Ma, Jack". A tail segment is the same case from the other side: assign reads it as suffixes, a maiden marker there included since #601, and since #603 a title word past the second comma as a title (rules.md#C2, so `Freiherr` below reads title) (`Jane Doe, PhD, Jr nee van Ma` reads suffix `PhD, Jr nee van Ma`, the marker an ordinary word, rules.md#M2), so group's two chain emitters are handed no report list there, after either comma — `John Smith, Jr., Freiherr von Richthofen` still chains `von` and reports only `comma-structure` (rules.md#C2, 2026-09-28). Read that scope narrowly: the comma settles nothing about a particle trailing the given name, so P6's attachment decides that fork on the same path and reports it (#405; in `post_rules` until #613, in `_assign` since, right after the walk that reads the given part), in the kind naming the reading it OVERRODE, which is the reading assign made and not the word's vocabulary: `SUFFIX_OR_NAME` where assign had read the run as a post-nominal (`vd`, `mc`), else `PARTICLE_OR_GIVEN` where the run holds an ambiguous particle (`van`, and `do`, which is in the suffix vocabulary too but in its AMBIGUOUS half, so no credential reading was overridden), else silence. The decision site also has the token index and the detail text in hand, which the tag scan would have to reconstruct -- all but the FIELD a report names, which a rule in post_rules can still change (H1's move behind a title, P1's family-first fold, P6's attachment, and under opt-in policies the patronymic rotations and `middle_as_family`'s fold): an emitter naming one passes the rest of its sentence as `field_tail` and assemble words the field from the final role (#626, where `Kim Min Do` under `FAMILY_FIRST` was told 'Do' was a middle name with 'Do' in the family). **If a fork's two branches are taken in DIFFERENT stages, every one of them needs the emitter** -- `PARTICLE_OR_GIVEN` is decided in `_assign` when the ambiguous particle stays a lone leading piece, in `_group` when something shifts it off the name's leading piece and the prefix chain claims it, and wherever P6's attachment takes a trailing particle into the family -- in `_assign` after a comma that names a family (since #613), in `post_rules` at the end of a name read family-first (#467) and after a comma with nothing before it, where H1 and M4 must read the name first -- so every one of them reports; for two years only the first did. What can still do the shifting is narrow, and #367 is why: a plain title no longer can (`Dr. Van Johnson` reads as `Van Johnson` does and reports from `_assign`), so the `_group` emitter needs a word that is BOTH a title and a particle — measured, `TITLES ∩ particles_ambiguous` is `{freiherr, st}` in the default vocabulary (`do` left TITLES in #296's audit; decisions.md's Excluded block records the before and after), plus any overlap a caller's config creates — standing ahead of the chained particle as the LEADING NAME word. Titles may precede it, so `Dr. St van Johnson` reaches the emitter and `St van Johnson` does too; a given name may not, so `Jan Freiherr von Richthofen` does not reach it while `Freiherr von Richthofen` and `Dr. Freiherr von Richthofen` do. Two shapes that look like they should reach it and do NOT, both measured by stepping `STAGES` and watching where `ambiguities` grows: `Dr. Do van Johnson` and `Do St Johnson` report from `assign`, not `group`, because `do` is no longer a title and so stays the leading name piece assign reports on — a both-vocabulary word CHAINED (`Jan St Johnson`) reports nothing at all. When checking whether that emitter is dead, a both-vocabulary word in the leading name position is the thing to look for, and the answer is that it is not dead. The stage-ownership map in `tests/v2/pipeline/test_state.py` must list `ambiguities` for each such stage, and it passes vacuously until a case row exercises the path, so add the row too. Report BOTH directions of a two-way fork — "John Smith MA" (read as a suffix) and "Jack MA" (read as the family name) are equally guesses. Every kind needs a trigger in `tests/v2/test_contracts.py::_AMBIGUITY_TRIGGERS` (an explicit `None`, strict-xfail, while reserved), and case-table rows pin expected kinds exactly, so a new emitter shows up in both immediately. **Pin the decision, not the vocabulary**: the only titled-particle test used an UNAMBIGUOUS particle, so it walked the right code path and proved nothing about the branch under test -- two criticals passed 1539 tests. A row contrasting the two readings ("John Smith V" against "John Smith B") is what makes an emitter's absence meaningful. +- **Ambiguities are emitted at the DECISION site**: an `Ambiguity` records a fork the parse had to call, not a token that sits in an ambiguous vocabulary. Emit where the branch is taken — the trailing-suffix peel (read once in `_group` for a main segment since #614, its picks reported by `_assign` once the roles they took are known, except a pick the particle chain takes into the name, which the chain reports and the read then gives up), the delimiter escape's follow-up in `classify` — never by scanning for a `vocab:*-ambiguous` tag. The same tagged token is a genuine fork in one position and unremarkable in another (`do` mid-name in "Joao da Silva do Amaral de Souza" chooses nothing). **A branch that runs but changes nothing is not a decision either** -- the prefix chain's `merge_pieces(pieces, ptags, k, j)` executes even when `j == k + 1`, folding a piece into itself, and keying on "the code got here" reported a fork for all ambiguous particles on "Do Van Jr." (`Dr.` when that was written, before #367 made a plain title transparent and put the shape out of the loop's reach entirely), where the particle stayed a lone leading name piece — the GIVEN name under the default order, the family name under `FAMILY_FIRST` — and `_assign` reported the same token again. Check that the branch actually claimed something (`j > k + 1`, or a run of particles group merged into the piece ahead of S2's read, which the chain then claims nothing past, #625) before recording. Structure often settles the question before it arises, which is why `PARTICLE_OR_GIVEN` is not emitted on the `FAMILY_COMMA` path's WHOLLY-FAMILY read -- the comma fixed which piece is the family -- and `SUFFIX_OR_NAME` is not emitted for "Ma, Jack". A tail segment is the same case from the other side: assign reads it as suffixes, a maiden marker there included since #601, and since #603 a title word past the second comma as a title (rules.md#C2, so `Freiherr` below reads title) (`Jane Doe, PhD, Jr nee van Ma` reads suffix `PhD, Jr nee van Ma`, the marker an ordinary word, rules.md#M2), so group's two chain emitters are handed no report list there, after either comma — `John Smith, Jr., Freiherr von Richthofen` still chains `von` and reports only `comma-structure` (rules.md#C2, 2026-09-28). Read that scope narrowly: the comma settles nothing about a particle trailing the given name, so P6's attachment decides that fork on the same path and reports it (#405; in `post_rules` until #613, in `_assign` since, right after the walk that reads the given part), in the kind naming the reading it OVERRODE, which is the reading assign made and not the word's vocabulary: `SUFFIX_OR_NAME` where assign had read the run as a post-nominal (`vd`, `mc`), else `PARTICLE_OR_GIVEN` where the run holds an ambiguous particle (`van`, and `do`, which is in the suffix vocabulary too but in its AMBIGUOUS half, so no credential reading was overridden), else silence. The decision site also has the token index and the detail text in hand, which the tag scan would have to reconstruct -- all but the FIELD a report names, which a rule in post_rules can still change (H1's move behind a title, P1's family-first fold, P6's attachment, and under opt-in policies the patronymic rotations and `middle_as_family`'s fold): an emitter naming one passes the rest of its sentence as `field_tail` and assemble words the field from the final role (#626, where `Kim Min Do` under `FAMILY_FIRST` was told 'Do' was a middle name with 'Do' in the family). **If a fork's two branches are taken in DIFFERENT stages, every one of them needs the emitter** -- `PARTICLE_OR_GIVEN` is decided in `_assign` when the ambiguous particle stays a lone leading piece, in `_group` when something shifts it off the name's leading piece and the prefix chain claims it, and wherever P6's attachment takes a trailing particle into the family -- in `_assign` after a comma that names a family (since #613), in `post_rules` at the end of a name read family-first (#467) and after a comma with nothing before it, where H1 and M4 must read the name first -- so every one of them reports; for two years only the first did. What can still do the shifting is narrow, and #367 is why: a plain title no longer can (`Dr. Van Johnson` reads as `Van Johnson` does and reports from `_assign`), so the `_group` emitter needs a word that is BOTH a title and a particle — measured, `TITLES ∩ particles_ambiguous` is `{freiherr, st}` in the default vocabulary (`do` left TITLES in #296's audit; decisions.md's Excluded block records the before and after), plus any overlap a caller's config creates — standing ahead of the chained particle as the LEADING NAME word. Titles may precede it, so `Dr. St van Johnson` reaches the emitter and `St van Johnson` does too; a given name may not, so `Jan Freiherr von Richthofen` does not reach it while `Freiherr von Richthofen` and `Dr. Freiherr von Richthofen` do. Two shapes that look like they should reach it and do NOT, both measured by stepping `STAGES` and watching where `ambiguities` grows: `Dr. Do van Johnson` and `Do St Johnson` report from `assign`, not `group`, because `do` is no longer a title and so stays the leading name piece assign reports on — a both-vocabulary word CHAINED (`Jan St Johnson`) reports nothing at all. When checking whether that emitter is dead, a both-vocabulary word in the leading name position is the thing to look for, and the answer is that it is not dead. The stage-ownership map in `tests/v2/pipeline/test_state.py` must list `ambiguities` for each such stage, and it passes vacuously until a case row exercises the path, so add the row too. Report BOTH directions of a two-way fork — "John Smith MA" (read as a suffix) and "Jack MA" (read as the family name) are equally guesses. Every kind needs a trigger in `tests/v2/test_contracts.py::_AMBIGUITY_TRIGGERS` (an explicit `None`, strict-xfail, while reserved), and case-table rows pin expected kinds exactly, so a new emitter shows up in both immediately. **Pin the decision, not the vocabulary**: the only titled-particle test used an UNAMBIGUOUS particle, so it walked the right code path and proved nothing about the branch under test -- two criticals passed 1539 tests. A row contrasting the two readings ("John Smith V" against "John Smith B") is what makes an emitter's absence meaningful. - **A kind is worth adding only if a reader would hesitate too**: the test is not "does the code take a branch" but whether a person reading that input would genuinely be unsure. "Smith, John V" reads as a middle initial to anyone -- the comma settles it -- so reporting it would be noise that teaches callers to ignore the field, which costs more than the missing report. Reachability of the second branch is necessary, not sufficient. Prefer leaving a fork silent and documenting the omission over emitting on input nobody finds ambiguous. - **Parser owns config-dependent conveniences**: `Parser.matches`/`Parser.capitalized`/`Parser.revise` exist because the `ParsedName` equivalents fall back to DEFAULT config for str/omitted arguments (documented loudly in both docstrings). `revise` harvests tokens from a full sub-parse of each replacement value (tags kept minus `FOLDED_TAG`, roles forced, the R1 entry pass `suffix_entries` re-run over the forced state so a suffix value's entries follow its own commas, ambiguities discarded); the merge tail is shared with `replace()` via `ParsedName._with_field_tokens`. `Parser.capitalized` delegates through `name.capitalized(self.lexicon)` specifically so `_parser` never imports `_render` — keep it that way. - **Per-word vocabulary fields warn on multi-word entries** (`_normset`/`_normpairs` via `_warn_dead_entry`, UserWarning, never a raise — see the given_name_titles Gotcha for why raising is wrong). `given_name_titles` is the one multi-word-matched field and is exempt; `_edit` passes `warn=False` (add() warns once via the new instance's `__post_init__`; remove() stores nothing). The default vocabulary and every locale pack must stay warning-free (`test_default_lexicon_builds_warning_free`, `test_pack_vocabulary_entries_are_single_words`). diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 2d406dc4..eaca6b59 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -763,18 +763,22 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f RESIDUALS, the paths that read as before #614: a segment of fewer than three pieces (no join reaches it); a run with a name piece inside it (an H5 title in front of a name word the peel declined, `John Dr. G.J.`, `John van V Dr. V`), which is not split -- group re-asks and assign reads for itself; and a name group's joins leave without a name word in front of the run, which assign reads again so it keeps one -- the same procedure as before #614, not always the same answer, the Freiherr garbage above being the measured case. MEASURED 2026-10-07 against 0c54662b (#618's merge), py3.11, `nameparser.__file__` asserted on each side; #619, merged between, moved no reading. The two #617 harnesses as decisions.md#P3's 2026-10-06 entry records them (the fingerprint, 78,324 parses; the connective grid, 485,300), and an S2 stress grid built for this entry: the 22 heads `John`, `John Smith`, `John Q. Smith`, `Mary Ann Smith`, `John van Smith`, `Juan de la Vega`, `John van`, `anh van`, `Jan Freiherr von Berg`, `Freiherr von Berg`, `abdul salam`, `abdul rahman al-said`, `Juan y Garcia`, `Josep Carod i Rovira`, `Jane Doe nee van der Berg`, `Dr. John Smith`, `JOHN SMITH`, `john smith`, `Smith`, `de Mesnil`, `Kim Min Jun`, `Ortega y Gasset`, each followed by every one- and two-word tail over the 35 words `PhD MA Ma DO Do do vd VD mc Mc MD Jr Jr. Sr III V V. VI I X.Y.Z. XYZ Ph. D. Esq. Prof. Dr. Sir MBA RN G.J. Ed van and Jones Ed. ba` (`Ph. D.` one word of the list); then 25,000 draws seeded `random.seed(614)`, each `choice(heads)`, `randint(3, 5)` and that many `choice(words)`, joined by spaces; then every head before a comma with every two-word pair behind it, and `Smith, ` and ` , John` for every word; deduplicated, under four policies (default, both family-first orders, strict comma suffixes) -- 81,138 names, 324,552 parses, comparing the seven roles, every ambiguity's kind and detail, `initials()` and `capitalized()`. Fingerprint: 6 parses move, `John van der Berg Prof.` under six policies, and nothing else. Connective grid: 1,248 move, no corpus or case name among them: 1,212 the title class, and 36 garbage shapes holding a title or a credential run behind a leading particle (`de der Dr. MA PhD Vega y Lopez` reads suffix 'PhD Vega y Lopez', #602's run, where 0c54662b kept every word the family), 8 of them moving a report's grouping alone (below). Stress grid: 8,472, no corpus or case name; 48 move reports alone, 20 the suffix wording above and 28 the grouping of an absorbed word's report where the run the view stops unit-making at is not the one the peel reads (`anh van Do Sr DO Jones MBA` reports 'DO' and 'Jones' where 0c54662b reported 'DO Jones'); restricted to a head and one tail word under the default policy, 12 parses, the six particle-chain heads with `Dr.` and `Prof.`, all the title class. (At bfa798fb, before the second review, these were 30, 2,936 and 20,532, the slice 33.) Whole tail readings per corpus name (`tail_reading` calls, 1,536 names): 1,125 read once, 56 twice, none three times, against 1,095, 55 and 30 on 0c54662b; peel passes, which count H5's fixed point's own re-peels, 1,215 names at one against 583. Five differential gates exit 0 with the one move classified `fix(#614)` at every baseline. COST, `tools/perf/call_count.py` on each interpreter against adf6da88 (#619's merge): `parse` 317 -> 308 on 3.11 and 296 -> 290 on 3.12-3.15, `facade` 354 -> 345 and 333 -> 327 (decisions.md#parse-cost, this date). Per name on 3.11, warm: the Vega name 376 -> 368, and every segment of fewer than three pieces unchanged, never reaching the read (`John Smith` 152, `Smith, John` 161, `John Smith, PhD` 206). -- 2026-10-08 (Derek), #620 — S2'S READ COUNTS THE CHAIN'S UNITS OVER THE PIECES AS WRITTEN, RATHER THAN MERGING A COPY OF THEM. The #614 read above counted through a merged copy of the pieces, a second statement of P2's chain with its own stop list, and its reviews found three defects that were that copy's: an index applied across the two layouts (`Freiherr St van Berg MA`), a unit formed inside #602's run hiding a title (`John Smith MD van Secretary Jones`), and a shape-only numeral folded into a unit (`Juan de la Vega VI`), each patched in the copy. Now `_pieces._chain_units` marks each piece by the unit the chain will put it in -- OPENS, JOINED (a name word the chain joins to the run in front of it) or BOUND (a particle inside the run, a word the run took) -- asking `chain_run_end`, the chain's run end, which `_group`'s chain now calls too. The peel's words to spare (`second_unit`) and #602's run start count in those marks and the peel stops at any piece the chain joins into a run, which is a name word whichever mark it carries; everything else reads the pieces as written, so nothing is hidden, and the copy, its view index and its credential-run gate are gone. +- 2026-10-08 (Derek), #620 — S2'S READ COUNTS THE CHAIN'S UNITS OVER THE PIECES AS WRITTEN, RATHER THAN MERGING A COPY OF THEM. The #614 read above counted through a merged copy of the pieces, a second statement of P2's chain with its own stop list, and its reviews found three defects that were that copy's: an index applied across the two layouts (`Freiherr St van Berg MA`), a unit formed inside #602's run hiding a title (`John Smith MD van Secretary Jones`), and a shape-only numeral folded into a unit (`Juan de la Vega VI`), each patched in the copy. Now `_pieces._chain_units` marks each piece by the unit the chain will put it in -- OPENS, JOINED (a name word the chain joins to the run in front of it) or BOUND (a particle inside the run, a word the run took; retired by #625, this section, 2026-10-10, group merging the run first) -- asking `chain_run_end`, the chain's run end, which `_group`'s chain now calls too. The peel's words to spare (`second_unit`) and #602's run start count in those marks and the peel stops at any piece the chain joins into a run, which is a name word whichever mark it carries; everything else reads the pieces as written, so nothing is hidden, and the copy, its view index and its credential-run gate are gone. WHAT THE COUNT KEEPS. A word the reading weighs -- a suffix word of either kind, the ambiguous class, a word in it by shape, an initial, a roman numeral by shape, a period-marked title word (`_weighed`) -- keeps a unit of its own though the chain may join it, as a position always counted it; only the plain words behind a particle run are JOINED, and a unit with nothing past its opener moves no name start into the titles (SUPERSEDED IN PART by #624, decisions.md#P2, 2026-10-10: no unit opens inside the titles, so the gate is gone). The first prototype marked every word the chain reaches JOINED, by the chain's own rule alone, and the review found it wrong in two shapes the earlier count read right: `Freiherr von Berg MA X.Y.Z.` kept 'MA X.Y.Z.' in the family (a member the chain would join counted as no word in front of the next), and `St St VI` and `Freiherr St MA` read title 'St St', 'Freiherr St' with the numeral or acronym the whole family (a unit of words the read then takes, which the chain never builds). Both masters read all three as they read now. (#624 since reads `St St VI` and `Freiherr St MA` as title 'St St' / 'Freiherr St' again, by a different route: the chain and the count now share one leading position, the last title-particle, so neither builds the unit -- decisions.md#P2, 2026-10-10.) The copy's exclusion list had been carrying the count positions always kept -- a member in front counts -- and the marks keep it without the copy. The numeral fork still reads the piece in front of the numeral as written -- reading the unit's first word broke rules.md#P2's `John van der J. V` -- so `John van B and Smith X` reads family 'van B and Smith X' again, the reading #614's copy had moved. - ALSO COUNTED IN UNITS: #602's run start, a particle opening a chain run counting once and its bound particle not at all (`John der la Jr. Prof. Smith` keeps its run, `van la Smith Secretary Jr. Dr. Smith` starts none, as on master). And `credential_run`, on the read's path, gathers adjacent particles into the one absorbed piece the chain makes of them (`Dr. John Smith Esq. RN Mc Mc` reports 'Mc Mc'); M2's clause-free view reports word by word as it always has. + ALSO COUNTED IN UNITS: #602's run start, a particle opening a chain run counting once and its bound particle not at all (`John der la Jr. Prof. Smith` keeps its run, `van la Smith Secretary Jr. Dr. Smith` starts none, as on master). And `credential_run`, on the read's path, gathers adjacent particles into the one absorbed piece the chain makes of them (`Dr. John Smith Esq. RN Mc Mc` reports 'Mc Mc'; SUPERSEDED by #625, this section, 2026-10-10: group merges the run before the read, which absorbs it as one piece); M2's clause-free view reports word by word as it always has. DECLINED, measured: a units test in `tail_reading`'s resume condition. A splice lowers a count only by taking out titles that open units of their own, and with fewer than two units in front of it every piece there is in the first piece's chain run, which the titles behind it join; fuzzed over 372,330 names, the test changed nothing. CLASSIFICATION, found writing this entry's rows: eight case rows #614 added as `parity` did not match 1.4.0 -- they had been compared against master -- and are relabelled with the change each reading came from, bisected (#424, #289, #296, #516, #602 twice; R2 and P2 for two that read so since the 2.0 pipeline). MEASURED 2026-10-08 against 4278693d (#621's merge), py3.11, `nameparser.__file__` asserted, with the #614 harnesses and recipe above: fingerprint 0 moves; connective grid 0; S2 stress grid 12, three garbage names under four policies, none in the head-plus-one-word slice -- `Freiherr von Berg Dr. X.Y.Z. VD` reads title 'Freiherr Dr.', suffix 'X.Y.Z. VD' (H5's trailing title, the class #614 moved; master family 'von Berg Dr. X.Y.Z.'), `Jan Freiherr von Berg VD V and ba I` reads family 'VD V and ba I' as adf6da88 did (master suffix 'I'), and `de Mesnil Ma do mc and Ph. D.` reports the 'do' the peel took beside adf6da88's 'Ma'. Five gates exit 0. Frames unchanged, 308 on 3.11 and 289 on 3.12-3.15 (decisions.md#parse-cost). - /SIMPLIFY (2026-10-08), measured. Each pass of `tail_reading` had rescanned for the second unit from the front, a C-level quadratic no frame guard saw: `Freiherr von` + `Berg `*n + `MA Dr. `*n read 5.77x for 4x the input against master's 4.08x; the lookup is now made once per read and handed to every pass (4.03x), and `test_the_second_unit_is_found_once_per_read` counts it. `_weighed` is pinned to the reading by `test_a_word_the_read_takes_is_never_one_the_count_joined` (no piece marked JOINED is ever in the run the read takes; 320 grid names fail with the test answering False). The gathering of an absorbed particle run reads the BOUND marks, and #602's run start one condition. Grids byte-identical, frames unchanged. Left for a follow-up, each moving readings: merging the tail's particle runs in group before the read, which would retire the BOUND stop and the gathering (80 grid names move, `Dr. mc mc`, and the chain's particle report needs re-keying); and the name start decided once, the item below. + /SIMPLIFY (2026-10-08), measured. Each pass of `tail_reading` had rescanned for the second unit from the front, a C-level quadratic no frame guard saw: `Freiherr von` + `Berg `*n + `MA Dr. `*n read 5.77x for 4x the input against master's 4.08x; the lookup is now made once per read and handed to every pass (4.03x), and `test_the_second_unit_is_found_once_per_read` counts it. `_weighed` is pinned to the reading by `test_a_word_the_read_takes_is_never_one_the_count_joined` (no piece marked JOINED is ever in the run the read takes; 320 grid names fail with the test answering False). The gathering of an absorbed particle run reads the BOUND marks, and #602's run start one condition. Grids byte-identical, frames unchanged. Left for a follow-up, each moving readings: merging the tail's particle runs in group before the read, which would retire the BOUND stop and the gathering (80 grid names move, `Dr. mc mc`, and the chain's particle report needs re-keying; done by #625, this section, 2026-10-10, where #624 had already moved those 80); and the name start decided once, the item below. LEFT OPEN, by Derek's scoping: reading a family comma's given part the same way, and deciding the name's start once for a title-particle head (#620's related paths, which change readings). The second was settled by #624 (decisions.md#P2, 2026-10-10). - 2026-10-10 (Derek), #627 — A LONE PARTICLE WITH THE RUN RIGHT BEHIND IT COUNTS AS A NAME WORD. #602's run start needs two name words before the credential, and "a lone particle is not a name word for the count", for `de Mesnil`, where the particle and the word behind it are one surname. Since #620 the count is taken in the chain's units, so a particle that opens a unit counts once (`van der`, `von Berg`) and a P3-joined phrase is one piece (`von und zu`); the lone particle stayed uncounted only where nothing joins it, which is exactly where a credential stands behind it. So `John von PhD Jones` read middle 'von PhD', family 'Jones' -- the chain later taking the credential as the particle's surname -- while `John van der PhD Jones` and `John von und zu PhD Jones` read family 'van der' / 'von und zu', suffix 'PhD Jones', and the comma spelling `Smith, John von PhD Jones` already read suffix 'PhD Jones'. The issue's question was which way to make them agree. DECIDED (Derek): the particle stops at the credential, so it is a name word of its own and the rest is suffix: `John von PhD Jones` → family 'von', suffix 'PhD Jones'. The input is malformed (a particle is not written before a post-nominal), so the decision is whichever reading the shared count already gives the other forms, not a special case for this one: the existing clause gains its own limit, a lone particle counting when the run starts right behind it, in `run_start`'s one loop. The caller without the chain's units (`trailing_start`) shares the expression so the two cannot disagree, but no reading shows that half: it changes the position it returns for `Smith, John von PhD Jones` (2 where master gave 4), so group's chain takes its stopped path and asks `trailing_start` again, reaching the same reading at 16 more frames (339 → 355 on py3.11; a rare shape no frame guard sees, decided for the shared expression over a per-caller copy); a mutant ignoring the flag there passed the suite with the grid byte-identical (#627's reviews). The 1.4.0 and 2.3.0 wheels read middle 'von PhD', as they read `von und zu PhD`; the joined forms changed this cycle with #602. Declined: counting every unit, a lone particle included. It moved 2,836 S2-grid parses rather than 2,432, the extra ones all `de Mesnil `, where it counted the surname as two words, and it split the two callers: a leading `van` counted before another particle (`van la Smith Jr. Jones` → suffix 'Jr. Jones') but not before a name word (`van Smith Jr. Jones`), and the case row `van la Smith Secretary Jr. Dr. Smith` caught it. MEASURED 2026-10-10 against 195c4081 (#630's merge), py3.11, `nameparser.__file__` asserted on each side. S2 stress grid (the #614 entry's recipe above, 324,552 parses): 2,432 role moves, no report-only move; every move is a lone particle -- in that grid `van`, or a dual `do`/`mc`/`vd` in any case -- with an unambiguous credential right behind it becoming the family (under the family-first orders the given name, the family staying first), the run taking the rest. The grid's comma shapes do not move; the review found the same move in shapes the grid lacks, as intended: other particles (`von`, `de`, `du`, `la`, `bin`), and a head before a suffix comma (`John von PhD Jones, Jr.` → family 'von', suffix 'PhD Jones, Jr.'). Fingerprint of every `tools/differential/corpus*.jsonl` name and `tests/v2/cases.py` text at 195c4081 under the three orders (1,925 texts, the seven fields and every report): 0 moves; this change adds one text, its own example, which moves. The five differential gates exit 0. Held by the case row `John von PhD Jones` (`fix(#627)`) against the row `van la Smith Secretary Jr. Dr. Smith`, whose lone particle has a name word behind it and still does not count, and the row `Freiherr von vd PhD Jones`, where the particle standing right before the credential is one the run took and the flag is not asked of it (a mutant asking it of every piece passed the suite until that row). +- 2026-10-10 (Derek), #625 — THE TAIL'S PARTICLE RUNS ARE MERGED IN GROUP BEFORE S2'S READ, SO THE READ NEEDS NO BOUND MARK AND NO GATHERING. Since #614 split the trailing run off ahead of the chain, the read sees particles the chain has not merged yet, and #620 re-created the chain's particle-to-particle merge in the reader: `_chain_units` marked a particle standing straight behind another BOUND, the peel stopped at it, and `credential_run` gathered an absorbed run of them into one report. Now group merges each run of two or more adjacent particles past the chain's leading position (`_pieces.merge_particle_runs`, at the trailing site, after P3's joins and before the read) into the piece the chain would open its run with, so the read counts, stops at and absorbs the run as one piece, and the BOUND mark, the peel's stop for it and the gathering are gone. The merged piece keeps `prefix` and drops `title`: a P3 join carries the title tag in from a word in both vocabularies (`St und Berg`), and on the merged run it read the run as a leading title. The title run and the lead are read once, before the merge, and handed to the read, which had asked both again. The chain's PARTICLE_OR_GIVEN report keyed on its run claiming something past the particle (`j > k + 1`), which a merged run has done already though the chain then claims nothing more (`Freiherr von der`); it now keys on that or on the run being one the merge made (`premerged`) -- not on the piece holding several tokens, which a P3-joined phrase does too, and which moved 11,358 reports on the grid below (`Dr. St Do und`). The rule text is unchanged: rules.md#S2 already counts a particle run, the duals straight behind a particle included, as the one name word the chain makes. What grows is code, not concepts: the read loses a mark and two branches and re-creates nothing of the chain, group gains the merge (about 45 more lines than it removes across `_pieces.py` and `_group.py`, docstrings included). + DECIDED (Derek): retiring the two reader-side mechanisms is worth the moves -- every one is garbage input, now read consistently, and none is a released reading restored. + MEASURED 2026-10-10 against 505d6dec (#632's merge), py3.11, `nameparser.__file__` asserted on each side. A particle-run grid -- `{head} {run} {tail}` for every head in (none), `Dr.`, `Freiherr`, `St`, `John`, `Dr. John`, `Freiherr von`, `John Smith`, `Sir`, `Freiherr St`; every run of one to three words over `van der de la von und mc Mc vd do Do St du`; every tail in (none), `Berg`, `MA`, `PhD`, `Ma`, `Jr.`, `Berg MA`, `Prof.`, `Berg Ph. D.`, `V`, `VI`, `X.Y.Z.`, `née Smith`, `Berg Jr. PhD`, `Berg van`, `, Jr.`, `, John` (a comma tail written against the run) -- 395,148 texts, 1,185,444 parses under the three orders, comparing the seven fields and every report (kind, detail, tokens): 774 role moves (258 texts under each order), no report-only move. Every move holds `St und`, the joined piece carrying the title tag, which master merged into the particle run in front with the tag kept and read as a title (`Freiherr Do St und Berg MA` → title 'Freiherr Do St und Berg', family 'MA'); now title 'Freiherr', family 'Do St und Berg', suffix 'MA', the name part read alike with and without the trailing 'MA'. Not a released reading: 1.4.0 reads first 'Do St und Berg', last 'MA', and 2.3.0 family 'Do St und Berg MA', and over the 258 moved texts under the default order master's seven fields equal 2.3.0's on 152, the tree's on none. The S2 stress grid (#614's recipe, above): no role move, and one text's report under its four policies -- `de Mesnil Ma do mc and Ph. D.` no longer reports 'do': P3 joins `mc and Ph. D.` into one suffix-tagged piece, the merge folds 'do' into it, and the peel takes the piece whole rather than picking 'do', the report adf6da88 did not make either and #620 added (its entry above). The #624 title-particle head grid (decisions.md#P2, 2026-10-10): 0. Every corpus name and case text, and #624's comma-tail grid, under the three orders: 0. The five differential gates exit 0, their reports identical line for line to the tree before. The 2026-10-08 prototype's 80 moves (`Dr. mc mc`) are not in the population: #624 already reads `Dr. Mc Mc` as H3 states. Recorded negative controls (2026-10-10): with the merge off, 30 tests in tests/v2 fail (32 over the whole suite, measured at the review's fix), `test_an_absorbed_particle_run_reports_as_the_one_word_the_chain_makes` among them; with the report keyed on `j > k + 1` alone, the case rows `a_particle_takes_the_dual_behind_it_at_the_head` and `a_particle_run_before_the_run_stays_one_name_word`; with the merge starting at the segment's second piece rather than past the lead, 13; with `title` kept, only `test_a_merged_particle_run_is_a_particle_and_no_title`, the unit test added for it. Cost: +3 frames on the reference row, which holds `de la` (decisions.md#parse-cost, this date). + THE REVIEW (2026-10-10) found the change reaching further, every shape garbage, none given a row (Derek): no move goes the other way, and nothing moved over comma shapes (an empty or nickname-only head, a family comma's given part), connective-separated runs, runs beside a maiden marker or holding a nickname, bound given names, one-case names, or caller lexicons adding title or suffix words as particles. (1) The `St und` class is any P3 join whose left word is in both the title and particle vocabularies, `und`, `and` or `of` (`y` moves nothing: master already reads `Freiherr Do St y Berg MA` as now), a caller's title-particle included (`Freiherr Do Freiherr and Berg MA`: master title 'Freiherr Do Freiherr and Berg', family 'MA'; now title 'Freiherr', family 'Do Freiherr and Berg', suffix 'MA'). (2) A title-particle head, a particle run, then a title and a dotted word (`Freiherr van der Dr. G.J. MA`, 88 texts on its grid under each order): the run now counts as one name word, so 'G.J.' has none to spare, the read takes nothing and group's own chain takes 'Dr. G.J.' -- family 'van der Dr. G.J.', suffix 'MA', as the one-particle `Freiherr van Dr. G.J. MA` reads (master title 'Freiherr Dr.', family 'van der', suffix 'G.J. MA'). (3) H3's give-back behind a leading title where every word is suffix vocabulary, the last two a run of duals (`Dr. vd Jr. mc mc`, about 1,800 parses over two grids): the merged run is a particle piece and no suffix word, so the title stays -- title 'Dr.', family 'vd', suffix 'Jr. mc mc', where master gave 'Dr.' back as the given name with 'vd' still standing; `Dr. vd Jr. mc` still gives it back. The review also found two runs in one name unreached by any test (an off-by-one in the merge's bookkeeping passed the suite), now the rows `two_particle_runs_merge_each` and `a_second_run_of_duals_is_the_family` (parity) and a half of the merge's unit test, which also holds that a lone particle is no merged run; and the read's title run taken before the merge held only by the frame bands, now `test_the_read_takes_the_title_run_as_written_before_the_merge` (`Dr. mc Jr mc mc`, which with the run asked again after the merge, by the read or by group, reads family 'mc Jr mc mc'; the fix's own review found the group-side half unheld by a read-level test, so it pins the parse too). ### indic-honorifics — the renunciate class and the Indic honorific vocabulary (2026-09-06, #346/#344/#343) @@ -1614,6 +1618,7 @@ Every number below is a py3.11 measurement of 2026-08-31, recomputable with `uv - 2026-09-26 #546 -- the stages copy their state with `_state.copy_with` instead of `dataclasses.replace`, and every row moves DOWN. `dataclasses.replace` walks `fields()` and calls `__init__` on each copy, three frames on 3.11 and 3.12 and four from 3.13 where `copy_with` is one. One parse of the reference name makes 18 copies: six of the state, one each from tokenize, segment, classify, group, assign and post_rules, and twelve of single tokens, six each from classify and assign (`extract_delimited` returns the state unchanged when there is no delimiter, and `script_segment` returns early on ASCII input). That is the whole of the drop, 18 x 2 = 36 frames on 3.11 and 3.12 and 18 x 3 = 54 from 3.13 (recompute: wrap `copy_with` with a counter in the eight stage modules and parse the reference name). A field copy builds what `replace` builds for a dataclass that is decorated itself rather than inheriting the decoration, keeps the generated `__init__`, and has no `__post_init__` and no `init=False` field. `_copyable_fields` checks exactly those four, and `_COPY_FIELDS` runs it over `WorkToken`, `PendingAmbiguity` and `ParseState` at import, so a class that stops qualifying fails there and `copy_with` copies nothing else. `test_the_guard_refuses_a_class_a_field_copy_would_get_wrong` records, for each refused shape, what `replace` builds and what an unguarded copy would build instead. To mypy, `copy_with` is `from dataclasses import replace as copy_with`, which keeps the dataclass plugin's keyword and type checks at every call site; an assignment (`copy_with = dataclasses.replace`) would not, since the plugin keys on the callee's full name (measured: a misspelled field and a wrong-typed value both pass through the assignment and both fail through the import). Measured 2026-09-26 with `uv run python tools/perf/call_count.py --against e0f1a2f`, each row on its own interpreter, parse/facade: 3.11 406/443 → 370/407, 3.12 384/421 → 348/385, 3.13, 3.14 and 3.15 402/439 → 348/385. The rows drop by 40 and 58 rather than 36 and 54 because e0f1a2f already read 4 under every row, inside the band, and the new rows are set to what the harness reads now. `_LINK_BASELINE`'s 64-link clause reads 2587 → 2301 on 3.11. By stage (`--stages`, py3.11, ms per 1000 parses of the reference name): group 19.6 → 17.3, classify 13.4 → 9.8, assign 11.1 → 8.0, tokenize 8.1 → 6.6, post_rules 7.9 → 6.5, segment 2.6 → 1.7. BEHAVIOR IDENTICAL: the differential gate's report at all five baselines matches e0f1a2f's line for line apart from the path header. - 2026-09-29 #561 -- SUPERSEDED IN PART: the 2026-09-26 bullet above, whose four conditions ("decorated itself", "keeps the generated `__init__`", no `__post_init__`, no `init=False` field) were not sufficient. `replace` builds through `obj.__class__(...)`, and a field copy skips every hook that call runs; four shapes met all four conditions and still parted: a validating `__new__`, a validating `__setattr__` on a non-frozen dataclass and a validating metaclass `__call__`, each giving `ValueError` from `replace` where an unguarded field copy built `value=-1`, and an `InitVar` with no default, which `replace` demands (a `ValueError` through 3.12, a `TypeError` from 3.13) and a field copy never sees. #561's first tightening (80554266) found only the first three, and by swapping the class's OWN `__dataclass_params__` for an inherited lookup it also let through an `__init__` borrowed from another dataclass under that dataclass's name, which the 2026-09-26 guard had refused; review caught both. So the guard is now a list of conditions a class must MEET (`_state._copy_refusals`, one reason string per unmet condition): decorated as a dataclass itself, frozen, an `__init__` generated for it (compiled from `` and named `.__init__`) that takes exactly its `init` fields, the default metaclass, no `__new__` above `object`, no `__post_init__`, no `init=False` field. It is a TRIPWIRE for a realistic edit to the three pipeline classes and says so, not a proof against any class: a forged `__qualname__` on a borrowed `__init__` in a class that is also decorated itself, or a `__class__` property, still passes, and `copy_with` copies nothing but `WorkToken`, `PendingAmbiguity` and `ParseState`, which meet every condition on 3.11 to 3.15. `_UNGUARDED_EFFECT` records for each row what `replace` builds, what an unguarded copy builds, and which conditions refuse it; `test_every_guard_condition_alone_refuses_a_recorded_shape` asserts that each condition is the only refusal of some row, so dropping one fails a test rather than contradicting this paragraph (checked by removing the `__new__` condition in a scratch copy: two failures). Call counts unchanged, the guard running at import only. - 2026-10-10 #624 -- every row moves DOWN by 4: `parse` 308 → 304 on 3.11 and 289 → 285 on 3.12-3.15, `facade` 345 → 341 and 326 → 322, each measured on its own interpreter with `tools/perf/call_count.py` in isolated copies of 8e6e524a (#631's merge) and of the change, the import asserted (8e6e524a read exactly the recorded rows). `--modules` on 3.11: `_pieces.py` 51 → 49 and `_group.py` 25 → 23 -- group's leading-position scan, a generator asking `is_leading_title` per piece, gave way to `leading_titles`, which it already called, and `chain_lead`, which asks `is_prefix_piece` only of the titles. +- 2026-10-10 #625 -- every row moves UP by 3: `parse` 304 → 307 on 3.11 and 285 → 288 on 3.12-3.15, `facade` 341 → 344 and 322 → 325, each measured on its own interpreter with `tools/perf/call_count.py` in isolated copies of 505d6dec (#632's merge) and of the change, the import asserted (505d6dec read exactly the recorded rows). `--modules` on 3.11: `_pieces.py` 49 → 52, `_group.py` 23 → 23. The reference name holds a particle run (`de la`), so it pays the merge's call and one more `merge_pieces` and `joined_tags`, the chain then merging `de la` with `Vega` where it had merged the three words at once; the first cut asked the title run and the lead in the merge and again in the read, +11, and group now asks the title run once and hands both on. Every name of three or more words read at the trailing site pays the call, +1 (`John Q. Smith` 202 → 203 on 3.11, warm); a name with a particle run pays the 3 (`Juan de la Vega` 230 → 233). - 2026-10-06 #617 -- every row moves DOWN by 10: `parse` 374 → 364 on 3.11 and 353 → 343 on 3.12-3.15, `facade` 411 → 401 and 390 → 380, each measured on its own interpreter with `PYTHONPATH= pythonX.Y tools/perf/call_count.py` against 83f4e914 (every row had sat inside its band before, the 3.12-3.15 ones at +5 over a 348 baseline). `--modules` puts all ten in group and the piece layer (`_group.py` + `_pieces.py` 142 → 132, every other module unchanged): P3's connective joins moved out of `_group_segment` into `_pieces.join_connectives`, which asks `is_conj_piece`, `is_title_piece` and `is_prefix_piece` directly where the loops had asked through one-line closures over them, two frames per test becoming one. The link pin moves with it, 2,678 → 2,479 for the 64-link name on 3.11 (`_LINK_BASELINE`), and there for two reasons, counted per function in #617's review: the closures (`conj` 133 frames, `title` and `prefix` one each) and the single-letter test, which had joined each connective's text through a generator (65 frames, one per link, more than the pin's whole band), less the one `join_connectives` frame. A deliberate move, not drift: the change was a refactor (decisions.md#P3, 2026-10-06). The reference rows hold no comma; a comma part holding a connective now pays the shared loop (`John Smith, Mr. and Mrs.` 262 → 269 on 3.11, warm), and one holding none skips it and costs what it did (`Smith, John` 161), the review having found the first cut charging every family-comma parse two frames per word of the part. THEN 39 MORE on every row, in the branch's /simplify pass: group ran the rootname count (about five frames a piece, its only reader the carve-out) and the shared loop over every segment of three or more pieces, connective or not, and over the corpora and the case table under three orders 0 of the 3,153 group calls with no connective left free changed a piece. Group now skips both where its frozen-set walk, which already reads every piece's `conjunction` tag, finds no connective free, as `_comma._reading_pieces` does: `parse` 364 → 325 on 3.11 and 343 → 304 on 3.12-3.15, `facade` 401 → 362 and 380 → 341, `_group.py` + `_pieces.py` 132 → 93, measured on each interpreter as above; the link pin does not move, every link in it joining. Per name on 3.11, warm: the Vega name 436 → 386, `Smith, John Quincy Adams Bob` 326 → 297, `Smith, John` and `John Smith, Mr. and Mrs.` unchanged at 161 and 269. No output moved over the #617 grids (decisions.md#P3). - 2026-10-07 -- every row moves DOWN by 8 more: `_group_segment`'s last three one-line predicate closures, `prefix` over `is_prefix_piece`, `suffix` over `is_suffix_piece` and `marker` over `_is_maiden_marker_piece`, are inlined as direct calls, finishing what #617 did for `title`, `conj` and `merge`. Each closure was a frame of its own ahead of the predicate's, and the chain's two inner scans ask `prefix` once per piece they pass, so the cost grows with particle pieces (frame-budget question 1 in AGENTS.md). Measured with `PYTHONSAFEPATH=1 PYTHONPATH= pythonX.Y tools/perf/call_count.py`, each row on its own interpreter, parse/facade against 0c54662b: 3.11 325/362 → 317/354, 3.12, 3.13, 3.14 and 3.15 304/341 → 296/333. `--modules` on 3.11 puts all eight in `_group.py` (34 → 26), every other module unchanged. The whole of the drop is the closures, counted per function on 0c54662b (a profile hook counting `call` events whose code is named `prefix`, `suffix` or `marker` in `_group.py`): the reference name asked `prefix` 7 times and `suffix` once, `Juan` + `de la Vega ` × 8 + `Smith` asked 40 and 9 and reads 937 → 888 on 3.11 (919 → 870 on 3.12-3.15), and the 64-link name asked `prefix` once, so `_LINK_BASELINE` moves 2,479 → 2,478 on 3.11 and gains rows for 3.12-3.15 at 2,459 (2,460 on 0c54662b), each measured on its own interpreter and the first the pin has had beyond 3.11; `marker` is reached only by P5's bound-join decline and the reference name never asks it. BEHAVIOR IDENTICAL: the full suite passes, and the differential gate's report at all five baselines matches 0c54662b's line for line apart from the path header. - 2026-10-08 #620 -- no row moves: `parse` 308 on 3.11 and 289 on 3.12-3.15, `facade` 345 and 326, measured as the #614 bullet below on 4278693d (#621's merge) and on the change alike. The 3.12-3.15 pin read 290/327 until this date: #614's /simplify took one frame off those interpreters (289/326 on 4278693d, measured 2026-10-08) and the pin stayed, inside its band, so the #614 bullet's 290 is the pin's number, not that tree's. `--modules` on 3.11 unchanged. The merged copy and its per-word exclusion test are gone, and `_chain_units` asks `chain_run_end` per particle unit and `_weighed` per word behind one; the chain itself calls `chain_run_end` with the particle test inline where it had called `is_prefix_piece` per piece, which pays for the rest. decisions.md#S2, this date. diff --git a/docs/design/mechanisms.md b/docs/design/mechanisms.md index 8bb1a215..3ad4f6a2 100644 --- a/docs/design/mechanisms.md +++ b/docs/design/mechanisms.md @@ -29,7 +29,7 @@ Problem shape. A rule needs to know how words were JOINED (chained titles, parti ## UNIT-PARTITION — count the units the joining rules built -Problem shape. A rule counts "one name word", but the input holds words that another rule has already joined into one name — and the joining structure it would read has been merged away. Contract statement. Three rules build multi-word units: a particle chain (P2), a conjunction join (P3), and a bound given-name pair (P5). A rule that counts name words counts those units, and takes each whole or not at all — with one stated exception, rules.md#C1's count of the words before a comma, which counts particle surnames alone (decisions.md#C1, #575). How it works. group builds each join as a piece, but its own prefix chain then merges the joined piece into a longer one, so PIECES no longer carries the boundary — the units are rebuilt from the tags the vocabulary layer left (`particle`, `conjunction`, `vocab:bound-given`). The rebuild is RECURSIVE: what a conjunction or a bound word joins is the next UNIT, not the next word, and absorbing a single token instead strands a particle severed from the words it chains. Note the two joins arrive here for opposite reasons — the conjunction join was built and then swallowed, while the bound-given join was never built at all (P5 joins only at the first non-title piece), so restoring piece boundaries in group would fix the first and silently split the second. Lives in. nameparser/_pipeline/_vocab.py (`unit_ends`, the walk, shared since #575) and its two readers: _post_rules.py (`_units`, over classify's tags, for P1's fold) and rules.md#C1's counts before a comma: `_comma._whole_name`, over classify's tags (`surname_unit_facts`) since #613 moved the decision after classify, and `_vocab.name_word_count`, which builds its facts from the vocabulary (`surname_unit_tags`, held to classify's tags by an agreement test) -- a view script_segment needed while it asked a hand copy of C1 before classify ran, until #613's /simplify had it ask `_comma.decide` itself on a classified copy. A fourth walk is S2's read (#614): `_pieces._chain_units` marks each piece by the unit P2's chain WILL put it in -- opening a unit, a name word the chain joins, or a particle inside the run -- asking the chain's own `chain_run_end`, which `_group`'s chain calls too; a word the reading weighs (`_weighed`) keeps a unit of its own, as a position always counted it. The trailing peel's words to spare and #602's run start count in those marks, the peel stops at a particle inside a run, and everything else -- the numeral fork's look at the word in front, the H5 chain, the run's title test -- reads the pieces as written. It runs before the chain merges anything, which is why it cannot read the merged tags `unit_ends` walks; #614 had rebuilt the chain as a merged COPY instead, whose three defects in one review -- words the reading weighs hidden inside a unit, indices drifting from the pieces' -- were each patched in that copy, and #620 replaced it with the marks (decisions.md#S2). C1 counts particle chains only, reaching one word as the fold does: a bound pair builds a given name, and whether P3 joins a connective depends on the whole name the count is helping to decide, so neither is one surname before a comma (`_vocab.SURNAME_UNIT_TAGS`, decisions.md#C1). rules.md P1 and C1 are the counting rules, P2/P3/P5 the joining ones. Reach for it when. A rule says "one name word" and the input can contain a join — enumerate the joining rules out of rules.md rather than the ones you remember. +Problem shape. A rule counts "one name word", but the input holds words that another rule has already joined into one name — and the joining structure it would read has been merged away. Contract statement. Three rules build multi-word units: a particle chain (P2), a conjunction join (P3), and a bound given-name pair (P5). A rule that counts name words counts those units, and takes each whole or not at all — with one stated exception, rules.md#C1's count of the words before a comma, which counts particle surnames alone (decisions.md#C1, #575). How it works. group builds each join as a piece, but its own prefix chain then merges the joined piece into a longer one, so PIECES no longer carries the boundary — the units are rebuilt from the tags the vocabulary layer left (`particle`, `conjunction`, `vocab:bound-given`). The rebuild is RECURSIVE: what a conjunction or a bound word joins is the next UNIT, not the next word, and absorbing a single token instead strands a particle severed from the words it chains. Note the two joins arrive here for opposite reasons — the conjunction join was built and then swallowed, while the bound-given join was never built at all (P5 joins only at the first non-title piece), so restoring piece boundaries in group would fix the first and silently split the second. Lives in. nameparser/_pipeline/_vocab.py (`unit_ends`, the walk, shared since #575) and its two readers: _post_rules.py (`_units`, over classify's tags, for P1's fold) and rules.md#C1's counts before a comma: `_comma._whole_name`, over classify's tags (`surname_unit_facts`) since #613 moved the decision after classify, and `_vocab.name_word_count`, which builds its facts from the vocabulary (`surname_unit_tags`, held to classify's tags by an agreement test) -- a view script_segment needed while it asked a hand copy of C1 before classify ran, until #613's /simplify had it ask `_comma.decide` itself on a classified copy. A fourth walk is S2's read (#614): `_pieces._chain_units` marks each piece by the unit P2's chain WILL put it in -- opening a unit, or a name word the chain joins -- asking the chain's own `chain_run_end`, which `_group`'s chain calls too; a word the reading weighs (`_weighed`) keeps a unit of its own, as a position always counted it. The particles straight behind each other are merged ahead of the read, into the piece the chain would open its run with (`_pieces.merge_particle_runs`, #625), so a particle run is one piece the read takes whole or not at all, and no mark stands for a particle inside it. The trailing peel's words to spare and #602's run start count in those marks, the peel stops at a word the chain joins, and everything else -- the numeral fork's look at the word in front, the H5 chain, the run's title test -- reads the pieces as written. It runs before the chain joins name words to a particle run, which is why it cannot read the merged tags `unit_ends` walks; #614 had rebuilt the chain as a merged COPY instead, whose three defects in one review -- words the reading weighs hidden inside a unit, indices drifting from the pieces' -- were each patched in that copy, and #620 replaced it with the marks (decisions.md#S2). C1 counts particle chains only, reaching one word as the fold does: a bound pair builds a given name, and whether P3 joins a connective depends on the whole name the count is helping to decide, so neither is one surname before a comma (`_vocab.SURNAME_UNIT_TAGS`, decisions.md#C1). rules.md P1 and C1 are the counting rules, P2/P3/P5 the joining ones. Reach for it when. A rule says "one name word" and the input can contain a join — enumerate the joining rules out of rules.md rather than the ones you remember. ## MARK-DONT-STRIP — record the decision, keep the fact diff --git a/nameparser/_pipeline/_group.py b/nameparser/_pipeline/_group.py index 145343e5..5729aa56 100644 --- a/nameparser/_pipeline/_group.py +++ b/nameparser/_pipeline/_group.py @@ -65,8 +65,8 @@ from nameparser._pipeline._pieces import ( Peel, TailRead, chain_lead, chain_run_end, is_conj_piece, is_leading_title, is_prefix_piece, is_suffix_piece, is_title_piece, join_connectives, joined_tags, - leading_titles, merge_pieces, peel_walk, read_trailing_run, - tail_reading, trailing_candidates, trailing_start, + leading_titles, merge_particle_runs, merge_pieces, peel_walk, + read_trailing_run, tail_reading, trailing_candidates, trailing_start, trailing_start_past_titles, ) from nameparser._pipeline._state import ( @@ -785,8 +785,17 @@ def _group_segment(seg: tuple[int, ...], additional: int, # part past a comma and a segment of fewer than three pieces -- # which no join reaches anyway -- read as before, and so does a # run holding a name piece (_pieces.read_trailing_run). + # The particles straight behind each other are merged first, + # into the one piece the chain would open its run with, so the + # read counts and absorbs the run as that piece (#625); + # `premerged` keeps which runs the merge made, for the chain's + # report below. + premerged: Set[int] = frozenset() if site is ClauseSite.TRAILING: - found = read_trailing_run(pieces, ptags, tokens, one_case) + n = leading_titles(pieces, ptags, tokens) + premerged, lead = merge_particle_runs(pieces, ptags, tokens, n) + found = read_trailing_run(pieces, ptags, tokens, one_case, n, + lead) if found is not None: read, start = found tail_pieces, tail_ptags = pieces[start:], ptags[start:] @@ -964,10 +973,13 @@ def chain(tail: int) -> None: # piece -- nothing # was chained, and _assign reports that case instead. # Without this the two emitters both fire on the same token. + # A run merged ahead of the read (`premerged`, #625) has + # claimed the particles behind it already, so its chain + # is a decision though `j == k + 1` ('Freiherr von der'). # (Tag test first: it is a set lookup and almost no name # has an ambiguous particle, while is_leading_title is a # call per piece.) - if (j > k + 1 + if ((j > k + 1 or pieces[k][0] in premerged) and "vocab:particle-ambiguous" in tokens[pieces[k][0]].tags): while titled < k and is_leading_title( diff --git a/nameparser/_pipeline/_pieces.py b/nameparser/_pipeline/_pieces.py index 180537ea..8f24f7ec 100644 --- a/nameparser/_pipeline/_pieces.py +++ b/nameparser/_pipeline/_pieces.py @@ -1055,9 +1055,9 @@ def credential_at_the_given_slot( return anchored is not None and anchored() -#: `_chain_units`' marks: a piece opening a name unit, a name word the -#: chain joins to the run in front of it, a particle inside that run -OPENS, JOINED, BOUND = 1, 0, -1 +#: `_chain_units`' marks: a piece opening a name unit, and a name word +#: the chain joins to the run in front of it +OPENS, JOINED = 1, 0 def second_unit(rest: Sequence[int], units: Sequence[int]) -> int: @@ -1106,9 +1106,8 @@ def peel_trailing(rest: Sequence[int], pieces: Sequence[Sequence[int]], `units`, where given, marks each piece by the unit the chain will put it in (`_chain_units`, #620): the words to spare are counted in units, so a particle chain counts as the one name word P2 makes of - it, and a piece the chain joins into a run -- a particle inside - it, or a word the reading does not weigh -- is a name word the - walk stops at. The numeral fork still reads the piece in front of the + it, and a word the chain joins to a run without the reading + weighing it is a name word the walk stops at. The numeral fork still reads the piece in front of the numeral as written: an initial there keeps it a name word whatever unit the initial is in ('John van der J. V', rules.md#P2). """ @@ -1119,8 +1118,8 @@ def peel_trailing(rest: Sequence[int], pieces: Sequence[Sequence[int]], k = len(rest) if start is None else start while k > 0: piece = pieces[rest[k - 1]] - # a piece the chain joins into a run is a name word: a particle - # inside the run, or a word the reading does not weigh + # a word the chain joins to a run, and the reading does not + # weigh, is a name word if units is not None and units[rest[k - 1]] != OPENS: break if is_suffix_piece(piece, ptags[rest[k - 1]], tokens): @@ -1244,7 +1243,8 @@ def run_start(rest: Sequence[int], names: int, `units` (`_chain_units`, #620) counts over pieces the chain has not joined yet: a word the chain will join to the run in front of it is no name word of its own, and a particle opening such a run is - the one the run makes ('der la', one surname).""" + the one the run makes ('der la', merged into one piece ahead of + the read, one surname).""" if names < 3: return names core = 0 @@ -1344,8 +1344,7 @@ def has_name_content(piece: Sequence[int], def credential_run(rest: Sequence[int], peel: Peel, p: int, pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], - tokens: Sequence[WorkToken], - units: Sequence[int] | None = None) -> Peel: + tokens: Sequence[WorkToken]) -> Peel: """`peel` with its name count cut back to `p`, where #602's run starts (`run_start`, which the caller asks first so that a name with no run pays one frame, not two). @@ -1359,30 +1358,7 @@ def credential_run(rest: Sequence[int], peel: Peel, p: int, a run, leaves both lists without the title test's frame.""" run_titles: list[int] = [] absorbed: list[tuple[int, ...]] = [] - r = p + 1 - while r < peel.names: - q = rest[r] - # Two or more particles in a row are the one piece P2's chain - # makes of them wherever they stand, a dual among them ('van - # mc', 'Mc Mc') no post-nominal of its own: where the chain has - # run that piece is here already, and S2's read (#620) sees the - # pieces before it runs, so it gathers the particles its - # `units` mark BOUND behind the one that opens them. M2's - # clause-free view reads words before any join and reports them - # word by word, as it always has. - e = r + 1 - while (units is not None and e < peel.names - and units[rest[e]] == BOUND): - e += 1 - if e > r + 1: - run: list[int] = [] - for x in rest[r:e]: - run.extend(pieces[x]) - if has_name_content(run, tokens): - absorbed.append(tuple(run)) - r = e - continue - r += 1 + for q in rest[p + 1:peel.names]: if is_suffix_piece(pieces[q], ptags[q], tokens): continue if is_title_piece(pieces[q], ptags[q], tokens): @@ -1640,7 +1616,7 @@ def tail_reading(rest: list[int], pieces: Sequence[Sequence[int]], if kept == peeled.names: p = run_start(rest, peeled.names, pieces, ptags, tokens, units) return rest, (), (peeled if p == peeled.names else credential_run( - rest, peeled, p, pieces, ptags, tokens, units)) + rest, peeled, p, pieces, ptags, tokens)) # `rest[:hi]` stands as written; `behind` holds the peeled runs # the splices left after it, and `titled` the chained ones, each # back to front -- a pass's run goes in FRONT of what the pass @@ -1691,7 +1667,7 @@ def tail_reading(rest: list[int], pieces: Sequence[Sequence[int]], p = run_start(rest, final.names, pieces, ptags, tokens, units) return (rest, tuple(j for run in reversed(titled) for j in run), final if p == final.names else credential_run( - rest, final, p, pieces, ptags, tokens, units)) + rest, final, p, pieces, ptags, tokens)) class TailRead(NamedTuple): @@ -1743,26 +1719,78 @@ def chain_run_end(k: int, pieces: Sequence[Sequence[int]], return j +# rules.md#P2: "A particle joins the words after it into one name part" +# -- the particles straight behind it first, merged ahead of S2's read +# so the read counts and absorbs the run as the one piece the chain +# makes of it (#625) +def merge_particle_runs(pieces: list[list[int]], ptags: list[set[str]], + tokens: Sequence[WorkToken], + n: int) -> tuple[set[int], int | None]: + """Merge each run of two or more adjacent particles past the + chain's leading position into one piece, the one `chain_run_end` + would open its run with ('van der', 'van mc'). The merged piece is + a particle still (`prefix`), and no title: a P3 join can carry the + tag in from a word in both vocabularies ('St und Do'), and the + merged run would then read as a leading title. + + `n` is the end of the leading title run (`leading_titles`), read + before anything merges. Returns the first token of each merged + piece -- the chain's PARTICLE_OR_GIVEN report keys on a run that + claimed something past its particle, which a merged run has though + `chain_run_end` then claims nothing more -- and the leading + position (`chain_lead`) where it asked it, else None. + + The particle test is `is_prefix_piece`'s definition written + inline, and the leading position is asked only where two particles + stand together past the first piece (AGENTS.md's frame budget, + question 1).""" + prefix: list[bool] = [] + for k in range(len(pieces)): + prefix.append("prefix" in ptags[k] + or len(pieces[k]) == 1 + and "particle" in tokens[pieces[k][0]].tags) + merged: set[int] = set() + k = 1 + while k + 1 < len(prefix) and not (prefix[k] and prefix[k + 1]): + k += 1 + if k + 1 >= len(prefix): + return merged, None + lead = chain_lead(pieces, ptags, tokens, n) + k = lead + 1 + while k < len(prefix): + if prefix[k]: + q = k + 1 + while q < len(prefix) and prefix[q]: + q += 1 + if q > k + 1: + merged.add(pieces[k][0]) + merge_pieces(pieces, ptags, k, q, add={"prefix"}, + drop={"title"}) + del prefix[k + 1:q] + k += 1 + return merged, lead + + def _chain_units(pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], tokens: Sequence[WorkToken], - n: int) -> list[int] | None: + n: int, lead: int | None = None) -> list[int] | None: """The name units P2's chain will make of `pieces`, as a mark per - piece -- OPENS where a piece opens a unit, JOINED where the chain - will join a name word to the run in front of it and the reading - does not weigh the word (`_weighed`, which keeps a word of its own - what a position always counted as one), and BOUND for a - particle inside the run, a word the run took and so a name word - whatever else it is ('van mc', 'von vd': rules.md#S2, the words - both particles and suffix vocabulary standing straight behind a - particle). + piece -- OPENS where a piece opens a unit, and JOINED where the + chain will join a name word to the run in front of it and the + reading does not weigh the word (`_weighed`, which keeps a word of + its own what a position always counted as one). The particles + straight behind each other are one piece already + (`merge_particle_runs`, #625), so a particle run is one unit and a + dual inside it ('van mc', 'von vd') no word the read can take. `n` is the end of the leading title run, and the chain starts past its leading position (`chain_lead`, the one answer group's chain - takes too, #624), which opens no unit inside the titles: the name - starts at `n`. + takes too, #624; `lead` where the caller has asked it already), + which opens no unit inside the titles: the name starts at `n`. - Nothing is merged: the read counts with the flags and reads every - piece as written, so no word it weighs is hidden inside a unit + Nothing else is merged: the read counts with the flags and reads + every other piece as written, so no word it weighs is hidden + inside a unit (#620; the merged copy #614 read through hid a title inside a credential run and a shape-only numeral, and its indices drifted from the pieces'). None where no particle stands past the first @@ -1775,14 +1803,12 @@ def _chain_units(pieces: Sequence[Sequence[int]], if not any(prefix[1:]): return None units = [OPENS] * count - k = chain_lead(pieces, ptags, tokens, n) + 1 + k = (chain_lead(pieces, ptags, tokens, n) if lead is None + else lead) + 1 while k < count: if prefix[k]: j = chain_run_end(k, pieces, ptags, tokens, count) q = k + 1 - while q < j and prefix[q]: - units[q] = BOUND - q += 1 while q < j and not _weighed(pieces[q], tokens): units[q] = JOINED q += 1 @@ -1818,7 +1844,8 @@ def _weighed(piece: Sequence[int], tokens: Sequence[WorkToken]) -> bool: def read_trailing_run(pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], tokens: Sequence[WorkToken], - one_case: bool | None, + one_case: bool | None, n: int, + lead: int | None = None, ) -> tuple[TailRead, int] | None: """`tail_reading` read once over a segment's pieces as group holds them after rules.md#P3's connective joins and before the particle @@ -1831,11 +1858,20 @@ def read_trailing_run(pieces: Sequence[Sequence[int]], ('John Dr. G.J.'), the one shape whose run holds a name piece -- there group joins as before #614 and assign reads for itself. A segment the read takes nothing from reads as a run of no pieces, at - the end.""" - n = leading_titles(pieces, ptags, tokens) + the end. + + `n` is the end of the leading title run (`leading_titles`) and + `lead` the chain's leading position where the caller has asked it, + both read before `merge_particle_runs` merged the pieces: the merge + starts past the lead and drops the title tag, so it moves the lead + never, and the title run only where H3 had given a title back + before a run of particles that are suffix vocabulary, which the + read takes as written: 'Dr. mc Jr mc mc' reads family 'mc', suffix + 'Jr mc mc', where the run read after the merge would leave 'mc Jr + mc mc' in the family.""" if n == len(pieces): return None - units = _chain_units(pieces, ptags, tokens, n) + units = _chain_units(pieces, ptags, tokens, n, lead) rest, titled, peel = tail_reading(peel_walk(n, ptags), pieces, ptags, tokens, one_case, units) tail: set[int] = set() diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 85f9aaeb..bf9f13cd 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -9867,9 +9867,30 @@ def _check_cjk_shape_purity(self) -> None: "a title. #614's read merged 'van Secretary' into one " "piece and the title went into the suffix (its second " "review); since #620 the read counts the chain's units " - "and merges nothing, so the run's title test sees every " - "word as written (1.4.0 middle 'Smith MD', family 'van " + "rather than reading a merged copy, and only " + "particle runs are merged ahead of it (#625), so the " + "run's title test sees 'Secretary' as written (1.4.0 middle 'Smith MD', family 'van " "Secretary Jones'; bisected to #602's run)"), + Case("two_particle_runs_merge_each", + "John van der Smith de la Berg", + {"given": "John", "middle": "van der Smith", + "family": "de la Berg"}, + classification="parity", + notes="each particle run joins the words after it, the final " + "group the family (rules.md#P2); group merges both runs " + "ahead of S2's read (#625), whose bookkeeping across the " + "first merge no other row reached (#625's review: an " + "off-by-one there read family 'van der Smith de la " + "Berg', the q + 1 one)"), + Case("a_second_run_of_duals_is_the_family", + "Anna van der Berg van mc", + {"given": "Anna", "middle": "van der Berg", + "family": "van mc"}, + classification="parity", + notes="'van mc' is one particle run, the dual a particle the " + "run took and no suffix (rules.md#P2, #S2); merged " + "ahead of the read after a first run (#625's review: an " + "off-by-one left 'mc' a suffix)"), Case("a_second_particle_run_is_a_unit_of_its_own", "Jan van Berg de Mc", {"given": "Jan", "middle": "van Berg", "family": "de Mc"}, diff --git a/tests/v2/pipeline/test_group.py b/tests/v2/pipeline/test_group.py index 129287f6..8e4a05b1 100644 --- a/tests/v2/pipeline/test_group.py +++ b/tests/v2/pipeline/test_group.py @@ -1981,21 +1981,25 @@ def test_an_absorbed_particle_run_reports_as_the_one_word_the_chain_makes( ) -> None: """A run of particles a credential run absorbs is reported as the one piece P2's chain makes of it, a dual among them included, as it was - before #614 split the run off ahead of the chain (#620). Only - particles that stand together: the merged 'Ph. D.' between two - is outside the walk and still parts them. M2's clause-free view - reads its words before any join and reports them one by one. - Recorded negative controls, #620 (2026-10-08): with the gathering - off 'Mc Mc' went unreported; with it on in M2's view the clause - reported 'mc van'; without the adjacency test 'vd van'.""" + before #614 split the run off ahead of the chain (#620): group + merges the particles straight behind each other before the read + (#625), which had gathered them out of `_chain_units`' marks + instead. Only particles that stand together: the merged 'Ph. D.' + between two parts them. M2's clause-free view reads its words + before any join and reports them one by one. Recorded negative + controls, #620 (2026-10-08): with the gathering off 'Mc Mc' went + unreported; with it on in M2's view the clause reported 'mc van'; + without the adjacency test 'vd van'. #625 (2026-10-10): with the + merge off this fails, the read taking the pieces one by one.""" def reported(text: str) -> list[list[str]]: return [[t.text for t in a.tokens] for a in parse(text).ambiguities if a.kind is AmbiguityKind.SUFFIX_OR_NAME] assert reported("Dr. John Smith Esq. RN Mc Mc") == [["Mc", "Mc"]] assert reported("Mary Ann Smith Dr. MD vd Ph. D. van") == [["van"]] assert reported("Jane Doe nee van der Berg Jr mc van") == [["van"]] - # a P3-joined phrase is a particle too, and the gathering runs on - # every pass of the read, after the H5 chain's splice as before it + # a P3-joined phrase is a particle too, and the merged run is one + # piece on every pass of the read, after the H5 chain's splice as + # before it assert reported("John Smith Jr. von und zu Mc") == [ ["von", "und", "zu", "Mc"]] assert reported("John Smith PhD van mc Dr.") == [["van", "mc"]] diff --git a/tests/v2/pipeline/test_pieces.py b/tests/v2/pipeline/test_pieces.py index 1f6998aa..ae2c063f 100644 --- a/tests/v2/pipeline/test_pieces.py +++ b/tests/v2/pipeline/test_pieces.py @@ -22,9 +22,10 @@ from nameparser._pipeline._pieces import ( _anchors, _numeral_behind_the_initial_veto, anchor_in_reach, credential_anchors, - BOUND, JOINED, OPENS, TailRead, _chain_units, + JOINED, OPENS, TailRead, _chain_units, credential_at_the_given_slot, is_leading_title, - join_connectives, leading_titles, read_trailing_run, + join_connectives, leading_titles, merge_particle_runs, + read_trailing_run, own_words, peel_trailing, peel_walk, segment_suffix_reading, tail_reading, trailing_candidates, trailing_titles, ) @@ -1071,12 +1072,26 @@ def test_only_groups_joins_derive_a_prefix() -> None: def _read(text: str) -> tuple[ParseState, tuple[TailRead, int] | None]: """S2's trailing read (#614) over a whole name's words as classify - tagged them, one piece each.""" + tagged them, one piece each, the particle runs merged as group + merges them ahead of the read (#625).""" state = _state_through("classify", text) + tokens = list(state.tokens) + pieces = [[i] for i in state.segments[0]] + ptags: list[set[str]] = [set() for _ in pieces] + n = leading_titles(pieces, ptags, tokens) + _merged, lead = merge_particle_runs(pieces, ptags, tokens, n) + return state, read_trailing_run(pieces, ptags, tokens, state.one_case, + n, lead) + + +def _merged_pieces(state: ParseState) -> tuple[list[list[int]], + list[set[str]]]: pieces = [[i] for i in state.segments[0]] ptags: list[set[str]] = [set() for _ in pieces] - return state, read_trailing_run(pieces, ptags, list(state.tokens), - state.one_case) + tokens = list(state.tokens) + merge_particle_runs(pieces, ptags, tokens, + leading_titles(pieces, ptags, tokens)) + return pieces, ptags def _texts(state: ParseState, indices: Set[int]) -> list[str]: @@ -1117,25 +1132,25 @@ def test_the_trailing_read_declines_where_there_is_nothing_to_split() -> None: def _units(text: str) -> list[tuple[str, int]]: - """`_chain_units` over a whole name's words, one piece each, as - (word, mark) pairs.""" + """`_chain_units` over a whole name's pieces as group holds them at + the read, as (piece text, mark) pairs.""" state = _state_through("classify", text) - pieces = [[i] for i in state.segments[0]] - ptags: list[set[str]] = [set() for _ in pieces] + pieces, ptags = _merged_pieces(state) tokens = list(state.tokens) units = _chain_units(pieces, ptags, tokens, leading_titles(pieces, ptags, tokens)) assert units is not None - return [(tokens[p[0]].text, u) for p, u in zip(pieces, units)] + return [(" ".join(tokens[i].text for i in p), u) + for p, u in zip(pieces, units)] def test_the_read_marks_the_units_the_chain_will_make() -> None: - """#620: a particle opens a unit, a name word the chain joins to it - is JOINED, and a particle inside the run is BOUND -- a word the run - took, which the peel does not weigh ('van mc'); a suffix piece ends - the run and stays a unit of its own.""" + """#620: a particle opens a unit, and a name word the chain joins to + it is JOINED; a suffix piece ends the run and stays a unit of its + own. The particles behind it are in its piece already, a dual among + them no word of its own ('van mc', #625).""" marks = _units("John van mc Berg PhD") - assert marks == [("John", OPENS), ("van", OPENS), ("mc", BOUND), + assert marks == [("John", OPENS), ("van mc", OPENS), ("Berg", JOINED), ("PhD", OPENS)] @@ -1150,6 +1165,75 @@ def test_the_chain_opens_no_unit_inside_the_titles() -> None: assert [u for _, u in marks] == [OPENS, OPENS, OPENS] +def test_a_merged_particle_run_is_a_particle_and_no_title() -> None: + """#625: the particles straight behind each other past the chain's + lead merge into one piece that keeps `prefix` and drops `title`. + P3's join carries the title tag in from a word in both vocabularies + ('St und Berg'), and kept on the merged run it read the run as a + leading title: 'Freiherr Do St und Berg MA' read title 'Freiherr Do + St und Berg' with the tag kept, family 'Do St und Berg' without it. + The particle leading the name stays out of the merge (P4), so 'de + la van Vega' merges 'la van' alone; two runs in one name merge + each, and a lone particle is no run ('John van Smith de la + Berg').""" + state = _state_through("classify", "Freiherr Do St und Berg MA") + tokens = list(state.tokens) + pieces = [[i] for i in state.segments[0]] + ptags: list[set[str]] = [set() for _ in pieces] + join_connectives(pieces, ptags, tokens, letter_stays=False, + titles_only=False) + assert "title" in ptags[2] + merged, lead = merge_particle_runs(pieces, ptags, tokens, + leading_titles(pieces, ptags, tokens)) + assert lead == 0 + assert [" ".join(tokens[i].text for i in p) for p in pieces] \ + == ["Freiherr", "Do St und Berg", "MA"] + assert "prefix" in ptags[1] and "title" not in ptags[1] + assert merged == {pieces[1][0]} + state = _state_through("classify", "de la van Vega") + tokens = list(state.tokens) + pieces = [[i] for i in state.segments[0]] + ptags = [set() for _ in pieces] + merged, lead = merge_particle_runs(pieces, ptags, tokens, 0) + assert lead == 0 + assert [" ".join(tokens[i].text for i in p) for p in pieces] \ + == ["de", "la van", "Vega"] + for text, want in ( + ("John van der Smith de la Berg", + ["John", "van der", "Smith", "de la", "Berg"]), + ("John van Smith de la Berg", + ["John", "van", "Smith", "de la", "Berg"])): + state = _state_through("classify", text) + tokens = list(state.tokens) + pieces = [[i] for i in state.segments[0]] + ptags = [set() for _ in pieces] + merged, _lead = merge_particle_runs(pieces, ptags, tokens, 0) + assert [" ".join(tokens[i].text for i in p) for p in pieces] \ + == want + assert merged == {p[0] for p in pieces if len(p) > 1} + + +def test_the_read_takes_the_title_run_as_written_before_the_merge( +) -> None: + """#625: group hands the read the title run read before the merge. + Merging 'mc mc' makes it no suffix piece, which would undo H3's + give-back of 'Dr.' and move the name's start; the read keeps the + run as written, so the credential run behind the first 'mc' stays + the read's. Recorded negative controls (2026-10-10): with the read + asking `leading_titles` again after the merge, the read takes + nothing here and the parse reads family 'mc Jr mc mc'; with group + handing the read the run asked after the merge, the read helper + below cannot see it and the parse assertion fails the same way.""" + state, found = _read("Dr. mc Jr mc mc") + assert found is not None + read, start = found + assert start == 2 + assert _texts(state, read.tail) == ["Jr", "mc", "mc"] + name = parse("Dr. mc Jr mc mc") + assert (name.title, name.family, name.suffix) == ("Dr.", "mc", + "Jr mc mc") + + _WEIGHED_GRID_HEADS = ("John van", "Freiherr von", "anh van", "Jan de la", "John van Berg", "Freiherr von Berg") _WEIGHED_GRID_WORDS = ("MA", "Ma", "ma", "DO", "do", "vd", "mc", "V", "VI", @@ -1168,8 +1252,7 @@ def _weighed_grid_violations() -> list[str]: state, found = _read(text) if found is None: continue - pieces = [[i] for i in state.segments[0]] - ptags: list[set[str]] = [set() for _ in pieces] + pieces, ptags = _merged_pieces(state) tokens = list(state.tokens) units = _chain_units(pieces, ptags, tokens, leading_titles(pieces, ptags, tokens)) diff --git a/tests/v2/test_benchmark.py b/tests/v2/test_benchmark.py index ab59d64d..c368da5e 100644 --- a/tests/v2/test_benchmark.py +++ b/tests/v2/test_benchmark.py @@ -90,11 +90,11 @@ #: #614's /simplify on and were corrected in #620, which moved no row #: (decisions.md#parse-cost). _CALL_BASELINE = { - (3, 11): {"parse": 304, "facade": 341}, - (3, 12): {"parse": 285, "facade": 322}, - (3, 13): {"parse": 285, "facade": 322}, - (3, 14): {"parse": 285, "facade": 322}, - (3, 15): {"parse": 285, "facade": 322}, + (3, 11): {"parse": 307, "facade": 344}, + (3, 12): {"parse": 288, "facade": 325}, + (3, 13): {"parse": 288, "facade": 325}, + (3, 14): {"parse": 288, "facade": 325}, + (3, 15): {"parse": 288, "facade": 325}, } _BAND = 0.02 diff --git a/tests/v2/test_cases.py b/tests/v2/test_cases.py index 86dd55a2..4608fb01 100644 --- a/tests/v2/test_cases.py +++ b/tests/v2/test_cases.py @@ -177,8 +177,8 @@ def test_the_field_sweep_sees_the_claims_it_checks() -> None: _FIELD_CLAIM.search(a.detail) is not None for case, order in _SWEPT_ROWS for a in _swept(case.id, order).ambiguities) - assert claims == 1149, ( - f"the field sweep checks {claims} claims, recorded as 1149 on " + assert claims == 1150, ( + f"the field sweep checks {claims} claims, recorded as 1150 on " f"2026-10-10")