diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 33729fff..2d406dc4 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -103,6 +103,11 @@ Open: [#360](https://github.com/derek73/python-nameparser/issues/360) which part - 2026-08-22 #424 — the chain stops where assign's trailing peel begins. The chain ran "until a trailing suffix begins" and asked with the suffix-piece test, which vetoes a bare `V` as an initial — the #401 question at a third site — so `John van der Berg V` read family 'van der Berg V' (1.4.0 read it so too: shipped since 1.x), and `John van der Berg Ma` read family 'van der Berg Ma' where 1.4.0 read suffix 'Ma' — a 2.0 regression on S2's other fork. Both now stop where `_trailing_start` says the run begins: the S2 peel (`_peel_trailing` over `_peel_walk`, shared with assign and P5's reserve since #425) read over the pieces as they stand from the first name piece, once per segment, kept as a length from the end that the chain's merges ahead of it do not move (the benchmark's `particles` shape is the guard). The chain takes both forks — its merges leave the acronym at least the three pieces the fork counted, so assign peels it after as the walk did before — where the maiden walk takes only the numeral (the M2 entry says why). The fork reads the piece before the numeral as it stands, so `John van der J. V` keeps family 'van der J. V', as `John J. V` reads the V as a name; and a numeral with a suffix behind it is not last in the walk, so `John van der Berg V Jr` keeps family 'van der Berg V'. No differential corpus name had either shape; the rules examples carry them — the numeral against every baseline, the acronym against 2.0.0 and 2.1.0. Two more sites were asking assign's question with a test of their own, both found by the reviews of this change. The walk's starting point and P4's leading-particle scan asked group's `title()`, which does not see H2's unlisted abbreviations — assign peels those as titles all the same — so `Xyz. van Johnson` chained where `Dr. van Johnson` did not (#367 had keyed the exception on the first piece of the name, by group's test), and `Xyz. van Berg MA` read family 'van Berg MA' on every 2.x tree — the chain had taken the acronym, as it took every acronym; the draft that stopped before it left assign two pieces where the fork had counted three, and MA became the family. Both ask assign's own test now — `_is_leading_title` and the period-abbreviation pattern moved down into group, with `_leading_titles` as the one definition of where the name begins: the chain's walk, the bound join's first name piece (the second docs review found that one still on group's test — `Xyz. abdul John Smith` joined nothing where `Dr. abdul John Smith` read given 'abdul John') and, through `_is_leading_title`, the scan and the PARTICLE_OR_GIVEN report — so `Xyz. van …` reads as `Dr. van …` does. `Esq. van Gogh`, a corpus name, moves with it (1.4.0 read family 'van Gogh' for both, as #367's rule records; its own rule at every baseline, and one for the P4 example). A particle of the unambiguous suffix vocabulary too (vd, mc) is a suffix piece to the peel: where it opens the trailing run the chain stops before it as before any suffix word and the peel takes it — `John Smith Mc V` reads suffix 'Mc, V' where master and 1.4.0 read family 'Mc V', a class the second docs review found unrecorded (2,554 of its 117,306 constructed names, an eighteen-word pool that puts the words in every position) — and where it continues a prefix run, the run takes it as a particle, so `John van Mc` keeps family 'van Mc', every baseline's reading and the one P6 chose for the shape after a comma. Both are Accepted under P2. The fork's count was the last thing the chain asked on its own authority, and the reviews found it wrong behind a title-and-particle word: #367 stops the leading-particle scan at `Freiherr`, `St`, `Do` — the name's own first piece — where assign's title peel steps over them, so the chain takes the name's first word, and the acronym the fork counted with three pieces met assign with two: `Freiherr von Berg MA` read given 'von Berg', family 'MA' (1.4.0's reading, as it happens; master read family 'von Berg MA'). So the chain asks the peel again over the pieces it leaves — `chain(tail)`, then `_trailing_start` over what it built, and a second run without the stop where the verdict changed, the snapshot being one copy per segment with a trailing run — and takes what assign will not peel: family 'von Berg MA'. The numeral cannot flip, a chain group never being initial-shaped, so `Freiherr von Richthofen V` keeps suffix 'V', and its chain, the one name piece left, reads as `Dr. Smith V` reads — where 1.4.0 and master read family 'von Richthofen V'; the code review asked for that boundary to be on record, and P2 carries it. #410 (2026-08-25) then moved the name half of both: a title with one name word behind it names the family whatever stands beside it, so the pair reads family 'von Richthofen' and family 'Smith' today, where each read `given` with no family when this entry was written. The boundary is unchanged — the chain still stops before the numeral, and the two names still read alike; only the field the one name piece lands in moved. This name is the fifth corpus name #410 moves, and the only one already classified, its ledger rule's `fields` spanning the narrower diff. - 2026-10-02 (Derek), #573 — the part before a family comma: a particle that is suffix vocabulary too heads the name word behind it in any case. Decided with P6's third exception and S2's never-given lead; the record is the P6 section's entry of the same date. - 2026-10-07 #614 — CROSS-REFERENCE: the chain's stop is the trailing run S2 reads once, after P3's joins and before the chain is made, counting the chain's own run -- a particle run, the words both particles and suffix vocabulary behind it included, and the plain words after it -- as the one name word it will make; the re-ask after the chain's own merges (the 'Freiherr von Berg Ma' rollback above) runs only where the run was not read there. decisions.md#S2, this date; rules.md#P2's statement amended. +- 2026-10-10 (Derek), #624 — WHERE THE NAME STARTS IS DECIDED ONCE, AND OF TWO TITLE-PARTICLES THE SECOND LEADS. `TITLES ∩ particles_ambiguous` is `{freiherr, st}`. Such a word in the leading titles is the chain's leading position (P4), so the particle behind it is inside a name (`Freiherr von Berg` → family 'von Berg'). Group's chain found that position with a scan of its own that stopped at the FIRST such word, and chained the second as a surname particle, while assign read both as titles and the trailing read's unit count (`_chain_units`, #620) opened a unit inside the titles to follow the chain: `Freiherr St John Smith MA` read title 'Freiherr', family 'St John Smith', with a title-or-name report saying it was "read as a family name by convention", where `Dr. St John Smith` and `Sir St John Smith` read title 'Dr. St' / 'Sir St', given 'John', family 'Smith'. + DECIDED (Derek: consistency, one shared answer, no special case for garbage): `_pieces.chain_lead` gives the leading position once -- the LAST title in assign's own title run (`leading_titles`) that is also a particle, else the first piece past the run -- and group's chain and the read's count both take it. The run is read as WRITTEN, before H3's give-back (Derek, the review's option (a)): a title the run hands back to the name because only suffix words follow it is still a title to the chain, so the particle behind it is the name's leading piece and chains nothing (P4). `leading_titles` stops at a piece that is itself a leading title only where it gave that piece back or at the segment's last piece (a title needing a following piece), so `chain_lead` recognises the give-back from `n` alone, the last piece ruled out by its position, behind the same inline tag test, and no second title walk is needed. No particle stands between the lead and the run's end, so the chain opens no unit inside the titles: the read's `at` and its `q > k + 1` gate are gone, and group's chain skips every piece up to the lead (`k <= leading`, equivalent to the old `k == leading` wherever the lead was the first title-particle, since no particle stood before it). `Freiherr St John Smith MA` → title 'Freiherr St', given 'John', family 'Smith', suffix 'MA'; `Freiherr St van Berg MA` → title 'Freiherr St', family 'van Berg'. + NOT DONE: the issue also proposed retiring assign's re-read fallback (`read = None` where group's joins left no name word past the titles). Measured, it still does work this change does not touch: disabled, sixteen or seventeen more tests fail depending on how it is disabled, most of them as an IndexError, so it guards totality and not only a reading -- among them the case row `a_title_read_before_the_chain_can_strand_the_name` (`Freiherr von Berg Dr. and Ed. Prof.`, where a CONNECTIVE join makes the name a title), the locale non-interference tests and two case-repair property tests. It stays. + MEASURED 2026-10-10 against 8e6e524a (#631's merge), py3.11, `nameparser.__file__` asserted on each side. A title-particle head grid -- every sequence of one to four words over `Freiherr St Dr. Sir Xyz. von van de do Mc John Smith MA PhD VI Prof. née Jr. and`, plus `Smith, ` and `, Jr.` over `Freiherr St Dr. von John`, 137,810 names under the three orders, comparing the seven fields and every report: 5,964 role moves, no report-only move. 5,868 have two title-particles in the leading title run (a period-shaped word such as `Jr.` and a connective title such as `and Freiherr` counting as titles there), all now titles -- where H3 then gives the last back, the word behind it is the post-nominal H3 saw rather than a word that title chained (`Dr. Freiherr St Mc` → title 'Dr. Freiherr', family 'St', suffix 'Mc'; master family 'St Mc'). 96 are H3's give-back before a run of dual particles, read as H3 states (`Dr. Mc Mc` → given 'Dr.', suffix 'Mc Mc'; master given 'Dr.', middle 'Mc', family 'Mc', the read having counted the second 'Mc' as bound into a run the chain never built). The first draft read the give-back as the chain's lead and gave `Dr. Mc Mc` title 'Dr.', family 'Mc Mc', chaining the name's leading particle behind a plain title against P4, and with a caller lexicon (`esq` added as an ambiguous particle) `Dr. esq vd` chained 'esq' and reported it; the review caught both, and option (a) is what this entry records. A comma-tail grid, `{head} {w1 … wk}` for k from 1 to 3 over the heads `John Smith, Jr.,`, `Smith, John,`, `Smith, Jr.,`, `,,`, `Smith,,`, `John Smith, PhD, Jr.,` and the words `Dr. Dr Attorney`, `Secretary of State` (one word of the list), `Lt.Gov. Freiherr St von`, `van der` (one word), `Bart Jones Jr. PhD MA Ma X.Y.Z. nee VI V 'Bob' (Bob) Prof.`, `Mr. and Mrs.` (one word), `I y` -- 97,650 texts, three orders: 1,368 parses on 456 texts move, every one a part past the second comma led by two title-particles, now all titles as C2 reads a title word there (`,, Freiherr Freiherr Bart` → title 'Freiherr Freiherr', suffix 'Bart'; master title 'Freiherr', suffix 'Freiherr Bart'). The S2 stress grid (#614's recipe): 0 moves. Every corpus name and case text at 8e6e524a under the three orders: only the case rows #624 rewrites move. The five differential gates exit 0. Cost: every frame baseline moves DOWN by 4 (decisions.md#parse-cost, this date). + ROWS: five rows pinned the old decision and now read the new one (`fix(#624)`, each note giving the earlier reading): `Freiherr St van Berg MA`, `Freiherr Freiherr Prof do`, `St St VI`, `Freiherr St MA`, `Freiherr St Prof.`; `Freiherr St John Smith MA`, the issue's example, is added, and `Dr. Mc Mc` pins the give-back half (the review found that half unpinned, a revert passing the suite). The marker row that held group's particle-or-given emitter on `St St née` moves to `St van née`, the same structure with a chained particle that is not a title (parity), since `St St née` no longer reaches that emitter. ### P6 — the trailing orphan particle @@ -753,19 +758,19 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f THE FIRST DESIGN AND THE REVIEW THAT REVERSED IT. Derek decided first that the read WINS OVER EVERY JOIN, connectives included -- prototyped both ways in a scratch copy -- and that implementation passed the suite, the gates and the fingerprint. #614's PR review then found it broke what the corpus never tests: a salutation alone read its connective as the family name (`Mr. and Mrs.` gave title 'Mr. Mrs.', family 'and' -- the H5 chain took 'Mrs.' before P3 could join), a joined leading salutation counted as three words (`Mr. and Mrs. Jr` gave family 'and', `Mr. and Mrs. Do` a spurious report), and the read counted a particle chain's or a connective join's words one by one where rules.md#P3 says a joined part is one name word wherever another rule counts (`Freiherr von Berg X.Y.Z.` and `Juan y Garcia X.Y.Z.` took the acronym as a credential against every release, while `Freiherr Berg X.Y.Z.` kept it the family). DECIDED (Derek, 2026-10-07): connectives join first, and the read counts the chain's units -- which returns every one of those to master's reading, and most of the first design's 'tail beats connectives' class with them. A second-round finding the counting view itself made: folding the plain words after a particle into its unit swallowed a shape-only roman numeral (`Juan de la Vega VI` read family 'de la Vega VI'), the #610 class AGENTS.md's spelling sweep records, so a word the numeral fork reads by shape stays out of the unit. DECIDED BEFORE THE REVERSAL AND STANDING: (1) the read wins over the particle chain and the bound-given join; (2) the consequences below, accepted as a class. WHAT MOVES. One realistic reading: a trailing period-marked title behind a particle chain is the title, `John van der Berg Prof.` reading title 'Prof.', family 'van der Berg' (every release 1.4.0-2.3.0 and 0c54662b: family 'van der Berg Prof.'), which retires the cost rules.md#H5 recorded as accepted; the six heads of the stress grid below that end in a particle chain move the same way with `Dr.` or `Prof.` behind them. The garbage `Freiherr von Berg Dr. and Ed. Prof.` reads title 'Freiherr von Berg Dr. and Ed.', family 'Prof.' (0c54662b family 'von Berg Dr. and Ed. Prof.'): the title is read off before the chain, which then joins a piece carrying the joined title and reads as a leading title, so assign reads the name again. P5 no longer re-reads the joined view: `abdul J. V` and `Sir abdul J. V` read given 'abdul J.', family 'V' -- 1.4.0's reading restored; 2.0 and 2.1 read suffix 'V', 2.2 and 2.3 kept 'J.' a middle name -- the initial in front suppressing the numeral fork in the one reading taken, three name words standing and the join leaving two; and `Sheik abdul Ma` reports the pick the read weighed (decisions.md#P5, this date). Reports: a pick the particle chain takes into the name is reported by the chain, in its own words, and group takes it out of the read's picks so assign does not report it again (#614's /simplify: assign first skipped a pick it found already reported, inferring group's decision from group's output), so `Jan vd Ma`, `John de Ma`, `John van der Berg Ma` and `Nguyen Van Lac` report as on 0c54662b. The second review found why it must be the chain: assign words a pick by the role it holds before post_rules, and under a leading title-particle that is the given role the fold then moves, so with assign reporting `Freiherr von Berg Ma` told the caller 'read as a given name' of a word in the family. The chain also reports a member the read never weighed, one it takes short of the run, its scan ending at the next particle (`Freiherr von Berg Ed Do` reports 'Ed'). One report moves for the better: a pick the run keeps is reported as the suffix it reads as, `Juan y Garcia Mc Do MD van Ed` reporting 'Ed' 'read as a suffix' where 0c54662b said the particle chain took it into the name, of a word in the suffix. - THE SECOND REVIEW (bfa798fb, 2026-10-07) found four defects in the counting view and the report hand-off, every one on a shape the corpus, the gates and the first stress grid held no instance of, and each now has a case row that dies without its fix. (1) The read started at the end of the leading titles counted in PIECES, applied to the view: a title that is also a particle opens a unit inside the titles, so the index landed past the name and `Freiherr St van Berg MA` read family 'St van Berg MA' (0c54662b suffix 'MA'); the view now returns where the name starts in its own indices -- the first unit opened at or before the titles' end, since a unit opened inside them makes a name of them even where it ends short (`Freiherr Freiherr Prof do` lost its report on 'do' until that second half). (2) Assign reported the chain's picks, in the wrong field's words, above. (3) The view made units INSIDE #602's credential run, which hid a title word from the run's title test: `John Smith MD van Secretary Jones` lost title 'Secretary' to the suffix. Past a run starter a unit is now the bare particle run, as the chain's own merge leaves it. (4) A `title` piece-tag exclusion in `_counted_as_one` made garbage differ from 0c54662b (`Freiherr von Mr. and Mrs. X.Y.Z.`) and guarded nothing (3) did not; removed rather than pinned. + THE SECOND REVIEW (bfa798fb, 2026-10-07) found four defects in the counting view and the report hand-off, every one on a shape the corpus, the gates and the first stress grid held no instance of, and each now has a case row that dies without its fix. (1) The read started at the end of the leading titles counted in PIECES, applied to the view: a title that is also a particle opens a unit inside the titles, so the index landed past the name and `Freiherr St van Berg MA` read family 'St van Berg MA' (0c54662b suffix 'MA'); the view now returns where the name starts in its own indices -- the first unit opened at or before the titles' end, since a unit opened inside them makes a name of them (SUPERSEDED IN PART by #624, decisions.md#P2, 2026-10-10: the chain now opens no unit inside the titles) even where it ends short (`Freiherr Freiherr Prof do` lost its report on 'do' until that second half). (2) Assign reported the chain's picks, in the wrong field's words, above. (3) The view made units INSIDE #602's credential run, which hid a title word from the run's title test: `John Smith MD van Secretary Jones` lost title 'Secretary' to the suffix. Past a run starter a unit is now the bare particle run, as the chain's own merge leaves it. (4) A `title` piece-tag exclusion in `_counted_as_one` made garbage differ from 0c54662b (`Freiherr von Mr. and Mrs. X.Y.Z.`) and guarded nothing (3) did not; removed rather than pinned. NOT THE SAME READ AS M2'S: the maiden take reads its own clause-free view of the words before any join, word by word, as rules.md#M2 states, and its Accepted clause already covers what that does where a particle chain would take the run; S2's read counts the chain's units after P3's joins. Two reads of one name answer different questions -- where the clause ends, where the name's run starts -- and are kept apart rather than made one here. RESIDUALS, the paths that read as before #614: a segment of fewer than three pieces (no join reaches it); a run with a name piece inside it (an H5 title in front of a name word the peel declined, `John Dr. G.J.`, `John van V Dr. V`), which is not split -- group re-asks and assign reads for itself; and a name group's joins leave without a name word in front of the run, which assign reads again so it keeps one -- the same procedure as before #614, not always the same answer, the Freiherr garbage above being the measured case. MEASURED 2026-10-07 against 0c54662b (#618's merge), py3.11, `nameparser.__file__` asserted on each side; #619, merged between, moved no reading. The two #617 harnesses as decisions.md#P3's 2026-10-06 entry records them (the fingerprint, 78,324 parses; the connective grid, 485,300), and an S2 stress grid built for this entry: the 22 heads `John`, `John Smith`, `John Q. Smith`, `Mary Ann Smith`, `John van Smith`, `Juan de la Vega`, `John van`, `anh van`, `Jan Freiherr von Berg`, `Freiherr von Berg`, `abdul salam`, `abdul rahman al-said`, `Juan y Garcia`, `Josep Carod i Rovira`, `Jane Doe nee van der Berg`, `Dr. John Smith`, `JOHN SMITH`, `john smith`, `Smith`, `de Mesnil`, `Kim Min Jun`, `Ortega y Gasset`, each followed by every one- and two-word tail over the 35 words `PhD MA Ma DO Do do vd VD mc Mc MD Jr Jr. Sr III V V. VI I X.Y.Z. XYZ Ph. D. Esq. Prof. Dr. Sir MBA RN G.J. Ed van and Jones Ed. ba` (`Ph. D.` one word of the list); then 25,000 draws seeded `random.seed(614)`, each `choice(heads)`, `randint(3, 5)` and that many `choice(words)`, joined by spaces; then every head before a comma with every two-word pair behind it, and `Smith, ` and ` , John` for every word; deduplicated, under four policies (default, both family-first orders, strict comma suffixes) -- 81,138 names, 324,552 parses, comparing the seven roles, every ambiguity's kind and detail, `initials()` and `capitalized()`. Fingerprint: 6 parses move, `John van der Berg Prof.` under six policies, and nothing else. Connective grid: 1,248 move, no corpus or case name among them: 1,212 the title class, and 36 garbage shapes holding a title or a credential run behind a leading particle (`de der Dr. MA PhD Vega y Lopez` reads suffix 'PhD Vega y Lopez', #602's run, where 0c54662b kept every word the family), 8 of them moving a report's grouping alone (below). Stress grid: 8,472, no corpus or case name; 48 move reports alone, 20 the suffix wording above and 28 the grouping of an absorbed word's report where the run the view stops unit-making at is not the one the peel reads (`anh van Do Sr DO Jones MBA` reports 'DO' and 'Jones' where 0c54662b reported 'DO Jones'); restricted to a head and one tail word under the default policy, 12 parses, the six particle-chain heads with `Dr.` and `Prof.`, all the title class. (At bfa798fb, before the second review, these were 30, 2,936 and 20,532, the slice 33.) Whole tail readings per corpus name (`tail_reading` calls, 1,536 names): 1,125 read once, 56 twice, none three times, against 1,095, 55 and 30 on 0c54662b; peel passes, which count H5's fixed point's own re-peels, 1,215 names at one against 583. Five differential gates exit 0 with the one move classified `fix(#614)` at every baseline. COST, `tools/perf/call_count.py` on each interpreter against adf6da88 (#619's merge): `parse` 317 -> 308 on 3.11 and 296 -> 290 on 3.12-3.15, `facade` 354 -> 345 and 333 -> 327 (decisions.md#parse-cost, this date). Per name on 3.11, warm: the Vega name 376 -> 368, and every segment of fewer than three pieces unchanged, never reaching the read (`John Smith` 152, `Smith, John` 161, `John Smith, PhD` 206). - 2026-10-08 (Derek), #620 — S2'S READ COUNTS THE CHAIN'S UNITS OVER THE PIECES AS WRITTEN, RATHER THAN MERGING A COPY OF THEM. The #614 read above counted through a merged copy of the pieces, a second statement of P2's chain with its own stop list, and its reviews found three defects that were that copy's: an index applied across the two layouts (`Freiherr St van Berg MA`), a unit formed inside #602's run hiding a title (`John Smith MD van Secretary Jones`), and a shape-only numeral folded into a unit (`Juan de la Vega VI`), each patched in the copy. Now `_pieces._chain_units` marks each piece by the unit the chain will put it in -- OPENS, JOINED (a name word the chain joins to the run in front of it) or BOUND (a particle inside the run, a word the run took) -- asking `chain_run_end`, the chain's run end, which `_group`'s chain now calls too. The peel's words to spare (`second_unit`) and #602's run start count in those marks and the peel stops at any piece the chain joins into a run, which is a name word whichever mark it carries; everything else reads the pieces as written, so nothing is hidden, and the copy, its view index and its credential-run gate are gone. - WHAT THE COUNT KEEPS. A word the reading weighs -- a suffix word of either kind, the ambiguous class, a word in it by shape, an initial, a roman numeral by shape, a period-marked title word (`_weighed`) -- keeps a unit of its own though the chain may join it, as a position always counted it; only the plain words behind a particle run are JOINED, and a unit with nothing past its opener moves no name start into the titles. The first prototype marked every word the chain reaches JOINED, by the chain's own rule alone, and the review found it wrong in two shapes the earlier count read right: `Freiherr von Berg MA X.Y.Z.` kept 'MA X.Y.Z.' in the family (a member the chain would join counted as no word in front of the next), and `St St VI` and `Freiherr St MA` read title 'St St', 'Freiherr St' with the numeral or acronym the whole family (a unit of words the read then takes, which the chain never builds). Both masters read all three as they read now. The copy's exclusion list had been carrying the count positions always kept -- a member in front counts -- and the marks keep it without the copy. The numeral fork still reads the piece in front of the numeral as written -- reading the unit's first word broke rules.md#P2's `John van der J. V` -- so `John van B and Smith X` reads family 'van B and Smith X' again, the reading #614's copy had moved. + WHAT THE COUNT KEEPS. A word the reading weighs -- a suffix word of either kind, the ambiguous class, a word in it by shape, an initial, a roman numeral by shape, a period-marked title word (`_weighed`) -- keeps a unit of its own though the chain may join it, as a position always counted it; only the plain words behind a particle run are JOINED, and a unit with nothing past its opener moves no name start into the titles (SUPERSEDED IN PART by #624, decisions.md#P2, 2026-10-10: no unit opens inside the titles, so the gate is gone). The first prototype marked every word the chain reaches JOINED, by the chain's own rule alone, and the review found it wrong in two shapes the earlier count read right: `Freiherr von Berg MA X.Y.Z.` kept 'MA X.Y.Z.' in the family (a member the chain would join counted as no word in front of the next), and `St St VI` and `Freiherr St MA` read title 'St St', 'Freiherr St' with the numeral or acronym the whole family (a unit of words the read then takes, which the chain never builds). Both masters read all three as they read now. (#624 since reads `St St VI` and `Freiherr St MA` as title 'St St' / 'Freiherr St' again, by a different route: the chain and the count now share one leading position, the last title-particle, so neither builds the unit -- decisions.md#P2, 2026-10-10.) The copy's exclusion list had been carrying the count positions always kept -- a member in front counts -- and the marks keep it without the copy. The numeral fork still reads the piece in front of the numeral as written -- reading the unit's first word broke rules.md#P2's `John van der J. V` -- so `John van B and Smith X` reads family 'van B and Smith X' again, the reading #614's copy had moved. ALSO COUNTED IN UNITS: #602's run start, a particle opening a chain run counting once and its bound particle not at all (`John der la Jr. Prof. Smith` keeps its run, `van la Smith Secretary Jr. Dr. Smith` starts none, as on master). And `credential_run`, on the read's path, gathers adjacent particles into the one absorbed piece the chain makes of them (`Dr. John Smith Esq. RN Mc Mc` reports 'Mc Mc'); M2's clause-free view reports word by word as it always has. DECLINED, measured: a units test in `tail_reading`'s resume condition. A splice lowers a count only by taking out titles that open units of their own, and with fewer than two units in front of it every piece there is in the first piece's chain run, which the titles behind it join; fuzzed over 372,330 names, the test changed nothing. CLASSIFICATION, found writing this entry's rows: eight case rows #614 added as `parity` did not match 1.4.0 -- they had been compared against master -- and are relabelled with the change each reading came from, bisected (#424, #289, #296, #516, #602 twice; R2 and P2 for two that read so since the 2.0 pipeline). MEASURED 2026-10-08 against 4278693d (#621's merge), py3.11, `nameparser.__file__` asserted, with the #614 harnesses and recipe above: fingerprint 0 moves; connective grid 0; S2 stress grid 12, three garbage names under four policies, none in the head-plus-one-word slice -- `Freiherr von Berg Dr. X.Y.Z. VD` reads title 'Freiherr Dr.', suffix 'X.Y.Z. VD' (H5's trailing title, the class #614 moved; master family 'von Berg Dr. X.Y.Z.'), `Jan Freiherr von Berg VD V and ba I` reads family 'VD V and ba I' as adf6da88 did (master suffix 'I'), and `de Mesnil Ma do mc and Ph. D.` reports the 'do' the peel took beside adf6da88's 'Ma'. Five gates exit 0. Frames unchanged, 308 on 3.11 and 289 on 3.12-3.15 (decisions.md#parse-cost). /SIMPLIFY (2026-10-08), measured. Each pass of `tail_reading` had rescanned for the second unit from the front, a C-level quadratic no frame guard saw: `Freiherr von` + `Berg `*n + `MA Dr. `*n read 5.77x for 4x the input against master's 4.08x; the lookup is now made once per read and handed to every pass (4.03x), and `test_the_second_unit_is_found_once_per_read` counts it. `_weighed` is pinned to the reading by `test_a_word_the_read_takes_is_never_one_the_count_joined` (no piece marked JOINED is ever in the run the read takes; 320 grid names fail with the test answering False). The gathering of an absorbed particle run reads the BOUND marks, and #602's run start one condition. Grids byte-identical, frames unchanged. Left for a follow-up, each moving readings: merging the tail's particle runs in group before the read, which would retire the BOUND stop and the gathering (80 grid names move, `Dr. mc mc`, and the chain's particle report needs re-keying); and the name start decided once, the item below. - LEFT OPEN, by Derek's scoping: reading a family comma's given part the same way, and deciding the name's start once for a title-particle head (#620's related paths, which change readings). + LEFT OPEN, by Derek's scoping: reading a family comma's given part the same way, and deciding the name's start once for a title-particle head (#620's related paths, which change readings). The second was settled by #624 (decisions.md#P2, 2026-10-10). - 2026-10-10 (Derek), #627 — A LONE PARTICLE WITH THE RUN RIGHT BEHIND IT COUNTS AS A NAME WORD. #602's run start needs two name words before the credential, and "a lone particle is not a name word for the count", for `de Mesnil`, where the particle and the word behind it are one surname. Since #620 the count is taken in the chain's units, so a particle that opens a unit counts once (`van der`, `von Berg`) and a P3-joined phrase is one piece (`von und zu`); the lone particle stayed uncounted only where nothing joins it, which is exactly where a credential stands behind it. So `John von PhD Jones` read middle 'von PhD', family 'Jones' -- the chain later taking the credential as the particle's surname -- while `John van der PhD Jones` and `John von und zu PhD Jones` read family 'van der' / 'von und zu', suffix 'PhD Jones', and the comma spelling `Smith, John von PhD Jones` already read suffix 'PhD Jones'. The issue's question was which way to make them agree. DECIDED (Derek): the particle stops at the credential, so it is a name word of its own and the rest is suffix: `John von PhD Jones` → family 'von', suffix 'PhD Jones'. The input is malformed (a particle is not written before a post-nominal), so the decision is whichever reading the shared count already gives the other forms, not a special case for this one: the existing clause gains its own limit, a lone particle counting when the run starts right behind it, in `run_start`'s one loop. The caller without the chain's units (`trailing_start`) shares the expression so the two cannot disagree, but no reading shows that half: it changes the position it returns for `Smith, John von PhD Jones` (2 where master gave 4), so group's chain takes its stopped path and asks `trailing_start` again, reaching the same reading at 16 more frames (339 → 355 on py3.11; a rare shape no frame guard sees, decided for the shared expression over a per-caller copy); a mutant ignoring the flag there passed the suite with the grid byte-identical (#627's reviews). The 1.4.0 and 2.3.0 wheels read middle 'von PhD', as they read `von und zu PhD`; the joined forms changed this cycle with #602. Declined: counting every unit, a lone particle included. It moved 2,836 S2-grid parses rather than 2,432, the extra ones all `de Mesnil `, where it counted the surname as two words, and it split the two callers: a leading `van` counted before another particle (`van la Smith Jr. Jones` → suffix 'Jr. Jones') but not before a name word (`van Smith Jr. Jones`), and the case row `van la Smith Secretary Jr. Dr. Smith` caught it. @@ -1608,6 +1613,7 @@ Every number below is a py3.11 measurement of 2026-08-31, recomputable with `uv - 2026-09-26 #546 -- the stages copy their state with `_state.copy_with` instead of `dataclasses.replace`, and every row moves DOWN. `dataclasses.replace` walks `fields()` and calls `__init__` on each copy, three frames on 3.11 and 3.12 and four from 3.13 where `copy_with` is one. One parse of the reference name makes 18 copies: six of the state, one each from tokenize, segment, classify, group, assign and post_rules, and twelve of single tokens, six each from classify and assign (`extract_delimited` returns the state unchanged when there is no delimiter, and `script_segment` returns early on ASCII input). That is the whole of the drop, 18 x 2 = 36 frames on 3.11 and 3.12 and 18 x 3 = 54 from 3.13 (recompute: wrap `copy_with` with a counter in the eight stage modules and parse the reference name). A field copy builds what `replace` builds for a dataclass that is decorated itself rather than inheriting the decoration, keeps the generated `__init__`, and has no `__post_init__` and no `init=False` field. `_copyable_fields` checks exactly those four, and `_COPY_FIELDS` runs it over `WorkToken`, `PendingAmbiguity` and `ParseState` at import, so a class that stops qualifying fails there and `copy_with` copies nothing else. `test_the_guard_refuses_a_class_a_field_copy_would_get_wrong` records, for each refused shape, what `replace` builds and what an unguarded copy would build instead. To mypy, `copy_with` is `from dataclasses import replace as copy_with`, which keeps the dataclass plugin's keyword and type checks at every call site; an assignment (`copy_with = dataclasses.replace`) would not, since the plugin keys on the callee's full name (measured: a misspelled field and a wrong-typed value both pass through the assignment and both fail through the import). Measured 2026-09-26 with `uv run python tools/perf/call_count.py --against e0f1a2f`, each row on its own interpreter, parse/facade: 3.11 406/443 → 370/407, 3.12 384/421 → 348/385, 3.13, 3.14 and 3.15 402/439 → 348/385. The rows drop by 40 and 58 rather than 36 and 54 because e0f1a2f already read 4 under every row, inside the band, and the new rows are set to what the harness reads now. `_LINK_BASELINE`'s 64-link clause reads 2587 → 2301 on 3.11. By stage (`--stages`, py3.11, ms per 1000 parses of the reference name): group 19.6 → 17.3, classify 13.4 → 9.8, assign 11.1 → 8.0, tokenize 8.1 → 6.6, post_rules 7.9 → 6.5, segment 2.6 → 1.7. BEHAVIOR IDENTICAL: the differential gate's report at all five baselines matches e0f1a2f's line for line apart from the path header. - 2026-09-29 #561 -- SUPERSEDED IN PART: the 2026-09-26 bullet above, whose four conditions ("decorated itself", "keeps the generated `__init__`", no `__post_init__`, no `init=False` field) were not sufficient. `replace` builds through `obj.__class__(...)`, and a field copy skips every hook that call runs; four shapes met all four conditions and still parted: a validating `__new__`, a validating `__setattr__` on a non-frozen dataclass and a validating metaclass `__call__`, each giving `ValueError` from `replace` where an unguarded field copy built `value=-1`, and an `InitVar` with no default, which `replace` demands (a `ValueError` through 3.12, a `TypeError` from 3.13) and a field copy never sees. #561's first tightening (80554266) found only the first three, and by swapping the class's OWN `__dataclass_params__` for an inherited lookup it also let through an `__init__` borrowed from another dataclass under that dataclass's name, which the 2026-09-26 guard had refused; review caught both. So the guard is now a list of conditions a class must MEET (`_state._copy_refusals`, one reason string per unmet condition): decorated as a dataclass itself, frozen, an `__init__` generated for it (compiled from `` and named `.__init__`) that takes exactly its `init` fields, the default metaclass, no `__new__` above `object`, no `__post_init__`, no `init=False` field. It is a TRIPWIRE for a realistic edit to the three pipeline classes and says so, not a proof against any class: a forged `__qualname__` on a borrowed `__init__` in a class that is also decorated itself, or a `__class__` property, still passes, and `copy_with` copies nothing but `WorkToken`, `PendingAmbiguity` and `ParseState`, which meet every condition on 3.11 to 3.15. `_UNGUARDED_EFFECT` records for each row what `replace` builds, what an unguarded copy builds, and which conditions refuse it; `test_every_guard_condition_alone_refuses_a_recorded_shape` asserts that each condition is the only refusal of some row, so dropping one fails a test rather than contradicting this paragraph (checked by removing the `__new__` condition in a scratch copy: two failures). Call counts unchanged, the guard running at import only. +- 2026-10-10 #624 -- every row moves DOWN by 4: `parse` 308 → 304 on 3.11 and 289 → 285 on 3.12-3.15, `facade` 345 → 341 and 326 → 322, each measured on its own interpreter with `tools/perf/call_count.py` in isolated copies of 8e6e524a (#631's merge) and of the change, the import asserted (8e6e524a read exactly the recorded rows). `--modules` on 3.11: `_pieces.py` 51 → 49 and `_group.py` 25 → 23 -- group's leading-position scan, a generator asking `is_leading_title` per piece, gave way to `leading_titles`, which it already called, and `chain_lead`, which asks `is_prefix_piece` only of the titles. - 2026-10-06 #617 -- every row moves DOWN by 10: `parse` 374 → 364 on 3.11 and 353 → 343 on 3.12-3.15, `facade` 411 → 401 and 390 → 380, each measured on its own interpreter with `PYTHONPATH= pythonX.Y tools/perf/call_count.py` against 83f4e914 (every row had sat inside its band before, the 3.12-3.15 ones at +5 over a 348 baseline). `--modules` puts all ten in group and the piece layer (`_group.py` + `_pieces.py` 142 → 132, every other module unchanged): P3's connective joins moved out of `_group_segment` into `_pieces.join_connectives`, which asks `is_conj_piece`, `is_title_piece` and `is_prefix_piece` directly where the loops had asked through one-line closures over them, two frames per test becoming one. The link pin moves with it, 2,678 → 2,479 for the 64-link name on 3.11 (`_LINK_BASELINE`), and there for two reasons, counted per function in #617's review: the closures (`conj` 133 frames, `title` and `prefix` one each) and the single-letter test, which had joined each connective's text through a generator (65 frames, one per link, more than the pin's whole band), less the one `join_connectives` frame. A deliberate move, not drift: the change was a refactor (decisions.md#P3, 2026-10-06). The reference rows hold no comma; a comma part holding a connective now pays the shared loop (`John Smith, Mr. and Mrs.` 262 → 269 on 3.11, warm), and one holding none skips it and costs what it did (`Smith, John` 161), the review having found the first cut charging every family-comma parse two frames per word of the part. THEN 39 MORE on every row, in the branch's /simplify pass: group ran the rootname count (about five frames a piece, its only reader the carve-out) and the shared loop over every segment of three or more pieces, connective or not, and over the corpora and the case table under three orders 0 of the 3,153 group calls with no connective left free changed a piece. Group now skips both where its frozen-set walk, which already reads every piece's `conjunction` tag, finds no connective free, as `_comma._reading_pieces` does: `parse` 364 → 325 on 3.11 and 343 → 304 on 3.12-3.15, `facade` 401 → 362 and 380 → 341, `_group.py` + `_pieces.py` 132 → 93, measured on each interpreter as above; the link pin does not move, every link in it joining. Per name on 3.11, warm: the Vega name 436 → 386, `Smith, John Quincy Adams Bob` 326 → 297, `Smith, John` and `John Smith, Mr. and Mrs.` unchanged at 161 and 269. No output moved over the #617 grids (decisions.md#P3). - 2026-10-07 -- every row moves DOWN by 8 more: `_group_segment`'s last three one-line predicate closures, `prefix` over `is_prefix_piece`, `suffix` over `is_suffix_piece` and `marker` over `_is_maiden_marker_piece`, are inlined as direct calls, finishing what #617 did for `title`, `conj` and `merge`. Each closure was a frame of its own ahead of the predicate's, and the chain's two inner scans ask `prefix` once per piece they pass, so the cost grows with particle pieces (frame-budget question 1 in AGENTS.md). Measured with `PYTHONSAFEPATH=1 PYTHONPATH= pythonX.Y tools/perf/call_count.py`, each row on its own interpreter, parse/facade against 0c54662b: 3.11 325/362 → 317/354, 3.12, 3.13, 3.14 and 3.15 304/341 → 296/333. `--modules` on 3.11 puts all eight in `_group.py` (34 → 26), every other module unchanged. The whole of the drop is the closures, counted per function on 0c54662b (a profile hook counting `call` events whose code is named `prefix`, `suffix` or `marker` in `_group.py`): the reference name asked `prefix` 7 times and `suffix` once, `Juan` + `de la Vega ` × 8 + `Smith` asked 40 and 9 and reads 937 → 888 on 3.11 (919 → 870 on 3.12-3.15), and the 64-link name asked `prefix` once, so `_LINK_BASELINE` moves 2,479 → 2,478 on 3.11 and gains rows for 3.12-3.15 at 2,459 (2,460 on 0c54662b), each measured on its own interpreter and the first the pin has had beyond 3.11; `marker` is reached only by P5's bound-join decline and the reference name never asks it. BEHAVIOR IDENTICAL: the full suite passes, and the differential gate's report at all five baselines matches 0c54662b's line for line apart from the path header. - 2026-10-08 #620 -- no row moves: `parse` 308 on 3.11 and 289 on 3.12-3.15, `facade` 345 and 326, measured as the #614 bullet below on 4278693d (#621's merge) and on the change alike. The 3.12-3.15 pin read 290/327 until this date: #614's /simplify took one frame off those interpreters (289/326 on 4278693d, measured 2026-10-08) and the pin stayed, inside its band, so the #614 bullet's 290 is the pin's number, not that tree's. `--modules` on 3.11 unchanged. The merged copy and its per-word exclusion test are gone, and `_chain_units` asks `chain_run_end` per particle unit and `_weighed` per word behind one; the chain itself calls `chain_run_end` with the particle test inline where it had called `is_prefix_piece` per piece, which pays for the rest. decisions.md#S2, this date. diff --git a/docs/design/rules.md b/docs/design/rules.md index fcf449a3..99da49cf 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -243,8 +243,8 @@ H4. Rationale: this is a name parser, not a title parser. Handed a vocabulary reports `title-or-name` as well, the fork there being whether the title word inside the unit is a title at all. Usually a join (P3), and a particle chain (P4) is the same shape and - reports too — `St St née` reads family "St née" with `st` title - vocabulary inside it. The + reports too — `St van Bishop` reads family "van Bishop" with + `bishop` title vocabulary inside it. The clause reaches EVERY join whose non-leading member is TITLES vocabulary, not the one word that prompted it: `Smith and King`, `John and King`, `Smith and Bishop` and `John of Judge` all @@ -679,12 +679,20 @@ P4. Rationale: a particle links forward from inside a name; at the nothing (the title is not a name word), and why "Van Johnson" is a given-name reading at all. An unlisted abbreviation before the particle is as transparent as a listed title, since assign - reads it as one (H2). + reads it as one (H2). A word in the leading titles that is both a + title and a particle puts the name's leading position on itself, + so the particle behind it is inside a name ('Freiherr von Berg'); + of several such words the last does, so no particle chains inside + the titles (#624). The titles are those written, before H3 gives + the last one back to the name, so each such word reads as a title + unless H3 gives it back. "Van Johnson" → given="Van" "Sir de Mesnil" → pieces=[["Sir"], ["de"], ["Mesnil"]] "Xyz. van Johnson" → given="van" "John van der Berg" → pieces=[["John"], ["van", "der", "Berg"]] · boundary - history: decisions.md#P2 · interacts: P1, P5, H2 · implemented: nameparser/_pipeline/_group.py + "Freiherr St John Smith" → title="Freiherr St" + "Freiherr St John Smith" → given="John" + history: decisions.md#P2 · interacts: P1, P5, H2, H3 · implemented: nameparser/_pipeline/_group.py, nameparser/_pipeline/_pieces.py P5. Rationale: some given-name words are incomplete alone — "abdul" is a bound form that the next word completes. diff --git a/nameparser/_pipeline/_assign.py b/nameparser/_pipeline/_assign.py index eae0ffad..04b90f88 100644 --- a/nameparser/_pipeline/_assign.py +++ b/nameparser/_pipeline/_assign.py @@ -134,8 +134,9 @@ def _absorbed(piece: Sequence[int], # rules.md#H2: "an abbreviation opening the part of the name that # carries the given name — the whole name, or the part after a # family comma — reads as a title even when unlisted" -- the count is -# _pieces.leading_titles since #424 (its test, is_leading_title, is -# the leading-particle scan's too); the roles are set here. +# _pieces.leading_titles since #424 (the chain's leading position is +# read off the same run, _pieces.chain_lead, #624); the roles are set +# here. def _peel_leading_titles(pieces: tuple[tuple[int, ...], ...], ptags: tuple[frozenset[str], ...], tokens: list[WorkToken]) -> int: diff --git a/nameparser/_pipeline/_group.py b/nameparser/_pipeline/_group.py index 11ecc9e3..145343e5 100644 --- a/nameparser/_pipeline/_group.py +++ b/nameparser/_pipeline/_group.py @@ -63,7 +63,7 @@ from nameparser._lexicon import _run_addresses_by_given from nameparser._pipeline._pieces import ( - Peel, TailRead, chain_run_end, is_conj_piece, is_leading_title, is_prefix_piece, + Peel, TailRead, chain_lead, chain_run_end, is_conj_piece, is_leading_title, is_prefix_piece, is_suffix_piece, is_title_piece, join_connectives, joined_tags, leading_titles, merge_pieces, peel_walk, read_trailing_run, tail_reading, trailing_candidates, trailing_start, @@ -819,18 +819,18 @@ def _group_segment(seg: tuple[int, ...], additional: int, # name text parsed two ways ("Van Johnson" -> given Van, family # Johnson; "Dr. Van Johnson" -> family "Van Johnson"). # - # "Title AND NOT prefix" rather than the plain "not a title" the - # rule is stated as, and the difference is not academic: `st`, - # `do` and `freiherr` are each BOTH a title and an ambiguous - # particle, so the plain test skipped over the very piece the - # exception exists to protect and "St John Smith" -- no title in - # front of it at all -- collapsed from title St, given John, - # family Smith into one given "St John Smith". A piece that - # could be the name's own first piece stops the scan; only a - # piece that can ONLY be a title is stepped over. + # A word that is BOTH a title and a particle (`st`, `freiherr`) + # in the title run leads in the run's place, and the last such + # word does (`chain_lead`, #624): a scan stepping over every + # title skipped the very piece the exception exists to protect, + # and "St John Smith" -- no title in front of it at all -- + # collapsed from title St, given John, family Smith into one + # given "St John Smith"; a scan stopping at the FIRST such word + # chained the second, so "Freiherr St John Smith" read family + # "St John Smith" while assign read both words as titles. # # Computed once, before the loop: every merge below starts at - # some k at or past this index, so no merge can move it. + # some k past this index, so no merge can move it. # # Suffix pieces are deliberately NOT skipped, and the reason is # what skipping them WOULD do rather than what it would cost. @@ -864,21 +864,20 @@ def _group_segment(seg: tuple[int, ...], additional: int, # take it" had to restate the consumer's condition and got it # wrong one suffix later (#417). # - # The `, 0` fallback is inert by construction rather than a - # default worth testing: it is reached only when every piece is - # a title and none is a prefix, and the loop below merges - # nothing unless some piece is a prefix. - # `is_title_piece` alone missed H2's unlisted abbreviations, which - # assign peels as titles all the same, so 'Xyz. van Johnson' - # chained where 'Dr. van Johnson' did not (#424 found it - # through the acronym fork: the chain had swallowed the given - # word and left assign two pieces where the fork counted - # three). The scan asks assign's own test. - leading = next((k for k in range(len(pieces)) - if not is_leading_title(pieces[k], ptags[k], - tokens) - or is_prefix_piece(pieces[k], ptags[k], tokens)), - 0) + # The leading position is asked of assign's own title run + # (`leading_titles`): `is_title_piece` alone missed H2's + # unlisted abbreviations, which assign peels as titles all the + # same, so 'Xyz. van Johnson' chained where 'Dr. van Johnson' + # did not (#424 found it through the acronym fork: the chain + # had swallowed the given word and left assign two pieces where + # the fork counted three). And `chain_lead` is the one answer + # the trailing read's count takes too: of two titles that are + # also particles the second leads (#624), where a scan of this + # loop's own had stopped at the first and chained the second, + # so 'Freiherr St John Smith' read family 'St John Smith' while + # assign read both words as titles. + name_start = leading_titles(pieces, ptags, tokens) + leading = chain_lead(pieces, ptags, tokens, name_start) # rules.md#P2: "a trailing suffix begins" -- where it begins # is read by assign's peel over the pieces as they stand # (#424), once per segment and kept as a length from the end, @@ -894,7 +893,6 @@ def _group_segment(seg: tuple[int, ...], additional: int, # takes both forks, and where no run was read ahead of it it # asks again after its merges whether the acronym still has # the pieces the fork counted (below). - name_start = leading_titles(pieces, ptags, tokens) # the run read above is already split off (#614), so the chain # stops at the end of what is left tail = (0 if read is not None @@ -914,7 +912,7 @@ def chain(tail: int) -> None: titled = 0 k = 0 while k < len(pieces): - if k == leading or not is_prefix_piece(pieces[k], + if k <= leading or not is_prefix_piece(pieces[k], ptags[k], tokens): k += 1 continue @@ -929,15 +927,15 @@ def chain(tail: int) -> None: # needs an emitter in each. # # Narrow, and #367 is why. `titled == k` says every - # piece ahead of this one is a title, and the - # loop skipped k == leading, so `leading` is STRICTLY - # before k -- and being before k it is one of those titles, - # while being `leading` it satisfies `not title or prefix`. - # For both, it must be a prefix as well: a word in both - # vocabularies (`st`, `do`, `freiherr` by default, or any - # overlap a caller configures). A plain title alone can no - # longer put a particle off the name's leading piece; it is - # stepped over and _assign reports the fork instead. + # piece ahead of this one is a title, and the loop + # skipped every piece up to `leading`, so `leading` is + # STRICTLY before k -- one of those titles, and, being + # before the title run's end, a prefix as well + # (`chain_lead`): a word in both vocabularies (`st`, + # `freiherr` by default, or any overlap a caller + # configures). A plain title alone can no longer put a + # particle off the name's leading piece; it is stepped + # over and _assign reports the fork instead. # # What that leaves is wider than one shape: any number of # plain title pieces, then a piece in BOTH vocabularies, diff --git a/nameparser/_pipeline/_pieces.py b/nameparser/_pipeline/_pieces.py index fcd05bf3..180537ea 100644 --- a/nameparser/_pipeline/_pieces.py +++ b/nameparser/_pipeline/_pieces.py @@ -140,6 +140,44 @@ def is_title_piece(piece: Sequence[int], ptags: Set[str], return len(piece) == 1 and "vocab:title" in tokens[piece[0]].tags +# rules.md#P4: "a particle in the name's leading position chains +# nothing" -- WHERE that position is, asked once, by group's chain and +# the trailing read's unit count (#624) +def chain_lead(pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], + tokens: Sequence[WorkToken], n: int) -> int: + """The particle chain's leading position, given `n`, the end of + the leading title run (`leading_titles`): the last title in the run + that is also a particle, else the first piece past the run. A title + is no name word, so the first piece past the run leads; a word in + both vocabularies ('Freiherr', 'St') leads in its stead and puts + the particle behind it inside a name ('Freiherr von Berg', 'St van + Johnson'), and of several such words the last leads, all being + titles ('Freiherr St John Smith' as 'Dr. St John Smith'). No + particle stands between the lead and the run's end, so the chain + opens no unit inside the titles. + + The run is read as walked, before H3's give-back: a title the run + hands back to the name because only suffix words follow it is + still a title to the chain, so the particle behind it is the + name's leading piece and chains nothing (P4: 'Dr. Mc Mc' keeps + 'Mc' and 'Mc' apart, as before #624). `leading_titles` stops at a + piece that is itself a leading title only where it gave that piece + back, before a suffix piece (the tags it tests inline first), or at + the segment's last piece, which a title may not be unless it is the + whole segment -- the `n + 1 < len(pieces)` test rules that one out. + The same tags go first here, so a name with no suffix behind its + titles pays no frame for the test.""" + if (n + 1 < len(pieces) + and ("suffix" in ptags[n + 1] + or "vocab:suffix" in tokens[pieces[n + 1][0]].tags) + and is_leading_title(pieces[n], ptags[n], tokens)): + n += 1 + for k in range(n - 1, -1, -1): + if is_prefix_piece(pieces[k], ptags[k], tokens): + return k + return n + + # A particle, or a piece a join made one: what group's prefix chain # (at its loop in _group_segment) chains and stops at, what P3's join # derives a prefix from, and what group's rootname count and leading @@ -304,8 +342,8 @@ def join_connectives(pieces: list[list[int]], ptags: list[set[str]], # still reaches it as `_pieces._PERIOD_ABBREV` -- an import binds the # same name here, so the sync test's target did not move. Out of # assign since #424 and in the piece layer since #439: the test is -# assign's, and group's leading-particle scan and trailing-run walk -# must start where assign starts. +# assign's, and the chain's leading position (`chain_lead`) and the +# trailing-run walk must start where assign starts. # rules.md#H2: "an abbreviation opening the part of the name that @@ -353,10 +391,10 @@ def leading_titles(pieces: Sequence[Sequence[int]], (rules.md#H3, decisions.md#H3 -- the block at the floor below carries the examples of each half, and the ordering its two inline tag reads were measured on). - One definition, read by assign (which sets the roles) and by the - chain's trailing-run walk; the leading-particle scan shares the - predicate, is_leading_title, but stops at a title-and-particle - word (P4, #367, #424).""" + One definition, read by assign (which sets the roles), by the + chain's trailing-run walk, and by `chain_lead`, which finds the + chain's leading position inside this run for group's chain and the + trailing read alike (P4, #367, #424, #624).""" n = 0 while n < len(pieces): if ((n + 1 < len(pieces) or len(pieces) == 1) @@ -1708,7 +1746,7 @@ def chain_run_end(k: int, pieces: Sequence[Sequence[int]], def _chain_units(pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], tokens: Sequence[WorkToken], - n: int) -> tuple[list[int] | None, int]: + n: int) -> list[int] | None: """The name units P2's chain will make of `pieces`, as a mark per piece -- OPENS where a piece opens a unit, JOINED where the chain will join a name word to the run in front of it and the reading @@ -1717,13 +1755,11 @@ def _chain_units(pieces: Sequence[Sequence[int]], particle inside the run, a word the run took and so a name word whatever else it is ('van mc', 'von vd': rules.md#S2, the words both particles and suffix vocabulary standing straight behind a - particle) -- and where the name starts. - `n` is the end of the leading title run; the chain's leading - position is the first title that is also a particle ('Freiherr - von vd', 'St van Mc'), else `n`, and the chain opens no unit there. - A unit the chain opens INSIDE the titles makes a name of them - ('Freiherr St van Berg MA', 'Freiherr Freiherr Prof do'), so the - first unit at or before `n` is where the name starts. + particle). + `n` is the end of the leading title run, and the chain starts past + its leading position (`chain_lead`, the one answer group's chain + takes too, #624), which opens no unit inside the titles: the name + starts at `n`. Nothing is merged: the read counts with the flags and reads every piece as written, so no word it weighs is hidden inside a unit @@ -1737,15 +1773,9 @@ def _chain_units(pieces: Sequence[Sequence[int]], and "particle" in tokens[pieces[k][0]].tags) for k in range(count)] if not any(prefix[1:]): - return None, n - leading = n - for k in range(n): - if prefix[k]: - leading = k - break + return None units = [OPENS] * count - at = n - k = leading + 1 + k = chain_lead(pieces, ptags, tokens, n) + 1 while k < count: if prefix[k]: j = chain_run_end(k, pieces, ptags, tokens, count) @@ -1756,14 +1786,10 @@ def _chain_units(pieces: Sequence[Sequence[int]], while q < j and not _weighed(pieces[q], tokens): units[q] = JOINED q += 1 - # a unit with nothing past its opener is no unit the count - # sees, and makes no name of the titles it opens inside - if q > k + 1 and k < at: - at = k k = j continue k += 1 - return units, at + return units def _weighed(piece: Sequence[int], tokens: Sequence[WorkToken]) -> bool: @@ -1773,9 +1799,8 @@ def _weighed(piece: Sequence[int], tokens: Sequence[WorkToken]) -> bool: word in it by shape, an initial, a roman numeral by shape, or a period-marked title word. A word the chain joins and the reading weighs is a word the reading may yet take ('Freiherr von Berg MA - X.Y.Z.' keeps suffix 'MA X.Y.Z.'), and one that opens a unit inside - the titles may leave the titles a title ('St St VI'), so it ends the - plain run the count folds into the particle (#620's review). A + X.Y.Z.' keeps suffix 'MA X.Y.Z.'), so it ends the plain run the + count folds into the particle (#620's review). A connective join is one name word, P3's own count. The tests inline: this asks once per word behind a particle run.""" if len(piece) > 1: @@ -1810,8 +1835,8 @@ def read_trailing_run(pieces: Sequence[Sequence[int]], n = leading_titles(pieces, ptags, tokens) if n == len(pieces): return None - units, at = _chain_units(pieces, ptags, tokens, n) - rest, titled, peel = tail_reading(peel_walk(at, ptags), pieces, ptags, + units = _chain_units(pieces, ptags, tokens, n) + rest, titled, peel = tail_reading(peel_walk(n, ptags), pieces, ptags, tokens, one_case, units) tail: set[int] = set() for k in rest[peel.names:]: diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 15bc9e79..85f9aaeb 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -7357,25 +7357,23 @@ def _check_cjk_shape_purity(self) -> None: "in the family name. Ungated, #399 moved 'van der' out " "to the middle and left family 'Nee'"), Case("maiden_marker_trailing_keeps_the_fork_report", - "St St née", - {"title": "St", "family": "St née"}, - ambiguities=("particle-or-given", "title-or-name"), + "St van née", + {"title": "St", "family": "van née"}, + ambiguities=("particle-or-given",), notes="'st' is both a title and an ambiguous particle (#367), " - "so this shape reaches group's PARTICLE_OR_GIVEN " - "emitter, which is guarded on the chain having merged " - "something. An ungated marker stop made the chain merge " - "nothing for a DIFFERENT reason than the guard assumes, " - "silencing the report while still deciding the fork -- " - "the shape A1 forbids and #405 closed at P6. Pinned because " - "removing a report a caller already sees is worse than " - "never emitting one. The second flag is H4's join " - "clause, gained in the #518 review round: the one name " - "unit is a particle CHAIN rather than a P3 join, and " - "carries `st` -- title vocabulary -- beside the word it " - "places, which is the same fork. The `particle` tag on " - "that word is what used to silence it, and a claimed " - "word decides the FIELD and not whether the word is a " - "title"), + "so it leads the chain and the ambiguous 'van' behind it " + "reaches group's PARTICLE_OR_GIVEN emitter, which is " + "guarded on the chain having merged something. An " + "ungated marker stop made the chain merge nothing for a " + "DIFFERENT reason than the guard assumes, silencing the " + "report while still deciding the fork -- the shape A1 " + "forbids and #405 closed at P6. Pinned because removing " + "a report a caller already sees is worse than never " + "emitting one. Until #624 the row was 'St St née', the " + "second 'St' chained; of two title-particles the second " + "now leads and is a title (title 'St St', family 'née', " + "no report), so the row moved to a chained particle " + "that is not a title"), Case("maiden_marker_particles_on_both_sides", "Anna von der Müller geb. von der Berg", {"given": "Anna", "family": "von der Müller", @@ -9836,32 +9834,29 @@ def _check_cjk_shape_purity(self) -> None: "review) (1.4.0 given 'Freiherr von vd', its all-particles " "guard reading the family as a given name; read so since the " "2.0 pipeline)"), - Case("a_second_title_particle_opens_the_name_inside_the_titles", + Case("a_second_title_particle_leads_the_chain", "Freiherr St van Berg MA", - {"title": "Freiherr", "family": "St van Berg", "suffix": "MA"}, - ambiguities=("particle-or-given", "title-or-name", - "suffix-or-name"), - classification="fix(#289)", - notes="'St' is a title and a particle, so it is both inside " - "the leading titles and the start of the chain's unit " - "'St van Berg'; the read must start where that unit " - "opens, not at the title run's end, which put it past " - "the name and kept 'MA' a name word (#614's second " - "review; since #620 the start is read off the pieces " - "themselves) (1.4.0 given 'St van Berg', family 'MA'; 2.0-2.3 " - "family 'St van Berg MA'; bisected to #289's case contrast)"), - Case("a_unit_opened_inside_the_titles_makes_them_a_name", - "Freiherr Freiherr Prof do", - {"title": "Freiherr", "given": "Freiherr Prof", "family": "do"}, + {"title": "Freiherr St", "family": "van Berg", "suffix": "MA"}, ambiguities=("particle-or-given", "suffix-or-name"), - classification="fix(P2)", - notes="the second 'Freiherr' opens the chain's unit 'Freiherr " - "Prof' inside the leading titles, which makes a name of " - "them though it ends short of the titles' end; the read " - "starts at that unit, so 'do' has a name word in front " - "of it and its pick is reported (#614's second review) (1.4.0 " - "given 'Freiherr Freiherr Prof', its particle chain taking the" - " leading title; read so since the 2.0 pipeline)"), + classification="fix(#624)", + notes="of two title-particles at the head the second leads the " + "chain (rules.md#P4), so both are titles and 'van Berg' " + "is the family, as 'Dr. St van Johnson' reads; the chain " + "and the read's count take that one answer (1.4.0 given " + "'St van Berg', family 'MA'; 2.0-2.3 family 'St van Berg " + "MA'; until #624 title 'Freiherr', family 'St van Berg', " + "the chain starting at the first)"), + Case("the_chain_opens_no_unit_inside_the_titles", + "Freiherr Freiherr Prof do", + {"title": "Freiherr Freiherr Prof", "family": "do"}, + ambiguities=(), + classification="fix(#624)", + notes="the second 'Freiherr' leads the chain, so no unit opens " + "inside the titles and the name is 'do' alone (1.4.0 " + "given 'Freiherr Freiherr Prof', its particle chain " + "taking the leading title; until #624 the second " + "'Freiherr' opened the unit 'Freiherr Prof' inside the " + "titles, given 'Freiherr Prof', family 'do')"), Case("a_title_word_inside_a_credential_run_stays_a_title", "John Smith MD van Secretary Jones", {"given": "John", "family": "Smith", "title": "Secretary", @@ -10060,24 +10055,23 @@ def _check_cjk_shape_purity(self) -> None: "Jones', family 'Smith')"), Case("a_numeral_alone_behind_two_title_particles_stays_a_name", "St St VI", - {"title": "St", "family": "St VI"}, - ambiguities=("particle-or-given", "title-or-name"), - classification="parity", - notes="the second 'St' opens no unit the count sees: the numeral " - "is a word the reading weighs, so the name starts past the " - "titles, one piece, and the fork has nothing in front of " - "it (#620's review: its first count read title 'St St', " - "family 'VI')"), + {"title": "St St", "family": "VI"}, + ambiguities=(), + classification="fix(#624)", + notes="the second 'St' leads the chain and is a title, so the " + "numeral is the one name word and the fork has nothing in " + "front of it (1.4.0 title 'St', family 'St VI'; #620's " + "first count read title 'St St', family 'VI' while the " + "chain took 'St VI', and its review pinned the chain's " + "answer; #624 gives both the count's)"), Case("an_acronym_alone_behind_two_title_particles_stays_a_name", "Freiherr St MA", - {"title": "Freiherr", "family": "St MA"}, - ambiguities=("particle-or-given", "suffix-or-name", - "title-or-name"), - classification="fix(P2)", - notes="as the row above, for the acronym (#620's review: its " - "first count read title 'Freiherr St', family 'MA', and " - "reported nothing; 1.4.0 given 'Freiherr St', family 'MA', " - "read so since the 2.0 pipeline)"), + {"title": "Freiherr St", "family": "MA"}, + ambiguities=(), + classification="fix(#624)", + notes="as the row above, for the acronym (1.4.0 given " + "'Freiherr St', family 'MA'; until #624 title 'Freiherr', " + "family 'St MA')"), Case("a_weighed_word_behind_a_particle_counts_as_a_word", "Freiherr von Berg MA X.Y.Z.", {"title": "Freiherr", "family": "von Berg", "suffix": "MA X.Y.Z."}, @@ -10101,14 +10095,38 @@ def _check_cjk_shape_purity(self) -> None: "adf6da88 family 'von J. ma')"), Case("a_period_title_alone_behind_two_title_particles_stays_a_name", "Freiherr St Prof.", - {"title": "Freiherr", "family": "St Prof."}, - ambiguities=("particle-or-given", "title-or-name"), - classification="fix(P2)", - notes="a period-marked title word is one the H5 chain weighs, so " - "'St' opens no unit and the titles stay titles; the chain " - "takes no title out of a name of one word (#620's review, " - "mutation; 1.4.0 given 'Freiherr St Prof.', read so since " - "the 2.0 pipeline)"), + {"title": "Freiherr St", "family": "Prof."}, + ambiguities=("title-or-name",), + classification="fix(#624)", + notes="as the rows above, for a period-marked title word: the " + "H5 chain takes no title out of a name of one word, so " + "'Prof.' is the name (1.4.0 given 'Freiherr St Prof.'; " + "until #624 title 'Freiherr', family 'St Prof.')"), + Case("a_title_given_back_still_leads_no_chain", + "Dr. Mc Mc", + {"given": "Dr.", "suffix": "Mc Mc"}, + ambiguities=("title-or-name",), + classification="fix(#624)", + notes="H3 gives 'Dr.' back to the name, everything behind it " + "being suffix vocabulary, and the chain reads the titles " + "as written: 'Mc' is the name's leading piece, chains " + "nothing (P4), and the two 'Mc' are the post-nominals H3 " + "saw (1.4.0 title 'Dr.', family 'Mc Mc'; until #624 given " + "'Dr.', middle 'Mc', family 'Mc', the read counting the " + "second 'Mc' as bound into a run the chain never built; " + "#624's first draft read title 'Dr.', family 'Mc Mc', " + "chaining through the give-back)"), + Case("a_second_title_particle_is_a_title", + "Freiherr St John Smith MA", + {"title": "Freiherr St", "given": "John", "family": "Smith", + "suffix": "MA"}, + ambiguities=("suffix-or-name",), + classification="fix(#624)", + notes="#624's example: as 'Dr. St John Smith' and 'Sir St John " + "Smith' read, the second title-particle is a title and " + "the name starts behind it (1.4.0 title 'Freiherr', given " + "'St John Smith', family 'MA'; until #624 family 'St John " + "Smith', with a title-or-name report naming it)"), Case("a_connective_join_behind_a_particle_is_one_word", "Freiherr von B and Smith ma", {"title": "Freiherr", "family": "von B and Smith ma"}, diff --git a/tests/v2/pipeline/test_group.py b/tests/v2/pipeline/test_group.py index 643cf998..129287f6 100644 --- a/tests/v2/pipeline/test_group.py +++ b/tests/v2/pipeline/test_group.py @@ -905,8 +905,8 @@ def test_the_chain_stops_before_the_numeral_assign_reads_as_the_suffix() -> None def test_the_chain_keeps_an_acronym_assign_will_not_peel() -> None: # Behind a word in both the title and particle vocabularies the - # leading-particle scan stops (P4, #367) before assign's title - # peel does, so the chain takes the name's first word: read over + # chain's leading position sits (P4, #367, `chain_lead`) inside + # assign's title run, so the chain takes the name's first word: read over # the pieces as they stand the acronym has three pieces to spare, # read over the pieces the chain leaves it has two, and assign # would make it the family ('Freiherr von Berg Ma' read given 'von diff --git a/tests/v2/pipeline/test_pieces.py b/tests/v2/pipeline/test_pieces.py index f857cc2e..4f045239 100644 --- a/tests/v2/pipeline/test_pieces.py +++ b/tests/v2/pipeline/test_pieces.py @@ -1116,17 +1116,17 @@ def test_the_trailing_read_declines_where_there_is_nothing_to_split() -> None: assert _read("John Dr. G.J.")[1] is None -def _units(text: str) -> tuple[list[tuple[str, int]], int]: +def _units(text: str) -> list[tuple[str, int]]: """`_chain_units` over a whole name's words, one piece each, as - (word, mark) pairs, and where the name starts.""" + (word, mark) pairs.""" state = _state_through("classify", text) pieces = [[i] for i in state.segments[0]] ptags: list[set[str]] = [set() for _ in pieces] tokens = list(state.tokens) - units, at = _chain_units(pieces, ptags, tokens, - leading_titles(pieces, ptags, tokens)) + units = _chain_units(pieces, ptags, tokens, + leading_titles(pieces, ptags, tokens)) assert units is not None - return [(tokens[p[0]].text, u) for p, u in zip(pieces, units)], at + return [(tokens[p[0]].text, u) for p, u in zip(pieces, units)] def test_the_read_marks_the_units_the_chain_will_make() -> None: @@ -1134,24 +1134,20 @@ def test_the_read_marks_the_units_the_chain_will_make() -> None: is JOINED, and a particle inside the run is BOUND -- a word the run took, which the peel does not weigh ('van mc'); a suffix piece ends the run and stays a unit of its own.""" - marks, at = _units("John van mc Berg PhD") + marks = _units("John van mc Berg PhD") assert marks == [("John", OPENS), ("van", OPENS), ("mc", BOUND), ("Berg", JOINED), ("PhD", OPENS)] - assert at == 0 - - -def test_a_unit_opened_inside_the_titles_starts_the_name() -> None: - """A title that is also a particle opens the chain's unit inside the - leading titles, and the name starts there, not past the titles. A - word the reading weighs ('MA') is a word of its own though the - chain may join it -- and a unit that is nothing else past its opener - leaves the titles a title ('Freiherr St MA'), #620's review.""" - marks, at = _units("Freiherr St van Berg MA") - assert [u for _, u in marks] == [OPENS, OPENS, BOUND, JOINED, OPENS] - assert at == 1 - marks, at = _units("Freiherr St MA") + + +def test_the_chain_opens_no_unit_inside_the_titles() -> None: + """#624: of two titles that are also particles the second leads the + chain (`chain_lead`), so no unit opens inside the titles: 'van' + opens the first, past them. A word the reading weighs ('MA') is a + word of its own though the chain may join it.""" + marks = _units("Freiherr St van Berg MA") + assert [u for _, u in marks] == [OPENS, OPENS, OPENS, JOINED, OPENS] + marks = _units("Freiherr St MA") assert [u for _, u in marks] == [OPENS, OPENS, OPENS] - assert at == 2 _WEIGHED_GRID_HEADS = ("John van", "Freiherr von", "anh van", "Jan de la", @@ -1175,8 +1171,8 @@ def _weighed_grid_violations() -> list[str]: pieces = [[i] for i in state.segments[0]] ptags: list[set[str]] = [set() for _ in pieces] tokens = list(state.tokens) - units, _ = _chain_units(pieces, ptags, tokens, - leading_titles(pieces, ptags, tokens)) + units = _chain_units(pieces, ptags, tokens, + leading_titles(pieces, ptags, tokens)) if units is None: continue read, _start = found diff --git a/tests/v2/test_benchmark.py b/tests/v2/test_benchmark.py index 43a26dd4..ab59d64d 100644 --- a/tests/v2/test_benchmark.py +++ b/tests/v2/test_benchmark.py @@ -90,11 +90,11 @@ #: #614's /simplify on and were corrected in #620, which moved no row #: (decisions.md#parse-cost). _CALL_BASELINE = { - (3, 11): {"parse": 308, "facade": 345}, - (3, 12): {"parse": 289, "facade": 326}, - (3, 13): {"parse": 289, "facade": 326}, - (3, 14): {"parse": 289, "facade": 326}, - (3, 15): {"parse": 289, "facade": 326}, + (3, 11): {"parse": 304, "facade": 341}, + (3, 12): {"parse": 285, "facade": 322}, + (3, 13): {"parse": 285, "facade": 322}, + (3, 14): {"parse": 285, "facade": 322}, + (3, 15): {"parse": 285, "facade": 322}, } _BAND = 0.02 diff --git a/tests/v2/test_cases.py b/tests/v2/test_cases.py index 4ad96674..86dd55a2 100644 --- a/tests/v2/test_cases.py +++ b/tests/v2/test_cases.py @@ -137,12 +137,14 @@ def test_a_report_names_the_field_its_word_lands_in( Checked over `_SWEPT_ROWS`, which the count test below walks too. - Negative control, measured 2026-10-08: this test over master's - parser (26cdb891) fails 14 parses, every one a report worded at - assign -- nine title-or-name join reports saying 'given' of a unit - H1 moved to the family behind a title (`Attorney General of - Minnesota`, `John of Prince Prof.`, `St St née`, `Freiherr von - Bishop X.Y.Z.` and five more rows, as declared), `Kim Min Do` + Negative control, measured 2026-10-08 and re-measured 2026-10-10 + over the rows after #624 (which rewrote six of them): this test + over master's parser (26cdb891) fails 14 parses, every one a report + worded at assign -- nine title-or-name join reports saying 'given' + of a unit H1 moved to the family behind a title (`Attorney General + of Minnesota`, `John of Prince Prof.`, `Freiherr St MA`, `Freiherr + von Bishop X.Y.Z.` and five more rows, as declared; on 2026-10-08 + `St St née` stood where `Freiherr St John Smith MA` does), `Kim Min Do` under FAMILY_FIRST ('middle', family, P6), `de Kim Ma` under FAMILY_FIRST ('middle', given, P1), and the opt-in movers' rows as declared: `Van Ivan Petrovich` and `Van Ali Veli oglu` ('given', @@ -175,8 +177,8 @@ def test_the_field_sweep_sees_the_claims_it_checks() -> None: _FIELD_CLAIM.search(a.detail) is not None for case, order in _SWEPT_ROWS for a in _swept(case.id, order).ambiguities) - assert claims == 1158, ( - f"the field sweep checks {claims} claims, recorded as 1158 on " + assert claims == 1149, ( + f"the field sweep checks {claims} claims, recorded as 1149 on " f"2026-10-10") diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index 8fd27e10..b823d8e9 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -4619,6 +4619,8 @@ def _claim(rule: dict) -> _Claim: _Claim(5, ('family', 'middle', 'suffix', 'title'), "3129cd9609b9", None), "fix(#627) a lone particle with the run right behind it is the surname": _Claim(1, ('family', 'middle', 'suffix'), "33096566ba7f", None), + "fix(#624) of two title-particles at the head the second leads the chain": + _Claim(1, ('family', 'given', 'title'), "c5b6da915be2", None), "fix(#274/#601/#602) a credential in the clause starts a run the take consumes": _Claim(1, ('family', 'maiden', 'middle', 'suffix'), "acdcc81cb17c", None), # 2026-10-04, #604: new, 3; the rules.md#P7 boundary and the @@ -5384,6 +5386,8 @@ def _claim(rule: dict) -> _Claim: _Claim(5, ('_ambiguities', 'family', 'middle', 'suffix', 'title'), "3129cd9609b9", None), "fix(#627) a lone particle with the run right behind it is the surname": _Claim(1, ('_ambiguities', 'family', 'middle', 'suffix'), "33096566ba7f", None), + "fix(#624) of two title-particles at the head the second leads the chain": + _Claim(1, ('_ambiguities', 'family', 'given', 'title'), "c5b6da915be2", None), "fix(#601/#602) a credential in the clause starts a run the take consumes": _Claim(1, ('_ambiguities', 'family', 'middle', 'suffix'), "acdcc81cb17c", None), "fix(#601) a marker behind a connective joins the connective run, and the run's initials follow": @@ -5863,6 +5867,8 @@ def _claim(rule: dict) -> _Claim: _Claim(5, ('_ambiguities', 'family', 'middle', 'suffix', 'title'), "3129cd9609b9", None), "fix(#627) a lone particle with the run right behind it is the surname": _Claim(1, ('_ambiguities', 'family', 'middle', 'suffix'), "33096566ba7f", None), + "fix(#624) of two title-particles at the head the second leads the chain": + _Claim(1, ('_ambiguities', 'family', 'given', 'title'), "c5b6da915be2", None), "fix(#601/#602) a credential in the clause starts a run the take consumes": _Claim(1, ('_ambiguities', 'family', 'middle', 'suffix'), "acdcc81cb17c", None), # 2026-10-04, #604: new, 3; the rules.md#P7 boundary and the @@ -6575,6 +6581,8 @@ def _claim(rule: dict) -> _Claim: _Claim(5, ('_ambiguities', 'family', 'middle', 'suffix', 'title'), "3129cd9609b9", None), "fix(#627) a lone particle with the run right behind it is the surname": _Claim(1, ('_ambiguities', 'family', 'middle', 'suffix'), "33096566ba7f", None), + "fix(#624) of two title-particles at the head the second leads the chain": + _Claim(1, ('_ambiguities', 'family', 'given', 'title'), "c5b6da915be2", None), "fix(#601/#602) a credential in the clause starts a run the take consumes": _Claim(1, ('_ambiguities', 'family', 'middle', 'suffix'), "acdcc81cb17c", None), "fix(#601) a marker behind a connective joins the connective run, and the run's initials follow": @@ -6907,6 +6915,8 @@ def _claim(rule: dict) -> _Claim: _Claim(4, ('_ambiguities', 'family', 'middle', 'suffix', 'title'), "e638a392a3a4", None), "fix(#627) a lone particle with the run right behind it is the surname": _Claim(1, ('_ambiguities', 'family', 'middle', 'suffix'), "33096566ba7f", None), + "fix(#624) of two title-particles at the head the second leads the chain": + _Claim(1, ('_ambiguities', 'family', 'given', 'title'), "c5b6da915be2", None), "fix(#601/#602) a credential in the clause starts a run the take consumes": _Claim(1, ('_ambiguities', 'family', 'middle', 'suffix'), "acdcc81cb17c", None), # 2026-10-04, #604: new, 3; the rules.md#P7 boundary and the diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index 8aa7b333..13da0796 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -84,6 +84,7 @@ "Eric H. Holder, Jr., Attorney General" "Eric H. Holder, Jr., Secretary of State" "Esq. Smith" +"Freiherr St John Smith" "Freiherr von Berg MA" "Freiherr von Berg, Ed" "Freiherr von Richthofen V" diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index 1b3eb399..81d81793 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -4835,6 +4835,14 @@ issue = "fix(#627) a lone particle with the run right behind it is the surname" name_regex = "^John von PhD Jones$" fields = ["family", "middle", "suffix"] +[[change]] +# rules.md#P4: "of several such words the last does, so no particle +# chains inside the titles" (#624, 2026-10-10): the second +# title-particle is a title and the name starts behind it. +issue = "fix(#624) of two title-particles at the head the second leads the chain" +name_regex = "^Freiherr St John Smith$" +fields = ["family", "given", "title"] + [[change]] # rules.md#M2 reads the clause-free name, whose credential starts #602's # run -- rules.md#S2: "A credential after the name core starts a run to diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index e0bdd718..f638ef33 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -3914,6 +3914,14 @@ issue = "fix(#627) a lone particle with the run right behind it is the surname" name_regex = "^John von PhD Jones$" fields = ["_ambiguities", "family", "middle", "suffix"] +[[change]] +# rules.md#P4: "of several such words the last does, so no particle +# chains inside the titles" (#624, 2026-10-10): the second +# title-particle is a title and the name starts behind it. +issue = "fix(#624) of two title-particles at the head the second leads the chain" +name_regex = "^Freiherr St John Smith$" +fields = ["_ambiguities", "family", "given", "title"] + [[change]] # rules.md#M2 reads the clause-free name, whose credential starts #602's # run -- rules.md#S2: "A credential after the name core starts a run to diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index fd2a396e..4c314d91 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -3859,6 +3859,14 @@ issue = "fix(#627) a lone particle with the run right behind it is the surname" name_regex = "^John von PhD Jones$" fields = ["_ambiguities", "family", "middle", "suffix"] +[[change]] +# rules.md#P4: "of several such words the last does, so no particle +# chains inside the titles" (#624, 2026-10-10): the second +# title-particle is a title and the name starts behind it. +issue = "fix(#624) of two title-particles at the head the second leads the chain" +name_regex = "^Freiherr St John Smith$" +fields = ["_ambiguities", "family", "given", "title"] + [[change]] # rules.md#M2 reads the clause-free name, whose credential starts #602's # run -- rules.md#S2: "A credential after the name core starts a run to diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index 17ffe6ed..f85fabe0 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -2357,6 +2357,14 @@ issue = "fix(#627) a lone particle with the run right behind it is the surname" name_regex = "^John von PhD Jones$" fields = ["_ambiguities", "family", "middle", "suffix"] +[[change]] +# rules.md#P4: "of several such words the last does, so no particle +# chains inside the titles" (#624, 2026-10-10): the second +# title-particle is a title and the name starts behind it. +issue = "fix(#624) of two title-particles at the head the second leads the chain" +name_regex = "^Freiherr St John Smith$" +fields = ["_ambiguities", "family", "given", "title"] + [[change]] # rules.md#M2 reads the clause-free name, whose credential starts #602's # run -- rules.md#S2: "A credential after the name core starts a run to diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index c033463d..ad18a89d 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -1643,6 +1643,14 @@ issue = "fix(#627) a lone particle with the run right behind it is the surname" name_regex = "^John von PhD Jones$" fields = ["_ambiguities", "family", "middle", "suffix"] +[[change]] +# rules.md#P4: "of several such words the last does, so no particle +# chains inside the titles" (#624, 2026-10-10): the second +# title-particle is a title and the name starts behind it. +issue = "fix(#624) of two title-particles at the head the second leads the chain" +name_regex = "^Freiherr St John Smith$" +fields = ["_ambiguities", "family", "given", "title"] + [[change]] # rules.md#M2 reads the clause-free name, whose credential starts #602's # run -- rules.md#S2: "A credential after the name core starts a run to