Repository navigation
Count a lone particle with the run behind it as a name word (#627) - #631
Merged
Merged
Conversation
#602's run starts at a credential with two name words in front of it, and a lone particle is not a name word for that count ('de Mesnil' is one surname). Since #620 the count is taken in the chain's units, so a particle that opens a unit counts ('van der', 'von Berg') and a joined phrase is one piece ('von und zu'): the lone particle stayed uncounted only where nothing joins it, which is where a credential stands right behind it. So 'John von PhD Jones' read middle 'von PhD' while 'John van der PhD Jones' and 'John von und zu PhD Jones' read suffix 'PhD Jones'. The clause gains its own limit in run_start's one loop, shared by both callers: a lone particle counts when the run starts right behind it, so it is the surname. 'John von PhD Jones' reads family 'von', suffix 'PhD Jones', as the comma spelling already did. Measured against 195c408: 2,432 of 324,552 S2-grid parses move, all that one shape; no corpus name or case text moves; the five gates exit 0. Counting every unit instead was declined: it also counted 'de Mesnil' as two words and split the callers on a leading particle. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #631 +/- ##
=======================================
Coverage 99.01% 99.01%
=======================================
Files 46 46
Lines 4486 4488 +2
=======================================
+ Hits 4442 4444 +2
Misses 44 44 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
- Case row `Freiherr von vd PhD Jones` (parity): a particle run is one name word, and the lone-particle flag is asked only of a piece that opens a unit. A mutant asking it of every piece passed the suite until this row; it now fails that row alone. - rules.md#S2 and the code comment: the particle "counts as a name word of its own" rather than "is the surname", which is true only in the default order. - decisions.md#S2: the units-less half of the clause moves no reading and is shared so the callers agree; the grid's limits (other particles and a head before a suffix comma move the same way); the fingerprint was taken over 195c408's texts; the new guard row. The #602 entry points at the qualification. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The five ledgers' fix(#627) comments quoted rules.md#S2 as "then it is the surname", which 8ea6865 reworded; they now quote "then it counts as a name word of its own". test_doc_citations scans only .py files, so nothing caught it (filed as a follow-up task). - decisions.md#S2: the units-less half does cost something -- group's chain takes its stopped path, +16 frames on `Smith, John von PhD Jones` (339 -> 355, py3.11) -- kept as the shared expression on a rare shape no guard sees. The heading and DECIDED sentence drop the order-specific "is the surname". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
derek73
added a commit
that referenced
this pull request
Oct 11, 2026
- chain_lead's docstring and the decisions.md #624 entry: leading_titles also stops at a segment's last piece, which the `n + 1 < len(pieces)` test rules out; P4 says each title-particle reads as a title unless H3 gives it back. - The five ledgers' fix(#624) comments quoted P4's first-draft wording, which 07957c5 replaced; they now quote the current text. The excerpt check still scans only .py files (the follow-up task filed on #631). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
derek73
added a commit
that referenced
this pull request
Oct 11, 2026
test_citations_are_verbatim_excerpts swept only *.py, so the `# rules.md#X: "..."` comments in tools/differential/expected_since_*.toml -- which AGENTS.md puts under the same excerpt discipline -- were checked by nothing; #631's review found five stale by hand. The sweep now reads the ledgers too. They are not "citing modules": test_implemented_matches_citing_modules still counts .py files only. Two scanner corrections the ledgers were the first input to reach: - a decisions.md#X citation was looked up in rules.md's statements (the regex accepted the doc name and dropped it), so a decisions.md#P6 quote was checked against rules.md#P6. It now must name a real decisions entry, and a quote OPENING it must be verbatim in that entry; a paraphrase pointer stays legal, the entry being a record rather than a statement. - an excerpt may elide with "[...]"; each fragment must be verbatim and in order. The fifteen failures (five citations, copied across ledgers): - H5 and S2 (1.4.0) quoted Accepted clauses; now quote statements that carry the same claim. - C1's "The part is read as its words stand" is now a longer sentence. - 2.2.0/2.3.0 S2 paraphrased with no quote; now quotes the sentence. - T3 paraphrased a mechanism that is not there: '王·Smith' is one word (the dot has a Latin neighbour) and '王Smith' reads the same; it now cites T3 and W4 for why no script order applies. Adds a reach assertion so a glob matching nothing cannot pass the excerpt test vacuously, and an elision-order test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
derek73
added a commit
that referenced
this pull request
Oct 11, 2026
derek73
added a commit
that referenced
this pull request
Oct 11, 2026
…ones (#632) * Check ledger rule citations as verbatim excerpts, fix the stale ones test_citations_are_verbatim_excerpts swept only *.py, so the `# rules.md#X: "..."` comments in tools/differential/expected_since_*.toml -- which AGENTS.md puts under the same excerpt discipline -- were checked by nothing; #631's review found five stale by hand. The sweep now reads the ledgers too. They are not "citing modules": test_implemented_matches_citing_modules still counts .py files only. Two scanner corrections the ledgers were the first input to reach: - a decisions.md#X citation was looked up in rules.md's statements (the regex accepted the doc name and dropped it), so a decisions.md#P6 quote was checked against rules.md#P6. It now must name a real decisions entry, and a quote OPENING it must be verbatim in that entry; a paraphrase pointer stays legal, the entry being a record rather than a statement. - an excerpt may elide with "[...]"; each fragment must be verbatim and in order. The fifteen failures (five citations, copied across ledgers): - H5 and S2 (1.4.0) quoted Accepted clauses; now quote statements that carry the same claim. - C1's "The part is read as its words stand" is now a longer sentence. - 2.2.0/2.3.0 S2 paraphrased with no quote; now quotes the sentence. - T3 paraphrased a mechanism that is not there: '王·Smith' is one word (the dot has a Latin neighbour) and '王Smith' reads the same; it now cites T3 and W4 for why no script order applies. Adds a reach assertion so a glob matching nothing cannot pass the excerpt test vacuously, and an elision-order test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Read every citation shape the tree writes; fix what that exposes The first commit checked only the colon form's first quote, and the review of #632 found the rest of the class: `rules.md#X -- "..."`, `rules.md#X's Accepted clause ("...")`, `rules.md#S2 consumes "..."` and a second quote chained by "and" all passed unread, 23 of them stale. The scanner now: - reads a quote after the ID in any of those shapes (_LEAD_RE), and every quote chained to it; a reference with no quote must still name a real rule, section, mechanism or decisions entry - quotes against a rule's statement AND its Accepted: clauses (example and pointer lines excluded), a section letter's Background (`rules.md#H Background`), a decisions entry's text - matches lowercase decisions keys, merges the `### differential- ledger, ...` arc headings onto their one key, rejects empty "[...]" fragments, reads `--` as the em dash and a nested ' as ", follows `#:` comment continuations and IDs wrapped at a hyphen - skips this module's own illustrative comments - records the retired excerpts as a negative control, and asserts the sweep reaches each shape (non-colon, chained, decisions, elided) The colon form alone still defines a citing module for implemented:. Fixed: P5 (the reserve's numeral clause and the join-after-the-run wording, #614), P2 (S2's once-read run), M2 (#601 replaced the walk with the clause-free trailing run; the 'Jane Doe Jr. nee Smith Ma' paragraph gains a dated note that #601 moved it to fix(#601/#602)), S2 ("OPENS THE STRING"), A1's segmenter clause, a decisions.md#S2 citation of rules.md#S2's Accepted clause, and four ledgers crediting AGENTS.md's "a reader would hesitate too" to rules.md#A1. Four ledgers' paraphrase of decisions.md#S2 now quotes its headline and names #544's second exception. With Accepted text quotable, the first commit's H5 and S2 rewrites go back to the original, more exact Accepted-clause quotes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Fix #632's second-round review: reach, controls, stale prose Scanner (tests/v2/test_doc_citations.py): - a section's Background is read whole, not its first line; a wrapped no-boundary: marker's continuation lines are no longer quotable - each citation records its lead, and the reach test asserts every lead shape is live ('s, parenthesis, dash, bare words) rather than any one of them; a new test pins what is quotable (an Accepted clause, a Background's wrapped line) and what is not (an example line, a no-boundary: continuation) - the docstring states the limits: case folds, and a quote more than four words from its ID and chained to nothing is not checked -- the measured alternative, every quote in the paragraph, flags 139, nearly all names written in double quotes - _RETIRED_EXCERPTS drops its P6 row (no version ever routed a decisions quote to a rule) and says which direction it guards Ledgers: fourteen stale quotes the check does not reach, found by the review -- C1's paired-initials sentences in all five (#563 made the speaking word an UNAMBIGUOUS suffix), N3 in three, P2 in one. Dated landing paragraphs for P2 (#424) and P5 (#425) now say they describe the rule as it then stood instead of quoting today's text as their reason; the M2 rule's opener names all four names #601 moved, not one. AGENTS.md: the count is measured, not summed from commit messages -- 27 stale quotes and 5 quoteless colon citations, the current check over e4b653a -- and the sentence states the four-word reach. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Name the count's comparator by its master commit e4b653a lived on #631's branch, which the squash merge leaves to be deleted; 8e6e524, master after #631, has the identical tree. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
#602's suffix run starts at a credential only with two name words in front of it, and "a lone particle is not a name word for the count" (rules.md#S2), because in
de Mesnilthe particle and the word behind it are one surname. Since #620 the count is taken in the particle chain's units, so a particle with something joined to it already counted. That left the rule firing in exactly one place: a bare particle with the credential right behind it, where nothing joins it.John von und zu PhD JonesJohn van der PhD JonesSmith, John von PhD JonesJohn von PhD JonesDecided by Derek: the particle stops at the known suffix. The input is malformed, so the reading is whichever one the shared count already gives the other forms, not a special case. The existing clause in
run_startgains its own limit, in the one loop both callers share (with the chain's units and without): a lone particle counts as a name word of its own when the run starts right behind it. A particle followed by a name word still counts once with that word (de Mesnil, and the existing rowvan la Smith Secretary Jr. Dr. Smith).Declined: deleting the clause outright, so that every unit counts. It also counted
de Mesnil <credentials>as two words, and it split the two callers: a leadingvancounted before another particle but not before a name word. The case row above caught it.Measured (decisions.md#S2, 2026-10-10)
Against master 195c408, py3.11,
nameparser.__file__asserted on each side:van, or a dualdo/mc/vdin any case) with an unambiguous credential right behind it becoming the family, and the run taking the rest. The grid's comma shapes don't move.von,de,du,la,bin) and for a head before a suffix comma (John von PhD Jones, Jr.→ family 'von', suffix 'PhD Jones, Jr.').von und zu PhD. The joined forms changed this cycle with Should a name word after a credential join the suffix?John Smith PhD Jonesreads middleSmith PhD#602.fix(#627)rule in each ledger.Smith, John161,Dr. Juan de la Vega III304 on py3.11), and the suite's frame guards pass unchanged. The moved shapes cost more:John von PhD Jones242 → 266, the run's own reading.trailing_start) shares the expression, so the two callers can't disagree. No reading shows that half, but it isn't free: forSmith, John von PhD Jonesit returns 2 where master returned 4, so group's chain takes its stopped path and askstrailing_startagain, reaching the same reading at 16 more frames (339 → 355). It's a rare garbage shape no guard sees, so the shared expression wins over a per-caller copy (AGENTS.md's frames rule).Changes
nameparser/_pipeline/_pieces.py:run_start's lone-particle clause and docstring.John von PhD Jones→ family 'von', suffix 'PhD Jones'.John Smith PhD Jonesreads middleSmith PhD#602 entry points at it.tests/v2/cases.py:a_lone_particle_before_the_run_is_the_surname(fix(#627)). The field sweep's recorded claim count moves from 1155 to 1158, the row's three orders.a_particle_run_before_the_run_stays_one_name_word(Freiherr von vd PhD Jones, parity): the flag is asked only of a piece that opens a unit. A variant asking it of every piece passed the whole suite until this row.tools/differential/: the regeneratedcorpus_rules.jsonland afix(#627)rule in all five ledgers.test_ledger_guards.pyrecords the new claims, plusfix(initials-per-word)'s reach growing by the new name (123 → 124).John Smith PhD Jonesreads middleSmith PhD#602), and the shape is malformed input.Review
abdul von PhD JonesandSir abdul von PhD Jonesmove consistently with the rule and are left without rows (malformed input)..py, filed as a follow-up task), the frames claims above, and the remaining "is the surname" phrasing in decisions.md.Closes #627
🤖 Generated with Claude Code