Skip to content

Count a lone particle with the run behind it as a name word (#627) - #631

Merged
derek73 merged 3 commits into
masterfrom
fix/issue-627-lone-particle-count
Oct 10, 2026
Merged

derek73 merged 3 commits into
masterfrom
fix/issue-627-lone-particle-count

Conversation

@derek73

@derek73 derek73 commented Oct 10, 2026 •

Copy link
Copy Markdown
Owner

Summary

#602's suffix run starts at a credential only with two name words in front of it, and "a lone particle is not a name word for the count" (rules.md#S2), because in de Mesnil the particle and the word behind it are one surname. Since #620 the count is taken in the particle chain's units, so a particle with something joined to it already counted. That left the rule firing in exactly one place: a bare particle with the credential right behind it, where nothing joins it.

Input Before Now
John von und zu PhD Jones family 'von und zu', suffix 'PhD Jones' same
John van der PhD Jones family 'van der', suffix 'PhD Jones' same
Smith, John von PhD Jones family 'von Smith', suffix 'PhD Jones' same
John von PhD Jones middle 'von PhD', family 'Jones' family 'von', suffix 'PhD Jones'

Decided by Derek: the particle stops at the known suffix. The input is malformed, so the reading is whichever one the shared count already gives the other forms, not a special case. The existing clause in run_start gains its own limit, in the one loop both callers share (with the chain's units and without): a lone particle counts as a name word of its own when the run starts right behind it. A particle followed by a name word still counts once with that word (de Mesnil, and the existing row van la Smith Secretary Jr. Dr. Smith).

Declined: deleting the clause outright, so that every unit counts. It also counted de Mesnil <credentials> as two words, and it split the two callers: a leading van counted before another particle but not before a name word. The case row above caught it.

Measured (decisions.md#S2, 2026-10-10)

Against master 195c408, py3.11, nameparser.__file__ asserted on each side:

  • S2 stress grid (the Rethink the credential-run rule (S2): read a name's trailing run once, over name units, and bind it before any join #614 entry's recipe, 324,552 parses): 2,432 role moves and no report-only move. Every move is a lone particle (in that grid van, or a dual do/mc/vd in any case) with an unambiguous credential right behind it becoming the family, and the run taking the rest. The grid's comma shapes don't move.
  • Beyond the grid: the review found the same move, as intended, for other particles (von, de, du, la, bin) and for a head before a suffix comma (John von PhD Jones, Jr. → family 'von', suffix 'PhD Jones, Jr.').
  • Fingerprint of every corpus name and case text at 195c408 under the three orders: 0 moves. This PR adds one text, its own example, which moves.
  • 1.4.0 and 2.3.0 wheels: both read middle 'von PhD', as they read von und zu PhD. The joined forms changed this cycle with Should a name word after a credential join the suffix? John Smith PhD Jones reads middle Smith PhD #602.
  • Five differential gates: all exit 0. The new rules.md example is explained by a fix(#627) rule in each ledger.
  • Frames: plain names are unchanged (Smith, John 161, Dr. Juan de la Vega III 304 on py3.11), and the suite's frame guards pass unchanged. The moved shapes cost more: John von PhD Jones 242 → 266, the run's own reading.
  • The caller without units (trailing_start) shares the expression, so the two callers can't disagree. No reading shows that half, but it isn't free: for Smith, John von PhD Jones it returns 2 where master returned 4, so group's chain takes its stopped path and asks trailing_start again, reaching the same reading at 16 more frames (339 → 355). It's a rare garbage shape no guard sees, so the shared expression wins over a per-caller copy (AGENTS.md's frames rule).

Changes

  • nameparser/_pipeline/_pieces.py: run_start's lone-particle clause and docstring.
  • rules.md#S2: one clause (the particle "counts as a name word of its own", true in every order), plus the example pair John von PhD Jones → family 'von', suffix 'PhD Jones'.
  • decisions.md#S2: new 2026-10-10 entry (decided, declined, measured). The Should a name word after a credential join the suffix? John Smith PhD Jones reads middle Smith PhD #602 entry points at it.
  • tests/v2/cases.py:
    • Row a_lone_particle_before_the_run_is_the_surname (fix(#627)). The field sweep's recorded claim count moves from 1155 to 1158, the row's three orders.
    • Row a_particle_run_before_the_run_stays_one_name_word (Freiherr von vd PhD Jones, parity): the flag is asked only of a piece that opens a unit. A variant asking it of every piece passed the whole suite until this row.
  • tools/differential/: the regenerated corpus_rules.jsonl and a fix(#627) rule in all five ledgers. test_ledger_guards.py records the new claims, plus fix(initials-per-word)'s reach growing by the new name (123 → 124).
  • No release-log bullet: the run reading is new in 2.4 (Should a name word after a credential join the suffix? John Smith PhD Jones reads middle Smith PhD #602), and the shape is malformed input.

Review

  • 3464370: the change.
  • 8ea6865: fixes the first review: the guard row, "counts as a name word of its own" for the order-specific "is the surname", and the decisions entry's grid limits, units-less half and fingerprint population. abdul von PhD Jones and Sir abdul von PhD Jones move consistently with the rule and are left without rows (malformed input).
  • e4b653a: fixes the fix commit's review: the five ledger comments quoted the old rules.md wording ("then it is the surname"; the excerpt check scans only .py, filed as a follow-up task), the frames claims above, and the remaining "is the surname" phrasing in decisions.md.

Closes #627

🤖 Generated with Claude Code

#602's run starts at a credential with two name words in front of it,
and a lone particle is not a name word for that count ('de Mesnil' is
one surname). Since #620 the count is taken in the chain's units, so a
particle that opens a unit counts ('van der', 'von Berg') and a joined
phrase is one piece ('von und zu'): the lone particle stayed uncounted
only where nothing joins it, which is where a credential stands right
behind it. So 'John von PhD Jones' read middle 'von PhD' while
'John van der PhD Jones' and 'John von und zu PhD Jones' read suffix
'PhD Jones'.

The clause gains its own limit in run_start's one loop, shared by both
callers: a lone particle counts when the run starts right behind it,
so it is the surname. 'John von PhD Jones' reads family 'von', suffix
'PhD Jones', as the comma spelling already did. Measured against
195c408: 2,432 of 324,552 S2-grid parses move, all that one shape; no
corpus name or case text moves; the five gates exit 0. Counting every
unit instead was declined: it also counted 'de Mesnil' as two words
and split the callers on a leading particle.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73 derek73 added this to the 2.4 milestone Oct 10, 2026
@derek73 derek73 added the bug label Oct 10, 2026
@derek73 derek73 self-assigned this Oct 10, 2026
@codecov

codecov Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 99.01%. Comparing base (195c408) to head (e4b653a).

Additional details and impacted files
@@           Coverage Diff           @@
##           master     #631   +/-   ##
=======================================
  Coverage   99.01%   99.01%           
=======================================
  Files          46       46           
  Lines        4486     4488    +2     
=======================================
+ Hits         4442     4444    +2     
  Misses         44       44           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

derek73 and others added 2 commits October 10, 2026 13:16
- Case row `Freiherr von vd PhD Jones` (parity): a particle run is one
  name word, and the lone-particle flag is asked only of a piece that
  opens a unit. A mutant asking it of every piece passed the suite
  until this row; it now fails that row alone.
- rules.md#S2 and the code comment: the particle "counts as a name word
  of its own" rather than "is the surname", which is true only in the
  default order.
- decisions.md#S2: the units-less half of the clause moves no reading
  and is shared so the callers agree; the grid's limits (other
  particles and a head before a suffix comma move the same way); the
  fingerprint was taken over 195c408's texts; the new guard row. The
  #602 entry points at the qualification.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The five ledgers' fix(#627) comments quoted rules.md#S2 as "then it
  is the surname", which 8ea6865 reworded; they now quote "then it
  counts as a name word of its own". test_doc_citations scans only .py
  files, so nothing caught it (filed as a follow-up task).
- decisions.md#S2: the units-less half does cost something -- group's
  chain takes its stopped path, +16 frames on
  `Smith, John von PhD Jones` (339 -> 355, py3.11) -- kept as the
  shared expression on a rare shape no guard sees. The heading and
  DECIDED sentence drop the order-specific "is the surname".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73 derek73 changed the title Count a lone particle with the run behind it as the surname (#627) Count a lone particle with the run behind it as a name word (#627) Oct 10, 2026
@derek73
derek73 merged commit 8e6e524 into master Oct 10, 2026
11 checks passed
derek73 added a commit that referenced this pull request Oct 11, 2026
- chain_lead's docstring and the decisions.md #624 entry: leading_titles
  also stops at a segment's last piece, which the `n + 1 < len(pieces)`
  test rules out; P4 says each title-particle reads as a title unless
  H3 gives it back.
- The five ledgers' fix(#624) comments quoted P4's first-draft wording,
  which 07957c5 replaced; they now quote the current text. The excerpt
  check still scans only .py files (the follow-up task filed on #631).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
derek73 added a commit that referenced this pull request Oct 11, 2026
test_citations_are_verbatim_excerpts swept only *.py, so the
`# rules.md#X: "..."` comments in tools/differential/expected_since_*.toml
-- which AGENTS.md puts under the same excerpt discipline -- were
checked by nothing; #631's review found five stale by hand. The sweep
now reads the ledgers too. They are not "citing modules":
test_implemented_matches_citing_modules still counts .py files only.

Two scanner corrections the ledgers were the first input to reach:
- a decisions.md#X citation was looked up in rules.md's statements
  (the regex accepted the doc name and dropped it), so a decisions.md#P6
  quote was checked against rules.md#P6. It now must name a real
  decisions entry, and a quote OPENING it must be verbatim in that
  entry; a paraphrase pointer stays legal, the entry being a record
  rather than a statement.
- an excerpt may elide with "[...]"; each fragment must be verbatim
  and in order.

The fifteen failures (five citations, copied across ledgers):
- H5 and S2 (1.4.0) quoted Accepted clauses; now quote statements that
  carry the same claim.
- C1's "The part is read as its words stand" is now a longer sentence.
- 2.2.0/2.3.0 S2 paraphrased with no quote; now quotes the sentence.
- T3 paraphrased a mechanism that is not there: '王·Smith' is one word
  (the dot has a Latin neighbour) and '王Smith' reads the same; it now
  cites T3 and W4 for why no script order applies.

Adds a reach assertion so a glob matching nothing cannot pass the
excerpt test vacuously, and an elision-order test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
derek73 added a commit that referenced this pull request Oct 11, 2026
e4b653a lived on #631's branch, which the squash merge leaves to be
deleted; 8e6e524, master after #631, has the identical tree.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
derek73 added a commit that referenced this pull request Oct 11, 2026
…ones (#632)

* Check ledger rule citations as verbatim excerpts, fix the stale ones

test_citations_are_verbatim_excerpts swept only *.py, so the
`# rules.md#X: "..."` comments in tools/differential/expected_since_*.toml
-- which AGENTS.md puts under the same excerpt discipline -- were
checked by nothing; #631's review found five stale by hand. The sweep
now reads the ledgers too. They are not "citing modules":
test_implemented_matches_citing_modules still counts .py files only.

Two scanner corrections the ledgers were the first input to reach:
- a decisions.md#X citation was looked up in rules.md's statements
  (the regex accepted the doc name and dropped it), so a decisions.md#P6
  quote was checked against rules.md#P6. It now must name a real
  decisions entry, and a quote OPENING it must be verbatim in that
  entry; a paraphrase pointer stays legal, the entry being a record
  rather than a statement.
- an excerpt may elide with "[...]"; each fragment must be verbatim
  and in order.

The fifteen failures (five citations, copied across ledgers):
- H5 and S2 (1.4.0) quoted Accepted clauses; now quote statements that
  carry the same claim.
- C1's "The part is read as its words stand" is now a longer sentence.
- 2.2.0/2.3.0 S2 paraphrased with no quote; now quotes the sentence.
- T3 paraphrased a mechanism that is not there: '王·Smith' is one word
  (the dot has a Latin neighbour) and '王Smith' reads the same; it now
  cites T3 and W4 for why no script order applies.

Adds a reach assertion so a glob matching nothing cannot pass the
excerpt test vacuously, and an elision-order test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Read every citation shape the tree writes; fix what that exposes

The first commit checked only the colon form's first quote, and the
review of #632 found the rest of the class: `rules.md#X -- "..."`,
`rules.md#X's Accepted clause ("...")`, `rules.md#S2 consumes "..."`
and a second quote chained by "and" all passed unread, 23 of them
stale. The scanner now:

- reads a quote after the ID in any of those shapes (_LEAD_RE), and
  every quote chained to it; a reference with no quote must still
  name a real rule, section, mechanism or decisions entry
- quotes against a rule's statement AND its Accepted: clauses (example
  and pointer lines excluded), a section letter's Background
  (`rules.md#H Background`), a decisions entry's text
- matches lowercase decisions keys, merges the `### differential-
  ledger, ...` arc headings onto their one key, rejects empty "[...]"
  fragments, reads `--` as the em dash and a nested ' as ", follows
  `#:` comment continuations and IDs wrapped at a hyphen
- skips this module's own illustrative comments
- records the retired excerpts as a negative control, and asserts the
  sweep reaches each shape (non-colon, chained, decisions, elided)

The colon form alone still defines a citing module for implemented:.

Fixed: P5 (the reserve's numeral clause and the join-after-the-run
wording, #614), P2 (S2's once-read run), M2 (#601 replaced the walk
with the clause-free trailing run; the 'Jane Doe Jr. nee Smith Ma'
paragraph gains a dated note that #601 moved it to fix(#601/#602)), S2
("OPENS THE STRING"), A1's segmenter clause, a decisions.md#S2 citation
of rules.md#S2's Accepted clause, and four ledgers crediting AGENTS.md's
"a reader would hesitate too" to rules.md#A1. Four ledgers' paraphrase
of decisions.md#S2 now quotes its headline and names #544's second
exception. With Accepted text quotable, the first commit's H5 and S2
rewrites go back to the original, more exact Accepted-clause quotes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Fix #632's second-round review: reach, controls, stale prose

Scanner (tests/v2/test_doc_citations.py):
- a section's Background is read whole, not its first line; a wrapped
  no-boundary: marker's continuation lines are no longer quotable
- each citation records its lead, and the reach test asserts every
  lead shape is live ('s, parenthesis, dash, bare words) rather than
  any one of them; a new test pins what is quotable (an Accepted
  clause, a Background's wrapped line) and what is not (an example
  line, a no-boundary: continuation)
- the docstring states the limits: case folds, and a quote more than
  four words from its ID and chained to nothing is not checked -- the
  measured alternative, every quote in the paragraph, flags 139, nearly
  all names written in double quotes
- _RETIRED_EXCERPTS drops its P6 row (no version ever routed a
  decisions quote to a rule) and says which direction it guards

Ledgers: fourteen stale quotes the check does not reach, found by the
review -- C1's paired-initials sentences in all five (#563 made the
speaking word an UNAMBIGUOUS suffix), N3 in three, P2 in one. Dated
landing paragraphs for P2 (#424) and P5 (#425) now say they describe
the rule as it then stood instead of quoting today's text as their
reason; the M2 rule's opener names all four names #601 moved, not one.

AGENTS.md: the count is measured, not summed from commit messages --
27 stale quotes and 5 quoteless colon citations, the current check
over e4b653a -- and the sentence states the four-word reach.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Name the count's comparator by its master commit

e4b653a lived on #631's branch, which the squash merge leaves to be
deleted; 8e6e524, master after #631, has the identical tree.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Should John von PhD Jones start a credential run as John von und zu PhD Jones does?

1 participant