Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENTS.md

Large diffs are not rendered by default.

11 changes: 8 additions & 3 deletions docs/design/decisions.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion docs/design/mechanisms.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ Problem shape. A rule needs to know how words were JOINED (chained titles, parti

## UNIT-PARTITION — count the units the joining rules built

Problem shape. A rule counts "one name word", but the input holds words that another rule has already joined into one name — and the joining structure it would read has been merged away. Contract statement. Three rules build multi-word units: a particle chain (P2), a conjunction join (P3), and a bound given-name pair (P5). A rule that counts name words counts those units, and takes each whole or not at all — with one stated exception, rules.md#C1's count of the words before a comma, which counts particle surnames alone (decisions.md#C1, #575). How it works. group builds each join as a piece, but its own prefix chain then merges the joined piece into a longer one, so PIECES no longer carries the boundary — the units are rebuilt from the tags the vocabulary layer left (`particle`, `conjunction`, `vocab:bound-given`). The rebuild is RECURSIVE: what a conjunction or a bound word joins is the next UNIT, not the next word, and absorbing a single token instead strands a particle severed from the words it chains. Note the two joins arrive here for opposite reasons — the conjunction join was built and then swallowed, while the bound-given join was never built at all (P5 joins only at the first non-title piece), so restoring piece boundaries in group would fix the first and silently split the second. Lives in. nameparser/_pipeline/_vocab.py (`unit_ends`, the walk, shared since #575) and its two readers: _post_rules.py (`_units`, over classify's tags, for P1's fold) and rules.md#C1's counts before a comma: `_comma._whole_name`, over classify's tags (`surname_unit_facts`) since #613 moved the decision after classify, and `_vocab.name_word_count`, which builds its facts from the vocabulary (`surname_unit_tags`, held to classify's tags by an agreement test) -- a view script_segment needed while it asked a hand copy of C1 before classify ran, until #613's /simplify had it ask `_comma.decide` itself on a classified copy. A fourth walk is S2's read (#614): `_pieces._chain_units` marks each piece by the unit P2's chain WILL put it in -- opening a unit, a name word the chain joins, or a particle inside the run -- asking the chain's own `chain_run_end`, which `_group`'s chain calls too; a word the reading weighs (`_weighed`) keeps a unit of its own, as a position always counted it. The trailing peel's words to spare and #602's run start count in those marks, the peel stops at a particle inside a run, and everything else -- the numeral fork's look at the word in front, the H5 chain, the run's title test -- reads the pieces as written. It runs before the chain merges anything, which is why it cannot read the merged tags `unit_ends` walks; #614 had rebuilt the chain as a merged COPY instead, whose three defects in one review -- words the reading weighs hidden inside a unit, indices drifting from the pieces' -- were each patched in that copy, and #620 replaced it with the marks (decisions.md#S2). C1 counts particle chains only, reaching one word as the fold does: a bound pair builds a given name, and whether P3 joins a connective depends on the whole name the count is helping to decide, so neither is one surname before a comma (`_vocab.SURNAME_UNIT_TAGS`, decisions.md#C1). rules.md P1 and C1 are the counting rules, P2/P3/P5 the joining ones. Reach for it when. A rule says "one name word" and the input can contain a join — enumerate the joining rules out of rules.md rather than the ones you remember.
Problem shape. A rule counts "one name word", but the input holds words that another rule has already joined into one name — and the joining structure it would read has been merged away. Contract statement. Three rules build multi-word units: a particle chain (P2), a conjunction join (P3), and a bound given-name pair (P5). A rule that counts name words counts those units, and takes each whole or not at all — with one stated exception, rules.md#C1's count of the words before a comma, which counts particle surnames alone (decisions.md#C1, #575). How it works. group builds each join as a piece, but its own prefix chain then merges the joined piece into a longer one, so PIECES no longer carries the boundary — the units are rebuilt from the tags the vocabulary layer left (`particle`, `conjunction`, `vocab:bound-given`). The rebuild is RECURSIVE: what a conjunction or a bound word joins is the next UNIT, not the next word, and absorbing a single token instead strands a particle severed from the words it chains. Note the two joins arrive here for opposite reasons — the conjunction join was built and then swallowed, while the bound-given join was never built at all (P5 joins only at the first non-title piece), so restoring piece boundaries in group would fix the first and silently split the second. Lives in. nameparser/_pipeline/_vocab.py (`unit_ends`, the walk, shared since #575) and its two readers: _post_rules.py (`_units`, over classify's tags, for P1's fold) and rules.md#C1's counts before a comma: `_comma._whole_name`, over classify's tags (`surname_unit_facts`) since #613 moved the decision after classify, and `_vocab.name_word_count`, which builds its facts from the vocabulary (`surname_unit_tags`, held to classify's tags by an agreement test) -- a view script_segment needed while it asked a hand copy of C1 before classify ran, until #613's /simplify had it ask `_comma.decide` itself on a classified copy. A fourth walk is S2's read (#614): `_pieces._chain_units` marks each piece by the unit P2's chain WILL put it in -- opening a unit, or a name word the chain joins -- asking the chain's own `chain_run_end`, which `_group`'s chain calls too; a word the reading weighs (`_weighed`) keeps a unit of its own, as a position always counted it. The particles straight behind each other are merged ahead of the read, into the piece the chain would open its run with (`_pieces.merge_particle_runs`, #625), so a particle run is one piece the read takes whole or not at all, and no mark stands for a particle inside it. The trailing peel's words to spare and #602's run start count in those marks, the peel stops at a word the chain joins, and everything else -- the numeral fork's look at the word in front, the H5 chain, the run's title test -- reads the pieces as written. It runs before the chain joins name words to a particle run, which is why it cannot read the merged tags `unit_ends` walks; #614 had rebuilt the chain as a merged COPY instead, whose three defects in one review -- words the reading weighs hidden inside a unit, indices drifting from the pieces' -- were each patched in that copy, and #620 replaced it with the marks (decisions.md#S2). C1 counts particle chains only, reaching one word as the fold does: a bound pair builds a given name, and whether P3 joins a connective depends on the whole name the count is helping to decide, so neither is one surname before a comma (`_vocab.SURNAME_UNIT_TAGS`, decisions.md#C1). rules.md P1 and C1 are the counting rules, P2/P3/P5 the joining ones. Reach for it when. A rule says "one name word" and the input can contain a join — enumerate the joining rules out of rules.md rather than the ones you remember.

## MARK-DONT-STRIP — record the decision, keep the fact

Expand Down
20 changes: 16 additions & 4 deletions nameparser/_pipeline/_group.py
Original file line number Diff line number Diff line change
Expand Up @@ -65,8 +65,8 @@
from nameparser._pipeline._pieces import (
Peel, TailRead, chain_lead, chain_run_end, is_conj_piece, is_leading_title, is_prefix_piece,
is_suffix_piece, is_title_piece, join_connectives, joined_tags,
leading_titles, merge_pieces, peel_walk, read_trailing_run,
tail_reading, trailing_candidates, trailing_start,
leading_titles, merge_particle_runs, merge_pieces, peel_walk,
read_trailing_run, tail_reading, trailing_candidates, trailing_start,
trailing_start_past_titles,
)
from nameparser._pipeline._state import (
Expand Down Expand Up @@ -785,8 +785,17 @@ def _group_segment(seg: tuple[int, ...], additional: int,
# part past a comma and a segment of fewer than three pieces --
# which no join reaches anyway -- read as before, and so does a
# run holding a name piece (_pieces.read_trailing_run).
# The particles straight behind each other are merged first,
# into the one piece the chain would open its run with, so the
# read counts and absorbs the run as that piece (#625);
# `premerged` keeps which runs the merge made, for the chain's
# report below.
premerged: Set[int] = frozenset()
if site is ClauseSite.TRAILING:
found = read_trailing_run(pieces, ptags, tokens, one_case)
n = leading_titles(pieces, ptags, tokens)
premerged, lead = merge_particle_runs(pieces, ptags, tokens, n)
found = read_trailing_run(pieces, ptags, tokens, one_case, n,
lead)
if found is not None:
read, start = found
tail_pieces, tail_ptags = pieces[start:], ptags[start:]
Expand Down Expand Up @@ -964,10 +973,13 @@ def chain(tail: int) -> None:
# piece -- nothing
# was chained, and _assign reports that case instead.
# Without this the two emitters both fire on the same token.
# A run merged ahead of the read (`premerged`, #625) has
# claimed the particles behind it already, so its chain
# is a decision though `j == k + 1` ('Freiherr von der').
# (Tag test first: it is a set lookup and almost no name
# has an ambiguous particle, while is_leading_title is a
# call per piece.)
if (j > k + 1
if ((j > k + 1 or pieces[k][0] in premerged)
and "vocab:particle-ambiguous"
in tokens[pieces[k][0]].tags):
while titled < k and is_leading_title(
Expand Down
Loading
Loading