Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 9 additions & 3 deletions docs/design/decisions.md

Large diffs are not rendered by default.

16 changes: 12 additions & 4 deletions docs/design/rules.md
Original file line number Diff line number Diff line change
Expand Up @@ -243,8 +243,8 @@ H4. Rationale: this is a name parser, not a title parser. Handed a
vocabulary reports `title-or-name` as well, the fork there being
whether the title word inside the unit is a title at all. Usually
a join (P3), and a particle chain (P4) is the same shape and
reports too — `St St née` reads family "St née" with `st` title
vocabulary inside it. The
reports too — `St van Bishop` reads family "van Bishop" with
`bishop` title vocabulary inside it. The
clause reaches EVERY join whose non-leading member is TITLES
vocabulary, not the one word that prompted it: `Smith and King`,
`John and King`, `Smith and Bishop` and `John of Judge` all
Expand Down Expand Up @@ -679,12 +679,20 @@ P4. Rationale: a particle links forward from inside a name; at the
nothing (the title is not a name word), and why "Van Johnson"
is a given-name reading at all. An unlisted abbreviation before
the particle is as transparent as a listed title, since assign
reads it as one (H2).
reads it as one (H2). A word in the leading titles that is both a
title and a particle puts the name's leading position on itself,
so the particle behind it is inside a name ('Freiherr von Berg');
of several such words the last does, so no particle chains inside
the titles (#624). The titles are those written, before H3 gives
the last one back to the name, so each such word reads as a title
unless H3 gives it back.
"Van Johnson" → given="Van"
"Sir de Mesnil" → pieces=[["Sir"], ["de"], ["Mesnil"]]
"Xyz. van Johnson" → given="van"
"John van der Berg" → pieces=[["John"], ["van", "der", "Berg"]] · boundary
history: decisions.md#P2 · interacts: P1, P5, H2 · implemented: nameparser/_pipeline/_group.py
"Freiherr St John Smith" → title="Freiherr St"
"Freiherr St John Smith" → given="John"
history: decisions.md#P2 · interacts: P1, P5, H2, H3 · implemented: nameparser/_pipeline/_group.py, nameparser/_pipeline/_pieces.py

P5. Rationale: some given-name words are incomplete alone — "abdul"
is a bound form that the next word completes.
Expand Down
5 changes: 3 additions & 2 deletions nameparser/_pipeline/_assign.py
Original file line number Diff line number Diff line change
Expand Up @@ -134,8 +134,9 @@ def _absorbed(piece: Sequence[int],
# rules.md#H2: "an abbreviation opening the part of the name that
# carries the given name — the whole name, or the part after a
# family comma — reads as a title even when unlisted" -- the count is
# _pieces.leading_titles since #424 (its test, is_leading_title, is
# the leading-particle scan's too); the roles are set here.
# _pieces.leading_titles since #424 (the chain's leading position is
# read off the same run, _pieces.chain_lead, #624); the roles are set
# here.
def _peel_leading_titles(pieces: tuple[tuple[int, ...], ...],
ptags: tuple[frozenset[str], ...],
tokens: list[WorkToken]) -> int:
Expand Down
72 changes: 35 additions & 37 deletions nameparser/_pipeline/_group.py
Original file line number Diff line number Diff line change
Expand Up @@ -63,7 +63,7 @@

from nameparser._lexicon import _run_addresses_by_given
from nameparser._pipeline._pieces import (
Peel, TailRead, chain_run_end, is_conj_piece, is_leading_title, is_prefix_piece,
Peel, TailRead, chain_lead, chain_run_end, is_conj_piece, is_leading_title, is_prefix_piece,
is_suffix_piece, is_title_piece, join_connectives, joined_tags,
leading_titles, merge_pieces, peel_walk, read_trailing_run,
tail_reading, trailing_candidates, trailing_start,
Expand Down Expand Up @@ -819,18 +819,18 @@ def _group_segment(seg: tuple[int, ...], additional: int,
# name text parsed two ways ("Van Johnson" -> given Van, family
# Johnson; "Dr. Van Johnson" -> family "Van Johnson").
#
# "Title AND NOT prefix" rather than the plain "not a title" the
# rule is stated as, and the difference is not academic: `st`,
# `do` and `freiherr` are each BOTH a title and an ambiguous
# particle, so the plain test skipped over the very piece the
# exception exists to protect and "St John Smith" -- no title in
# front of it at all -- collapsed from title St, given John,
# family Smith into one given "St John Smith". A piece that
# could be the name's own first piece stops the scan; only a
# piece that can ONLY be a title is stepped over.
# A word that is BOTH a title and a particle (`st`, `freiherr`)
# in the title run leads in the run's place, and the last such
# word does (`chain_lead`, #624): a scan stepping over every
# title skipped the very piece the exception exists to protect,
# and "St John Smith" -- no title in front of it at all --
# collapsed from title St, given John, family Smith into one
# given "St John Smith"; a scan stopping at the FIRST such word
# chained the second, so "Freiherr St John Smith" read family
# "St John Smith" while assign read both words as titles.
#
# Computed once, before the loop: every merge below starts at
# some k at or past this index, so no merge can move it.
# some k past this index, so no merge can move it.
#
# Suffix pieces are deliberately NOT skipped, and the reason is
# what skipping them WOULD do rather than what it would cost.
Expand Down Expand Up @@ -864,21 +864,20 @@ def _group_segment(seg: tuple[int, ...], additional: int,
# take it" had to restate the consumer's condition and got it
# wrong one suffix later (#417).
#
# The `, 0` fallback is inert by construction rather than a
# default worth testing: it is reached only when every piece is
# a title and none is a prefix, and the loop below merges
# nothing unless some piece is a prefix.
# `is_title_piece` alone missed H2's unlisted abbreviations, which
# assign peels as titles all the same, so 'Xyz. van Johnson'
# chained where 'Dr. van Johnson' did not (#424 found it
# through the acronym fork: the chain had swallowed the given
# word and left assign two pieces where the fork counted
# three). The scan asks assign's own test.
leading = next((k for k in range(len(pieces))
if not is_leading_title(pieces[k], ptags[k],
tokens)
or is_prefix_piece(pieces[k], ptags[k], tokens)),
0)
# The leading position is asked of assign's own title run
# (`leading_titles`): `is_title_piece` alone missed H2's
# unlisted abbreviations, which assign peels as titles all the
# same, so 'Xyz. van Johnson' chained where 'Dr. van Johnson'
# did not (#424 found it through the acronym fork: the chain
# had swallowed the given word and left assign two pieces where
# the fork counted three). And `chain_lead` is the one answer
# the trailing read's count takes too: of two titles that are
# also particles the second leads (#624), where a scan of this
# loop's own had stopped at the first and chained the second,
# so 'Freiherr St John Smith' read family 'St John Smith' while
# assign read both words as titles.
name_start = leading_titles(pieces, ptags, tokens)
leading = chain_lead(pieces, ptags, tokens, name_start)
# rules.md#P2: "a trailing suffix begins" -- where it begins
# is read by assign's peel over the pieces as they stand
# (#424), once per segment and kept as a length from the end,
Expand All @@ -894,7 +893,6 @@ def _group_segment(seg: tuple[int, ...], additional: int,
# takes both forks, and where no run was read ahead of it it
# asks again after its merges whether the acronym still has
# the pieces the fork counted (below).
name_start = leading_titles(pieces, ptags, tokens)
# the run read above is already split off (#614), so the chain
# stops at the end of what is left
tail = (0 if read is not None
Expand All @@ -914,7 +912,7 @@ def chain(tail: int) -> None:
titled = 0
k = 0
while k < len(pieces):
if k == leading or not is_prefix_piece(pieces[k],
if k <= leading or not is_prefix_piece(pieces[k],
ptags[k], tokens):
k += 1
continue
Expand All @@ -929,15 +927,15 @@ def chain(tail: int) -> None:
# needs an emitter in each.
#
# Narrow, and #367 is why. `titled == k` says every
# piece ahead of this one is a title, and the
# loop skipped k == leading, so `leading` is STRICTLY
# before k -- and being before k it is one of those titles,
# while being `leading` it satisfies `not title or prefix`.
# For both, it must be a prefix as well: a word in both
# vocabularies (`st`, `do`, `freiherr` by default, or any
# overlap a caller configures). A plain title alone can no
# longer put a particle off the name's leading piece; it is
# stepped over and _assign reports the fork instead.
# piece ahead of this one is a title, and the loop
# skipped every piece up to `leading`, so `leading` is
# STRICTLY before k -- one of those titles, and, being
# before the title run's end, a prefix as well
# (`chain_lead`): a word in both vocabularies (`st`,
# `freiherr` by default, or any overlap a caller
# configures). A plain title alone can no longer put a
# particle off the name's leading piece; it is stepped
# over and _assign reports the fork instead.
#
# What that leaves is wider than one shape: any number of
# plain title pieces, then a piece in BOTH vocabularies,
Expand Down
89 changes: 57 additions & 32 deletions nameparser/_pipeline/_pieces.py
Original file line number Diff line number Diff line change
Expand Up @@ -140,6 +140,44 @@ def is_title_piece(piece: Sequence[int], ptags: Set[str],
return len(piece) == 1 and "vocab:title" in tokens[piece[0]].tags


# rules.md#P4: "a particle in the name's leading position chains
# nothing" -- WHERE that position is, asked once, by group's chain and
# the trailing read's unit count (#624)
def chain_lead(pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]],
tokens: Sequence[WorkToken], n: int) -> int:
"""The particle chain's leading position, given `n`, the end of
the leading title run (`leading_titles`): the last title in the run
that is also a particle, else the first piece past the run. A title
is no name word, so the first piece past the run leads; a word in
both vocabularies ('Freiherr', 'St') leads in its stead and puts
the particle behind it inside a name ('Freiherr von Berg', 'St van
Johnson'), and of several such words the last leads, all being
titles ('Freiherr St John Smith' as 'Dr. St John Smith'). No
particle stands between the lead and the run's end, so the chain
opens no unit inside the titles.

The run is read as walked, before H3's give-back: a title the run
hands back to the name because only suffix words follow it is
still a title to the chain, so the particle behind it is the
name's leading piece and chains nothing (P4: 'Dr. Mc Mc' keeps
'Mc' and 'Mc' apart, as before #624). `leading_titles` stops at a
piece that is itself a leading title only where it gave that piece
back, before a suffix piece (the tags it tests inline first), or at
the segment's last piece, which a title may not be unless it is the
whole segment -- the `n + 1 < len(pieces)` test rules that one out.
The same tags go first here, so a name with no suffix behind its
titles pays no frame for the test."""
if (n + 1 < len(pieces)
and ("suffix" in ptags[n + 1]
or "vocab:suffix" in tokens[pieces[n + 1][0]].tags)
and is_leading_title(pieces[n], ptags[n], tokens)):
n += 1
for k in range(n - 1, -1, -1):
if is_prefix_piece(pieces[k], ptags[k], tokens):
return k
return n


# A particle, or a piece a join made one: what group's prefix chain
# (at its loop in _group_segment) chains and stops at, what P3's join
# derives a prefix from, and what group's rootname count and leading
Expand Down Expand Up @@ -304,8 +342,8 @@ def join_connectives(pieces: list[list[int]], ptags: list[set[str]],
# still reaches it as `_pieces._PERIOD_ABBREV` -- an import binds the
# same name here, so the sync test's target did not move. Out of
# assign since #424 and in the piece layer since #439: the test is
# assign's, and group's leading-particle scan and trailing-run walk
# must start where assign starts.
# assign's, and the chain's leading position (`chain_lead`) and the
# trailing-run walk must start where assign starts.


# rules.md#H2: "an abbreviation opening the part of the name that
Expand Down Expand Up @@ -353,10 +391,10 @@ def leading_titles(pieces: Sequence[Sequence[int]],
(rules.md#H3, decisions.md#H3 -- the block at the floor below
carries the examples of each half, and the ordering its two inline
tag reads were measured on).
One definition, read by assign (which sets the roles) and by the
chain's trailing-run walk; the leading-particle scan shares the
predicate, is_leading_title, but stops at a title-and-particle
word (P4, #367, #424)."""
One definition, read by assign (which sets the roles), by the
chain's trailing-run walk, and by `chain_lead`, which finds the
chain's leading position inside this run for group's chain and the
trailing read alike (P4, #367, #424, #624)."""
n = 0
while n < len(pieces):
if ((n + 1 < len(pieces) or len(pieces) == 1)
Expand Down Expand Up @@ -1708,7 +1746,7 @@ def chain_run_end(k: int, pieces: Sequence[Sequence[int]],
def _chain_units(pieces: Sequence[Sequence[int]],
ptags: Sequence[Set[str]],
tokens: Sequence[WorkToken],
n: int) -> tuple[list[int] | None, int]:
n: int) -> list[int] | None:
"""The name units P2's chain will make of `pieces`, as a mark per
piece -- OPENS where a piece opens a unit, JOINED where the chain
will join a name word to the run in front of it and the reading
Expand All @@ -1717,13 +1755,11 @@ def _chain_units(pieces: Sequence[Sequence[int]],
particle inside the run, a word the run took and so a name word
whatever else it is ('van mc', 'von vd': rules.md#S2, the words
both particles and suffix vocabulary standing straight behind a
particle) -- and where the name starts.
`n` is the end of the leading title run; the chain's leading
position is the first title that is also a particle ('Freiherr
von vd', 'St van Mc'), else `n`, and the chain opens no unit there.
A unit the chain opens INSIDE the titles makes a name of them
('Freiherr St van Berg MA', 'Freiherr Freiherr Prof do'), so the
first unit at or before `n` is where the name starts.
particle).
`n` is the end of the leading title run, and the chain starts past
its leading position (`chain_lead`, the one answer group's chain
takes too, #624), which opens no unit inside the titles: the name
starts at `n`.

Nothing is merged: the read counts with the flags and reads every
piece as written, so no word it weighs is hidden inside a unit
Expand All @@ -1737,15 +1773,9 @@ def _chain_units(pieces: Sequence[Sequence[int]],
and "particle" in tokens[pieces[k][0]].tags)
for k in range(count)]
if not any(prefix[1:]):
return None, n
leading = n
for k in range(n):
if prefix[k]:
leading = k
break
return None
units = [OPENS] * count
at = n
k = leading + 1
k = chain_lead(pieces, ptags, tokens, n) + 1
while k < count:
if prefix[k]:
j = chain_run_end(k, pieces, ptags, tokens, count)
Expand All @@ -1756,14 +1786,10 @@ def _chain_units(pieces: Sequence[Sequence[int]],
while q < j and not _weighed(pieces[q], tokens):
units[q] = JOINED
q += 1
# a unit with nothing past its opener is no unit the count
# sees, and makes no name of the titles it opens inside
if q > k + 1 and k < at:
at = k
k = j
continue
k += 1
return units, at
return units


def _weighed(piece: Sequence[int], tokens: Sequence[WorkToken]) -> bool:
Expand All @@ -1773,9 +1799,8 @@ def _weighed(piece: Sequence[int], tokens: Sequence[WorkToken]) -> bool:
word in it by shape, an initial, a roman numeral by shape, or a
period-marked title word. A word the chain joins and the reading
weighs is a word the reading may yet take ('Freiherr von Berg MA
X.Y.Z.' keeps suffix 'MA X.Y.Z.'), and one that opens a unit inside
the titles may leave the titles a title ('St St VI'), so it ends the
plain run the count folds into the particle (#620's review). A
X.Y.Z.' keeps suffix 'MA X.Y.Z.'), so it ends the plain run the
count folds into the particle (#620's review). A
connective join is one name word, P3's own count. The tests inline:
this asks once per word behind a particle run."""
if len(piece) > 1:
Expand Down Expand Up @@ -1810,8 +1835,8 @@ def read_trailing_run(pieces: Sequence[Sequence[int]],
n = leading_titles(pieces, ptags, tokens)
if n == len(pieces):
return None
units, at = _chain_units(pieces, ptags, tokens, n)
rest, titled, peel = tail_reading(peel_walk(at, ptags), pieces, ptags,
units = _chain_units(pieces, ptags, tokens, n)
rest, titled, peel = tail_reading(peel_walk(n, ptags), pieces, ptags,
tokens, one_case, units)
tail: set[int] = set()
for k in rest[peel.names:]:
Expand Down
Loading
Loading