Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@ Three committed contributor docs carry the parser's normative rules and their re

**Counting claims.** A bare count in prose is either an assertion or a liability, keyed by who observes its staleness: asserted counts (a test holds the number) fail CI at change time — the useful kind; dated snapshots ("51 sites at spec time") cannot go stale; standing present-tense prose counts are the forbidden class — promote to an assertion, add a date, or state the invariant and let a test count. After changing how many times something runs, sweep for counts, not for the thing's name.

**Release-log claims.** Quantified or universal behavior claims in release bullets must come from the differential gate's classified summary, be verified against rules.md examples, or -- for a view the gate cannot see -- carry a recompute recipe stored with the design entry the bullet cites; never write one from memory. The classified summary covers the CONTRACT tier plus whatever radar diffs a rule classifies; a radar corpus's unmatched diffs are listed under UNCLASSIFIED (radar) and are not in it, so a claim quantified from the summary alone is silent about them. The gate compares the seven role fields, `_ambiguities` and -- since #484 -- `initials()`, under the `_initials` pseudo-field; so an initials-only change DOES show in the classified summary. But that pseudo-field sees only the names whose initials moved WHILE EVERY ROLE AND EVERY AMBIGUITY KIND STAYED PUT -- main()'s roles-identical guard keeps it out of any diff a role or a report is already in -- so a count taken from it is a FLOOR on initials movement, not the population: measured 2026-09-02 at 2.1.0 → tree on the v2 surface, 83 of the 1120 compared ENTRIES (1116 distinct names; seven entries carry a declared order rather than the default, and three strings are compared under more than one) changed their `initials()` string and only 28 were visible under `_initials`, the other 55 having moved a role as well. A bullet about how many names' initials changed still needs the recompute recipe. `capitalized()` and any other render view stay invisible to the gate (decisions.md#R4, #R3), and for those the first two sources still cannot reach a claim: a gate run is byte-identical across the change, and an example line witnesses an output without counting anything. A recipe names the corpus files, the policy sweep, and -- the part that is easy to omit and fatal -- THE COMPARATOR, which must be something the shipped tree is not: #408's first recipe said to compare `initials()` against a folded-first partition, which is what `initials()` now IS, so it reproduced 0 where the bullet claimed 660 and was the only stated provenance for the number. Run the recipe as written before shipping the bullet. Cross-version numbers (a released wheel, the pre-change tree) are dated snapshots under Counting claims, since nothing in the repository re-runs them. Per-rule ledger toml comments asserting PARSER behavior cite rule IDs under the excerpt discipline; free prose is for ledger mechanics only (owned by tools/differential/README.md).
**Release-log claims.** Quantified or universal behavior claims in release bullets must come from the differential gate's classified summary, be verified against rules.md examples, or -- for a view the gate cannot see -- carry a recompute recipe stored with the design entry the bullet cites; never write one from memory. The classified summary covers the CONTRACT tier plus whatever radar diffs a rule classifies; a radar corpus's unmatched diffs are listed under UNCLASSIFIED (radar) and are not in it, so a claim quantified from the summary alone is silent about them. The gate compares the seven role fields, `_ambiguities` and -- since #484 -- `initials()`, under the `_initials` pseudo-field; so an initials-only change DOES show in the classified summary. But that pseudo-field sees only the names whose initials moved WHILE EVERY ROLE AND EVERY AMBIGUITY KIND STAYED PUT -- main()'s roles-identical guard keeps it out of any diff a role or a report is already in -- so a count taken from it is a FLOOR on initials movement, not the population: measured 2026-09-02 at 2.1.0 → tree on the v2 surface, 83 of the 1120 compared ENTRIES (1116 distinct names; seven entries carry a declared order rather than the default, and three strings are compared under more than one) changed their `initials()` string and only 28 were visible under `_initials`, the other 55 having moved a role as well. A bullet about how many names' initials changed still needs the recompute recipe. `capitalized()` and any other render view stay invisible to the gate (decisions.md#R4, #R3), and for those the first two sources still cannot reach a claim: a gate run is byte-identical across the change, and an example line witnesses an output without counting anything. A recipe names the corpus files, the policy sweep, and -- the part that is easy to omit and fatal -- THE COMPARATOR, which must be something the shipped tree is not: #408's first recipe said to compare `initials()` against a folded-first partition, which is what `initials()` now IS, so it reproduced 0 where the bullet claimed 660 and was the only stated provenance for the number. Run the recipe as written before shipping the bullet. Cross-version numbers (a released wheel, the pre-change tree) are dated snapshots under Counting claims, since nothing in the repository re-runs them. Per-rule ledger toml comments asserting PARSER behavior cite rule IDs under the excerpt discipline, and since 2026-10-10 tests/v2/test_doc_citations.py checks them as it checks code. A ledger comment is copied from ledger to ledger and reworded by nothing: measured 2026-10-10 by running the check as it now stands over the tree before it (8e6e524a, master after #631), 27 quotes were stale and 5 colon citations quoted nothing, beyond the five #631's review found by hand -- most in a shape the colon-only check could not read (`rules.md#X -- "..."`, `rules.md#X's Accepted clause ("...")`, a second quote chained by "and"). A rule's statement and its `Accepted:` clauses are quotable, its example lines are not; `rules.md#H` names the H section, whose Background is quotable; a decisions.md citation may paraphrase, but what it quotes must be verbatim; and a reference with no quote must still name something that exists. The check's reach is a quote within four words of the ID or chained to one that is: a quote further off reads exactly like a name written in double quotes, so it is NOT checked, and the second review of #632 found 14 such quotes stale by hand. Keep a quote beside its ID. Quotes of AGENTS.md itself are not checked at all. Free prose is for ledger mechanics only (owned by tools/differential/README.md).

**Writing the user docs (docs/*.rst) has its own AGENTS.md too.** `docs/AGENTS.md` carries the style the docs are written in — tables as indexes, subheadings per task, measured claims, recipe doctests — distilled from PR #588. It loads the same way as the docs/design/ one below; if your tool does not do nested discovery, read it before editing a `.rst` file under docs/.

Expand Down
6 changes: 3 additions & 3 deletions nameparser/_pipeline/_script_segment.py
Original file line number Diff line number Diff line change
Expand Up @@ -715,9 +715,9 @@ def _split_surname_site(state: ParseState) -> ParseState:
for j in state.segments[0]):
return state
# No try/except around the call: rules.md#A1's Accepted clause
# ("a user-supplied segmenter's own error propagates"). The two
# checks below are that same doctrine, curated,
# and they are where the line this module draws is easiest to state:
# ("a user-supplied segmenter's own error, which propagates"). The
# two checks below are that same doctrine, curated, and they are
# where the line this module draws is easiest to state:
# a PROTOCOL VIOLATION BY THE SEGMENTER AUTHOR RAISES, while an
# ADAPTER'S DEFENSE AGAINST ITS LIBRARY DECLINES. Both checks here
# are the first kind -- a wrong answer TYPE and an answer indexing
Expand Down
2 changes: 1 addition & 1 deletion tests/v2/pipeline/test_pieces.py
Original file line number Diff line number Diff line change
Expand Up @@ -426,7 +426,7 @@ def test_the_walks_own_leading_piece_never_anchors_what_follows_it(

def test_a_reserve_kept_leading_piece_beside_a_genuine_family_loss(
) -> None:
# decisions.md#S2's Accepted boundary ("an unambiguous suffix is
# rules.md#S2's Accepted clause ("an unambiguous suffix is
# consumed even when that leaves no family name at all", 'Smith
# Jr.' -> family='') applies just the same when the LEADING piece
# is itself listed suffix vocabulary rather than an ordinary name:
Expand Down
Loading
Loading