diff --git a/AGENTS.md b/AGENTS.md index 25727bf4..f283e3fe 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -316,7 +316,7 @@ The 2.0 rewrite lands as underscore-private modules alongside the v1 code. These - **Method organization**, fixed section order in every class: fields + `__post_init__` validation → alternative constructors → dunders (construction/equality → protocol → operators) → properties → public methods by concern (access → editing → comparison → rendering delegates) → private helpers last, except a helper serving exactly one section may sit at that section's head. Sanctioned deviation, facade layer only: `HumanName` and the shim `Constants` organize by v1 concern groups (`# -- render defaults --`, `# -- config / parsing --`, `# -- fields --`, ..., dunders and pickle last) — the classes mirror v1's own surface and die in 3.0; the canonical order still binds every core type. - **Validation is eager and fail-loud**: every `raise` states the offending value, the expected form, and the fix. Exception taxonomy: wrong type — including wrong element type inside a collection, bare `str` where an iterable of strings is expected, or a `Mapping` where a plain iterable is expected — raises `TypeError`; well-typed but unacceptable values raise `ValueError`; failed enum lookups stay `ValueError` for any input (stdlib `EnumType` precedent). **When the message hands the reader code to paste, that code has to survive a type checker** — nameparser ships `py.typed`. #337's segmenterless warning offered `Policy(segment_scripts=())`, an `arg-type` error, because these fields are annotated with what they STORE rather than everything the constructor accepts. Prefer the `frozenset()` / `()` spellings in messages and docstrings, and pin the offered spelling in a test — the warning tests matched on `ja_segmenter` and never checked the actionable half of the message. **A warning emitted in `Parser.__post_init__` needs `parser_for` to re-emit it from its own frame** (the `catch_warnings(record=True)` block at its return): `__post_init__`'s `stacklevel` is sized for direct `Parser(...)` construction, and through `parser_for`'s extra frame the default one-line rendering attributes the warning to the library's own `return Parser(...)` — the exact call the message tells the user to change becomes invisible. No single stacklevel serves both entry points; a new construction warning gets the re-emission for free, but a new CONSTRUCTION SITE for `Parser` inside this package needs its own re-emission or its callers get library-attributed warnings (#337 review). - **Guard, hint, and emit for the WHOLE family, and parametrize the test over it**: a check added to one member of a set belongs on all of it, and the test must sweep the family, not one example. This session shipped `_reject_str_and_mapping` on `Policy` but not `PolicyPatch`, the bytes decode hint on three of five config entry points, and a regex-sync roster missing four of its copies — each a separate follow-up bug that a `{class} × {field} × {bad-value}` parametrization would have caught and a per-example test hid. When you find you're guarding member N, grep for the other members first. -- **Ambiguities are emitted at the DECISION site**: an `Ambiguity` records a fork the parse had to call, not a token that sits in an ambiguous vocabulary. Emit where the branch is taken — the trailing-suffix peel in `_assign`, the delimiter escape's follow-up in `classify` — never by scanning for a `vocab:*-ambiguous` tag. The same tagged token is a genuine fork in one position and unremarkable in another (`do` mid-name in "Joao da Silva do Amaral de Souza" chooses nothing). **A branch that runs but changes nothing is not a decision either** -- the prefix chain's `merge(k, j)` executes even when `j == k + 1`, folding a piece into itself, and keying on "the code got here" reported a fork for all ambiguous particles on "Do Van Jr." (`Dr.` when that was written, before #367 made a plain title transparent and put the shape out of the loop's reach entirely), where the particle stayed a lone leading name piece — the GIVEN name under the default order, the family name under `FAMILY_FIRST` — and `_assign` reported the same token again. Check that the branch actually claimed something (`j > k + 1`) before recording. Structure often settles the question before it arises, which is why `PARTICLE_OR_GIVEN` is not emitted on the `FAMILY_COMMA` path's WHOLLY-FAMILY read -- the comma fixed which piece is the family -- and `SUFFIX_OR_NAME` is not emitted for "Ma, Jack". A tail segment is the same case from the other side: assign reads it wholly as suffixes outside any maiden clause standing in it (`Jane Doe, PhD, Jr nee van Ma` keeps maiden `van Ma`), so group's two chain emitters are handed no report list there, after either comma — `John Smith, Jr., Freiherr von Richthofen` still chains `von` and reports only `comma-structure` (rules.md#C2, 2026-09-28). Read that scope narrowly: the comma settles nothing about a particle trailing the given name, so P6's attachment in `post_rules` decides that fork on the same path and reports it (#405), in the kind naming the reading it OVERRODE, which is the reading assign made and not the word's vocabulary: `SUFFIX_OR_NAME` where assign had read the run as a post-nominal (`vd`, `mc`), else `PARTICLE_OR_GIVEN` where the run holds an ambiguous particle (`van`, and `do`, which is in the suffix vocabulary too but in its AMBIGUOUS half, so no credential reading was overridden), else silence. The decision site also has the token index and the detail text in hand, which the tag scan would have to reconstruct. **If a fork's two branches are taken in DIFFERENT stages, every one of them needs the emitter** -- `PARTICLE_OR_GIVEN` is decided in `_assign` when the ambiguous particle stays a lone leading piece, in `_group` when something shifts it off the name's leading piece and the prefix chain claims it, and in `post_rules` when P6's attachment takes a trailing particle into the family after a comma, so all three report; for two years only the first did. What can still do the shifting is narrow, and #367 is why: a plain title no longer can (`Dr. Van Johnson` reads as `Van Johnson` does and reports from `_assign`), so the `_group` emitter needs a word that is BOTH a title and a particle — measured, `TITLES ∩ particles_ambiguous` is `{freiherr, st}` in the default vocabulary (`do` left TITLES in #296's audit; decisions.md's Excluded block records the before and after), plus any overlap a caller's config creates — standing ahead of the chained particle as the LEADING NAME word. Titles may precede it, so `Dr. St van Johnson` reaches the emitter and `St van Johnson` does too; a given name may not, so `Jan Freiherr von Richthofen` does not reach it while `Freiherr von Richthofen` and `Dr. Freiherr von Richthofen` do. Two shapes that look like they should reach it and do NOT, both measured by stepping `STAGES` and watching where `ambiguities` grows: `Dr. Do van Johnson` and `Do St Johnson` report from `assign`, not `group`, because `do` is no longer a title and so stays the leading name piece assign reports on — a both-vocabulary word CHAINED (`Jan St Johnson`) reports nothing at all. When checking whether that emitter is dead, a both-vocabulary word in the leading name position is the thing to look for, and the answer is that it is not dead. The stage-ownership map in `tests/v2/pipeline/test_state.py` must list `ambiguities` for each such stage, and it passes vacuously until a case row exercises the path, so add the row too. Report BOTH directions of a two-way fork — "John Smith MA" (read as a suffix) and "Jack MA" (read as the family name) are equally guesses. Every kind needs a trigger in `tests/v2/test_contracts.py::_AMBIGUITY_TRIGGERS` (an explicit `None`, strict-xfail, while reserved), and case-table rows pin expected kinds exactly, so a new emitter shows up in both immediately. **Pin the decision, not the vocabulary**: the only titled-particle test used an UNAMBIGUOUS particle, so it walked the right code path and proved nothing about the branch under test -- two criticals passed 1539 tests. A row contrasting the two readings ("John Smith V" against "John Smith B") is what makes an emitter's absence meaningful. +- **Ambiguities are emitted at the DECISION site**: an `Ambiguity` records a fork the parse had to call, not a token that sits in an ambiguous vocabulary. Emit where the branch is taken — the trailing-suffix peel in `_assign`, the delimiter escape's follow-up in `classify` — never by scanning for a `vocab:*-ambiguous` tag. The same tagged token is a genuine fork in one position and unremarkable in another (`do` mid-name in "Joao da Silva do Amaral de Souza" chooses nothing). **A branch that runs but changes nothing is not a decision either** -- the prefix chain's `merge(k, j)` executes even when `j == k + 1`, folding a piece into itself, and keying on "the code got here" reported a fork for all ambiguous particles on "Do Van Jr." (`Dr.` when that was written, before #367 made a plain title transparent and put the shape out of the loop's reach entirely), where the particle stayed a lone leading name piece — the GIVEN name under the default order, the family name under `FAMILY_FIRST` — and `_assign` reported the same token again. Check that the branch actually claimed something (`j > k + 1`) before recording. Structure often settles the question before it arises, which is why `PARTICLE_OR_GIVEN` is not emitted on the `FAMILY_COMMA` path's WHOLLY-FAMILY read -- the comma fixed which piece is the family -- and `SUFFIX_OR_NAME` is not emitted for "Ma, Jack". A tail segment is the same case from the other side: assign reads it wholly as suffixes, a maiden marker there included since #601 (`Jane Doe, PhD, Jr nee van Ma` reads suffix `PhD, Jr nee van Ma`, the marker an ordinary word, rules.md#M2), so group's two chain emitters are handed no report list there, after either comma — `John Smith, Jr., Freiherr von Richthofen` still chains `von` and reports only `comma-structure` (rules.md#C2, 2026-09-28). Read that scope narrowly: the comma settles nothing about a particle trailing the given name, so P6's attachment in `post_rules` decides that fork on the same path and reports it (#405), in the kind naming the reading it OVERRODE, which is the reading assign made and not the word's vocabulary: `SUFFIX_OR_NAME` where assign had read the run as a post-nominal (`vd`, `mc`), else `PARTICLE_OR_GIVEN` where the run holds an ambiguous particle (`van`, and `do`, which is in the suffix vocabulary too but in its AMBIGUOUS half, so no credential reading was overridden), else silence. The decision site also has the token index and the detail text in hand, which the tag scan would have to reconstruct. **If a fork's two branches are taken in DIFFERENT stages, every one of them needs the emitter** -- `PARTICLE_OR_GIVEN` is decided in `_assign` when the ambiguous particle stays a lone leading piece, in `_group` when something shifts it off the name's leading piece and the prefix chain claims it, and in `post_rules` when P6's attachment takes a trailing particle into the family after a comma, so all three report; for two years only the first did. What can still do the shifting is narrow, and #367 is why: a plain title no longer can (`Dr. Van Johnson` reads as `Van Johnson` does and reports from `_assign`), so the `_group` emitter needs a word that is BOTH a title and a particle — measured, `TITLES ∩ particles_ambiguous` is `{freiherr, st}` in the default vocabulary (`do` left TITLES in #296's audit; decisions.md's Excluded block records the before and after), plus any overlap a caller's config creates — standing ahead of the chained particle as the LEADING NAME word. Titles may precede it, so `Dr. St van Johnson` reaches the emitter and `St van Johnson` does too; a given name may not, so `Jan Freiherr von Richthofen` does not reach it while `Freiherr von Richthofen` and `Dr. Freiherr von Richthofen` do. Two shapes that look like they should reach it and do NOT, both measured by stepping `STAGES` and watching where `ambiguities` grows: `Dr. Do van Johnson` and `Do St Johnson` report from `assign`, not `group`, because `do` is no longer a title and so stays the leading name piece assign reports on — a both-vocabulary word CHAINED (`Jan St Johnson`) reports nothing at all. When checking whether that emitter is dead, a both-vocabulary word in the leading name position is the thing to look for, and the answer is that it is not dead. The stage-ownership map in `tests/v2/pipeline/test_state.py` must list `ambiguities` for each such stage, and it passes vacuously until a case row exercises the path, so add the row too. Report BOTH directions of a two-way fork — "John Smith MA" (read as a suffix) and "Jack MA" (read as the family name) are equally guesses. Every kind needs a trigger in `tests/v2/test_contracts.py::_AMBIGUITY_TRIGGERS` (an explicit `None`, strict-xfail, while reserved), and case-table rows pin expected kinds exactly, so a new emitter shows up in both immediately. **Pin the decision, not the vocabulary**: the only titled-particle test used an UNAMBIGUOUS particle, so it walked the right code path and proved nothing about the branch under test -- two criticals passed 1539 tests. A row contrasting the two readings ("John Smith V" against "John Smith B") is what makes an emitter's absence meaningful. - **A kind is worth adding only if a reader would hesitate too**: the test is not "does the code take a branch" but whether a person reading that input would genuinely be unsure. "Smith, John V" reads as a middle initial to anyone -- the comma settles it -- so reporting it would be noise that teaches callers to ignore the field, which costs more than the missing report. Reachability of the second branch is necessary, not sufficient. Prefer leaving a fork silent and documenting the omission over emitting on input nobody finds ambiguous. - **Parser owns config-dependent conveniences**: `Parser.matches`/`Parser.capitalized`/`Parser.revise` exist because the `ParsedName` equivalents fall back to DEFAULT config for str/omitted arguments (documented loudly in both docstrings). `revise` harvests tokens from a full sub-parse of each replacement value (tags kept minus `FOLDED_TAG`, roles forced, the R1 entry pass `suffix_entries` re-run over the forced state so a suffix value's entries follow its own commas, ambiguities discarded); the merge tail is shared with `replace()` via `ParsedName._with_field_tokens`. `Parser.capitalized` delegates through `name.capitalized(self.lexicon)` specifically so `_parser` never imports `_render` — keep it that way. - **Per-word vocabulary fields warn on multi-word entries** (`_normset`/`_normpairs` via `_warn_dead_entry`, UserWarning, never a raise — see the given_name_titles Gotcha for why raising is wrong). `given_name_titles` is the one multi-word-matched field and is exempt; `_edit` passes `warn=False` (add() warns once via the new instance's `__post_init__`; remove() stores nothing). The default vocabulary and every locale pack must stay warning-free (`test_default_lexicon_builds_warning_free`, `test_pack_vocabulary_entries_are_single_words`). @@ -385,7 +385,7 @@ Add a dedicated `copy.deepcopy()` round-trip test for it too (see `test_regexes_ **`_normalize` must reach a fixed point** — storage and match-time share the one fold, and `Lexicon.__setstate__` re-validates, so a value that changes on re-normalization changes under its owner. `strip().strip(".")` alone is not idempotent (`'. a .'` → `' a '` → `'a'`). The loop is the fix; keep any new stripping inside it. **Anything built on `_normalize` must converge too** — `_fold_words` runs `_normalize` per word and DROPS the words that fold away (`_title_key` is that list space-joined, and `_run_addresses_by_given` reads the list itself, so its last-word arm is the last word of the FOLDED key by construction); keeping the empty slot stored `'lt .'` as `'lt '`, a key match-time can never rebuild (so the entry is silently inert) and `__setstate__` rejects on the next round-trip as "not written by this version". -**Perf regressions are caught by the scaling test, not the absolute-time ones** — `tests/v2/test_benchmark.py::test_parse_cost_grows_no_worse_than_linearly` times a repeated unit at n vs 4n over a table of shapes (one per pipeline inner loop) and bounds the ratio; the `_thousand_names` tests use constant-size, delimiter-free input and are structurally blind to a complexity regression. Two rules when touching it: calibrate `_MAX_RATIO` against the WEAKEST quadratic's signal (a mixed quadratic surfaces far below the textbook 16×, so the operating point `_BASE` matters more than the bound), and confirm a planted regression fails it across REPEATED runs — one failure is a coin-flip on a timing test. The shapes cover different dimensions (segment count only via `commas`, intra-piece accumulation only via `particles`/`conjunctions`, non-ASCII input only via `honorifics` — every other unit is pure ASCII, so `script_segment` returns at its bail and the CJK stages go unmeasured, M2's clause view only via `maiden_clause`, whose unit has to END on a class member: `MA nee ` holds the same two words, the peel stops at the trailing marker, and the shape reaches nothing — and connective RUN LENGTH only via `link_run`, whose unit has to be MIXED CASE and hold a connective of the generational class: `i und ` reads the letter as an initial and reaches nothing where `i Und ` reaches everything, and S2's anchor pass only via `credential_run`, whose unit has to hold a Title-case member behind an unambiguous credential: `PhD MA ` reads the same fields and never asks the pass, the capitals deciding each member first); measure before pruning one. **A shape whose input needs a PREFIX cannot be a `_SHAPES` row at all**, since that table repeats a unit and nothing else — a maiden clause needs a name word and a marker before the run it is about, and `"nee i Und " * n` reaches the clause rule not at all, measuring the identical ratio on a broken tree and a fixed one. That one is `test_a_clause_link_run_does_not_cost_quadratically`, which builds its own input and counts FRAMES. A prefixed shape whose cost is Python-level joins `_REREAD_SHAPES` beside it — (input builder, the one function the defect re-runs, reachability probe) rows counted in that function's frames at 8 units against 32 — which holds #558's trailing fixed point, its maiden-clause twin, and #559's particle chain, a shape that grows at both ends and so has no unit to repeat. A prefixed shape whose cost is C-level work — a list scanned by `in`, a tuple re-hashed as a dict key — emits no frame for that count to see, so it goes in `_PREFIXED_SHAPES` instead: (prefix, unit, reachability probe) rows timed on the clock like `_SHAPES`, at a base of their own recorded beside the table (#553, whose run shape read 5.97× once at base 800 on the pre-fix tree, under the bound; py3.11, 2026-09-28). **That table skips under a line tracer** (`sys.settrace`, coverage's core below py3.14), which slows every Python line and not the C-level scan, diluting the quadratic below the bound (5.46–5.52× broken at base 1600 under `coverage run` on py3.11, same date). The `sys.monitoring` core coverage uses from 3.14 does not dilute it (8.87–10.91× broken), and CI collects coverage in its 3.14 build job alone, so every CI job runs the table; `ja-extra` also sets `NAMEPARSER_REQUIRE_CLOCK_GUARDS` so that a line tracer there fails the rows instead of skipping them. **Coverage is collected on ONE interpreter, the one that uploads it** (`COVERAGE_ON` in `.github/workflows/python-package.yml`): the package has no version-conditional code, so one report stands for all of them, and under a line tracer the property grids in `tests/v2/test_properties.py` had cost the 3.12 job most of its ~12 minutes (measured 2026-09-29: the three #397 grids take 139s under coverage's settrace core on 3.12, 24s under its sys.monitoring core, 17s with none). **A shape the CLOCK cannot reach needs a FRAME-count guard instead**, which is the second scaling test in that file (`test_a_trailing_credential_run_does_not_cost_exponentially`, #531): where the defect is an exponential rather than a quadratic, the input length that separates the curves on a timing test does not finish, so the guard counts frames over 8 units against 16 and bounds THAT ratio. One pair does not see every curve, and the fix round for #531 measured why: at 2× the input the per-member LINEAR work swamps a quadratic (2.08× for a genuine one against 1.73× clean), so that pair guards the exponential alone and a second, longer pair — 16 against 64, where the same quadratic reads 7.42× against 3.53× clean — is what can see one. Assert them in that order: an exponential never returns from the longer run, so the cheap pair has to have failed first. Frame counts do not move under load, so this shape needs no repeated-run calibration — but it does need the same reachability assertion `_POLICY_SHAPES` rows carry, since the walk under measurement runs only while every unit still reads as a credential. A stage gated on an opt-in `Policy` field needs a `_POLICY_SHAPES` entry instead, since bare `parse()` never enters it — and that table's rows carry a **reachability probe** run before the measurement, because a precedence change can quietly stop the shape reaching the stage and leave a green test measuring a no-op (`_POLICY_SHAPES` is also asserted non-empty: an empty `parametrize` is a skip, not a failure, so deleting its last row would retire the guard silently). +**Perf regressions are caught by the scaling test, not the absolute-time ones** — `tests/v2/test_benchmark.py::test_parse_cost_grows_no_worse_than_linearly` times a repeated unit at n vs 4n over a table of shapes (one per pipeline inner loop) and bounds the ratio; the `_thousand_names` tests use constant-size, delimiter-free input and are structurally blind to a complexity regression. Two rules when touching it: calibrate `_MAX_RATIO` against the WEAKEST quadratic's signal (a mixed quadratic surfaces far below the textbook 16×, so the operating point `_BASE` matters more than the bound), and confirm a planted regression fails it across REPEATED runs — one failure is a coin-flip on a timing test. The shapes cover different dimensions (segment count only via `commas`, intra-piece accumulation only via `particles`/`conjunctions`, non-ASCII input only via `honorifics` — every other unit is pure ASCII, so `script_segment` returns at its bail and the CJK stages go unmeasured, M2's clause view only via `maiden_clause`, the one unit holding a marker (until #601 its order mattered, `MA nee ` reaching nothing; since #601 either order reaches the take) — and connective RUN LENGTH only via `link_run`, whose unit has to be MIXED CASE and hold a connective of the generational class: `i und ` reads the letter as an initial and reaches nothing where `i Und ` reaches everything, and S2's anchor pass only via `credential_run`, whose unit has to hold a Title-case member behind an unambiguous credential: `PhD MA ` reads the same fields and never asks the pass, the capitals deciding each member first); measure before pruning one. **A shape whose input needs a PREFIX cannot be a `_SHAPES` row at all**, since that table repeats a unit and nothing else — a maiden clause needs a name word and a marker before the run it is about, and `"nee i Und " * n` reaches the clause rule not at all, measuring the identical ratio on a broken tree and a fixed one. That one is `test_a_clause_link_run_does_not_cost_quadratically`, which builds its own input and counts FRAMES. A prefixed shape whose cost is Python-level joins `_REREAD_SHAPES` beside it — (input builder, the one function the defect re-runs, reachability probe) rows counted in that function's frames at 8 units against 32 — which holds #558's trailing fixed point, its maiden-clause twin, and #559's particle chain, a shape that grows at both ends and so has no unit to repeat. A prefixed shape whose cost is C-level work — a list scanned by `in`, a tuple re-hashed as a dict key — emits no frame for that count to see, so it goes in `_PREFIXED_SHAPES` instead: (prefix, unit, reachability probe) rows timed on the clock like `_SHAPES`, at a base of their own recorded beside the table (#553, whose run shape read 5.97× once at base 800 on the pre-fix tree, under the bound; py3.11, 2026-09-28). **That table skips under a line tracer** (`sys.settrace`, coverage's core below py3.14), which slows every Python line and not the C-level scan, diluting the quadratic below the bound (5.46–5.52× broken at base 1600 under `coverage run` on py3.11, same date). The `sys.monitoring` core coverage uses from 3.14 does not dilute it (8.87–10.91× broken), and CI collects coverage in its 3.14 build job alone, so every CI job runs the table; `ja-extra` also sets `NAMEPARSER_REQUIRE_CLOCK_GUARDS` so that a line tracer there fails the rows instead of skipping them. **Coverage is collected on ONE interpreter, the one that uploads it** (`COVERAGE_ON` in `.github/workflows/python-package.yml`): the package has no version-conditional code, so one report stands for all of them, and under a line tracer the property grids in `tests/v2/test_properties.py` had cost the 3.12 job most of its ~12 minutes (measured 2026-09-29: the three #397 grids take 139s under coverage's settrace core on 3.12, 24s under its sys.monitoring core, 17s with none). **A shape the CLOCK cannot reach needs a FRAME-count guard instead**, which is the second scaling test in that file (`test_a_trailing_credential_run_does_not_cost_exponentially`, #531): where the defect is an exponential rather than a quadratic, the input length that separates the curves on a timing test does not finish, so the guard counts frames over 8 units against 16 and bounds THAT ratio. One pair does not see every curve, and the fix round for #531 measured why: at 2× the input the per-member LINEAR work swamps a quadratic (2.08× for a genuine one against 1.73× clean), so that pair guards the exponential alone and a second, longer pair — 16 against 64, where the same quadratic reads 7.42× against 3.53× clean — is what can see one. Assert them in that order: an exponential never returns from the longer run, so the cheap pair has to have failed first. Frame counts do not move under load, so this shape needs no repeated-run calibration — but it does need the same reachability assertion `_POLICY_SHAPES` rows carry, since the walk under measurement runs only while every unit still reads as a credential. A stage gated on an opt-in `Policy` field needs a `_POLICY_SHAPES` entry instead, since bare `parse()` never enters it — and that table's rows carry a **reachability probe** run before the measurement, because a precedence change can quietly stop the shape reaching the stage and leave a green test measuring a no-op (`_POLICY_SHAPES` is also asserted non-empty: an empty `parametrize` is a skip, not a failure, so deleting its last row would retire the guard silently). **Expected-failure tests use `@pytest.mark.xfail`** — the conftest parametrized fixture breaks `@unittest.expectedFailure`; always use `@pytest.mark.xfail` instead. diff --git a/docs/design/AGENTS.md b/docs/design/AGENTS.md index ff0d0641..fb039a15 100644 --- a/docs/design/AGENTS.md +++ b/docs/design/AGENTS.md @@ -6,6 +6,8 @@ rules.md states intended parsing behavior and is NORMATIVE; decisions.md is the **Starting a design: pin the negative space first.** A spec's "X is not supported" is a claim about today's parser; hold it with a case row for X written BEFORE the implementation, so a row that passes unexpectedly kills the premise before any code exists. The 2026-07-30 comma-suffix spec did this for one item (`suffix_phrase_is_comma_scoped`) and not for its central premise, which is the half that cost a full plan (#291). +**Rethinking a rule that grew by patching.** Three signs together mean the next fix should be a rethink, not another patch: the rule works by scanning and stopping; each later fix added a stop, an exception or a check on one; and a guard exists whose job is to model what another stage would do. The third sign is the decisive one — a rule that must predict other rules' behavior to stay correct is defending an invariant it should be built from. #601 is the worked case. Five issues (#424, #533, #535, #538, #549) had each added a stop or a release check to M2, and its decisions.md section reached about 13,000 words. The rethink (mechanisms.md#READ-WITHOUT-THEN-BIND, #LICENSED-BY-NEIGHBOR) deleted the release machinery, cut the rule's statement to 883 words (counted 2026-10-04), and turned the invariant from a checked property into the construction. The rethink was drafted in rule shape and prototyped in seven variants behind one switch before anything landed (#601's comments); the variants that failed did so on measured counts, not on review opinion. Rule length alone is not the signal: S2 and R4 have each gathered 35 dated decisions.md entries, and R4's are mostly vocabulary and rendering cases rather than added stops (counted 2026-10-04). + **Landing a design.** The gitignored spec (docs/superpowers/specs/) is the working medium; it dies with the branch, and the docs are the record. Before a design PR merges, walk its spec (including amendments) and distill the durable residue: decisions made or reversed → decisions.md entries; proposals rejected with evidence → Declined:; vocabulary that must stay out → Excluded:; behavior the design settled → rules.md (with a deviates: marker if unshipped); reusable patterns → mechanisms.md; options weighed → a weighing entry. Then check the spec cites nothing the docs don't now carry — a spec section with no committed home when the PR merges is lost, not deferred (a 2026-08-16 sweep of eight weeks of specs recovered nine such items). The same-PR amendment rule above covers code-driven changes; this covers the design-driven ones. **A count in a dated entry is evidence, not a live fact.** decisions.md entries are snapshots by convention, so measurements belong in them — but a reader wanting TODAY's number must not have to trust the snapshot's date. Where an entry quotes something that drifts (vocabulary sizes, corpus counts, set compositions), give the one-liner that recomputes it, and phrase the argument so it survives the digits moving — "the two shares differ by orders of magnitude" outlives "58% vs 0.65%". A count that carries no argument is better deleted than dated: "over every name in the corpus, no prefilter" says what "over all 782 names" says, and cannot go stale. Do NOT reach for a test asserting the count — that is the constant-content pattern, and it fails on every legitimate vocabulary addition. #326 is the cautionary case: it quoted a vocabulary composition, carried a date, and was stale in five days. diff --git a/docs/design/decisions.md b/docs/design/decisions.md index d4f51e86..c12f5b14 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -243,6 +243,7 @@ the fullwidth-colon marker (旧姓:佐藤 arrives as one word; the head-peel q - 2026-09-26 #535 — THE MAIDEN WALK READS THE TRAILING TITLE CHAIN, AND EVERY STOP ASKS ONE RELEASE QUESTION OF THE SPAN IT GIVES UP (Derek, 2026-09-26). This resolves the NOT TRANSPARENT HERE clause of the 2026-09-19 #533 entry above, which recorded `Jane Doe nee Smith MA Prof.` and `Jane Doe nee Smith Prof. MA` disagreeing and deferred the pair as one decision about what "trailing" means inside a clause. BOTH HALVES WERE TAKEN TOGETHER: a trailing title ends the clause (`Jane Doe nee Smith Prof.` reads title 'Prof.', maiden 'Smith', where 2.3.0 read maiden 'Smith Prof.'), and a credential or numeral standing in front of that title gets the stop it gets with the title absent (`… Smith MA Prof.` gives suffix 'MA', `… Smith V Prof.` suffix 'V'). Only the first half would have left the pair disagreeing in the other direction; H5's transparency is a statement about the peel and the chain read together to their fixed point, so the walk now reads the end of the clause the way assign reads the end of the name — `tail_reading`, the shared predicate mechanisms.md#ONE-PREDICATE-PER-QUESTION names — and over the forms the transparency property test below exercises, the two spellings give one answer. They split where something in or ahead of the clause stands to take the title — for example a particle chain (ahead of the clause or inside it), a bound-given join, a name left with no name word, or, after a family comma, a title word in front of a class member with a title behind it, which the given part's chain stops short of exactly as it does with no marker (`Doe, Jane nee Smith Rev. MA Prof.` reads maiden 'Smith Rev.', title 'Prof.', suffix 'MA', where `… Rev. Prof. MA` reads title 'Rev. Prof.', maiden 'Smith', suffix 'MA'; bare `Doe, Jane Rev. MA Prof.` reads middle 'Rev.') — and the Accepted pairs recorded here and in rules.md#M2 are examples of that, not a complete list. `tests/v2/test_properties.py::test_a_trailing_title_is_transparent_to_the_maiden_clause` holds that as an invariant over the two inputs, over heads and runs where the credential is given up and nothing ahead would take the title, with its negative control at e0f1a2fa in its docstring. ACCEPTED, THE TRANSPARENCY BOUNDARY: where the clause KEEPS the credential the spellings differ, because a clause is one contiguous run and a title inside the kept text cannot leave without the words behind it — `Doe nee Smith ba Prof.` reads title 'Prof.', maiden 'Smith ba', and `Doe nee Smith Prof. ba` maiden 'Smith Prof. ba'; `abdul nee Smith MA Prof.` reads title 'Prof.', family 'abdul', maiden 'Smith MA', and `abdul nee Smith Prof. MA` given 'abdul', maiden 'Smith Prof. MA' (2.3.0 kept every word in all four). The report differs with them: `Doe, Dr. nee Smith MA` and `… Prof. MA` report `suffix-or-name` on the kept MA, `… MA Prof.` reports nothing, the kept member no longer ENDING the clause, which is what M2's report is asked of. rules.md#M2 carries the first pair as an Accepted example. ACCEPTED, FURTHER SPLITS, measured 2026-09-26. Where something ahead of the clause would take the title, the spellings differ even with the credential given up in both, because the first-suffix-word stop ends the clause before the title check runs and the title check alone declines: `Jane van der Berg nee Smith PhD Prof.` reads title 'Prof.', maiden 'Smith', suffix 'PhD', and `… Prof. PhD` maiden 'Smith Prof.', suffix 'PhD' — H5's accepted particle-chain boundary inherited, the bare `Jane van der Berg Smith PhD Prof.` / `… Prof. PhD` splitting the same way — and `Berg, abdul nee Smith PhD Prof.` (bound-given join) and `Doe, Dr. nee Smith PhD Prof.` (no name word left) split likewise; the parent d9d80492 and 2.3.0 read all of these as the tree does, so only the claim is new. And a released particle with a title behind it is withdrawn: `Jane Doe nee Smith DO Prof.` reads title 'Prof.', maiden 'Smith DO', where `… Prof. DO` reads title 'Prof.', maiden 'Smith', suffix 'DO' (2.3.0 maiden 'Smith DO Prof.' and 'Smith Prof. DO'); `do` likewise. A particle INSIDE the clause does it from the other side, its chain able to take the credential behind it: `Jane Doe nee Smith do MA Prof.` reads title 'Prof.', maiden 'Smith do MA', and `… do Prof. MA` title 'Prof.', maiden 'Smith do', suffix 'MA' (2.3.0 kept every word in both). rules.md#M2 carries a pair of each shape as Accepted examples. ACCEPTED, H5'S REACH: `Jane Doe nee Smith King.` reads title 'King.', maiden 'Smith' (2.3.0 maiden 'Smith King.'), H5's accepted `Mary Jane King.` cost now reaching the last word of a birth name as it reaches the last word of a current one; rules.md#M2 carries it as an Accepted example. READER SCOPE IS UNCHANGED: the chain is consulted only where a trailing rule reads the clause's words (no comma, the part before a suffix comma, the given part after a family comma); before a family comma and in a tail segment the walk reads the peel alone and the clause keeps the title (`Doe nee Smith Prof., Jane` keeps maiden 'Smith Prof.'). THE FIRST-WORD FLOOR MOVED INTO THE CHAIN. `trailing_titles` and `tail_reading` take a `floor` (1 everywhere else), and the walk passes the position just past the marker's first word, so the chain never takes that word: `Jane Doe nee King.` keeps its maiden name (TITLES holds borne surnames and the marker announced one), and `Jane Doe nee Prof. Dr.` keeps 'Prof.' and gives up 'Dr.'. The first draft applied the floor as a clamp on the title stop after the chain had run, and that is measurably wrong rather than merely inelegant: a chain allowed to take the first word has already spliced it out of the count the re-peel reads, so `Jane Doe nee King. ba` kept 'ba' in the clause where `Jane Doe nee Smith ba` gives it up. `test_a_title_first_word_counts_as_a_word` is the invariant (a title-vocabulary first word against an ordinary one, the credential behind it read alike) and its docstring records the clamp's failure count. ONE RELEASE CHECK FOR THREE STOPS. The numeral, the credential and the title stop each ask `_release_reads_off` whether the name the take would leave reads what they give up as titles or suffixes with no join below the take absorbing it — rules.md#M2's invariant, "a word the clause gives up reads as a post-nominal or the clause keeps it", which the title stop now answers to as well. It is asked of the whole SPAN a stop gives up, not of the stop's own word: a first draft that checked the word alone broke the invariant on 513 of a 5,198-parse sweep (measured 2026-09-26 on that draft, which is not in the tree to recompute), because a stop gives up every word behind it — `DOE NEE SMITH PROF. MA` stopped at the title and handed the MA to the family, where the name left standing reads MA as a name; it keeps maiden 'SMITH PROF. MA' now. The question is asked per reader, the way that reader reads: the TRAILING reader is assign's own `tail_reading` over the view; the GIVEN_SLOT reader has its count settled by the comma, so it is the writing alone with a name word ahead. The name-word-ahead half is what keeps `Dr. nee Jones Smith Prof.` whole — the take would leave `Dr. Prof.`, whose Prof. would be the family name. THE JOINS. A released span holding a title with a particle ahead of it declines, because P2's chain runs on over a trailing title (rules.md#H5's Accepted `John van der Berg Prof.`): `Jane van der Berg nee Smith Prof.` keeps maiden 'Smith Prof.', and a released particle with a title behind it is withdrawn for the same reason (`Jane Doe nee Smith MA do Prof.` keeps maiden 'Smith MA do'). The bound-given (P5) half of `_join_takes_the_member` is asked for the GIVEN_SLOT reader alone: P5 is `BoundJoin.LENIENT` only after a family comma, and before one its STRICT reserve already refuses a join that would change a suffix reading, so asking there over-declined — `abdul nee Smith V` kept 'V' in the clause, where the scoped check lets it go to suffix as 2.3.0 did. The consequence is recorded by a case row rather than prevented: `abdul nee Smith Dr.` reads title 'Dr.', family 'abdul', maiden 'Smith', which is how `abdul Dr.` reads bare (2.3.0 read given 'abdul', maiden 'Smith Dr.'). THREE NUMERAL-STOP READINGS MOVE WITH IT, decided in rather than deferred, since each is the same invariant broken at the stop this change was already rewriting. The numeral stop never asked the join question: `Berg, abdul nee Smith V` read given 'abdul V' at 2.2.0 and 2.3.0 (given 'abdul nee', suffix 'V' at 2.0.0 and 2.1.0, #411's reserve differing between those pairs before the walk runs) and reads maiden 'Smith V' now. After a family comma the given slot reads a lone numeral as a suffix only where the given part is the LAST comma part — assign's own two-segment condition from #144 — and the walk now asks it too (`tail_follows`): `Doe, Jane nee Smith V, PhD` read middle 'V' at 2.2.0, 2.3.0 and e0f1a2fa, an M2 violation predating this change that its first draft had extended to `… V Prof., PhD`, and it reads maiden 'Smith V' now, as 2.0.0 read it. The withdrawal reaches through a title behind the numeral, so `Doe, Jane nee Smith Prof. V, PhD` keeps maiden 'Smith Prof. V' where `Doe, Jane nee Smith V Prof., PhD` gives up the title — the same asymmetry the bare given slot has with no marker in it (`Doe, Jane Prof. V, PhD` against `Doe, Jane V Prof., PhD`). And the numeral stop reads FROM the marker, unlike the other two, so a numeral straight after the marker is not held to the first-word floor, and it may decline the clause; examples, not a rule over every shape: `Dr. nee V` and `Doe, J. nee V` keep maiden 'V' as 2.3.0 did, though `Doe, J. V` reads suffix 'V', and `Doe nee V, Jane` and `Smith, John, PhD nee V` read no clause at all (family 'Doe nee V', suffix 'PhD nee V', as at 2.3.0) — `Jane Doe nee V Prof.` read maiden 'V Prof.' at 2.3.0 and reads family 'nee', suffix 'V', title 'Prof.' now, as `Jane Doe nee V` reads plus the title. After a family comma with another comma part behind the given one, the same #144 condition that keeps `Doe, Jane nee Smith V, PhD` whole now keeps a lone numeral in the clause as well: `Doe, Jane nee V, PhD` read given 'Jane', middle 'nee V', suffix 'PhD' at 2.3.0 and at the parent d9d80492 and reads maiden 'V', suffix 'PhD' now, as 2.0.0 and 2.1.0 read it, `Doe, Jane nee V, Jr.` likewise — the marker had been read as a name word there at 2.2.0 and 2.3.0, and a case row records it. A LINK THE WALK STOPS AT ASKS THE RELEASE QUESTION TOO. Reading the end of the clause through the title chain makes the link exception refuse a link it used to join — in `Jane Doe nee Smith i DO Prof.` the DO is the peel's once the title is chained — and a stop at a link gives up the words behind it. Unchecked, that put 'Doe i' in the middle name and 'DO Prof.' in the family (e0f1a2fa read maiden 'Smith i DO Prof.'), and `Berg, abdul nee Smith i V Prof., MD` read middle 'V'. So a link the exception refuses asks `_release_reads_off` of the run it would give up — only where the exception, bounded by the peel over the words as WRITTEN, would have joined it (the refusal is the title chain's), where a trailing rule reads the clause, and where the link is not the first word after the marker (a first-word stop declines the clause and gives nothing up); where that fails the clause keeps the link and walks on. Scoped that narrowly because a first version asked it of every refused link and kept links the name left standing reads off: `Doe, Jane nee Smith i V` kept maiden 'Smith i' where every release gives up suffix 'i V' (23,472 of a 411,936-parse link grid moved clean readings, measured 2026-09-26). With it, the given-part model reads the lenient numeral in both of assign's passes — the literal last piece first, then the last one standing once the chain has taken the titles behind it — and only where no comma part follows (`Doe, Jane nee Smith i V Prof.` gives up 'i V' and the title, as bare `Doe, Jane i V Prof.` reads them; `… i V Prof., PhD` keeps maiden 'Smith i V'). The same model moves the credential stop where a lenient numeral follows the credential: `Doe, Jane nee Smith MA V` now gives up 'MA V' (suffix 'MA V', maiden 'Smith') as bare `Doe, Jane MA V` reads it, where d9d80492 and 2.3.0 read maiden 'Smith MA', suffix 'V'; `… MA V, PhD` keeps maiden 'Smith MA V'. Over that link grid (links i, y, e, and, i y, y i, and i × 8 heads × bodies × `, MD` tail × 3 cases × 3 name orders), against the tree before the link check: 0 new violations, 6,195 fixed, and the clean readings that move now match the bare given part's (`DOE, JANE NEE SMITH MA I` gives up 'MA I', as `DOE, JANE MA I` reads suffix 'MA I'). `Jane Doe nee Smith i DO Prof.` now reads title 'Prof.', maiden 'Smith i DO' and reports the kept DO, and `Berg, abdul nee Smith i V Prof., MD` maiden 'Smith i V Prof.', suffix 'MD'. A RELEASED TITLE THAT IS ALSO A PARTICLE IS KEPT AFTER A FAMILY COMMA, because P6 attaches a particle trailing the given part to the family: released by the title stop, 'St.' in `Doe, Jane nee Smith St.` reads family 'St. Doe', so `_join_takes_the_member` declines it and the name reads maiden 'Smith St.', as e0f1a2fa and 2.3.0 read it, and so does `Doe, Jane nee Smith MA St.`. Titles only — a credential that is also a particle ('DO') is the given slot's own lean, which the credential stop has already asked (#533). A FOURTH MEASUREMENT, against e0f1a2fa rather than d9d80492: the M2 invariant over a fuzz of twelve heads ({`Jane Doe`, `Doe, Jane`, `Doe, Prof.`, `Jane Doe, PhD`, `Doe, J.`, `Berg, Jane van der`, `Jane van der Berg`, `Berg, abdul`, `abdul Berg`, `J. Doe`, `Prof. Jane Doe`, `Dr.`}), ` nee Smith ` and every one-to-three-word sequence over {MA, Ma, V, Prof., St., King., PhD, ba, do, DO, Jr., i, van, y} holding at least one of Prof., St., King., then nothing or `, MD`, as written, upper- and lower-cased (107,352 parses, measured 2026-09-26 on the narrowed tree; the intermediate tree above read 2,228): 588 texts that violated it at e0f1a2fa no longer do, and 12 newly do, all `Dr. nee Smith i MA|V` followed by `Prof.`, `King.` or `St.`, with or without `, MD`. Each is the title-carrying twin of `Dr. nee Smith i MA` / `Dr. nee Smith i V`, which read family 'i' at e0f1a2fa already: the bare-title head leaves no name word, the unguarded stop's class (#548), which the title now reaches because it leaves the clause. ACCEPTED: `Jane Doe (nee Smith Prof.)` reads title 'Prof.', maiden 'Smith' — bracket content ending in a period is suffix-shaped (S1), so the brackets are dropped and the clause is read bare, outside the reach of the delimiter precedence; rules.md#M2 carries it as an Accepted example. MEASURED 2026-09-26, the tree against its parent d9d80492, over three generated grids, and these are dated snapshots. Recipe: grid A is each head in {`Doe, Jane`, `Doe, J.`, `Doe, Prof.`, `Berg, abdul`, `Doe, Jane van`, `Jane Doe`, `J. Doe`, `Dr.`, `abdul`, `Jane van der Berg`, `John`}, then ` nee `, then every sequence of one to three words (repetition allowed) from {Smith, Jones, Prof., Dr., Sir, MA, DO, PhD, V, III, i, do, van, Ma, M.A., King., Rev., ba}, then either nothing or `, PhD`, each text as written, upper-cased and lower-cased, duplicates removed — 327,936 texts; grid B, aimed at the bound-given join, is the same construction over heads {`Dr. abdul`, `abdul rahman`, `Dr. abdul rahman`, `Berg, Dr. abdul`, `Berg, abdul rahman`, `abu`, `Berg, abu`, `Mr. abu`, `Dr.`, `Berg, Dr.`, `abdul`, `Berg, abdul`}, words {Smith, Prof., MA, V, do, van, PhD, Ma, ba, III, Jr., M.A.} and tails {nothing, `, PhD`, `, Jr., MD`}, plus these fourteen texts, as written only: `Jane Doe nee Ph. D. Prof.`, `Jane Doe nee Ph. D. Smith Prof.`, `Jane Doe z domu King. ba`, `Jane Doe z domu Smith Prof.`, `Jane Doe nee King. Prof. ba`, `Jane Doe nee King. Prof.`, `Doe, Jane nee Smith V,`, `Doe, Jane nee Smith V, ` (with the trailing space), `Doe, Jane, nee Smith V`, `Doe, Jane nee Smith V Prof., PhD`, `Doe, Jane nee Smith Prof. V, PhD`, `Doe, Jane nee Smith MA, PhD`, `Doe, Jane nee Smith V (Jr.)`, `Doe, Jane nee Smith V "Bo"` — 173,057 texts. The invariant tested is M2's: in a parse with a non-empty maiden field, every token written after the ` nee ` marker is roled maiden, title or suffix (so the two `z domu` texts are parsed and not checked). Grid C, a four-word probe of the title chain after a family comma, is each head in {`Doe, Jane`, `Jane Doe`, `Doe, J.`, `John Smith`, `Doe, Jane Mary`}, then ` nee `, then every sequence of four words (repetition allowed) from {Smith, Jones, Prof., Dr., King., MA, PhD, V, Jr., ba, do}, then either nothing or `, PhD`, as written only — 146,410 texts. Per grid, the counts are: texts that newly violate the invariant at the tree, texts that violated it at the parent and no longer do, texts that violate it at both. Grid A: 0, 2,012 and 17,854; grid C: 0, 1,368 and 15,534; grid B: 0, 1,222 and 11,149 (all re-measured 2026-09-26 on the tree with the link and particle-title checks below, as narrowed; an intermediate tree whose link check fired on every refused link read grid A 0, 3,772 and 16,094, its extra 1,760 fixes being links the as-written reading refuses too, which are #548's class), one of those last being the check itself counting the quoted nickname in `Doe, Jane nee Smith V "Bo"`. In grids A and B, every other remaining violation has the clause ending immediately before a word of the unambiguous suffix vocabulary written the way the plain suffix-word stop takes it (PhD, III, M.A., Jr., and a lower-case i or v; never a bare capital, which reads as an initial), which is where that stop ends a clause without asking a release question. DEFERRED: the walk's plain suffix-word stop — the one ending the clause at the first suffix word, as distinct from the trailing numeral, credential and title stops — asks no release question at all, and where the words behind it cannot read as post-nominals they land in a name part: on degenerate heads the stop word itself does (`Dr. nee Smith PhD Prof.` reads family 'PhD', `Doe nee Smith Jr. Prof., Jane` family 'Doe Prof.'), and with an ordinary head the words behind it do (`Doe, Jane nee Smith PhD Smith` reads middle 'Smith'). Unchanged here, and what "the clause keeps it" should mean where the kept words would follow a credential the clause itself ended at is its own question; rules.md#M2 carries `Doe nee Smith Jr. Prof., Jane` as an Accepted example meanwhile. The trailing numeral's stop has the same gap before a family comma, where it is made over the peel alone with no release question: `Doe nee Smith V, Jane` reads family 'Doe V', maiden 'Smith' (so did 2.2.0, 2.3.0 and the parent d9d80492; 2.0.0 and 2.1.0 read maiden 'Smith V'), and rules.md#M2 carries it beside the other. Open: #548. COST, measured 2026-09-26 on CPython 3.11.16 as profiler call events in one `Parser.parse` after a warm-up parse, parent → tree: `John Smith` 171 unchanged, as are `Jane Doe Prof.` and `John Smith MA`; `Jane Doe nee Smith` 244 → 246; `Jane Doe nee Smith MA` 377 → 388; `Jane Doe nee Smith Prof.` 276 → 372, the one shape that now runs the chain and a release check it never ran. Recompute: count `sys.setprofile` call events around the second of two `Parser().parse` calls, with the tree and d9d80492 each first on `sys.path`. - 2026-09-27 (Derek), #544 — S2'S COMPANY CLAUSE DOES NOT REACH ACROSS A CLAUSE, recorded as an Accepted boundary rather than repaired. `Jane Doe Jr. nee Smith Ma` keeps maiden 'Smith Ma' and reports it, the clause's own words standing between 'Jr.' and 'Ma', where the clause-less `Jane Doe Jr. Ma` reads suffix 'Jr. Ma'. tests/v2/test_properties.py's clause-agreement walk pins the pairs that differ for exactly this reason as its `anchored_head` class — 810 of its 20,412 pairs, recorded 2026-09-27, 0 with the anchor off. Inside the clause the company reads as it does anywhere: `Jane Doe nee Smith PhD MEng` ends the clause at 'PhD' and reads suffix 'PhD MEng', maiden 'Smith'. +- 2026-10-04 (Derek), #601 — A MARKER COUNTS ONLY BEHIND A SURNAME, AND IT TAKES THE WORDS AFTER IT UP TO THE TRAILING RUN THE NAME WOULD END WITH IF THE CLAUSE WERE NOT WRITTEN. This SUPERSEDES the release machinery of the #533 and #535 entries above (the release question, the per-stop checks, the anchored-head boundary of the #544 entry) and the given-part and suffix-comma reach the 2026-07-03 rule had; the #399, #420, #434 and delimiter entries stand. The trigger was #548, `Dr. nee Smith PhD Prof.`, where 2.3.0 read family 'PhD', maiden 'Smith': a clause whose head was nothing but a title, and a release check asking whether each word it gave up could still read as a post-nominal. #548 was closed into this issue and #602 (S2's run) because both halves of its answer were rules rather than repairs. THE HEAD RULE (Derek): a marker is one only in the name before any comma or the surname part before a family comma, behind at least one name word past the leading title run, and not straight behind an unambiguous suffix word or a connective. Anywhere else it is an ordinary word — in the given part after a family comma (`Doe, Jane nee Smith` reads middle 'nee Smith', 1.4.0's reading) and in a part after a suffix comma (`Smith, John, PhD née Jones` reads suffix 'PhD née Jones', also 1.4.0's). Derek's framing was that a maiden marker always follows a family name, and that a marker following something the parser has decided is a title, suffix, particle or connective is not a marker. THE PARTICLE HALF WAS DROPPED BY MEASUREMENT: prototype variant 1 refused a marker behind particle vocabulary and so declined `Anh Do née Tran` and `Mai Le née Nguyen`, Do, Le and Van being particles that are also the commonest Vietnamese surnames; a particle counts as the surname (`Jane van nee Smith` reads family 'van', maiden 'Smith'). THE BOUND (Derek): the clause takes its words up to the run of post-nominals and titles that the trailing rules read at the end of the name with the clause removed, and the words it gives up take the roles that reading gives them — the TRAILING rule's reading of the clause-free name, not the full parse of it, so a head the full parse would read differently (H1's lone title, a P5 join) cannot change where the clause ends. The first word after the marker is always taken, the marker having announced a name — a lone roman numeral included, so `Jane Smith née V` reads maiden 'V' where the clause-free `Jane Smith V` reads suffix 'V', accepted rather than given a numeral exception (rules.md#M2's Accepted block) — unless it is an unambiguous suffix word, where the marker declines (`Jane Smith nee PhD` reads family 'nee', suffix 'PhD', as before). Before a family comma the clause keeps the 2026-07-03 stop at the first suffix word, since no trailing rule reads that part's end. VARIANTS, prototyped 2026-10-03 behind an environment switch on one tree (the numbers are in #601's comments): v2/v3 grew the clause until the existing release check accepted it and reached 0 invariant violations but inherited the check's caution about particles (`Anh Do geb. de la Vega PhD Prof.` kept 'PhD Prof.' in the maiden name); v4 read the run once over the name as written without binding it, and is the recorded negative control in tests/v2/test_properties.py (350 of the clause grid's 4,482 parses fail under it, measured 2026-10-04); v5/v6 bound that reading as roles but still counted the clause's words among the words to spare, so `John nee Smith ba` read 'ba' as a credential where `John ba` reads a name; v7, the clause-free reading bound as roles, was adopted. Against #535's grids v7 left 0 violations of the M2 invariant where the tree then had 15,586 (grid A, 327,936 texts) and 8,713 (grid B, 173,057), measured 2026-10-03 with an invariant check written to #535's recipe. WHAT THE PROTOTYPE MISSED, found implementing it: the numeral fork. A trailing roman numeral is a suffix only by its shape against the word before it, and v7 read the clause's last word without that test, so `John née Jones Smith VI` kept 'VI' in the maiden name; the take now asks `is_trailing_numeral_suffix` of the last piece, and cases.py's `a_numeral_the_clause_free_name_reads_ends_the_clause` pins it. And the head's title run is measured over the words before the marker, a lone title excepted: measured over the whole segment, `Lord Chancellor née Jones` counted 'Chancellor' as a title and refused the marker. TRIAGE, 2026-10-04: the prototype's 68 failing case rows (40 facade failures among them) were re-pinned as `fix(#601)` with their group named in the note, bar the head-rule bug above; tests/v2/test_parser.py's corpus-append invariant became two-sided (appending ` née X` moves no other field, OR the take records why it declined), with a named set for the suffixes that themselves hold a marker. DEGENERATE INPUT, NO READING DESIGNED FOR (Derek, 2026-10-04: garbage in, garbage out). Two names that read worse by eye moved, and they are recorded here so the move is not mistaken for a regression, not because either reading is wanted: `Berg, Jane van der nee Smith DO` reads middle 'van der nee Smith', family 'Berg' where 2.3.0 read family 'van der Berg', maiden 'Smith DO' — the marker in a given part being a word, the particles no longer trail the given name (P6); and `Jane Doe, PhD née Smith` reads given 'PhD', middle 'née Smith', family 'Jane Doe' where 2.3.0 read given 'Jane', family 'Doe', suffix 'PhD', maiden 'Smith', which 2.3.0 reached only through the clause taking the marker (the comma part opening with a credential is #603's shape). Measured 2026-10-04 against 2.3.0. COST: the take is one function and no release helper survives; the #533 review's comma-head grids could no longer reach a take at all and were replaced by counting heads (tests/v2/test_properties.py's `_maiden_clause_grid`, 4,482 parses). ### N1 — the default delimiter pairs @@ -722,6 +723,11 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `John Smith, Ph. D.` 254 → 254, `John Smith, CPA` 217 → 217, `John Smith, MD PhD` 260 → 260, `John Smith, MBA CPA` 261 → 261, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 266, and an all-caps record with a capitalized given name and a credential, `LLOYD WEBBER, ANDREW PhD`, 338 → 364, a constant +26 that does not grow with the name's length (`'LLOYD'*64 + ' WEBBER, ANDREW PhD'` 653 → 679; the second draft's per-character contrast test had made that +350). Past the all-caps check the contrast is linear in the words before the comma, about four frames a word (the generator, the case test's call, and the own-words walk's fold and marker test; `'GAULLE ' + 'ap '*k + 'GAULLE, CHARLES'` is +7, +26, +38, +54 over master at k = 0, 1, 4, 8), and constant in a word's letters (`'de ' + 'GAULLE'*k + ', CHARLES'` 253 → 275, 343 → 365 and 631 → 653 at k = 1, 16, 64), where the capital-then-lowercase draft grew a frame per letter; all paid only once a caps word is in hand. The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Every caller of the caps predicate asks first, in C, what it would decline anyway — alphabetic capitals, not a listed suffix word — so a listed credential never pays for the call (the review of the first fix found `John Smith, CPA` +4 and `John Smith, Ph. D.` +5 before these). The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. - 2026-10-02 (Derek), #573 — a never-given particle opening the given part after a family comma leaves no word to spare; the member behind it is the surname it heads, unless mixed-case capitals say credential. The record is the P6 section's entry of the same date. - 2026-10-02 (Derek), #578 — A PARTICLE BETWEEN A CREDENTIAL AND A TITLE-CASE WORD STAYS SILENT, AS IT READS. `Doe, Jane PhD vd Ma` reads given 'Jane', middle 'vd Ma', family 'Doe', suffix 'PhD', reporting nothing. No new rule decides that: S2's given-part slot already says a particle in front of a member takes the member out of the slot, and where the word behind reads as a name the two are one name, asked nothing and reporting nothing — `vd` reads as `van` does (`Doe, Jane PhD van Ma`, the same reading), and S2's own boundary example is `Doe, John van Ma`. The writing is malformed whichever reading is meant: a Title-case `Ma` behind a lowercase `vd` behind a `PhD` is neither a Dutch listing nor a credential run as anyone writes one. A reader leaning anywhere leans suffix (Derek: both words follow a known credential, and the writer likely did not know how suffixes are capitalized), but that is a guess about a writer's mistake, not a convention the writing carries. Declined: (a) reading `vd Ma` as suffixes, which needs code to make a credential speak across a particle — against P2 (a particle joins the word behind it) and #554 (`vd` speaks for nothing behind it); (b) reporting the fork, which needs an emitter on the family-comma path for a decision the particle chain makes before assign knows which read the part gets (AGENTS.md: prefer leaving a fork silent and documenting the omission over machinery for input nobody reads with confidence). Where the writing does speak the parser already follows it: `Doe, Jane PhD vd MA` reads suffix 'PhD vd MA' and reports (#573). Pinned by the case row `a_particle_after_a_credential_heads_a_title_case_name`. +- 2026-10-04 (Derek), #602 — A CREDENTIAL AFTER THE NAME CORE STARTS A RUN TO THE END OF ITS PART. An unambiguous Latin suffix word of two or more letters standing after two name words with no comma (a lone particle not counting toward the two), or after the given word in the part after a family comma, makes every later word of that part a suffix, except a title word, which reads as a title: `John Smith PhD Jones` → suffix 'PhD Jones', `Eric H. Holder Jr. Attorney General` → suffix 'Jr.', title 'Attorney General'. A name word the run absorbs is reported `suffix-or-name`, a reader hesitating over `Jones` there. Before this the three comma shapes disagreed — middle 'Smith PhD' with no comma, middle 'Jones' and suffix 'PhD' after a family comma, and suffix 'Jr., Jones' after a suffix comma (C2's tail) — and the no-comma reading made a credential a middle name, the one reading nobody intends. The suffix-comma path already read it this way, and the comma spelling of the Holder name always did. + WHAT STARTS THE RUN is one predicate, `_pieces.starts_a_credential_run`, read by the trailing peel (`credential_run`, applied to every finished peel, so assign, P5's reserve, P2's chain stop and the maiden walk see one answer) and by the given-part walk. EXCLUDED, each with its reason: a TITLE WORD — 746 title words are in no suffix set (measured 2026-10-04: `len(L.titles - L.suffix_acronyms - L.suffix_words)` for the default lexicon), many of them surnames (`King`, `Bishop`, `Judge`), so `Mary Jane King Smith` keeps family 'Smith'; a title is read from the end of a name only behind a comma or with a period (H5, Derek). The AMBIGUOUS CLASS, which carries no evidence of which it is. A SINGLE LETTER, in any case: a lowercase `i` or `v` carries no `initial` tag, and found by the v1 off-switch test, `Josep Carod i Rovira` with `i` removed from the conjunctions read suffix 'i Rovira'. A CONNECTIVE. A word that is also PARTICLE or BOUND GIVEN-NAME vocabulary — found by the suite, not foreseen: `abd` is in the credential list (ABD) and heads `Abd Allah`, and `Mohamed Ali Abd Allah` read suffix 'Abd Allah' until it was excluded. The NON-LATIN HONORIFIC words, an honorific followed by a name word being a different question from a credential followed by one. A word that is also a title (`MD`, `Ms`, `Sr.`) does start the run: position decides, and after the name core it is the suffix. + DECLINED: lifting the credential out — `John Smith PhD Jones` as given 'John', middle 'Smith', family 'Jones', suffix 'PhD', the family-comma path's old reading. It keeps every name word in a name field, but it reads a credential as something written in the middle of a name, which is the reading the rule rejects. + REVERSED BY THIS ENTRY rather than edited: the #578 bullet above (`Doe, Jane PhD vd Ma` now reads suffix 'PhD vd Ma', the credential starting the given part's run), and two of the three shapes #544's ACCEPTED LIMITS kept out of the company's reach — a title standing between the credential and the member, and a member a particle chain had taken. The case rows `a_title_between_degree_and_member_keeps_it_a_name` and `a_member_the_particle_chain_took_is_out_of_reach` keep their ids and now pin the run. The third, the merged `Ph. D.`, stays out of reach: that piece is outside the walk the run is read over. + MEASURED 2026-10-04, every name in `tools/differential/corpus*.jsonl` parsed on this tree and on aa401d80 with `nameparser.__file__` asserted on each side: 7 of 1,514 move, four of them this entry's own new rules.md#S2 examples. The three already in the corpus: `Jane van der Berg née Jr Jones` (suffix 'Jr Jones', where the marker's declined clause had left middle 'van der Berg née Jr'), and `John Doe, MD - PhD - FACS` and `Smith, MD - PhD - FACS`, whose dash inside the given part, a word there under every policy (#549), is now in the run: suffix 'PhD - FACS' where they read middle '-', suffix 'PhD, FACS'. The comma-part question — whether a part OPENING with a credential makes the comma a suffix comma (`John Smith, PhD Jones` still reads family 'John Smith') — is #603's. ### indic-honorifics — the renunciate class and the Indic honorific vocabulary (2026-09-06, #346/#344/#343) diff --git a/docs/design/mechanisms.md b/docs/design/mechanisms.md index 3a061b0e..218be198 100644 --- a/docs/design/mechanisms.md +++ b/docs/design/mechanisms.md @@ -45,6 +45,14 @@ Problem shape. Where should a new "recognize X" behavior live? Contract statemen "Esq. Smith" reads title, not suffix. How it works. The two layers compose without ordering bugs because the positional layer never overrides a vocabulary claim (rule O4 is the positional layer's contract). The exception is still ONE, re-checked 2026-09-08 against the trailing title run (rule H5), which reads title vocabulary in the trailing slot and is NOT a second one: it runs after the suffix peel rather than before it, so a post-nominal keeps its claim — `John Smith Esq.` reads suffix and `John Smith Prof. Jr.` reads suffix `Jr.`, both measured. That rule is also the worked answer to the Reach-for-it line below: it wanted a word's identity and its position at once and was split by SLOT, the front asking about shape (H2) and the back about vocabulary (H5), two questions with two criteria rather than one rule fighting both layers. Lives in. _classify/_group (vocabulary side), _assign (positional side). Reach for it when. A proposed rule wants a word's identity AND its position at once — split it, or it will fight both layers. +## READ-WITHOUT-THEN-BIND — read the name without the construct, then fix that reading + +Problem shape. An optional construct sits inside a name — a maiden clause, a trailing title — and the requirement is that writing it changes nothing else about the name. Checked after the fact, that invariant has to be defended at every place the construct can end, and each defense has to model what later stages would do to the words it gives up. Contract statement. Read the name as if the construct were not written, using the rule that reads that position, not the whole parse. The construct takes only what that reading leaves. Where later stages could still move what the reading decided, give those words their roles at the take so no later join can reach them. The invariant then holds by construction and is not a property each stop must defend. How it works. M2 before #601 scanned forward, stopped at suffix-looking words, and kept a release check asking at each stop whether the name left standing would read what was given up as post-nominals. That check modelled the P2 chain, the P5 bound join and P6's attachment, and #548 was it failing at the one stop that never asked. #601 replaced the stops and the check with the clause-free reading plus the bind: a net 689 lines out of `_group.py` (+210/−899 in the #601 commit), and #535's grids went from 15,586 and 8,713 invariant violations to 0 (prototype measurement, 2026-10-03, #601's comments). The first half already appears twice. H5's trailing title is TRANSPARENT: "what stands once the chain is taken reads exactly as it would read written without the title, plus the title". P3 asks its word count and its one-case question of the name's OWN words, so a clause beside the name changes neither. M2 is the first to bind as well, because its construct sits in front of the run it reads and the joins come later. Two limits, both measured on #601. Read with the TRAILING rule, not the full parse: a head the full parse reads differently (H1's lone title, a P5 join) would otherwise move where the construct ends. And binding means the run can read differently from the same words with no construct (`JANE Q. DOE GEB. LE DO DO` gives suffix `DO DO`, while the full parse of `JANE Q. DOE DO DO` chains them into the family), which rules.md#M2 accepts in so many words. Lives in. nameparser/_pipeline/_group.py (`_maiden_take`: the view, `tail_reading` over it, and the role-tagged pieces `group()` applies); nameparser/_pipeline/_pieces.py (`tail_reading`, H5's chain); the P3 own-word test in `_group.py`. Reach for it when. A rule is growing stops, exceptions or a release check whose job is to keep one construct from disturbing the rest of the name, and above all when that check has started to model another stage. Ask what the name reads as without the construct, and make that the rule. + +## LICENSED-BY-NEIGHBOR — a structural word is structural only beside the right neighbor + +Problem shape. A word in a structural vocabulary — a marker announcing a maiden name, a connective joining surnames — turns up where what it would announce makes no sense: behind a title alone, behind a credential, at the end of a name. Designing a reading for each such shape is endless, and the shapes are mostly junk. Contract statement. Make the vocabulary claim conditional on the neighbor: the word is structural only where the word beside it has the role the structure needs. Anywhere else it is an ordinary word and goes to the positional layer (TWO-LAYER-ASSIGN), with no reading designed for the junk. How it works. M2 (#601): a marker counts only behind a name word of the part holding the family name, not behind the leading title run, an unambiguous suffix word or a connective. So `Dr. nee Smith` gives first `nee`, last `Smith`, and `Jane Doe PhD nee Smith` gives suffix `PhD nee Smith`, neither designed. P3 is the same shape for the connectives that are also generational vocabulary: `i` joins only where "a name word stands on each side of it". The condition must be checked against real names, not just against the vocabulary. #601's first cut also refused a marker behind a particle and broke `Mai Le née Nguyen` and `Anh Do née Tran` — Le, Do and Van are particle vocabulary and among the commonest Vietnamese surnames. Lives in. nameparser/_pipeline/_group.py (`_maiden_take`'s head rule; `_between_name_words` for P3). Reach for it when. A structural word is producing bad readings in positions its structure does not fit. Make its claim depend on the neighbor before adding a reading or a stop for each position. + ## STATE-OFFSET-CHANNELS — early facts ride the state Problem shape. A fact known during tokenization matters to a much later stage. Contract statement. A pre-token fact is recorded as offsets on the ParseState (comma_offsets, interpunct_offsets) and consulted later by position — or by presence alone where the fact is name-level, as both interpunct consumers do, or as a single boolean for the whole name, as `ParseState.one_case` is (#289/#516: the written-case fact three later sites read — the trailing suffix slot, the post-comma given slot, and the tail-segment reading — none of which may disagree about it) — rather than re-derived from text. How it works. The offsets survive every intermediate stage untouched; #298's transcription marker rides this channel from tokenize to order resolution (rules T3/W4). `one_case` differs from the offset channels in WHO writes it: it is computed once by whichever of segment and classify needs it first (classify always, segment only where a comma form could turn it on), recorded as `None` for "not asked yet", and never recomputed once set (decisions.md#S2). Lives in. nameparser/_pipeline/_state.py, produced in _tokenize (the offsets) or in segment/classify (`one_case`). Reach for it when. You are about to re-scan the original string in a late stage to rediscover something tokenize already knew. diff --git a/docs/design/rules.md b/docs/design/rules.md index 35892d88..a53f5b42 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -332,9 +332,9 @@ H5. Rationale: a word abbreviated with a period at the END of a name puts the word out of reach. A particle chain (P2) has already taken the trailing word into the family name, and no title word is standing in the trailing slot at all. A maiden clause is not - such a join: the clause's walk reads the end of the name through - this chain (M2), so a trailing title ends the clause and is a - title, where the name the take leaves reads it as a title (M2). + such a join: it ends where the run the end of the name reads + without it begins (M2), so where a trailing rule reads the part + a trailing title ends the clause and is a title. "John van der Berg Prof." → family="van der Berg Prof." "Mary Smith née Jones Prof." → title="Prof." Accepted: what the chain leaves is also what counts as a name @@ -761,7 +761,7 @@ P5. Rationale: some given-name words are incomplete alone — "abdul" "abd Berg née Jones" → family="Berg" "abd Allah Smith née Jones" → given="abd Allah" "abd née Jones" → given="abd" - "Berg, abd née Jones" → suffix="abd" + "Berg, abd née Jones" → given="abd" · boundary Accepted: after a family comma the join stands though, unjoined, assign would read the word it takes as the suffix — the family is fixed there and the pair is the given; and a bare ambiguous @@ -963,7 +963,7 @@ P6. Rationale: a particle ending the name has nothing to link negative-control sweep pinning the disagreeing set the precedence bullet above names. A change that breaks one side of that pair should expect that test, not this file, to say so first. - history: decisions.md#P6 · interacts: A1, C1, P1, S2, P5, M2 · implemented: nameparser/_pipeline/_post_rules.py + history: decisions.md#P6 · interacts: A1, C1, P1, S2, P5, M2 · implemented: nameparser/_pipeline/_assign.py, nameparser/_pipeline/_post_rules.py P7. Rationale: a one-letter particle is spelled with the same letter as an initial, and the period is what tells them apart. An @@ -1133,6 +1133,23 @@ S2. Rationale: generational suffixes and credentials are recognized after a family comma that the company leaves holding no name word, which reads wholly as the credential run. A member its own capitals already made the credential reports nothing new. + A credential after the name core starts a run to the end of its + part (#602). An unambiguous Latin suffix word of two or more + letters, standing after two name words where no comma divides + the name or before a suffix comma, or after the given word in + the part after a family comma, makes every later word of that + part a suffix, except a title word, which reads as a title, and + in that given part the particles ending it, behind any + post-nominals, which P6 attaches to the family — other than a + lone member of the ambiguous class, which reads as the + credential. A lone particle is + not a name word for the count, 'de Mesnil' being one surname. A title word + never starts the run, and neither does a member of the ambiguous + class, a single letter, a connective, a word that is also + particle or bound given-name vocabulary, or a non-Latin honorific + word: a credential is the writer's mark that the post-nominals + have begun, and those carry no such mark. A name word the run + absorbs is reported `suffix-or-name`. An unlisted word joins this same ambiguous class by SHAPE where the caller asks for it. Two or more period-separated chunks is one such shape, admitted by default (S3); an unlisted all-caps @@ -1209,6 +1226,16 @@ S2. Rationale: generational suffixes and credentials are recognized "II Van Johnson" → given="II" · boundary "Sir Ph. D. Van Johnson" → suffix="Ph. D." · boundary "Ph. D., John" → family="Ph. D." · boundary + "John Smith PhD Jones" → suffix="PhD Jones" + "John Smith PhD Jones" → ambiguities=("suffix-or-name",) + "Eric H. Holder Jr. Attorney General" → title="Attorney General" + "John Smith PhD Prof. Ma" → suffix="PhD Ma" + "Smith, John PhD Jones" → suffix="PhD Jones" + "Smith, John PhD de Jr." → family="de Smith" · boundary + "Mary Jane King Smith" → family="Smith" + "Josep Carod i Rovira" → family="Carod i Rovira" + "Mohamed Ali Abd Allah" → family="Allah" + "John PhD Smith" → family="Smith" · boundary Accepted: a title before the credential keeps it a credential, and the name loses its surname exactly as it did before this clause — `Sir Ph. D. Van Johnson` reads given `Van Johnson` with @@ -1224,18 +1251,19 @@ S2. Rationale: generational suffixes and credentials are recognized speak for it ends the name, words to spare or not, and whatever stands in front of it is name text. Where a credential does stand in front, the company above decides instead, except in - three shapes that keep it out of reach, each of which reads the - member as the name: a credential written split across two words - with no comma after the name, a title standing between the - credential and the member, and a member a particle chain (P2) - has already taken. Each reading of the three is older than the - company clause, so they are witnessed by tests/v2/cases.py's - a_split_degree_in_front_is_out_of_the_walk, + one shape that keeps it out of reach and reads the member as the + name: a credential written split across two words with no comma + after the name. That reading is older than the company clause, + so it is witnessed by tests/v2/cases.py's + a_split_degree_in_front_is_out_of_the_walk rather than by a line + here, which would bring the name into the corpus that enforces + it at released baselines for a reading this clause did not make. + A title standing between the credential and the member, and a + member a particle chain (P2) had taken, were two more such shapes + until the credential's run (#602) took both in: the rows a_title_between_degree_and_member_keeps_it_a_name and - a_member_the_particle_chain_took_is_out_of_reach rather than by - lines here, which would bring each name into the corpus that - enforces them at released baselines for a reading this clause - did not make. + a_member_the_particle_chain_took_is_out_of_reach keep their ids + and now pin the run. "Jack Wei Ma" → family="Ma" "Jack Wei Ma" → ambiguities=("suffix-or-name",) "abdul Smith Jr Ma" → suffix="Jr Ma" @@ -1259,7 +1287,7 @@ S2. Rationale: generational suffixes and credentials are recognized and unchanged (decisions.md#v1-xfail-triage: `king` stays a title, for the addressing forms). "Dr Jr" → suffix="Jr" - history: decisions.md#S2 · interacts: H1, H2, H3, H5, C1, C2, S3, P1, P2, P3, P5, P6, M2 · implemented: nameparser/_pipeline/_classify.py, nameparser/_pipeline/_group.py, nameparser/_pipeline/_pieces.py, nameparser/_pipeline/_vocab.py + history: decisions.md#S2 · interacts: H1, H2, H3, H5, C1, C2, S3, P1, P2, P3, P5, P6, M2 · implemented: nameparser/_pipeline/_assign.py, nameparser/_pipeline/_classify.py, nameparser/_pipeline/_group.py, nameparser/_pipeline/_pieces.py, nameparser/_pipeline/_vocab.py S3. Rationale: credentials are often written run together with periods; the chunks between the periods are what carry the @@ -1415,92 +1443,38 @@ M1. Rationale: an enclosure the caller has declared to mean maiden M2. Rationale: a maiden marker announces that what follows it is the former family name; the marker is an announcement, not a name. - A recognized maiden marker standing after at least one name - word takes the words after it — up to any suffix word, or the - trailing roman numeral assign reads as the suffix (S2), both as - written and as the take would leave the name, the word before - the numeral being then the word before the marker, or a - trailing word of the ambiguous credential class (S2), asked - that same double way and stopping the take only where the rule - reading the name left standing reads the word as the - credential, and never the first word after the marker — as the - maiden name, and - the marker itself is dropped. Where a trailing rule reads the - words, a trailing title ends it too: a period-marked word the - trailing title chain takes (H5), read together with the trailing - suffix run to where neither takes more, so a credential or a - numeral in front of the title stops the take exactly as it does - with the title absent. Neither the credential stop nor the title - stop takes the first word after the marker; the numeral stop is - not held to that, and a numeral there may decline the clause. - One suffix word does not stop it. Where such a word is also a - connective standing between two name words of the clause (P3), - a link inside the birth name does not end it, and the words on - both sides of the link are the maiden name. A link with the - marker on one side of it, or with the trailing run on the - other, is joining nothing there and ends the clause like any - other suffix word. Where a trailing rule reads the words, a link - that ends the clause only because a trailing title is read as - one — a link that, read over the words as written, would join — - gives up the words behind it only where the name left standing - reads them as post-nominals or titles; otherwise the clause keeps - the link and runs on. The link first after the marker is not - asked: stopping there declines the clause. - A separator the caller declared, standing in a trailing suffix - part, ends the clause as a comma would (C1): the words beyond it - are the next part's, and a marker with one straight before or - after it reads as it would with a comma typed there. - The trailing numeral, credential and title stops are each asked - TWICE for one reason: the - count of words to spare includes the very words the marker - removes, so a reading taken over the name as written can be - wrong about the name the take would leave. WHICH rule does the - reading depends on where the clause stands. Where the clause is - in the part a trailing rule reads — a name with no comma, and - the part before a SUFFIX comma, which that rule reads the same - way — that rule is the reader. After a family comma it is the - reading the end of the given part takes, where the comma has - already settled the count and the writing decides alone. A lone - numeral reads as a suffix there only where no comma part follows - the given one, exactly as it does outside a marker's clause; with - one behind it the numeral stays name text and the clause keeps - it. Before a family comma, and in a part after a second one, no - trailing rule reads those words at all: the clause keeps them and - says nothing about them. - That the credential stop spares the first word after the marker - is a deliberate divergence from what certain suffix vocabulary - gets in the same position, where the marker declines and stays - an ordinary word. The marker announces a name, and this class is - the one carrying no evidence of which it is: its members are - borne surnames as well as credentials, and nobody writes a - credential straight after the marker, so a lone member reads as - the name it was announced to be. - A member ENDING a clause that some rule reads is reported where - the clause KEEPS it (S2); one the clause gives up is reported - where the reading that took it reports, so no word is reported - twice and none goes unreported. - That holds because a word the clause gives up reads as a - post-nominal or the clause keeps it. The stop is right only - where the released word ends the parse in the SUFFIX, so a stop - that would put it in a name part is no stop and the clause keeps - the word. What the take LEAVES BEHIND decides that, not the - clause as written. Two shapes leave nothing that could read the - word as a credential. A part whose other words are all - post-nominals or titles has no name word left for a trailing - slot to be the end of, and a part of nothing but credentials is - read whole and asked nothing. And a join reached below the take - — a particle chain (P2), or a bound given-name join (P5) — can - absorb the released word into a name part before any trailing - rule sees it, which would carry a word of the BIRTH name into - the current one. In both the clause keeps the word, and reports - it as it reports every member it keeps. - Where a trailing rule reads the words, the trailing numeral, the - trailing credential and the trailing title each give up a run - only where the name the take leaves - reads the WHOLE run, not only the word the stop is made at, as - titles or post-nominals; otherwise the clause keeps it. After a - family comma a released title that is also a particle is kept, - since the particle attachment (P6) would carry it into the family. + A recognized maiden marker counts only in the part of the name + that holds the family name — the whole name where no comma + divides it or before a suffix comma, the part before a family + comma — and only with a + name word straight before it. A word of the leading title run is + no name word there, a lone title included; neither is a word of + the unambiguous suffix vocabulary, unless it opens the name (S2), + nor a connective. A particle + is one: with nothing after it to attach to, it is no particle + there. In the given part after a family comma, in a part after a + suffix comma, and behind a title, a suffix word or a connective, + the marker is an ordinary word; a name standing somewhere before + it is not enough. + A marker that counts takes the words after it as the maiden name, + and is itself dropped. It takes them up to the trailing run of + post-nominals and titles that the end of the name reads as if the + clause were not written — the trailing rules (S2, H5) asked over + the words before the marker and the post-nominal or title words + that end it, before any join — and always takes the first word + after it, the marker having announced a name. A suffix word + straight after the marker announces nothing, and the marker is + then a word. The run so found ends the name: its words are the + suffixes and titles that reading makes them, and no join reaches + into it. Before a family comma no trailing rule reads the part, so + the clause ends at the first suffix word, and the words from it + on read as that part reads them. + Where a trailing rule reads the part, a member of the ambiguous + class (S2) that ends the maiden name is reported where the clause + keeps it, and one in the run is reported where the run is read, + so no word is reported twice and none goes unreported. Before a + family comma no rule reads it, and the member the clause keeps + there is reported nowhere. Delimiters outrank every reading inside them. Where a recognized marker stands inside a delimited clause, the whole span is the maiden name whatever its last word is, and whether or not the @@ -1527,134 +1501,87 @@ M2. Rationale: a maiden marker announces that what follows it is the instead of absorbing it; a marker left as a word bounds nothing. Where the bound leaves a family of nothing but particles, they are not in particle position and read as ordinary words (R2). - "Jane Smith née Jones" → maiden="Jones" - "Jane née Jones Smith" → maiden="Jones Smith" - "Jane Smith née Jones PhD" → suffix="PhD" - "John née Jones Smith V" → maiden="Jones Smith" - "John née Jones Smith V" → suffix="V" - "Jane Smith née V" → suffix="V" - "Dr. nee V" → maiden="V" · boundary - "J. née Jones Smith V" → maiden="Jones Smith V" · boundary - "Jane née Jones J. V" → maiden="Jones J. V" · boundary - "Jane Doe nee Smith MA" → maiden="Smith" - "Jane Doe nee Smith MA" → suffix="MA" - "Jane Doe nee Smith Ma" → maiden="Smith Ma" · boundary - "Jane Doe nee MA" → maiden="MA" · boundary - "Jane Doe nee MA Smith" → maiden="MA Smith" · boundary - "John née Jones Smith MA" → maiden="Jones Smith" - "Doe, Dr. nee Smith MA" → maiden="Smith MA" · boundary - "Berg, abdul nee Jones MA" → maiden="Jones MA" · boundary - "Jane Doe nee Smith DO DO" → maiden="Smith DO DO" · boundary - "Jane Doe nee Puig i Soler" → maiden="Puig i Soler" - "Jane Doe nee Puig i" → maiden="Puig" · boundary - "Jane Doe nee Puig i" → suffix="i" · boundary - "Jane Doe nee Smith i DO Prof." → maiden="Smith i DO" - "Doe, Jane nee Smith St." → maiden="Smith St." - "Smith, John, PhD née Puig Mr. - i Soler" extra_suffix_delimiters-dash → maiden="Puig Mr." - "Smith, John, PhD née Puig - i Soler" extra_suffix_delimiters-dash → maiden="Puig" - "Jane Doe nee Smith Prof." → maiden="Smith" - "Jane Doe nee Smith Prof." → title="Prof." - "Jane Doe nee Smith MA Prof." → suffix="MA" - "Jane Doe nee Smith Prof. MA" → maiden="Smith" - "Jane Doe nee Smith V Prof." → suffix="V" - "Jane Doe nee King." → maiden="King." · boundary - "Jane Doe nee Prof. Dr." → maiden="Prof." · boundary - "Dr. nee Jones Smith Prof." → maiden="Jones Smith Prof." · boundary - "Doe nee Smith Prof., Jane" → maiden="Smith Prof." · boundary - "Jane van der Berg nee Smith Prof." → maiden="Smith Prof." · boundary - "Berg, abdul nee Smith V" → maiden="Smith V" · boundary - "Jane Doe nee King. ba" → suffix="ba" · boundary - "Doe, Jane nee Smith V, PhD" → maiden="Smith V" - "Jane Doe (nee Smith MA)" → maiden="Smith MA" - "Jane Doe (nee Smith Ma)" → maiden="Smith Ma" - "Jane Doe (nee Smith) MA" → suffix="MA" - "Jones née" → family="née" · boundary - "née Jones" → family="Jones" · boundary + "Jane Smith née Jones" → maiden="Jones" + "Maria López née García Pérez" → maiden="García Pérez" + "Jane Smith née Jones PhD" → suffix="PhD" + "John née Jones Smith V" → maiden="Jones Smith" + "John née Jones Smith V" → suffix="V" + "Doe nee Smith, Jane" → maiden="Smith" + "Doe, Jane nee Smith" → maiden="" + "Doe, Jane nee Smith" → middle="nee Smith" + "Jane Smith, née Jones" → maiden="" + "Smith, John, PhD née Jones" → maiden="" + "Jane Doe PhD nee Smith" → maiden="" + "Jane and née Jones" → maiden="" + "Dr. nee Smith PhD Prof." → maiden="" · boundary + "Dr. nee Smith PhD Prof." → family="Smith" · boundary + "Mai Le née Nguyen" → maiden="Nguyen" + "abd née Jones" → maiden="Jones" + "Jane née Jr y Jones" → maiden="" + "Jane Doe nee Smith MA" → maiden="Smith" + "Jane Doe nee Smith MA" → suffix="MA" + "Jane Doe nee Smith Ma" → maiden="Smith Ma" · boundary + "Jane Doe nee MA" → maiden="MA" · boundary + "Jane Doe nee MA Smith" → maiden="MA Smith" · boundary + "John née Jones Smith Ma" → maiden="Jones Smith Ma" + "John nee Prof. ba MA" → maiden="Prof. ba" · boundary + "Jane Doe nee Puig i Soler" → maiden="Puig i Soler" + "Jane Doe nee Puig i" → maiden="Puig" · boundary + "Jane Doe nee Puig i" → suffix="i" · boundary + "Jane Doe nee Smith Prof." → maiden="Smith" + "Jane Doe nee Smith Prof." → title="Prof." + "Jane Doe nee Smith MA Prof." → suffix="MA" + "Jane Doe nee Smith Prof. MA" → maiden="Smith" + "Jane Doe nee Smith V Prof." → suffix="V" + "Jane Doe nee King." → maiden="King." · boundary + "Jane Doe nee Prof. Dr." → maiden="Prof." · boundary + "Jane van der Berg nee Smith Prof." → title="Prof." + "Jane van der Berg nee Prof. King. MA" → maiden="Prof." · boundary + "Jane Doe nee Smith PhD Smith" → maiden="Smith" + "Jane Doe nee Smith PhD Smith" → suffix="PhD Smith" + "Doe nee Smith Jr. Prof., Jane" → family="Doe Prof." · boundary + "Jane Doe (nee Smith MA)" → maiden="Smith MA" + "Jane Doe (nee Smith Ma)" → maiden="Smith Ma" + "Jane Doe (nee Smith) MA" → suffix="MA" + "Jones née" → family="née" · boundary + "née Jones" → family="Jones" · boundary "Jane van der Berg née Jones" → maiden="Jones" - "Jane de la née Jones" → family="de la" - "Jane van der Berg née" → family="van der Berg née" - "Jane van der Berg née PhD" → family="van der Berg née" + "Jane de la née Jones" → maiden="Jones" + "Jane de la née Jones" → family="de la" + "Jane van der Berg née" → family="van der Berg née" + "Jane van der Berg née PhD" → family="van der Berg née" "Jane van der Berg née y Jones" → maiden="y Jones" - "van der Berg, abdul née Jones" → maiden="Jones" - "Maria Kowalska z domu Nowak" → maiden="Nowak" - "Anna z Nowak" → family="Nowak" · boundary - "Anna z (domu) Nowak" → family="Nowak" · boundary + "Maria Kowalska z domu Nowak" → maiden="Nowak" + "Anna z Nowak" → family="Nowak" · boundary + "Anna z (domu) Nowak" → family="Nowak" · boundary Accepted: the fullwidth-colon spelling arrives as one word, so the marker inside it goes unrecognized; #317 tracks whether it should peel. - "山田 花子 旧姓:佐藤" → maiden="" - Accepted: a marker straight after a comma is post-comma given - text, not a marker. - "Jane Smith, née Jones" → maiden="" - Accepted: the marker reads the words as written, so a suffix word - inside the maiden name ends it even where a connective beside it - would have bound the two into one name word (P3); the connective - then builds a family name out of what is left. A suffix word - that IS the connective is the exception stated above, and ends - the clause only where it joins nothing. - "Jane née Jr y Jones" → maiden="" - Accepted: a bare acronym the reading declines is maiden text all - the same — the writing decides this one (S2), and the count such - a reading needs is taken over the name the take would leave - rather than over the words as they stand. - "John née Jones Smith Ma" → maiden="Jones Smith Ma" + "山田 花子 旧姓:佐藤" → maiden="" + Accepted: a given name with no current surname beside a maiden + name reads by M4's lone-word rule, the lone name word before the + marker being the family. A writer giving a given name with its + birth name writes the clause delimited. + "Jane née Jones Smith" → maiden="Jones Smith" · boundary + Accepted: the run is what the trailing rules read, before any + join — not the whole reading of the clause-free name. Where a + particle chain would take the words of the run in that name, the + run keeps them suffixes and titles. + "Jane Doe nee Smith DO DO" → suffix="DO DO" · boundary + Accepted: the first word after the marker is the maiden name + whatever it is unless it is a suffix word, and a lone roman + numeral is not one — it is shaped like an initial — so it is the + maiden name where the clause-free name would read it as the + suffix. Decided with #601 (decisions.md#M2, 2026-10-04). + "Jane Smith née V" → maiden="V" · boundary Accepted: bracket content ending in a period is suffix-shaped (M3), so the brackets are dropped (S1) and the clause is read as if written bare — the delimiter precedence above does not reach - it, and a trailing title inside gives itself up as the bare - clause's does. - "Jane Doe (nee Smith Prof.)" → title="Prof." - Accepted: the title is transparent only where the clause gives - the credential up. A clause is one run of words, so where it - KEEPS the credential a title written behind it can still leave, - while one written in front of it cannot leave without the words - behind it, and the two spellings differ. - "Doe nee Smith ba Prof." → maiden="Smith ba" - "Doe nee Smith Prof. ba" → maiden="Smith Prof. ba" - Accepted: where the clause gives the credential up in both - spellings, the title still leaves only where nothing ahead stands - to take it. A particle chain ahead runs on over a trailing title - (H5), and a bound given-name join or a name left with no name - word would absorb it too, so a title in front of the credential - stays in the clause while one behind it, which the first suffix - word has already cut off, leaves — the split H5 already accepts - for the same name written without a marker. - "Jane van der Berg nee Smith PhD Prof." → title="Prof." - "Jane van der Berg nee Smith Prof. PhD" → maiden="Smith Prof." - Accepted: a released particle with a title behind it is withdrawn, - because the particle chain would run on over the title (P2, H5); - so the particle written in front of the title stays in the clause, - while written behind it, it leaves with the title. A particle - standing INSIDE the clause, ahead of the credential, does the - same from the other side: its chain would take the credential, - so the clause keeps the credential, and a title can leave only - from behind it. - "Jane Doe nee Smith DO Prof." → maiden="Smith DO" - "Jane Doe nee Smith Prof. DO" → suffix="DO" - "Jane Doe nee Smith do MA Prof." → maiden="Smith do MA" - "Jane Doe nee Smith do Prof. MA" → suffix="MA" - Accepted: H5's reach into the ordinary surnames the title - vocabulary holds reaches the end of a clause as it reaches the end - of a name: a period written behind one ends the clause as a title - and takes the word out of the birth name. - "Jane Doe nee Smith King." → title="King." - Accepted: the stop at the first suffix word asks no question of - what it gives up, so a clause that ends there can hand a word - behind it to a name part, where the invariant stated above says - the clause keeps it. Before a family comma, where the statement - above says no trailing rule reads the words and the clause keeps - them, the trailing numeral's stop is still made, asked of the - peel alone with no such question, and a lone numeral written - there goes to the family. Open: #548. - "Doe nee Smith Jr. Prof., Jane" → family="Doe Prof." - "Doe nee Smith V, Jane" → family="Doe V" - Accepted: a credential written in front of the marker speaks for - no word of the clause (S2's company), the clause's own words - standing between the two, so the clause keeps a member its - writing declines even where the same name written without the - clause reads that member as the credential. - "Jane Doe Jr. nee Smith Ma" → maiden="Smith Ma" - "Jane Doe Jr. Ma" → suffix="Jr. Ma" · boundary + it. + "Jane Doe (nee Smith Prof.)" → title="Prof." + Accepted: a marker behind a credential is an ordinary word, and + the credential's run (S2) takes it in with everything after it. + "Jane Doe Jr. nee Smith" → suffix="Jr. nee Smith" · boundary history: decisions.md#M2 · interacts: P2, P3, P5, P6, R1, R2, M1, S1, S2, H1, H5, C1 · implemented: nameparser/_pipeline/_group.py M3. Rationale: an enclosure says nothing about whether it means @@ -2015,13 +1942,12 @@ C1. Rationale: a credential run after the comma means the name is in here. It is also the one consequence of this question Latin script cannot witness, nothing there being glued to the end of a name — which is why this clause carries no example of its own. - Accepted: a maiden marker with a declared delimiter straight - before or after it takes no clause, the same words written with - a comma there taking none either: a marker opening a trailing - part has nothing ahead of it to follow, and one ending a part has - nothing behind it to take (M2). Releases 2.0 through 2.3 read - past the delimiter and took the clause; the comma reading is - accepted as the cost of one rule for a declared separator. + Accepted: a maiden marker in a part a declared delimiter + separates takes no clause, as no marker in a part after a suffix + comma does, the marker being an ordinary word there (M2). + Releases 2.0 through 2.3 read past the delimiter and took the + clause; the comma reading is accepted as the cost of one rule + for a declared separator. "Smith, John, MD - née Jones Smith" extra_suffix_delimiters-dash → suffix="MD, née Jones Smith" Accepted: a delimiter core the policy names (T1) is a word here, not structure — v1 applied the delimiter to the suffix-comma @@ -2065,8 +1991,8 @@ C2. Rationale: text beyond the recognized comma parts should be A part the parse consumes wholly as suffixes raises no report about reading a word of it as a name, whether it is the part after a suffix comma (C1) or a part beyond the second: nothing - in it is read as one outside a maiden clause standing in it - (M2), so a particle chain run over it (P2) reports neither a + in it is read as one, a maiden marker there being an ordinary + word (M2), so a particle chain run over it (P2) reports neither a particle chained onto a name word nor an acronym taken into the name. What such a part reports is its own — C1's flip and this rule's flag. diff --git a/docs/release_log.rst b/docs/release_log.rst index b3020b77..405b4c3f 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -30,15 +30,13 @@ Release Log - **Fix a credential ending the given part of a family-comma listing being read as a middle name in silence.** ``HumanName("Doe, John MA")`` gives first ``John``, last ``Doe``, suffix ``MA``, where 2.0 through 2.3 gave middle ``MA`` -- and 1.4.0 gave the suffix, so this restores v1's reading for that half. The comma has already named the family and the first word after it is the given name, so the words-to-spare count that governs the comma-less form is satisfied by construction and the writing decides alone: ``Doe, John Ma`` keeps middle ``Ma``, written the way a name is written, and ``Doe, John Ed`` keeps middle ``Ed``. Either reading is now reported, and the report belongs to the SPELLING rather than to the fields -- a declined name re-rendered without its comma, ``John Ma Doe``, re-parses to those same three fields and reports nothing, the word no longer standing where the question is asked. A name word behind the credential still ends its reach and stays silent -- ``Doe, John MA Smith`` gives middle ``MA Smith`` and reports nothing -- while a credential run or a trailing title is transparent to it: ``Doe, John MA PhD`` gives suffix ``MA PhD`` and ``Doe, John MA Prof.`` gives title ``Prof.`` with suffix ``MA``. Two second-order movements an upgrader may see, both consequences of the word leaving the given part rather than of this rule reaching further: ``Doe, John Prof. MA`` now gives title ``Prof.`` where it gave middle ``Prof. MA``, the trailing-title chain reaching a word the credential used to hide; and ``Doe, John van MA`` gives last ``van Doe`` with suffix ``MA`` where it gave middle ``van MA``, the surname-particle rule reaching a particle the same way. A name written wholly in one case says nothing either way and takes the credential, which is what 1.4.0 read: ``DOE, MARY JO MA``, ``doe, john ma``, ``田中, 太郎 MA`` and ``김, 민준 MA`` all give a suffix. The unlisted dotted spelling moves with them without the parity claim -- ``Doe, John X.Y.Z.`` gives suffix ``X.Y.Z.`` where 1.4.0 and 2.3.0 both gave a middle name -- to match the comma-less ``John Doe X.Y.Z.``. One word is carved out: ``do`` is the only member of this class that is also a surname particle, so capitals decide it and, with nothing in front of it, the particle reading keeps every other spelling (a degree in front is the other exception, the next-but-one entry). ``Doe, John DO`` gives suffix ``DO``, while ``Doe, John do``, ``Doe, John Do``, ``DOE, JOHN DO`` and ``doe, john do`` are unchanged and keep the particle-or-given report they already had. In a name written wholly in one case the two cannot be told apart, so ``SMITH, JOHN DO`` keeps last ``DO SMITH`` as ``NASCIMENTO, EDSON ARANTES DO`` does -- right about the Portuguese record, wrong about the osteopath, and the report is how a caller finds the second. See the ``S2`` and ``P6`` entries of ``docs/design/decisions.md`` (closes #531) - - **Fix a maiden marker's clause swallowing a trailing credential in silence.** ``HumanName("Jane Doe nee Smith MA")`` gives maiden ``Smith`` with suffix ``MA``, where 2.0 through 2.3 gave maiden ``Smith MA`` and said nothing; 1.4.0 read the ``MA`` as a suffix too. ``Doe, Jane nee Smith MA`` moves with it, and so do the one-case spellings ``JANE DOE NEE SMITH MA`` and ``jane doe nee smith ma``. The words a marker takes now end where a trailing credential begins, which is what the marker's other two stops -- a suffix word, a trailing roman numeral -- have always done. Until this release it was the last trailing position in the library where a word of the ambiguous credential class was read without a report, and it was order-sensitive besides: ``Jane Doe nee Smith MA PhD`` gave maiden ``Smith MA`` while ``Jane Doe nee Smith PhD MA`` gave maiden ``Smith``, so whether the word was read at all depended on which side of the unambiguous credential the writer put it. Both now give maiden ``Smith``, with suffix ``MA PhD`` and ``PhD MA``. The writing still decides, exactly as it does for the same word ending a name with no clause: ``Jane Doe nee Smith Ma`` keeps maiden ``Smith Ma``, and ``Jane Doe nee Yo-Yo Ma`` keeps a two-word birth surname whole. The one member of this class that is also a surname particle keeps the carve-out it has outside a clause -- ``Doe, Jane nee Smith DO`` gives suffix ``DO`` while ``Doe, Jane nee Smith do`` and ``Doe, Jane nee Smith Do`` keep maiden ``Smith do`` and ``Smith Do``, and the comma-less ``Jane Doe nee Smith do`` gives suffix ``do`` as ``John Doe do`` does. Either reading is now reported, and there is no third: a word the clause gives up reads as a post-nominal, or the clause keeps it and says so. A name word behind the credential ends its reach and stays silent -- ``Jane Doe nee MA Smith`` gives maiden ``MA Smith`` and reports nothing -- and this stop never takes the first word after the marker, whatever its writing says: ``Jane Doe nee MA`` keeps maiden ``MA`` and reports, the marker having announced a name where there would otherwise be none, and ``Jane Doe nee MA PhD`` keeps it too. That differs on purpose from what a certain post-nominal gets there, ``Jane Smith nee PhD`` and ``Jane Smith nee V`` leaving the marker standing as an ordinary word as before. Where no trailing rule reads the clause's tail nothing is decided and the clause keeps every word: ``Smith nee Jones MA, Jane`` and ``Smith, John, Jr nee Jones MA`` both keep maiden ``Jones MA``, unchanged and with no ``suffix-or-name`` report. A title written behind or in front of the credential usually does not change how the credential is read; where something in or ahead of the clause could take the title it can -- see the trailing-title bullet below. The clause also keeps a word it cannot promise a credential reading for, which is where three shapes that look like they should move do not. Where the part the word would land in holds no name of its own there is nothing to read it as a credential, so ``Doe, Dr. nee Smith MA`` and ``Jane Doe, Jr nee Smith MA`` both keep maiden ``Smith MA`` and report. Where a join would swallow it first the same applies, and it is the birth name that would lose the word: ``Berg, abdul nee Jones MA`` keeps maiden ``Jones MA`` rather than reading first ``abdul MA``, and ``Berg, Jane van der nee Smith DO`` keeps maiden ``Smith DO`` rather than letting the particle chain carry the ``DO`` into last ``van der DO Berg``. Each of those reads as 2.3.0 read it. Delimiters settle the question outright and always did: ``HumanName("Jane Doe (nee Smith MA)")`` keeps the whole span as the maiden name and reports nothing, the writer having drawn the boundary, while ``Jane Doe (nee Smith) MA`` gives suffix ``MA`` for the word left outside it. One name is a restoration rather than a change: ``John Smith nee Jones R.A.I.`` gives suffix ``R.A.I.`` again, as 2.3.0 read it, this unreleased cycle having moved it into the maiden name when the unlisted-dotted reading above took the word out of the certain-suffix class. See the ``M2`` and ``S2`` entries of ``docs/design/decisions.md`` (closes #533) + - **Change where a maiden marker takes a maiden name: only behind a surname, and only up to the credentials and titles the name ends with.** ``HumanName("Jane Doe nee Smith")``, ``Mai Le née Nguyen`` and ``Doe nee Smith, Jane`` give maiden ``Smith`` and ``Nguyen`` as before, but a marker counts only in the name before any comma or in the surname part before a family comma, and only behind a name word. Anywhere else it is an ordinary word, as in 1.4.0: ``Doe, Jane nee Smith`` gives middle ``nee Smith``, ``Smith, John, PhD née Jones`` gives suffix ``PhD née Jones``, and ``Dr. nee Smith`` gives first ``nee``, last ``Smith``, where 2.0 through 2.3 gave maiden ``Smith``, ``Jones`` and ``Smith`` -- the last with no name at all. A credential in front of the marker ends the name (see the credential-run entry below), so ``Jane Doe PhD nee Smith`` gives suffix ``PhD nee Smith`` where 2.0 through 2.3 gave suffix ``PhD``, maiden ``Smith``. A particle or a lone word is a surname here: ``Jane van nee Smith`` gives last ``van``, maiden ``Smith``, and ``Jane Smith née V`` gives last ``Smith``, maiden ``V``, as 2.0 and 2.1 read it, where 2.2 and 2.3 gave last ``née``, suffix ``V``. The words the marker takes now end where the run of post-nominals and titles the name would end with if the clause were not written begins, and the words it gives up read as they would there: ``Jane Doe nee Smith MA`` gives maiden ``Smith``, suffix ``MA``, where 2.0 through 2.3 gave maiden ``Smith MA`` and said nothing, as do the one-case ``JANE DOE NEE SMITH MA`` and ``jane doe nee smith ma``; ``Jane Doe nee Smith MA PhD`` gives suffix ``MA PhD`` where 2.3.0 gave maiden ``Smith MA``, suffix ``PhD``, so the two orders of the credentials now agree; ``Jane Doe nee Smith DO DO`` gives suffix ``DO DO``; ``Jane Doe nee Smith Prof.``, ``Mary Smith née Jones Prof.`` and ``Jane van der Berg nee Smith Prof.`` give title ``Prof.``; ``Jane Doe nee Smith King.`` gives title ``King.``; ``Jane Doe nee Smith MA Prof.`` and ``Jane Doe nee Smith Prof. MA`` both give title ``Prof.``, suffix ``MA``; and ``Jane Doe nee Smith V Prof.`` gives suffix ``V`` -- each where 2.3.0 kept every word in the maiden name. A credential the clause gives up or keeps at that boundary is reported as ``suffix-or-name``. The writing still decides, as it does at the end of a name with no clause: ``Jane Doe nee Smith Ma`` keeps maiden ``Smith Ma``, and ``Jane Doe nee Yo-Yo Ma`` keeps a two-word birth surname whole. The first word after the marker is always taken, the marker having announced a name: ``Jane Doe nee MA`` and ``Jane Doe nee King.`` keep it, and ``Jane Doe nee Prof. Dr.`` gives maiden ``Prof.``, title ``Dr.``, where 2.3.0 gave maiden ``Prof. Dr.``. Before a family comma the clause keeps every word, as before: ``Doe nee Smith Prof., Jane`` keeps maiden ``Smith Prof.``. Delimiters settle the question outright: ``Jane Doe (nee Smith MA)`` keeps the whole span, and brackets around a clause ending in a period are dropped as in 2.3.0, so ``Jane Doe (nee Smith Prof.)`` gives title ``Prof.``. The name #548 reported, ``Dr. nee Smith PhD Prof.``, gives title ``Dr. Prof.``, first ``nee``, last ``Smith``, suffix ``PhD``, where 2.3.0 gave last ``PhD``, maiden ``Smith``. ``John Smith nee Jones R.A.I.`` gives suffix ``R.A.I.``, as 2.3.0 read it. The rule replaces the clause rules this cycle's #533 and #535 had added, and the readings those changes gave here are superseded. See the ``M2`` entry of ``docs/design/decisions.md`` (closes #601) - **Fix a comma part read wholly as suffixes reporting a particle in it as chained onto a name.** ``parse("John Smith, Jr., Freiherr von Richthofen").ambiguities`` names ``comma-structure`` alone, where 2.0 through 2.3 also named ``particle-or-given`` for ``von`` -- a word the same parse had put in the suffix, so the report described a reading it never made. No field moves. A part the parser consumes as suffixes, after a suffix comma or past the second comma, reports what the part is and nothing about its words as names, and the particle chain's credential-acronym report this release adds (above) keeps the same bound: ``John Smith, PhD Do Ma`` reports the comma's decision once and not again for ``Ma``. See the 2026-09-28 bullet of the ``C1`` entry in ``docs/design/decisions.md`` - - **Fix a trailing title after a maiden marker being read as part of the maiden name.** ``HumanName("Jane Doe nee Smith Prof.")`` gives title ``Prof.`` with maiden ``Smith``, where 2.0 through 2.3 gave maiden ``Smith Prof.``, and ``Mary Smith née Jones Prof.`` moves the same way. A credential or roman numeral in front of the title is read as it is with the title absent, so the two spellings ``Jane Doe nee Smith MA Prof.`` and ``Jane Doe nee Smith Prof. MA`` now agree -- title ``Prof.``, maiden ``Smith``, suffix ``MA`` -- where 2.3.0 kept every word in the maiden name (``Smith MA Prof.`` and ``Smith Prof. MA``), and ``Jane Doe nee Smith V Prof.`` gives suffix ``V``. Where the clause keeps the word in front of the title they still differ, since a title cannot leave without the words behind it: ``Doe nee Smith ba Prof.`` gives title ``Prof.`` with maiden ``Smith ba``, while ``Doe nee Smith Prof. ba`` keeps maiden ``Smith Prof. ba`` as 2.3.0 kept all three words in both. They can also differ where something in or ahead of the clause could take the title -- a particle, a bound given name, or a name left with no name word: ``Jane Doe nee Smith do MA Prof.`` keeps maiden ``Smith do MA`` while ``Jane Doe nee Smith do Prof. MA`` gives maiden ``Smith do``, suffix ``MA``, both with title ``Prof.``. A period behind an ordinary surname that is also a title makes it one here as it does at the end of any name: ``Jane Doe nee Smith King.`` gives title ``King.`` with maiden ``Smith``, where 2.3.0 gave maiden ``Smith King.``. The given part after a family comma reads it the same way: ``Doe, Jane nee Smith MA Prof.`` gives title ``Prof.`` too. A title straight after the marker stays the maiden name, the marker having announced one: ``Jane Doe nee King.`` keeps maiden ``King.``, and ``Jane Doe nee Prof. Dr.`` keeps maiden ``Prof.`` and gives title ``Dr.``, where 2.3.0 gave maiden ``Prof. Dr.``. Where the title itself would end the clause, the clause keeps it wherever giving it up would put it in a name part, and each of these reads as 2.3.0 read it: before a family comma (``Doe nee Smith Prof., Jane`` keeps maiden ``Smith Prof.``), where no name word would be left in front of it (``Dr. nee Jones Smith Prof.``), and behind a particle whose chain would take it (``Jane van der Berg nee Smith Prof.``). Brackets around a clause ending in a period are dropped before any of this is read, as in 2.3.0, so ``Jane Doe (nee Smith Prof.)`` gives title ``Prof.`` too. One reading moves because of the name the title is now read against: ``abdul nee Smith Dr.`` gives title ``Dr.`` with last ``abdul``, as ``abdul Dr.`` does, where 2.3.0 gave first ``abdul`` with maiden ``Smith Dr.``. See the ``M2`` entry of ``docs/design/decisions.md`` (closes #535) + - **A credential after the name core starts a suffix run to the end of its part.** ``HumanName("John Smith PhD Jones")`` gives suffix ``PhD Jones`` and reports ``suffix-or-name`` for the name word it took, where 2.3.0 gave middle ``Smith PhD``, last ``Jones``; ``Smith, John PhD Jones`` gives suffix ``PhD Jones`` where 2.3.0 gave middle ``Jones``, suffix ``PhD``. A title in the run is a title, so ``Eric H. Holder Jr. Attorney General`` gives title ``Attorney General``, suffix ``Jr.``, as the comma spelling ``Eric H. Holder, Jr. Attorney General`` already read, where 2.3.0 gave middle ``H. Holder Jr. Attorney``, last ``General``. Only an unambiguous credential or generational word starts the run, and only behind the name core -- two name words with no comma, or the given part after one. A credential that is also a name does not start one, nor does a surname particle or a single letter, so ``John Smith MA Jones`` and ``Mohamed Ali Abd Allah`` read as before, and a bare title word starts nothing: ``Mary Jane King Smith`` keeps middle ``Jane King``. See the ``S2`` entry of ``docs/design/decisions.md`` (closes #602) - - **Fix a roman numeral ending a maiden clause landing in the first or middle name.** ``HumanName("Berg, abdul nee Smith V")`` gives first ``abdul`` with maiden ``Smith V``, where 2.2 and 2.3 gave first ``abdul V`` -- a word of the birth name joined into the current given name -- and 2.0 and 2.1 gave first ``abdul nee``. ``Doe, Jane nee Smith V, PhD`` gives maiden ``Smith V``, as 2.0 and 2.1 did, where 2.2 and 2.3 gave middle ``V``: after a family comma a lone numeral ending the given part is a suffix only when no further comma part follows it, and the clause now asks that as the given part itself does. Without the credential tail the numeral still goes to the suffix (``Doe, Jane nee Smith V`` gives maiden ``Smith``, suffix ``V``). A numeral straight after the marker can leave the marker an ordinary word, as ``Jane Smith née V`` does, and where it does a title behind the numeral no longer changes that: ``Jane Doe nee V Prof.`` gives title ``Prof.``, suffix ``V``, middle ``Doe``, last ``nee``, where 2.3.0 gave maiden ``V Prof.``. Elsewhere the clause keeps it -- ``Dr. nee V`` keeps maiden ``V`` as before -- and ``Doe, Jane nee V, PhD`` now gives maiden ``V``, suffix ``PhD``, as 2.0 and 2.1 did, where 2.2 and 2.3 gave middle ``nee V`` -- the same further-comma-part condition as ``Doe, Jane nee Smith V, PhD`` above -- with ``Doe, Jane nee V, Jr.`` moving the same way. See the ``M2`` entry of ``docs/design/decisions.md`` (#535) - - - **Add the Catalan and Polish surname link.** ``parse("Josep Carod i Rovira")`` gives family ``Carod i Rovira``, where every release from 1.4.0 through 2.3.0 gave middle ``Carod i`` with family ``Rovira``; ``Josep Lluis Carod i Rovira`` gives middle ``Lluis`` with that same family; and ``Carod i Rovira, Josep`` gives it too, where they read family ``Carod Rovira`` and took the link into ``suffix`` as a generation marker. ``i`` is connective vocabulary now, the way ``y`` already was, and a connective counts as a name word wherever the three-word carve-out counts them -- whatever else the vocabulary says the word is, which matters here because ``i`` is also the roman numeral. A connective that is also generational vocabulary joins only where a name word stands on each side of it, so ``John Quincy Smith i`` keeps suffix ``i``, ``Josep Lluis Carod i III`` keeps suffix ``i III``, and the two-word ``Carod i`` keeps its generation reading. Written wholly in one case the letter reads as an initial and says so: ``JOSEP CAROD I ROVIRA`` and ``josep carod i rovira`` keep the fields they had and gain a ``conjunction-or-initial`` report, which a one-case name gains wherever a bare ``i`` or ``I`` stands among the name's own words -- a letter inside a maiden clause is read by the clause's rules and stays silent, as ``e`` already was -- and in an all-lower name that reading can move a field, each such name now reading as its all-caps twin already did (``parse("john smith i jr")`` gives middle ``smith``, family ``i`` and suffix ``jr`` where it gave family ``smith`` and suffix ``i jr``). Case repair follows the reading: a lower-case ``i`` the parse read as the generation is still title-cased by ``capitalize(force=True)`` (``Carod i`` gives ``Carod I``, as every release did), while one standing among the name words keeps its lower case as ``y`` always has (``Carod i Rovira`` gives ``Carod i Rovira``, where ``Carod I Rovira`` was the pre-2.4 answer). A link inside a maiden clause stays in the birth name, which no release read that way: ``HumanName("Jane Doe nee Puig i Soler")`` gives maiden ``Puig i Soler`` with last ``Doe``, where 2.0 through 2.3 ended the birth name at the link and gave maiden ``Puig`` with middle ``Doe i``, last ``Soler`` -- and the same words would have joined into last ``Doe i Soler`` under the change above, carrying a word of the birth name into the current surname. ``Doe, Jane nee Puig i Soler`` and ``Jane Doe née Kowalska i Nowak`` move with it, as does the all-lower ``jane doe nee puig i soler``; the ``y`` spelling always read this way and is untouched. The link still has to be joining: ``Jane Doe nee Puig i`` keeps maiden ``Puig`` with suffix ``i``, and ``Jane Doe nee Puig i III`` suffix ``i III``. A caller with Catalan or Polish data removes the entry from ``conjunctions_ambiguous`` and gets the join in the one-case names too; a caller who wants none of this removes ``i`` from ``conjunctions``, which restores every prior FIELD and every prior report, with two readings it does not restore and cannot: a letter the two vocabularies disagree about being an initial reads as one here and as the generation there, and case repair leaves a connective the parse placed among the NAME words in lower case where the off switch title-cases it -- ``parse("Dr. John i Smith").capitalized(force=True)`` keeps ``i`` where the off switch gives ``Dr. John I Smith``, and ``Carod i Rovira`` and ``Josep i Rovira`` are the same shape. Those two are the whole of what the switch does not undo, and ``tests/v2/test_properties.py`` states them as its invariants' only exemptions. A delimiter the caller declares through ``Policy(extra_suffix_delimiters=...)`` parts a trailing suffix part as a comma does (see the suffix-delimiter entry below), so no link joins across one: under ``(" - ",)``, ``Smith, John, PhD née Puig Mr. - i Soler`` keeps maiden ``Puig Mr.``, ``Smith, John, PhD née Puig - i Soler`` keeps maiden ``Puig`` with suffix ``PhD, i Soler``, and ``Smith, John, PhD - i Soler`` keeps suffix ``PhD, i Soler``, all as 2.3.0 read them. The default policy declares no such delimiter. See the ``P3`` and ``M2`` entries of ``docs/design/decisions.md`` (closes #397, closes #538) + - **Add the Catalan and Polish surname link.** ``parse("Josep Carod i Rovira")`` gives family ``Carod i Rovira``, where every release from 1.4.0 through 2.3.0 gave middle ``Carod i`` with family ``Rovira``; ``Josep Lluis Carod i Rovira`` gives middle ``Lluis`` with that same family; and ``Carod i Rovira, Josep`` gives it too, where they read family ``Carod Rovira`` and took the link into ``suffix`` as a generation marker. ``i`` is connective vocabulary now, the way ``y`` already was, and a connective counts as a name word wherever the three-word carve-out counts them -- whatever else the vocabulary says the word is, which matters here because ``i`` is also the roman numeral. A connective that is also generational vocabulary joins only where a name word stands on each side of it, so ``John Quincy Smith i`` keeps suffix ``i``, ``Josep Lluis Carod i III`` keeps suffix ``i III``, and the two-word ``Carod i`` keeps its generation reading. Written wholly in one case the letter reads as an initial and says so: ``JOSEP CAROD I ROVIRA`` and ``josep carod i rovira`` keep the fields they had and gain a ``conjunction-or-initial`` report, which a one-case name gains wherever a bare ``i`` or ``I`` stands among the name's own words -- a letter inside a maiden clause is read by the clause's rules and stays silent, as ``e`` already was -- and in an all-lower name that reading can move a field, each such name now reading as its all-caps twin already did (``parse("john smith i jr")`` gives middle ``smith``, family ``i`` and suffix ``jr`` where it gave family ``smith`` and suffix ``i jr``). Case repair follows the reading: a lower-case ``i`` the parse read as the generation is still title-cased by ``capitalize(force=True)`` (``Carod i`` gives ``Carod I``, as every release did), while one standing among the name words keeps its lower case as ``y`` always has (``Carod i Rovira`` gives ``Carod i Rovira``, where ``Carod I Rovira`` was the pre-2.4 answer). A link inside a maiden clause stays in the birth name, which no release read that way: ``HumanName("Jane Doe nee Puig i Soler")`` gives maiden ``Puig i Soler`` with last ``Doe``, where 2.0 through 2.3 ended the birth name at the link and gave maiden ``Puig`` with middle ``Doe i``, last ``Soler`` -- and the same words would have joined into last ``Doe i Soler`` under the change above, carrying a word of the birth name into the current surname. ``Jane Doe née Kowalska i Nowak`` moves with it, as does the all-lower ``jane doe nee puig i soler``; the ``y`` spelling always read this way and is untouched. The link still has to be joining: ``Jane Doe nee Puig i`` keeps maiden ``Puig`` with suffix ``i``, and ``Jane Doe nee Puig i III`` suffix ``i III``. A caller with Catalan or Polish data removes the entry from ``conjunctions_ambiguous`` and gets the join in the one-case names too; a caller who wants none of this removes ``i`` from ``conjunctions``, which restores every prior FIELD and every prior report, with two readings it does not restore and cannot: a letter the two vocabularies disagree about being an initial reads as one here and as the generation there, and case repair leaves a connective the parse placed among the NAME words in lower case where the off switch title-cases it -- ``parse("Dr. John i Smith").capitalized(force=True)`` keeps ``i`` where the off switch gives ``Dr. John I Smith``, and ``Carod i Rovira`` and ``Josep i Rovira`` are the same shape. Those two are the whole of what the switch does not undo, and ``tests/v2/test_properties.py`` states them as its invariants' only exemptions. A delimiter the caller declares through ``Policy(extra_suffix_delimiters=...)`` parts a trailing suffix part as a comma does (see the suffix-delimiter entry below), so no link joins across one: under ``(" - ",)``, ``Smith, John, PhD - i Soler`` keeps suffix ``PhD, i Soler``, as 2.3.0 read it. The default policy declares no such delimiter. See the ``P3`` and ``M2`` entries of ``docs/design/decisions.md`` (closes #397, closes #538) - **Fix a connective contributing no initial even where it is joining nothing.** ``parse("Juan de y").initials()`` gives ``J. y.``, where every release gave ``J.`` while ``family_base`` said ``y`` -- two views of one parse disagreeing about one token. A connective contributes nothing where it is JOINING, and initials like any other name word where its part holds nothing else for it to join. One rule for all three groups, so ``John and Jane Smith`` gives ``J. J. S.`` where 2.0 through 2.3 gave ``J. a. J. S.`` and 1.4.0 the run-together ``J a J. S.``, ``Duke of Edinburgh`` gives ``D. E.`` where 2.0 through 2.3 gave ``D. o. E.`` and 1.4.0 ``D o E.``, and ``John & Jane`` gives ``J. J.``. The question is asked of the whole part and never of a word count, so ``Jon Dough and`` has base ``Dough and`` and keeps ``J. D.``, and ``Juan Velasquez y Garcia`` keeps ``J. V. G.``. ``HumanName.initials()`` moves with the core -- over the differential corpora the two surfaces move on the same names and give the same values, reading one mark. Two names come back into 1.4.0 parity rather than away from it: ``JUAN Y GARCIA`` and ``محمد و علي`` both give the answer 1.4.0 gave. Parsing got cheaper by the same change -- the marks come off one pass instead of two, six fewer Python frames per name on 3.11. Two limits carried over from the 2.4 facade fix above: case repair still keeps such a connective lower-case, so ``initials()`` and ``capitalize()`` disagree about it on purpose, and a name restored from a pickle or a copy, or built from keyword fields, carries no tags and takes the older reading. See the ``R3`` entry of ``docs/design/decisions.md`` (closes #461) @@ -70,7 +68,7 @@ Release Log - **Fix a katakana name typed with a separate voicing mark being read family-first.** ``HumanName("ア゙イ タロウ")`` gives first ``ア゙イ``, last ``タロウ``, as ``アイ タロウ`` does, where 2.1 through 2.3 gave last ``ア゙イ``, first ``タロウ``. A dakuten or handakuten after a kana that has no precomposed voiced form (``ア゙``, ``ン゙``), or the spacing ``゛`` and ``゜``, was read as hiragana, which made the name Japanese by script and turned it around; the mark now belongs to the kana it follows. It flipped the rest of the name too: ``マイケル ア゙イ`` gives first ``マイケル`` where it gave last ``マイケル``. Hiragana names and kanji-and-kana names with such a mark read as before. See the ``W4`` entry of ``docs/design/decisions.md`` (closes #596) - - **Fix a declared suffix delimiter being taken into a joined name part instead of separating suffixes.** ``HumanName("Smith, John, PhD - and MD", suffix_delimiter=" - ").suffix`` is ``PhD, and MD``, where 2.0 through 2.3 gave ``PhD - and MD``; ``Smith, John, Puig - y Soler`` gives ``Puig, y Soler`` where they gave ``Puig - y Soler``. Both are 1.4.0's answers again. A delimiter declared through ``suffix_delimiter`` or ``Policy(extra_suffix_delimiters=...)`` now separates a trailing suffix part as a comma typed in its place would, giving the same fields: a connective beside it never joins across it, a maiden clause ends at it, and the delimiter itself is dropped, where 2.0 through 2.3 dropped only a delimiter standing alone and kept one a connective had joined. **A maiden marker with the delimiter straight before or after it no longer takes a maiden name**, as it takes none after a typed comma: ``Smith, John, MD - née Jones Smith`` gives suffix ``MD, née Jones Smith`` and no maiden, where 2.0 through 2.3 gave maiden ``Jones Smith``. Only the parts after a suffix comma are affected -- the credentials of ``Name, PhD`` or of ``Family, Given, PhD``; in the name before the first comma, and in the given-name part of ``Family, Given``, the delimiter is still a word, as in 1.4.0. The default policy declares no delimiter, so nothing changes without one. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #549) + - **Fix a declared suffix delimiter being taken into a joined name part instead of separating suffixes.** ``HumanName("Smith, John, PhD - and MD", suffix_delimiter=" - ").suffix`` is ``PhD, and MD``, where 2.0 through 2.3 gave ``PhD - and MD``; ``Smith, John, Puig - y Soler`` gives ``Puig, y Soler`` where they gave ``Puig - y Soler``. Both are 1.4.0's answers again. A delimiter declared through ``suffix_delimiter`` or ``Policy(extra_suffix_delimiters=...)`` now separates a trailing suffix part as a comma typed in its place would, giving the same fields: a connective beside it never joins across it, a maiden clause ends at it, and the delimiter itself is dropped, where 2.0 through 2.3 dropped only a delimiter standing alone and kept one a connective had joined. A maiden marker in a part the delimiter separates takes no maiden name, as no marker after a suffix comma does (see the maiden-marker entry above): ``Smith, John, MD - née Jones Smith`` gives suffix ``MD, née Jones Smith`` and no maiden, where 2.0 through 2.3 gave maiden ``Jones Smith``. Only the parts after a suffix comma are affected -- the credentials of ``Name, PhD`` or of ``Family, Given, PhD``; in the name before the first comma, and in the given-name part of ``Family, Given``, the delimiter is still a word, as in 1.4.0. The default policy declares no delimiter, so nothing changes without one. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #549) - **Fix halfwidth corner brackets not being read as a nickname.** ``HumanName("山田 「タロー」 タロウ")`` gives nickname ``タロー``, last ``山田``, first ``タロウ``, where every release gave middle ``「タロー」`` and first ``山田`` (the order moves with the halfwidth katakana change above; the bracket would otherwise have blocked it). The halfwidth ``「」`` are the corner brackets of legacy JIS X 0201 data, the same punctuation as ``「」``, and are now a default nickname pair in both APIs: ``DEFAULT_NICKNAME_DELIMITERS`` gains ``("「", "」")`` and the 1.x ``nickname_delimiters`` gains the key ``halfwidth_corner_brackets``. They are not limited to Japanese text: ``John 「Jack」 Smith`` gives nickname ``Jack`` where it gave middle ``「Jack」``. A ``Constants`` restored from a pickle keeps the keys it was saved with, as it did when 2.0 added the other typographic pairs. See the ``N1`` entry of ``docs/design/decisions.md`` (closes #597) diff --git a/nameparser/_parser.py b/nameparser/_parser.py index 5405f337..faf8e85d 100644 --- a/nameparser/_parser.py +++ b/nameparser/_parser.py @@ -182,8 +182,12 @@ def revise(self, name: ParsedName, **fields: str) -> ParsedName: ``revise(n, suffix="MD PhD")`` is one entry and a name's rendered suffix revises back to itself wherever the value's words read as the whole name read them -- the honorific peel - above is the one corpus exception of 368 suffix-bearing names, - 2026-09-06 (#511). A delimiter the policy names through + above was the one corpus exception of 368 suffix-bearing names, + 2026-09-06 (#511). A second came with #601/#602: a suffix + holding a maiden marker the whole name left a word, whose + value read on its own takes the clause and drops the marker + ('Jane Doe PhD nee Smith' renders 'PhD nee Smith', which + revises to 'PhD Smith'), three corpus names on 2026-10-04. A delimiter the policy names through ``extra_suffix_delimiters`` parts a value only where the value's own words, read as a name, give it a tail segment for the core to be dropped on; a run of post-nominals has none, diff --git a/nameparser/_pipeline/_assign.py b/nameparser/_pipeline/_assign.py index e7d3ed66..beaea73b 100644 --- a/nameparser/_pipeline/_assign.py +++ b/nameparser/_pipeline/_assign.py @@ -76,9 +76,10 @@ ) from nameparser._pipeline._pieces import ( anchor_in_reach, credential_at_the_given_slot, given_slot_anchors, - is_lone_never_given_particle, is_suffix_piece, leading_titles, - listed_lean, peel_walk, - segment_suffix_reading, tail_reading, trailing_titles, + _NOT_A_RUN_START, has_name_content, is_lone_never_given_particle, + is_suffix_piece, is_title_piece, is_wholly_particle, leading_titles, + listed_lean, peel_walk, segment_suffix_reading, + starts_a_credential_run, tail_reading, trailing_titles, ) from nameparser._pipeline._state import ( AMBIGUOUS_ACRONYM_TAG, ParseState, PendingAmbiguity, Structure, @@ -93,6 +94,17 @@ def _set_roles(tokens: list[WorkToken], piece: tuple[int, ...], tokens[i] = copy_with(tokens[i], role=role) +def _absorbed(piece: Sequence[int], + tokens: Sequence[WorkToken]) -> PendingAmbiguity: + """rules.md#S2's report for a name word #602's run absorbs, one + spelling for the no-comma and the given-part run.""" + text = " ".join(tokens[i].text for i in piece) + return PendingAmbiguity( + AmbiguityKind.SUFFIX_OR_NAME, + f"{text!r} follows a credential, so it reads as part of the " + f"suffix run; it may be a name word", tuple(piece)) + + #: Tags that say the word's own reading was claimed before position #: could speak, so O5's convention decided nothing. Built from M4's #: `_NEVER_FLIPPED` pair rather than respelling it: the two sets answer @@ -311,6 +323,13 @@ def _assign_main(seg_idx: int, state: ParseState, state.one_case) for piece_idx in titled_tail: _set_roles(tokens, pieces[piece_idx], Role.TITLE) + # rules.md#S2: "except a title word, which reads as a title" -- the + # credential run (#602) is placed here, so a name word it absorbed + # is reported here too + for piece_idx in peeled.run_titles: + _set_roles(tokens, pieces[piece_idx], Role.TITLE) + for piece in peeled.absorbed: + ambiguities.append(_absorbed(piece, tokens)) if peeled.numeral is not None: # a trailing single letter is a name part unless it happens # to be a roman numeral -- and V/X/I are ordinary middle @@ -322,6 +341,9 @@ def _assign_main(seg_idx: int, state: ParseState, f"letter there would be a middle initial", peeled.numeral)) name_pieces, suffix_pieces = rest[:peeled.names], rest[peeled.names:] + if peeled.run_titles: + suffix_pieces = [q for q in suffix_pieces + if q not in peeled.run_titles] if peeled.names == 0: # everything suffix-shaped after titles: first one is the name name_pieces, suffix_pieces = suffix_pieces[:1], suffix_pieces[1:] @@ -899,13 +921,8 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: # not merely a stray report # (decisions.md#S2, 2026-09-18). # - # A FUNCTION since #533, not two conditions - # written to match: the maiden walk's - # second check asks this same question of - # the name a take would leave, and the - # drift would have been silent -- each - # site's own tests would have gone on - # passing (mechanisms.md + # A FUNCTION since #533 rather than a + # condition written here (mechanisms.md # #ONE-PREDICATE-PER-QUESTION). The call # costs one frame PER MEMBER asked at this # slot, not one per name: against @@ -1047,11 +1064,85 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: # memo, not a second spelling of the question, and it # cannot have gone stale because nothing the chain does # moved the end of the name. + # rules.md#S2: "or after the given word in the part after a + # family comma" -- the given part's own credential run + # (#602), the comma having fixed the family. The tag tests are inline ahead of the + # predicate, as in `_pieces.credential_run`, so a name word + # pays no frame: this loop runs on every family-comma name. + sticky_from = len(pieces) + for m in range(n + 1, len(pieces)): + tags = tokens[pieces[m][0]].tags + if (m not in titled_idx and "vocab:suffix" in tags + and tags.isdisjoint(_NOT_A_RUN_START) + and starts_a_credential_run(pieces[m], ptags[m], + tokens)): + sticky_from = m + break + # The particle tail P6 will attach is not the run's to take + # (rules.md#P6: "a particle ending the name attaches to that + # family name", looking past the post-nominals behind it): + # the wholly-particle pieces ending the part, behind any run + # words, found as P6 finds them. They are left to the walk + # below and reach P6 with the role they had before #602. + # Absorbing them reported a suffix reading P6 then overrode, + # and P6 reported the override as a declined post-nominal + # ('Smith, John PhD de', 'Smith, John PhD de Jr.'). A + # particle P6 will NOT attach -- one with a credential + # behind it and another particle past that, 'Smith, John + # PhD de PhD van' -- stays in the run, as does a lone member + # of the ambiguous credential class (`do`): read as the + # credential, it is the word P6's #531 exception keeps out + # of the attachment, so the run and P6 agree on it. + # `pieces[p6_lo:p6_hi]` is that tail: back past the run + # words behind it, then over the particles, stopping at a + # class member + p6_lo = p6_hi = len(pieces) + if sticky_from < len(pieces): + while (p6_hi > sticky_from + and not is_wholly_particle(pieces[p6_hi - 1], + tokens)): + p6_hi -= 1 + p6_lo = p6_hi + while (p6_lo > sticky_from + and is_wholly_particle(pieces[p6_lo - 1], tokens) + and not (len(pieces[p6_lo - 1]) == 1 + and not tokens[pieces[p6_lo - 1][0]].tags + .isdisjoint(_AMBIGUOUS_CREDENTIAL_TAGS))): + p6_lo -= 1 for m in range(n + 1, len(pieces)): if m in titled_idx: continue suffix_here = (reads_as_a_suffix(m, titled_idx) if titled_idx else m not in walkable) + if m >= sticky_from and not p6_lo <= m < p6_hi: + # inside the run: a title word reads as a title, + # every other word as a suffix, and a word the walk + # would have kept as a name is reported, here where + # the run absorbs it + if (m > sticky_from + and not is_suffix_piece(pieces[m], ptags[m], + tokens) + and is_title_piece(pieces[m], ptags[m], + tokens)): + _set_roles(tokens, pieces[m], Role.TITLE) + continue + piece = pieces[m] + if (len(piece) == 1 + and not tokens[piece[0]].tags.isdisjoint( + _AMBIGUOUS_CREDENTIAL_TAGS)): + # a class member inside the run is still the + # fork #531's report names, now read as the + # credential, as #544's anchor reads one + ambiguities.append(PendingAmbiguity( + AmbiguityKind.SUFFIX_OR_NAME, + f"{tokens[piece[0]].text!r} ending the given " + f"part is also an ordinary name word; read " + f"as a credential", (piece[0],))) + elif not suffix_here and has_name_content(piece, + tokens): + ambiguities.append(_absorbed(piece, tokens)) + _set_roles(tokens, piece, Role.SUFFIX) + continue _set_roles(tokens, pieces[m], Role.SUFFIX if suffix_here else Role.MIDDLE) # #531's report, and the FIFTH SUFFIX_OR_NAME site in diff --git a/nameparser/_pipeline/_group.py b/nameparser/_pipeline/_group.py index 31535fe9..bdaa01b6 100644 --- a/nameparser/_pipeline/_group.py +++ b/nameparser/_pipeline/_group.py @@ -5,7 +5,9 @@ the only stage after tokenize that reads it). Produces: pieces + piece_tags per segment (runs of token indices -- tokens are NEVER joined into strings: the anti-#100 invariant); maiden -tail tokens get role=MAIDEN; marker tokens land in dropped. +tail tokens get role=MAIDEN, and the trailing run a maiden take gives +up gets its SUFFIX and TITLE roles here (#601), so no join can reach +it; marker tokens land in dropped. Reads: token tags (from classify), Lexicon.given_name_titles (the P5 licence, #369) and Policy.extra_suffix_delimiters, whose delimiter-core tokens part a tail segment as a comma would and are @@ -41,22 +43,23 @@ import bisect from collections.abc import Iterable, Sequence, Set from enum import IntEnum -from typing import Literal, assert_never +from typing import assert_never from nameparser._lexicon import _run_addresses_by_given from nameparser._pipeline._pieces import ( - anchor_in_reach, credential_at_the_given_slot, given_slot_anchors, is_leading_title, is_suffix_piece, is_title_piece, - is_trailing_title_word, - Peel, leading_titles, peel_trailing, peel_walk, tail_reading, + is_trailing_title_word, starts_a_credential_run, + leading_titles, peel_walk, tail_reading, trailing_start, trailing_start_past_titles, ) from nameparser._pipeline._state import ( - AMBIGUOUS_ACRONYM_TAG, ParseState, PendingAmbiguity, Structure, + ParseState, PendingAmbiguity, Structure, WorkToken, _AMBIGUOUS_CREDENTIAL_TAGS, copy_with, ) from nameparser._pipeline._vocab import D, PH -from nameparser._pipeline._vocab import delimiter_cores +from nameparser._pipeline._vocab import ( + delimiter_cores, is_trailing_numeral_suffix, +) from nameparser._types import AmbiguityKind, Role # the credential-pair regexes live in _vocab, whose own @@ -67,48 +70,41 @@ Piece = list[int] #: What the marker pass took out of a segment: the marker's TOKEN -#: indices (more than one where the entry is a phrase, 'z domu') and -#: the maiden-name pieces, themselves token indices, to take -#: Role.MAIDEN. `_group_segment` produces it. -MaidenTake = tuple[Piece, list[Piece]] +#: indices (more than one where the entry is a phrase, 'z domu'), and +#: every other piece it took with the role that piece takes -- the +#: maiden name's Role.MAIDEN, and the trailing run's SUFFIX or TITLE +#: (#601). `_group_segment` produces it. +MaidenTake = tuple[Piece, list[tuple[Piece, Role]]] #: What `_maiden_take` answers with, one index space further out: PIECE -#: indices into its own `pieces` argument -- the marker's pieces, then -#: the maiden name's. `_group_segment` resolves them to the token -#: indices MaidenTake declares, which is the only reason the two are -#: different types rather than one name used twice. -MaidenIndices = tuple[list[int], list[int]] - - -class TailReader(IntEnum): - """Which rule reads the words the maiden walk would leave standing - at the end of this segment -- the reader every stop's release - check (`_release_reads_off`) has to ask, since a stop is only right - where that reader takes what the stop gives up (rules.md#M2, #533, - #535). It also decides whether the walk reads the trailing title - chain at all: TRAILING and GIVEN_SLOT read the end of the clause - through `tail_reading`, NONE through the peel alone. - - NONE is a statement and not a default: before a family comma the - words are the family the comma already named, and a tail segment - is read as credentials whole, so no trailing rule is consulted - there and the clause keeps what it has -- a stop would hand a word - to `family` rather than to `suffix`. +#: indices into its own `pieces` argument. `_group_segment` resolves +#: them to the token indices MaidenTake declares, which is the only +#: reason the two are different types rather than one name used twice. +MaidenIndices = tuple[list[int], list[tuple[int, Role]]] + + +class ClauseSite(IntEnum): + """Where a maiden clause stands, which decides whether a marker + counts there at all and which rule reads the end of the part + (rules.md#M2, #601). + + A marker counts only in the part holding the family name, so + GIVEN_SLOT and TAIL are statements, not defaults: a marker there is + an ordinary word. FAMILY_PART is the part before a family comma, + where a marker counts but no trailing rule reads the words after + the clause. A CLOSED set: `_maiden_take` dispatches on it exhaustively, ONCE - (`typing.assert_never`), and every later site there branches on - the narrowed reading that dispatch binds; `_release_reads_off` - dispatches on the two readers that reach it the same way. So a - fourth member is a type error until it is given a reading. group() is the one - place (structure, segment index) is mapped onto it, and - tests/v2/pipeline/test_group.py's - `test_the_reader_is_pinned_to_the_structure_it_is_read_from` - pins that mapping -- by watching the call, not by restating it -- - along with this enum's size. Change either and that test says so. + (`typing.assert_never`). group() is the one place (structure, + segment index) is mapped onto it, and tests/v2/pipeline/ + test_group.py's `test_the_site_is_pinned_to_the_structure_it_is_ + read_from` pins that mapping -- by watching the call, not by + restating it -- along with this enum's size. """ - NONE = 0 # FAMILY_COMMA segment 0, and every tail segment - TRAILING = 1 # S2 peel + H5 chain: NO_COMMA, SUFFIX_COMMA seg 0 - GIVEN_SLOT = 2 # #531's reading + H5 chain: FAMILY_COMMA seg 1 + FAMILY_PART = 0 # FAMILY_COMMA segment 0: a marker counts, nothing trails + TRAILING = 1 # NO_COMMA, SUFFIX_COMMA segment 0: S2 peel + H5 chain + GIVEN_SLOT = 2 # FAMILY_COMMA segment 1: a marker is an ordinary word + TAIL = 3 # every tail segment: a marker is an ordinary word class BoundJoin(IntEnum): @@ -127,11 +123,10 @@ class BoundJoin(IntEnum): # rules.md#S2: "a trailing word of the suffix vocabulary reads as a # suffix" -- group does not decide that; it stops before whatever -# trailing_start says the run is, so the chain and the maiden walk -# end where assign's peel begins (#424). The maiden walk reads the -# title-aware `tail_reading` instead wherever a trailing rule reads -# the clause (#535), so there it ends where assign's peel and title -# chain together begin. +# trailing_start says the run is, so the chain ends where assign's +# peel begins (#424). The maiden take reads the title-aware +# `tail_reading` over the clause-free view instead (#601), so the +# clause ends where assign's peel and title chain together begin. # rules.md#P2: "a particle joins the words after it into one name # part, the join running until the next particle starts a group of # its own, a trailing suffix begins" -- and on to the maiden marker @@ -146,12 +141,9 @@ def _is_prefix_piece(piece: Sequence[int], ptags: Set[str], return len(piece) == 1 and "particle" in tokens[piece[0]].tags -# rules.md#M2: "a recognized maiden marker standing after at least one -# name word takes the words after it" -- up to any suffix word, the -# trailing numeral or credential assign reads as the suffix, or (where -# a trailing rule reads the clause) a trailing title H5's chain takes, -# as the maiden name, the marker itself dropped (history: -# decisions.md#M2) +# rules.md#M2: "A marker that counts takes the words after it as the +# maiden name, and is itself dropped" -- where a marker counts, and up +# to what, are `_maiden_take`'s (history: decisions.md#M2) # # A marker piece is a LONE marker -- M2's own "standing as a word of # its own". The consumer runs before every join but the Ph. D. merge @@ -234,816 +226,167 @@ def _marker_run_pieces(pieces: Sequence[Sequence[int]], tokens[pieces[k][0]].tags for k in range(m + 1, len(pieces))) -# rules.md#M2: "a word the clause gives up reads as a post-nominal or -# the clause keeps it" -- the two halves of that, asked of the VIEW -# the take would leave. Split out of `_maiden_take` because each -# models a LATER decision about the released word, and the stop is -# only right where that decision goes the way the release assumed -# (#533 review). -def _a_name_word_ahead(view: Sequence[Sequence[int]], - view_tags: Sequence[Set[str]], - tokens: Sequence[WorkToken], - at: int) -> bool: - """Whether the view holds a name word BEFORE the member at `at`. - - #531's reading is the given part's own trailing slot, and a slot - needs a part: the take removes the clause, so where every piece - ahead of the member is a title or a suffix the member is the only - name word left and there is no trailing slot for it to end. - Asked of the view rather than of the segment because the segment - still has the maiden name in it -- that is exactly the word the - release takes away ('Doe, Prof. nee Smith A.B.' left 'Prof. A.B.', - whose A.B. is the GIVEN name and no credential at all).""" - return any(not is_title_piece(view[q], view_tags[q], tokens) - and not is_suffix_piece(view[q], view_tags[q], tokens) - for q in range(at)) - - -def _join_takes_the_member(view: Sequence[Sequence[int]], - view_tags: Sequence[Set[str]], - tokens: Sequence[WorkToken], - at: int, reader: TailReader) -> bool: - """Whether a join BELOW this pass would absorb the member at `at`. - - A released word only reads as the credential while it is still the - lone piece assign's peel looks at; inside a joined piece it is a - word of the name the join built, and the release has moved it from - the birth name into the CURRENT one -- #424's class of failure, - now from the other side ('Berg, Jane van der nee Smith DO' read - family 'van der DO Berg'). - - Three joins can reach it, and each is modelled by the shape it - needs rather than by running it: - - * P2's chain, where the member is particle vocabulary standing - behind a particle piece, beside another released one ('Jane Doe - nee Smith DO DO' read family 'DO DO'), or with a title behind - it, the chain running on over a trailing title (rules.md#H5's - Accepted 'John van der Berg Prof.'; 'Jane Doe nee Smith MA do - Prof.'). - * P6's attachment, after a family comma, where the member is a - TITLE that is also particle vocabulary: P6 attaches a particle - trailing the given part to the family ('Doe, Jane nee Smith St.' - would read family 'St. Doe'). Titles only -- a credential that - is also a particle ('DO') is the given slot's own lean, which - the credential stop has already asked (#533). - * P5's bound-given join, where the member is the word after the - bound one ('Berg, abdul nee Jones MA' read given 'abdul MA'). - Asked for the GIVEN_SLOT reader alone: that join is - `BoundJoin.LENIENT` only after a family comma, and before one - the reserve is `BoundJoin.STRICT`, which reads the same peel and - chain the release was just asked of and so never joins a span - that reading covered (rules.md#P5). - - All three over-decline rather than predict: a shape that only - MIGHT join keeps its word in the maiden name, which is the - conservative direction M2's invariant asks for.""" - member = tokens[view[at][0]] - # P6's half (see the docstring): a released title that is also a - # particle, after a family comma. - if (reader is TailReader.GIVEN_SLOT and "particle" in member.tags - and is_title_piece(view[at], view_tags[at], tokens)): +#: the tags `_clause_tail_word` admits a lone word on +_CLAUSE_TAIL_TAGS = _AMBIGUOUS_CREDENTIAL_TAGS | {"vocab:suffix"} + + +# rules.md#M2: "It takes them up to the trailing run of post-nominals +# and titles that the end of the name reads as if the clause were not +# written" (#601). A word that may belong to that run -- the view +# decides which of them it actually takes. +def _clause_tail_word(piece: Sequence[int], ptags: Set[str], + tokens: Sequence[WorkToken]) -> bool: + """Suffix vocabulary, a title word H5's chain takes from the end + (period-marked: a BARE title word ending a name is a name word, and + pulled into the view it would only inflate the count of words to + spare), or a member of the ambiguous class: a word the end of the + clause-free name might read as a post-nominal or a title. Being one + is no answer -- `_maiden_take` reads the view to find out which of + them the trailing rule actually takes. A bare title inside #602's + credential run reaches the view another way, through the run's own + start.""" + if ("suffix" in ptags or is_suffix_piece(piece, ptags, tokens) + or is_trailing_title_word(piece, ptags, tokens)): return True - if "particle" in member.tags: - if at and _is_prefix_piece(view[at - 1], view_tags[at - 1], tokens): - return True - # a title behind it: the chain runs on over it (docstring) - for q in range(at + 1, len(view)): - if len(view[q]) == 1 and "particle" in tokens[view[q][0]].tags: - return True - if is_title_piece(view[q], view_tags[q], tokens): - return True - # rules.md#P5: "a word the peel reads as a suffix unjoined must - # read so joined, or the join declines" -- why the STRICT reserve - # needs no model here and only the GIVEN_SLOT reader asks. P5 - # joins the first non-title piece to the one after it, so the - # member is at risk exactly where it IS the one after it. The tag - # pair is tested before `leading_titles` is asked, which keeps the - # ordinary credential release from paying that frame. - return (reader is TailReader.GIVEN_SLOT - and at > 0 and len(view[at - 1]) == 1 - and "vocab:bound-given" in tokens[view[at - 1][0]].tags - and leading_titles(view, view_tags, tokens) == at - 1) - - -# rules.md#M2: "a word the clause gives up reads as a post-nominal or -# the clause keeps it" -- asked once, here, for every stop the walk -# makes: the numeral, the credential and the title (#535). A stop -# gives up the whole SPAN behind the word it stops at, so the question -# is asked of the span and not of the word: 'DOE NEE SMITH PROF. MA' -# stopped at the title and handed the MA behind it to the family, -# where the name left standing ('DOE PROF. MA') reads MA as a name. -def _release_reads_off(view: Sequence[Sequence[int]], - view_tags: Sequence[Set[str]], - tokens: Sequence[WorkToken], - start: int, end: int, at: int, - reader: Literal[TailReader.TRAILING, - TailReader.GIVEN_SLOT], - one_case: bool | None, - reading: tuple[list[int], tuple[int, ...], Peel] - | None = None, - *, tail_follows: bool) -> bool: - """Whether the name the take would leave (`view`) reads every - piece in `start`..`end` as a title or a suffix, and no join below - this pass absorbs any piece from `at` on. - - `at` is the first released piece. `start` is where coverage is - asked from, which is `at` itself except where the caller has - already asked the first piece its own question (the numeral fork - and the given-slot credential test each read their word in a way - the span test does not, and pass `at + 1`). `end` is where it - stops: `len(view)`, except for the title stop when a numeral or - credential stop stands behind it -- that stop has already had its - own span asked, in the way ITS word reads: the span check does not - ask the given slot's lenient numeral itself -- the numeral fork has - ('Doe, Jane nee Smith Prof. V' gives the V up there) -- though the - GIVEN_SLOT title chain below counts it as the first pass's suffix, - so the title in front of it is asked only of itself. - - Each reader is asked the way it READS, because the release is only - right where that reader places the span: - - * TRAILING -- the no-comma path and the part before a suffix - comma -- is assign's own reading, `tail_reading` over the view, - the S2 peel and the H5 chain to their fixed point. A piece is - covered where that reading took it as a title or peeled it as a - suffix, or where group flagged it a credential outright. - * GIVEN_SLOT -- after a family comma -- has words to spare by - construction, so the question is the writing alone: a name word - must stand ahead (`_a_name_word_ahead`), and every piece of the - span must be a suffix piece, a class member #531's reading takes - as the credential, or a word H5's chain takes as a title. - - Then the joins. A title in the span with a particle ahead of it is - taken by P2's chain, which runs on over a trailing title - (rules.md#H5's Accepted 'John van der Berg Prof.'), and every piece - of the span is asked `_join_takes_the_member`. Both decline rather - than predict, which is the conservative direction M2 asks for. - - NONE never reaches this, and the annotation says so: no trailing - rule reads those words, so no stop is made for this to check. - - `reading` is the TRAILING reader's `tail_reading` of this same - view, where the caller has one in hand (the numeral fork reads it - first), so the view is read once rather than twice. - """ - if reader is TailReader.GIVEN_SLOT: - if not _a_name_word_ahead(view, view_tags, tokens, at): - return False - # The given part's title chain is read the way assign reads it - # after a family comma: over the pieces its FIRST suffix pass - # leaves, from the end. That pass takes a class member as the - # credential only where every piece behind it is taken too, so - # a member with a title behind it is still a name word there - # and the chain stops at it -- 'Doe, Jane Dr. MA Prof.' reads - # middle 'Dr.', and a clause that gave 'Dr. MA Prof.' up put - # 'Dr.' in the middle name (#535). `chain_ok[q]` says whether - # the chain, walking from the end, is still running at `q`. - # The one suffix these passes read which `is_suffix_piece` does - # not is the lenient numeral (#144): a suffix word the initial - # veto refuses, read as a suffix only where no comma part - # follows the given one (`tail_follows`, assign's own - # two-segment condition) and only where it is the given part's - # LAST word -- the literal last piece in the first pass, and - # the last one standing once the chain has taken the titles - # behind it in the second ('Doe, Jane i V Prof.' reads suffix - # 'i V', title 'Prof.', as 'Doe, Jane nee Smith i V Prof.' must - # be able to give them up). - last = len(view) - 1 - chain_ok = [False] * len(view) - members_ok = running = True - # #544: the anchors `credential_at_the_given_slot` may ask for, - # over the view as the take would leave it -- the reading - # assign's given slot makes of the same pieces, from past the - # leading title run. One forward pass for both loops below, - # run only the first time a member's writing leaves the - # question open; each call site hands over a lambda, so no - # frame is spent building the question either. - anchors: list[bool] | None = None - - def anchored_at(q: int) -> bool: - nonlocal anchors - if anchors is None: - # the reach test first, and only before the pass - # exists (`_pieces.anchor_in_reach`) - if not anchor_in_reach(range(q - 1, -1, -1), view, - view_tags, tokens): - return False - anchors = given_slot_anchors( - view, view_tags, tokens, - leading_titles(view, view_tags, tokens)) - return anchors[q] - for q in range(last, -1, -1): - piece = view[q] - if (is_suffix_piece(piece, view_tags[q], tokens) - or (q == last and not tail_follows - and len(piece) == 1 - and "vocab:suffix" in tokens[piece[0]].tags)): - chain_ok[q] = running - continue - if (members_ok and len(piece) == 1 - and AMBIGUOUS_ACRONYM_TAG in tokens[piece[0]].tags - and credential_at_the_given_slot( - tokens[piece[0]], one_case, - lambda: anchored_at(q))): - chain_ok[q] = running - continue - members_ok = False - if not is_trailing_title_word(piece, view_tags[q], tokens): - running = False - chain_ok[q] = running - titles_behind = [False] * (len(view) + 1) - titles_behind[len(view)] = True - for q in range(last, -1, -1): - titles_behind[q] = (titles_behind[q + 1] and chain_ok[q] - and is_trailing_title_word( - view[q], view_tags[q], tokens)) - for q in range(start, end): - piece = view[q] - if is_suffix_piece(piece, view_tags[q], tokens): - continue - if (not tail_follows and titles_behind[q + 1] - and len(piece) == 1 - and "vocab:suffix" in tokens[piece[0]].tags): - continue - if is_trailing_title_word(piece, view_tags[q], tokens): - if chain_ok[q]: - continue - return False - if (len(piece) == 1 - and AMBIGUOUS_ACRONYM_TAG in tokens[piece[0]].tags - and credential_at_the_given_slot( - tokens[piece[0]], one_case, - lambda: anchored_at(q))): - continue - return False - elif reader is TailReader.TRAILING: - rest, chained, peeled = reading or tail_reading( - peel_walk(leading_titles(view, view_tags, tokens), view_tags), - view, view_tags, tokens, one_case) - covered = set(chained) - covered.update(rest[peeled.names:]) - if not all(q in covered or "suffix" in view_tags[q] - for q in range(start, end)): - return False - else: - assert_never(reader) - if (any(is_title_piece(view[q], view_tags[q], tokens) - for q in range(at, len(view))) - and any(_is_prefix_piece(view[q], view_tags[q], tokens) - for q in range(at))): - return False - return not any(_join_takes_the_member(view, view_tags, tokens, q, - reader) - for q in range(at, len(view))) - - -# rules.md#M2: "a link inside the birth name does not end it" -- the -# one shape the walk below steps over rather than stopping at. -# rules.md#P3: "A connective that is also generational vocabulary -# joins only where a name word stands on each side of it" is the -# reason, and the class test is that rule's own. The take runs BEFORE -# every join, so the link is still a piece of its own here and the -# question is asked of the pieces as classify left them -- the same -# inputs `_group_segment`'s `frozen` loop gives -# `_between_name_words`, which is why this calls that predicate -# rather than restating the class -# (mechanisms.md#ONE-PREDICATE-PER-QUESTION). -def _link_joins_inside_the_clause(k: int, lo: int, hi: int, - pieces: Sequence[Sequence[int]], - ptags: Sequence[Set[str]], - tokens: Sequence[WorkToken], - beside: list[_Beside]) -> bool: - """Whether the suffix piece at `k` is a connective PLACED TO JOIN - between two name words of the clause `lo`..`hi`. - - The caller asks `is_suffix_piece` first and consults this only - where the answer was yes, so the generational half of "also - generational vocabulary" is settled and the connective half is - what is left to ask. A lone link is therefore no link at all - ('Jane Doe nee Puig i' keeps maiden 'Puig' and suffix 'i'), and - neither is one standing before the generation or the credential a - clause ends with ('... nee Puig i III', '... i MA'): `hi` is where - assign's peel begins, so those stand at or past it and - `_between_name_words` refuses them by bound. - - Defined here, beside its one caller, and forward-referencing the - two predicates it is built out of: `_is_conj_piece` and - `_between_name_words` are the JOIN's, further down this module, and - moving them up to meet this would say they belonged to the clause. - - `beside` is the caller's memo cell, filled on the first CONNECTIVE - piece the walk reaches rather than on the first suffix piece, so - a clause ending at an ordinary credential never builds it at all - and a clause holding a RUN of links builds it once ('Jane Doe nee - Puig i i i ... Soler'). Filled here rather than at the call site - so the laziness costs no frame of its own, and one fill serves - the whole walk because that walk mutates neither `pieces` nor - `ptags` (#397 second review, the run fix).""" - if not _is_conj_piece(pieces[k], ptags[k], tokens): - return False - if not beside: - beside.append(_run_neighbours(pieces, ptags, tokens)) - return _between_name_words(k, lo, hi, pieces, ptags, tokens, beside[0]) + return (len(piece) == 1 + and not tokens[piece[0]].tags.isdisjoint(_CLAUSE_TAIL_TAGS)) + def _maiden_take(pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], tokens: Sequence[WorkToken], one_case: bool | None, - reader: TailReader, + site: ClauseSite, ambiguities: list[PendingAmbiguity], - *, tail_follows: bool, ) -> MaidenIndices | None: """The PIECE indices the marker pass removes, split the way - MaidenIndices declares them: the MARKER's pieces (one, or several - for a phrase entry like 'z domu') and the maiden name's. None when - the pass declines. Piece indices into `pieces`, not the TOKEN - indices of the MaidenTake `_group_segment` builds out of them -- - the two are different index spaces and both are named types so a - reader never has to guess which one an annotation means. - - Split here rather than by the caller because `run` is known here - and nowhere else -- it comes off the continuation tags, which are - read to find the walk's starting point in the first place. - - Computed before any join (the Ph. D. merge aside), so "up to any - trailing suffix" means the first suffix WORD after the marker: a - connective beside that suffix cannot un-suffix it first. 'Jane - Smith née Jr y Jones' declines where the joined reading took 'Jr y - Jones' as the maiden name, and 'Jane Smith née Jones Jr y Smith' - takes only 'Jones'. The one reading the order costs; M2's Accepted - row and decisions.md#M2 (#420) record it. - - "Up to any trailing suffix" also means up to a trailing CREDENTIAL - since #533, where the rule that reads the name left standing reads - the word as one -- the TRAILING peel with no comma, the given - part's own slot after a family comma, and nobody before that comma - or past a second one, which is what `reader` says. A member - standing alone after the marker is never given up: the marker - announces a name, and the rule gives a word up only where a maiden - name is left standing. Nor is one given up to a reader that will - not be there to read it, or to a JOIN that runs before the reader - does: rules.md#M2's invariant is that a released word ends the - parse suffix-roled, and the shared release check - (`_release_reads_off`) is what makes the stop conservative enough - to hold it (#533 review). - - And up to a trailing TITLE since #535, where the reader has a - trailing rule: the walk reads the end of the name through H5's - chain as assign does, so the title ends the clause and a - credential or numeral in front of it gets the stops it gets with - the title absent. All three stops -- the numeral, the credential - and the title -- ask one question of the whole span they give up - (`_release_reads_off`), and so does a link the walk stops at. Only - the credential and title stops spare the first word after the - marker. The numeral stop does not: it reads FROM the marker by - design, so a numeral standing straight after it is not held to - the first-word floor and may decline the clause -- 'Jane Smith née - V' (rules.md#M2) stays a marker with nothing behind it and the - name has no maiden clause at all -- while 'Dr. nee V' keeps maiden - 'V'. Those are examples, not a rule over every shape: the numeral - fork and the reader decide it, and 'Doe, J. nee V' keeps maiden - 'V' though 'Doe, J. V' reads suffix 'V'. - - A tail segment's delimiter cores never reach this walk: group() - cuts the segment at them first and hands each part over on its - own, as the comma twin's segments would be (#549). So a core is - neither a word the marker can take ('PhD née - Jones' reads as - 'PhD née, Jones') nor the name word M2 needs ahead of the marker. + MaidenIndices declares them: the marker's pieces (one, or several + for a phrase entry like 'z domu'), then the maiden name's and the + trailing run's, each with the role it takes. None where the marker is + an ordinary word. Piece indices into `pieces`, not the TOKEN + indices `_group_segment` builds out of them. + + Three questions, in order (rules.md#M2, #601): + + * WHERE A MARKER COUNTS. Only in the part holding the family name + -- never in the given part after a family comma, never in a tail + part -- and only with a name word straight before it: not a word + of the leading title run (H3's give-back included, so 'Lord + Chancellor née Jones' counts; a lone title is a title), not a + suffix word unless it opens the name (S2), not a connective. A + particle counts: with nothing after it to attach to it is no + particle there ('Mai Le née Nguyen'). + * WHAT IT TAKES. Every word after it up to the trailing run the + name reads as if the clause were not written, and always the + first word after it. The run is read over the VIEW -- the head + plus the vocabulary-shaped tail, and anything from the first + credential that starts a run (S2, #602) -- by `tail_reading`, so + the count of words to spare is the clause-free name's. Before a + family comma no trailing rule reads the part, and the family + part takes a suffix word from anywhere in it: the clause stops + at the first one and hands the rest back unbound. + * THE RUN IS BOUND. Where a trailing rule reads the part, the take + consumes the run and returns the role the reading gave each + piece, so no join below this pass can reach into it. + + Computed before any join (the Ph. D. merge aside), as the marker + pass always has been (#420): the joins see a name without the + clause in it. """ + if site is ClauseSite.GIVEN_SLOT or site is ClauseSite.TAIL: + return None m = next((v for v in range(1, len(pieces)) if _is_maiden_marker_piece(pieces[v], tokens)), None) if m is None: return None # the marker may be a phrase, in which case it is several pieces - run = _marker_run_pieces(pieces, tokens, m) - # "up to any trailing suffix": a suffix WORD anywhere after the - # marker ends the maiden name, and so does the trailing numeral as - # assign will read it, which the suffix-piece test does not see - # (#424): 'John née Jones Smith V' took the V as maiden text. Both - # forks now -- the acronym one since #533, asked the same double - # way. Read from the MARKER, not after it: a - # numeral straight after the marker then has the piece before it - # the fork wants, and 'Jane Smith née V' declines like 'Jane Smith - # née PhD' -- nothing after the marker but a suffix, so the marker - # stays a word -- as 1.4.0 read it. - # - # `one_case` is LIVE at these sites since #533 and was not before: - # the numeral fork is decided before the peel ever reads a lean, - # so the fact reached only the bare-acronym fork, which the - # numeral-only reading discarded. The acronym fork is asked now, - # so the writing decides here as it decides at the trailing slot - # of a name. Measured 2026-09-19 with a runtime wrapper forcing - # this function's `one_case` argument to None, over the population - # and the six policies decisions.md#S2's 2026-09-18 recipe names - # -- 36 of 10,752 parses move (1,792 names), on 6 distinct names - # ('Doe, Jane nee Smith DO', 'Doe, Jane nee Smith Ma', 'Jane Doe - # nee Smith Ma', 'Jane Doe nee Smith Ma JD', 'Jane Doe nee Yo-Yo - # Ma', 'John née Jones Smith MA'). THE PAIR IS THE FINDING: over - # the same corpus WITHOUT this change's own rows it moves 0, so - # the plumbing was live either way and the corpus simply held no - # name that could show it -- the blindness mechanisms.md's corpus - # field note asks to be measured before any "N names move" is - # written down. - # - # The chain-tail measure below (`tail`, and the re-peel after the - # chain) is the opposite, and the 2026-09-18 sweep over the - # pre-change 9,852 says so: dropping it moves 18 of them, on - # 'John van der Berg Ma', 'John de Ma' and 'Freiherr von Berg MA' - # under every one of the six policies. - rest = peel_walk(m, ptags) - # #535: where a trailing rule reads these words, the walk reads the - # end of the name as that rule does -- the S2 peel and the H5 title - # chain to their fixed point (`tail_reading`), so a title behind - # the clause no longer hides the credential or numeral in front of - # it, and the title is itself a stop (below). `rest` comes back - # with the chained pieces SPLICED OUT, which is what keeps - # `rest[-1]` the numeral and `rest[peeled.names]` the first piece - # the peel took. Where no rule reads them (NONE) there is no chain - # to consult and the peel alone stands, as before. - # THE reader dispatch: exhaustive here, once, so a fourth - # `TailReader` member is a type error at this line; every later - # site branches on `reads`, which is the reader where a trailing - # rule reads these words and None where none does. - reads: Literal[TailReader.TRAILING, TailReader.GIVEN_SLOT] | None - # the walk as written, before any chain splices it: the link check - # below re-asks the link exception with its pre-#535 bound - written = rest - if reader is TailReader.NONE: - reads = None - chained: tuple[int, ...] = () - peeled = peel_trailing(rest, pieces, ptags, tokens, one_case) - elif reader is TailReader.TRAILING or reader is TailReader.GIVEN_SLOT: - reads = reader - # the FIRST-WORD FLOOR, as the chain's own: `rest` opens with - # the marker run, and the word after it stays the maiden name - # whatever it is, so the chain may not take it -- taken, it - # left the count the re-peel reads, and 'Jane Doe nee King. ba' - # kept 'ba' in the clause where 'Jane Doe nee Smith ba' gives - # it up (#535 review). A marker run with nothing after it has - # no floor to set; the walk below declines it anyway. - first = m + run if m + run < len(pieces) else None - floor = rest.index(first) + 1 if first in rest else 1 - rest, chained, peeled = tail_reading(rest, pieces, ptags, tokens, - one_case, floor) - else: - assert_never(reader) - trailing = rest[-1] if peeled.numeral is not None else len(pieces) - # The fork reads the piece before the numeral, and the take - # REMOVES that piece: afterwards assign sees the piece before the - # marker there, and if that is initial-shaped the fork will not - # fire -- a walk that stopped anyway handed the V to the family - # ('J. née Jones Smith V'). So the numeral must read as the suffix - # as the take would leave the name too, and the question is asked - # the way P5's reserve asks it (#425): the peel is run over the - # VIEW the take would leave, not one condition of it. Checking the - # preceding piece alone misses a title before the marker, where - # 'Dr. née Jones Smith V' leaves the numeral as assign's whole - # rest and no fork fires at all. - if trailing < len(pieces): - left = [i for i in range(len(pieces)) if i < m or i >= trailing] - view = [pieces[i] for i in left] - view_tags = [ptags[i] for i in left] - # The same peel pair as above, read over the view -- and, for - # the NONE reader, only `Peel.numeral` off it (every other - # reader reads the view through `tail_reading` and asks the - # shared release check below), because the bare-acronym fork - # COUNTS pieces and this view no longer holds the pieces it - # counted - # ('John née Jones Smith Ma' peeled over the pieces as written - # reads the acronym as a credential with words to spare, and - # once 'Jones Smith' has left it is the family of what - # remains). The acronym fork builds a view of its own below. - view_rest = peel_walk(leading_titles(view, view_tags, tokens), - view_tags) - if reads is None: - kept = peel_trailing(view_rest, view, view_tags, tokens, - one_case).numeral is None - else: - # the view read through the chain too, and the span behind - # the numeral asked the shared release question -- which - # includes the joins, the half this fork never asked: - # 'Berg, abdul nee Smith V' handed the V to the bound-given - # join and read given 'abdul V' (2.2 and 2.3; 2.0 and 2.1 - # read given 'abdul nee', suffix 'V' -- #411's own reserve - # differs there too, before #535 ever runs) - # - # After a family comma the given slot reads a lone numeral - # as a suffix only where the given part is the LAST comma - # part (#144, and assign's own two-segment condition): with - # a credential tail behind it the numeral is a middle - # initial there, so a clause may not give it up -- - # 'Doe, Jane nee Smith V, PhD' read middle 'V' (#535 - # review). - view_reading = tail_reading(view_rest, view, view_tags, - tokens, one_case) - view_peel = view_reading[2] - at = left.index(trailing) - # `tail_follows` is only ever true for the GIVEN_SLOT - # reader: group() computes it from that reader. - kept = (view_peel.numeral is None - or tail_follows - or not _release_reads_off( - view, view_tags, tokens, at + 1, len(view), at, - reads, one_case, view_reading, - tail_follows=tail_follows)) - if kept: - trailing = len(pieces) - # #533: the ACRONYM fork, asked the way the numeral is -- the peel - # over the pieces as they stand, then again over the name the take - # would leave, and a stop only where both read the word as the - # credential. `peeled.names` is a COUNT of positions in `rest`, so - # `rest[peeled.names]` is the first piece the peel took and the - # class member it stopped at; no walk of our own is needed. One - # peel, one view built once per take: O(pieces) for the take, not - # per member, and no re-entrancy -- the predicate never calls the - # walk that calls it. - if (reads is not None and peeled.names < len(rest) - and m + run + 1 < len(pieces)): - # THE FIRST-WORD FLOOR: the stop never takes the FIRST word - # after the marker -- a class member standing alone there - # stays the maiden name. A clamp rather than a veto: where the - # peel consumed that word AND words behind it, only the first - # stays ('Doe, J. nee MA ba' keeps maiden 'MA' and reads - # suffix 'ba'; a veto handed 'ba' back to the clause too). The - # clamped piece may then be no member at all, and the test - # below declines -- which changes nothing, the walk stopping - # at that suffix word of its own accord. - stop = max(rest[peeled.names], m + run + 1) - head = pieces[stop] - # `len(head) == 1` is DEFENSIVE and measured inert - # (2026-09-19) rather than unreachable: it asks a LONE piece's - # question, and the answer below reads `head[0]` as if the - # piece were the word. The multi-token piece it keeps out is - # the Ph. D. merge above, whose `suffix` ptag stops `peel_walk` - # ever returning it -- so the peel half of `stop` cannot be it, - # but the FIRST-WORD FLOOR is an index rather than a walk and - # does reach the merged piece whenever that piece is the second - # word after the marker ('BERG, ABDUL Z DOMU MA PH. D.'). - # - # THE NEGATIVE CONTROL, measured 2026-09-19 over a population - # built to HOLD that shape -- 4,224 names (twelve heads x four - # markers x 23 bodies, each also with a ', MD' tail and in - # upper and lower case) under six policies and two lexicons, - # the default and one listing `ph` ambiguous, 50,688 parses. A - # probe that fires wherever the tag test admits a head this - # length test then DECLINES -- the only sites where dropping it - # could matter -- fires 1,440 times, against 25,920 reaches of - # this site and 19,296 tag admissions. Dropping it is - # byte-identical all the same -- fields, ambiguities and every - # token's role and tags -- over 905,796 parses (that population - # plus the review's 142,518-name corpus under the six - # policies). What it buys is the price, which is real where the - # count is not: with `ph` listed, dropping it takes 'BERG, - # ABDUL Z DOMU MA PH. D.' from 438 frames to 453. Kept for the - # reason `_assign.previous_kept` is: an inert branch is cheaper - # than a question asked of the wrong shape, and the three - # sibling sites (`_pieces.segment_suffix_reading`, the - # GIVEN_SLOT branch of `_release_reads_off` above, and the - # emitter at the end of this function) each pair a length - # test with a tag test the same way. - # - # The tag is the CLASS the rule is stated in terms of, and it - # is not redundant with the walk the way the length test is: - # 'J. née Jones Smith V' reaches here on a piece - # `is_suffix_piece` REFUSES for being initial-shaped, and what - # declines it is the view check rather than the walk. Its own - # control is a price too, not a count: dropping it runs the - # view machinery over every ordinary credential the peel took - # -- 'Jane Doe nee Smith PhD' 341 -> 353 frames and 'Doe, Jane - # nee Smith PhD' 367 -> 372, counted per `Parser.parse` the way - # tests/v2/test_benchmark counts them. No test pins those - # numbers: `_CALL_BASELINE` is per-interpreter and per entry - # point, and a row for one name would have to be guessed for - # the four interpreters only CI runs. - # - # And `stop < trailing`, DEFENSIVE: a stop at or past the - # numeral stop would move the clause's end the wrong way. The - # walked piece is bounded by `trailing`, but the FLOOR is an - # index, and the title chain's splice can leave the numeral - # stop in front of it ('Jane Doe nee V Prof.' reaches `stop > - # trailing` above, though its head is no class member and the - # tag test below declines it anyway). Measured 2026-09-26: no - # input on the maiden grids has a class-member head with `stop - # >= trailing`, and dropping the guard changes no reading there -- - # kept because a caller's vocabulary can list a title as a - # class member. - if (stop < trailing and len(head) == 1 - and AMBIGUOUS_ACRONYM_TAG in tokens[head[0]].tags): - left = [i for i in range(len(pieces)) if i < m or i >= stop] - view = [pieces[i] for i in left] - view_tags = [ptags[i] for i in left] - # where the member stands in that view: everything before - # the marker, then the run the take would leave behind. - # The question is whether the reader takes THIS piece, not - # whether it takes something -- a suffix word behind the - # member answers yes to the weaker question while the - # member itself reads as the family name ('JOHN NEE JONES - # SMITH MA PHD' left 'JOHN MA PHD', whose MA is the family). - at = left.index(stop) - if reads is TailReader.GIVEN_SLOT: - # after a family comma the words to spare are there by - # construction, so the reader is #531's -- the member's - # own reading, asked through the one predicate that - # owns it, and that rule's FLOOR: the member ends the - # given part only where every piece behind it reads as - # a suffix too ('Doe, Jane nee Smith MA do' keeps - # maiden 'Smith MA do', the trailing particle not - # being a credential, so the clause keeps both words - # rather than handing one of them to the current - # name's middle -- which is what the clause-less - # 'Doe, Jane MA do' does with them, middle 'MA' and - # family 'do Doe'). The member is asked here, the span - # behind it and the name word ahead of it by the - # shared check below. Anchored (#544) as assign's given - # slot anchors it: over the view up to the member, from - # past the leading title run. - takes = credential_at_the_given_slot( - tokens[head[0]], one_case, - lambda: anchor_in_reach( - range(at - 1, -1, -1), view, view_tags, tokens) - and given_slot_anchors( - view, view_tags, tokens, - leading_titles(view, view_tags, tokens), - at + 1)[at]) - start = at + 1 - else: - # TRAILING: the peel over the view IS the member's - # question here, so the shared check asks it from the - # member itself: covered means the reading of the name - # left standing took this piece as the credential, so - # one `tail_reading` answers both questions. - takes = True - start = at - # rules.md#M2: "a word the clause gives up reads as a - # post-nominal or the clause keeps it" -- of the whole span - # the stop gives up, and with the joins below this pass - # asked too (#533 review, #535). - if takes and _release_reads_off(view, view_tags, tokens, - start, len(view), at, reads, - one_case, - tail_follows=tail_follows): - trailing = stop - # rules.md#M2 (#535): a trailing title the H5 chain takes ends the - # clause too -- asked as the other two stops are, over the name the - # take would leave (`_release_reads_off`), because the chain needs - # a name word to stand behind and 'Dr. nee Jones Smith Prof.' - # leaves 'Dr. Prof.', whose Prof. would be the family name. The - # FIRST-WORD FLOOR is the chain's own (`floor` above): a title - # standing straight after the marker stays the maiden name, the - # marker having announced one ('Jane Doe nee King.'), and where - # titles follow it only the ones behind it go ('Jane Doe nee Prof. - # Dr.' reads maiden 'Prof.', title 'Dr.'). - if chained and reads is not None: - stop = min(chained) - if stop < trailing: - left = [i for i in range(len(pieces)) if i < m or i >= stop] - view = [pieces[i] for i in left] - view_tags = [ptags[i] for i in left] - at = left.index(stop) - end = (left.index(trailing) if trailing < len(pieces) - else len(view)) - if _release_reads_off(view, view_tags, tokens, at, end, at, - reads, one_case, - tail_follows=tail_follows): - trailing = stop - # The walk starts past the WHOLE marker: a phrase's second word is - # the marker, not the first word it takes. With nothing behind the - # marker at all there is no clause to walk and no first word to - # bound it with, so the decline the `j <= m + run` test below - # reaches is taken here instead -- `lo` would have no piece to name - # ('Jane van der Berg née'). - if m + run >= len(pieces): + lo = m + _marker_run_pieces(pieces, tokens, m) + if lo >= len(pieces): return None - # rules.md#M2: "a link inside the birth name does not end it" -- - # the clause's OWN bounds for the link exception, which are not the - # segment's. `lo` is the first piece after the marker run, so the - # marker is never the name word on a link's left. No delimiter - # core stands anywhere in the clause: a tail segment's cores are - # cut out before this walk (#549), and a dash where no tail - # segment holds it is an ordinary word at either policy ('John PhD - # née - i Jones' keeps maiden '- i Jones' configured or not, where - # 'Smith, John, PhD née - i Jones' configured takes no clause). - # `peel_start` is where assign's trailing run begins -- over the - # pieces as written for a NONE reader (`trailing_start`'s whole - # answer, read off the peel pair above rather than re-running it), - # and for any other reader `tail_reading`'s title-aware answer - # (`rest[peeled.names]`, `rest` having been read through the H5 - # chain, so a chained title is not in it). Either way the - # generation or credential a clause ends with is never the name - # word on a link's right ('... nee Puig i III', '... i MA', whose - # MA carries no `vocab:suffix` tag for the piece test to refuse it - # by). - # - # It can stand PAST `trailing`: the title stop sets `trailing` to - # a piece the chain read past ('Jane Doe nee Smith Prof.' reaches - # here with `peel_start` 5 and `trailing` 4). That is harmless: the - # walk below never reaches `trailing`, so the link exception is - # never asked about a piece at or past it, and a bound past the - # walk's own end cannot let a link join across anything the walk - # visits. A dated snapshot from before #535, measured 2026-09-20 - # with a probe here over the whole suite: 93,408 reaches of this - # site, `peel_start > trailing` 0 of them. - lo = m + run - peel_start = (rest[peeled.names] if peeled.names < len(rest) - else len(pieces)) - j = lo - # The link exception's memo cell, filled inside the predicate on - # the first connective it is asked about (see its docstring). - beside: list[_Beside] = [] - while j < trailing: - if (not is_suffix_piece(pieces[j], ptags[j], tokens) - or _link_joins_inside_the_clause(j, lo, peel_start, - pieces, ptags, tokens, - beside)): - j += 1 - continue - # rules.md#M2: "a word the clause gives up reads as a - # post-nominal or the clause keeps it" -- asked here of a LINK - # the exception refused only BECAUSE the title chain moved its - # bound, and of no other stop. The exception reads the end of - # the name through the chain where a trailing rule reads the - # clause, so a link can now stop the walk where, bounded by the - # peel over the words as written, it joined ('Jane Doe nee - # Smith i DO Prof.': the DO is the peel's once the title is - # chained), and that new stop gives up the words behind it - # too. Where the name left standing would not read that run as - # post-nominals and titles, the clause keeps the link as it - # kept it before -- otherwise 'Doe i' became a middle name - # (#535). A link the as-written bound refuses too stops as it - # always did ('Doe, Jane nee Smith i V' gives suffix 'i V'), - # and a suffix word that is no link is #548's question. So is - # a link standing FIRST after the marker: stopping there - # declines the clause outright, which gives nothing up. The - # as-written peel is read only here, on a refused link. - if (j > m + run and reads is not None - and _is_conj_piece(pieces[j], ptags[j], tokens)): - as_written = peel_trailing(written, pieces, ptags, tokens, - one_case) - written_start = (written[as_written.names] - if as_written.names < len(written) - else len(pieces)) - if _link_joins_inside_the_clause(j, lo, written_start, - pieces, ptags, tokens, - beside): - left = [i for i in range(len(pieces)) if i < m or i >= j] - view = [pieces[i] for i in left] - view_tags = [ptags[i] for i in left] - at = left.index(j) - if not _release_reads_off(view, view_tags, tokens, at, - len(view), at, reads, one_case, - tail_follows=tail_follows): - j += 1 - continue - break - # j == m + run means nothing followed the marker but a suffix, so - # the pass declines and the marker stays ordinary words - # (rules.md#M2). - if j <= m + run: + before = m - 1 + lead = (1 if m == 1 and is_title_piece(pieces[0], ptags[0], tokens) + else leading_titles(pieces[:m], ptags[:m], tokens)) + if (before < lead + or (before > lead + and is_suffix_piece(pieces[before], ptags[before], tokens)) + or _is_conj_piece(pieces[before], ptags[before], tokens)): return None - # #533, mechanisms.md#AMBIGUITY-AT-THE-DECISION-SITE: "Emit at the - # site that takes the branch, not where an ambiguous tag sits" -- - # the walk that KEPT the word is where the fork was called, so it - # is where the report is raised. The LAST word of the maiden name - # is the one the trailing rule was asked about: everything behind - # it read as a suffix (that is what let the peel reach it), and a - # member the reading TOOK is not in the maiden name any more -- - # assign reports that one where it peels it, so no token is ever - # reported twice. A member with a name word behind it was never - # asked and stays silent, which rules.md#A1's hesitating reader is - # the reason for rather than the accident of. - # - # Gated on the reader for the same reason the walk is: where no - # trailing rule reads these words, nothing was decided and nothing - # may report. Gated on EITHER tag, as the chain emitter's is, so a - # by-shape member reports with the dotted switch off -- classify - # writes the shape tag there while the class does not admit it, - # which is the one place a declined fork can be recorded. - # - # `last[0]` is the whole word and no `len(last) == 1` guards it, - # unlike the three sibling sites: here the invariant IS structural - # and a test would be inert by construction rather than by - # measurement. The only multi-token piece this pass can see is the - # Ph. D. merge, `merge` gives that piece the `suffix` ptag - # unconditionally, and `is_suffix_piece` answers yes off that ptag - # alone -- so the walk just above ENDS before it, and the last - # maiden piece is never it whatever a caller's vocabulary says. - # Measured too, both ways: a probe on `len(last) != 1` here fired - # 0 times over 1,760,904 parses, and deleting the test is - # byte-identical (fields, ambiguities, token roles and tags) over - # the 905,796-parse oracle at the same frame counts. - last = pieces[j - 1] + # a suffix word straight after the marker: nothing was announced + if is_suffix_piece(pieces[lo], ptags[lo], tokens): + return None + if site is ClauseSite.FAMILY_PART: + end = next((j for j in range(lo + 1, len(pieces)) + if "suffix" in ptags[j] + or is_suffix_piece(pieces[j], ptags[j], tokens)), + len(pieces)) + return list(range(m, lo)), [(j, Role.MAIDEN) + for j in range(lo, end)] + if site is not ClauseSite.TRAILING: + assert_never(site) + # the last piece is a candidate on the numeral fork's own SHAPE test + # too, which asks no vocabulary ('VI' is in no suffix list, and + # 'John Smith VI' reads it as the suffix all the same) + c = len(pieces) + while c - 1 > lo and ( + _clause_tail_word(pieces[c - 1], ptags[c - 1], tokens) + or (c == len(pieces) and len(pieces[c - 1]) == 1 + and is_trailing_numeral_suffix( + tokens[pieces[c - 1][0]].text, + tokens[pieces[c - 2][0]].text))): + c -= 1 + # the tag test is the predicate's own necessary half, inline so a + # name word in the clause pays no frame + c = next((j for j in range(lo + 1, c) + if ("suffix" in ptags[j] + or "vocab:suffix" in tokens[pieces[j][0]].tags) + and starts_a_credential_run(pieces[j], ptags[j], tokens)), c) + end = len(pieces) + released: list[tuple[int, Role]] = [] + # nothing behind the clause that a trailing rule could take: the + # view would be the head alone, and the whole rest is the clause + if c < len(pieces): + left = list(range(m)) + list(range(c, len(pieces))) + view = [pieces[i] for i in left] + vtags = [ptags[i] for i in left] + vrest, vchained, vpeel = tail_reading( + peel_walk(leading_titles(view, vtags, tokens), vtags), + view, vtags, tokens, one_case) + titles = set(vchained) | set(vpeel.run_titles) + in_run = titles | set(vrest[vpeel.names:]) | { + q for q in range(m, len(view)) if "suffix" in vtags[q]} + e = len(view) + while e - 1 >= m and (e - 1) in in_run: + e -= 1 + end = c + (e - m) + released = [(j, Role.TITLE if m + (j - c) in titles + else Role.SUFFIX) for j in range(end, len(pieces))] + # mechanisms.md#AMBIGUITY-AT-THE-DECISION-SITE: "Emit at the + # site that takes the branch, not where an ambiguous tag sits" + # -- the take consumed the run, so the forks the reading called + # on it are reported here, where assign will never see them + consumed = {i for j in range(end, len(pieces)) for i in pieces[j]} + for piece in (*vpeel.picks, *vpeel.absorbed, + *(() if vpeel.numeral is None + else (vpeel.numeral,))): + if piece[0] in consumed: + ambiguities.append(PendingAmbiguity( + AmbiguityKind.SUFFIX_OR_NAME, + f"{tokens[piece[0]].text!r} after the maiden name " + f"reads as a suffix; it may be a name word", + tuple(piece))) + # ...and the member the clause KEEPS as its last word: the reading + # left it a name word, which is a fork too + last = pieces[end - 1] word = tokens[last[0]] - if (reads is not None - and not word.tags.isdisjoint(_AMBIGUOUS_CREDENTIAL_TAGS)): + if not word.tags.isdisjoint(_AMBIGUOUS_CREDENTIAL_TAGS): ambiguities.append(PendingAmbiguity( AmbiguityKind.SUFFIX_OR_NAME, - f"{word.text!r} ending the maiden name is also " - f"a post-nominal; the maiden marker's clause keeps it " - f"rather than reading it as one", - tuple(last))) - return list(range(m, m + run)), list(range(m + run, j)) + f"{word.text!r} ending the maiden name is also a " + f"post-nominal; the maiden marker's clause keeps it rather " + f"than reading it as one", tuple(last))) + return (list(range(m, lo)), + [(j, Role.MAIDEN) for j in range(lo, end)] + released) #: What `_between_name_words` reads instead of walking: two arrays over @@ -1102,9 +445,8 @@ def _run_neighbours(pieces: Sequence[Sequence[int]], nee Smith`, none of which reaches this at all. Called only where a generational connective was found in the - segment (the `frozen` loop) or where a clause's walk reached a - CONNECTIVE suffix piece (`_link_joins_inside_the_clause`), so no - ordinary name pays for it at all. + segment (the `frozen` loop), so no ordinary name pays for it at + all. Dropping either pass's `not` fails test_a_connective_piece_counts_toward_the_carve_outs_total @@ -1174,12 +516,10 @@ def _is_rootname(piece: Sequence[int], ptags: Set[str], # clause, its second sentence included, because that sentence is what # this predicate answering `False` means. Reading it as a generation # is the CALLER's half: `_group_segment`'s `frozen` set is where a -# connective this refuses is placed as the generation, and -# `_link_joins_inside_the_clause` is the maiden walk's. -# `Sequence[Sequence[int]]` rather than `Sequence[Piece]`, widened -# when the maiden walk became a second caller: this reads a piece and -# never edits one, and `_maiden_take` holds its pieces at the wider -# type the stage's entry point hands it. +# connective this refuses is placed as the generation. +# `Sequence[Sequence[int]]` rather than `Sequence[Piece]`: this reads a +# piece and never edits one (widened when the maiden walk was a +# second caller, before #601). def _between_name_words(k: int, lo: int, hi: int, pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], @@ -1266,9 +606,7 @@ def _group_segment(seg: tuple[int, ...], additional: int, opens_the_name: bool = False, *, one_case: bool | None, - reader: TailReader, - maiden_ambiguities: list[PendingAmbiguity], - tail_follows: bool = False, + site: ClauseSite, ) -> tuple[list[Piece], list[set[str]], MaidenTake | None]: pieces: list[Piece] = [[i] for i in seg] ptags: list[set[str]] = [set() for _ in seg] @@ -1277,22 +615,8 @@ def _group_segment(seg: tuple[int, ...], additional: int, # list) suppresses reporting -- see group() for when that applies. if ambiguities is None: ambiguities = [] - # The maiden walk's own channel, and REQUIRED beside `reader` for - # the same reason: the one production caller answers both off the - # segment's structure, and a default would be this module guessing - # what that caller already knows. group() passes `None` on the - # first for the chain emitter after a family comma -- the comma - # fixed the family, so that fork is settled -- and in a tail - # segment, which assign reads wholly as suffixes outside a maiden - # clause standing in it. #533's fork is - # neither: a credential ending a maiden clause is a question the - # comma settles nothing about, which is why the two channels are - # two parameters. They are given the SAME list wherever nothing is - # suppressed, which is every segment that is neither after a - # family comma nor a tail; what the split buys is the other case, - # where `None` on the first must not reach the second -- a maiden - # channel defaulting to whatever the first was would let a caller - # passing `ambiguities=None` silence both. + # The maiden take reports on this same channel: it reports only at + # the TRAILING site, which group() never suppresses. def title(k: int) -> bool: return is_title_piece(pieces[k], ptags[k], tokens) @@ -1404,11 +728,9 @@ def merge(lo: int, hi: int, add: Set[str] = frozenset(), else: k += 1 - # rules.md#M2: "a recognized maiden marker standing after at least - # one name word takes the words after it" -- up to any suffix - # word, or the trailing numeral assign reads as the suffix, as the - # maiden name, the marker itself dropped - # (history: decisions.md#M2) -- the marker pass (#274), and it runs + # rules.md#M2: "A marker that counts takes the words after it as + # the maiden name, and is itself dropped" (history: + # decisions.md#M2) -- the marker pass (#274), and it runs # BEFORE every join below. Each join rule asks a question about the # name -- how many words it has (P3's carve-out, P5's reserve), # what the word after a particle or a bound word is -- and the @@ -1426,13 +748,12 @@ def merge(lo: int, hi: int, add: Set[str] = frozenset(), # returns what it took, and group() records the drop and the roles. taken: MaidenTake | None = None take = _maiden_take(pieces, ptags, tokens, one_case, - reader, maiden_ambiguities, - tail_follows=tail_follows) + site, ambiguities) if take is not None: - marker_ks, maiden_ks = take + marker_ks, removed = take taken = ([i for k in marker_ks for i in pieces[k]], - [pieces[k] for k in maiden_ks]) - for k in reversed(marker_ks + maiden_ks): + [(pieces[k], role) for k, role in removed]) + for k in reversed(marker_ks + [k for k, _ in removed]): del pieces[k] del ptags[k] @@ -2005,34 +1326,34 @@ def group(state: ParseState) -> ParseState: # Suppressed after a family comma for the same reason _assign # suppresses it there: the family name is already fixed, so # there is no fork left to report. Suppressed in a tail segment - # as well, after either comma: assign reads that segment, - # outside a maiden clause standing in it ('Jane Doe, PhD, Jr - # nee van Ma' keeps maiden 'van Ma'), wholly as suffixes, so a + # as well, after either comma: assign reads that segment + # wholly as suffixes, a marker standing in it being an ordinary + # word (rules.md#M2, #601), so a # chain report there -- a particle # chained onto a name piece, or an acronym taken into the name # -- names a reading the parse never takes. rules.md#C2: "a part # the parse consumes wholly as suffixes raises no report about # reading a word of it as a name" tail = tail_start is not None and seg_idx >= tail_start - # #533: which rule reads what the maiden walk would leave, off - # the three facts already in hand here. A tail segment is read - # as credentials whole and segment 0 of a family comma is the - # family the comma named, so neither consults a trailing rule. + # #601: where a maiden clause would stand, off the three facts + # already in hand here -- which decides whether a marker counts + # at all and which rule reads the end of the part. if tail: - reader = TailReader.NONE + site = ClauseSite.TAIL elif family_comma: - reader = (TailReader.GIVEN_SLOT if seg_idx == 1 - else TailReader.NONE) + site = (ClauseSite.GIVEN_SLOT if seg_idx == 1 + else ClauseSite.FAMILY_PART) else: - reader = TailReader.TRAILING + site = ClauseSite.TRAILING # rules.md#C1: "a delimiter the policy declares parts a # trailing suffix part as a comma would" (#549): the segment is # grouped as the parts between its cores, each read exactly as # the comma twin's own segment would be, and the cores are # dropped (v1 expand_suffix_delimiter parity, #206), where # post_rules' entry pass reads them back as entry boundaries. - # No join, maiden walk or link search reaches across a core, - # because none of them ever sees one. A one-token segment that + # No join or link search reaches across a core, because none + # of them ever sees one, and a maiden marker in a tail part is + # an ordinary word (#601). A one-token segment that # is its core keeps it (v1 expand() splits within a part, never # erases a lone part); a segment of several cores and nothing # else drops them all, as the #206 drop did. @@ -2068,22 +1389,21 @@ def group(state: ParseState) -> ParseState: state.lexicon.given_name_titles, opens_the_name=(seg_idx == 0 and not family_comma), one_case=state.one_case, - reader=reader, - maiden_ambiguities=ambiguities, - tail_follows=(reader is TailReader.GIVEN_SLOT - and len(state.segments) > 2)) + site=site) pieces.extend(part_pieces) ptags.extend(part_ptags) # the marker is dropped and the maiden name's tokens become # MAIDEN (#274); which pieces those are was settled in # _group_segment, before the joins if taken is not None: - marker_piece, maiden_pieces = taken + marker_piece, removed_pieces = taken dropped.extend(marker_piece) - for piece in maiden_pieces: + # the maiden name, and the trailing run the take + # consumed with the roles the clause-free name's + # reading gave it (#601) + for piece, role in removed_pieces: for i in piece: - tokens[i] = copy_with( - tokens[i], role=Role.MAIDEN) + tokens[i] = copy_with(tokens[i], role=role) # Group used to decide the suffix ENTRY, marking a continuation # token "joined" between pieces off segment SHAPE -- `tail` by # index, ORed since #429 with segment_suffix_reading's diff --git a/nameparser/_pipeline/_pieces.py b/nameparser/_pipeline/_pieces.py index 24b0634f..e8f166f4 100644 --- a/nameparser/_pipeline/_pieces.py +++ b/nameparser/_pipeline/_pieces.py @@ -260,6 +260,39 @@ def is_suffix_piece(piece: Sequence[int], ptags: Set[str], return "vocab:suffix" in tags and "initial" not in tags +# rules.md#S2: "A title word never starts the run, and neither does a +# member of the ambiguous class, a single letter, a connective" -- what +# starts #602's credential run, one predicate for the trailing peel and +# the given-part walk (mechanisms.md#ONE-PREDICATE-PER-QUESTION) +def starts_a_credential_run(piece: Sequence[int], ptags: Set[str], + tokens: Sequence[WorkToken]) -> bool: + """Whether one piece starts #602's run: an unambiguous LATIN + suffix word of two or more letters -- a credential acronym or a + generational word -- that no other vocabulary claims as a name. + + `is_suffix_piece` already refuses the ambiguous class and an + initial-shaped letter. Refused here: any single letter (a lowercase + `i` or `v` mid-name is an initial or a connective far more often + than a numeral, and carries no `initial` tag), a connective, a + word that is also bound given-name or particle vocabulary (`abd` + heads `Abd Allah`; `vd`, `mc`), and the non-Latin honorific words. + """ + if len(piece) != 1 or not is_suffix_piece(piece, ptags, tokens): + return False + tok = tokens[piece[0]] + if not tok.tags.isdisjoint(_NOT_A_RUN_START): + return False + # C-level, no generator frame per character + letters = "".join(filter(str.isalpha, tok.text.lower())) + return len(letters) >= 2 and letters.isascii() + + +#: tags of suffix words another vocabulary claims as a name word or a +#: joiner, which therefore never start #602's run +_NOT_A_RUN_START = frozenset({"conjunction", "particle", + "vocab:bound-given"}) + + def is_lone_never_given_particle(piece: Sequence[int], tokens: Sequence[WorkToken]) -> bool: """A piece that is one never-given particle (rules.md#P1's fold @@ -399,8 +432,7 @@ def given_slot_anchors(pieces: Sequence[Sequence[int]], skipped piece. `start` is where the leading title run ends, so a title/suffix dual standing in it anchors nothing ('Smith, MD MA Ma'), and the piece at `start` is the given name the reserve keeps. - The one home of that query for assign's given slot, group's - GIVEN_SLOT reader, and the maiden clause's take.""" + The one home of that query for assign's given slot.""" stop = len(pieces) if end is None else end order: Sequence[int] = range(start, stop) if skip: @@ -614,6 +646,11 @@ class Peel(NamedTuple): numeral: tuple[int, ...] | None picks: tuple[tuple[int, ...], ...] anchors: list[bool] | None + #: #602 (`credential_run`): the pieces in rest[names:] that are + #: title words inside the credential run, and the name pieces the + #: run absorbed, which assign reports + run_titles: tuple[int, ...] = () + absorbed: tuple[tuple[int, ...], ...] = () # rules.md#S2: "a trailing word of the suffix vocabulary reads as a @@ -649,20 +686,18 @@ def trailing_start(start: int, pieces: Sequence[Sequence[int]], or a bare acronym with words to spare, into the family or the maiden name. - Both forks, always. A caller that needs one of them alone -- the - maiden walk re-asking the numeral over the view its take would - leave, where the acronym fork's piece COUNT no longer describes - the name -- calls the `peel_walk` + `peel_trailing` pair this - wraps and reads the half it wants (#533). A `numeral_only` flag - lived here for that one caller and cost it a frame. The maiden - walk asks that pair only for the NONE reader; every other reader - reads the same pieces through `tail_reading`, which runs this - pair's peel and H5's title chain to a fixed point. So this + Both forks, always, and #602's credential run on top + (`credential_run`), so P2's chain stops where the run starts. + Readers that also need H5's title chain read the same pieces + through `tail_reading`, which runs this pair's peel and the chain + to a fixed point -- the maiden take among them (#601). So this function's answer is the whole answer only where no trailing title chain also reads the pieces.""" rest = peel_walk(start, ptags) - peeled = peel_trailing(rest, pieces, ptags, tokens, one_case) - return rest[peeled.names] if peeled.names < len(rest) else len(pieces) + names = run_start(rest, peel_trailing(rest, pieces, ptags, tokens, + one_case).names, + pieces, ptags, tokens) + return rest[names] if names < len(rest) else len(pieces) # #289/#516: the "listed member, not by-shape" test both @@ -678,18 +713,13 @@ def trailing_start(start: int, pieces: Sequence[Sequence[int]], # keeps a non-member piece ("Smith, John"'s "John") from ever making # the call at all. # -# A THIRD caller since #533 -- credential_at_the_given_slot just +# A further caller since #533 -- credential_at_the_given_slot just # below -- deliberately does NOT pre-check: it owns #531's reading # and leaves membership to its own callers (its docstring says so), -# and the frame argument holds transitively because both of them ask +# and the frame argument holds transitively because its caller asks # inline -- `AMBIGUOUS_ACRONYM_TAG in tok.tags` after a -# `len(piece) == 1` at assign's given-part trailing slot, and the -# same pair in `_group.py`'s `_release_reads_off` (its GIVEN_SLOT -# branch) and in the acronym fork of `_maiden_take` (#535 review -# folded the view check's own `all(...)` into the shared release -# check, but the pre-check pair travelled with it rather than moving -# into this function). So no non-member piece reaches this function -# down that route either. +# `len(piece) == 1` at assign's given-part trailing slot. So no +# non-member piece reaches this function down that route either. def listed_lean(token: WorkToken, one_case: bool | None) -> Lean | None: """`ambiguous_lean` for a LISTED bare-ambiguous token, or None if the token is not tagged a listed member, is admitted by SHAPE @@ -722,12 +752,8 @@ def credential_at_the_given_slot( asked last, so a member the writing already settles never pays for the walk. - Called from assign's walk over the given part, from - `_release_reads_off`'s GIVEN_SLOT branch (the shared release check - every maiden-walk stop asks, rules.md#M2), and from the acronym - fork of `_maiden_take`, asked there of the member itself rather - than the span around it. One function rather than a condition - written at each + Called from assign's walk over the given part, as one function + rather than a condition written at the site (mechanisms.md#ONE-PREDICATE-PER-QUESTION). It is a text-and-tags question, which is what puts it in this module rather than beside any caller. @@ -875,6 +901,94 @@ def peel_trailing(rest: Sequence[int], pieces: Sequence[Sequence[int]], return Peel(k, numeral, tuple(picks), anchors) +# rules.md#S2: "A credential after the name core starts a run to the +# end of its part" -- applied to a finished peel (#602), so every +# reader of the trailing run (assign, P5's reserve, P2's chain stop, +# the maiden take) sees one answer. +def run_start(rest: Sequence[int], names: int, + pieces: Sequence[Sequence[int]], + ptags: Sequence[Set[str]], + tokens: Sequence[WorkToken]) -> int: + """Where #602's run starts in `rest[:names]`: the position of the + first credential that starts one (`starts_a_credential_run`) with + two name pieces in front of it, or `names` where none does -- with + no comma the family has to exist first. A lone particle does not + count toward the two, 'de Mesnil' being one surname. The position + alone, for readers that need only where the name ends + (`trailing_start`); `credential_run` builds the rest of the answer. + + The tag tests are INLINE ahead of the predicate, so a name word -- + every piece of an ordinary name -- and a connective or particle + never pay the predicate's frame: this runs on every parse, once + per piece, and a clause holding a run of links ('Puig i i i ... + Soler') otherwise paid two frames per link (decisions.md#parse-cost). + The predicate repeats the refusal; the inline copy is only the + cheap half of the same question. Under three names no run can + start, two core pieces having to stand in front of it.""" + if names < 3: + return names + core = 0 + for p in range(names): + q = rest[p] + tags = tokens[pieces[q][0]].tags + if (core >= 2 and "vocab:suffix" in tags + and tags.isdisjoint(_NOT_A_RUN_START) + and starts_a_credential_run(pieces[q], ptags[q], tokens)): + return p + # a lone particle is not yet a name: 'de Mesnil' is one surname + if not (len(pieces[q]) == 1 and "particle" in tags): + core += 1 + return names + + +def is_wholly_particle(piece: Sequence[int], + tokens: Sequence[WorkToken]) -> bool: + """Whether every token of a piece is particle vocabulary -- the + unit rules.md#P6 attaches after a family comma.""" + if len(piece) == 1: + return "particle" in tokens[piece[0]].tags + return all("particle" in tokens[i].tags for i in piece) + + +def has_name_content(piece: Sequence[int], + tokens: Sequence[WorkToken]) -> bool: + """Whether a piece holds a letter or digit: a piece with none is + no name word (rules.md#A2), so a run that takes it in absorbs no + name and reports nothing. C-level, no frame per character.""" + return any(map(str.isalnum, "".join([tokens[i].text for i in piece]))) + + +def credential_run(rest: Sequence[int], peel: Peel, p: int, + pieces: Sequence[Sequence[int]], + ptags: Sequence[Set[str]], + tokens: Sequence[WorkToken]) -> Peel: + """`peel` with its name count cut back to `p`, where #602's run + starts (`run_start`, which the caller asks first so that a name + with no run pays one frame, not two). + + Every word from the credential on is then in the suffix run, + except a title word, which is recorded in `run_titles` for assign + to read as a title; a name word the run absorbs is recorded in + `absorbed` for assign to report -- not a word with no letter or + digit in it, and not a member the peel already reported as a pick. + One pass, the suffix test first: a credential, the common word in + a run, leaves both lists without the title test's frame.""" + run_titles: list[int] = [] + absorbed: list[tuple[int, ...]] = [] + for r in range(p + 1, peel.names): + q = rest[r] + if is_suffix_piece(pieces[q], ptags[q], tokens): + continue + if is_title_piece(pieces[q], ptags[q], tokens): + run_titles.append(q) + continue + piece = tuple(pieces[q]) + if piece not in peel.picks and has_name_content(piece, tokens): + absorbed.append(piece) + return Peel(p, peel.numeral, peel.picks, peel.anchors, + tuple(run_titles), tuple(absorbed)) + + # rules.md#H5: "only a word the vocabulary knows as a title is one, # and a bare title word is a name word" # -- the trailing run's own predicate. NOT is_leading_title: @@ -887,7 +1001,6 @@ def peel_trailing(rest: Sequence[int], pieces: Sequence[Sequence[int]], def trailing_titles(rest: Sequence[int], pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], tokens: Sequence[WorkToken], - floor: int = 1, end: int | None = None) -> int: """How many pieces of `rest` the trailing title chain LEAVES standing: `rest[:kept]` are the name pieces and `rest[kept:]` the @@ -899,11 +1012,8 @@ def trailing_titles(rest: Sequence[int], pieces: Sequence[Sequence[int]], not claim, and in `tail_reading` the leftovers of whichever peel is current. `end`, where given, reads `rest[:end]` without the copy, for `tail_reading`'s passes. - Floor: `floor` leading positions of `rest` are never taken -- - 1 by default, so one name piece stands and a name is never all - title; the maiden walk passes the position just past the marker's - first word, so the chain never takes that word (rules.md#M2's - first-word floor, `tail_reading` passes it through). An empty + Floor: the first position of `rest` is never taken, so one name + piece stands and a name is never all title. An empty `rest` returns 0, which is what leaves assign's bare-suffix carve-out reached exactly as before. @@ -927,7 +1037,7 @@ def trailing_titles(rest: Sequence[int], pieces: Sequence[Sequence[int]], stops (decisions.md#parse-cost). """ k = len(rest) if end is None else end - while k > floor: + while k > 1: idx = rest[k - 1] piece = pieces[idx] # no #323 veto on the shape here, unlike is_leading_title's: @@ -946,8 +1056,8 @@ def trailing_titles(rest: Sequence[int], pieces: Sequence[Sequence[int]], # rules.md#H5: "successive single words that wear the abbreviation # shape and are title vocabulary chain into the title from the end" # -- one word's half of that, for a caller asking it of a piece the -# chain did not walk to: the maiden walk's release check, after a -# family comma (#535). `trailing_titles` asks the same three tests +# chain did not walk to: the maiden take, choosing the words the +# clause-free view is read over (#601). `trailing_titles` asks the same three tests # INLINE, in the same order, for the frame budget its docstring # states; keep the two in step. def is_trailing_title_word(piece: Sequence[int], ptags: Set[str], @@ -969,7 +1079,6 @@ def tail_reading(rest: list[int], pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], tokens: Sequence[WorkToken], one_case: bool | None, - floor: int = 1, ) -> tuple[list[int], tuple[int, ...], Peel]: """The S2 peel and the H5 chain read together to a FIXED POINT: peel, chain, splice the chained pieces out, peel again over what @@ -1001,8 +1110,8 @@ def tail_reading(rest: list[int], pieces: Sequence[Sequence[int]], One function for assign's placement, group's bound-given reserve (P5), which must count the name words assign will leave, and the - maiden walk and its release check (rules.md#M2), which must end - the clause where assign's reading of the name will begin. Deriving that agreement twice is what left the two + maiden take (rules.md#M2), which must end the clause where the + reading of the name will begin. Deriving that agreement twice is what left the two disagreeing at S2's bare-ambiguous reserve: 'abdul rahman MA' declined the join and 'abdul rahman MA Prof.' took it (mechanisms.md#ONE-PREDICATE-PER-QUESTION). @@ -1012,14 +1121,11 @@ def tail_reading(rest: list[int], pieces: Sequence[Sequence[int]], walk it built. Every caller reads the one returned here instead, which is the one the final peel partitions. - `floor` is `trailing_titles`' own: how many leading positions of - `rest` the chain may never take. 1 everywhere but the maiden walk, - whose `rest` opens with the marker and whose FIRST word after it - stays the maiden name whatever it is (rules.md#M2) -- a chain that - took that word spliced it out of the count the re-peel reads, and - 'Jane Doe nee King. ba' lost the suffix 'Jane Doe nee Smith ba' - keeps. The splice only ever removes positions at or - past the floor, so the floor names the same pieces every pass. + The chain's floor is `trailing_titles`' own, one name piece left + standing. + + #602's credential run is applied to the final peel on the way out + (`credential_run`), so every reader sees the run where it starts. Linear in the walk (#558). Re-peeling the whole walk every pass re-read each suffix the passes before had already peeled, so a @@ -1048,10 +1154,12 @@ def tail_reading(rest: list[int], pieces: Sequence[Sequence[int]], collected back to front and joined once. """ peeled = peel_trailing(rest, pieces, ptags, tokens, one_case) - kept = trailing_titles(rest, pieces, ptags, tokens, floor, - peeled.names) + kept = trailing_titles(rest, pieces, ptags, tokens, + end=peeled.names) if kept == peeled.names: - return rest, (), peeled + p = run_start(rest, peeled.names, pieces, ptags, tokens) + return rest, (), (peeled if p == peeled.names else credential_run( + rest, peeled, p, pieces, ptags, tokens)) # `rest[:hi]` stands as written; `behind` holds the peeled runs # the splices left after it, and `titled` the chained ones, each # back to front -- a pass's run goes in FRONT of what the pass @@ -1083,12 +1191,15 @@ def tail_reading(rest: list[int], pieces: Sequence[Sequence[int]], peeled = peel_trailing(rest, pieces, ptags, tokens, one_case) picks = list(peeled.picks) numeral = peeled.numeral - kept = trailing_titles(rest, pieces, ptags, tokens, floor, - peeled.names) + kept = trailing_titles(rest, pieces, ptags, tokens, + end=peeled.names) if behind: rest = rest[:hi] + [j for run in reversed(behind) for j in run] + final = Peel(peeled.names, numeral, tuple(picks), None) + p = run_start(rest, final.names, pieces, ptags, tokens) return (rest, tuple(j for run in reversed(titled) for j in run), - Peel(peeled.names, numeral, tuple(picks), None)) + final if p == final.names else credential_run( + rest, final, p, pieces, ptags, tokens)) # rules.md#H5: "the title is TRANSPARENT to the suffix reading: where diff --git a/nameparser/_pipeline/_state.py b/nameparser/_pipeline/_state.py index ebe646ef..2b48585d 100644 --- a/nameparser/_pipeline/_state.py +++ b/nameparser/_pipeline/_state.py @@ -140,7 +140,8 @@ class ParseState: pieces, still as sub-slices of the original, and every later index in the segment runs shifts by n); classify -> token tags AND one_case; group -> pieces/piece_tags/dropped AND maiden token - roles; + roles, plus the SUFFIX/TITLE roles of the run a maiden take gives + up (#601); assign -> the remaining token roles AND `order`, the effective order it read them under; post_rules -> roles again, and the ambiguity P6's attachment reports. diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 4ac2a3de..057a05e8 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -946,15 +946,17 @@ def _check_cjk_shape_purity(self) -> None: shape=1), Case("a_title_between_degree_and_member_keeps_it_a_name", "John Smith PhD Prof. Ma", - {"given": "John", "middle": "Smith PhD Prof.", "family": "Ma"}, - classification="fix(#289)", - ambiguities=("suffix-or-name",), - notes="rules.md#S2's Accepted limit: a title standing between " - "the credential and the member ends the run, so the " - "member's writing decides and the degree and title are " - "name text. Unchanged by #544; 2.3.0 read title 'Prof.', " - "suffix 'PhD Ma', and 1.4.0 middle 'Smith PhD', last " - "'Prof.', suffix 'Ma'"), + {"title": "Prof.", "given": "John", "family": "Smith", + "suffix": "PhD Ma"}, + classification="fix(#602)", + ambiguities=("suffix-or-name",), + notes="rules.md#S2 (#602): the credential after two name words " + "starts a run to the end of the part, the title inside it " + "reads as a title, and the member is in the run. The id " + "names the reading #544 pinned (middle 'Smith PhD Prof.', " + "family 'Ma'), which #602 retires; 2.3.0 read this same " + "way, and 1.4.0 middle 'Smith PhD', last 'Prof.', suffix " + "'Ma'"), Case("a_split_degree_in_front_is_out_of_the_walk", "John Smith Ph. D. MEng", {"given": "John", "middle": "Smith", "family": "MEng", @@ -970,12 +972,81 @@ def _check_cjk_shape_purity(self) -> None: "suffix 'Ph. D. MEng', and 1.4.0 'Ph. D., MEng'"), Case("a_member_the_particle_chain_took_is_out_of_reach", "John Smith PhD Do Do", - {"given": "John", "middle": "Smith PhD", "family": "Do Do"}, - notes="rules.md#S2's Accepted limit: 'Do' is particle " - "vocabulary too, and P2's chain joins 'Do Do' into one " - "piece before the peel, so no lone member stands behind " - "the degree. Unchanged by #544; 1.4.0 and 2.3.0 read " - "the same"), + {"given": "John", "family": "Smith", "suffix": "PhD Do Do"}, + classification="fix(#602)", + ambiguities=("suffix-or-name",), + notes="rules.md#S2 (#602): the credential starts the run and " + "P2's chain stops where it starts, so the 'Do Do' the " + "chain joined is absorbed and reported. The id names " + "#544's Accepted limit, which #602 retires; 1.4.0 and " + "2.3.0 read middle 'Smith PhD', family 'Do Do'"), + # #602: a credential after the name core starts a run to the end of + # its part (rules.md#S2). The movers, then the exclusions that keep + # a run from starting -- each a negative control for the movers. + Case("a_credential_starts_a_run_to_the_end_of_the_part", + "John Smith PhD Jones", + {"given": "John", "family": "Smith", "suffix": "PhD Jones"}, + classification="fix(#602)", + ambiguities=("suffix-or-name",), + notes="the credential is the writer's mark that the " + "post-nominals have begun; the name word behind it is " + "absorbed and reported. 2.3.0 read middle 'Smith PhD', " + "family 'Jones'"), + Case("a_generational_word_starts_the_run_too", "John Smith Jr. Jones", + {"given": "John", "family": "Smith", "suffix": "Jr. Jones"}, + classification="fix(#602)", + ambiguities=("suffix-or-name",), + notes="generational words start the run as credentials do " + "(decided on #602). 2.3.0 read middle 'Smith Jr.'"), + Case("a_title_inside_the_run_is_a_title", + "Eric H. Holder Jr. Attorney General", + {"title": "Attorney General", "given": "Eric", "middle": "H.", + "family": "Holder", "suffix": "Jr."}, + classification="fix(#602)", + notes="the credential starts the run and the title words " + "inside it read as titles, as the comma spelling 'Eric " + "H. Holder, Jr. Attorney General' always read. 2.3.0 " + "read middle 'H. Holder Jr. Attorney', family 'General'"), + Case("the_given_part_reads_the_run", "Smith, John PhD Jones", + {"given": "John", "family": "Smith", "suffix": "PhD Jones"}, + classification="fix(#602)", + ambiguities=("suffix-or-name",), + notes="after a family comma one given word is the core. 2.3.0 " + "read middle 'Jones', suffix 'PhD'"), + Case("a_title_inside_the_given_parts_run_is_a_title", + "Holder, Eric Jr. Attorney General", + {"title": "Attorney General", "given": "Eric", "family": "Holder", + "suffix": "Jr."}, + classification="fix(#602)"), + Case("a_title_word_does_not_start_the_run", "Mary Jane King Smith", + {"given": "Mary", "middle": "Jane King", "family": "Smith"}, + notes="746 title words are in no suffix set, many of them " + "surnames (measured 2026-10-04), so a title word never " + "starts the run: a title is read from the end of a name " + "only behind a comma or with a period (H5)"), + Case("a_joined_connective_is_no_run", "Josep Carod i Rovira", + {"given": "Josep", "family": "Carod i Rovira"}, + notes="'i' is a suffix word and a connective; a connective " + "never starts the run"), + Case("a_numeral_the_clause_free_name_reads_ends_the_clause", + "John née Jones Smith VI", + {"family": "John", "suffix": "VI", "maiden": "Jones Smith"}, + ambiguities=("suffix-or-name",), + notes="#601: the take finds a trailing numeral by the numeral " + "fork's SHAPE test, which asks no vocabulary -- 'vi' is in " + "no suffix list -- so the clause ends where 'John VI' would " + "read the suffix. #601's prototype tested vocabulary only " + "and kept 'VI' in the maiden name; a unit test under a " + "reduced lexicon caught it"), + Case("a_bound_given_word_is_no_run", "Mohamed Ali Abd Allah", + {"given": "Mohamed", "middle": "Ali Abd", "family": "Allah"}, + notes="'abd' is in the credential list (ABD) and heads 'Abd " + "Allah'; a word another vocabulary claims as a name never " + "starts the run"), + Case("one_name_word_is_no_core", "John PhD Smith", + {"given": "John", "middle": "PhD", "family": "Smith"}, + notes="with no comma two name words must stand in front of the " + "credential, the family having to exist first"), Case("a_particle_in_front_of_a_member_anchors_nothing", "Jan vd Ma", {"given": "Jan", "family": "vd Ma"}, @@ -1000,10 +1071,13 @@ def _check_cjk_shape_purity(self) -> None: # end. Decided as it reads, silent: no code special-cases it. Case("a_particle_after_a_credential_heads_a_title_case_name", "Doe, Jane PhD vd Ma", - {"given": "Jane", "middle": "vd Ma", "family": "Doe", - "suffix": "PhD"}, - classification="fix(#289)", - notes="S2: a particle in front of a member takes it out of " + {"given": "Jane", "family": "Doe", "suffix": "PhD vd Ma"}, + classification="fix(#602)", + ambiguities=("suffix-or-name",), + notes="#602: the credential starts the given part's run, so " + "'vd' and 'Ma' are in it, the member reported. Before " + "#602, and the reading the id names: " + "S2: a particle in front of a member takes it out of " "the given part's trailing slot, and where the word " "behind reads as a name the two are one name, asked " "nothing and reporting nothing -- 'vd' reads as 'van' " @@ -1267,17 +1341,19 @@ def _check_cjk_shape_purity(self) -> None: "whole. 2.3.0 read given 'von'"), Case("the_anchor_does_not_reach_across_a_maiden_clause", "Jane Doe Jr. nee Smith Ma", - {"given": "Jane", "family": "Doe", "suffix": "Jr.", - "maiden": "Smith Ma"}, - classification="fix(#533)", - ambiguities=("suffix-or-name",), - notes="rules.md#M2's boundary: the clause's name words stand " - "between 'Jr.' and 'Ma', so the credential in front " - "speaks for nothing past the marker and the clause keeps " - "its member, where 'Jane Doe Jr. Ma' reads suffix 'Jr. " - "Ma'. Unchanged by #544; 2.3.0 read the same fields " - "without the report, and 1.4.0 had no maiden routing", - shape=1), + {"given": "Jane", "family": "Doe", "suffix": "Jr. nee Smith Ma"}, + shape=1, + classification="fix(#601)", + ambiguities=("suffix-or-name", "suffix-or-name", "suffix-or-name"), + notes="#601: a marker behind a title or a suffix word is an ordinary " + "word (rules.md#M2). Before #601: given 'Jane', family 'Doe', " + "suffix 'Jr.', maiden 'Smith Ma'. As pinned then: rules.md#M2's " + "boundary: the clause's name words stand between 'Jr.' and " + "'Ma', so the credential in front speaks for nothing past the " + "marker and the clause keeps its member, where 'Jane Doe Jr. " + "Ma' reads suffix 'Jr. Ma'. Unchanged by #544; 2.3.0 read the " + "same fields without the report, and 1.4.0 had no maiden " + "routing"), Case("a_degree_after_a_one_word_family_anchors_the_member", "Smith, PhD MEng", {"family": "Smith", "suffix": "PhD MEng"}, @@ -1342,40 +1418,43 @@ def _check_cjk_shape_purity(self) -> None: shape=2), Case("the_anchor_reads_inside_a_clause_after_a_comma", "Doe, Jane nee Smith PhD MEng", - {"given": "Jane", "family": "Doe", "suffix": "PhD MEng", - "maiden": "Smith"}, - classification="fix(#544)", - ambiguities=("suffix-or-name",), - notes="the maiden walk's given-slot reader asks the same " - "company over the name the take would leave. 2.3.0 " - "read the same; #540 had read middle 'MEng', and 1.4.0 " - "had no maiden routing", - shape=2), + {"given": "Jane", "middle": "nee Smith", "family": "Doe", "suffix": "PhD MEng"}, + shape=2, + classification="fix(#601)", + ambiguities=("suffix-or-name",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', suffix 'PhD MEng', maiden 'Smith'. As pinned then: the " + "maiden walk's given-slot reader asks the same company over the " + "name the take would leave. 2.3.0 read the same; #540 had read " + "middle 'MEng', and 1.4.0 had no maiden routing"), Case("the_clause_gives_up_a_member_a_later_suffix_anchors", "Doe, Jane nee Smith MA Jr Ed", - {"given": "Jane", "family": "Doe", "suffix": "MA Jr Ed", - "maiden": "Smith"}, - classification="fix(#544)", + {"given": "Jane", "middle": "nee Smith", "family": "Doe", "suffix": "MA Jr Ed"}, + classification="fix(#601)", ambiguities=("suffix-or-name", "suffix-or-name"), - notes="the release check the maiden walk asks at the member " - "it stops on reads the span behind it the way assign's " - "given slot does, anchors included: 'Jr' speaks for the " - "Title-case 'Ed', so every word behind 'MA' reads as a " - "suffix and the clause gives 'MA' up. 2.3.0 read " - "middle 'Ed', suffix 'Jr', maiden 'Smith MA', and 1.4.0 " - "had no maiden routing"), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', suffix 'MA Jr Ed', maiden 'Smith'. As pinned then: the " + "release check the maiden walk asks at the member it stops on " + "reads the span behind it the way assign's given slot does, " + "anchors included: 'Jr' speaks for the Title-case 'Ed', so " + "every word behind 'MA' reads as a suffix and the clause gives " + "'MA' up. 2.3.0 read middle 'Ed', suffix 'Jr', maiden 'Smith " + "MA', and 1.4.0 had no maiden routing"), Case("the_clause_gives_up_a_title_a_later_suffix_anchors_past", "Doe, Jane nee Smith Dr. Jr Ma", - {"title": "Dr.", "given": "Jane", "family": "Doe", - "suffix": "Jr Ma", "maiden": "Smith"}, - classification="fix(#544)", - ambiguities=("suffix-or-name",), - notes="the same release check at a title the walk stops on: " - "the given part's title chain runs from the end through " - "'Ma' only because 'Jr' anchors it, and so reaches " - "'Dr.', which the clause gives up. 2.3.0 read middle " - "'Ma', suffix 'Jr', maiden 'Smith Dr.', and 1.4.0 had " - "no maiden routing"), + {"title": "Dr.", "given": "Jane", "middle": "nee Smith", "family": "Doe", "suffix": "Jr Ma"}, + classification="fix(#601)", + ambiguities=("suffix-or-name",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: title 'Dr.', given " + "'Jane', family 'Doe', suffix 'Jr Ma', maiden 'Smith'. As " + "pinned then: the same release check at a title the walk stops " + "on: the given part's title chain runs from the end through " + "'Ma' only because 'Jr' anchors it, and so reaches 'Dr.', which " + "the clause gives up. 2.3.0 read middle 'Ma', suffix 'Jr', " + "maiden 'Smith Dr.', and 1.4.0 had no maiden routing"), Case("a_dual_opening_the_given_part_is_a_title_there", "Smith, Ms Ma", {"title": "Ms", "given": "Ma", "family": "Smith"}, @@ -1796,6 +1875,48 @@ def _check_cjk_shape_purity(self) -> None: "tag and stays inside P6's run -- byte-identical before " "and after #531. Narrowing that condition to the " "ambiguous tag is what buys it"), + Case("a_trailing_particle_is_not_the_credential_runs_to_take", + "Smith, John PhD de", + {"given": "John", "family": "de Smith", "suffix": "PhD"}, + classification="fix(#379)", + notes="P6 over S2's given-part run (#602): a wholly-" + "particle piece is left to the walk and P6 attaches it, " + "so neither stage reports. The run had absorbed 'de' and " + "reported it, and P6 then reported the attachment as a " + "declined post-nominal -- two reports for a reading " + "neither stage kept (code review of #601/#602, " + "2026-10-04). 2.3.0 read this way; 1.4.0 middle 'de'"), + Case("a_particle_between_credentials_is_not_the_runs_to_take", + "Smith, John PhD de Jr.", + {"given": "John", "family": "de Smith", "suffix": "PhD, Jr."}, + classification="fix(#379)", + notes="P6: the run it looks past the post-nominal to " + "find is the same wholly-particle piece, so S2's run " + "leaves it alone wherever it stands. 2.3.0 read this " + "way; 1.4.0 middle 'de'"), + Case("a_particle_p6_does_not_attach_stays_in_the_run", + "Smith, John PhD de PhD van", + {"given": "John", "family": "van Smith", + "suffix": "PhD de PhD"}, + classification="fix(#602)", + ambiguities=("suffix-or-name", "particle-or-given"), + notes="P6 attaches only the particles ending the part, so the " + "run keeps 'de', a credential standing behind it and " + "another particle past that, and reports it as a word it " + "absorbed (fix-commit review of #601/#602, 2026-10-04). " + "2.3.0 read middle 'de', family 'van Smith', suffix " + "'PhD, PhD', the particle left a silent middle name " + "between two credentials"), + Case("a_trailing_particle_chain_keeps_its_own_report", + "Smith, John PhD Jr. de la", + {"given": "John", "family": "de la Smith", + "suffix": "PhD Jr."}, + classification="fix(#379)", + ambiguities=("particle-or-given",), + notes="P6: the attachment reports the ambiguous " + "particle it overrode, as it did before #602, rather than " + "a post-nominal the run had manufactured. 2.3.0 read this " + "way; 1.4.0 middle 'de la', suffix 'PhD, Jr.'"), Case("tussenvoegsel_behind_a_post_nominal", "Berg, Jan van Jr.", {"given": "Jan", "family": "van Berg", "suffix": "Jr."}, classification="fix(#379)", @@ -4008,15 +4129,17 @@ def _check_cjk_shape_purity(self) -> None: shape=1), Case("the_link_stays_in_the_birth_name_after_a_family_comma_too", "Doe, Jane nee Puig i Soler", - {"given": "Jane", "family": "Doe", "maiden": "Puig i Soler"}, - classification="fix(#397)", - notes="the comma form, and a different reader rather than a " - "second member of one shape: segment 1 is read at the " - "GIVEN slot, so the leak landed in given 'Jane i " - "Soler' rather than in the family. 2.0.0 through 2.3.0 " - "read maiden 'Puig' with middle 'Soler' and suffix " - "'i'. 1.4.0 read middle 'nee Puig Soler', suffix 'i'", - shape=2), + {"given": "Jane", "middle": "nee Puig i Soler", "family": "Doe"}, + shape=2, + classification="fix(#601)", + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', maiden 'Puig i Soler'. As pinned then: the comma form, " + "and a different reader rather than a second member of one " + "shape: segment 1 is read at the GIVEN slot, so the leak landed " + "in given 'Jane i Soler' rather than in the family. 2.0.0 " + "through 2.3.0 read maiden 'Puig' with middle 'Soler' and " + "suffix 'i'. 1.4.0 read middle 'nee Puig Soler', suffix 'i'"), Case("a_link_with_nothing_on_its_right_still_ends_the_clause", "Jane Doe nee Puig i", {"given": "Jane", "family": "Doe", "suffix": "i", @@ -4544,23 +4667,23 @@ def _check_cjk_shape_purity(self) -> None: shape=2), Case("a_maiden_clause_takes_the_member_with_it", "Doe, Jane nee Smith MA", - {"given": "Jane", "family": "Doe", "suffix": "MA", - "maiden": "Smith"}, - classification="fix(#533)", - ambiguities=("suffix-or-name",), - notes="the row that named the silence, now naming the " - "reading. The maiden marker no longer claims a " - "trailing credential: the words it takes end where a " - "trailing credential begins, and after a family comma " - "the reader of what is left standing is the given " - "part's own trailing slot (#531), which reads 'MA' as " - "the credential. 1.4.0 had no maiden routing and read " - "middle 'nee Smith', suffix 'MA', so the SUFFIX is " - "1.4.0 parity and the maiden field is not; 2.0.0 " - "through 2.3.0 read maiden 'Smith MA' in silence " - "(measured 2026-09-19). Keeping the id: it is the same " - "question, answered the other way", - shape=2), + {"given": "Jane", "middle": "nee Smith", "family": "Doe", "suffix": "MA"}, + shape=2, + classification="fix(#601)", + ambiguities=("suffix-or-name",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', suffix 'MA', maiden 'Smith'. As pinned then: the row " + "that named the silence, now naming the reading. The maiden " + "marker no longer claims a trailing credential: the words it " + "takes end where a trailing credential begins, and after a " + "family comma the reader of what is left standing is the given " + "part's own trailing slot (#531), which reads 'MA' as the " + "credential. 1.4.0 had no maiden routing and read middle 'nee " + "Smith', suffix 'MA', so the SUFFIX is 1.4.0 parity and the " + "maiden field is not; 2.0.0 through 2.3.0 read maiden 'Smith " + "MA' in silence (measured 2026-09-19). Keeping the id: it is " + "the same question, answered the other way"), # ---- #533: the maiden clause's trailing credential ------------- # The rule: the words a maiden marker takes end where a trailing # credential begins, and a member of the ambiguous credential @@ -4817,14 +4940,16 @@ def _check_cjk_shape_purity(self) -> None: shape=1), Case("a_trailing_title_after_a_family_comma_clause", "Doe, Jane nee Smith MA Prof.", - {"title": "Prof.", "given": "Jane", "family": "Doe", - "suffix": "MA", "maiden": "Smith"}, - classification="fix(#535)", - ambiguities=("suffix-or-name",), - notes="the given part's own slot reads the chain too " - "('Doe, Jane MA Prof.' gives title and suffix). 2.3.0 " - "read maiden 'Smith MA Prof.'", - shape=2), + {"title": "Prof.", "given": "Jane", "middle": "nee Smith", "family": "Doe", "suffix": "MA"}, + shape=2, + classification="fix(#601)", + ambiguities=("suffix-or-name",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: title 'Prof.', given " + "'Jane', family 'Doe', suffix 'MA', maiden 'Smith'. As pinned " + "then: the given part's own slot reads the chain too ('Doe, " + "Jane MA Prof.' gives title and suffix). 2.3.0 read maiden " + "'Smith MA Prof.'"), Case("the_first_word_after_the_marker_stays_even_as_a_title", "Jane Doe nee King.", {"given": "Jane", "family": "Doe", "maiden": "King."}, @@ -4844,11 +4969,14 @@ def _check_cjk_shape_purity(self) -> None: "read maiden 'Prof. Dr.'"), Case("a_title_stop_that_leaves_no_name_word_is_no_stop", "Dr. nee Jones Smith Prof.", - {"title": "Dr.", "maiden": "Jones Smith Prof."}, - notes="the view check: the take would leave 'Dr. Prof.', " - "where H5's chain has no name word to stand behind and " - "'Prof.' would be the family name, so the clause keeps " - "it. Unchanged from 2.3.0"), + {"title": "Dr. Prof.", "given": "nee", "middle": "Jones", "family": "Smith"}, + classification="fix(#601)", + notes="#601: a marker behind a title or a suffix word is an ordinary " + "word (rules.md#M2). Before #601: title 'Dr.', maiden 'Jones " + "Smith Prof.'. As pinned then: the view check: the take would " + "leave 'Dr. Prof.', where H5's chain has no name word to stand " + "behind and 'Prof.' would be the family name, so the clause " + "keeps it. Unchanged from 2.3.0"), Case("no_title_stop_before_a_family_comma", "Doe nee Smith Prof., Jane", {"given": "Jane", "family": "Doe", "maiden": "Smith Prof."}, @@ -4858,20 +4986,26 @@ def _check_cjk_shape_purity(self) -> None: "title. Unchanged from 2.3.0"), Case("a_particle_ahead_would_chain_the_released_title", "Jane van der Berg nee Smith Prof.", - {"given": "Jane", "family": "van der Berg", - "maiden": "Smith Prof."}, - notes="P2's chain runs on over a trailing title (rules.md#H5's " - "Accepted 'John van der Berg Prof.'), so releasing the " - "title would put it in the family; the clause keeps it. " - "Unchanged from 2.3.0"), + {"title": "Prof.", "given": "Jane", "family": "van der Berg", "maiden": "Smith"}, + classification="fix(#601)", + notes="#601: the clause ends at the clause-free name's trailing run, " + "which the take consumes (rules.md#M2). Before #601: given " + "'Jane', family 'van der Berg', maiden 'Smith Prof.'. As pinned " + "then: P2's chain runs on over a trailing title (rules.md#H5's " + "Accepted 'John van der Berg Prof.'), so releasing the title " + "would put it in the family; the clause keeps it. Unchanged " + "from 2.3.0"), Case("a_numeral_the_bound_join_would_take_stays_maiden", "Berg, abdul nee Smith V", - {"given": "abdul", "family": "Berg", "maiden": "Smith V"}, - classification="fix(#535)", - notes="the numeral stop now asks M2's join question too: " - "released, the V was taken by the bound-given join " - "after the comma (P5) and read given 'abdul V' at " - "2.3.0 -- a word of the birth name in the current one"), + {"given": "abdul", "middle": "nee Smith", "family": "Berg", "suffix": "V"}, + classification="fix(#601)", + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'abdul', " + "family 'Berg', maiden 'Smith V'. As pinned then: the numeral " + "stop now asks M2's join question too: released, the V was " + "taken by the bound-given join after the comma (P5) and read " + "given 'abdul V' at 2.3.0 -- a word of the birth name in the " + "current one"), Case("a_period_final_bracket_reads_as_the_bare_clause", "Jane Doe (nee Smith Prof.)", {"title": "Prof.", "given": "Jane", "family": "Doe", @@ -4894,63 +5028,69 @@ def _check_cjk_shape_purity(self) -> None: "reading #533 gave stands"), Case("a_given_slot_numeral_with_a_credential_tail_stays", "Doe, Jane nee Smith V, PhD", - {"given": "Jane", "family": "Doe", "suffix": "PhD", - "maiden": "Smith V"}, - classification="fix(#535)", - notes="rules.md#M2's given-slot reader declines the numeral " - "where a third comma part follows (#144's own " - "condition, asked of the maiden clause too): the given " - "part is not the LAST comma part, so the release is " - "withdrawn and the clause keeps 'V'. 2.3.0 read middle " - "'V', maiden 'Smith', suffix 'PhD' -- an M2 violation " - "predating #535"), + {"given": "Jane", "middle": "nee Smith V", "family": "Doe", "suffix": "PhD"}, + classification="fix(#601)", + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', suffix 'PhD', maiden 'Smith V'. As pinned then: " + "rules.md#M2's given-slot reader declines the numeral where a " + "third comma part follows (#144's own condition, asked of the " + "maiden clause too): the given part is not the LAST comma part, " + "so the release is withdrawn and the clause keeps 'V'. 2.3.0 " + "read middle 'V', maiden 'Smith', suffix 'PhD' -- an M2 " + "violation predating #535"), Case("a_lone_numeral_before_a_credential_tail_stays_maiden", "Doe, Jane nee V, PhD", - {"given": "Jane", "family": "Doe", "suffix": "PhD", - "maiden": "V"}, - classification="fix(#535)", - notes="the numeral stop reads FROM the marker and is not held " - "to the first-word floor; here the given slot does not " - "read a lone numeral as a suffix with a third comma part " - "behind it (#144's condition, asked of the clause since " - "#535), so the clause keeps 'V', as 2.0.0 and 2.1.0 read " - "it. 2.2.0, 2.3.0 and the #538 commit d9d80492 read " - "middle 'nee V', the marker a name word; 'Doe, Jane nee " - "V, Jr.' moves the same way"), + {"given": "Jane", "middle": "nee V", "family": "Doe", "suffix": "PhD"}, + classification="fix(#601)", + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', suffix 'PhD', maiden 'V'. As pinned then: the numeral " + "stop reads FROM the marker and is not held to the first-word " + "floor; here the given slot does not read a lone numeral as a " + "suffix with a third comma part behind it (#144's condition, " + "asked of the clause since #535), so the clause keeps 'V', as " + "2.0.0 and 2.1.0 read it. 2.2.0, 2.3.0 and the #538 commit " + "d9d80492 read middle 'nee V', the marker a name word; 'Doe, " + "Jane nee V, Jr.' moves the same way"), Case("the_given_title_chain_stops_at_a_member_with_a_title_behind", "Doe, Jane nee Smith Rev. MA Prof.", - {"title": "Prof.", "given": "Jane", "family": "Doe", - "suffix": "MA", "maiden": "Smith Rev."}, - classification="fix(#535)", - ambiguities=("suffix-or-name",), - notes="after a family comma the given part's title chain runs " - "from the end over what its first suffix pass leaves, " - "and that pass leaves 'MA' a name word while a title " - "stands behind it -- 'Doe, Jane Dr. MA Prof.' reads " - "middle 'Dr.' -- so the chain stops at 'MA' and 'Rev.' " - "stays in the clause; only 'MA' and 'Prof.' leave. 2.3.0 " - "and the #538 commit d9d80492 read maiden 'Smith Rev. " - "MA Prof.'"), + {"title": "Prof.", "given": "Jane", "middle": "nee Smith Rev.", "family": "Doe", "suffix": "MA"}, + classification="fix(#601)", + ambiguities=("suffix-or-name",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: title 'Prof.', given " + "'Jane', family 'Doe', suffix 'MA', maiden 'Smith Rev.'. As " + "pinned then: after a family comma the given part's title chain " + "runs from the end over what its first suffix pass leaves, and " + "that pass leaves 'MA' a name word while a title stands behind " + "it -- 'Doe, Jane Dr. MA Prof.' reads middle 'Dr.' -- so the " + "chain stops at 'MA' and 'Rev.' stays in the clause; only 'MA' " + "and 'Prof.' leave. 2.3.0 and the #538 commit d9d80492 read " + "maiden 'Smith Rev. MA Prof.'"), Case("a_lenient_numeral_leaves_with_the_title_in_front", "Doe, Jane nee Smith Prof. V", - {"title": "Prof.", "given": "Jane", "family": "Doe", - "suffix": "V", "maiden": "Smith"}, - classification="fix(#535)", - notes="the lenient trailing numeral (#144) is a suffix of the " - "given part's first pass once the numeral stop has asked " - "it, so the title in front of it leaves the clause too, " - "as 'Doe, Jane Prof. V' reads title 'Prof.', suffix 'V'. " - "2.3.0 read maiden 'Smith Prof.', suffix 'V'"), + {"title": "Prof.", "given": "Jane", "middle": "nee Smith", "family": "Doe", "suffix": "V"}, + classification="fix(#601)", + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: title 'Prof.', given " + "'Jane', family 'Doe', suffix 'V', maiden 'Smith'. As pinned " + "then: the lenient trailing numeral (#144) is a suffix of the " + "given part's first pass once the numeral stop has asked it, so " + "the title in front of it leaves the clause too, as 'Doe, Jane " + "Prof. V' reads title 'Prof.', suffix 'V'. 2.3.0 read maiden " + "'Smith Prof.', suffix 'V'"), Case("a_title_behind_that_numeral_still_stays", "Doe, Jane nee Smith V Prof., PhD", - {"title": "Prof.", "given": "Jane", "family": "Doe", - "suffix": "PhD", "maiden": "Smith V"}, - classification="fix(#535)", - notes="the same third-comma-part decline reaches through the " - "title chain: the numeral stays maiden text and the " - "title behind it still leaves the clause. 2.3.0 read " - "maiden 'Smith V Prof.', suffix 'PhD', with no title at " - "all"), + {"title": "Prof.", "given": "Jane", "middle": "nee Smith V", "family": "Doe", "suffix": "PhD"}, + classification="fix(#601)", + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: title 'Prof.', given " + "'Jane', family 'Doe', suffix 'PhD', maiden 'Smith V'. As " + "pinned then: the same third-comma-part decline reaches through " + "the title chain: the numeral stays maiden text and the title " + "behind it still leaves the clause. 2.3.0 read maiden 'Smith V " + "Prof.', suffix 'PhD', with no title at all"), Case("the_bound_given_join_is_the_given_slots_alone", "abdul nee Smith V", {"given": "abdul", "suffix": "V", "maiden": "Smith"}, @@ -4972,79 +5112,88 @@ def _check_cjk_shape_purity(self) -> None: "maiden 'Smith Dr.'"), Case("a_numeral_straight_after_the_marker_declines_through_a_title", "Jane Doe nee V Prof.", - {"title": "Prof.", "given": "Jane", "middle": "Doe", - "family": "nee", "suffix": "V"}, - classification="fix(#535)", - ambiguities=("suffix-or-name",), - notes="the numeral stop reads FROM the marker by design, " - "unlike the credential and title stops, so a numeral " - "standing straight after it is not held to the " - "first-word floor, and here it declines the clause -- " - "as 'Jane " - "Smith née V' does bare (rules.md#M2) -- and the title " - "behind the numeral then reads the declined name as " - "'Jane Doe nee V' plus the title would. 2.3.0 read " - "maiden 'V Prof.'"), + {"title": "Prof.", "given": "Jane", "family": "Doe", "maiden": "V"}, + classification="fix(#601)", + notes="#601: the first word after the marker is the maiden name (open " + "on #601) (rules.md#M2). Before #601: title 'Prof.', given " + "'Jane', middle 'Doe', family 'nee', suffix 'V'. As pinned " + "then: the numeral stop reads FROM the marker by design, unlike " + "the credential and title stops, so a numeral standing straight " + "after it is not held to the first-word floor, and here it " + "declines the clause -- as 'Jane Smith née V' does bare " + "(rules.md#M2) -- and the title behind the numeral then reads " + "the declined name as 'Jane Doe nee V' plus the title would. " + "2.3.0 read maiden 'V Prof.'"), Case("a_title_behind_a_tail_bound_numeral_keeps_both", "Doe, Jane nee Smith Prof. V, PhD", - {"given": "Jane", "family": "Doe", "suffix": "PhD", - "maiden": "Smith Prof. V"}, - classification="fix(#535)", - notes="NOT the transparent twin of 'Doe, Jane nee Smith V " - "Prof., PhD' (title 'Prof.', maiden 'Smith V'): the " - "third comma part's credential tail withdraws the " - "numeral's release, so 'V' stays a name word, and the " - "title stop is then asked and declines -- the given " - "part's chain runs from the end and stops at 'V', so " - "it never reaches 'Prof.', and the title stays maiden " - "text in front of the numeral. The " - "same asymmetry is the bare given slot's own, with no " - "marker in it -- 'Doe, Jane Prof. V, PhD' reads middle " - "'Prof. V' (no title) where 'Doe, Jane V Prof., PhD' " - "reads title 'Prof.', middle 'V'. 2.3.0 read middle " - "'V', maiden 'Smith Prof.'"), + {"given": "Jane", "middle": "nee Smith Prof. V", "family": "Doe", "suffix": "PhD"}, + classification="fix(#601)", + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', suffix 'PhD', maiden 'Smith Prof. V'. As pinned then: " + "NOT the transparent twin of 'Doe, Jane nee Smith V Prof., PhD' " + "(title 'Prof.', maiden 'Smith V'): the third comma part's " + "credential tail withdraws the numeral's release, so 'V' stays " + "a name word, and the title stop is then asked and declines -- " + "the given part's chain runs from the end and stops at 'V', so " + "it never reaches 'Prof.', and the title stays maiden text in " + "front of the numeral. The same asymmetry is the bare given " + "slot's own, with no marker in it -- 'Doe, Jane Prof. V, PhD' " + "reads middle 'Prof. V' (no title) where 'Doe, Jane V Prof., " + "PhD' reads title 'Prof.', middle 'V'. 2.3.0 read middle 'V', " + "maiden 'Smith Prof.'"), Case("a_link_the_walk_stops_at_gives_up_only_what_reads_off", "Jane Doe nee Smith i DO Prof.", - {"title": "Prof.", "given": "Jane", "family": "Doe", - "maiden": "Smith i DO"}, - classification="fix(#397/#535)", - ambiguities=("suffix-or-name",), - notes="with the title chained, the DO is the trailing peel's, " - "so the link exception refuses 'i' and the walk would " - "stop there -- and a stop at a link gives up the words " - "behind it. The name left standing ('Jane Doe i DO " - "Prof.') does not read 'i DO' as post-nominals, so the " - "clause keeps the link and the DO (reporting the kept " - "credential), and only the title leaves. 2.3.0 read " - "middle 'Doe i', family 'DO Prof.', maiden 'Smith' ('i' " - "was a plain suffix word there); the #538 commit " + {"title": "Prof.", "given": "Jane", "family": "Doe", "suffix": "i DO", "maiden": "Smith"}, + classification="fix(#601)", + ambiguities=("suffix-or-name",), + notes="#601: the clause ends at the clause-free name's trailing run, " + "which the take consumes (rules.md#M2). Before #601: title " + "'Prof.', given 'Jane', family 'Doe', maiden 'Smith i DO'. As " + "pinned then: with the title chained, the DO is the trailing " + "peel's, so the link exception refuses 'i' and the walk would " + "stop there -- and a stop at a link gives up the words behind " + "it. The name left standing ('Jane Doe i DO Prof.') does not " + "read 'i DO' as post-nominals, so the clause keeps the link and " + "the DO (reporting the kept credential), and only the title " + "leaves. 2.3.0 read middle 'Doe i', family 'DO Prof.', maiden " + "'Smith' ('i' was a plain suffix word there); the #538 commit " "d9d80492 read maiden 'Smith i DO Prof.'"), Case("a_link_whose_run_would_land_in_a_name_part_stays", "Berg, abdul nee Smith i V Prof., MD", - {"given": "abdul", "family": "Berg", "suffix": "MD", - "maiden": "Smith i V Prof."}, - classification="fix(#397)", - notes="the given-slot twin of the row above: giving up " - "'i V Prof.' would leave the V a middle name behind the " - "bound pair, so the clause keeps the whole run. 2.3.0 " - "read middle 'V', title 'Prof.', suffix 'i, MD', maiden " - "'Smith'; the #538 commit d9d80492 read as this does"), + {"title": "Prof.", "given": "abdul", "middle": "nee Smith V", "family": "Berg", "suffix": "i, MD"}, + classification="fix(#601)", + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'abdul', " + "family 'Berg', suffix 'MD', maiden 'Smith i V Prof.'. As " + "pinned then: the given-slot twin of the row above: giving up " + "'i V Prof.' would leave the V a middle name behind the bound " + "pair, so the clause keeps the whole run. 2.3.0 read middle " + "'V', title 'Prof.', suffix 'i, MD', maiden 'Smith'; the #538 " + "commit d9d80492 read as this does"), Case("a_released_particle_title_is_kept_after_a_family_comma", "Doe, Jane nee Smith St.", - {"given": "Jane", "family": "Doe", "maiden": "Smith St."}, - classification="fix(#274)", - notes="'St.' is a title and a particle. Released after a " - "family comma, P6 would attach it to the family ('Doe, " - "Jane St.' reads family 'St. Doe'), a word of the birth " - "name carried into the current one, so the clause keeps " - "it. Unchanged from 2.3.0"), + {"given": "Jane", "middle": "nee Smith", "family": "St. Doe"}, + classification="fix(#601)", + ambiguities=("particle-or-given",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', maiden 'Smith St.'. As pinned then: 'St.' is a title " + "and a particle. Released after a family comma, P6 would attach " + "it to the family ('Doe, Jane St.' reads family 'St. Doe'), a " + "word of the birth name carried into the current one, so the " + "clause keeps it. Unchanged from 2.3.0"), Case("a_released_particle_title_behind_a_credential_is_kept", "Doe, Jane nee Smith MA St.", - {"given": "Jane", "family": "Doe", "maiden": "Smith MA St."}, - classification="fix(#274)", - notes="the same guard with a credential in front of the " - "title: releasing 'MA St.' would hand 'St.' to the " - "family, so the clause keeps both. Unchanged from 2.3.0"), + {"given": "Jane", "middle": "nee Smith", "family": "St. Doe", "suffix": "MA"}, + classification="fix(#601)", + ambiguities=("particle-or-given", "suffix-or-name"), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', maiden 'Smith MA St.'. As pinned then: the same guard " + "with a credential in front of the title: releasing 'MA St.' " + "would hand 'St.' to the family, so the clause keeps both. " + "Unchanged from 2.3.0"), Case("a_split_credential_behind_the_member_counts_as_released", "Jane Doe nee Smith MA Ph. D.", {"given": "Jane", "family": "Doe", "suffix": "MA Ph. D.", @@ -5057,12 +5206,14 @@ def _check_cjk_shape_purity(self) -> None: "whole. 2.3.0 read maiden 'Smith MA', suffix 'Ph. D.'"), Case("a_split_credential_behind_the_member_after_a_family_comma", "Doe, Jane nee Smith MA Ph. D.", - {"given": "Jane", "family": "Doe", "suffix": "MA Ph. D.", - "maiden": "Smith"}, - classification="fix(#533)", - ambiguities=("suffix-or-name",), - notes="the given-slot twin of the row above. 2.3.0 read " - "maiden 'Smith MA', suffix 'Ph. D.'"), + {"given": "Jane", "middle": "nee Smith", "family": "Doe", "suffix": "MA Ph. D."}, + classification="fix(#601)", + ambiguities=("suffix-or-name",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', suffix 'MA Ph. D.', maiden 'Smith'. As pinned then: the " + "given-slot twin of the row above. 2.3.0 read maiden 'Smith " + "MA', suffix 'Ph. D.'"), Case("a_connective_behind_the_member_stops_the_peel", "Jane Doe nee Smith MA y", {"given": "Jane", "family": "Doe", "maiden": "Smith MA y"}, @@ -5210,54 +5361,61 @@ def _check_cjk_shape_purity(self) -> None: # ---- #533 after a family comma: #531's slot is the reader ------ Case("the_comma_reader_declines_the_title_cased_member", "Doe, Jane nee Smith Ma", - {"given": "Jane", "family": "Doe", "maiden": "Smith Ma"}, - classification="fix(#533)", - ambiguities=("suffix-or-name",), - notes="after a family comma the comma has already settled " - "the count, so the reader is #531's slot and the " - "writing decides alone -- Title case in a mixed-case " - "name keeps the word. Reported either way. 1.4.0 read " - "suffix 'Ma'", - shape=2), + {"given": "Jane", "middle": "nee Smith Ma", "family": "Doe"}, + shape=2, + classification="fix(#601)", + ambiguities=("suffix-or-name",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', maiden 'Smith Ma'. As pinned then: after a family comma " + "the comma has already settled the count, so the reader is " + "#531's slot and the writing decides alone -- Title case in a " + "mixed-case name keeps the word. Reported either way. 1.4.0 " + "read suffix 'Ma'"), Case("the_comma_reader_takes_the_all_lower_member", "Doe, Jane nee Smith ma", - {"given": "Jane", "family": "Doe", "suffix": "ma", - "maiden": "Smith"}, - classification="fix(#533)", - ambiguities=("suffix-or-name",), - notes="the row that makes the COMMA reader load-bearing " - "rather than decorative: the count-based reader " - "declines an all-lower member with two pieces to " - "spare, and #531's declines nothing but a particle. " - "Measured -- this name moves under the comma reader " - "and would not under the count. 'Doe, Jane ma' already " - "reads suffix 'ma', which is what the clause form now " - "agrees with", - shape=2), + {"given": "Jane", "middle": "nee Smith", "family": "Doe", "suffix": "ma"}, + shape=2, + classification="fix(#601)", + ambiguities=("suffix-or-name",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', suffix 'ma', maiden 'Smith'. As pinned then: the row " + "that makes the COMMA reader load-bearing rather than " + "decorative: the count-based reader declines an all-lower " + "member with two pieces to spare, and #531's declines nothing " + "but a particle. Measured -- this name moves under the comma " + "reader and would not under the count. 'Doe, Jane ma' already " + "reads suffix 'ma', which is what the clause form now agrees " + "with"), Case("the_do_pair_after_a_comma_keeps_the_particle_spelling", "Doe, Jane nee Smith do", - {"given": "Jane", "family": "Doe", "maiden": "Smith do"}, - classification="fix(#533)", - ambiguities=("suffix-or-name",), - notes="'do' is the one class member that is also particle " - "vocabulary, so after a comma the clause's trailing " - "slot, #531's slot and P6's attachment all want it -- " - "and the reading is #531's, unchanged and shared " - "through one predicate. Lean None plus a particle tag " - "means P6's word, so the clause keeps it and reports " - "the fork it consulted", - shape=2), + {"given": "Jane", "middle": "nee Smith", "family": "do Doe"}, + shape=2, + classification="fix(#601)", + ambiguities=("particle-or-given",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', maiden 'Smith do'. As pinned then: 'do' is the one " + "class member that is also particle vocabulary, so after a " + "comma the clause's trailing slot, #531's slot and P6's " + "attachment all want it -- and the reading is #531's, unchanged " + "and shared through one predicate. Lean None plus a particle " + "tag means P6's word, so the clause keeps it and reports the " + "fork it consulted"), Case("the_do_pair_after_a_comma_reads_the_capitals", "Doe, Jane nee Smith DO", - {"given": "Jane", "family": "Doe", "suffix": "DO", - "maiden": "Smith"}, - classification="fix(#533)", - ambiguities=("suffix-or-name",), - notes="the same word where the capitals speak: a positive " - "credential lean is the one spelling that outranks " - "P6's attachment (decisions.md#S2, 2026-09-18), so the " - "clause gives the word up. 1.4.0 read suffix 'DO'", - shape=2), + {"given": "Jane", "middle": "nee Smith", "family": "Doe", "suffix": "DO"}, + shape=2, + classification="fix(#601)", + ambiguities=("suffix-or-name",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', suffix 'DO', maiden 'Smith'. As pinned then: the same " + "word where the capitals speak: a positive credential lean is " + "the one spelling that outranks P6's attachment " + "(decisions.md#S2, 2026-09-18), so the clause gives the word " + "up. 1.4.0 read suffix 'DO'"), Case("the_no_comma_do_reads_by_the_count_instead", "Jane Doe nee Smith do", {"given": "Jane", "family": "Doe", "suffix": "do", @@ -5282,35 +5440,38 @@ def _check_cjk_shape_purity(self) -> None: shape=1), Case("the_comma_floor_keeps_a_member_a_particle_follows", "Doe, Jane nee Smith MA do", - {"given": "Jane", "family": "Doe", "maiden": "Smith MA do"}, - classification="fix(#533)", - ambiguities=("suffix-or-name",), - notes="where #531's FLOOR earns its place, and a check that " - "asked only about 'MA' got this wrong: the take would " - "leave 'Jane MA do', where 'do' does not read as a " - "suffix (P6 keeps it) so #531's slot reads 'MA' as a " - "MIDDLE name -- releasing it from the clause would " - "move a word from one person's name into another's. " - "With the floor the clause keeps 'MA do' whole, and " - "reports the 'do' it kept. rules.md#S2's own 'Doe, " - "John MA do' clause is the reading this rests on", - shape=2), + {"given": "Jane", "middle": "nee Smith MA", "family": "do Doe"}, + shape=2, + classification="fix(#601)", + ambiguities=("particle-or-given",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', maiden 'Smith MA do'. As pinned then: where #531's " + "FLOOR earns its place, and a check that asked only about 'MA' " + "got this wrong: the take would leave 'Jane MA do', where 'do' " + "does not read as a suffix (P6 keeps it) so #531's slot reads " + "'MA' as a MIDDLE name -- releasing it from the clause would " + "move a word from one person's name into another's. With the " + "floor the clause keeps 'MA do' whole, and reports the 'do' it " + "kept. rules.md#S2's own 'Doe, John MA do' clause is the " + "reading this rests on"), Case("the_clamp_never_takes_the_first_word_after_the_marker", "Doe, J. nee MA ba", - {"given": "J.", "family": "Doe", "suffix": "ba", - "maiden": "MA"}, - classification="fix(#533)", + {"given": "J.", "middle": "nee", "family": "Doe", "suffix": "MA ba"}, + shape=2, + classification="fix(#601)", ambiguities=("suffix-or-name", "suffix-or-name"), - notes="the floor is a CLAMP, not a veto, and this row is " - "why. The peel takes 'ba' and then 'MA', so the first " - "piece the peel took IS the only maiden word; a veto " - "that cancelled the stop whenever nothing would be " - "left handed 'ba' back to the clause too, giving " - "maiden 'MA ba' where 'Doe, J. ba' reads suffix 'ba'. " - "Clamped to the piece after the marker, the clause " - "keeps 'MA' and gives 'ba' up. TWO reports: the walk's " - "for the word it kept, assign's for the word it took", - shape=2), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'J.', family " + "'Doe', suffix 'ba', maiden 'MA'. As pinned then: the floor is " + "a CLAMP, not a veto, and this row is why. The peel takes 'ba' " + "and then 'MA', so the first piece the peel took IS the only " + "maiden word; a veto that cancelled the stop whenever nothing " + "would be left handed 'ba' back to the clause too, giving " + "maiden 'MA ba' where 'Doe, J. ba' reads suffix 'ba'. Clamped " + "to the piece after the marker, the clause keeps 'MA' and gives " + "'ba' up. TWO reports: the walk's for the word it kept, " + "assign's for the word it took"), Case("the_clamps_control_without_the_clause", "Doe, J. ba", {"given": "J.", "family": "Doe", "suffix": "ba"}, @@ -5324,49 +5485,54 @@ def _check_cjk_shape_purity(self) -> None: shape=2), Case("the_clause_leaves_a_middle_initial_alone", "Doe, Jane Q. nee Smith MA", - {"given": "Jane", "middle": "Q.", "family": "Doe", - "suffix": "MA", "maiden": "Smith"}, - classification="fix(#533)", - ambiguities=("suffix-or-name",), - notes="the member leaves the clause and 'Q.' stays the " - "middle initial it always was -- the stop reaches the " - "clause's trailing word, not the name in front of it", - shape=2), + {"given": "Jane", "middle": "Q. nee Smith", "family": "Doe", "suffix": "MA"}, + shape=2, + classification="fix(#601)", + ambiguities=("suffix-or-name",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', middle " + "'Q.', family 'Doe', suffix 'MA', maiden 'Smith'. As pinned " + "then: the member leaves the clause and 'Q.' stays the middle " + "initial it always was -- the stop reaches the clause's " + "trailing word, not the name in front of it"), Case("a_no_name_segment_leaves_the_clause_nobody_to_read_it", "Doe, Dr. nee Smith MA", - {"title": "Dr.", "family": "Doe", "maiden": "Smith MA"}, - classification="parity", - ambiguities=("suffix-or-name",), - notes="rules.md#M2's invariant, and the row that used to " - "record the opposite: an earlier round of #533 read " - "suffix 'MA' here and said so in silence. Once 'Smith' " - "leaves with the marker, segment 1 is 'Dr. MA' -- a " - "no-name segment, which the credential-run gate reads " - "whole without ever reaching #531's emitter, so the " - "released word would have landed in `given`, not in " - "`suffix`. With no name word ahead of it the member is " - "no trailing word of a given part, the walk declines, " - "and the clause keeps it. The REPORT survives the " - "decline: the emitter asks whether a trailing rule " - "reads these words at all, which after a family comma " - "it does. 1.4.0 read given 'nee', middle 'Smith', " - "suffix 'MA'", - shape=2), + {"title": "Dr.", "given": "nee", "middle": "Smith", "family": "Doe", "suffix": "MA"}, + shape=2, + classification="fix(#601)", + ambiguities=("suffix-or-name",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: title 'Dr.', family " + "'Doe', maiden 'Smith MA'. As pinned then: rules.md#M2's " + "invariant, and the row that used to record the opposite: an " + "earlier round of #533 read suffix 'MA' here and said so in " + "silence. Once 'Smith' leaves with the marker, segment 1 is " + "'Dr. MA' -- a no-name segment, which the credential-run gate " + "reads whole without ever reaching #531's emitter, so the " + "released word would have landed in `given`, not in `suffix`. " + "With no name word ahead of it the member is no trailing word " + "of a given part, the walk declines, and the clause keeps it. " + "The REPORT survives the decline: the emitter asks whether a " + "trailing rule reads these words at all, which after a family " + "comma it does. 1.4.0 read given 'nee', middle 'Smith', suffix " + "'MA'"), Case("the_bound_given_join_would_take_the_released_member", "Berg, abdul nee Jones MA", - {"given": "abdul", "family": "Berg", "maiden": "Jones MA"}, - classification="parity", - ambiguities=("suffix-or-name",), - notes="the other half of M2's invariant: P5's LENIENT " - "post-comma join runs BELOW the marker pass and would " - "swallow the released 'MA' into the bound-given pair " - "before assign could read it -- 'abdul MA' as the " - "given name, which is where the clause-less control " - "below genuinely puts it. A word joined away is a word " - "the clause gave up for nothing, so the walk declines " - "and keeps it. An earlier round of #533 released it " - "and read given 'abdul MA' in silence", - shape=2), + {"given": "abdul", "middle": "nee Jones", "family": "Berg", "suffix": "MA"}, + shape=2, + classification="fix(#601)", + ambiguities=("suffix-or-name",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'abdul', " + "family 'Berg', maiden 'Jones MA'. As pinned then: the other " + "half of M2's invariant: P5's LENIENT post-comma join runs " + "BELOW the marker pass and would swallow the released 'MA' into " + "the bound-given pair before assign could read it -- 'abdul MA' " + "as the given name, which is where the clause-less control " + "below genuinely puts it. A word joined away is a word the " + "clause gave up for nothing, so the walk declines and keeps it. " + "An earlier round of #533 released it and read given 'abdul MA' " + "in silence"), Case("the_bound_given_joins_control_without_the_clause", "Berg, abdul MA", {"given": "abdul MA", "family": "Berg"}, @@ -5379,33 +5545,38 @@ def _check_cjk_shape_purity(self) -> None: shape=2), Case("a_particle_chain_would_take_the_released_member", "Berg, Jane van der nee Smith DO", - {"given": "Jane", "family": "van der Berg", - "maiden": "Smith DO"}, - classification="parity", - ambiguities=("particle-or-given", "suffix-or-name"), - notes="M2's invariant against P2 rather than P5. 'DO' is " - "particle vocabulary standing behind a particle piece, " - "so the chain below this pass would absorb it into the " - "family -- a word crossing from the BIRTH name into " - "the current one, which is #424's failure from the " - "other side. An earlier round of #533 released it and " - "read family 'van der DO Berg' in silence, and the " - "clause-less 'Berg, Jane van der DO' reads that way " - "for its own reasons and is unchanged", - shape=2), + {"given": "Jane", "middle": "van der nee Smith", "family": "Berg", "suffix": "DO"}, + shape=2, + classification="fix(#601)", + ambiguities=("suffix-or-name",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Degenerate input, pinned to " + "detect change rather than for a reading anyone wants " + "(decisions.md#M2). Before #601: given 'Jane', family " + "'van der Berg', maiden 'Smith DO'. As pinned then: M2's " + "invariant against P2 rather than P5. 'DO' is particle " + "vocabulary standing behind a particle piece, so the chain " + "below this pass would absorb it into the family -- a word " + "crossing from the BIRTH name into the current one, which is " + "#424's failure from the other side. An earlier round of #533 " + "released it and read family 'van der DO Berg' in silence, and " + "the clause-less 'Berg, Jane van der DO' reads that way for its " + "own reasons and is unchanged"), Case("two_released_particle_members_would_chain_each_other", "Jane Doe nee Smith DO DO", - {"given": "Jane", "family": "Doe", "maiden": "Smith DO DO"}, - classification="parity", - ambiguities=("suffix-or-name",), - notes="the no-comma spelling of the same decline, where what " - "would do the joining is the OTHER released member: " - "the trailing peel reads both 'DO's as credentials, " - "but the moment they are out of the clause the first " - "is a non-leading particle and chains the second into " - "family 'DO DO'. An earlier round of #533 read exactly " - "that, and in silence", - shape=2), + {"given": "Jane", "family": "Doe", "suffix": "DO DO", "maiden": "Smith"}, + shape=2, + classification="fix(#601)", + ambiguities=("suffix-or-name", "suffix-or-name"), + notes="#601: the clause ends at the clause-free name's trailing run, " + "which the take consumes (rules.md#M2). Before #601: given " + "'Jane', family 'Doe', maiden 'Smith DO DO'. As pinned then: " + "the no-comma spelling of the same decline, where what would do " + "the joining is the OTHER released member: the trailing peel " + "reads both 'DO's as credentials, but the moment they are out " + "of the clause the first is a non-leading particle and chains " + "the second into family 'DO DO'. An earlier round of #533 read " + "exactly that, and in silence"), Case("the_default_vocabularys_own_corpus_mover", "John née Jones Smith Ma", {"family": "John", "maiden": "Jones Smith Ma"}, @@ -5437,16 +5608,18 @@ def _check_cjk_shape_purity(self) -> None: shape=2), Case("a_particle_member_declines_on_its_lean_after_a_comma", "Doe, Jane nee Smith Do", - {"given": "Jane", "family": "Doe", "maiden": "Smith Do"}, - classification="fix(#533)", - ambiguities=("suffix-or-name",), - notes="#531's reading at the given slot, reached through a " - "clause: a member that is also particle vocabulary is " - "the credential on a POSITIVE lean alone, and " - "Title-case inside a mixed-case name is not one. So " - "the clause keeps it and says so. The caps spelling " - "'Doe, Jane nee Smith DO' is the other direction", - shape=2), + {"given": "Jane", "middle": "nee Smith", "family": "Do Doe"}, + shape=2, + classification="fix(#601)", + ambiguities=("particle-or-given",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', maiden 'Smith Do'. As pinned then: #531's reading at " + "the given slot, reached through a clause: a member that is " + "also particle vocabulary is the credential on a POSITIVE lean " + "alone, and Title-case inside a mixed-case name is not one. So " + "the clause keeps it and says so. The caps spelling 'Doe, Jane " + "nee Smith DO' is the other direction"), Case("a_numeral_between_the_clause_and_the_member", "Jane Doe nee Smith V MA", {"given": "Jane", "family": "Doe", "suffix": "MA", @@ -5524,49 +5697,53 @@ def _check_cjk_shape_purity(self) -> None: shape=2), Case("no_trailing_rule_reads_a_third_comma_part", "Smith, John, Jr nee Jones MA", - {"given": "John", "family": "Smith", "suffix": "Jr", - "maiden": "Jones MA"}, - classification="fix(#274)", + {"given": "John", "family": "Smith", "suffix": "Jr nee Jones MA"}, + classification="fix(#601)", ambiguities=("comma-structure",), - notes="the other NONE reader, and the row that killed the " - "first prototype: a segment past the second comma is " - "read as credentials whole, so no trailing rule is " - "consulted, the clause keeps 'MA' -- and nothing " - "reports, because nothing was decided. Untagged: shape " - "2 is a TWO-part listing. Unchanged from 2.0.0", - ), + notes="#601: a marker in a part after a suffix comma is an ordinary " + "word (rules.md#M2). Before #601: given 'John', family 'Smith', " + "suffix 'Jr', maiden 'Jones MA'. As pinned then: the other NONE " + "reader, and the row that killed the first prototype: a segment " + "past the second comma is read as credentials whole, so no " + "trailing rule is consulted, the clause keeps 'MA' -- and " + "nothing reports, because nothing was decided. Untagged: shape " + "2 is a TWO-part listing. Unchanged from 2.0.0"), # ---- #533: the policy sweep, all core-only ----------------------- Case("the_clause_reads_the_same_under_the_strict_comma_knob", "Doe, Jane nee Smith MA", - {"given": "Jane", "family": "Doe", "suffix": "MA", - "maiden": "Smith"}, + {"given": "Jane", "middle": "nee Smith", "family": "Doe", "suffix": "MA"}, policy=Policy(lenient_comma_suffixes=False), - classification="fix(#533)", + classification="fix(#601)", ambiguities=("suffix-or-name",), - notes="the knob governs the LENIENT trailing predicate, " - "which this slot does not inherit, so the reading is " - "the default's"), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', suffix 'MA', maiden 'Smith'. As pinned then: the knob " + "governs the LENIENT trailing predicate, which this slot does " + "not inherit, so the reading is the default's"), Case("the_clause_reads_the_same_under_family_first", "Doe, Jane nee Smith MA", - {"given": "Jane", "family": "Doe", "suffix": "MA", - "maiden": "Smith"}, + {"given": "Jane", "middle": "nee Smith", "family": "Doe", "suffix": "MA"}, policy=Policy(name_order=FAMILY_FIRST), - classification="fix(#533)", - ambiguities=("suffix-or-name",), - notes="name_order does not enter it: the peel is " - "order-independent and the comma has already named the " - "family, so all three orders move the same names " - "(measured over the whole corpus under six policies). " - "UNTAGGED, and a shape 4 tag would be wrong -- this is " - "a comma listing, not the family-first arrangement"), + classification="fix(#601)", + ambiguities=("suffix-or-name",), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', suffix 'MA', maiden 'Smith'. As pinned then: name_order " + "does not enter it: the peel is order-independent and the comma " + "has already named the family, so all three orders move the " + "same names (measured over the whole corpus under six " + "policies). UNTAGGED, and a shape 4 tag would be wrong -- this " + "is a comma listing, not the family-first arrangement"), Case("the_clause_reads_the_same_under_ff_given_last", "Doe, Jane nee Smith MA", - {"given": "Jane", "family": "Doe", "suffix": "MA", - "maiden": "Smith"}, + {"given": "Jane", "middle": "nee Smith", "family": "Doe", "suffix": "MA"}, policy=Policy(name_order=FAMILY_FIRST_GIVEN_LAST), - classification="fix(#533)", + classification="fix(#601)", ambiguities=("suffix-or-name",), - notes="the third order, for the same reason as the row above"), + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', family " + "'Doe', suffix 'MA', maiden 'Smith'. As pinned then: the third " + "order, for the same reason as the row above"), Case("a_leading_title_takes_the_slot_the_member_would_have_had", "Doe, Dr. MA Smith", {"title": "Dr.", "given": "MA", "middle": "Smith", @@ -6211,12 +6388,16 @@ def _check_cjk_shape_purity(self) -> None: "is dropped from consumed tail segments (pinned live " "2026-07-16)"), Case("suffix_delimiter_detection", "Doe, John RN - CRNA", - {"given": "John", "middle": "-", "family": "Doe", - "suffix": "RN, CRNA"}, + {"given": "John", "family": "Doe", "suffix": "RN - CRNA"}, policy=_SD, - notes="the delimiter fires only at suffix sites; the stray " - "token keeps its per-piece walk role (v1 parity, pinned " - "live 2026-07-16)"), + classification="fix(#602)", + notes="the delimiter fires only at suffix sites, so in the " + "given part the dash stays a word -- and since #602 a " + "word after the credential is in the run: one entry " + "'RN - CRNA', the dash unreported, holding no letter or " + "digit (rules.md#A2). Before #602 the stray token kept " + "its per-piece walk role, middle '-', suffix 'RN, CRNA' " + "(v1 parity, pinned live 2026-07-16)"), Case("suffix_delimiter_suffix_comma", "John Smith, RN - CRNA", {"given": "John", "family": "Smith", "suffix": "RN, CRNA"}, policy=_SD, @@ -6283,11 +6464,17 @@ def _check_cjk_shape_purity(self) -> None: "2.3.0 read maiden 'Jones', stepping past the core"), Case("suffix_delimiter_core_that_survives_is_a_boundary_too", "Smith, MD - PhD - FACS", - {"title": "MD", "given": "-", "middle": "-", "family": "Smith", - "suffix": "PhD, FACS"}, + {"title": "MD", "given": "-", "family": "Smith", + "suffix": "PhD - FACS"}, policy=_SD, - classification="fix(#436/#437)", - notes="the other half, and the one a dropped-core test alone " + classification="fix(#602)", + notes="#602 moved this row: the dash after 'PhD' is in the " + "credential run now, one entry 'PhD - FACS'. The parting " + "half it pinned is held by test_post_rules' " + "test_a_surviving_name_word_parts_two_entries on a name " + "word that survives ('Smith, Jane abd Jones PhD'). As " + "pinned under #436/#437: " + "the other half, and the one a dropped-core test alone " "would miss. After a FAMILY comma segment 1 is not a " "tail, so nothing drops the cores and the dashes stand " "as ordinary name words between the two post-nominals. " @@ -6785,26 +6972,32 @@ def _check_cjk_shape_purity(self) -> None: "makes 'van der Jones' a single piece"), Case("bound_given_join_sees_only_the_surviving_name", "abd née Jones Jr Smith Berg", - {"given": "abd", "middle": "Jr Smith", "family": "Berg", - "maiden": "Jones"}, - classification="fix(#418)", - notes="#411's shape, re-pinned twice. The maiden walk stops " - "at the inner suffix and takes only 'Jones'; with the " - "marker pass ahead of the joins, P5 then sees 'abd Jr " - "Smith Berg' and reads it exactly as it reads that " - "name written alone. Under #411 the join declined here " - "because the piece it would absorb was the marker; " - "under #420 the marker was gone before P5 looked and " - "the join took the suffix instead, given 'abd Jr'; " - "since #421 the join declines a suffix piece as it " - "declines a marker, so it takes nothing and 'Jr Smith' " - "is the middle name, as for 'John Jr Smith Berg'"), + {"given": "abd", "maiden": "Jones Jr Smith Berg"}, + classification="fix(#601)", + notes="#601: a credential in the clause starts a run the take " + "consumes (with #602) (rules.md#M2). Before #601: given 'abd', " + "middle 'Jr Smith', family 'Berg', maiden 'Jones'. As pinned " + "then: #411's shape, re-pinned twice. The maiden walk stops at " + "the inner suffix and takes only 'Jones'; with the marker pass " + "ahead of the joins, P5 then sees 'abd Jr Smith Berg' and reads " + "it exactly as it reads that name written alone. Under #411 the " + "join declined here because the piece it would absorb was the " + "marker; under #420 the marker was gone before P5 looked and " + "the join took the suffix instead, given 'abd Jr'; since #421 " + "the join declines a suffix piece as it declines a marker, so " + "it takes nothing and 'Jr Smith' is the middle name, as for " + "'John Jr Smith Berg'"), Case("bound_given_join_takes_a_chain_carrying_a_declined_marker", "Abd van der Berg née Jr Jones", - {"given": "Abd van der Berg née", "middle": "Jr", - "family": "Jones"}, - classification="fix(#417)", - notes="the field-level face of #417. The consumer declines " + {"given": "Abd", "family": "van der Berg née", + "suffix": "Jr Jones"}, + classification="fix(#602)", + ambiguities=("suffix-or-name",), + notes="#602: 'Jr' after the name core starts the run, so P5's " + "reserve sees no name word to spare and the bound word " + "stands alone. Before #602 this pinned " + "the field-level face of #417 (given 'Abd van der Berg " + "née', middle 'Jr', family 'Jones'). The consumer declines " "(a suffix follows the marker), the particle chain " "takes the declined marker as the word M2 says it is, " "and P5 then joins the bound word to the chain -- a " @@ -6831,16 +7024,17 @@ def _check_cjk_shape_purity(self) -> None: "after it, it is just a word)"), Case("bound_given_join_declines_leaving_the_suffix_reading", "Berg, abd née Jones", - {"family": "Berg", "suffix": "abd", "maiden": "Jones"}, - classification="fix(#411)", - notes="'abd' is the one word in both the bound-given and the " - "suffix vocabulary, and P5 says the suffix reading " - "wins in the given slot after a family comma. With the " - "join declining, that reading is what is left -- so " - "the name has no given name at all, matching how " - "'Berg, abd' alone has always parsed. Pinned because " - "it is the shape where the declining join changes most " - "and it reads alarmingly"), + {"given": "abd", "middle": "née Jones", "family": "Berg"}, + classification="fix(#601)", + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: family 'Berg', " + "suffix 'abd', maiden 'Jones'. As pinned then: 'abd' is the one " + "word in both the bound-given and the suffix vocabulary, and P5 " + "says the suffix reading wins in the given slot after a family " + "comma. With the join declining, that reading is what is left " + "-- so the name has no given name at all, matching how 'Berg, " + "abd' alone has always parsed. Pinned because it is the shape " + "where the declining join changes most and it reads alarmingly"), Case("bound_given_marker_immediately_after_the_bound_word", "abd née Jones", {"given": "abd", "maiden": "Jones"}, @@ -6870,19 +7064,19 @@ def _check_cjk_shape_purity(self) -> None: "vocabulary coverage, not a second code path"), Case("bound_given_join_no_longer_swallows_a_marker", "van der Berg, abdul née Jones", - {"given": "abdul", "family": "van der Berg", - "maiden": "Jones"}, - classification="fix(#411)", - notes="was the second of #412's two join-swallows and is now " - "fixed as a side effect of #411, which is why the row " - "is renamed rather than deleted. P5's join used to " - "merge 'abdul' with the marker before M2's bound could " - "see a lone marker piece. #411 made the join decline by " - "counting the reserve without the words the maiden name " - "takes; since #420 the marker pass takes 'née Jones' " - "before P5 looks, leaving 'abdul' alone with nothing to " - "join. P3's connective join was the " - "last to go, under #412 -- " + {"given": "abdul", "middle": "née Jones", "family": "van der Berg"}, + classification="fix(#601)", + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'abdul', " + "family 'van der Berg', maiden 'Jones'. As pinned then: was the " + "second of #412's two join-swallows and is now fixed as a side " + "effect of #411, which is why the row is renamed rather than " + "deleted. P5's join used to merge 'abdul' with the marker " + "before M2's bound could see a lone marker piece. #411 made the " + "join decline by counting the reserve without the words the " + "maiden name takes; since #420 the marker pass takes 'née " + "Jones' before P5 looks, leaving 'abdul' alone with nothing to " + "join. P3's connective join was the last to go, under #412 -- " "connective_join_never_reaches_a_taken_marker"), Case("maiden_marker_ahead_of_a_conjunction", "Jane née and Jones Smith", @@ -6907,15 +7101,16 @@ def _check_cjk_shape_purity(self) -> None: "records the measurement"), Case("maiden_marker_after_particles_in_a_comma_segment", "Smith, Jane van der Berg née Jones", - {"given": "Jane", "middle": "van der Berg", - "family": "Smith", "maiden": "Jones"}, - classification="fix(#399)", - notes="the listing form, where the chain and the marker are " - "both on the given side of the comma. Before #399 the " - "marker and the maiden name stayed in the middle name " - "('van der Berg née Jones'). Distinct from M2's " - "remaining Accepted note, which is about a marker " - "standing straight AFTER the comma"), + {"given": "Jane", "middle": "van der Berg née Jones", "family": "Smith"}, + classification="fix(#601)", + notes="#601: a marker in the given part after a family comma is an " + "ordinary word (rules.md#M2). Before #601: given 'Jane', middle " + "'van der Berg', family 'Smith', maiden 'Jones'. As pinned " + "then: the listing form, where the chain and the marker are " + "both on the given side of the comma. Before #399 the marker " + "and the maiden name stayed in the middle name ('van der Berg " + "née Jones'). Distinct from M2's remaining Accepted note, which " + "is about a marker standing straight AFTER the comma"), Case("bound_given_reserve_excludes_the_maiden_name", "Abd Berg née Jones", {"given": "Abd", "family": "Berg", "maiden": "Jones"}, diff --git a/tests/v2/pipeline/test_group.py b/tests/v2/pipeline/test_group.py index dc22180d..0d939e21 100644 --- a/tests/v2/pipeline/test_group.py +++ b/tests/v2/pipeline/test_group.py @@ -11,13 +11,12 @@ from nameparser._pipeline._classify import classify from nameparser._pipeline._extract import extract_delimited, _maiden_marked from nameparser._pipeline._group import ( - TailReader, _group_segment, _release_reads_off, group, - marker_run_length, + ClauseSite, _group_segment, group, marker_run_length, ) from nameparser._pipeline._script_segment import script_segment from nameparser._pipeline._segment import segment from nameparser._pipeline._state import ( - ParseState, PendingAmbiguity, Structure, WorkToken, + ParseState, Structure, WorkToken, ) from nameparser._pipeline._tokenize import tokenize from nameparser._pipeline._vocab import maiden_marker_run @@ -206,10 +205,13 @@ def test_maiden_marker_consumes_tail() -> None: def test_maiden_marker_stops_at_suffix() -> None: + # the take consumes the trailing run it ends at, roled (#601): the + # PhD leaves the pieces with the clause, already a suffix out = _grouped("Jane Smith née Jones PhD") maiden = [t.text for t in out.tokens if t.role is Role.MAIDEN] assert maiden == ["Jones"] - assert _piece_texts(out)[0][-1] == "PhD" + assert [t.text for t in out.tokens if t.role is Role.SUFFIX] == ["PhD"] + assert _piece_texts(out) == [["Jane", "Smith"]] def test_leading_marker_is_not_consumed() -> None: @@ -451,9 +453,12 @@ def test_where_the_marker_lands_when_the_consumer_declines() -> None: ["Jane", "van der Berg née", "Jr", "Jones"]] # ...and a chain carrying a declined marker is a name piece like # any other to the bound-given join (P5 declines a marker "standing - # as a word of its own", and this one is not) + # as a word of its own", and this one is not) -- but since #602 the + # 'Jr' behind it starts a credential run to the end of the part + # (rules.md#S2), so the reserve finds no name word to spare and the + # join declines for that reason instead assert _piece_texts(_grouped("abdul van der Berg née Jr Jones")) == [ - ["abdul van der Berg née", "Jr", "Jones"]] + ["abdul", "van der Berg née", "Jr", "Jones"]] # consumer declines, no chain: the marker stands as its own piece assert _piece_texts(_grouped("Jane Smith née")) == [ ["Jane", "Smith", "née"]] @@ -473,12 +478,14 @@ def test_the_connective_carveout_counts_the_surviving_name() -> None: # words, so the connective joins assert _piece_texts(_grouped("juan y garcia née")) == [ ["juan y garcia", "née"]] - # the three-piece gate ahead of the count has the same exposure: - # taken on the list as written, 'Jane and née Jones' is four - # pieces, the joins run, and 'and' takes 'Jane' -- the #418 - # empty family one gate earlier. Two pieces remain, so no join. + # a marker straight after a connective is an ordinary word (#601's + # head rule), so 'Jane and née Jones' groups as 'Jane and Zee + # Jones' does. Until #601 the marker was taken there and the #418 + # exposure was the three-piece gate counting it. assert _piece_texts(_grouped("Jane and née Jones")) == [ - ["Jane", "and"]] + ["Jane and née", "Jones"]] + assert _piece_texts(_grouped("Jane and Zee Jones")) == [ + ["Jane and Zee", "Jones"]] _DASH = Policy(extra_suffix_delimiters=frozenset({" - "})) @@ -498,15 +505,13 @@ def test_a_delimiter_core_in_a_suffix_tail_is_not_maiden_text() -> None: assert not any(t.role is Role.MAIDEN for t in twin.tokens) -def test_a_trailing_core_is_cut_before_the_walk() -> None: - # A core standing last is cut off before the walk (#549), so it - # does not make the numeral "not last": the V is the suffix and the - # marker takes 'Jones' alone. (#424's test review found a surviving - # mutant in the walk's core skip; #549 deleted the skip, so the - # site that mutant lived in is gone.) +def test_a_marker_in_a_suffix_part_is_a_word() -> None: + # A marker counts only in the part holding the family name + # (rules.md#M2, #601), so in a tail part it takes nothing, core or + # no core. This was the #549 test that a trailing core is cut + # before the walk, and the walk no longer runs in a tail part. out = _grouped("Smith, John, PhD née Jones V -", policy=_DASH) - assert [t.text for t in out.tokens if t.role is Role.MAIDEN] == [ - "Jones"] + assert not any(t.role is Role.MAIDEN for t in out.tokens) def test_a_core_is_screened_before_the_marker_looks_for_a_word_ahead() -> None: @@ -530,11 +535,12 @@ def test_no_join_reaches_a_taken_marker() -> None: # connective after the marker, behind a chain (#412's headline) assert _piece_texts(_grouped("Jane van der Berg née y Jones")) == [ ["Jane", "van der Berg"]] - # connective before the marker: 'Smith and' is what 'Jane Smith - # and' alone reads too, so the odd-looking family is consistency, - # not an artifact + # connective before the marker: the marker is an ordinary word + # there (#601), and groups as any word would -- 'Jane Smith and + # Zee Jones' reads the same. Until #601 the marker was taken and + # this read 'Jane', 'Smith and'. assert _piece_texts(_grouped("Jane Smith and née Jones")) == [ - ["Jane", "Smith and"]] + ["Jane", "Smith and née", "Jones"]] # marker-headed: the connective is the first word the marker takes assert _piece_texts(_grouped("Jane née and Jones Smith")) == [ ["Jane"]] @@ -939,10 +945,14 @@ def test_the_maiden_walk_stops_before_the_numeral_too() -> None: # suffix", asked with the same test: 'John née Jones Smith V' took # the V into the maiden name. Read by the peel from the marker on, # the V is the suffix, and the walk stops before it. + # Since #601 the take consumes the numeral it ends at, already a + # suffix -- and finds it by the fork's SHAPE test, which asks no + # vocabulary: this module's lexicon lists no 'v' at all. out = _grouped("John née Jones Smith V") assert [t.text for t in out.tokens if t.role is Role.MAIDEN] == \ ["Jones", "Smith"] - assert _piece_texts(out) == [["John", "V"]] + assert [t.text for t in out.tokens if t.role is Role.SUFFIX] == ["V"] + assert _piece_texts(out) == [["John"]] def test_the_maiden_walk_keeps_the_acronym_its_writing_declines( @@ -989,18 +999,13 @@ def test_the_clause_stops_before_a_credential_the_reader_takes( def test_the_clause_never_gives_up_the_first_word_after_the_marker( ) -> None: - """Option 1's floor, as a CLAMP. A member standing alone after the - marker stays the maiden name; where the peel consumed that word - AND words behind it, only the first stays -- a veto that cancelled - the stop outright handed the words behind it back to the clause - too.""" + """The first word after the marker is the maiden name whatever it + is, the marker having announced one. Since #601 that word is never + in the view the trailing run is read over, so it cannot be taken + -- the CLAMP this test pinned, for a peel that consumed it and the + words behind it, has nothing left to clamp.""" out = _grouped("Jane Doe née MA", lexicon=_AMBIGUOUS_LEX) assert _maiden_texts(out) == ["MA"] - out = _grouped("Doe, J. née MA ba", - lexicon=_AMBIGUOUS_LEX.add( - suffix_acronyms={"ba"}, - suffix_acronyms_ambiguous={"ba"})) - assert _maiden_texts(out) == ["MA"] def test_the_clamped_stop_may_land_on_no_member_and_declines() -> None: @@ -1037,12 +1042,13 @@ def test_the_reader_is_none_in_the_family_segment() -> None: assert _suffix_forks(out) == [] -def test_the_reader_is_none_in_a_third_comma_part() -> None: - """A segment past the second comma is read as credentials whole, - so no trailing rule is consulted there either.""" +def test_a_marker_in_a_third_comma_part_is_a_word() -> None: + """A segment past the second comma is a tail part, read as + credentials whole, and a marker there takes nothing (#601). Until + #601 the clause was taken whole there, maiden 'Jones MA'.""" out = _grouped("Smith, John, Jr née Jones MA", lexicon=_AMBIGUOUS_LEX.add(suffix_words={"jr"})) - assert _maiden_texts(out) == ["Jones", "MA"] + assert _maiden_texts(out) == [] assert _suffix_forks(out) == [] @@ -1072,63 +1078,13 @@ def test_the_emitter_reports_a_by_shape_member_too() -> None: assert len(_suffix_forks(out)) == 1 -def test_the_maiden_report_survives_the_family_comma_suppression( -) -> None: - """group() passes `None` for the chain emitter, deliberately -- - the comma fixed the family. The maiden fork is not that fork, so - it travels on its own channel and reports after a comma too.""" - out = _grouped("Doe, Jane née Smith Ma", lexicon=_AMBIGUOUS_LEX) - assert _maiden_texts(out) == ["Smith", "Ma"] - assert len(_suffix_forks(out)) == 1 - - -def test_the_two_ambiguity_channels_route_independently() -> None: - """Both channels are REQUIRED arguments, and they are two so that - silencing one never silences the other. - - `reader` and `maiden_ambiguities` have no defaults: the one - production caller answers both off the segment's structure, and a - default would be this module guessing what that caller knows. The - routing is what the split buys -- the same list in both slots is - one channel, two lists are two, and #533's review found the - earlier spelling defaulting the maiden channel to whatever the - first was, so `ambiguities=None` silenced both. - """ - state = classify(segment(tokenize(extract_delimited(ParseState( - original="Jane Doe née Smith Ma", lexicon=_AMBIGUOUS_LEX, - policy=Policy()))))) - general: list[PendingAmbiguity] = [] - maiden: list[PendingAmbiguity] = [] - _group_segment(state.segments[0], 0, state.tokens, - ambiguities=general, one_case=state.one_case, - reader=TailReader.TRAILING, - maiden_ambiguities=maiden) - assert [a.kind for a in maiden] == [AmbiguityKind.SUFFIX_OR_NAME] - assert general == [] - # the general channel suppressed, the maiden one still speaks -- - # which is exactly what group() does after a family comma - only_maiden: list[PendingAmbiguity] = [] - _group_segment(state.segments[0], 0, state.tokens, - ambiguities=None, one_case=state.one_case, - reader=TailReader.TRAILING, - maiden_ambiguities=only_maiden) - assert [a.kind for a in only_maiden] == [AmbiguityKind.SUFFIX_OR_NAME] - # and NONE is the reader that silences the maiden channel itself, - # because nothing was decided there - silent: list[PendingAmbiguity] = [] - _group_segment(state.segments[0], 0, state.tokens, - ambiguities=None, one_case=state.one_case, - reader=TailReader.NONE, maiden_ambiguities=silent) - assert silent == [] - - -def test_an_unmapped_reader_is_a_loud_failure_rather_than_a_default( +def test_an_unmapped_site_is_a_loud_failure_rather_than_a_default( ) -> None: """The exhaustive dispatch, exercised. - `_maiden_take` ends its reader branch with `assert_never`, which - makes a fourth `TailReader` member a mypy error at this site - rather than a silent fall-through to one of the three readings. + `_maiden_take` ends its site dispatch with `assert_never`, which + makes a fifth `ClauseSite` member a mypy error at this site + rather than a silent fall-through to one of the four readings. At RUNTIME that line is unreachable by construction, so it is reached here the only way it can be -- with a value outside the enum -- both to pin the loudness and to keep the line from being @@ -1140,61 +1096,48 @@ def test_an_unmapped_reader_is_a_loud_failure_rather_than_a_default( with pytest.raises(AssertionError): _group_segment(state.segments[0], 0, state.tokens, ambiguities=[], one_case=state.one_case, - reader=cast(TailReader, 99), - maiden_ambiguities=[]) + site=cast(ClauseSite, 99), + ) -def test_the_release_check_fails_loudly_on_an_unmapped_reader( -) -> None: - """`_release_reads_off` dispatches on the two readers that reach it - and ends with `assert_never`, as `_maiden_take`'s one dispatch - does. Unreachable at runtime by construction, so reached here with - a value outside the enum -- the loudness pinned, and the line kept - from being an uncovered statement.""" - with pytest.raises(AssertionError): - _release_reads_off([[0]], [set()], [], 0, 1, 0, - cast(TailReader, 99), None, # type: ignore[arg-type] - tail_follows=False) - - -def test_the_reader_is_pinned_to_the_structure_it_is_read_from( +def test_the_site_is_pinned_to_the_structure_it_is_read_from( monkeypatch: pytest.MonkeyPatch, ) -> None: - """`TailReader` is a closed set, and group() maps (structure, + """`ClauseSite` is a closed set, and group() maps (structure, segment index) onto it in one place. Pinned here by WATCHING that mapping rather than restating it -- a restatement passes when the code changes under it, which is the shape of vacuous guard AGENTS.md warns about. `_maiden_take` dispatches on the enum - exhaustively (`assert_never`), so a fourth member with no row + exhaustively (`assert_never`), so a fifth member with no row here is a type error rather than a silent default. """ - assert len(TailReader) == 3 + assert len(ClauseSite) == 4 want = { # no comma: the whole name, read by the S2 peel - (Structure.NO_COMMA, 0): TailReader.TRAILING, + (Structure.NO_COMMA, 0): ClauseSite.TRAILING, # suffix comma: segment 0 is the name, the rest is the - # credential run and is read whole - (Structure.SUFFIX_COMMA, 0): TailReader.TRAILING, - (Structure.SUFFIX_COMMA, 1): TailReader.NONE, - (Structure.SUFFIX_COMMA, 2): TailReader.NONE, - # family comma: segment 0 is the family the comma named, - # segment 1 is the given part with its own trailing slot - # (#531), and a third part is a credential run again - (Structure.FAMILY_COMMA, 0): TailReader.NONE, - (Structure.FAMILY_COMMA, 1): TailReader.GIVEN_SLOT, - (Structure.FAMILY_COMMA, 2): TailReader.NONE, + # credential run, where a marker is an ordinary word + (Structure.SUFFIX_COMMA, 0): ClauseSite.TRAILING, + (Structure.SUFFIX_COMMA, 1): ClauseSite.TAIL, + (Structure.SUFFIX_COMMA, 2): ClauseSite.TAIL, + # family comma: segment 0 is the family the comma named, where + # a marker counts; segment 1 is the given part, where it does + # not (#601); and a third part is a credential run again + (Structure.FAMILY_COMMA, 0): ClauseSite.FAMILY_PART, + (Structure.FAMILY_COMMA, 1): ClauseSite.GIVEN_SLOT, + (Structure.FAMILY_COMMA, 2): ClauseSite.TAIL, } texts = ("Jane Doe née Smith MA", "Jane Doe née Smith, MD, PhD", "Doe, Jane née Smith MA, MD") - seen: dict[tuple[Structure, int], TailReader] = {} + seen: dict[tuple[Structure, int], ClauseSite] = {} real = _group_module._group_segment def spy(seg: tuple[int, ...], additional: int, tokens: Sequence[WorkToken], *args: object, **kwargs: object) -> object: seen[(state.structure, len(seen_order))] = cast( - TailReader, kwargs["reader"]) + ClauseSite, kwargs["site"]) seen_order.append(seg) return real(seg, additional, tokens, *args, **kwargs) # type: ignore[arg-type] @@ -1210,13 +1153,16 @@ def spy(seg: tuple[int, ...], additional: int, def test_a_marker_followed_only_by_the_numeral_is_just_a_word() -> None: - # The peel is read from the marker, so 'née V' is two pieces and - # the fork fires on the V: nothing follows the marker but a - # suffix, the pass declines, and the marker stays a word -- as - # for 'Jane Smith née PhD', and as 1.4.0 read it (suffix 'V'). + # A word straight after the marker is the maiden name unless it + # is an unambiguous suffix piece, and a lone 'V' is not one -- it + # is initial-shaped -- so since #601 'Jane Smith née V' reads + # maiden 'V', where the old walk read the numeral fork from the + # marker and declined (suffix 'V', as 1.4.0 read it). An open + # question on #601; 'Jane Smith née PhD' still declines. out = _grouped("Jane Smith née V") + assert [t.text for t in out.tokens if t.role is Role.MAIDEN] == ["V"] + out = _grouped("Jane Smith née PhD") assert [t.text for t in out.tokens if t.role is Role.MAIDEN] == [] - assert _piece_texts(out) == [["Jane", "Smith", "née", "V"]] def test_the_walk_stops_only_where_the_numeral_survives_the_take() -> None: @@ -1231,13 +1177,11 @@ def test_the_walk_stops_only_where_the_numeral_survives_the_take() -> None: out = _grouped("J. née Jones Smith V") assert [t.text for t in out.tokens if t.role is Role.MAIDEN] == \ ["Jones", "Smith", "V"] - # and it is the whole peel that is re-asked, not one condition of - # it: a title before the marker is peeled by assign first, leaving - # the numeral as the whole rest, where no fork fires (the code - # review found the first re-ask handing the V to the given name) + # A title straight before the marker leaves it an ordinary word + # since #601 ('Dr. Zee Jones Smith V' groups the same); until then + # the whole peel was re-asked over 'Dr. V' and the clause kept the V. out = _grouped("Dr. née Jones Smith V") - assert [t.text for t in out.tokens if t.role is Role.MAIDEN] == \ - ["Jones", "Smith", "V"] + assert not any(t.role is Role.MAIDEN for t in out.tokens) def test_an_unlisted_abbreviation_is_as_transparent_as_a_title() -> None: @@ -1726,15 +1670,6 @@ def test_the_clause_link_arm_is_what_moves_it_not_the_vocabulary( assert _maiden_texts(out) == ["Puig", "i", "Soler"] -def test_the_clause_link_survives_a_family_comma() -> None: - # the same walk under the GIVEN_SLOT reader, which is a different - # branch of the take rather than a second member of one shape: - # before the exception the leak landed in the given part. - out = _grouped("Doe, Jane née Puig i Soler", lexicon=_LINK_LEX) - assert _maiden_texts(out) == ["Puig", "i", "Soler"] - assert _piece_texts(out) == [["Doe"], ["Jane"]] - - def test_a_clause_link_runs_twice_over() -> None: # every link of the clause is asked, not just the first: the walk # steps over each one it finds between two birth-name words. @@ -1780,16 +1715,16 @@ def test_an_ambiguous_credential_on_the_right_is_refused_by_bound( def test_a_suffix_word_that_is_no_connective_still_ends_the_clause( ) -> None: - # the recorded negative control for the CLASS half of the - # exception, the shape 'Juan Garcia Lopez y' is for the join: 'jr' - # stands between two birth-name words and is not a connective at - # all, so the exception is never asked and the clause ends at it - # as it always did. Drop the connective conjunct and this row - # reads maiden 'Puig jr Soler' -- measured. Identical under both - # lexicons, the letter deciding nothing here. + # 'jr' stands between two birth-name words and is not a + # connective, so it ends the clause -- and since #601 and #602 it + # starts a credential run in the clause-free name ('Jane Doe jr + # Soler'), which the take consumes, 'Soler' included. Until then + # the clause ended at 'jr' and handed 'jr Soler' to the name. out = _grouped("Jane Doe née Puig jr Soler", lexicon=_LINK_LEX) assert _maiden_texts(out) == ["Puig"] - assert _piece_texts(out) == [["Jane", "Doe", "jr", "Soler"]] + assert [t.text for t in out.tokens if t.role is Role.SUFFIX] == [ + "jr", "Soler"] + assert _piece_texts(out) == [["Jane", "Doe"]] def test_the_marker_is_not_the_name_word_on_the_links_left() -> None: @@ -1805,49 +1740,6 @@ def test_the_marker_is_not_the_name_word_on_the_links_left() -> None: assert _maiden_texts(plain) == ["i", "Soler"] -def test_a_core_straight_after_the_marker_leaves_it_nothing_to_take( -) -> None: - # A delimiter core is TAIL-segment structure that group() cuts the - # segment at before this pass (#549), so it is never a word of the - # clause, and a marker with a core straight behind it ends its part - # with nothing to take, as 'PhD née, i Jones' does. - out = _grouped("Smith, John, PhD née - i Jones", policy=_DASH, - lexicon=_LINK_LEX) - assert _maiden_texts(out) == [] - # and the control that says the CORE is doing it: with no - # delimiter configured the dash is an ordinary word, so the link - # has a name word on its left and the clause keeps the run. - plain = _grouped("Smith, John, PhD née - i Jones", lexicon=_LINK_LEX) - assert _maiden_texts(plain) == ["-", "i", "Jones"] - - -def test_a_core_ends_a_maiden_clause_as_a_comma_would() -> None: - """rules.md's C1: "a delimiter the policy declares parts a trailing - suffix part as a comma would" (#549), so a maiden clause ends at a - core exactly as it ends at the comma written in its place, and a - link beside the core has no name word on that side. - - This reverses #538, which read the clause as the text written - WITHOUT the core: 'Puig - i Soler' kept maiden 'Puig i Soler' - there. RECORDED NEGATIVE CONTROL: at 061f02da the third assertion - read ["Puig", "i", "Soler"]. - """ - out = _grouped("Smith, John, PhD née Puig Mr. - i Soler", - policy=_DASH, lexicon=_LINK_LEX) - assert _maiden_texts(out) == ["Puig", "Mr."] - twin = _grouped("Smith, John, PhD née Puig Mr., i Soler", - lexicon=_LINK_LEX) - assert _maiden_texts(twin) == ["Puig", "Mr."] - # between two NAME words the core ends the clause too, as the - # comma does: the link and the word behind it are the next part's. - between = _grouped("Smith, John, PhD née Puig - i Soler", - policy=_DASH, lexicon=_LINK_LEX) - assert _maiden_texts(between) == ["Puig"] - twin = _grouped("Smith, John, PhD née Puig, i Soler", - lexicon=_LINK_LEX) - assert _maiden_texts(twin) == ["Puig"] - - def test_a_core_beside_a_connective_in_a_credential_tail_is_dropped( ) -> None: """No connective joins across a core, generational or ordinary, and @@ -1994,42 +1886,6 @@ def test_a_frozen_link_is_still_absorbed_by_a_neighbours_join() -> None: assert out.suffix == "" -def test_the_title_stop_needs_a_name_word_left_standing() -> None: - """rules.md#M2 (#535): the title stop is asked over the name the - take would leave. 'Dr. nee Jones Smith Prof.' would leave 'Dr. - Prof.', where H5's chain has no name word to stand behind and the - title would read as the family name -- so the clause keeps it.""" - out = _grouped("Dr. nee Jones Smith Prof.", lexicon=Lexicon.default()) - assert _maiden_texts(out) == ["Jones", "Smith", "Prof."] - # the control: with a name word ahead of the marker the same clause - # gives the title up - ok = _grouped("Jane Doe nee Jones Smith Prof.", lexicon=Lexicon.default()) - assert _maiden_texts(ok) == ["Jones", "Smith"] - - -def test_a_released_title_behind_a_particle_is_withdrawn() -> None: - """P2's chain runs on over a trailing title (rules.md#H5 Accepted), - so a title the clause gives up with a particle ahead of it in the - remaining name would join the family; the release is withdrawn.""" - out = _grouped("Jane van der Berg nee Smith Prof.", lexicon=Lexicon.default()) - assert _maiden_texts(out) == ["Smith", "Prof."] - out = _grouped("Jane Doe nee Smith MA do Prof.", lexicon=Lexicon.default()) - assert _maiden_texts(out) == ["Smith", "MA", "do"] - - -def test_the_numeral_stop_asks_the_join_question() -> None: - """The numeral fork never asked whether a join below the take - would absorb the word it releases: after a family comma the - bound-given join took the V ('Berg, abdul nee Smith V' read given - 'abdul V' at 2.2.0 and 2.3.0). It asks now, through the shared - check.""" - out = _grouped("Berg, abdul nee Smith V", lexicon=Lexicon.default()) - assert _maiden_texts(out) == ["Smith", "V"] - # the control: with no bound word the numeral is released as before - ok = _grouped("Berg, Jane nee Smith V", lexicon=Lexicon.default()) - assert _maiden_texts(ok) == ["Smith"] - - def test_the_first_word_floor_holds_a_title_out_of_the_chain() -> None: """A title straight after the marker stays the maiden name: the floor is the title chain's own, so the chain never takes that word @@ -2051,98 +1907,3 @@ def test_the_floor_keeps_the_first_word_in_the_chains_count() -> None: assert _maiden_texts(ok) == ["Smith"] -def test_a_given_slot_numeral_with_a_credential_tail_stays() -> None: - """rules.md#M2/#144 (#535 review): after a family comma the given - slot reads a lone numeral as a suffix only where the given part is - the LAST comma part -- a third comma part behind it withdraws the - release, so 'Doe, Jane nee Smith V, PhD' keeps the V where - 'Doe, Jane nee Smith V' (no tail) gives it up.""" - out = _grouped("Doe, Jane nee Smith V, PhD", lexicon=Lexicon.default()) - assert _maiden_texts(out) == ["Smith", "V"] - ok = _grouped("Doe, Jane nee Smith V", lexicon=Lexicon.default()) - assert _maiden_texts(ok) == ["Smith"] - - -def test_the_bound_given_half_of_the_join_model_is_the_given_slots_alone() -> None: - """rules.md#P5 (#535 review): the LENIENT bound-given join only - applies after a family comma; before one the STRICT reserve - declines a join that would change a suffix reading, so a numeral - the peel already reads as a suffix stays given up -- 'abdul nee - Smith V' (no comma) keeps the release where 'Berg, abdul nee - Smith V' (after a comma) withdraws it.""" - out = _grouped("abdul nee Smith V", lexicon=Lexicon.default()) - assert _maiden_texts(out) == ["Smith"] - ok = _grouped("Berg, abdul nee Smith V", lexicon=Lexicon.default()) - assert _maiden_texts(ok) == ["Smith", "V"] - - -def test_the_given_title_chain_stops_at_a_member_with_a_title_behind() -> None: - """rules.md#M2 after a family comma (#535): the given part's title - chain is read from the end over what its first suffix pass leaves, - and that pass takes a class member only where every piece behind - it is taken too -- so 'MA' with 'Prof.' behind it is a name word - there and the chain stops at it ('Doe, Jane Dr. MA Prof.' reads - middle 'Dr.'). A clause that released 'Rev.' as a title would put - it in the middle name; it keeps it instead, and only the title - behind the member and the member itself leave. And the lenient - trailing numeral (#144) counts as that pass's suffix once the - numeral stop has asked it, so 'Prof. V' leaves the clause whole, - as 'Doe, Jane Prof. V' reads.""" - rev = _grouped("Doe, Jane nee Smith Rev. MA Prof.", - lexicon=Lexicon.default()) - assert _maiden_texts(rev) == ["Smith", "Rev."] - num = _grouped("Doe, Jane nee Smith Prof. V", lexicon=Lexicon.default()) - assert _maiden_texts(num) == ["Smith"] - - -def test_a_link_the_walk_stops_at_gives_up_only_a_run_that_reads_off( -) -> None: - """rules.md#M2 (#535): with the title chained, the link exception - can refuse a link it used to join, and a stop at a link gives up - the words behind it. 'i DO Prof.' left standing behind 'Jane Doe' - does not read as post-nominals, so the clause keeps the link and - the DO ('Doe i' was the middle name without the check); 'i MA - Prof.' does, so there the link and the credential leave.""" - kept = _grouped("Jane Doe nee Smith i DO Prof.", - lexicon=Lexicon.default()) - assert _maiden_texts(kept) == ["Smith", "i", "DO"] - given_up = _grouped("Jane Doe nee Smith i MA Prof.", - lexicon=Lexicon.default()) - assert _maiden_texts(given_up) == ["Smith"] - - -def test_only_a_link_the_title_chain_refused_asks_the_release_question( -) -> None: - """The link check is asked only where reading the title chain made - the link exception refuse a link it joined over the words as - written. A link refused either way stops as it always did: 'Doe, - Jane nee Smith i V' gives up 'i V' (a check that fired here kept a - dangling 'i' in the birth name). And the given part's lenient - numeral counts as a suffix where only chained titles stand behind - it, as bare 'Doe, Jane i V Prof.' reads suffix 'i V' -- but not - where another comma part follows, where it is a middle initial.""" - plain = _grouped("Doe, Jane nee Smith i V", lexicon=Lexicon.default()) - assert _maiden_texts(plain) == ["Smith"] - titled = _grouped("Doe, Jane nee Smith i V Prof.", - lexicon=Lexicon.default()) - assert _maiden_texts(titled) == ["Smith"] - tail = _grouped("Doe, Jane nee Smith i V Prof., PhD", - lexicon=Lexicon.default()) - assert _maiden_texts(tail) == ["Smith", "i", "V"] - - -def test_a_released_particle_title_is_kept_after_a_family_comma() -> None: - """rules.md#M2 with P6 (#535): after a family comma a released - title that is also a particle would be attached to the family by - P6 ('Doe, Jane St.' reads family 'St. Doe'), so the clause keeps - it. A title that is no particle still leaves, and so does the - credential DO, which the given slot's own lean reads (#533); with - no comma P6 does not run and 'St.' leaves as a title.""" - kept = _grouped("Doe, Jane nee Smith St.", lexicon=Lexicon.default()) - assert _maiden_texts(kept) == ["Smith", "St."] - title = _grouped("Doe, Jane nee Smith Prof.", lexicon=Lexicon.default()) - assert _maiden_texts(title) == ["Smith"] - credential = _grouped("Doe, Jane nee Smith DO", lexicon=Lexicon.default()) - assert _maiden_texts(credential) == ["Smith"] - no_comma = _grouped("Jane Doe nee Smith St.", lexicon=Lexicon.default()) - assert _maiden_texts(no_comma) == ["Smith"] diff --git a/tests/v2/pipeline/test_pieces.py b/tests/v2/pipeline/test_pieces.py index 306447e9..8e8eb247 100644 --- a/tests/v2/pipeline/test_pieces.py +++ b/tests/v2/pipeline/test_pieces.py @@ -551,8 +551,12 @@ def test_a_multi_letter_numeral_anchors_like_any_suffix_word() -> None: # anchors exactly as 'Jr' does. n = parse("John Smith PhD III Ma") assert (n.family, n.suffix) == ("Smith", "PhD III Ma") - n = parse("John Smith PhD V Ma") - assert (n.middle, n.family, n.suffix) == ("Smith PhD V", "Ma", "") + # without the degree in front: the single letter is a middle + # initial and anchors nothing. With it, since #602, the degree + # starts a run to the end of the part and 'V Ma' are inside it + # ('John Smith PhD V Ma' reads suffix 'PhD V Ma'). + n = parse("John Smith V Ma") + assert (n.middle, n.family, n.suffix) == ("Smith V", "Ma", "") def test_the_anchor_pass_starts_past_the_leading_title_run() -> None: @@ -838,3 +842,47 @@ def test_a_title_the_fixed_point_splices_out_keeps_no_peel_pick() -> None: assert (name.title, name.given, name.family, name.suffix) == ( "Ma.", "John", "Smith", "PhD Jr.") assert name.ambiguities == () + + +# rules.md#S2 (#602): a credential after the name core starts a run to +# the end of its part. The exclusion tests pass before the change too: +# they are the negative controls that make the first two meaningful. +def test_a_credential_after_two_name_words_starts_the_run() -> None: + name = parse("John Smith PhD Jones") + assert (name.given, name.family, name.suffix, name.middle) == ( + "John", "Smith", "PhD Jones", "") + + +def test_a_title_word_inside_the_run_reads_as_a_title() -> None: + name = parse("Eric H. Holder Jr. Attorney General") + assert (name.title, name.suffix, name.family) == ( + "Attorney General", "Jr.", "Holder") + + +def test_a_title_word_never_starts_the_run() -> None: + # 746 title words are in no suffix set, many of them surnames + assert parse("Mary Jane King Smith").family == "Smith" + + +def test_the_ambiguous_class_and_initials_do_not_start_the_run() -> None: + assert parse("John Smith MA Jones").family == "Jones" + assert parse("John Smith V Jones").family == "Jones" + + +def test_a_joined_connective_does_not_start_the_run() -> None: + assert parse("Josep Carod i Rovira").suffix == "" + + +def test_one_name_word_before_the_credential_is_no_core() -> None: + assert parse("John PhD Smith").family == "Smith" + + +def test_the_given_part_after_a_family_comma_reads_the_run_too() -> None: + name = parse("Smith, John PhD Jones") + assert (name.given, name.family, name.suffix, name.middle) == ( + "John", "Smith", "PhD Jones", "") + + +def test_a_title_inside_the_given_parts_run_is_a_title() -> None: + name = parse("Holder, Eric Jr. Attorney General") + assert (name.title, name.suffix) == ("Attorney General", "Jr.") diff --git a/tests/v2/pipeline/test_post_rules.py b/tests/v2/pipeline/test_post_rules.py index d3abeb34..c2e4077b 100644 --- a/tests/v2/pipeline/test_post_rules.py +++ b/tests/v2/pipeline/test_post_rules.py @@ -28,11 +28,16 @@ given_name_titles=frozenset({"sir"}), particles=frozenset({"de", "der", "ibn", "la", "van", "vd"}), particles_ambiguous=frozenset({"la", "van"}), - suffix_words=frozenset({"jr"}), + # `씨` is a suffix word that starts no credential run (rules.md#S2, + # #602: an honorific word never does), the filler a test needs + # where a credential would now take the rest of the part + suffix_words=frozenset({"jr", "씨"}), # `vd` mirrors its shipped dual membership -- particle AND # UNAMBIGUOUS suffix acronym -- which is what puts a name on P6's - # suffix arm below. - suffix_acronyms=frozenset({"md", "vd"}), + # suffix arm below. `ma` mirrors the ambiguous class, which starts + # no run either. + suffix_acronyms=frozenset({"md", "vd", "ma"}), + suffix_acronyms_ambiguous=frozenset({"ma"}), conjunctions=frozenset({"y"}), bound_given_names=frozenset({"abdul"}), ) @@ -499,11 +504,15 @@ def test_no_leftover_is_re_laid_out() -> None: # out. The whole role map is asserted because that draft failed in # two directions: it lost a given name outright on "van Berg Jan de" # and promoted a post-nominal into the given slot here. A test on - # `GIVEN != "Jr."` alone passes on an EMPTY given, which is the + # `GIVEN != "Ma"` alone passes on an EMPTY given, which is the # first of those. - out = _parsed("Berg Jan Jr. de", Policy(name_order=FAMILY_FIRST)) + # + # The post-nominal is the ambiguous `Ma`, which stays a middle + # name. It was `Jr.` until #602, which made a credential after the + # name core start a run to the end of the part, taking `de` in. + out = _parsed("Berg Jan Ma de", Policy(name_order=FAMILY_FIRST)) assert _by_role(out, Role.GIVEN) == "Jan" - assert _by_role(out, Role.MIDDLE) == "Jr." + assert _by_role(out, Role.MIDDLE) == "Ma" assert _by_role(out, Role.FAMILY) == "Berg de" assert _folded(out) == "de" @@ -810,8 +819,8 @@ def test_middle_as_family_folds_a_middle_the_old_reach_never_left() -> None: @pytest.mark.parametrize("policy,given,middle", [ - (Policy(name_order=FAMILY_FIRST), "van Berg", "MD Juan"), - (Policy(name_order=FAMILY_FIRST_GIVEN_LAST), "Juan", "van Berg MD"), + (Policy(name_order=FAMILY_FIRST), "van Berg", "씨 Juan"), + (Policy(name_order=FAMILY_FIRST_GIVEN_LAST), "Juan", "van Berg 씨"), ]) def test_a_suffix_word_mid_run_ends_the_chain( policy: Policy, given: str, middle: str) -> None: @@ -819,8 +828,10 @@ def test_a_suffix_word_mid_run_ends_the_chain( # stop _group's own chain uses (`not prefix(j) and not suffix(j)`). # Without it the whole leftover is ONE unit and lands in one field. # Three leftover units here, so this is also the only place the two - # orders are pinned apart at a count other than two. - out = _parsed("de Mesnil van Berg MD Juan", policy) + # orders are pinned apart at a count other than two. The suffix word + # is an honorific because a credential here would start #602's run + # and take `Juan` with it (rules.md#S2); it was `MD` until then. + out = _parsed("de Mesnil van Berg 씨 Juan", policy) assert _by_role(out, Role.FAMILY) == "de Mesnil" assert _by_role(out, Role.GIVEN) == given assert _by_role(out, Role.MIDDLE) == middle @@ -1006,9 +1017,16 @@ def test_a_bare_marker_maiden_clause_does_not_part_the_run() -> None: this renders 'MD, PhD'. The bracketed twin above drops its marker too, so that mutant fails both; what separates the two tests is the role test, which sees 'Jones' in both and 'nee' in neither - (measured 2026-09-06).""" - assert _entry_tags("Smith, MD nee Jones PhD") == [ - ("MD", False), ("PhD", True)] + (measured 2026-09-06). + + Since #601 a bare marker counts only behind a name word, so the + clause stands between two post-nominals only inside #602's + credential run: 'Jr.' starts it and takes 'Smith' in, the marker + standing after that name word. The fixture was 'Smith, MD nee Jones + PhD' until then, a marker in the given part after a family comma, + which is an ordinary word now.""" + assert _entry_tags("Jane Doe Jr. Smith nee Jones PhD") == [ + ("Jr.", False), ("Smith", True), ("PhD", True)] def test_a_marker_the_policy_names_as_a_delimiter_parts_the_run() -> None: @@ -1020,10 +1038,14 @@ def test_a_marker_the_policy_names_as_a_delimiter_parts_the_run() -> None: the reachable case, and it parts by the policy's own declaration (measured 2026-09-06).""" nee = Policy(extra_suffix_delimiters=frozenset({" nee "})) - assert _entry_tags("John Doe, MD nee Jones PhD", nee) == [ - ("MD", False), ("PhD", False)] - assert _entry_tags("John Doe, MD nee Jones PhD") == [ - ("MD", False), ("PhD", True)] + # the fixture puts the clause inside #602's credential run, the one + # place a bare clause stands between two post-nominals since #601 + # (the bare test above says why); it was 'John Doe, MD nee Jones + # PhD' until then + assert _entry_tags("Jane Doe Jr. Smith nee Jones PhD", nee) == [ + ("Jr.", False), ("Smith", True), ("PhD", False)] + assert _entry_tags("Jane Doe Jr. Smith nee Jones PhD") == [ + ("Jr.", False), ("Smith", True), ("PhD", True)] def test_a_dropped_core_parts_two_entries() -> None: @@ -1040,15 +1062,19 @@ def test_a_dropped_core_parts_two_entries() -> None: def test_a_surviving_name_word_parts_two_entries() -> None: """The `parted` conjunct, kept-core half -- the one a dropped-core - test alone cannot reach. After a FAMILY comma segment 1 is no - tail, so nothing drops the dashes and they stand as name words - between the post-nominals. Drop the conjunct and this renders - 'PhD FACS' under EVERY policy, the bare one included, where the - dash is no delimiter at all.""" + test alone cannot reach: a name word standing between two + post-nominals parts their entries. `abd` is suffix vocabulary that + starts no credential run (it is bound given-name vocabulary too, + rules.md#S2), so `Jones` behind it survives as a name word before + the credential. Drop the conjunct and this renders 'abd PhD'. + + The fixture was 'Smith, MD - PhD - FACS' until #602, whose dashes + stood as name words between the post-nominals; a credential now + starts a run to the end of the part, taking them in.""" for policy in (Policy(), Policy(extra_suffix_delimiters=frozenset({" - "}))): - assert _entry_tags("Smith, MD - PhD - FACS", policy) == [ - ("PhD", False), ("FACS", False)] + assert _entry_tags("Smith, Jane abd Jones PhD", policy) == [ + ("abd", False), ("PhD", False)] def test_the_pass_runs_after_roles_are_settled() -> None: diff --git a/tests/v2/test_benchmark.py b/tests/v2/test_benchmark.py index ac3ca0c2..f68c61a1 100644 --- a/tests/v2/test_benchmark.py +++ b/tests/v2/test_benchmark.py @@ -218,13 +218,13 @@ def test_a_thousand_names_still_parse_in_reasonable_time( # first piece is a bound given-name word, # so the reserve's per-piece question is # asked over the whole run (#401) -# M2 clause VIEW maiden_clause ONLY -- the one unit that -# reaches the #533 acronym fork, which -# needs a maiden marker AND a class member -# ending the string. 'MA nee ' has both -# words and reaches nothing: the peel -# stops at the trailing marker, so the -# ORDER inside the unit is the shape +# M2 clause VIEW maiden_clause ONLY -- the one unit +# holding a maiden marker, so the take's +# clause-free view over a run of class +# members (#601). Until #601 it reached +# the #533 acronym fork, which only this +# order did; since then 'MA nee ' reaches +# the take too (measured 2026-10-04) # connective RUN LENGTH link_run ONLY -- P3's both-sides # condition walks the run of connectives # beside a link, and no other unit here @@ -639,7 +639,8 @@ def test_a_trailing_credential_run_does_not_cost_exponentially() -> None: # same reason the one above needed a second: `_SHAPES` repeats a unit # and nothing else, so it cannot express "a name word, a marker, then # a long run" -- and a maiden clause needs exactly that prefix. -# Measured: `"nee i Und " * n` never reaches the clause link exception +# Measured: `"nee i Und " * n` never reached the clause link exception +# (retired by #601; `_clause_run` says what the guard measures now) # at all (the take declines, and b9ed1429 and this tree measure the # identical 3.92/4.07/4.12 on it), which is the silent-no-op a # reachability probe exists to catch. So this guard builds its own @@ -674,23 +675,31 @@ def _clause_run(members: int) -> str: """A maiden clause whose birth name is a RUN of links. Mixed case on purpose: written wholly in one case the letter reads - as an initial (rules.md#P3's marked subset) and the clause's link - exception is never asked, so the guard would measure nothing. + as an initial (rules.md#P3's marked subset). Since #601 the clause + keeps the run because nothing in it ends the clause, so this + measures the take's walk over a long clause; until then it measured + the clause's link exception, which #601 retired. """ return "Jane Doe nee Puig " + "i " * members + "Soler" +def _name_link_run(members: int) -> str: + """A NAME whose family is a run of links, the shape P3's frozen + loop asks `_group._between_name_words` about once per link. Mixed + case on purpose, for the reason `_clause_run` gives.""" + return "Jane Puig " + "i " * members + "Soler" + + def test_a_clause_link_run_does_not_cost_quadratically() -> None: if sys.getprofile() is not None: pytest.skip("a profile hook is already installed; this test owns it") small_text = _clause_run(_CLAUSE_RUN_SMALL) large_text = _clause_run(_CLAUSE_RUN_LARGE) # REACHABILITY, the probe every shape in this file carries: the - # walk under measurement runs only while the clause KEEPS the run, - # which is rules.md#M2's link exception. End the clause at the - # first link instead and the guard measures a walk that no longer - # happens, at a comfortable ratio, forever. Asked at both sizes, - # the run length being what this varies. + # take walks the clause only while the clause KEEPS the run. End + # the clause at the first link instead and the guard measures a + # walk that no longer happens, at a comfortable ratio, forever. + # Asked at both sizes, the run length being what this varies. for text, members in ((small_text, _CLAUSE_RUN_SMALL), (large_text, _CLAUSE_RUN_LARGE)): assert parse(text).maiden == " ".join( @@ -762,9 +771,10 @@ def test_the_paired_initials_title_scan_does_not_cost_quadratically() -> None: # tail #558: `tail_reading` re-peeled the whole walk once per # title the H5 chain took, asking `listed_lean` of every # member it had already peeled. -# clause #558 again, through the maiden walk, which runs that fixed -# point three times (the clause's take, its release check, -# and `trailing_start_past_titles`). +# clause #558 again, through the maiden take, which reads that +# fixed point over the clause-free view (until #601 three +# times: the take, its release check, and +# `trailing_start_past_titles`). # chain #559: the particle chain asked "is every piece ahead of # this one a title?" afresh at every chain site, so leading # titles x particle sites `is_leading_title` calls. @@ -871,6 +881,16 @@ def test_a_fixed_point_does_not_reread_what_it_has_read( # `_CALL_BASELINE` uses, and a bespoke one here would need its own # argument every time the shape moved. # +# RE-POINTED 2026-10-04 (#601), and why: the pin measured a link run +# in a MAIDEN CLAUSE, whose link exception asked `_between_name_words` +# once per link. #601 retired the exception -- the clause keeps its +# links because nothing in it ends the clause -- and the clause's +# links leave with it before P3's frozen loop runs, so that shape +# reached the function 0 times (measured). The same run written in +# the NAME reaches it 64 times, one per link, and a per-link +# regression moves it as it moved the clause: +64 against a 2% band +# of 53.6 at the 2,678 frames recorded below. +# # ONE ROW, py3.11, and an unknown interpreter SKIPS rather than fails # -- which is where this parts company with `_check_budget`, whose # table carries every interpreter CI runs and so can afford to fail on @@ -882,9 +902,10 @@ def test_a_fixed_point_does_not_reread_what_it_has_read( # interpreter and add the row. # # Lowered 2026-09-26 from 2587 with `_CALL_BASELINE` above, for the -# same reason (decisions.md#parse-cost). +# same reason (decisions.md#parse-cost). 2,678 from 2026-10-04, the +# re-pointed name-level shape (above), measured on py3.11. _LINK_BASELINE = { - (3, 11): 2301, + (3, 11): 2678, } #: The same +-2% `_CALL_BASELINE` uses, and for the same reason: frame #: counts are deterministic for a given tree and interpreter, so the @@ -906,21 +927,21 @@ def test_a_link_costs_what_it_is_pinned_at() -> None: if version not in _LINK_BASELINE: pytest.skip( f"no link baseline for Python {version[0]}.{version[1]}; " - f"measure `_clause_run({_CLAUSE_RUN_LARGE})` on this " + f"measure `_name_link_run({_CLAUSE_RUN_LARGE})` on this " f"interpreter and add the row to _LINK_BASELINE") - text = _clause_run(_CLAUSE_RUN_LARGE) + text = _name_link_run(_CLAUSE_RUN_LARGE) # REACHABILITY, the probe every shape in this file carries: the - # frames counted are the clause walk's, and they are only there - # while the clause KEEPS the run (rules.md#M2's link exception). - # End the clause at the first link and this measures a name that - # no longer holds 64 links, comfortably inside the band forever. - assert parse(text).maiden == " ".join( + # frames counted are P3's link test's, and they are only there + # while every link joins the family. A link that stopped joining + # would leave a name that no longer holds 64 links in one run, + # comfortably inside the band forever. + assert parse(text).family == " ".join( ["Puig"] + ["i"] * _CLAUSE_RUN_LARGE + ["Soler"]) baseline = _LINK_BASELINE[version] actual = _frames_for(text) low, high = baseline * (1 - _LINK_BAND), baseline * (1 + _LINK_BAND) assert low <= actual <= high, ( - f"a maiden clause holding {_CLAUSE_RUN_LARGE} links costs " + f"a name holding {_CLAUSE_RUN_LARGE} links costs " f"{actual} frames on Python {version[0]}.{version[1]}, band " f"{low:.0f}-{high:.0f} around a baseline of {baseline}. Growth " f"and shrinkage are both signals, and one frame per link is " diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index ae5c0698..694258df 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -850,8 +850,6 @@ def test_case_shape_ids_exist_in_the_inventory() -> None: ("田中さん, Jr. Ph. D.", "Smith, Jr. Ph. D. MD"), "fix(#411)": ("abdul née Jones", "Smith, abdul Rahman", "Jane Smith ABD", "van der Berg, abd née Jones"), - "fix(#411/S2)": ("van der Berg, abdul née Jones", "Berg, abd Rahman", - "Jane Smith ABD"), "fix(#367) a title no longer displaces a leading never-given particle": ("John Sir de Mesnil", "Smith, Sir de Vaux", "Sir Smith"), "fix(#272/#308)": ("田中·太郎 김씨", "A·B 씨", "山田·花子"), @@ -1498,14 +1496,6 @@ def test_case_shape_ids_exist_in_the_inventory() -> None: "fix(#274/#533) the clause keeps the credential a join or a missing reader would have taken": ("Berg, abdul MA", "Doe, Dr. MA", "Berg, Jane van der DO", "Doe, Jane MA do", "Jane Doe nee Smith MA"), - # Same, at the 2.x baselines where the decline shows as a report - # rather than as a role move. - "fix(#411/#533) the bound-given pair after a comma": - ("Berg, abdul MA", "Berg, abdul nee Jones Ma", - "Berg, abdul nee Jones"), - "fix(#399/#533) a maiden marker bounds the particle chain, and the clause keeps": - ("Berg, Jane van der DO", "Berg, Jane van der nee Smith Do", - "Berg, Jane van der nee Smith"), # The delimited rule must not reach the undelimited spellings, # which is the whole of what it says: the boundary is the writer's. "fix(#335/#533) a marker-led bracketed clause is the maiden name whatever its last word is": @@ -1536,11 +1526,8 @@ def test_case_shape_ids_exist_in_the_inventory() -> None: "fix(#533) the clause gives the credential up, and the run it joins renders as the writer spaced it": ("Jane Doe nee Smith PhD MA", "Jane Doe nee Smith MA", "Jane Doe nee Smith Ma", "Doe, J. nee MA", "Doe, J. ba"), - # The compound rules are case-SENSITIVE by construction, so each + # The compound rule is case-SENSITIVE by construction, so it # takes the other spellings of its own words as probes. - "fix(#445/#533) the clause gives the credential up, and the one name word it leaves is the family": - ("John née Jones Smith Ma", "JOHN NEE JONES SMITH MA PHD", - "John née Jones Smith", "Jane Doe nee Smith MA"), "fix(#445/#533) a one-case clause keeps the credential, and the one name word it leaves is the family": ("John nee Jones Smith MA PHD", "John née Jones Smith Ma", "JOHN NEE JONES SMITH MA", "John née Jones Smith MA"), @@ -1582,34 +1569,10 @@ def test_case_shape_ids_exist_in_the_inventory() -> None: # relabeled -- same probes, same reason. "fix(#274/#535) a trailing title ends the maiden clause": ("Dr. Jane Doe nee Smith Prof.", "Jane Doe nee Smith Prof. Dr."), - # The lone-marker shape's own probe, at 1.4.0: a name word standing - # ahead of the marker or behind the title must not match either -- - # 'Dr. nee Jones Smith Prof.' is the boundary where only a TITLE - # precedes the marker, so no name word is left standing there. - "fix(#274) a trailing title after the marker with nothing else in the name": - ("Jane Dr. nee Jones Smith Prof.", "Dr. nee Jones Smith Prof. Jr."), # The 1.4.0-only Dr.-postnominal rule's own probe: a name word # standing ahead of the marker or behind the title must not match. "fix(#274/#296/#535) a trailing title after a dropped postnominal Dr., with no maiden reading either": ("Dr. Jane Doe nee Prof. Dr.", "Jane Doe nee Prof. Dr. Jr."), - # The numeral rule's own probe: a name word ahead of the released - # numeral must not match either. Used at 2.2.0/2.3.0, where the - # rule is #535-only. - "fix(#535) the numeral stop asks the join question": - ("Berg, Jane nee Smith V", "Dr. Berg, abdul nee Smith V"), - # This baseline's copy, shared with 1.4.0 (both read 'Berg, abdul - # nee Smith V' identically -- #274 moves nothing here, so 1.4.0 - # carries the same two-issue label rather than a third): the - # bound-given reserve (#411) already decides the shape before - # #535 exists, so the issue is relabeled -- same probes, same - # reason. - "fix(#411/#535) the numeral stop asks the join question": - ("Berg, Jane nee Smith V", "Dr. Berg, abdul nee Smith V"), - # The #399 sibling's own probe, at 2.0.0/2.1.0: a name word - # standing ahead of the marker or behind the title must not match. - "fix(#399) a maiden marker bounds the particle chain that swallowed it, two trailing words": - ("Dr. Jane van der Berg nee Smith Prof.", - "Jane van der Berg nee Smith Prof. Jr."), # The dropped-Dr.-postnominal rule's own probe, at 2.0.0/2.1.0. "fix(#296/#535) a trailing title after a dropped postnominal Dr.": ("Dr. Jane Doe nee Prof. Dr.", "Jane Doe nee Prof. Dr. Jr."), @@ -1617,10 +1580,9 @@ def test_case_shape_ids_exist_in_the_inventory() -> None: # 2.0.0-2.3.0. "fix(#533/#535) a title in front of the credential the clause gives up": ("Dr. Jane Doe nee Smith Prof. MA", "Jane Doe nee Smith Prof. MA Jr."), - # The given-slot-numeral-with-a-tail rule's own probe, at 2.2.0 - # and 2.3.0. - "fix(#535) a given-slot numeral with a credential tail stays": - ("Dr. Doe, Jane nee Smith V, PhD", "Doe, Jane nee Smith V, PhD Jr."), + # (Nine probes left on 2026-10-04 with the rules they pinned, which + # #601 deleted: the lone-marker, numeral-stop, #399-sibling, + # given-slot-numeral, bound-given-after-a-comma and #411/S2 rules.) } @@ -2314,9 +2276,6 @@ class _LatinCopy(NamedTuple): # no reader ahead of the member, or a join below the marker pass # that would take it -- and no vocabulary decides that. The # delimited trio is the same, keyed on the writer's brackets. - frozenset({"Berg, Jane van der nee Smith DO", "Berg, abdul nee Jones MA", - "Doe, Dr\\. nee Smith MA", "Doe, Jane nee Smith Do", - "Jane Doe nee MA PhD"}), frozenset({"Jane Doe \\(nee Smith MA\\)", "Jane Doe \\(nee Smith Ma\\)", "Jane Doe \\(nee Smith\\) MA"}), # rules.md#M2's two delimiter examples (added with #538, read by @@ -2331,9 +2290,6 @@ class _LatinCopy(NamedTuple): # name. Names the doc chose, not a copy of the particle wordlist, # so there is nothing for the alternation to drift from. frozenset({"Juan Ó Pérez", "SEÁN Ó MURCHÚ", "ina binti navalamar"}), - frozenset({"Smith, John, PhD née Puig Mr\\. - i Soler", - "Smith, John, PhD née Puig - i Soler", - "Smith, John, MD - née Jones Smith"}), # fix(#400)'s two openings: start-of-name or just after a family # comma. `abd` joins forward on the given side wherever that side # begins, and the alternation is over ANCHORS, not over words -- @@ -2370,42 +2326,6 @@ class _LatinCopy(NamedTuple): # Prof.', a rules.md#M2 Accepted example: at every earlier # baseline its diff has #342's cause too ('ba' was a plain suffix # word there) and it sits in a joint-labelled rule instead. - frozenset({"Doe, Jane nee Smith MA Prof\\.", - "Jane Doe \\x28nee Smith Prof\\.\\x29", - "Jane Doe nee Smith King\\.", - "Jane Doe nee Smith MA Prof\\.", - "Jane Doe nee Smith Ma Prof\\.", - "Jane Doe nee Smith Prof\\.", - "Jane Doe nee Smith Prof\\. MA", - "Jane Doe nee Smith V Prof\\.", - "Mary Smith née Jones Prof\\."}), - frozenset({"Doe, Jane nee Smith MA Prof\\.", - "Jane Doe \\x28nee Smith Prof\\.\\x29", - "Jane Doe nee Smith King\\.", - "Jane Doe nee Smith MA Prof\\.", - "Jane Doe nee Smith Ma Prof\\.", - "Jane Doe nee Smith Prof\\.", - "Jane Doe nee Smith V Prof\\.", - "Mary Smith née Jones Prof\\."}), - frozenset({"Doe, Jane nee Smith MA Prof\\.", - "Jane Doe \\x28nee Smith Prof\\.\\x29", - "Jane Doe nee Prof\\. Dr\\.", - "Jane Doe nee Smith King\\.", - "Jane Doe nee Smith MA Prof\\.", - "Jane Doe nee Smith Ma Prof\\.", - "Jane Doe nee Smith Prof\\.", - "Jane Doe nee Smith V Prof\\.", - "Mary Smith née Jones Prof\\."}), - frozenset({"Doe nee Smith ba Prof\\.", - "Doe, Jane nee Smith MA Prof\\.", - "Jane Doe \\x28nee Smith Prof\\.\\x29", - "Jane Doe nee Prof\\. Dr\\.", - "Jane Doe nee Smith King\\.", - "Jane Doe nee Smith MA Prof\\.", - "Jane Doe nee Smith Ma Prof\\.", - "Jane Doe nee Smith Prof\\.", - "Jane Doe nee Smith V Prof\\.", - "Mary Smith née Jones Prof\\."}), # #436/#437's Latin movers, one corpus name per alternative -- a # list of names, not a copy of any wordlist, so there is no # vocabulary for it to drift from. The rule's subject is the @@ -2878,70 +2798,19 @@ class _LatinCopy(NamedTuple): # Smith Prof. MA' -- the title chain now ends the clause before # this rule's credential reading is asked, so the name moved to # its own fix(#535) rule: - frozenset({r"Doe, J\. nee MA ba", r"Doe, Jane Q\. nee Smith MA", - "Doe, Jane nee Smith DO", "Doe, Jane nee Smith MA", - "Doe, Jane nee Smith ma", "JANE DOE NEE SMITH MA", - "JANE DOE NEE YO-YO MA", r"Jane Doe geb\. Smith MA", - "Jane Doe nee Smith MA", "Jane Doe nee Smith MA JD", - "Jane Doe nee Smith MA PhD", "Jane Doe nee Smith Ma JD", - "Jane Doe nee Smith V MA", - "Jane Doe nee Smith do", - "John née Jones Smith MA", "Maria Kowalska z domu Nowak MA", - "jane doe nee smith ma"}), # and at 2.3.0, seventeen for the same reason plus one: 'Jane Doe # nee King. ba' joins here too (the review's own floor boundary, # this baseline never having split its credential at all). - frozenset({r"Doe, J\. nee MA ba", r"Doe, Jane Q\. nee Smith MA", - "Doe, Jane nee Smith DO", "Doe, Jane nee Smith MA", - "Doe, Jane nee Smith ma", "JANE DOE NEE SMITH MA", - "JANE DOE NEE YO-YO MA", r"Jane Doe geb\. Smith MA", - r"Jane Doe nee King\. ba", - "Jane Doe nee Smith MA", "Jane Doe nee Smith MA JD", - "Jane Doe nee Smith MA PhD", "Jane Doe nee Smith Ma JD", - "Jane Doe nee Smith V MA", - "Jane Doe nee Smith do", - "John née Jones Smith MA", "Maria Kowalska z domu Nowak MA", - "jane doe nee smith ma"}), # and at 2.0.0 and 2.1.0, thirteen for the same reason. - frozenset({r"Doe, Jane Q\. nee Smith MA", "Doe, Jane nee Smith DO", - "Doe, Jane nee Smith MA", "Doe, Jane nee Smith ma", - "JANE DOE NEE SMITH MA", "JANE DOE NEE YO-YO MA", - r"Jane Doe geb\. Smith MA", "Jane Doe nee Smith MA", - "Jane Doe nee Smith MA JD", "Jane Doe nee Smith MA PhD", - "Jane Doe nee Smith Ma JD", - "Jane Doe nee Smith V MA", - "Jane Doe nee Smith do", - "jane doe nee smith ma"}), # The names the clause KEEPS and now reports, at 2.3.0 (eight), # plus 'Doe nee Smith Prof. ba' (2026-09-26), rules.md#M2's # Accepted example of a title in front of a kept credential, # which #535 moves nothing on, - frozenset({"Berg, Jane van der nee Smith DO", "Berg, abdul nee Jones MA", - r"Doe nee Smith Prof\. ba", r"Doe, Dr\. nee Smith MA", - "Doe, Jane nee Smith Do", "Doe, Jane nee Smith MA do", - "Doe, Jane nee Smith Ma", "Doe, Jane nee Smith do", - "JOHN NEE JONES SMITH MA PHD", r"Jane Doe Jr\. nee Smith Ma", - "Jane Doe nee MA", "Jane Doe nee MA PhD", - "Jane Doe nee Smith DO DO", "Jane Doe nee Smith Ma", - "Jane Doe nee Yo-Yo Ma", "John née Jones Smith Ma"}), # and at 2.2.0 (nine): 'Jane Doe nee King. ba' joins here (the # review's own floor boundary; unlike at 2.3.0, this baseline # already split the credential, so only the report is new). - frozenset({"Berg, Jane van der nee Smith DO", "Berg, abdul nee Jones MA", - r"Doe, Dr\. nee Smith MA", "Doe, Jane nee Smith Do", - "Doe, Jane nee Smith MA do", "Doe, Jane nee Smith Ma", - "Doe, Jane nee Smith do", "JOHN NEE JONES SMITH MA PHD", - r"Jane Doe Jr\. nee Smith Ma", r"Jane Doe nee King\. ba", - "Jane Doe nee MA", "Jane Doe nee MA PhD", - "Jane Doe nee Smith DO DO", "Jane Doe nee Smith Ma", - "Jane Doe nee Yo-Yo Ma", "John née Jones Smith Ma"}), # and at 2.0.0 and 2.1.0 (nine): the review added 'Jane Doe nee # King. ba' beside 'Doe, J. nee MA ba', the same FLOOR boundary. - frozenset({r"Doe, Dr\. nee Smith MA", r"Doe, J\. nee MA ba", - "Doe, Jane nee Smith Ma", r"Jane Doe Jr\. nee Smith Ma", - r"Jane Doe nee King\. ba", "Jane Doe nee MA", - "Jane Doe nee MA PhD", "Jane Doe nee Smith DO DO", - "Jane Doe nee Smith Ma", "Jane Doe nee Yo-Yo Ma"}), # The two names where the released member lands somewhere other # than the trailing peel. One set, shared by the 1.4.0 rule and # the 2.2.0/2.3.0 one; at 2.0.0 and 2.1.0 the rule holds one name @@ -2954,16 +2823,6 @@ class _LatinCopy(NamedTuple): # narrowed fix(#424/#445) anchor, which holds the two spellings # whose reading that rule describes and no longer reaches the # third by (?i). - frozenset({"Doe, Jane nee Smith MA do", "Doe, Jane nee Smith Ma", - "Doe, Jane nee Smith do", r"Jane Doe Jr\. nee Smith Ma", - "Jane Doe nee MA", "Jane Doe nee Smith Ma", - "Jane Doe nee Yo-Yo Ma"}), - frozenset({r"Doe, Dr\. nee Smith MA", r"Doe, J\. nee MA ba", - "Doe, Jane nee Smith Ma", "Jane Doe nee MA", - "Jane Doe nee MA PhD", "Jane Doe nee Smith DO DO", - "Jane Doe nee Smith Ma", "Jane Doe nee Yo-Yo Ma"}), - frozenset({r"Doe, J\. nee MA ba", "Jane Doe nee Smith MA JD", - "Jane Doe nee Smith MA PhD", "Jane Doe nee Smith Ma JD"}), frozenset({"JOHN NEE JONES SMITH MA PHD", "John n[ée]e Jones Smith Ma"}), # 2026-09-20, #397: the #436/#437 rule's set grows at both the @@ -3008,9 +2867,6 @@ class _LatinCopy(NamedTuple): # last two are rules.md#M2's two delimiter examples (added with #538), # which the corpus parses at the DEFAULT policy and so for this # rule's own sentence rather than for #538's. - frozenset({"Doe, Jane nee Puig i Soler", "Jane Doe nee Puig i Soler", - "Smith, John, PhD née Puig Mr\\. - i Soler", - "Smith, John, PhD née Puig - i Soler"}), # #397's join and its one-case report, one corpus name per # alternative -- lists of names, not copies of any wordlist. The # join's subject is a SHAPE the vocabulary participates in at one @@ -3102,8 +2958,6 @@ class _LatinCopy(NamedTuple): # ended (fix(#274/#436/#437)), and the dotted spelling of a # two-token name (fix(suffix-routing)), whose members are # anchored on the name word and so copy no vocabulary. - frozenset({"Doe, Jane nee Smith PhD MEng", "Jane Doe nee Smith PhD MA", - "Jane Doe nee Smith PhD MEng"}), frozenset({r"jack\s+m\.a\.", r"wang\s+m\.eng\."}), # 2026-10-01, #564: the comma caps rule, literal names that copy no # set, at 1.4.0 and in the 2.x copy. @@ -3143,6 +2997,95 @@ class _LatinCopy(NamedTuple): frozenset({"Jane Doe, MS LAc", "John Smith, MD MEng", "John Smith, MEng PhD", "John Smith, PhD MEng", "john smith, phd meng"}), + # #601/#602 (2026-10-04): the literal-anchored rules those two + # changes added or narrowed, one alternative per corpus name -- + # lists of names, not copies of any wordlist. + frozenset({'Berg, Jane van der nee Smith DO', 'Berg, abd née Jones', + 'Berg, abdul nee Jones MA', 'Doe, Dr\\. nee Smith MA', + 'Doe, J\\. nee MA ba', 'Doe, Jane Q\\. nee Smith MA', + 'Doe, Jane nee Puig i Soler', 'Doe, Jane nee Smith', + 'Doe, Jane nee Smith DO', 'Doe, Jane nee Smith Do', + 'Doe, Jane nee Smith MA', 'Doe, Jane nee Smith MA Prof\\.', + 'Doe, Jane nee Smith MA do', 'Doe, Jane nee Smith Ma', + 'Doe, Jane nee Smith PhD MEng', 'Doe, Jane nee Smith do', + 'Doe, Jane nee Smith ma'}), + frozenset({'Doe nee Smith Prof\\. ba', 'JOHN NEE JONES SMITH MA PHD', + 'Jane Doe nee MA', 'Jane Doe nee MA PhD', + 'Jane Doe nee Smith Ma', 'Jane Doe nee Yo-Yo Ma', + 'John née Jones Smith Ma'}), + frozenset({'Doe nee Smith ba Prof\\.', + 'Jane Doe \\x28nee Smith Prof\\.\\x29', + 'Jane Doe nee Prof\\. Dr\\.', 'Jane Doe nee Smith King\\.', + 'Jane Doe nee Smith MA Prof\\.', + 'Jane Doe nee Smith Ma Prof\\.', 'Jane Doe nee Smith Prof\\.', + 'Jane Doe nee Smith V Prof\\.', 'Mary Smith née Jones Prof\\.'}), + frozenset({'Doe, Jane nee Puig i Soler', 'Doe, Jane nee Smith Do', + 'Doe, Jane nee Smith MA Prof\\.', 'Doe, Jane nee Smith MA do', + 'Doe, Jane nee Smith Ma', 'Doe, Jane nee Smith do'}), + frozenset({'Dr\\. nee Smith PhD Prof\\.', 'Jane Doe Jr\\. nee Smith', + 'Jane Doe Jr\\. nee Smith Ma', 'Jane Doe PhD nee Smith'}), + frozenset({'Dr\\. nee Smith PhD Prof\\.', 'Jane Doe Jr\\. nee Smith', + 'Jane Doe Jr\\. nee Smith Ma', 'Jane Doe PhD nee Smith', + 'Jane and née Jones'}), + frozenset({'Eric H\\. Holder Jr\\. Attorney General', + 'Jane van der Berg née Jr Jones', + 'John Smith PhD Jones', 'John Smith PhD Prof\\. Ma', + 'Smith, John PhD Jones'}), + frozenset({'Eric H\\. Holder Jr\\. Attorney General', + 'Jane van der Berg née Jr Jones', + 'John Smith PhD Jones', 'Smith, John PhD Jones'}), + frozenset({'JANE DOE NEE SMITH MA', 'JANE DOE NEE YO-YO MA', + 'Jane Doe geb\\. Smith MA', 'Jane Doe nee King\\. ba', + 'Jane Doe nee Smith MA', 'Jane Doe nee Smith MA JD', + 'Jane Doe nee Smith MA PhD', 'Jane Doe nee Smith Ma JD', + 'Jane Doe nee Smith V MA', 'Jane Doe nee Smith do', + 'John née Jones Smith MA', 'Maria Kowalska z domu Nowak MA', + 'jane doe nee smith ma'}), + frozenset({'JANE DOE NEE SMITH MA', 'JANE DOE NEE YO-YO MA', + 'Jane Doe geb\\. Smith MA', 'Jane Doe nee Smith MA', + 'Jane Doe nee Smith MA JD', 'Jane Doe nee Smith MA PhD', + 'Jane Doe nee Smith Ma JD', 'Jane Doe nee Smith V MA', + 'Jane Doe nee Smith do', 'John née Jones Smith MA', + 'Maria Kowalska z domu Nowak MA', 'jane doe nee smith ma'}), + frozenset({'JANE DOE NEE SMITH MA', 'JANE DOE NEE YO-YO MA', + 'Jane Doe geb\\. Smith MA', 'Jane Doe nee Smith MA', + 'Jane Doe nee Smith MA JD', 'Jane Doe nee Smith MA PhD', + 'Jane Doe nee Smith Ma JD', 'Jane Doe nee Smith V MA', + 'Jane Doe nee Smith do', 'jane doe nee smith ma'}), + frozenset({'JOHN NEE JONES SMITH MA PHD', 'Jane Doe nee King\\. ba', + 'Jane Doe nee MA', 'Jane Doe nee MA PhD', + 'Jane Doe nee Smith Ma', 'Jane Doe nee Yo-Yo Ma', + 'John née Jones Smith Ma'}), + frozenset({'Jane Doe \\x28nee Smith Prof\\.\\x29', + 'Jane Doe nee Prof\\. Dr\\.', 'Jane Doe nee Smith King\\.', + 'Jane Doe nee Smith MA Prof\\.', + 'Jane Doe nee Smith Ma Prof\\.', 'Jane Doe nee Smith Prof\\.', + 'Jane Doe nee Smith V Prof\\.', 'Mary Smith née Jones Prof\\.'}), + frozenset({'Jane Doe \\x28nee Smith Prof\\.\\x29', + 'Jane Doe nee Smith King\\.', 'Jane Doe nee Smith MA Prof\\.', + 'Jane Doe nee Smith Ma Prof\\.', 'Jane Doe nee Smith Prof\\.', + 'Jane Doe nee Smith Prof\\. MA', 'Jane Doe nee Smith V Prof\\.', + 'Mary Smith née Jones Prof\\.'}), + frozenset({'Jane Doe \\x28nee Smith Prof\\.\\x29', + 'Jane Doe nee Smith King\\.', 'Jane Doe nee Smith MA Prof\\.', + 'Jane Doe nee Smith Ma Prof\\.', 'Jane Doe nee Smith Prof\\.', + 'Jane Doe nee Smith V Prof\\.', 'Mary Smith née Jones Prof\\.'}), + frozenset({'Jane Doe nee King\\. ba', 'Jane Doe nee MA', + 'Jane Doe nee MA PhD', 'Jane Doe nee Smith Ma', + 'Jane Doe nee Yo-Yo Ma'}), + frozenset({'Jane Doe nee MA', 'Jane Doe nee Smith Ma', + 'Jane Doe nee Yo-Yo Ma'}), + frozenset({'Jane Doe nee Puig i Soler', + 'Smith, John, PhD née Puig - i Soler', + 'Smith, John, PhD née Puig Mr\\. - i Soler'}), + frozenset({'Jane Doe nee Smith DO DO', + 'Jane van der Berg nee Prof\\. King\\. MA', + 'Jane van der Berg nee Smith Prof\\.', 'John nee Prof\\. ba MA'}), + frozenset({'Jane Doe nee Smith MA JD', 'Jane Doe nee Smith MA PhD', + 'Jane Doe nee Smith Ma JD'}), + frozenset({'Jane Doe nee Smith PhD MA', 'Jane Doe nee Smith PhD MEng'}), + frozenset({'Smith, John, MD - née Jones Smith', + 'Smith, John, PhD née Jones'}), }) def _unjustified_reach(name_regex: str, members: set[str]) -> list[str]: @@ -3595,8 +3538,9 @@ def _claim(rule: dict) -> _Claim: # 2026-09-27, #544: 11 -> 17; gains 'Doe, Jane PhD MEng', # 'Jane Doe Jr. Ma', 'John Smith PhD Ed Ma', 'John Smith PhD # MEng' and 2 more. + # 2026-10-04, #601/#602: 17 -> 16; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#436/#437) a space-separated post-nominal run renders with spaces, not commas": - _Claim(17, ('suffix',), "dd1fbf116ab6", None), + _Claim(16, ('suffix',), "b171e45499c3", None), # #346's alternation. Four corpus names, `family` and # `given` together: the fold moves both roles at once, so a # widening taking one alone would change the roles here @@ -3703,8 +3647,9 @@ def _claim(rule: dict) -> _Claim: # 2026-10-03, #549: 108 -> 109, 'Smith, John, MD - née Jones # Smith', rules.md#C1's Accepted example. Reach -- it carries a # marker. + # 2026-10-04, #601/#602: 109 -> 98; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#274) maiden markers consumed": - _Claim(109, ('family', 'maiden', 'middle'), "87d3acaaea68", None), + _Claim(98, ('family', 'maiden', 'middle'), "7b41ba0f0b38", None), # 2026-09-19, #533: 5 -> 6, the same one new corpus name # '田中 太郎 旧姓 佐藤 MA' as the CJK rule above. "fix(cjk-maiden-marker) maiden marker consumed, compounding with the CJK order flip": @@ -3747,6 +3692,8 @@ def _claim(rule: dict) -> _Claim: # boundary, which attaches as this rule reads. Reach, verified. "fix(#379) a tussenvoegsel after a family comma attaches to the family": _Claim(32, ('family', 'middle'), "3e3ff4af0060", None), + "fix(#379) a tussenvoegsel behind a post-nominal after a family comma attaches": + _Claim(1, ('family', 'middle'), "f24e5eee3cc1", None), # 2026-10-01, #554: 2 -> 3, 'John Smith, Jr vd', #554's # rules.md#C1 example. Reach, verified by name. # 2026-10-02, #573: 3 -> 4, 'DOE, JANE PHD VD', P6's @@ -3868,6 +3815,9 @@ def _claim(rule: dict) -> _Claim: # name by name. # 2026-10-02, #573: 430 -> 439, its nine rules.md#P2, # P6 and S2 comma examples. Reach, verified name by name. + # 2026-10-04, #601/#602: 443 -> 438; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. + # 2026-10-04, #601/#602 review: 438 -> 439, rules.md#S2's new + # example 'Smith, John PhD de Jr.'. Reach, verified by name. "fix(comma-family) lone post-comma piece routes to suffix/title, not first": # 2026-10-01, #564: 423 -> 430, 'John Smith, XYZ', 'Smith, # XYZ', 'García Márquez, MJ', 'García Márquez, MJ PhD', @@ -3878,11 +3828,20 @@ def _claim(rule: dict) -> _Claim: # 'Smith, John, PhD - and MD' and 'Smith, John, MD - née Jones # Smith', #549's rules.md#C1 examples. Reach -- each is a # comma name. + # 2026-10-04, #602: 442 -> 443, 'Smith, John PhD Jones', the + # rules.md#S2 example of the given part's credential run. + # Reach -- a comma name. # 2026-10-04, #604: 442 -> 443, 'Pérez, Juan Ó.', #604's # rules.md#P7 comma example. Reach. + # 2026-10-04, merging #604 into the #601/#602 branch: 440, + # recomputed over the merged corpus -- #604's P7 example and + # #602's two S2 examples in, the marker names #601 moved to + # its own rules out. Reach. # 2026-10-04, #606: 443 -> 444, 'Greve, Anna', #606's # rules.md#H1 comma example. Reach. - _Claim(444, ('given', 'suffix', 'title'), "f2ad75955602", None), + # 2026-10-04, merging #606 in as well: 441 over the merged + # corpus. Reach. + _Claim(441, ('given', 'suffix', 'title'), "a3028e60b1e7", None), "fix(comma-family) a comma followed only by titles keeps the given/family split": _Claim(2, ('family', 'given'), "5bd9c6d96c38", None), "fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example": @@ -3988,6 +3947,9 @@ def _claim(rule: dict) -> _Claim: # name by name. # 2026-10-02, #573: 430 -> 439, its nine rules.md#P2, # P6 and S2 comma examples. Reach, verified name by name. + # 2026-10-04, #601/#602: 443 -> 438; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. + # 2026-10-04, #601/#602 review: 438 -> 439, rules.md#S2's new + # example 'Smith, John PhD de Jr.'. Reach, verified by name. "fix(comma-precomma-family) pre-comma run reads as family, not given": # 2026-10-01, #564: 423 -> 430, 'John Smith, XYZ', 'Smith, # XYZ', 'García Márquez, MJ', 'García Márquez, MJ PhD', @@ -3996,9 +3958,16 @@ def _claim(rule: dict) -> _Claim: # name by name. # 2026-10-03, #549: 439 -> 442, the same three #549 examples. # Reach. + # 2026-10-04, #602: 442 -> 443, 'Smith, John PhD Jones'. Reach. # 2026-10-04, #604: 442 -> 443, 'Pérez, Juan Ó.'. Reach. + # 2026-10-04, merging #604 into the #601/#602 branch: 440, + # recomputed over the merged corpus -- #604's P7 example and + # #602's two S2 examples in, the marker names #601 moved to + # its own rules out. Reach. # 2026-10-04, #606: 443 -> 444, 'Greve, Anna'. Reach. - _Claim(444, ('family', 'given'), "f2ad75955602", None), + # 2026-10-04, merging #606 in as well: 441 over the merged + # corpus. Reach. + _Claim(441, ('family', 'given'), "a3028e60b1e7", None), # 2026-10-01, #575: new, 4; 'De La Cruz, Ed', 'Freiherr von # Berg, Ed', 'Van Buren, Ed', 'de la Cruz, Ma'. "fix(#575) a particle surname before a comma is one name word": @@ -4095,12 +4064,11 @@ def _claim(rule: dict) -> _Claim: # 2026-09-26, #535: 2 -> 3, 'Berg, abdul nee Smith V' joining # the corpus -- the same bound-given/maiden shape. Reach, not # explanation. + # 2026-10-04, #601/#602: 3 -> 1; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#411) the bound-given reserve stops counting words the maiden name takes": - _Claim(3, ('given', 'maiden', 'middle'), "9305df1c05c5", None), + _Claim(1, ('given', 'middle'), "431d3dd24c14", None), "fix(#400/#274) bound-given join and maiden consumption in one name": _Claim(1, ('family', 'given', 'maiden', 'middle'), "6bed6d349342", None), - "fix(#411/S2) a declining bound-given join leaves the suffix reading after a family comma": - _Claim(1, ('given', 'maiden', 'middle', 'suffix'), "0f8ed9db0a32", None), "fix(#418) the connective carve-out counts the name the maiden clause leaves behind": _Claim(1, ('family', 'given', 'maiden', 'middle'), "7923e6d3c5a7", None), "fix(credential-pair-order) a split credential and a suffix render in written order": @@ -4151,8 +4119,9 @@ def _claim(rule: dict) -> _Claim: _Claim(1, ('family', 'given'), "ee4339908f4d", None), "fix(#367) a title no longer displaces a leading never-given particle": _Claim(1, ('family', 'given'), "db724fb9c779", None), + # 2026-10-04, #601/#602: 7 -> 6; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#445) a maiden marker makes the lone name word the family": - _Claim(7, ('family', 'given', 'maiden', 'middle'), "3de9ef12b4a8", None), + _Claim(6, ('family', 'given', 'maiden', 'middle'), "3a1de67a470f", None), "fix(N3) a nickname-led name with a trailing suffix keeps the suffix in `suffix`": _Claim(1, ('family', 'suffix'), "570f265a2f46", None), # #451's four replacements for the fields-only catch-all, whose @@ -4271,8 +4240,9 @@ def _claim(rule: dict) -> _Claim: # 2026-10-01, #575: 104 -> 105, 'Ortega y Gasset, Ed'. Reach. # 2026-10-03, #549: 105 -> 106, 'Smith, John, PhD - and MD'. # Reach. + # 2026-10-04, #601/#602: 106 -> 107; 'Jane and née Jones', a new rules.md#M2 example, is in reach -- a marker behind a connective is a word in the run since #601. "fix(initials-per-word) a connective run initials each word (facade, since 2.0.0)": - _Claim(106, ('_initials',), "8b55345f9ad4", ('DEFAULT',)), + _Claim(107, ('_initials',), "9a8658d181fb", ('DEFAULT',)), # 2026-09-19, #533: 41 -> 43. Two new corpus names opening # with a bound-given word, 'Berg, abdul MA' and 'Berg, abdul # nee Jones MA' -- the P5 pair this change added to record @@ -4280,8 +4250,9 @@ def _claim(rule: dict) -> _Claim: # 2026-09-26, #535: 43 -> 44, 'Berg, abdul nee Smith V' -- # another opening bound-given word. Reach again. # 2026-10-01, #575: 44 -> 45, 'abdul Salam, Ed'. Reach. + # 2026-10-04, #601/#602: 45 -> 43; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(initials-per-word) a bound-given run initials each word (facade, since 2.0.0)": - _Claim(45, ('_initials',), "690d9625f120", ('DEFAULT',)), + _Claim(43, ('_initials',), "de26d994148e", ('DEFAULT',)), # 2026-09-18: 109 -> 110. One new corpus name, # 'john van der berg ma' -- rules.md#P2's one-case contrast, # and a particle chain like every other member. @@ -4300,8 +4271,9 @@ def _claim(rule: dict) -> _Claim: # Ed', 'Freiherr von Berg, Ed'). Reach, verified name by name. # 2026-10-02, #573: 124 -> 125, 'SMITH, VD DE LA MA' -- a # particle chain, 'VD DE LA'. Reach, verified name by name. + # 2026-10-04, #601/#602: 125 -> 123; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(initials-per-word) a particle chain inside a name part initials each word (facade, since 2.0.0)": - _Claim(125, ('_initials',), 'c1fe82860c6c', ('DEFAULT',)), + _Claim(123, ('_initials',), "d69d11083dba", ('DEFAULT',)), # 2026-09-23, #459: 18 -> 19, 'john smith ph. d.', rules.md#R4's # two-token line. Reach, verified name by name. "fix(initials-per-word) the Ph. D. merge initials each word (facade, since 2.0.0)": @@ -4466,8 +4438,9 @@ def _claim(rule: dict) -> _Claim: # word up, which is the rule below. # 2026-09-27, #544: 6 -> 7; gains 'Jane Doe Jr. nee Smith Ma', # a case row #544 added whose roles every 2.x release reads. + # 2026-10-04, #601/#602: 7 -> 3; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#274/#424) accepted: a maiden clause keeps a trailing credential v1 read as a post-nominal": - _Claim(7, ('family', 'maiden', 'middle', 'suffix'), "2d5c58b81dba", None), + _Claim(3, ('family', 'maiden', 'middle', 'suffix'), "bd025e92d67b", None), # One corpus name, four roles. The rule is a compound of two # released changes and #533 moves nothing on it, so a growth # here is this rule reaching a name whose clause the @@ -4475,20 +4448,16 @@ def _claim(rule: dict) -> _Claim: # 2026-09-27, #544: 1 -> 3; gains 'Doe, Jane nee Smith PhD # MEng' and 'Jane Doe nee Smith PhD MEng', case rows #544 added # whose roles every 2.x release reads -- the same two changes. + # 2026-10-04, #601/#602: 3 -> 2; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#274/#436/#437) a clause the unambiguous credential ended, and the run it left renders with spaces": - _Claim(3, ('family', 'maiden', 'middle', 'suffix'), "60850ce87c18", None), + _Claim(2, ('family', 'maiden', 'middle', 'suffix'), "45b66dfef88c", None), # Four corpus names, four roles. The `suffix` is where the # released member lands and `maiden` is what it left, so a # widening that took one without the other would change the # roles here before it reached the gate. + # 2026-10-04, #601/#602: 4 -> 3; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#533) the clause gives the credential up, and the run it joins renders as the writer spaced it": - _Claim(4, ('family', 'maiden', 'middle', 'suffix'), "3b2a2ee29ece", None), - # One corpus name. `suffix` is NOT in the 1.4.0 roles and - # that is the point of the row: v1 read the member as a - # post-nominal and so does the tree, which is what #533 - # restored here. - "fix(#445/#533) the clause gives the credential up, and the one name word it leaves is the family": - _Claim(1, ('family', 'given', 'maiden', 'middle'), "6403e56cea29", None), + _Claim(3, ('family', 'maiden', 'middle', 'suffix'), "19d1d3c15d86", None), # One corpus name. The by-shape member has no lean to read, # so a growth here is the rule reaching a LISTED member -- # a different reading under this rule's sentence. @@ -4500,8 +4469,9 @@ def _claim(rule: dict) -> _Claim: # reached a name whose member moved. "fix(#434/#533) a marker PHRASE takes the maiden name, and its clause ends at the credential": _Claim(1, ('family', 'maiden', 'middle'), "bf8359f65c8a", None), + # 2026-10-04, #601/#602: 5 -> 1; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#274/#533) the clause keeps the credential a join or a missing reader would have taken": - _Claim(5, ('family', 'given', 'maiden', 'middle', 'suffix'), 'ae9d39617af1', None), + _Claim(1, ('family', 'maiden', 'middle', 'suffix'), "13f391350f24", None), "fix(#335/#533) a marker-led bracketed clause is the maiden name whatever its last word is": _Claim(3, ('maiden', 'nickname'), 'cc1045ecdc05', None), # 2026-09-20, #397 and #461's three new rules at this @@ -4530,36 +4500,9 @@ def _claim(rule: dict) -> _Claim: ('DEFAULT',)), "fix(#461) a connective holding its part alone contributes an initial": _Claim(15, ('_initials',), "4d436d1ebeca", ('DEFAULT',)), - "fix(#274/#397) a link inside a maiden clause stays in the birth name, after a family comma": - _Claim(1, ('maiden', 'middle', 'suffix'), 'ad2442b1b12a', - None), "fix(#274/#436/#437) a clause the generational link ended, and the run it left renders with spaces": _Claim(1, ('family', 'maiden', 'middle', 'suffix'), '7d64445a7252', None), - # 2026-09-22, #397 follow-up: new rule, one literal name -- - # rules.md#M2's `deviates: #538` example, which this baseline - # reads as one long suffix. - # 2026-09-26, #538: 1 -> 2, rules.md#M2's second live example, - # 'Smith, John, PhD née Puig - i Soler', joining the same - # alternation for the same reason. - # 2026-10-03, #549: 2 -> 3, rules.md#C1's Accepted example, - # 'Smith, John, MD - née Jones Smith', joining it too. - "fix(#274/#397) a maiden clause inside a suffix-comma tail leaves the suffix field, link and all": - _Claim(3, ('maiden', 'suffix'), 'c1572309c3f2', None), - # New rule (#274): one corpus name, 'Dr. nee Jones Smith - # Prof.' -- only a title precedes the marker, so no name word - # is left standing ahead of it. Its diff from this baseline is - # #274's alone (v1 has no maiden reading), verified against - # the parent tree d9d80492: before #535 ran, it already read - # title 'Dr.', maiden 'Jones Smith Prof.', identically to - # HEAD, so #535 moves nothing here. `given` moves because - # 'nee' is v1's given text; the sibling '#274 maiden markers - # consumed' rule above does not declare it, which is why this - # name needs its own rule rather than joining that one's - # alternation. - "fix(#274) a trailing title after the marker with nothing else in the name": - _Claim(1, ('family', 'given', 'maiden', 'middle'), - 'b6132e37d7a8', None), # New rule (#274/#296/#535): one corpus name, 'Jane Doe nee # Prof. Dr.' -- moved out of the nine-name list below (review): # this baseline still has 'dr' in the suffix vocabulary @@ -4589,61 +4532,9 @@ def _claim(rule: dict) -> _Claim: # moves `title` (and `suffix` for the credential-bearing ones # among the seven; 1.4.0 is compared on the facade, where # `_ambiguities` does not exist). + # 2026-10-04, #601/#602: 9 -> 7; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#274/#535) a trailing title ends the maiden clause": - _Claim(9, ('family', 'maiden', 'middle', 'suffix', 'title'), - '6be225d75666', None), - # New rule (#411/#535): one corpus name, 'Berg, abdul nee - # Smith V' -- the numeral-join literal, relabelled from a bare - # '#535' rule (review): this baseline and 2.0.0 read the name - # IDENTICALLY (given 'abdul nee', middle 'Smith', family - # 'Berg', suffix 'V'), so #274 (no maiden reading at all) - # moves nothing here and the label is the same two-issue one - # 2.0.0/2.1.0 carry, not a third. Two causes move it at this - # baseline: #411, the bound-given reserve, which the parent - # tree d9d80492 shows already deciding `given`/`maiden` before - # #535 ran; and #535, which moves the released V from given - # back into maiden. - "fix(#411/#535) the numeral stop asks the join question": - _Claim(1, ('given', 'maiden', 'middle', 'suffix'), - 'c8d3b254b804', None), - # 2026-09-26: rules.md#M2's Accepted examples added 'Jane Doe - # nee Smith King.' to the title-stop rule, and the - # two rules below hold the Accepted 'ba' pair, one name each. - "fix(#274/#342/#445/#535) a title behind a credential the clause keeps still leaves it": - _Claim(1, ('family', 'given', 'maiden', 'middle', 'title'), - 'b8aa7d51a16c', None), - "fix(#274/#342/#445/#533) a title in front of a credential the clause keeps stays in it": - _Claim(1, ('family', 'given', 'maiden', 'middle', 'suffix'), - 'cb346cb8419a', None), - # 2026-09-26: one-name rules for rules.md#M2's second-round - # Accepted examples (the particle pair, the particle-chain - # title pair, and 'Dr. nee V' at 1.4.0). - "fix(#274/#535) a particle in front of a trailing title stays in the clause": - _Claim(1, ('family', 'maiden', 'middle', 'title'), - 'a1484fe026d1', None), - "fix(#274/#535) a particle behind a trailing title leaves the clause with it": - _Claim(1, ('family', 'maiden', 'middle', 'title'), - '97e10b3d2aa8', None), - "fix(#274/#316) a trailing title behind a maiden clause's first suffix word": - _Claim(1, ('family', 'maiden', 'middle', 'suffix', 'title'), - 'd58b847181aa', None), - "fix(#274) a numeral straight after the marker stays maiden where nothing reads it as a suffix": - _Claim(1, ('family', 'given', 'maiden'), - '4ff67af4187f', None), - # 2026-09-26: one-name rules for rules.md#M2's - # Accepted examples (the particle-inside-the-clause pair, and - # 'Doe nee Smith V, Jane' at 2.0.0/2.1.0). - "fix(#274/#535) a particle inside the clause keeps the credential in front of a trailing title": - _Claim(1, ('family', 'maiden', 'middle', 'title'), - '647659c5eafc', None), - "fix(#274/#535) a particle inside the clause lets a credential behind the title go": - _Claim(1, ('family', 'maiden', 'middle', 'title'), - '8db3f5842115', None), - # 2026-09-26: the one-name rule for rules.md#M2's example - # of a link the clause stops at and keeps. - "fix(#274/#397/#535) a link the clause stops at keeps a run that would not read off": - _Claim(1, ('family', 'maiden', 'middle', 'title'), - '1b74e094fed9', None), + _Claim(7, ('family', 'maiden', 'middle', 'suffix', 'title'), "5ffb52f287a9", None), # 2026-09-27, #544: new, 1; gains 'John Smith, X.Y.Z. MA'. "fix(#544) a run of ambiguous members after a suffix comma reads by the name-word count": _Claim(1, ('family', 'given', 'suffix'), "23c792cefe18", ('DEFAULT',)), @@ -4663,6 +4554,19 @@ def _claim(rule: dict) -> _Claim: _Claim(2, ('given', 'title'), "8cdafcf56c45", None), "fix(#597) halfwidth corner brackets enclose a nickname": _Claim(3, ('family', 'given', 'middle', 'nickname'), "d1b9b0addff3", None), + # 2026-10-04, #601/#602: new rules. + "fix(#601) a marker in the given part after a family comma is an ordinary word": + _Claim(6, ('family', 'middle', 'suffix', 'title'), "007582298dfc", None), + "fix(#601/#602) a marker behind a title, a suffix word or a connective is an ordinary word": + _Claim(4, ('family', 'middle', 'suffix', 'title'), "456e8a63ed75", None), + "fix(#274/#601) the clause ends at the clause-free name's trailing run, which the take consumes": + _Claim(4, ('family', 'given', 'maiden', 'middle', 'suffix', 'title'), "d15d58da6be1", None), + "fix(#274/#601) the first word after the marker is the maiden name": + _Claim(1, ('family', 'maiden', 'middle', 'suffix'), "aaf53040b071", None), + "fix(#602) a credential after the name core starts a run to the end of its part": + _Claim(5, ('family', 'middle', 'suffix', 'title'), "3129cd9609b9", None), + "fix(#274/#601/#602) a credential in the clause starts a run the take consumes": + _Claim(1, ('family', 'maiden', 'middle', 'suffix'), "acdcc81cb17c", None), # 2026-10-04, #604: new, 3; the rules.md#P7 boundary and the # two rules.md#R4 case-repair examples. "fix(#604) the Irish and Malay patronymic particles": @@ -4715,8 +4619,9 @@ def _claim(rule: dict) -> _Claim: # case rows. Roles unmoved, `suffix` alone as before. # 2026-09-27, #544: 15 -> 18; gains 'Jane Doe Jr. Ma', 'John # Smith PhD Ed Ma', 'John Smith PhD Ma'. + # 2026-10-04, #601/#602: 18 -> 17; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#436/#437) a space-separated post-nominal run renders with spaces, not commas": - _Claim(18, ('suffix',), "60ec2103c480", None), + _Claim(17, ('suffix',), "8f6fa4cee3c3", None), # #449's six rules, second in every 2.x ledger. The # alternation reached twenty-two corpus names at landing # (twenty-three since #544, 2026-09-27) and @@ -4840,6 +4745,8 @@ def _claim(rule: dict) -> _Claim: # Smith, Jr do', #554's rules.md#C1 examples. Reach, verified # name by name. _Claim(32, ('_ambiguities', 'family', 'middle'), "3e3ff4af0060", None), + "fix(#379) a tussenvoegsel behind a post-nominal after a family comma attaches": + _Claim(1, ('family', 'middle'), "f24e5eee3cc1", None), # 2026-09-18: 126 -> 131. Five corpus names arrived with # #289/#516's own case rows -- 'J.씨', 'John Smith 田.中.', # '毛泽东, MA', '田中 太郎, MA', '마틴 킹, MA' -- all of them @@ -4903,8 +4810,9 @@ def _claim(rule: dict) -> _Claim: # boundary, which attaches as this rule reads. Reach, verified. "fix(#380) a trailing vd after a family comma is the tussenvoegsel, not a post-nominal": _Claim(4, ('_ambiguities', 'family', 'suffix'), "25798c4e2a22", None), + # 2026-10-04, #601/#602: 6 -> 7; 'Mai Le née Nguyen', a new rules.md#M2 example, is in reach -- at this baseline the particle chain swallowed its marker, the #399 shape. "fix(#399) a maiden marker bounds the particle chain that swallowed it": - _Claim(6, ('family', 'maiden'), "89e1f4afdd2a", ('DEFAULT',)), + _Claim(7, ('family', 'maiden'), "0a306405fb9a", ('DEFAULT',)), "fix(#360) mc moved into the never-given particles, so it folds into the family": _Claim(1, ('_ambiguities', 'family', 'given'), "ee4339908f4d", None), "fix(#400) abd joins the word after it as one given name": @@ -4921,8 +4829,9 @@ def _claim(rule: dict) -> _Claim: # 2026-09-26, #535: 2 -> 3, 'Berg, abdul nee Smith V' joining # the corpus -- the same bound-given/maiden shape. Reach, not # explanation. + # 2026-10-04, #601/#602: 3 -> 1; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#411) the bound-given reserve stops counting words the maiden name takes": - _Claim(3, ('given', 'maiden', 'middle'), "9305df1c05c5", None), + _Claim(1, ('given', 'maiden', 'middle'), "431d3dd24c14", None), "fix(#412) a connective join no longer absorbs the maiden marker beside it": _Claim(2, ('family', 'maiden'), "51c0eb36b5c5", None), "fix(#418) the connective carve-out counts the name the maiden clause leaves behind": @@ -5019,12 +4928,11 @@ def _claim(rule: dict) -> _Claim: # case-insensitively and correctly, the two being one # given->family shape written two ways. Reach, not # explanation. + # 2026-10-04, #601/#602: 7 -> 5; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#445) a maiden marker makes the lone name word the family": - _Claim(7, ('_ambiguities', 'family', 'given'), "73c2f924d257", None), + _Claim(5, ('_ambiguities', 'family', 'given'), "2648aa85b6b4", None), "fix(#445) the lone name word beside a marker a connective join no longer absorbs": _Claim(1, ('family', 'given', 'maiden', 'middle'), "52544a41dd62", None), - "fix(#424) a marker followed only by the numeral is just a word": - _Claim(1, ('_ambiguities', 'family', 'maiden', 'middle', 'suffix'), "aaf53040b071", None), "fix(#369) the bound given-name join takes a particle-and-bound word, so no fork is reported": _Claim(1, ('_ambiguities',), "81cf02ffdb33", None), "fix(#360) ste moved into the never-given particles with mc": @@ -5239,8 +5147,9 @@ def _claim(rule: dict) -> _Claim: # LEAVING this rule's alternation -- the title chain now ends # the clause before the credential is asked, so the name has # its own fix(#535) rule instead. + # 2026-10-04, #601/#602: 14 -> 10; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#533) a credential ending a maiden clause reads as a credential": - _Claim(14, ('_ambiguities', 'maiden', 'suffix'), 'e5cd5fb89c8b', None), + _Claim(10, ('_ambiguities', 'maiden', 'suffix'), "e05c5164668c", None), # The declining half: fourteen corpus names at 2.2.0, eight # at 2.0.0 and 2.1.0. `_ambiguities` alone, so a role # appearing here is this rule reaching a name whose clause @@ -5251,8 +5160,9 @@ def _claim(rule: dict) -> _Claim: # only the report. # 2026-09-27, #544: 9 -> 10; gains 'Jane Doe Jr. nee Smith # Ma'. + # 2026-10-04, #601/#602: 10 -> 4; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#533) the maiden clause reports the credential it keeps": - _Claim(10, ('_ambiguities',), "7c30ed523806", None), + _Claim(4, ('_ambiguities',), "c57e8eb30830", None), # One corpus name. The by-shape member has no lean to read, # so a growth here is the rule reaching a LISTED member -- # a different reading under this rule's sentence. @@ -5265,12 +5175,6 @@ def _claim(rule: dict) -> _Claim: # being one. "fix(#533) restores the suffix reading a 2.4 retag had moved into the maiden name": _Claim(1, ('_ambiguities',), "0a7295a0cb19", None), - # One corpus name. `suffix` is NOT in the 1.4.0 roles and - # that is the point of the row: v1 read the member as a - # post-nominal and so does the tree, which is what #533 - # restored here. - "fix(#445/#533) the clause gives the credential up, and the one name word it leaves is the family": - _Claim(1, ('_ambiguities', 'family', 'given', 'maiden', 'suffix'), "6403e56cea29", None), # One corpus name, at 2.0.0 and 2.1.0 only. Case-sensitive by # construction, so a growth here is the anchor having picked # up the mixed-case spelling, which reads the other way. @@ -5289,10 +5193,6 @@ def _claim(rule: dict) -> _Claim: _Claim(1, ('_ambiguities', 'maiden', 'suffix'), "170a53c37765", None), "fix(#335/#533) a marker-led bracketed clause is the maiden name whatever its last word is": _Claim(3, ('_ambiguities', 'maiden', 'nickname'), 'cc1045ecdc05', None), - "fix(#399/#533) a maiden marker bounds the particle chain, and the clause keeps the credential the chain would have taken": - _Claim(1, ('_ambiguities', 'family', 'maiden', 'middle'), 'fa3fe4878b47', None), - "fix(#411/#533) the bound-given pair after a comma, and the clause it keeps its credential in": - _Claim(1, ('_ambiguities', 'given', 'maiden', 'middle'), '431d3dd24c14', None), # 2026-09-20, #397 and #461's four new rules, all literal # alternations of exactly the names each explains -- a count # that moves means the alternation grew, and _MUST_NOT_MATCH @@ -5320,21 +5220,9 @@ def _claim(rule: dict) -> _Claim: # 2026-09-26, #538: 3 -> 4, rules.md#M2's second live example, # 'Smith, John, PhD née Puig - i Soler', joining the same # alternation for the same reason. Verified name by name. + # 2026-10-04, #601/#602: 4 -> 1; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#397) a link inside a maiden clause stays in the birth name": - _Claim(4, ('family', 'maiden', 'middle', 'suffix'), - '4556aa8f93e5', ('DEFAULT',)), - # New rule (#399): one corpus name, 'Jane van der Berg nee - # Smith Prof.' -- a sibling of the general #399 rule above, - # for the two-trailing-word shape its own anchor - # (`\\sn[ée]e\\s+\\S+$`) cannot reach: per rules.md#P2, "a - # maiden marker takes the words after it (M2), or the name - # ends", and the chain here swallows both trailing words. Its - # diff from this baseline is #399's alone, verified against - # the parent tree d9d80492: before #535 ran, it already read - # family 'van der Berg', maiden 'Smith Prof.', identically to - # HEAD. - "fix(#399) a maiden marker bounds the particle chain that swallowed it, two trailing words": - _Claim(1, ('family', 'maiden'), '6b2d1fa71195', None), + _Claim(1, ('family', 'maiden', 'middle'), "4e62e18718a4", ('DEFAULT',)), # New rule (#296/#535): one corpus name, 'Jane Doe nee Prof. # Dr.'. This baseline still has 'dr' in the suffix vocabulary # (#296 had not run yet), so it reads suffix 'Dr.'. The parent @@ -5360,59 +5248,9 @@ def _claim(rule: dict) -> _Claim: # this baseline has an earlier cause too. `title`, `maiden` # and `suffix` move for the credential-bearing ones, and # `_ambiguities` for the declined-member one. + # 2026-10-04, #601/#602: 8 -> 6; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#535) a trailing title ends the maiden clause": - _Claim(8, ('_ambiguities', 'maiden', 'suffix', 'title'), - '3216299d78a6', None), - # New rule (#411/#535): one corpus name, 'Berg, abdul nee - # Smith V', relabelled from a bare '#535' rule (review): #411 - # (the bound-given reserve) already decides `given`/`maiden` - # at this baseline before #535 exists, verified against the - # parent tree d9d80492; #535 then moves the released V from - # given back into maiden. - "fix(#411/#535) the numeral stop asks the join question": - _Claim(1, ('given', 'maiden', 'middle', 'suffix'), - 'c8d3b254b804', None), - # 2026-09-26: rules.md#M2's Accepted examples added 'Jane Doe - # nee Smith King.' to the title-stop rule, and the - # two rules below hold the Accepted 'ba' pair, one name each. - "fix(#342/#445/#535) a title behind a credential the clause keeps still leaves it": - _Claim(1, ('_ambiguities', 'family', 'given', 'maiden', 'middle', - 'title'), 'b8aa7d51a16c', None), - "fix(#342/#445/#533) a title in front of a credential the clause keeps stays in it": - _Claim(1, ('_ambiguities', 'family', 'given', 'maiden', 'suffix'), - 'cb346cb8419a', None), - # 2026-09-26: one-name rules for rules.md#M2's second-round - # Accepted examples (the particle pair, the particle-chain - # title pair, and 'Dr. nee V' at 1.4.0). - "fix(#535) a particle in front of a trailing title stays in the clause": - _Claim(1, ('_ambiguities', 'maiden', 'title'), - 'a1484fe026d1', None), - "fix(#533/#535) a particle behind a trailing title leaves the clause with it": - _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), - '97e10b3d2aa8', None), - "fix(#316/#399) a trailing title behind a maiden clause's first suffix word": - _Claim(1, ('family', 'maiden', 'middle', 'suffix', 'title'), - 'd58b847181aa', None), - "fix(#399) a maiden marker bounds the particle chain ahead of a title the clause keeps": - _Claim(1, ('family', 'maiden'), - '58cf2330fed6', None), - # 2026-09-26: one-name rules for rules.md#M2's - # Accepted examples (the particle-inside-the-clause pair, and - # 'Doe nee Smith V, Jane' at 2.0.0/2.1.0). - "fix(#535) a particle inside the clause keeps the credential in front of a trailing title": - _Claim(1, ('_ambiguities', 'maiden', 'title'), - '647659c5eafc', None), - "fix(#533/#535) a particle inside the clause lets a credential behind the title go": - _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), - '8db3f5842115', None), - "fix(#424) accepted: before a family comma the numeral the walk gives up goes to the family": - _Claim(1, ('family', 'maiden'), - '339279b78f2f', None), - # 2026-09-26: the one-name rule for rules.md#M2's example - # of a link the clause stops at and keeps. - "fix(#397/#535) a link the clause stops at keeps a run that would not read off": - _Claim(1, ('_ambiguities', 'family', 'maiden', 'middle', 'title'), - '1b74e094fed9', None), + _Claim(6, ('_ambiguities', 'maiden', 'suffix', 'title'), "3e41b335baa4", None), # 2026-09-27, #544: new, 5; gains 'Jane Doe, MS LAc', 'John # Smith, MD MEng', 'John Smith, MEng PhD', 'John Smith, PhD # MEng' and 1 more. @@ -5455,6 +5293,21 @@ def _claim(rule: dict) -> _Claim: _Claim(2, ('_ambiguities', 'given', 'title'), "8cdafcf56c45", None), "fix(#597) halfwidth corner brackets enclose a nickname": _Claim(3, ('family', 'given', 'middle', 'nickname'), "d1b9b0addff3", None), + # 2026-10-04, #601/#602: new rules. + "fix(#601) a marker in the given part after a family comma is an ordinary word": + _Claim(17, ('_ambiguities', 'family', 'given', 'maiden', 'middle', 'suffix', 'title'), "c7ed65a885ac", None), + "fix(#601) a marker in a part after a suffix comma is an ordinary word": + _Claim(2, ('maiden', 'suffix'), "59dfc0c40e36", None), + "fix(#601/#602) a marker behind a title, a suffix word or a connective is an ordinary word": + _Claim(5, ('_ambiguities', 'family', 'given', 'maiden', 'suffix', 'title'), "1c9432e08a21", None), + "fix(#601) the clause ends at the clause-free name's trailing run, which the take consumes": + _Claim(4, ('_ambiguities', 'family', 'given', 'maiden', 'suffix', 'title'), "d15d58da6be1", None), + "fix(#602) a credential after the name core starts a run to the end of its part": + _Claim(5, ('_ambiguities', 'family', 'middle', 'suffix', 'title'), "3129cd9609b9", None), + "fix(#601/#602) a credential in the clause starts a run the take consumes": + _Claim(1, ('_ambiguities', 'family', 'middle', 'suffix'), "acdcc81cb17c", None), + "fix(#601) a marker behind a connective joins the connective run, and the run's initials follow": + _Claim(1, ('_initials',), "b3a89ea7087c", None), # 2026-10-04, #604: new, 3; the rules.md#P7 boundary and the # two rules.md#R4 case-repair examples. "fix(#604) the Irish and Malay patronymic particles": @@ -5511,8 +5364,9 @@ def _claim(rule: dict) -> _Claim: # case rows. Roles unmoved, `suffix` alone as before. # 2026-09-27, #544: 15 -> 18; gains 'Jane Doe Jr. Ma', 'John # Smith PhD Ed Ma', 'John Smith PhD Ma'. + # 2026-10-04, #601/#602: 18 -> 17; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#436/#437) a space-separated post-nominal run renders with spaces, not commas": - _Claim(18, ('suffix',), "60ec2103c480", None), + _Claim(17, ('suffix',), "8f6fa4cee3c3", None), # #449's six rules, second in every 2.x ledger. The # alternation reached twenty-two corpus names at landing # (twenty-three since #544, 2026-09-27) and @@ -5768,8 +5622,9 @@ def _claim(rule: dict) -> _Claim: # LEAVING this rule's alternation -- the title chain now ends # the clause before the credential is asked, so the name has # its own fix(#535) rule instead. + # 2026-10-04, #601/#602: 17 -> 11; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#533) a credential ending a maiden clause reads as a credential": - _Claim(17, ('_ambiguities', 'maiden', 'suffix'), '6bc9b2ac9772', None), + _Claim(11, ('_ambiguities', 'maiden', 'suffix'), "9f217cc8fda4", None), # The declining half: fourteen corpus names at this baseline, # eight at 2.0.0 and 2.1.0. `_ambiguities` alone, so a role # appearing here is this rule reaching a name whose clause @@ -5780,16 +5635,9 @@ def _claim(rule: dict) -> _Claim: # only the report. # 2026-09-27, #544: 15 -> 16; gains 'Jane Doe Jr. nee Smith # Ma'. + # 2026-10-04, #601/#602: 16 -> 6; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#533) the maiden clause reports the credential it keeps": - _Claim(16, ('_ambiguities',), "de33054cd9e5", None), - # New rule (#535): one corpus name, 'Doe, Jane nee Smith V, - # PhD' -- after a family comma the given slot reads a lone - # numeral as a suffix only where the given part is the LAST - # comma part (#144), asked of the maiden clause too since this - # review; a third comma part behind it withdraws the release, - # moving `middle` and `maiden`. - "fix(#535) a given-slot numeral with a credential tail stays": - _Claim(1, ('maiden', 'middle'), 'c668f28295a7', None), + _Claim(6, ('_ambiguities',), "fe166ed5ead1", None), # One corpus name. The by-shape member has no lean to read, # so a growth here is the rule reaching a LISTED member -- # a different reading under this rule's sentence. @@ -5840,9 +5688,9 @@ def _claim(rule: dict) -> _Claim: # 2026-09-26, #538: 3 -> 4, rules.md#M2's second live example, # 'Smith, John, PhD née Puig - i Soler', joining the same # alternation for the same reason. Verified name by name. + # 2026-10-04, #601/#602: 4 -> 1; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#397) a link inside a maiden clause stays in the birth name": - _Claim(4, ('family', 'maiden', 'middle', 'suffix'), - '4556aa8f93e5', ('DEFAULT',)), + _Claim(1, ('family', 'maiden', 'middle'), "4e62e18718a4", ('DEFAULT',)), # New rule (#533/#535): one corpus name, 'Jane Doe nee Smith # Prof. MA'. #533 already gives 'MA' up to the credential # reading and reports the fork (`suffix`, `maiden`, @@ -5859,49 +5707,9 @@ def _claim(rule: dict) -> _Claim: # literal list, minus the one name above whose diff from this # baseline has an earlier cause too. `title`, `maiden`, # `suffix` and `_ambiguities` move between them. + # 2026-10-04, #601/#602: 9 -> 7; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#535) a trailing title ends the maiden clause": - _Claim(9, ('_ambiguities', 'maiden', 'suffix', 'title'), - '0b7fa8ce0b7b', None), - # New rule (#535): one corpus name, 'Berg, abdul nee Smith V', - # the numeral-join literal. #535 is the sole cause here (the - # parent tree d9d80492 already reads it identically to this - # baseline before #535 ran). - "fix(#535) the numeral stop asks the join question": - _Claim(1, ('given', 'maiden'), 'c8d3b254b804', None), - # 2026-09-26: rules.md#M2's Accepted examples added 'Jane Doe - # nee Smith King.' to the title-stop rule, and the - # two rules below hold the Accepted 'ba' pair, one name each. - "fix(#342/#535) a title behind a credential the clause keeps still leaves it": - _Claim(1, ('_ambiguities', 'family', 'given', 'maiden', 'middle', - 'title'), 'b8aa7d51a16c', None), - "fix(#342/#533) a title in front of a credential the clause keeps stays in it": - _Claim(1, ('_ambiguities', 'maiden', 'suffix'), 'cb346cb8419a', None), - # 2026-09-26: one-name rules for rules.md#M2's second-round - # Accepted examples (the particle pair, the particle-chain - # title pair, and 'Dr. nee V' at 1.4.0). - "fix(#535) a particle in front of a trailing title stays in the clause": - _Claim(1, ('_ambiguities', 'maiden', 'title'), - 'a1484fe026d1', None), - "fix(#533/#535) a particle behind a trailing title leaves the clause with it": - _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), - '97e10b3d2aa8', None), - "fix(#316) a trailing title behind a maiden clause's first suffix word": - _Claim(1, ('family', 'middle', 'suffix', 'title'), - 'd58b847181aa', None), - # 2026-09-26: one-name rules for rules.md#M2's - # Accepted examples (the particle-inside-the-clause pair, and - # 'Doe nee Smith V, Jane' at 2.0.0/2.1.0). - "fix(#535) a particle inside the clause keeps the credential in front of a trailing title": - _Claim(1, ('_ambiguities', 'maiden', 'title'), - '647659c5eafc', None), - "fix(#533/#535) a particle inside the clause lets a credential behind the title go": - _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), - '8db3f5842115', None), - # 2026-09-26: the one-name rule for rules.md#M2's example - # of a link the clause stops at and keeps. - "fix(#397/#535) a link the clause stops at keeps a run that would not read off": - _Claim(1, ('_ambiguities', 'family', 'maiden', 'middle', 'title'), - '1b74e094fed9', None), + _Claim(7, ('_ambiguities', 'maiden', 'suffix', 'title'), "0155334273c4", None), # 2026-09-27, #544: new, 5; gains 'Jane Doe, MS LAc', 'John # Smith, MD MEng', 'John Smith, MEng PhD', 'John Smith, PhD # MEng' and 1 more. @@ -5949,6 +5757,21 @@ def _claim(rule: dict) -> _Claim: _Claim(3, ('family', 'given'), "d822332b50f9", None), "fix(#597) halfwidth corner brackets enclose a nickname": _Claim(3, ('family', 'given', 'middle', 'nickname'), "d1b9b0addff3", None), + # 2026-10-04, #601/#602: new rules. + "fix(#601) a marker in the given part after a family comma is an ordinary word": + _Claim(17, ('_ambiguities', 'family', 'given', 'maiden', 'middle', 'suffix', 'title'), "c7ed65a885ac", None), + "fix(#601) a marker in a part after a suffix comma is an ordinary word": + _Claim(2, ('maiden', 'suffix'), "59dfc0c40e36", None), + "fix(#601/#602) a marker behind a title, a suffix word or a connective is an ordinary word": + _Claim(5, ('_ambiguities', 'family', 'given', 'maiden', 'suffix', 'title'), "1c9432e08a21", None), + "fix(#601) the clause ends at the clause-free name's trailing run, which the take consumes": + _Claim(4, ('_ambiguities', 'maiden', 'suffix', 'title'), "d15d58da6be1", None), + "fix(#601) the first word after the marker is the maiden name": + _Claim(1, ('_ambiguities', 'family', 'maiden', 'middle', 'suffix'), "aaf53040b071", None), + "fix(#602) a credential after the name core starts a run to the end of its part": + _Claim(5, ('_ambiguities', 'family', 'middle', 'suffix', 'title'), "3129cd9609b9", None), + "fix(#601/#602) a credential in the clause starts a run the take consumes": + _Claim(1, ('_ambiguities', 'family', 'middle', 'suffix'), "acdcc81cb17c", None), # 2026-10-04, #604: new, 3; the rules.md#P7 boundary and the # two rules.md#R4 case-repair examples. "fix(#604) the Irish and Malay patronymic particles": @@ -6001,8 +5824,9 @@ def _claim(rule: dict) -> _Claim: # case rows. Roles unmoved, `suffix` alone as before. # 2026-09-27, #544: 15 -> 18; gains 'Jane Doe Jr. Ma', 'John # Smith PhD Ed Ma', 'John Smith PhD Ma'. + # 2026-10-04, #601/#602: 18 -> 17; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#436/#437) a space-separated post-nominal run renders with spaces, not commas": - _Claim(18, ('suffix',), "60ec2103c480", None), + _Claim(17, ('suffix',), "8f6fa4cee3c3", None), # #449's six rules, second in every 2.x ledger. The # alternation reached twenty-two corpus names at landing # (twenty-three since #544, 2026-09-27) and @@ -6134,6 +5958,8 @@ def _claim(rule: dict) -> _Claim: # Smith, Jr do', #554's rules.md#C1 examples. Reach, verified # name by name. _Claim(32, ('_ambiguities', 'family', 'middle'), "3e3ff4af0060", None), + "fix(#379) a tussenvoegsel behind a post-nominal after a family comma attaches": + _Claim(1, ('family', 'middle'), "f24e5eee3cc1", None), "fix(#424) an unlisted abbreviation is as transparent as a listed title to the leading particle": _Claim(1, ('_ambiguities', 'family', 'given'), "ca7b37af6cf8", None), "fix(#367) a title no longer displaces a leading particle out of the leading position": @@ -6144,8 +5970,9 @@ def _claim(rule: dict) -> _Claim: # boundary, which attaches as this rule reads. Reach, verified. "fix(#380) a trailing vd after a family comma is the tussenvoegsel, not a post-nominal": _Claim(4, ('_ambiguities', 'family', 'suffix'), "25798c4e2a22", None), + # 2026-10-04, #601/#602: 6 -> 7; 'Mai Le née Nguyen', a new rules.md#M2 example, is in reach -- at this baseline the particle chain swallowed its marker, the #399 shape. "fix(#399) a maiden marker bounds the particle chain that swallowed it": - _Claim(6, ('family', 'maiden'), "89e1f4afdd2a", ('DEFAULT',)), + _Claim(7, ('family', 'maiden'), "0a306405fb9a", ('DEFAULT',)), "fix(#360) mc moved into the never-given particles, so it folds into the family": _Claim(1, ('_ambiguities', 'family', 'given'), "ee4339908f4d", None), "fix(#400) abd joins the word after it as one given name": @@ -6160,8 +5987,9 @@ def _claim(rule: dict) -> _Claim: # 2026-09-26, #535: 2 -> 3, 'Berg, abdul nee Smith V' joining # the corpus -- the same bound-given/maiden shape. Reach, not # explanation. + # 2026-10-04, #601/#602: 3 -> 1; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#411) the bound-given reserve stops counting words the maiden name takes": - _Claim(3, ('given', 'maiden', 'middle'), "9305df1c05c5", None), + _Claim(1, ('given', 'maiden', 'middle'), "431d3dd24c14", None), "fix(#412) a connective join no longer absorbs the maiden marker beside it": _Claim(2, ('family', 'maiden'), "51c0eb36b5c5", None), "fix(#418) the connective carve-out counts the name the maiden clause leaves behind": @@ -6258,12 +6086,11 @@ def _claim(rule: dict) -> _Claim: # case-insensitively and correctly, the two being one # given->family shape written two ways. Reach, not # explanation. + # 2026-10-04, #601/#602: 7 -> 5; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#445) a maiden marker makes the lone name word the family": - _Claim(7, ('_ambiguities', 'family', 'given'), "73c2f924d257", None), + _Claim(5, ('_ambiguities', 'family', 'given'), "2648aa85b6b4", None), "fix(#445) the lone name word beside a marker a connective join no longer absorbs": _Claim(1, ('family', 'given', 'maiden', 'middle'), "52544a41dd62", None), - "fix(#424) a marker followed only by the numeral is just a word": - _Claim(1, ('_ambiguities', 'family', 'maiden', 'middle', 'suffix'), "aaf53040b071", None), "fix(#369) the bound given-name join takes a particle-and-bound word, so no fork is reported": _Claim(1, ('_ambiguities',), "81cf02ffdb33", None), "fix(#360) ste moved into the never-given particles with mc": @@ -6463,8 +6290,9 @@ def _claim(rule: dict) -> _Claim: # LEAVING this rule's alternation -- the title chain now ends # the clause before the credential is asked, so the name has # its own fix(#535) rule instead. + # 2026-10-04, #601/#602: 14 -> 10; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#533) a credential ending a maiden clause reads as a credential": - _Claim(14, ('_ambiguities', 'maiden', 'suffix'), 'e5cd5fb89c8b', None), + _Claim(10, ('_ambiguities', 'maiden', 'suffix'), "e05c5164668c", None), # The declining half: fourteen corpus names at 2.2.0, eight # at 2.0.0 and 2.1.0. `_ambiguities` alone, so a role # appearing here is this rule reaching a name whose clause @@ -6475,8 +6303,9 @@ def _claim(rule: dict) -> _Claim: # only the report. # 2026-09-27, #544: 9 -> 10; gains 'Jane Doe Jr. nee Smith # Ma'. + # 2026-10-04, #601/#602: 10 -> 4; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#533) the maiden clause reports the credential it keeps": - _Claim(10, ('_ambiguities',), "7c30ed523806", None), + _Claim(4, ('_ambiguities',), "c57e8eb30830", None), # One corpus name. The by-shape member has no lean to read, # so a growth here is the rule reaching a LISTED member -- # a different reading under this rule's sentence. @@ -6489,12 +6318,6 @@ def _claim(rule: dict) -> _Claim: # being one. "fix(#533) restores the suffix reading a 2.4 retag had moved into the maiden name": _Claim(1, ('_ambiguities',), "0a7295a0cb19", None), - # One corpus name. `suffix` is NOT in the 1.4.0 roles and - # that is the point of the row: v1 read the member as a - # post-nominal and so does the tree, which is what #533 - # restored here. - "fix(#445/#533) the clause gives the credential up, and the one name word it leaves is the family": - _Claim(1, ('_ambiguities', 'family', 'given', 'maiden', 'suffix'), "6403e56cea29", None), # One corpus name, at 2.0.0 and 2.1.0 only. Case-sensitive by # construction, so a growth here is the anchor having picked # up the mixed-case spelling, which reads the other way. @@ -6519,10 +6342,6 @@ def _claim(rule: dict) -> _Claim: _Claim(1, ('_ambiguities', 'maiden', 'suffix'), "6bab87214ddf", None), "fix(#335/#533) a marker-led bracketed clause is the maiden name whatever its last word is": _Claim(3, ('_ambiguities', 'maiden', 'nickname'), 'cc1045ecdc05', None), - "fix(#399/#533) a maiden marker bounds the particle chain, and the clause keeps the credential the chain would have taken": - _Claim(1, ('_ambiguities', 'family', 'maiden', 'middle'), 'fa3fe4878b47', None), - "fix(#411/#533) the bound-given pair after a comma, and the clause it keeps its credential in": - _Claim(1, ('_ambiguities', 'given', 'maiden', 'middle'), '431d3dd24c14', None), # 2026-09-20, #397 and #461's four new rules, all literal # alternations of exactly the names each explains -- a count # that moves means the alternation grew, and _MUST_NOT_MATCH @@ -6550,21 +6369,9 @@ def _claim(rule: dict) -> _Claim: # 2026-09-26, #538: 3 -> 4, rules.md#M2's second live example, # 'Smith, John, PhD née Puig - i Soler', joining the same # alternation for the same reason. Verified name by name. + # 2026-10-04, #601/#602: 4 -> 1; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#397) a link inside a maiden clause stays in the birth name": - _Claim(4, ('family', 'maiden', 'middle', 'suffix'), - '4556aa8f93e5', ('DEFAULT',)), - # New rule (#399): one corpus name, 'Jane van der Berg nee - # Smith Prof.' -- a sibling of the general #399 rule above, - # for the two-trailing-word shape its own anchor - # (`\\sn[ée]e\\s+\\S+$`) cannot reach: per rules.md#P2, "a - # maiden marker takes the words after it (M2), or the name - # ends", and the chain here swallows both trailing words. Its - # diff from this baseline is #399's alone, verified against - # the parent tree d9d80492: before #535 ran, it already read - # family 'van der Berg', maiden 'Smith Prof.', identically to - # HEAD. - "fix(#399) a maiden marker bounds the particle chain that swallowed it, two trailing words": - _Claim(1, ('family', 'maiden'), '6b2d1fa71195', None), + _Claim(1, ('family', 'maiden', 'middle'), "4e62e18718a4", ('DEFAULT',)), # New rule (#296/#535): one corpus name, 'Jane Doe nee Prof. # Dr.'. This baseline still has 'dr' in the suffix vocabulary # (#296 had not run yet), so it reads suffix 'Dr.'. The parent @@ -6590,59 +6397,9 @@ def _claim(rule: dict) -> _Claim: # this baseline has an earlier cause too. `title`, `maiden` # and `suffix` move for the credential-bearing ones, and # `_ambiguities` for the declined-member one. + # 2026-10-04, #601/#602: 8 -> 6; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#535) a trailing title ends the maiden clause": - _Claim(8, ('_ambiguities', 'maiden', 'suffix', 'title'), - '3216299d78a6', None), - # New rule (#411/#535): one corpus name, 'Berg, abdul nee - # Smith V', relabelled from a bare '#535' rule (review): #411 - # (the bound-given reserve) already decides `given`/`maiden` - # at this baseline before #535 exists, verified against the - # parent tree d9d80492; #535 then moves the released V from - # given back into maiden. - "fix(#411/#535) the numeral stop asks the join question": - _Claim(1, ('given', 'maiden', 'middle', 'suffix'), - 'c8d3b254b804', None), - # 2026-09-26: rules.md#M2's Accepted examples added 'Jane Doe - # nee Smith King.' to the title-stop rule, and the - # two rules below hold the Accepted 'ba' pair, one name each. - "fix(#342/#445/#535) a title behind a credential the clause keeps still leaves it": - _Claim(1, ('_ambiguities', 'family', 'given', 'maiden', 'middle', - 'title'), 'b8aa7d51a16c', None), - "fix(#342/#445/#533) a title in front of a credential the clause keeps stays in it": - _Claim(1, ('_ambiguities', 'family', 'given', 'maiden', 'suffix'), - 'cb346cb8419a', None), - # 2026-09-26: one-name rules for rules.md#M2's second-round - # Accepted examples (the particle pair, the particle-chain - # title pair, and 'Dr. nee V' at 1.4.0). - "fix(#535) a particle in front of a trailing title stays in the clause": - _Claim(1, ('_ambiguities', 'maiden', 'title'), - 'a1484fe026d1', None), - "fix(#533/#535) a particle behind a trailing title leaves the clause with it": - _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), - '97e10b3d2aa8', None), - "fix(#316/#399) a trailing title behind a maiden clause's first suffix word": - _Claim(1, ('family', 'maiden', 'middle', 'suffix', 'title'), - 'd58b847181aa', None), - "fix(#399) a maiden marker bounds the particle chain ahead of a title the clause keeps": - _Claim(1, ('family', 'maiden'), - '58cf2330fed6', None), - # 2026-09-26: one-name rules for rules.md#M2's - # Accepted examples (the particle-inside-the-clause pair, and - # 'Doe nee Smith V, Jane' at 2.0.0/2.1.0). - "fix(#535) a particle inside the clause keeps the credential in front of a trailing title": - _Claim(1, ('_ambiguities', 'maiden', 'title'), - '647659c5eafc', None), - "fix(#533/#535) a particle inside the clause lets a credential behind the title go": - _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), - '8db3f5842115', None), - "fix(#424) accepted: before a family comma the numeral the walk gives up goes to the family": - _Claim(1, ('family', 'maiden'), - '339279b78f2f', None), - # 2026-09-26: the one-name rule for rules.md#M2's example - # of a link the clause stops at and keeps. - "fix(#397/#535) a link the clause stops at keeps a run that would not read off": - _Claim(1, ('_ambiguities', 'family', 'maiden', 'middle', 'title'), - '1b74e094fed9', None), + _Claim(6, ('_ambiguities', 'maiden', 'suffix', 'title'), "3e41b335baa4", None), # 2026-09-27, #544: new, 5; gains 'Jane Doe, MS LAc', 'John # Smith, MD MEng', 'John Smith, MEng PhD', 'John Smith, PhD # MEng' and 1 more. @@ -6689,6 +6446,21 @@ def _claim(rule: dict) -> _Claim: _Claim(3, ('family', 'given'), "d822332b50f9", None), "fix(#597) halfwidth corner brackets enclose a nickname": _Claim(3, ('family', 'given', 'middle', 'nickname'), "d1b9b0addff3", None), + # 2026-10-04, #601/#602: new rules. + "fix(#601) a marker in the given part after a family comma is an ordinary word": + _Claim(17, ('_ambiguities', 'family', 'given', 'maiden', 'middle', 'suffix', 'title'), "c7ed65a885ac", None), + "fix(#601) a marker in a part after a suffix comma is an ordinary word": + _Claim(2, ('maiden', 'suffix'), "59dfc0c40e36", None), + "fix(#601/#602) a marker behind a title, a suffix word or a connective is an ordinary word": + _Claim(5, ('_ambiguities', 'family', 'given', 'maiden', 'suffix', 'title'), "1c9432e08a21", None), + "fix(#601) the clause ends at the clause-free name's trailing run, which the take consumes": + _Claim(4, ('_ambiguities', 'family', 'given', 'maiden', 'suffix', 'title'), "d15d58da6be1", None), + "fix(#602) a credential after the name core starts a run to the end of its part": + _Claim(5, ('_ambiguities', 'family', 'middle', 'suffix', 'title'), "3129cd9609b9", None), + "fix(#601/#602) a credential in the clause starts a run the take consumes": + _Claim(1, ('_ambiguities', 'family', 'middle', 'suffix'), "acdcc81cb17c", None), + "fix(#601) a marker behind a connective joins the connective run, and the run's initials follow": + _Claim(1, ('_initials',), "b3a89ea7087c", None), # 2026-10-04, #604: new, 3; the rules.md#P7 boundary and the # two rules.md#R4 case-repair examples. "fix(#604) the Irish and Malay patronymic particles": @@ -6862,24 +6634,18 @@ def _claim(rule: dict) -> _Claim: # 2.2.0, this baseline never splits the credential at all, so # it joins the movers here instead of the declining half # below; the whole diff is #533's and #535 moves nothing. + # 2026-10-04, #601/#602: 18 -> 11; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#533) a credential ending a maiden clause reads as a credential": - _Claim(18, ('_ambiguities', 'maiden', 'suffix'), '537d3586cd79', None), + _Claim(11, ('_ambiguities', 'maiden', 'suffix'), "9f217cc8fda4", None), # The declining half: fourteen corpus names at this baseline, # eight at 2.0.0 and 2.1.0. `_ambiguities` alone, so a role # appearing here is this rule reaching a name whose clause # gave the member up. # 2026-09-27, #544: 15 -> 16; gains 'Jane Doe Jr. nee Smith # Ma'. + # 2026-10-04, #601/#602: 16 -> 6; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#533) the maiden clause reports the credential it keeps": - _Claim(16, ('_ambiguities',), "b34e76bea15b", None), - # New rule (#535): one corpus name, 'Doe, Jane nee Smith V, - # PhD' -- after a family comma the given slot reads a lone - # numeral as a suffix only where the given part is the LAST - # comma part (#144), asked of the maiden clause too since this - # review; a third comma part behind it withdraws the release, - # moving `middle` and `maiden`. - "fix(#535) a given-slot numeral with a credential tail stays": - _Claim(1, ('maiden', 'middle'), 'c668f28295a7', None), + _Claim(6, ('_ambiguities',), "fe166ed5ead1", None), # One corpus name. The by-shape member has no lean to read, # so a growth here is the rule reaching a LISTED member -- # a different reading under this rule's sentence. @@ -6930,9 +6696,9 @@ def _claim(rule: dict) -> _Claim: # 2026-09-26, #538: 3 -> 4, rules.md#M2's second live example, # 'Smith, John, PhD née Puig - i Soler', joining the same # alternation for the same reason. Verified name by name. + # 2026-10-04, #601/#602: 4 -> 1; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#397) a link inside a maiden clause stays in the birth name": - _Claim(4, ('family', 'maiden', 'middle', 'suffix'), - '4556aa8f93e5', ('DEFAULT',)), + _Claim(1, ('family', 'maiden', 'middle'), "4e62e18718a4", ('DEFAULT',)), # New rule (#533/#535): one corpus name, 'Jane Doe nee Smith # Prof. MA'. #533 already gives 'MA' up to the credential # reading and reports the fork (`suffix`, `maiden`, @@ -6949,38 +6715,9 @@ def _claim(rule: dict) -> _Claim: # literal list, minus the one name above whose diff from this # baseline has an earlier cause too. `title`, `maiden`, # `suffix` and `_ambiguities` move between them. + # 2026-10-04, #601/#602: 10 -> 7; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. "fix(#535) a trailing title ends the maiden clause": - _Claim(10, ('_ambiguities', 'maiden', 'suffix', 'title'), - '67b72e8f082f', None), - # New rule (#535): one corpus name, 'Berg, abdul nee Smith V', - # the numeral-join literal. #535 is the sole cause here (the - # parent tree d9d80492 already reads it identically to this - # baseline before #535 ran). - "fix(#535) the numeral stop asks the join question": - _Claim(1, ('given', 'maiden'), 'c8d3b254b804', None), - # 2026-09-26: one-name rules for rules.md#M2's second-round - # Accepted examples (the particle pair, the particle-chain - # title pair, and 'Dr. nee V' at 1.4.0). - "fix(#535) a particle in front of a trailing title stays in the clause": - _Claim(1, ('_ambiguities', 'maiden', 'title'), - 'a1484fe026d1', None), - "fix(#533/#535) a particle behind a trailing title leaves the clause with it": - _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), - '97e10b3d2aa8', None), - # 2026-09-26: one-name rules for rules.md#M2's - # Accepted examples (the particle-inside-the-clause pair, and - # 'Doe nee Smith V, Jane' at 2.0.0/2.1.0). - "fix(#535) a particle inside the clause keeps the credential in front of a trailing title": - _Claim(1, ('_ambiguities', 'maiden', 'title'), - '647659c5eafc', None), - "fix(#533/#535) a particle inside the clause lets a credential behind the title go": - _Claim(1, ('_ambiguities', 'maiden', 'suffix', 'title'), - '8db3f5842115', None), - # 2026-09-26: the one-name rule for rules.md#M2's example - # of a link the clause stops at and keeps. - "fix(#397/#535) a link the clause stops at keeps a run that would not read off": - _Claim(1, ('_ambiguities', 'family', 'maiden', 'middle', 'title'), - '1b74e094fed9', None), + _Claim(7, ('_ambiguities', 'maiden', 'suffix', 'title'), "0155334273c4", None), # 2026-09-27, #544: new, 5; gains 'Jane Doe, MS LAc', 'John # Smith, MD MEng', 'John Smith, MEng PhD', 'John Smith, PhD # MEng' and 1 more. @@ -7026,6 +6763,21 @@ def _claim(rule: dict) -> _Claim: _Claim(3, ('family', 'given'), "d822332b50f9", None), "fix(#597) halfwidth corner brackets enclose a nickname": _Claim(3, ('_ambiguities', 'family', 'given', 'middle', 'nickname'), "d1b9b0addff3", None), + # 2026-10-04, #601/#602: new rules. + "fix(#601) a marker in the given part after a family comma is an ordinary word": + _Claim(17, ('_ambiguities', 'family', 'given', 'maiden', 'middle', 'suffix', 'title'), "c7ed65a885ac", None), + "fix(#601) a marker in a part after a suffix comma is an ordinary word": + _Claim(2, ('maiden', 'suffix'), "59dfc0c40e36", None), + "fix(#601/#602) a marker behind a title, a suffix word or a connective is an ordinary word": + _Claim(5, ('_ambiguities', 'family', 'given', 'maiden', 'suffix'), "1c9432e08a21", None), + "fix(#601) the clause ends at the clause-free name's trailing run, which the take consumes": + _Claim(4, ('_ambiguities', 'maiden', 'suffix', 'title'), "d15d58da6be1", None), + "fix(#601) the first word after the marker is the maiden name": + _Claim(1, ('_ambiguities', 'family', 'maiden', 'middle', 'suffix'), "aaf53040b071", None), + "fix(#602) a credential after the name core starts a run to the end of its part": + _Claim(4, ('_ambiguities', 'family', 'middle', 'suffix', 'title'), "e638a392a3a4", None), + "fix(#601/#602) a credential in the clause starts a run the take consumes": + _Claim(1, ('_ambiguities', 'family', 'middle', 'suffix'), "acdcc81cb17c", None), # 2026-10-04, #604: new, 3; the rules.md#P7 boundary and the # two rules.md#R4 case-repair examples. "fix(#604) the Irish and Malay patronymic particles": @@ -8778,12 +8530,9 @@ def test_a_rule_reaching_no_corpus_name_says_why_it_is_kept() -> None: # declarations say the contests are INTENDED, and these rows # say how much corpus they cover, so a rule quietly growing # into a fourth pair fails here. - ("fix(#274/#533) the clause keeps the credential a join or a missing reader would have taken", - "fix(comma-precomma-family) pre-comma run reads as family, not given", 4), - ("fix(#274/#533) the clause keeps the credential a join or a missing reader would have taken", - "fix(#411) the bound-given reserve stops counting words the maiden name takes", 1), - ("fix(#274/#533) the clause keeps the credential a join or a missing reader would have taken", - "fix(#379) a tussenvoegsel after a family comma attaches to the family", 2), + # The three #274/#533 pairs left on 2026-10-04 (#601): the four + # given-part names they were about lost their clause, and the + # rule its three exemptions. ("fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example", "fix(comma-precomma-family) pre-comma run reads as family, not given", 2), ("fix(#296) a credential-only comma string reads a name and its postnominal", @@ -8801,8 +8550,8 @@ def test_a_rule_reaching_no_corpus_name_says_why_it_is_kept() -> None: "fix(suffix-routing) a two-token name ending in the suffix word jr keeps it in `suffix`", 2), ("fix(#400/#274) bound-given join and maiden consumption in one name", "fix(#400) abd joins the word after it as one given name", 1), - ("fix(#411/S2) a declining bound-given join leaves the suffix reading after a family comma", - "fix(#400) abd joins the word after it as one given name", 1), + # ('fix(#411/S2) ...' against fix(#400) left with its rule on + # 2026-10-04, #601: 'Berg, abd née Jones' has no clause now.) ("fix(#272/#308) nakaguro division and a glued hangul honorific in one name", "fix(cjk-glued-honorific-peel) glued honorific peels into suffix", 1), ("fix(nickname-typographic-pairs) two typographic quote spans read as one nickname set", diff --git a/tests/v2/test_parser.py b/tests/v2/test_parser.py index 06b31e25..340b1b13 100644 --- a/tests/v2/test_parser.py +++ b/tests/v2/test_parser.py @@ -570,10 +570,12 @@ def test_the_bound_given_join_leaves_a_suffix_where_it_stands() -> None: n = parse("abdul Ph. D. Smith Berg") assert (n.given, n.middle, n.family, n.suffix) == \ ("abdul", "Smith", "Berg", "Ph. D.") - # after a family comma the decline holds under the LENIENT reserve + # after a family comma the decline holds under the LENIENT reserve, + # and since #602 the 'Jr' starts the given part's credential run, + # taking 'Smith' with it (rules.md#S2) n = parse("Berg, abdul Jr Smith") assert (n.given, n.middle, n.family, n.suffix) == \ - ("abdul", "Smith", "Berg", "Jr") + ("abdul", "", "Berg", "Jr Smith") def test_the_reserve_declines_and_assign_reads_the_unjoined_pieces() -> None: @@ -727,62 +729,6 @@ def test_the_chain_and_the_walk_stop_where_the_peel_begins() -> None: assert (n.maiden, n.suffix) == ("Smith V", "MA") -def test_a_post_nominal_head_leaves_the_clause_nobody_to_read_it( -) -> None: - """#533 review: the clause KEEPS the member, and says so. - - After a family comma the given-slot reader takes the member only - where the take leaves a given part for it to end. Where the part - before the marker is nothing but post-nominals, it does not: the - take would leave a segment of credentials, which is read whole - and asked nothing, and the released word would land in `given` - rather than in `suffix`. So the walk declines and the word stays - in the maiden name -- rules.md#M2's invariant, which replaced the - ACCEPTED silent mover an earlier round of this branch shipped - here (it read maiden 'Smith', suffix 'Jr MA'). - - The clause still REPORTS, and that is the measurement this test - exists for: the emitter is gated on there being a trailing rule - at all, not on the view check the rule then fails, so a fork - called the conservative way is still a fork the caller hears - about. - """ - n = parse("Jane Doe, Jr nee Smith MA") - assert (n.given, n.family, n.maiden, n.suffix) == ( - "Jane", "Doe", "Smith MA", "Jr") - assert [a.kind for a in n.ambiguities] == [ - AmbiguityKind.SUFFIX_OR_NAME] - # 2f57ff21 read it this way too: the review round restored the - # parent reading rather than inventing a third one - assert n.maiden == "Smith MA" - # a post-nominal, not only a generational word, heads it the same - for head in ("III", "PhD"): - n = parse(f"Jane Doe, {head} nee Smith MA") - assert (n.maiden, n.suffix) == ("Smith MA", head) - assert [a.kind for a in n.ambiguities] == [ - AmbiguityKind.SUFFIX_OR_NAME] - # and a TITLE heads it the same way, which is the spelling - # `AmbiguityKind.SUFFIX_OR_NAME`'s third position was written as - n = parse("Doe, Dr. nee Smith MA") - assert (n.title, n.family, n.maiden) == ("Dr.", "Doe", "Smith MA") - assert [a.kind for a in n.ambiguities] == [ - AmbiguityKind.SUFFIX_OR_NAME] - # the clause-less control is what the member WOULD have read as, - # and the difference is the point: without the clause there is a - # given name in front of the member and the slot exists - control = parse("Doe, Dr. Smith MA") - assert (control.given, control.suffix) == ("Smith", "MA") - # the KEPT direction reported before this round too -- the - # clause's own emitter is what raises it, and it never needed the - # given slot - for text, maiden in (("Jane Doe, Jr nee Smith Ma", "Smith Ma"), - ("Jane Doe, Jr nee MA", "MA")): - n = parse(text) - assert n.maiden == maiden - assert [a.kind for a in n.ambiguities] == [ - AmbiguityKind.SUFFIX_OR_NAME] - - def test_the_numeral_fork_fires_on_the_last_piece_only() -> None: # The shared peel's own contract, pinned at the stage that owns # it: a numeral with a suffix behind it is a name word, for an @@ -1271,13 +1217,21 @@ def test_revise_reads_a_glued_honorific_on_its_own() -> None: #: join. Named here so the guard below fails on a NEW exception and #: not on the known one. _HONORIFIC_PEEL = frozenset({"김민준씨, J.씨", "김민준씨., J.씨"}) +#: The second known limit (2026-10-04, #601/#602): a suffix that holds +#: a maiden marker the whole name left a word -- in a tail part, or +#: taken into #602's credential run behind a credential -- and that +#: the value's own sub-parse consumes mid-value, as revise() documents. +_MARKER_IN_SUFFIX = frozenset({"Jane Doe PhD nee Smith", + "Smith, John, MD - née Jones Smith", + "Smith, John, PhD née Jones"}) def _suffix_bearing_corpus_names() -> list[str]: from ._differential_fixtures import _CORPUS_NAMES p = Parser() return [n for n in _CORPUS_NAMES - if p.parse(n).suffix and n not in _HONORIFIC_PEEL] + if p.parse(n).suffix and n not in _HONORIFIC_PEEL + and n not in _MARKER_IN_SUFFIX] def test_the_known_round_trip_exceptions_are_one_limit() -> None: @@ -1301,6 +1255,19 @@ def ends_in_a_tail_without_being_one(name: str) -> bool: assert all(ends_in_a_tail_without_being_one(n) for n in _HONORIFIC_PEEL) assert not [n for n in _suffix_bearing_corpus_names() if ends_in_a_tail_without_being_one(n)] + # the marker limit, characterized the same way: the parsed suffix, + # read ON ITS OWN, takes a maiden clause -- its marker the whole + # name left a word, and the value's sub-parse consumes. Not every + # suffix holding a marker word: 'Jane Doe Jr. nee Smith' renders + # 'Jr. nee Smith', whose own parse reads 'Jr.' as a leading title + # (H2), so the marker stays a word there too and the value + # round-trips (measured 2026-10-04). + def takes_a_clause_on_its_own(name: str) -> bool: + return bool(p.parse(p.parse(name).suffix).maiden) + + assert all(takes_a_clause_on_its_own(n) for n in _MARKER_IN_SUFFIX) + assert not [n for n in _suffix_bearing_corpus_names() + if takes_a_clause_on_its_own(n)] def test_the_suffix_bearing_corpus_is_not_empty() -> None: @@ -1802,7 +1769,8 @@ def test_the_clause_free_corpus_is_not_empty() -> None: def test_a_maiden_clause_changes_nothing_else(name: str) -> None: """The grouping rules count and join only the words that remain once the marker and the maiden name leave (rules.md#M2, #418), so - appending a clause adds a maiden name and moves no other field. + appending a clause adds a maiden name and moves no other field -- + wherever the appended marker counts at all. Over the corpus rather than by example, because the defect was an appended-clause shape on names that are otherwise ordinary ('John @@ -1811,47 +1779,59 @@ def test_a_maiden_clause_changes_nothing_else(name: str) -> None: Before the marker pass moved ahead of the joins, seven of these names failed this. - One skip, and it is M2's own boundary: a corpus name that parses - to no name word ('', '(', '()') gives the appended marker nothing - to stand behind, so M2 leaves it a word. A name that is only a - NICKNAME is in that class too, which is why the test below asks - about five fields and not about `nickname`: '(Bud)' parses to a - nickname alone, and '(Bud) née Jones' reads given 'née', family - 'Jones' and no maiden at all -- the marker stayed a word, so - there is no clause to assert. A title-only or suffix-only name is - NOT in that class -- 'Coach née Jones' reads maiden 'Jones' -- - and is asserted like any other. #410's lone-residual shape used - to be skipped here too -- 'Dr. Jane' read family 'Jane' and 'Dr. - Jane née Smith' given 'Jane' -- and no longer moves, so the - assertion now covers every name that reaches the marker with - something to stand behind. + Whether the marker counts is M2's head rule (#601): the word + straight before it -- the base name's last word, a nickname aside + -- must be a name word, which a word of the leading title run, a + word of the unambiguous suffix vocabulary and a connective are + not. Where the clause is taken the assertion is the one above; + where it is not ('Smith Jr.', 'Coach', 'John Smith PhD') the word + before the marker must be one of those refusals, so neither half + of the split can pass vacuously. A corpus name that parses to no + name word ('', '(', '()') has nothing before the marker and is + skipped. One class of name moves ON PURPOSE, and is asserted MOVING rather than skipped: rules.md#M4 makes the clause decide a name that holds exactly one name word, because a marker announces a former - surname and only means anything beside a current one. Fourteen - of these names are in that class ('Smith', 'Smith Jr.', "'Smitty' - Jones Jr.", 'John V', 'de' ...), and the flip they assert is + surname and only means anything beside a current one. The flip is exactly given -> family with every other field standing still -- which is the whole of what #445 changed, checked over the corpus - rather than at the six rows cases.py carries. + rather than at the rows cases.py carries. The `flips` predicate below is a second reading of M4's guard, and that is a maintenance cost taken deliberately rather than the silent-drift hazard it resembles: it fails LOUDLY in both directions -- narrow M4 and the flip assertion fails, widen it and the stands-still assertion does. What it buys is the carve-out - witness. Review offered the cheaper form, asserting only that - (given, family) is either unchanged or moved wholesale, with no - predicate at all; under that form a name whose word the vocabulary - claims would be free to flip, and dropping the `vocab:bound-given` - carve-out would stop failing here on 'abdul'. The corpus is the - only place that name is asked. + witness: dropping the `vocab:bound-given` carve-out would stop + failing here on 'abdul', and the corpus is the only place that + name is asked. """ base = parse(name) if not (base.given or base.middle or base.family or base.title or base.suffix): pytest.skip("nothing before the marker at all: M2 leaves it a word") + with_clause = parse(name + " née Jones") + # The head rule, read both ways. Where the appended marker takes + # nothing, the word straight before it -- the base name's last + # word, a nickname aside -- must be one of the rule's refusals: a + # title, a word of the UNAMBIGUOUS suffix vocabulary (an ambiguous + # member or a numeral the peel happens to take is none) or the + # merged credential's continuation, or a connective. A decline for + # any other reason fails, and so does a taken clause that moves a + # field, below. + words = [t for t in base.tokens if t.role is not Role.NICKNAME] + last = words[-1] if words else None + if with_clause.maiden == "": + assert last is not None and ( + last.role is Role.TITLE + or ("vocab:suffix" in last.tags and "initial" not in last.tags) + or (last.role is Role.SUFFIX and "joined" in last.tags) + or "conjunction" in last.tags), ( + f"{name!r}: the appended marker took nothing, but the word " + f"before it, {last.text if last else None!r}, is none of " + f"rules.md#M2's refusals") + return # M4's guard, read off the base parse: one GIVEN token, no other # name word, no title (a titled name is H1's), and neither # carve-out tag. Tokens rather than fields because the rule counts @@ -1861,7 +1841,6 @@ def test_a_maiden_clause_changes_nothing_else(name: str) -> None: flips = (len(givens) == 1 and not base.middle and not base.family and not base.title and not ({"initial", "vocab:bound-given"} & givens[0].tags)) - with_clause = parse(name + " née Jones") assert with_clause.maiden == "Jones" moved = ("title", "middle", "suffix", "nickname") if flips else ( "title", "given", "middle", "family", "suffix", "nickname") diff --git a/tests/v2/test_properties.py b/tests/v2/test_properties.py index 7e353de4..8eceadfe 100644 --- a/tests/v2/test_properties.py +++ b/tests/v2/test_properties.py @@ -288,9 +288,18 @@ def test_the_comma_agreement_exceptions_are_all_still_exceptions( #: is the recorded negative control. The one-case class's 1,026 is #: unmoved by the three heads #544 added ('Jane Doe PhD', 'Doe, Jane #: PhD', 'PhD'), all three being mixed case. -_MAIDEN_ANCHORED_HEAD_EXCEPTIONS = 810 +#: Re-recorded 2026-10-04 for #602: 1,134, every one of the 324 new +#: members under the `nodot` policy -- a by-shape dotted member +#: ('X.Y.Z.', 'R.A.I.' in each case) behind one of the three +#: credential heads, 108 each, which the head's credential run (#602) +#: took in. Then EMPTIED the same day by #601, and kept at 0 rather +#: than deleted, so a member arriving fails here: a marker behind a +#: credential is an ordinary word now (rules.md#M2), the credential's +#: run takes the whole clause in, and the clause form reads the member +#: as a credential exactly as the plain form does. +_MAIDEN_ANCHORED_HEAD_EXCEPTIONS = 0 _MAIDEN_ANCHORED_HEAD_DIGEST = ( - "c5a8f5277e49fe583e0edc834346f623dd02459845ccd6ab9325ecfe7a167c7c") + "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855") _MAIDEN_AGREEMENT_EXCEPTIONS = 1026 _MAIDEN_AGREEMENT_DIGEST = ( "4b70727a2633fea1a9c219173d48b223cc0866a5d340b69a9eb4f14fc386ae6a") @@ -544,7 +553,13 @@ def test_a_maiden_clause_does_not_change_how_a_trailing_word_reads( either one alone removed, 0; 0 here (measured 2026-09-28). """ members = ("ba", "do", "ed", "jd", "ma", "x.y.z.", "r.a.i.") - heads = ("Jane Doe", "Doe, Jane", "John", "J.", "Dr.", "Jane", + # 'Dr.' left the heads with #601: a marker behind a lone title is + # an ordinary word, so its pairs compared a name holding one more + # word to spare against one without it (648 pairs disagreed, every + # one on that head, measured 2026-10-04). Every other head agrees + # -- the comma heads, whose marker is a word too, and the heads + # ending in a credential, whose run (#602) takes the clause in. + heads = ("Jane Doe", "Doe, Jane", "John", "J.", "Jane", "Jane van der Berg", "JANE DOE", "jane doe", "DOE, JANE", "doe, jane", "Jane Q. Doe", "Doe, Dr. Jane", "Doe, J.", "Smith, Jane", "Jane Doe Jr.", "Jane Doe PhD", @@ -682,17 +697,19 @@ def audit(label: str, text: str, name: ParsedName) -> None: def _maiden_clause_grid() -> list[tuple[str, Parser, str]]: """(name, parser, policy label) for the M2 release grid. - Rich enough to hold every shape the #533 review found: title-led - and post-nominal-led comma heads, particle heads, a bound-given - head, by-shape and caps-on members, two adjacent particle members, - mixed/ALL-CAPS/lower writing, comma and no-comma, four markers, - and the default policy beside each 2.4 switch. + Heads where a marker counts (rules.md#M2, #601): particle heads, a + bound-given head and pair, a title-led head, a lone name word and a + Vietnamese head whose surname is particle vocabulary, with by-shape + and caps-on members, two adjacent particle members, mixed/ALL-CAPS/ + lower writing, a suffix-comma tail beside none, four markers, and + the default policy beside each 2.4 switch. Until #601 it also held + the #533 review's comma heads ('Doe, Prof.', 'Berg, abdul', ...), + where a marker is an ordinary word now; a grid of them could no + longer reach the take at all. """ - heads = ("Jane Doe", "Doe, Jane", "Doe, Prof.", "Doe, Dr.", - "Jane Doe, Jr", "Jane Doe, PhD", "Doe, J.", "Doe, PhD", - "Berg, Jane van der", "Jane van der Berg", "Berg, abdul", - "abdul Berg", "J. Doe", "Doe", "Prof. Jane Doe", - "Doe, Jane van der", "Doe, Sir") + heads = ("Jane Doe", "Jane van der Berg", "abdul Berg", "J. Doe", + "Doe", "Prof. Jane Doe", "Jane Q. Doe", "abdul rahman", + "Mai Le") bodies = ("Smith", "Smith MA", "Smith Ma", "Smith ma", "Smith A.B.", "Smith X.Y.Z.", "Smith XYZ", "Smith ba", "Smith DO", "Smith Do", "Smith do", "Smith MA XYZ", "Smith DO DO", @@ -763,9 +780,16 @@ def test_a_word_the_clause_gives_up_lands_in_suffix() -> None: other direction. The grid is the pin, and it is a grid rather than a name list so - that it could fail. At d97d3eb7 -- the commit this review round - started from -- it fails on 790 of its 8,466 parses, 290 distinct - names, covering every shape the review reported: 'Doe, Prof. nee + that it could fail. Since #601 the invariant holds by construction + -- the take hands the run to the suffix and title roles itself -- + so the RECORDED NEGATIVE CONTROL is the design that construction + replaced: prototype variant 4 of #601 (the head rule, the run read + once over the name as written, and NOT bound) fails 350 of this + grid's 4,482 parses, 512 tokens across 122 distinct texts; this + tree fails none (measured 2026-10-04). The history, over the + pre-#601 grid of 17 heads and 8,466 parses: at d97d3eb7 -- the + commit the #533 review round started from -- it failed on 790 of + them, 290 distinct names, covering every shape the review reported: 'Doe, Prof. nee Smith A.B.' reading given 'A.B.' (32 parses, and 'X.Y.Z.' another 32), 'DOE, PROF. NEE SMITH MA' reading given 'MA' (206, the largest class, with 'ba' at 96), 'Berg, Jane van der nee Smith @@ -842,25 +866,22 @@ def test_a_trailing_title_is_transparent_to_the_maiden_clause() -> None: def test_a_title_the_clause_gives_up_lands_in_title() -> None: - """rules.md#M2's invariant for the title stop (#535): a + """rules.md#M2's invariant for a title in the run (#535, #601): a period-marked title word written after the marker ends the parse in the maiden name or in the title field, never in a name part. - Over heads reaching all three readers and the guards: no comma - (TRAILING), the given part after a family comma (GIVEN_SLOT), a - tail segment behind a suffix comma ('Jane Doe, PhD', NONE), a bare - title head (the view check), a particle head (the chain), a - bound-given head (P5). No head puts the clause before a family - comma, the other place NONE reads. It holds at e0f1a2fa too, - where the clause kept every title, so its control is a mutation: - RECORDED NEGATIVE CONTROL, re-measured 2026-09-26 after the link - and particle-title bodies joined: the title stop's release check - replaced by `if True:` fails 83 tokens across 78 of these 240 - texts (54 across 49 of the first 160). 240 parses, about 0.03s on - CPython 3.11 (measured the same day). + Over heads where a marker counts (#601): a plain head, an initial, + a particle head (the chain a released title would join) and a + bound-given head (P5). The comma heads and the bare title head the + #535 version held put the marker where it is an ordinary word now, + so they asked nothing of a clause. RECORDED NEGATIVE CONTROL, + measured 2026-10-04: prototype variant 4 of #601 (the run read + once over the name as written, and not bound) fails 35 tokens + across 32 of these 120 texts; this tree none. The #535 figure, + over that version's 240 texts: the title stop's release check + replaced by `if True:` failed 83 tokens across 78 of them. """ - heads = ("Jane Doe", "Doe, Jane", "Doe, Prof.", "Dr.", "J.", - "Jane van der Berg", "Berg, abdul", "Jane Doe, PhD") + heads = ("Jane Doe", "J.", "Jane van der Berg", "abdul Berg") bodies = ("Smith Prof.", "Smith MA Prof.", "Smith Prof. MA", "Smith V Prof.", "Smith Ma Prof.", "Smith Prof. Dr.", "Prof.", "Prof. Dr.", "Smith King.", "Smith MA do Prof.", @@ -894,16 +915,17 @@ def test_a_title_first_word_counts_as_a_word() -> None: clauses cannot pass by agreeing about nothing. An invariant over two INPUTS (docs/design/AGENTS.md axis 11), so - it consults no rule statement. RECORDED NEGATIVE CONTROL, measured - 2026-09-26 on a copy of this tree with the floor removed entirely - (`tail_reading` handed 1 in `_maiden_take`): every one of the 30 - pairs fails, on the head check alone -- the chain takes 'King.', - the clause declines, and the suffixes still agree (0 of 30 - disagree), which is why the head check is here. This replaces a - figure recorded for a clamp-based variant, which a later attempt - could not reproduce from its description. + it consults no rule statement. The second form puts a suffix comma + after the clause, where a trailing rule still reads it; until #601 + it was a family-comma head, where the marker is an ordinary word + now. RECORDED NEGATIVE CONTROL, measured 2026-10-04: with the + first word after the marker let into the view `_maiden_take` reads + the run over (its two `lo + 1` bounds made `lo`), every one of the + 30 pairs fails, on the head check alone -- the suffixes still + agree, which is why the head check is here. The same profile was + recorded on 2026-09-26 for the floor `tail_reading` then carried. """ - forms = ("Jane Doe nee {} {}", "Doe, Jane nee {} {}") + forms = ("Jane Doe nee {} {}", "Jane Doe nee {} {}, MD") runs = ("ba", "MA", "V", "PhD", "MA JD") heads = ("King.", "Smith") failures = [] @@ -1120,17 +1142,34 @@ def test_no_two_ambiguities_name_the_same_token_span() -> None: assert not failures, ( f"{len(failures)} parse(s) report one token twice:\n" + "\n".join(failures[:15])) - assert multi == 4320, ( + # 4,239 since 2026-10-04 (#601/#602), from 4,320: over the member + # grid 2,052 multi-report parses left -- every one on the heads + # 'Doe, Jane van' and 'Doe, Jane van der', whose clause in the + # given part no longer takes, and with it the maiden report beside + # P6's -- and 1,965 arrived, 1,134 on 'Jane Doe Jr.', whose + # credential now starts a run that takes the clause in and reports + # what it absorbs, the rest on heads whose marker is an ordinary + # word now; the tails loop below adds 6. + assert multi == 4239, ( f"{multi} of these parses carry more than one report, recorded " - f"as 4320 on 2026-09-19. A parse with one report cannot fail " + f"as 4239 on 2026-10-04. A parse with one report cannot fail " f"the check above, so this is what keeps the grid honest: move " f"the number deliberately, and never to 0") - assert both == 1944, ( + # 2,052 since 2026-10-04 (#601/#602), from 1,944: 216 parses left, + # every one on a comma head whose clause in the given part no + # longer takes ('Doe, Jane', 'Doe, Jane van', ...), and 324 + # arrived, on 'Dr.' and 'J.' -- where the marker is an ordinary word + # behind the title, or the clause's own run is read over the + # clause-free view -- and on 'Jane Doe Jr.', whose credential run + # reports what it absorbs. The run's reports moved from assign's + # peel to the take that consumes the run (#601), so the pair is + # the maiden take's member report against the run's. + assert both == 2052, ( f"{both} parse(s) carry TWO suffix-or-name reports, recorded " - f"as 1944 on 2026-09-19. This is the pair the test is named " - f"for -- the maiden emitter against assign's trailing peel -- " - f"and it was 0 for the whole grid until the two-member tails " - f"were added. Never re-record it as 0") + f"as 2052 on 2026-10-04. This is the pair the test is named " + f"for -- the clause's kept member against the trailing run's " + f"reading -- and it was 0 for the whole grid until the " + f"two-member tails were added. Never re-record it as 0") @pytest.mark.parametrize("text", _FORK_CORPUS) diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index f9ba1ea3..6bc24104 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -30,8 +30,6 @@ "Berg, Jan vd" "Berg, abd née Jones" "Berg, abdul V" -"Berg, abdul nee Jones MA" -"Berg, abdul nee Smith V" "Berg, abdul van" "Berg, abdul vd" "Carod i" @@ -46,14 +44,9 @@ "De La Cruz, M.J. PhD" "Del Toro" "Doe nee Smith Jr. Prof., Jane" -"Doe nee Smith Prof. ba" -"Doe nee Smith Prof., Jane" -"Doe nee Smith V, Jane" -"Doe nee Smith ba Prof." -"Doe, Dr. nee Smith MA" +"Doe nee Smith, Jane" "Doe, Jane PhD MEng" -"Doe, Jane nee Smith St." -"Doe, Jane nee Smith V, PhD" +"Doe, Jane nee Smith" "Doe, Jane, and Jr." "Doe, John DO" "Doe, John DO Ed" @@ -84,9 +77,9 @@ "Dr. abdul salam" "Dr. juan garcia" "Dr. med. univ. Margit Popp, MSc" -"Dr. nee Jones Smith Prof." -"Dr. nee V" +"Dr. nee Smith PhD Prof." "Duke of Edinburgh" +"Eric H. Holder Jr. Attorney General" "Esq. Smith" "Freiherr von Berg MA" "Freiherr von Berg, Ed" @@ -137,29 +130,23 @@ "Jane Doe (nee Smith Ma)" "Jane Doe (nee Smith Prof.)" "Jane Doe (nee Smith) MA" -"Jane Doe Jr. Ma" -"Jane Doe Jr. nee Smith Ma" +"Jane Doe Jr. nee Smith" +"Jane Doe PhD nee Smith" "Jane Doe nee King." -"Jane Doe nee King. ba" "Jane Doe nee MA" "Jane Doe nee MA Smith" "Jane Doe nee Prof. Dr." "Jane Doe nee Puig i" "Jane Doe nee Puig i Soler" "Jane Doe nee Smith DO DO" -"Jane Doe nee Smith DO Prof." -"Jane Doe nee Smith King." "Jane Doe nee Smith MA" "Jane Doe nee Smith MA Prof." "Jane Doe nee Smith Ma" +"Jane Doe nee Smith PhD Smith" "Jane Doe nee Smith Prof." -"Jane Doe nee Smith Prof. DO" "Jane Doe nee Smith Prof. MA" "Jane Doe nee Smith V Prof." "Jane Doe nee Smith X.Y.Z." -"Jane Doe nee Smith do MA Prof." -"Jane Doe nee Smith do Prof. MA" -"Jane Doe nee Smith i DO Prof." "Jane Doe, MS LAc" "Jane Smith (Nee)" "Jane Smith (Nee) (Jones)" @@ -171,13 +158,12 @@ "Jane Smith née Jones PhD" "Jane Smith née V" "Jane Smith, née Jones" +"Jane and née Jones" "Jane de la née Jones" -"Jane née Jones J. V" "Jane née Jones Smith" "Jane née Jr y Jones" -"Jane van der Berg nee Smith PhD Prof." +"Jane van der Berg nee Prof. King. MA" "Jane van der Berg nee Smith Prof." -"Jane van der Berg nee Smith Prof. PhD" "Jane van der Berg née" "Jane van der Berg née Jones" "Jane van der Berg née PhD" @@ -189,6 +175,7 @@ "John . Smith" "John Doctor Smith" "John Ma" +"John PhD Smith" "John Prof. MA" "John Quincy Smith i" "John Smith" @@ -204,7 +191,9 @@ "John Smith Msc.Ed." "John Smith PhD" "John Smith PhD Ed Ma" +"John Smith PhD Jones" "John Smith PhD MEng" +"John Smith PhD Prof. Ma" "John Smith Prof." "John Smith Prof. Dr." "John Smith Prof. Jr." @@ -243,7 +232,7 @@ "John Smith, XYZ" "John Smith, vd Ma" "John and Jane Smith" -"John née Jones Smith MA" +"John nee Prof. ba MA" "John née Jones Smith Ma" "John née Jones Smith V" "John van Buren, Ed" @@ -287,19 +276,23 @@ "Lord Chancellor" "MD DDS" "MSc Dr. med. univ." +"Mai Le née Nguyen" "Mari' Aube'" "Maria Kowalska (z domu Nowak)" "Maria Kowalska (z domu)" "Maria Kowalska z domu Nowak" "Maria Luisa y de la Cruz" +"Maria López née García Pérez" "Marquess of Bath" "Mary Beth Smith" "Mary Jane King" +"Mary Jane King Smith" "Mary Jane King." "Mary Smith née Jones Prof." "Mc Donald" "Md Abdul Karim" "Mesnil de" +"Mohamed Ali Abd Allah" "Morse, Det. Insp. Jane" "Mr. Jack Jill" "Mr. Jack and Jill" @@ -363,14 +356,15 @@ "Smith, Esq." "Smith, John" "Smith, John PhD I." +"Smith, John PhD Jones" +"Smith, John PhD de Jr." "Smith, John Prof." "Smith, John V" "Smith, John V." "Smith, John, MD - née Jones Smith" "Smith, John, PhD - and MD" "Smith, John, PhD - i Soler" -"Smith, John, PhD née Puig - i Soler" -"Smith, John, PhD née Puig Mr. - i Soler" +"Smith, John, PhD née Jones" "Smith, John, Puig - i Soler" "Smith, John, and" "Smith, Jr." @@ -486,7 +480,6 @@ "van Gogh" "van der Berg, MA" "van der Berg, PhD" -"van der Berg, abdul née Jones" "wang meng" "Ó. Pérez" "Иван Петрович Абрамович" diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index b8dc85f9..203308b3 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -383,6 +383,27 @@ DO' actually landing here. """ [[change]] +# One name, 'Jane Doe nee MA PhD', where the marker's clause takes +# words v1 read as a middle name and a suffix and keeps the member +# standing straight after the marker. Against 1.4.0 the visible diff +# is #274's consumption -- v1 had no maiden reading at all, so +# `given`/`middle`/`family`/`suffix` move as `maiden` fills -- and the +# member stays the maiden name because rules.md#M2's take "always +# takes the first word after it, the marker having announced a name"; +# only the unambiguous 'PhD' behind it leaves, which is why `suffix` +# narrows rather than empties. +# +# Four more names sat here until #601 (2026-10-04): 'Doe, Dr. nee +# Smith MA', 'Berg, abdul nee Jones MA', 'Berg, Jane van der nee Smith +# DO' and 'Doe, Jane nee Smith Do'. Each puts the marker in the given +# part after a family comma, where it is an ordinary word since #601 +# -- as v1 read it -- so their clause is gone, and with it the three +# `precedes_narrower` exemptions this rule declared over the comma, +# #411 and #379 rules for them. +# +# `_ambiguities` is NOT declared and cannot be at this baseline: the +# v1 surface has no such field, and the report this name carries is +# recorded in the 2.x ledgers where the comparison can see it. issue = "fix(#274/#533) the clause keeps the credential a join or a missing reader would have taken" # Five names where the marker's clause takes words v1 read as a # middle name and a suffix, and KEEPS the credential at the end of @@ -413,59 +434,8 @@ issue = "fix(#274/#533) the clause keeps the credential a join or a missing read # v1 surface has no such field, and the reports these names now # carry are recorded in the 2.x ledgers where the comparison can see # them. -name_regex = "^(?:Berg, Jane van der nee Smith DO|Berg, abdul nee Jones MA|Doe, Dr\\. nee Smith MA|Doe, Jane nee Smith Do|Jane Doe nee MA PhD)$" -fields = ["family", "given", "maiden", "middle", "suffix"] - -[[change.precedes_narrower]] -issue = "fix(comma-precomma-family) pre-comma run reads as family, not given" -why = """ -LIVE, and new with the #533 review's corpus rows. Four of the five -names here have a family comma, so the pre-comma rule's {given, -family} is a strict subset of these fields and nothing but file order -keeps them here. - -The discriminator is WHAT MOVED the given name. The precomma rule -describes 2.x splitting a pre-comma run that 1.4 read as one first -name; on these four the pre-comma run is a single word 1.4 and the -tree both read as the family, and what empties `given` is the MARKER -leaving with the word after it -- 'Doe, Dr. nee Smith MA' reads given -'nee' at 1.4.0 because v1 had no marker reading at all. Classified by -the precomma rule instead, each would be attributed to a comma -behavior none of them exercises. -""" - -[[change.precedes_narrower]] -issue = "fix(#411) the bound-given reserve stops counting words the maiden name takes" -why = """ -LIVE, and new with the #533 review's corpus rows. 'Berg, abdul nee -Jones MA' is reached by both and #411's fields are a strict subset. - -The discriminator is that #411's rule is about the RESERVE and this -name's diff is about the CREDENTIAL: the reserve half is real and is -why `given` reads 'abdul' rather than 'abdul nee', but the `suffix` -1.4.0 read is emptied by the clause KEEPING the 'MA', which #411 has -nothing to say about and whose field it does not declare. The review -round is what settled that reading -- an earlier round released the -word into P5's join and read given 'abdul MA' -- so a rule naming -#411 alone would attribute this change's decline to the reserve. -""" - -[[change.precedes_narrower]] -issue = "fix(#379) a tussenvoegsel after a family comma attaches to the family" -why = """ -LIVE, and new with the #533 review's corpus rows. 'Berg, Jane van der -nee Smith DO' and 'Doe, Jane nee Smith Do' each carry a maiden marker -AND a tussenvoegsel-shaped word after a family comma, and the -tussenvoegsel rule's fields are a strict subset of these. - -The discriminator is the one fix(#274)'s own block above this file -gives for the same pair: what 1.4.0 read as middle text plus a -trailing post-nominal, the tree reads as a maiden name, and the -marker leaving with its words is what every field of these diffs is -about. On 'Berg, Jane van der nee Smith DO' the tussenvoegsel is not -even the moved word -- 'van der' attaches to the family either way, -and the 'DO' the clause KEEPS is what empties `suffix`. -""" +name_regex = "^Jane Doe nee MA PhD$" +fields = ["family", "maiden", "middle", "suffix"] [[change]] issue = "fix(#335/#533) a marker-led bracketed clause is the maiden name whatever its last word is" @@ -988,42 +958,6 @@ No _CROSS_RULE_WINNERS row for this name; the reason rests on the two rules' own prose, which is unambiguous here -- one names two behaviours meeting, the other one behaviour.""" -[[change]] -issue = "fix(#411/S2) a declining bound-given join leaves the suffix reading after a family comma" -# 'Berg, abd nee Jones': `abd` is the one shipped word in BOTH the -# bound-given and the suffix vocabulary. With #411's join declining -- -# it will not absorb a marker -- what is left in the given slot after a -# family comma is S2's suffix reading, which rules.md#P5 already states -# wins there. So the name has no given name at all, exactly as -# 'Berg, abd' alone has always parsed. -# -# Four fields move, which is why neither fix(#411) nor fix(#274) -# claims it: the extra `suffix` breaks their subset test. That is the -# gate behaving correctly -- the name arrived UNEXPLAINED and is read -# once here. -# -# `abd` and not an alternation: no other shipped bound-given word is -# also suffix vocabulary, so no other word reaches this reading. -name_regex = "(?i),\\s*abd\\s+n[eé]e\\b" -fields = ["given", "middle", "suffix", "maiden"] - -[[change.precedes_narrower]] -issue = "fix(#400) abd joins the word after it as one given name" -why = """ -this rule describes the join DECLINING, which is the negation of what -fix(#400) describes, and then describes what a declining join leaves -in the given slot after a family comma: S2's suffix reading, with -'abd' ending up in `suffix` and the name having no given name at all. -Two rules cannot both be right about the same word, and the one -asserting that it joins is not the one to explain a name where it did -not. - -LATENT: {given, middle, suffix, maiden} is two roles wider than -fix(#400)'s {given, middle}, so the subset test excludes it at any -position -- this rule's comment already reads the same arithmetic as -why neither fix(#411) nor fix(#274) claims the name. No -_CROSS_RULE_WINNERS row; the prose carries it.""" - [[change]] issue = "fix(#411) the bound-given reserve stops counting words the maiden name takes" # 'van der Berg, abdul nee Jones': P5 reserves a name word so the join @@ -1058,7 +992,7 @@ issue = "fix(#411) the bound-given reserve stops counting words the maiden name # an author to expect. 'abdel', 'abdal' and the Arabic spelling do # behave like 'abdul'; only this name is in the corpora. name_regex = "(?i),\\s*abdul\\s+n[eé]e\\b" -fields = ["given", "middle", "maiden"] +fields = ["given", "middle"] [[change]] issue = "fix(#400) abd joins the word after it as one given name" @@ -4325,7 +4259,7 @@ issue = "fix(#274/#424) accepted: a maiden clause keeps a trailing credential v1 # for no word of the clause". #544 decided that boundary and moves no role # on the name, so the 1.4.0 diff is this rule's, the Title-case member's # writing declining it as it does for 'Jane Doe nee Smith Ma'. -name_regex = "^(?:Doe, Jane nee Smith MA do|Doe, Jane nee Smith Ma|Doe, Jane nee Smith do|Jane Doe Jr\\. nee Smith Ma|Jane Doe nee MA|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma)$" +name_regex = "^(?:Jane Doe nee MA|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma)$" fields = ["family", "maiden", "middle", "suffix"] [[change]] @@ -4349,7 +4283,7 @@ issue = "fix(#274/#436/#437) a clause the unambiguous credential ended, and the # tree before #544 read the Title-case MEng as a name word, and #544's # company clause (rules.md#S2) put the run back, so against this # baseline #544 moves no role on either name. -name_regex = "^(?:Doe, Jane nee Smith PhD MEng|Jane Doe nee Smith PhD MA|Jane Doe nee Smith PhD MEng)$" +name_regex = "^(?:Jane Doe nee Smith PhD MA|Jane Doe nee Smith PhD MEng)$" fields = ["family", "maiden", "middle", "suffix"] [[change]] @@ -4376,30 +4310,9 @@ issue = "fix(#533) the clause gives the credential up, and the run it joins rend # 'Jane Doe nee Smith MA JD' read differently from each other -- the # capitals are the evidence the lean reads -- so an (?i) anchor here # would stand ready to explain either with the other's reasoning. -name_regex = "^(?:Doe, J\\. nee MA ba|Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD)$" +name_regex = "^(?:Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD)$" fields = ["family", "maiden", "middle", "suffix"] -[[change]] -issue = "fix(#445/#533) the clause gives the credential up, and the one name word it leaves is the family" -# 'John née Jones Smith MA', whose diff has three causes from 1.4.0 -# and one rule has to explain the whole of it. v1 read given 'John', -# middle 'née Jones', last 'Smith', suffix 'MA'; the tree reads -# family 'John', suffix 'MA', maiden 'Jones Smith'. The marker takes -# the words after it (#274), rules.md#M4 makes the one name word the -# clause leaves the family (#445), and #533 is what takes the MA back -# out of the maiden name -- at 2f57ff21 this read maiden 'Jones Smith -# MA' with no suffix at all, so `suffix` was in the diff then and is -# not now. -# -# Its own rule because 'fix(#424/#445) accepted: the maiden walk -# keeps a bare acronym, and the lone name word is the family' above -# says the OPPOSITE of what this name now does, and until this commit -# its (?i) anchor claimed both spellings. That rule is narrowed to -# the spellings whose reading it describes; this one carries the -# spelling that gives the member up. -name_regex = "^John née Jones Smith MA$" -fields = ["family", "given", "maiden", "middle"] - [[change]] issue = "fix(#533) accepted: an unlisted dotted acronym ending a maiden clause leaves v1's family name" # 'Jane Doe nee Smith X.Y.Z.': v1 read middle 'Doe nee Smith', last @@ -4417,25 +4330,6 @@ issue = "fix(#533) accepted: an unlisted dotted acronym ending a maiden clause l name_regex = "^Jane Doe nee Smith X\\.Y\\.Z\\.$" fields = ["family", "maiden", "middle", "suffix"] -[[change]] -issue = "fix(#274) a trailing title after the marker with nothing else in the name" -# 'Dr. nee Jones Smith Prof.' moves the same fields at this baseline -# whether or not #535 exists: v1 has no maiden reading at all, so -# 'nee' stays given text and 'Jones Smith Prof.' stays split across -# middle/family -- the whole diff is #274's (maiden markers -# consumed), and #535 moves NOTHING here (verified: the parent tree -# d9d80492, before #535, already reads title 'Dr.', maiden 'Jones -# Smith Prof.', identically to HEAD). Not folded into the '#274 -# maiden markers consumed' rule above because that rule's `fields` -# lack `given`, which only this name's shape moves. Only a TITLE -# precedes the marker here, so no name word is left standing ahead of -# it, and it is this -# absence, not the marker's own reading, that this rule's SLOT is -# about. Literal-anchored for the reason fix(#533)'s rules give: that -# SLOT is not a vocabulary a regex could describe more broadly. -name_regex = "^Dr\\. nee Jones Smith Prof\\.$" -fields = ["family", "given", "maiden", "middle"] - [[change]] issue = "fix(#274/#296/#535) a trailing title after a dropped postnominal Dr., with no maiden reading either" # 'Jane Doe nee Prof. Dr.' moves at this baseline for two reasons v1 @@ -4464,8 +4358,11 @@ issue = "fix(#274/#535) a trailing title ends the maiden clause" # {maiden, middle, suffix} instead -- `family` there already matches, # the comma having fixed it. #535 additionally moves `title` (and # `suffix` for the credential-bearing ones among the -# seven). rules.md#M2: "Where a trailing rule reads the words, a -# trailing title ends it too" -- the walk reads the end of the name +# seven). rules.md#M2: "It takes them up to the trailing run +# of post-nominals and titles that the end of the name reads as if +# the clause were not written" (#601; until #601 this cited the +# sentence it replaced, "a trailing title ends it too") -- the end of +# the name is read # through the trailing title chain -- rules.md#H5: "successive # single words that wear the abbreviation shape and are title # vocabulary chain into the title from the end" -- so the title @@ -4475,137 +4372,9 @@ issue = "fix(#274/#535) a trailing title ends the maiden clause" # fix(#533)'s rules give: the subject is a SLOT, and a regex would # claim the names whose view check keeps the title as readily as the # ones that release it. -name_regex = "^(?:Doe, Jane nee Smith MA Prof\\.|Jane Doe \\x28nee Smith Prof\\.\\x29|Jane Doe nee Smith King\\.|Jane Doe nee Smith MA Prof\\.|Jane Doe nee Smith Ma Prof\\.|Jane Doe nee Smith Prof\\.|Jane Doe nee Smith Prof\\. MA|Jane Doe nee Smith V Prof\\.|Mary Smith née Jones Prof\\.)$" +name_regex = "^(?:Jane Doe \\x28nee Smith Prof\\.\\x29|Jane Doe nee Smith King\\.|Jane Doe nee Smith MA Prof\\.|Jane Doe nee Smith Ma Prof\\.|Jane Doe nee Smith Prof\\.|Jane Doe nee Smith Prof\\. MA|Jane Doe nee Smith V Prof\\.|Mary Smith née Jones Prof\\.)$" fields = ["middle", "family", "title", "suffix", "maiden"] -[[change]] -issue = "fix(#274/#342/#445/#535) a title behind a credential the clause keeps still leaves it" -# rules.md#M2's Accepted transparency boundary, the half that gives -# the title up: the clause keeps 'ba' (one name word left, none to -# spare) and the title behind it still ends the clause, so title -# 'Prof.', family 'Doe', maiden 'Smith ba'. Four causes together, -# checked against the parent tree d9d80492 (family 'Doe', maiden -# 'Smith ba Prof.', no report): before #535, v1 has no maiden reading (#274), #342 marked 'ba' an ambiguous acronym, so the clause no longer stops at it as a plain suffix word, and #445 makes the lone name word beside a maiden clause the family; -# #535 then reads the title off the end of the clause. -# Literal-anchored for the SLOT reason fix(#533)'s rules give. -name_regex = "^Doe nee Smith ba Prof\\.$" -fields = ["title", "given", "middle", "family", "maiden"] - -[[change]] -issue = "fix(#274/#342/#445/#533) a title in front of a credential the clause keeps stays in it" -# rules.md#M2's Accepted transparency boundary, the half that keeps -# the title: a title in FRONT of the kept 'ba' cannot leave without -# it, so the clause keeps 'Smith Prof. ba' and reports the kept -# credential. The parent tree d9d80492 reads it exactly as HEAD -# does, so #535 moves nothing here: v1 has no maiden reading (#274), #342 marked 'ba' an ambiguous acronym, so the clause no longer stops at it as a plain suffix word, and #445 makes the lone name word beside a maiden clause the family, and #533 reports -# the credential the clause keeps. -# Literal-anchored for the SLOT reason fix(#533)'s rules give. -name_regex = "^Doe nee Smith Prof\\. ba$" -fields = ["given", "middle", "family", "suffix", "maiden"] - -[[change]] -issue = "fix(#274/#535) a particle in front of a trailing title stays in the clause" -# rules.md#M2's Accepted pair on a released particle: the span 'DO -# Prof.' holds a particle with a title behind it, and P2's chain would -# run on over the title, so the particle's release is withdrawn and -# the clause keeps 'Smith DO' while the title stop still gives 'Prof.' -# up. The parent tree d9d80492 reads maiden 'Smith DO Prof.', so the -# move is #535's together with #274 (v1 had no maiden reading at -# all). Literal-anchored for the SLOT reason fix(#533)'s rules give. -name_regex = "^Jane Doe nee Smith DO Prof\\.$" -fields = ["title", "middle", "family", "maiden"] - -[[change]] -issue = "fix(#274/#535) a particle behind a trailing title leaves the clause with it" -# The other half of that pair: written behind the title the particle -# is released with nothing behind it, so the title and 'DO' both leave -# the clause. v1 had no maiden reading (#274); #535 reads the title -# off the clause, as for 'Jane Doe nee Smith Prof. MA', which this -# ledger also labels #274/#535. Literal-anchored for the SLOT reason -# fix(#533)'s rules give. -name_regex = "^Jane Doe nee Smith Prof\\. DO$" -fields = ["title", "middle", "family", "maiden"] - -[[change]] -issue = "fix(#274/#316) a trailing title behind a maiden clause's first suffix word" -# rules.md#M2's Accepted pair on a title something ahead would take: -# the first-suffix-word stop ends the clause at 'PhD', so 'Prof.' -# stands behind the clause and reads as a trailing title. The parent -# tree d9d80492 and 2.3.0 read it as HEAD does, so #535 moves nothing; -# v1 had no maiden reading (#274) and read no trailing title (#316). -# Literal-anchored: one corpus name, the doc's own example. -name_regex = "^Jane van der Berg nee Smith PhD Prof\\.$" -fields = ["title", "middle", "family", "suffix", "maiden"] - -[[change]] -issue = "fix(#274) a numeral straight after the marker stays maiden where nothing reads it as a suffix" -# rules.md#M2's boundary: a numeral straight after the marker declines -# the clause only where the name left standing reads it as a suffix, -# and 'Dr.' alone reads nothing, so the clause keeps 'V'. Every 2.x -# baseline reads it so; v1 had no maiden reading (#274) and read given -# 'nee', family 'V'. Literal-anchored: one corpus name, the doc's own -# example. -name_regex = "^Dr\\. nee V$" -fields = ["given", "family", "maiden"] - -[[change]] -issue = "fix(#274/#535) a particle inside the clause keeps the credential in front of a trailing title" -# rules.md#M2's Accepted pair on a particle INSIDE the clause: the -# particle's chain would take the credential behind it, so the clause -# keeps 'Smith do MA' and only the title behind the credential leaves. -# The parent tree d9d80492 reads maiden 'Smith do MA Prof.' with no -# report, and v1 had no maiden reading at all (#274), so the move is -# #535's together with #274. Literal-anchored: one corpus name, the -# doc's own example. -name_regex = "^Jane Doe nee Smith do MA Prof\\.$" -fields = ["title", "middle", "family", "maiden"] - -[[change]] -issue = "fix(#274/#535) a particle inside the clause lets a credential behind the title go" -# The other half of that pair: with the title in front of 'MA', the -# credential is released and the title leaves with it, so maiden -# 'Smith do', suffix 'MA', title 'Prof.'. v1 had no maiden reading -# (#274); #535 reads the title off the clause, as for 'Jane Doe nee -# Smith Prof. MA', which this ledger also labels #274/#535. -# Literal-anchored: one corpus name, the doc's own example. -name_regex = "^Jane Doe nee Smith do Prof\\. MA$" -fields = ["title", "middle", "family", "maiden"] - -[[change]] -issue = "fix(#274/#397/#535) a link the clause stops at keeps a run that would not read off" -# rules.md#M2: a link that ends the clause gives up the words behind -# it only where the name left standing reads them as post-nominals or -# titles. With the title chained the link exception refuses 'i', but -# 'Jane Doe i DO Prof.' does not read 'i DO' off, so the clause keeps -# the link and the DO (reporting the kept credential) and only 'Prof.' -# leaves. Checked against the parent tree d9d80492, which read maiden -# 'Smith i DO Prof.': v1 had no maiden reading (#274) and no Catalan -# link, 'i' being a plain suffix word, and #397 joins the link inside -# the clause; #535 then reads the title off. Literal-anchored: one -# corpus name, the doc's own example. -name_regex = "^Jane Doe nee Smith i DO Prof\\.$" -fields = ["title", "middle", "family", "maiden"] - -[[change]] -issue = "fix(#411/#535) the numeral stop asks the join question" -# 'Berg, abdul nee Smith V' reads IDENTICALLY at this baseline and at -# 2.0.0 -- given 'abdul nee', middle 'Smith', family 'Berg', suffix -# 'V' -- so #274 (v1 has no maiden reading) moves nothing here; this -# name simply never exercises a maiden reading before #411 lands. -# Two reasons together move it from this baseline to HEAD: the -# bound-given reserve (#411) decides how much of the tail the given -# join takes before #535 ever asks anything -- rules.md#P5: "The -# marker and the words it will take are not among the words to -# spare: they leave the name, so counting them asks the question -# about a name that will not exist" -- and the parent tree d9d80492, -# before #535, already reads given 'abdul V', maiden 'Smith', -# differing from this baseline; #535 is what moves the released V -# from given back into maiden (`given` 'abdul V' -> 'abdul', `maiden` -# 'Smith' -> 'Smith V'). Literal-anchored for the same SLOT reason -# fix(#533)'s rules give. -name_regex = "^(?:Berg, abdul nee Smith V)$" -fields = ["given", "middle", "suffix", "maiden"] - [[change]] issue = "fix(#434/#533) a marker PHRASE takes the maiden name, and its clause ends at the credential" # 'Maria Kowalska z domu Nowak MA'. fix(#274)'s regex is the @@ -4646,20 +4415,6 @@ fields = ["family", "maiden", "middle"] # bullet carries the measurement. # --------------------------------------------------------------- -[[change]] -issue = "fix(#274/#397) a link inside a maiden clause stays in the birth name, after a family comma" -# 'Doe, Jane nee Puig i Soler'. v1 read middle 'nee Puig Soler' with -# suffix 'i' -- the marker an ordinary word and the link a generation. -# The tree reads given 'Jane', family 'Doe', maiden 'Puig i Soler': -# the marker consumed (#274) and the link kept inside the clause it -# stands in (#397). `suffix` is in the list because v1 had one and the -# tree does not, which is the link's half. -# -# Literal, one name. The shape would be "a maiden marker and a bare -# letter", which reaches every clause name in the corpora. -name_regex = "^Doe, Jane nee Puig i Soler$" -fields = ["maiden", "middle", "suffix"] - [[change]] issue = "fix(#274/#436/#437) a clause the generational link ended, and the run it left renders with spaces" # 'Jane Doe nee Puig i III'. Two changes meet in the v1 diff: the @@ -4678,47 +4433,6 @@ issue = "fix(#274/#436/#437) a clause the generational link ended, and the run i name_regex = "^Jane Doe nee Puig i III$" fields = ["family", "maiden", "middle", "suffix"] -[[change]] -issue = "fix(#274/#397) a maiden clause inside a suffix-comma tail leaves the suffix field, link and all" -# 'Smith, John, PhD née Puig Mr. - i Soler', which arrives from -# rules.md#M2 rather than from a report: it is one of rules.md#M2's -# delimiter examples, added with #538 (under a configured ' - ' it -# reads maiden 'Puig Mr.', by the comma reading since #549). This corpus does not configure that delimiter, -# so what the gate sees is the DEFAULT facade reading, where the dash -# is an ordinary name word like any other. -# -# v1 had no maiden markers at all, so everything behind the suffix -# comma stayed in the suffix: suffix 'PhD née Puig Mr. - i Soler' -# against the tree's 'PhD', maiden '' against 'Puig Mr. - i Soler' -# (measured 2026-09-22). One rule for the whole of it -- the marker -# leaving the name (#274) and, inside what it takes, the link kept in -# the clause it stands in (#397), which is what puts 'i Soler' on the -# maiden side of the same two fields. -# -# Not a widening of the fix(#274/#397) comma rule above: that one is -# about a FAMILY comma, where v1 left a middle name behind and the -# field list says so. This tail leaves v1 a suffix and nothing else. -# -# Literal names, for the reason those rules give: the shape would -# be "a maiden marker and a bare letter", which reaches every clause -# name in the corpora. 'Smith, John, PhD née Puig - i Soler' joined -# this alternation with #538 (2026-09-26): it is rules.md#M2's second -# live example, and it reaches the DEFAULT facade the same way its -# neighbour does -- no delimiter configured here either, so the dash -# is an ordinary name word. Measured at this baseline: suffix -# 'PhD née Puig - i Soler' -> 'PhD', maiden '' -> 'Puig - i Soler' -- -# the same shape as the neighbour's own reading with 'Mr.' dropped -# from it; the tree's reading (suffix 'PhD', maiden 'Puig - i Soler') -# is unchanged by the fix, and only the corpus name is new. -# 'Smith, John, MD - née Jones Smith' joined with #549 (2026-10-03): -# rules.md#C1's Accepted example of a marker beside a configured -# delimiter, reaching the DEFAULT facade with the dash an ordinary -# word again. Measured at this baseline: suffix 'MD - née Jones Smith' -# -> 'MD -', maiden '' -> 'Jones Smith' -- the marker half of this -# rule alone, there being no link in it. -name_regex = "^(?:Smith, John, PhD née Puig Mr\\. - i Soler|Smith, John, PhD née Puig - i Soler|Smith, John, MD - née Jones Smith)$" -fields = ["maiden", "suffix"] - [[change]] issue = "change(suffix-acronym-collisions) ph leaves the acronym set" # 'John Smith Ph.': 'ph' left SUFFIX_ACRONYMS with #459's all-caps @@ -4995,6 +4709,91 @@ issue = "fix(#597) halfwidth corner brackets enclose a nickname" name_regex = "^(?:山田 「タロー」 タロウ|山田「タロー」太郎|John 「Jack」 Smith)$" fields = ["given", "middle", "family", "nickname"] +[[change]] +# rules.md#M2: "In the given part after a family comma, in a part after a +# suffix comma, and behind a title, a suffix word or a connective, the +# marker is an ordinary word" (#601, 2026-10-04) -- which is how 1.4.0, +# with no maiden reading at all, read it too, so `maiden` is not in +# this diff. What differs is the given part's own reading of the +# clause's words since 1.4.0 -- the link join (#397: 'Puig i Soler'), +# P6's attachment of a trailing particle to the family (#405: 'Do +# Doe'), the capitals lean (#289: 'Ma' a name), and the trailing title +# chain (H5: 'Prof.') -- each as the same text with the marker replaced +# by an ordinary word reads it. These six sat in the #274/#533, +# #274/#424, #274/#535 and #274/#397 maiden rules until #601 left them +# no clause. +issue = "fix(#601) a marker in the given part after a family comma is an ordinary word" +name_regex = "^(?:Doe, Jane nee Puig i Soler|Doe, Jane nee Smith Do|Doe, Jane nee Smith MA Prof\\.|Doe, Jane nee Smith MA do|Doe, Jane nee Smith Ma|Doe, Jane nee Smith do)$" +fields = ["family", "middle", "suffix", "title"] + +[[change]] +# rules.md#M2: "In the given part after a family comma, in a part after a +# suffix comma, and behind a title, a suffix word or a connective, the +# marker is an ordinary word" (#601, 2026-10-04). Where a credential +# stands before the marker it also starts #602's run -- rules.md#S2: "A +# credential after the name core starts a run to the end of its part" -- +# which takes the marker and the words after it in. +issue = "fix(#601/#602) a marker behind a title, a suffix word or a connective is an ordinary word" +name_regex = "^(?:Dr\\. nee Smith PhD Prof\\.|Jane Doe Jr\\. nee Smith|Jane Doe Jr\\. nee Smith Ma|Jane Doe PhD nee Smith)$" +fields = ["family", "middle", "suffix", "title"] + +[[change]] +# rules.md#M2: "It takes them up to the trailing run of post-nominals and +# titles that the end of the name reads as if the clause were not +# written" (#601, 2026-10-04), and "The run so found ends the name: its +# words are the suffixes and titles that reading makes them, and no join +# reaches into it". The old walk's release check kept these words in the +# clause where a join it modelled might take them; the run is bound now. +# Against 1.4.0, which had no maiden reading, `maiden` fills +# as well (#274). +issue = "fix(#274/#601) the clause ends at the clause-free name's trailing run, which the take consumes" +name_regex = "^(?:Jane Doe nee Smith DO DO|Jane van der Berg nee Prof\\. King\\. MA|Jane van der Berg nee Smith Prof\\.|John nee Prof\\. ba MA)$" +fields = ["family", "given", "maiden", "middle", "suffix", "title"] + +[[change]] +# rules.md#M2's take "always takes the first word after it, the marker +# having announced a name" (#601, 2026-10-04): a lone roman numeral is no +# suffix word, so it is the maiden name, where the old walk read the +# numeral fork from the marker and declined. Decided with #601 (decisions.md#M2). +# Against 1.4.0, which had no maiden reading, `maiden` fills +# as well (#274). +issue = "fix(#274/#601) the first word after the marker is the maiden name" +name_regex = "^Jane Smith née V$" +fields = ["family", "maiden", "middle", "suffix"] + +[[change]] +# rules.md#S2: "A credential after the name core starts a run to the end +# of its part" (#602, 2026-10-04): every later word of the part is a +# suffix except a title word, which reads as a title, and an absorbed +# name word is reported. +# 'Jane van der Berg née Jr Jones' (radar): the marker declines on +# the suffix word after it (rules.md#M2) and stays a word, so 'Jr' +# stands behind a name core and its run takes 'Jones' (2026-10-04). +issue = "fix(#602) a credential after the name core starts a run to the end of its part" +name_regex = "^(?:Eric H\\. Holder Jr\\. Attorney General|Jane van der Berg née Jr Jones|John Smith PhD Jones|John Smith PhD Prof\\. Ma|Smith, John PhD Jones)$" +fields = ["family", "middle", "suffix", "title"] + +[[change]] +# rules.md#M2 reads the clause-free name, whose credential starts #602's +# run -- rules.md#S2: "A credential after the name core starts a run to +# the end of its part" -- and the take consumes that run: maiden 'Smith', +# suffix 'PhD Smith' (2026-10-04). +# Against 1.4.0, which had no maiden reading, `maiden` fills +# as well (#274). +issue = "fix(#274/#601/#602) a credential in the clause starts a run the take consumes" +name_regex = "^Jane Doe nee Smith PhD Smith$" +fields = ["family", "maiden", "middle", "suffix"] + +[[change]] +# rules.md#P6: "a particle ending the name attaches to that family name", +# and a post-nominal written behind the particle does not end the name +# for this purpose. rules.md#S2's example for the given-part run leaving +# the particle to P6 (#601/#602 review, 2026-10-04); the attachment +# itself landed in 2.2 (#379). +issue = "fix(#379) a tussenvoegsel behind a post-nominal after a family comma attaches" +name_regex = "^Smith, John PhD de Jr\\.$" +fields = ["middle", "family"] + # --------------------------------------------------------------- # #604: THE IRISH AND MALAY PATRONYMIC PARTICLES. # 'ó' and 'ní' join the never-given particles, 'ua', 'binti' and diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 9133273f..2ccc6d6b 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -671,6 +671,7 @@ fields = ["given", "family", "_ambiguities"] [[change]] issue = "fix(#411) the bound-given reserve stops counting words the maiden name takes" +dormant = "no corpus name since #601 (2026-10-04): the rules.md#M2 examples that carried this shape left the doc with the walk they illustrated. The reserve still counts only the words the clause leaves (the take runs before every join), so the rule is kept for the next name that shows it" # 'van der Berg, abdul nee Jones': P5 reserves a name word so the join # always leaves a family name behind, but the reserve was counted while # the marker and the maiden name were still pieces -- group's marker @@ -1178,17 +1179,6 @@ issue = "fix(#424/#445) the maiden walk stops before the trailing numeral, and t name_regex = "(?i)^john\\s+n[ée]e\\s+jones\\s+smith\\s+v$" fields = ["given", "family", "suffix", "maiden", "_ambiguities"] -[[change]] -issue = "fix(#424) a marker followed only by the numeral is just a word" -# 'Jane Smith née V': the peel is read from the marker, so the V has -# the piece before it the fork wants and reads as the suffix; nothing -# follows the marker but a suffix, the pass declines, and the marker -# stays a word -- as for 'Jane Smith née PhD'. maiden 'V' -> middle -# 'Smith', family 'née', suffix 'V', which is 1.4.0's reading, so -# this rule has no 1.4.0 twin. A rules.md example. -name_regex = "(?i)^jane\\s+smith\\s+n[ée]e\\s+v$" -fields = ["middle", "family", "suffix", "maiden", "_ambiguities"] - [[change]] issue = "fix(#369) the bound given-name join takes a particle-and-bound word, so no fork is reported" # 'Sheik Abu Bakar': the fields are byte-identical to this baseline @@ -2946,7 +2936,7 @@ issue = "fix(#533) a credential ending a maiden clause reads as a credential" # Literal-anchored for the reason fix(#531)'s rules give: the subject # is a SLOT, so a regex for it would claim the names whose WRITING # declines the word as readily as the ones it takes. -name_regex = "^(?:Doe, Jane Q\\. nee Smith MA|Doe, Jane nee Smith DO|Doe, Jane nee Smith MA|Doe, Jane nee Smith ma|JANE DOE NEE SMITH MA|JANE DOE NEE YO-YO MA|Jane Doe geb\\. Smith MA|Jane Doe nee Smith MA|Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD|Jane Doe nee Smith V MA|Jane Doe nee Smith do|jane doe nee smith ma)$" +name_regex = "^(?:JANE DOE NEE SMITH MA|JANE DOE NEE YO-YO MA|Jane Doe geb\\. Smith MA|Jane Doe nee Smith MA|Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD|Jane Doe nee Smith V MA|Jane Doe nee Smith do|jane doe nee smith ma)$" fields = ["_ambiguities", "maiden", "suffix"] [[change]] @@ -2992,45 +2982,9 @@ issue = "fix(#533) the maiden clause reports the credential it keeps" # half at rules.md#M2's new Accepted boundary: the credential in front # of the marker speaks for no member of the clause, the Title-case 'Ma' # is kept, and the clause reports it. -name_regex = "^(?:Doe, Dr\\. nee Smith MA|Doe, J\\. nee MA ba|Doe, Jane nee Smith Ma|Jane Doe Jr\\. nee Smith Ma|Jane Doe nee King\\. ba|Jane Doe nee MA|Jane Doe nee MA PhD|Jane Doe nee Smith DO DO|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma)$" +name_regex = "^(?:Jane Doe nee King\\. ba|Jane Doe nee MA|Jane Doe nee MA PhD|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma)$" fields = ["_ambiguities"] -[[change]] -issue = "fix(#411/#533) the bound-given pair after a comma, and the clause it keeps its credential in" -# 'Berg, abdul nee Jones MA'. Two changes in one name and neither is -# enough on its own, which is why it is a rule rather than an -# alternative in either. The clause consumption (#274) and the -# reserve that stops counting the words it takes (#411) are what move -# `given`, `middle` and `maiden` against this baseline; the report is -# #533's. -# -# What #533 does here is NOT a role move, and the comment is the only -# place that can say so: an earlier round of this branch released the -# 'MA' from the clause and let P5's lenient post-comma join swallow -# it into the pair, reading given 'abdul MA' in silence. The review -# measured that as M2's invariant broken -- a word the clause gave up -# landing in a NAME part -- so the stop is declined, the clause keeps -# 'Jones MA', and the fork is reported instead. -name_regex = "^Berg, abdul nee Jones MA$" -fields = ["_ambiguities", "given", "maiden", "middle"] - -[[change]] -issue = "fix(#399/#533) a maiden marker bounds the particle chain, and the clause keeps the credential the chain would have taken" -# 'Berg, Jane van der nee Smith DO'. The marker bounds the particle -# join arriving from its left (#399, rules.md#M2), which is what -# moves `middle`, `family` and `maiden` against this baseline, and -# the chain still reports its own PARTICLE_OR_GIVEN. -# -# The SUFFIX_OR_NAME beside it is #533's, and it is a report without -# a role move on purpose. 'DO' is particle vocabulary, so releasing -# it from the clause would put it back in reach of the very chain the -# marker just bounded -- family 'van der DO Berg', a word crossing -# from the birth name into the current one, which an earlier round of -# this branch did in silence. rules.md#M2's invariant declines the -# stop, the clause keeps 'Smith DO', and the fork is reported. -name_regex = "^Berg, Jane van der nee Smith DO$" -fields = ["_ambiguities", "family", "maiden", "middle"] - [[change]] issue = "fix(#335/#533) a marker-led bracketed clause is the maiden name whatever its last word is" # 'Jane Doe (nee Smith MA)', 'Jane Doe (nee Smith Ma)', 'Jane Doe @@ -3088,23 +3042,6 @@ issue = "fix(#533) restores the suffix reading a 2.4 retag had moved into the ma name_regex = "^John Smith nee Jones R\\.A\\.I\\.$" fields = ["_ambiguities"] -[[change]] -issue = "fix(#399) a maiden marker bounds the particle chain that swallowed it, two trailing words" -# 'Jane van der Berg nee Smith Prof.' moves {family, maiden} at this -# baseline whether or not #535 exists -- verified against the parent -# tree d9d80492, which already reads family 'van der Berg', maiden -# 'Smith Prof.' identically to HEAD, before #535 touched anything. -# It is the #399 rule's own shape: rules.md#P2: "a maiden marker -# takes the words after it (M2), or the name ends" -- but the chain -# here swallows the marker and BOTH trailing words before the maiden -# consumer could see any of them, and the general rule's anchor takes -# exactly one word after the marker (`\\sn[ée]e\\s+\\S+$`), so this -# name's two-word clause needs its own literal rather than widening -# that regex's reach past what its own review measured. #535 moves -# nothing here. -name_regex = "^Jane van der Berg nee Smith Prof\\.$" -fields = ["family", "maiden"] - [[change]] issue = "fix(#296/#535) a trailing title after a dropped postnominal Dr." # 'Jane Doe nee Prof. Dr.' moves {suffix, title} at this baseline for @@ -3138,8 +3075,11 @@ fields = ["_ambiguities", "maiden", "suffix", "title"] [[change]] issue = "fix(#535) a trailing title ends the maiden clause" -# rules.md#M2: "Where a trailing rule reads the words, a trailing -# title ends it too" -- the walk reads the end of the name through +# rules.md#M2: "It takes them up to the trailing run +# of post-nominals and titles that the end of the name reads as if +# the clause were not written" (#601; until #601 this cited the +# sentence it replaced, "a trailing title ends it too") -- the end of +# the name is read through # the trailing title chain -- rules.md#H5: "successive single words # that wear the abbreviation shape and are title vocabulary chain # into the title from the end" -- so the title leaves the clause and @@ -3153,170 +3093,9 @@ issue = "fix(#535) a trailing title ends the maiden clause" # each because the parent tree d9d80492 already differed from this # baseline before #535 touched it; the seven names here are ones # where #535 is the sole cause, verified the same way. -name_regex = "^(?:Doe, Jane nee Smith MA Prof\\.|Jane Doe \\x28nee Smith Prof\\.\\x29|Jane Doe nee Smith King\\.|Jane Doe nee Smith MA Prof\\.|Jane Doe nee Smith Ma Prof\\.|Jane Doe nee Smith Prof\\.|Jane Doe nee Smith V Prof\\.|Mary Smith née Jones Prof\\.)$" +name_regex = "^(?:Jane Doe \\x28nee Smith Prof\\.\\x29|Jane Doe nee Smith King\\.|Jane Doe nee Smith MA Prof\\.|Jane Doe nee Smith Ma Prof\\.|Jane Doe nee Smith Prof\\.|Jane Doe nee Smith V Prof\\.|Mary Smith née Jones Prof\\.)$" fields = ["_ambiguities", "maiden", "suffix", "title"] -[[change]] -issue = "fix(#342/#445/#535) a title behind a credential the clause keeps still leaves it" -# rules.md#M2's Accepted transparency boundary, the half that gives -# the title up: the clause keeps 'ba' (one name word left, none to -# spare) and the title behind it still ends the clause, so title -# 'Prof.', family 'Doe', maiden 'Smith ba'. Three causes together, -# checked against the parent tree d9d80492 (family 'Doe', maiden -# 'Smith ba Prof.', no report): before #535, #342 marked 'ba' an ambiguous acronym, so the clause no longer stops at it as a plain suffix word, and #445 makes the lone name word beside a maiden clause the family; -# #535 then reads the title off the end of the clause. -# Literal-anchored for the SLOT reason fix(#533)'s rules give. -name_regex = "^Doe nee Smith ba Prof\\.$" -fields = ["title", "given", "middle", "family", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#342/#445/#533) a title in front of a credential the clause keeps stays in it" -# rules.md#M2's Accepted transparency boundary, the half that keeps -# the title: a title in FRONT of the kept 'ba' cannot leave without -# it, so the clause keeps 'Smith Prof. ba' and reports the kept -# credential. The parent tree d9d80492 reads it exactly as HEAD -# does, so #535 moves nothing here: #342 marked 'ba' an ambiguous acronym, so the clause no longer stops at it as a plain suffix word, and #445 makes the lone name word beside a maiden clause the family, and #533 reports -# the credential the clause keeps. -# Literal-anchored for the SLOT reason fix(#533)'s rules give. -name_regex = "^Doe nee Smith Prof\\. ba$" -fields = ["given", "family", "suffix", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#535) a particle in front of a trailing title stays in the clause" -# rules.md#M2's Accepted pair on a released particle: the span 'DO -# Prof.' holds a particle with a title behind it, and P2's chain would -# run on over the title, so the particle's release is withdrawn and -# the clause keeps 'Smith DO' while the title stop still gives 'Prof.' -# up. The parent tree d9d80492 reads maiden 'Smith DO Prof.' as this -# baseline does, so #535 is the cause alone. Literal-anchored for the -# SLOT reason fix(#533)'s rules give. -name_regex = "^Jane Doe nee Smith DO Prof\\.$" -fields = ["title", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#533/#535) a particle behind a trailing title leaves the clause with it" -# The other half of that pair: written behind the title the particle -# is released with nothing behind it, so the title and 'DO' both leave -# the clause. Two causes together, checked against the parent tree -# d9d80492 (maiden 'Smith Prof.', suffix 'DO', reported): #533 already -# gives 'DO' up and reports the fork; #535 then reads the title in -# front of it off the clause, as for 'Jane Doe nee Smith Prof. MA'. -# Literal-anchored for the SLOT reason fix(#533)'s rules give. -name_regex = "^Jane Doe nee Smith Prof\\. DO$" -fields = ["title", "suffix", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#316/#399) a trailing title behind a maiden clause's first suffix word" -# rules.md#M2's Accepted pair on a title something ahead would take: -# the first-suffix-word stop ends the clause at 'PhD', so 'Prof.' -# stands behind the clause and reads as a trailing title. The parent -# tree d9d80492 and 2.3.0 read it as HEAD does, so #535 moves nothing; -# two causes, #399 (this baseline's particle chain swallowed the -# marker, middle 'van der Berg nee Smith PhD') and #316 (no trailing -# title reading here, family 'Prof.'). Literal-anchored: one corpus -# name, the doc's own example. -name_regex = "^Jane van der Berg nee Smith PhD Prof\\.$" -fields = ["title", "middle", "family", "suffix", "maiden"] - -[[change]] -issue = "fix(#399) a maiden marker bounds the particle chain ahead of a title the clause keeps" -# The other half of that pair: the title in front of 'PhD' stays in -# the clause, the particle chain ahead being able to take it -# (rules.md#H5's 'John van der Berg Prof.'). The parent tree d9d80492 -# reads it as HEAD does, so #535 moves nothing; #399 is the cause here -# -- this baseline's particle chain swallowed the marker and read -# family 'van der Berg nee Smith Prof.'. Literal-anchored: one corpus -# name, the doc's own example. -name_regex = "^Jane van der Berg nee Smith Prof\\. PhD$" -fields = ["family", "maiden"] - -[[change]] -issue = "fix(#535) a particle inside the clause keeps the credential in front of a trailing title" -# rules.md#M2's Accepted pair on a particle INSIDE the clause: the -# particle's chain would take the credential behind it, so the clause -# keeps 'Smith do MA' and only the title behind the credential leaves. -# The parent tree d9d80492 reads maiden 'Smith do MA Prof.' with no -# report, as this baseline does, so #535 is the whole cause. -# Literal-anchored: one corpus name, the doc's own example. -name_regex = "^Jane Doe nee Smith do MA Prof\\.$" -fields = ["title", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#533/#535) a particle inside the clause lets a credential behind the title go" -# The other half of that pair: with the title in front of 'MA', the -# credential is released and the title leaves with it, so maiden -# 'Smith do', suffix 'MA', title 'Prof.'. Two causes together, checked -# against the parent tree d9d80492 (maiden 'Smith do Prof.', suffix -# 'MA', reported): #533 already gives 'MA' up and reports the fork; -# #535 then reads the title in front of it off the clause. -# Literal-anchored: one corpus name, the doc's own example. -name_regex = "^Jane Doe nee Smith do Prof\\. MA$" -fields = ["title", "suffix", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#424) accepted: before a family comma the numeral the walk gives up goes to the family" -# rules.md#M2's Accepted row under #548: before a family comma the -# trailing numeral's stop is made over the peel alone, with no release -# question, so a lone numeral behind the clause goes to the family -# ('Doe V'). This baseline read maiden 'Smith V'; 2.2.0, 2.3.0 and the -# parent tree d9d80492 all read family 'Doe V', so the move predates -# #535 and is #424's -- the walk stopping before the trailing numeral -# -- recorded as accepted while #548 is open. Literal-anchored: one -# corpus name, the doc's own example. -name_regex = "^Doe nee Smith V, Jane$" -fields = ["family", "maiden"] - -[[change]] -issue = "fix(#397/#535) a link the clause stops at keeps a run that would not read off" -# rules.md#M2: a link that ends the clause gives up the words behind -# it only where the name left standing reads them as post-nominals or -# titles. With the title chained the link exception refuses 'i', but -# 'Jane Doe i DO Prof.' does not read 'i DO' off, so the clause keeps -# the link and the DO (reporting the kept credential) and only 'Prof.' -# leaves. Checked against the parent tree d9d80492, which read maiden -# 'Smith i DO Prof.': this baseline read 'i' as a plain suffix word -# that ended the clause, and #397 joins the link inside the clause; -# #535 then reads the title off. Literal-anchored: one corpus name, -# the doc's own example. -name_regex = "^Jane Doe nee Smith i DO Prof\\.$" -fields = ["title", "middle", "family", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#411/#535) the numeral stop asks the join question" -# 'Berg, abdul nee Smith V' moves at this baseline for two reasons -# together: #411 (the bound-given reserve) already decides how much -# of the tail the given join takes before #535 exists -- rules.md#P5: -# "The marker and the words it will take are not among the words to -# spare: they leave the name, so counting them asks the question -# about a name that will not exist" -- verified against the parent -# tree d9d80492, which already reads given 'abdul V', maiden 'Smith', -# differing from this baseline's given 'abdul nee', middle 'Smith', -# suffix 'V'. #535 then moves the released V from given back into -# maiden (given 'abdul V' -> 'abdul', maiden 'Smith' -> 'Smith V'). -# Literal-anchored for the same SLOT reason fix(#533)'s rules give. -name_regex = "^(?:Berg, abdul nee Smith V)$" -fields = ["given", "middle", "suffix", "maiden"] - -[[change]] -issue = "fix(#445/#533) the clause gives the credential up, and the one name word it leaves is the family" -# 'John née Jones Smith MA', whose diff has TWO causes from here and -# one rule has to explain the whole of it -- the shape the -# fix(#424/#445) rules above already record for the numeral and for -# the Title-cased acronym. #445 moves the one name word the clause -# leaves into the family (given 'John' -> family 'John'), which this -# baseline has never read, and #533 takes the MA out of the maiden -# name and into `suffix`. Attributing either to the other would be -# false, and a rule declaring only one pair of fields would explain -# nothing while the name went UNEXPLAINED. -# -# Case-SENSITIVE, unlike the fix(#445) rule above, and measured: the -# capitals are what decide here, so 'John née Jones Smith Ma' keeps -# its member in the clause and 'John née Jones Smith MA' gives it up. -# An (?i) anchor would stand ready to explain either name's diff with -# the other's reasoning. -name_regex = "^John née Jones Smith MA$" -fields = ["_ambiguities", "family", "given", "maiden", "suffix"] - [[change]] issue = "fix(#445/#533) a one-case clause keeps the credential, and the one name word it leaves is the family" # 'JOHN NEE JONES SMITH MA PHD'. One case, so no lean; a suffix word @@ -3608,8 +3387,8 @@ issue = "fix(#397) a link inside a maiden clause stays in the birth name" # maiden 'Puig -' -> 'Puig - i Soler' -- the same shape as the # neighbour's own reading with 'Mr.' dropped from it, unchanged by # the fix; only the corpus name is new. -name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler|Smith, John, PhD née Puig - i Soler)$" -fields = ["family", "maiden", "middle", "suffix"] +name_regex = "^(?:Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler|Smith, John, PhD née Puig - i Soler)$" +fields = ["family", "maiden", "middle"] orders = ["DEFAULT"] [[change]] @@ -4062,6 +3841,95 @@ issue = "fix(#597) halfwidth corner brackets enclose a nickname" name_regex = "^(?:山田 「タロー」 タロウ|山田「タロー」太郎|John 「Jack」 Smith)$" fields = ["given", "middle", "family", "nickname"] +[[change]] +# rules.md#M2: "In the given part after a family comma, in a part after a +# suffix comma, and behind a title, a suffix word or a connective, the +# marker is an ordinary word" (#601, 2026-10-04). The clause these names +# carried at this baseline is gone: `maiden` empties and its words read +# where the given part reads any word. What else differs is the given +# part's own reading of those words -- the spaced credential run +# (#436/#437), the link join (#397), P6's attachment of a trailing +# particle to the family (#405), the capitals lean (#289) and the +# trailing title chain (H5) -- each as the same text with the marker +# replaced by an ordinary word reads it. +issue = "fix(#601) a marker in the given part after a family comma is an ordinary word" +name_regex = "^(?:Berg, Jane van der nee Smith DO|Berg, abd née Jones|Berg, abdul nee Jones MA|Doe, Dr\\. nee Smith MA|Doe, J\\. nee MA ba|Doe, Jane Q\\. nee Smith MA|Doe, Jane nee Puig i Soler|Doe, Jane nee Smith|Doe, Jane nee Smith DO|Doe, Jane nee Smith Do|Doe, Jane nee Smith MA|Doe, Jane nee Smith MA Prof\\.|Doe, Jane nee Smith MA do|Doe, Jane nee Smith Ma|Doe, Jane nee Smith PhD MEng|Doe, Jane nee Smith do|Doe, Jane nee Smith ma)$" +fields = ["_ambiguities", "family", "given", "maiden", "middle", "suffix", "title"] + +[[change]] +# rules.md#M2: "In the given part after a family comma, in a part after a +# suffix comma, and behind a title, a suffix word or a connective, the +# marker is an ordinary word" (#601, 2026-10-04): a tail part is read +# wholly as suffixes, the marker and its words with it. +issue = "fix(#601) a marker in a part after a suffix comma is an ordinary word" +name_regex = "^(?:Smith, John, MD - née Jones Smith|Smith, John, PhD née Jones)$" +fields = ["maiden", "suffix"] + +[[change]] +# rules.md#M2: "In the given part after a family comma, in a part after a +# suffix comma, and behind a title, a suffix word or a connective, the +# marker is an ordinary word" (#601, 2026-10-04). Where a credential +# stands before the marker it also starts #602's run -- rules.md#S2: "A +# credential after the name core starts a run to the end of its part" -- +# which takes the marker and the words after it in. +issue = "fix(#601/#602) a marker behind a title, a suffix word or a connective is an ordinary word" +name_regex = "^(?:Dr\\. nee Smith PhD Prof\\.|Jane Doe Jr\\. nee Smith|Jane Doe Jr\\. nee Smith Ma|Jane Doe PhD nee Smith|Jane and née Jones)$" +fields = ["_ambiguities", "family", "given", "maiden", "suffix", "title"] + +[[change]] +# rules.md#M2: "It takes them up to the trailing run of post-nominals and +# titles that the end of the name reads as if the clause were not +# written" (#601, 2026-10-04), and "The run so found ends the name: its +# words are the suffixes and titles that reading makes them, and no join +# reaches into it". The old walk's release check kept these words in the +# clause where a join it modelled might take them; the run is bound now. +issue = "fix(#601) the clause ends at the clause-free name's trailing run, which the take consumes" +name_regex = "^(?:Jane Doe nee Smith DO DO|Jane van der Berg nee Prof\\. King\\. MA|Jane van der Berg nee Smith Prof\\.|John nee Prof\\. ba MA)$" +fields = ["_ambiguities", "family", "given", "maiden", "suffix", "title"] + +[[change]] +# rules.md#S2: "A credential after the name core starts a run to the end +# of its part" (#602, 2026-10-04): every later word of the part is a +# suffix except a title word, which reads as a title, and an absorbed +# name word is reported. +# 'Jane van der Berg née Jr Jones' (radar): the marker declines on +# the suffix word after it (rules.md#M2) and stays a word, so 'Jr' +# stands behind a name core and its run takes 'Jones' (2026-10-04). +issue = "fix(#602) a credential after the name core starts a run to the end of its part" +name_regex = "^(?:Eric H\\. Holder Jr\\. Attorney General|Jane van der Berg née Jr Jones|John Smith PhD Jones|John Smith PhD Prof\\. Ma|Smith, John PhD Jones)$" +fields = ["_ambiguities", "family", "middle", "suffix", "title"] + +[[change]] +# rules.md#M2 reads the clause-free name, whose credential starts #602's +# run -- rules.md#S2: "A credential after the name core starts a run to +# the end of its part" -- and the take consumes that run: maiden 'Smith', +# suffix 'PhD Smith' (2026-10-04). +issue = "fix(#601/#602) a credential in the clause starts a run the take consumes" +name_regex = "^Jane Doe nee Smith PhD Smith$" +fields = ["_ambiguities", "family", "middle", "suffix"] + +[[change]] +# The initials half of 'Jane and née Jones' (#601, 2026-10-04): a +# marker behind a connective is an ordinary word -- rules.md#M2: "In +# the given part after a family comma, in a part after a suffix comma, +# and behind a title, a suffix word or a connective, the marker is an +# ordinary word" -- so it joins the connective run 'Jane and née', and +# the run initials each of its words but the connective (R3). The +# roles' half is fix(#601/#602)'s marker-behind-a-connective rule above. +issue = "fix(#601) a marker behind a connective joins the connective run, and the run's initials follow" +name_regex = "^Jane and née Jones$" +fields = ["_initials"] + +[[change]] +# rules.md#P6: "a particle ending the name attaches to that family name", +# and a post-nominal written behind the particle does not end the name +# for this purpose. rules.md#S2's example for the given-part run leaving +# the particle to P6 (#601/#602 review, 2026-10-04); the attachment +# itself landed in 2.2 (#379). +issue = "fix(#379) a tussenvoegsel behind a post-nominal after a family comma attaches" +name_regex = "^Smith, John PhD de Jr\\.$" +fields = ["middle", "family"] + # --------------------------------------------------------------- # #604: THE IRISH AND MALAY PATRONYMIC PARTICLES. # 'ó' and 'ní' join the never-given particles, 'ua', 'binti' and diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index 9caff2e6..3fe80b3b 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -845,17 +845,6 @@ issue = "fix(#424/#445) the maiden walk stops before the trailing numeral, and t name_regex = "(?i)^john\\s+n[ée]e\\s+jones\\s+smith\\s+v$" fields = ["given", "family", "suffix", "maiden", "_ambiguities"] -[[change]] -issue = "fix(#424) a marker followed only by the numeral is just a word" -# 'Jane Smith née V': the peel is read from the marker, so the V has -# the piece before it the fork wants and reads as the suffix; nothing -# follows the marker but a suffix, the pass declines, and the marker -# stays a word -- as for 'Jane Smith née PhD'. maiden 'V' -> middle -# 'Smith', family 'née', suffix 'V', which is 1.4.0's reading, so -# this rule has no 1.4.0 twin. A rules.md example. -name_regex = "(?i)^jane\\s+smith\\s+n[ée]e\\s+v$" -fields = ["middle", "family", "suffix", "maiden", "_ambiguities"] - [[change]] issue = "fix(#369) the bound given-name join takes a particle-and-bound word, so no fork is reported" # 'Sheik Abu Bakar': the fields are byte-identical to this baseline @@ -1162,6 +1151,7 @@ fields = ["given", "family", "_ambiguities"] [[change]] issue = "fix(#411) the bound-given reserve stops counting words the maiden name takes" +dormant = "no corpus name since #601 (2026-10-04): the rules.md#M2 examples that carried this shape left the doc with the walk they illustrated. The reserve still counts only the words the clause leaves (the take runs before every join), so the rule is kept for the next name that shows it" # 'van der Berg, abdul nee Jones': P5 reserves a name word so the join # always leaves a family name behind, but the reserve was counted while # the marker and the maiden name were still pieces -- group's marker @@ -2834,7 +2824,7 @@ issue = "fix(#533) a credential ending a maiden clause reads as a credential" # Literal-anchored for the reason fix(#531)'s rules give: the subject # is a SLOT, so a regex for it would claim the names whose WRITING # declines the word as readily as the ones it takes. -name_regex = "^(?:Doe, Jane Q\\. nee Smith MA|Doe, Jane nee Smith DO|Doe, Jane nee Smith MA|Doe, Jane nee Smith ma|JANE DOE NEE SMITH MA|JANE DOE NEE YO-YO MA|Jane Doe geb\\. Smith MA|Jane Doe nee Smith MA|Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD|Jane Doe nee Smith V MA|Jane Doe nee Smith do|jane doe nee smith ma)$" +name_regex = "^(?:JANE DOE NEE SMITH MA|JANE DOE NEE YO-YO MA|Jane Doe geb\\. Smith MA|Jane Doe nee Smith MA|Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD|Jane Doe nee Smith V MA|Jane Doe nee Smith do|jane doe nee smith ma)$" fields = ["_ambiguities", "maiden", "suffix"] [[change]] @@ -2880,45 +2870,9 @@ issue = "fix(#533) the maiden clause reports the credential it keeps" # half at rules.md#M2's new Accepted boundary: the credential in front # of the marker speaks for no member of the clause, the Title-case 'Ma' # is kept, and the clause reports it. -name_regex = "^(?:Doe, Dr\\. nee Smith MA|Doe, J\\. nee MA ba|Doe, Jane nee Smith Ma|Jane Doe Jr\\. nee Smith Ma|Jane Doe nee King\\. ba|Jane Doe nee MA|Jane Doe nee MA PhD|Jane Doe nee Smith DO DO|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma)$" +name_regex = "^(?:Jane Doe nee King\\. ba|Jane Doe nee MA|Jane Doe nee MA PhD|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma)$" fields = ["_ambiguities"] -[[change]] -issue = "fix(#411/#533) the bound-given pair after a comma, and the clause it keeps its credential in" -# 'Berg, abdul nee Jones MA'. Two changes in one name and neither is -# enough on its own, which is why it is a rule rather than an -# alternative in either. The clause consumption (#274) and the -# reserve that stops counting the words it takes (#411) are what move -# `given`, `middle` and `maiden` against this baseline; the report is -# #533's. -# -# What #533 does here is NOT a role move, and the comment is the only -# place that can say so: an earlier round of this branch released the -# 'MA' from the clause and let P5's lenient post-comma join swallow -# it into the pair, reading given 'abdul MA' in silence. The review -# measured that as M2's invariant broken -- a word the clause gave up -# landing in a NAME part -- so the stop is declined, the clause keeps -# 'Jones MA', and the fork is reported instead. -name_regex = "^Berg, abdul nee Jones MA$" -fields = ["_ambiguities", "given", "maiden", "middle"] - -[[change]] -issue = "fix(#399/#533) a maiden marker bounds the particle chain, and the clause keeps the credential the chain would have taken" -# 'Berg, Jane van der nee Smith DO'. The marker bounds the particle -# join arriving from its left (#399, rules.md#M2), which is what -# moves `middle`, `family` and `maiden` against this baseline, and -# the chain still reports its own PARTICLE_OR_GIVEN. -# -# The SUFFIX_OR_NAME beside it is #533's, and it is a report without -# a role move on purpose. 'DO' is particle vocabulary, so releasing -# it from the clause would put it back in reach of the very chain the -# marker just bounded -- family 'van der DO Berg', a word crossing -# from the birth name into the current one, which an earlier round of -# this branch did in silence. rules.md#M2's invariant declines the -# stop, the clause keeps 'Smith DO', and the fork is reported. -name_regex = "^Berg, Jane van der nee Smith DO$" -fields = ["_ambiguities", "family", "maiden", "middle"] - [[change]] issue = "fix(#335/#533) a marker-led bracketed clause is the maiden name whatever its last word is" # 'Jane Doe (nee Smith MA)', 'Jane Doe (nee Smith Ma)', 'Jane Doe @@ -2976,26 +2930,6 @@ issue = "fix(#533) restores the suffix reading a 2.4 retag had moved into the ma name_regex = "^John Smith nee Jones R\\.A\\.I\\.$" fields = ["_ambiguities"] -[[change]] -issue = "fix(#445/#533) the clause gives the credential up, and the one name word it leaves is the family" -# 'John née Jones Smith MA', whose diff has TWO causes from here and -# one rule has to explain the whole of it -- the shape the -# fix(#424/#445) rules above already record for the numeral and for -# the Title-cased acronym. #445 moves the one name word the clause -# leaves into the family (given 'John' -> family 'John'), which this -# baseline has never read, and #533 takes the MA out of the maiden -# name and into `suffix`. Attributing either to the other would be -# false, and a rule declaring only one pair of fields would explain -# nothing while the name went UNEXPLAINED. -# -# Case-SENSITIVE, unlike the fix(#445) rule above, and measured: the -# capitals are what decide here, so 'John née Jones Smith Ma' keeps -# its member in the clause and 'John née Jones Smith MA' gives it up. -# An (?i) anchor would stand ready to explain either name's diff with -# the other's reasoning. -name_regex = "^John née Jones Smith MA$" -fields = ["_ambiguities", "family", "given", "maiden", "suffix"] - [[change]] issue = "fix(#445/#533) a one-case clause keeps the credential, and the one name word it leaves is the family" # 'JOHN NEE JONES SMITH MA PHD'. One case, so no lean; a suffix word @@ -3097,23 +3031,6 @@ fields = ["_ambiguities", "maiden", "suffix"] # criterion and its measured blast radius. # --------------------------------------------------------------- -[[change]] -issue = "fix(#399) a maiden marker bounds the particle chain that swallowed it, two trailing words" -# 'Jane van der Berg nee Smith Prof.' moves {family, maiden} at this -# baseline whether or not #535 exists -- verified against the parent -# tree d9d80492, which already reads family 'van der Berg', maiden -# 'Smith Prof.' identically to HEAD, before #535 touched anything. -# It is the #399 rule's own shape: rules.md#P2: "a maiden marker -# takes the words after it (M2), or the name ends" -- but the chain -# here swallows the marker and BOTH trailing words before the maiden -# consumer could see any of them, and the general rule's anchor takes -# exactly one word after the marker (`\\sn[ée]e\\s+\\S+$`), so this -# name's two-word clause needs its own literal rather than widening -# that regex's reach past what its own review measured. #535 moves -# nothing here. -name_regex = "^Jane van der Berg nee Smith Prof\\.$" -fields = ["family", "maiden"] - [[change]] issue = "fix(#296/#535) a trailing title after a dropped postnominal Dr." # 'Jane Doe nee Prof. Dr.' moves {suffix, title} at this baseline for @@ -3147,8 +3064,11 @@ fields = ["_ambiguities", "maiden", "suffix", "title"] [[change]] issue = "fix(#535) a trailing title ends the maiden clause" -# rules.md#M2: "Where a trailing rule reads the words, a trailing -# title ends it too" -- the walk reads the end of the name through +# rules.md#M2: "It takes them up to the trailing run +# of post-nominals and titles that the end of the name reads as if +# the clause were not written" (#601; until #601 this cited the +# sentence it replaced, "a trailing title ends it too") -- the end of +# the name is read through # the trailing title chain -- rules.md#H5: "successive single words # that wear the abbreviation shape and are title vocabulary chain # into the title from the end" -- so the title leaves the clause and @@ -3162,150 +3082,9 @@ issue = "fix(#535) a trailing title ends the maiden clause" # each because the parent tree d9d80492 already differed from this # baseline before #535 touched it; the seven names here are ones # where #535 is the sole cause, verified the same way. -name_regex = "^(?:Doe, Jane nee Smith MA Prof\\.|Jane Doe \\x28nee Smith Prof\\.\\x29|Jane Doe nee Smith King\\.|Jane Doe nee Smith MA Prof\\.|Jane Doe nee Smith Ma Prof\\.|Jane Doe nee Smith Prof\\.|Jane Doe nee Smith V Prof\\.|Mary Smith née Jones Prof\\.)$" +name_regex = "^(?:Jane Doe \\x28nee Smith Prof\\.\\x29|Jane Doe nee Smith King\\.|Jane Doe nee Smith MA Prof\\.|Jane Doe nee Smith Ma Prof\\.|Jane Doe nee Smith Prof\\.|Jane Doe nee Smith V Prof\\.|Mary Smith née Jones Prof\\.)$" fields = ["_ambiguities", "maiden", "suffix", "title"] -[[change]] -issue = "fix(#342/#445/#535) a title behind a credential the clause keeps still leaves it" -# rules.md#M2's Accepted transparency boundary, the half that gives -# the title up: the clause keeps 'ba' (one name word left, none to -# spare) and the title behind it still ends the clause, so title -# 'Prof.', family 'Doe', maiden 'Smith ba'. Three causes together, -# checked against the parent tree d9d80492 (family 'Doe', maiden -# 'Smith ba Prof.', no report): before #535, #342 marked 'ba' an ambiguous acronym, so the clause no longer stops at it as a plain suffix word, and #445 makes the lone name word beside a maiden clause the family; -# #535 then reads the title off the end of the clause. -# Literal-anchored for the SLOT reason fix(#533)'s rules give. -name_regex = "^Doe nee Smith ba Prof\\.$" -fields = ["title", "given", "middle", "family", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#342/#445/#533) a title in front of a credential the clause keeps stays in it" -# rules.md#M2's Accepted transparency boundary, the half that keeps -# the title: a title in FRONT of the kept 'ba' cannot leave without -# it, so the clause keeps 'Smith Prof. ba' and reports the kept -# credential. The parent tree d9d80492 reads it exactly as HEAD -# does, so #535 moves nothing here: #342 marked 'ba' an ambiguous acronym, so the clause no longer stops at it as a plain suffix word, and #445 makes the lone name word beside a maiden clause the family, and #533 reports -# the credential the clause keeps. -# Literal-anchored for the SLOT reason fix(#533)'s rules give. -name_regex = "^Doe nee Smith Prof\\. ba$" -fields = ["given", "family", "suffix", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#535) a particle in front of a trailing title stays in the clause" -# rules.md#M2's Accepted pair on a released particle: the span 'DO -# Prof.' holds a particle with a title behind it, and P2's chain would -# run on over the title, so the particle's release is withdrawn and -# the clause keeps 'Smith DO' while the title stop still gives 'Prof.' -# up. The parent tree d9d80492 reads maiden 'Smith DO Prof.' as this -# baseline does, so #535 is the cause alone. Literal-anchored for the -# SLOT reason fix(#533)'s rules give. -name_regex = "^Jane Doe nee Smith DO Prof\\.$" -fields = ["title", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#533/#535) a particle behind a trailing title leaves the clause with it" -# The other half of that pair: written behind the title the particle -# is released with nothing behind it, so the title and 'DO' both leave -# the clause. Two causes together, checked against the parent tree -# d9d80492 (maiden 'Smith Prof.', suffix 'DO', reported): #533 already -# gives 'DO' up and reports the fork; #535 then reads the title in -# front of it off the clause, as for 'Jane Doe nee Smith Prof. MA'. -# Literal-anchored for the SLOT reason fix(#533)'s rules give. -name_regex = "^Jane Doe nee Smith Prof\\. DO$" -fields = ["title", "suffix", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#316/#399) a trailing title behind a maiden clause's first suffix word" -# rules.md#M2's Accepted pair on a title something ahead would take: -# the first-suffix-word stop ends the clause at 'PhD', so 'Prof.' -# stands behind the clause and reads as a trailing title. The parent -# tree d9d80492 and 2.3.0 read it as HEAD does, so #535 moves nothing; -# two causes, #399 (this baseline's particle chain swallowed the -# marker, middle 'van der Berg nee Smith PhD') and #316 (no trailing -# title reading here, family 'Prof.'). Literal-anchored: one corpus -# name, the doc's own example. -name_regex = "^Jane van der Berg nee Smith PhD Prof\\.$" -fields = ["title", "middle", "family", "suffix", "maiden"] - -[[change]] -issue = "fix(#399) a maiden marker bounds the particle chain ahead of a title the clause keeps" -# The other half of that pair: the title in front of 'PhD' stays in -# the clause, the particle chain ahead being able to take it -# (rules.md#H5's 'John van der Berg Prof.'). The parent tree d9d80492 -# reads it as HEAD does, so #535 moves nothing; #399 is the cause here -# -- this baseline's particle chain swallowed the marker and read -# family 'van der Berg nee Smith Prof.'. Literal-anchored: one corpus -# name, the doc's own example. -name_regex = "^Jane van der Berg nee Smith Prof\\. PhD$" -fields = ["family", "maiden"] - -[[change]] -issue = "fix(#535) a particle inside the clause keeps the credential in front of a trailing title" -# rules.md#M2's Accepted pair on a particle INSIDE the clause: the -# particle's chain would take the credential behind it, so the clause -# keeps 'Smith do MA' and only the title behind the credential leaves. -# The parent tree d9d80492 reads maiden 'Smith do MA Prof.' with no -# report, as this baseline does, so #535 is the whole cause. -# Literal-anchored: one corpus name, the doc's own example. -name_regex = "^Jane Doe nee Smith do MA Prof\\.$" -fields = ["title", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#533/#535) a particle inside the clause lets a credential behind the title go" -# The other half of that pair: with the title in front of 'MA', the -# credential is released and the title leaves with it, so maiden -# 'Smith do', suffix 'MA', title 'Prof.'. Two causes together, checked -# against the parent tree d9d80492 (maiden 'Smith do Prof.', suffix -# 'MA', reported): #533 already gives 'MA' up and reports the fork; -# #535 then reads the title in front of it off the clause. -# Literal-anchored: one corpus name, the doc's own example. -name_regex = "^Jane Doe nee Smith do Prof\\. MA$" -fields = ["title", "suffix", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#424) accepted: before a family comma the numeral the walk gives up goes to the family" -# rules.md#M2's Accepted row under #548: before a family comma the -# trailing numeral's stop is made over the peel alone, with no release -# question, so a lone numeral behind the clause goes to the family -# ('Doe V'). This baseline read maiden 'Smith V'; 2.2.0, 2.3.0 and the -# parent tree d9d80492 all read family 'Doe V', so the move predates -# #535 and is #424's -- the walk stopping before the trailing numeral -# -- recorded as accepted while #548 is open. Literal-anchored: one -# corpus name, the doc's own example. -name_regex = "^Doe nee Smith V, Jane$" -fields = ["family", "maiden"] - -[[change]] -issue = "fix(#397/#535) a link the clause stops at keeps a run that would not read off" -# rules.md#M2: a link that ends the clause gives up the words behind -# it only where the name left standing reads them as post-nominals or -# titles. With the title chained the link exception refuses 'i', but -# 'Jane Doe i DO Prof.' does not read 'i DO' off, so the clause keeps -# the link and the DO (reporting the kept credential) and only 'Prof.' -# leaves. Checked against the parent tree d9d80492, which read maiden -# 'Smith i DO Prof.': this baseline read 'i' as a plain suffix word -# that ended the clause, and #397 joins the link inside the clause; -# #535 then reads the title off. Literal-anchored: one corpus name, -# the doc's own example. -name_regex = "^Jane Doe nee Smith i DO Prof\\.$" -fields = ["title", "middle", "family", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#411/#535) the numeral stop asks the join question" -# 'Berg, abdul nee Smith V' moves at this baseline for two reasons -# together: #411 (the bound-given reserve) already decides how much -# of the tail the given join takes before #535 exists -- rules.md#P5: -# "The marker and the words it will take are not among the words to -# spare: they leave the name, so counting them asks the question -# about a name that will not exist" -- verified against the parent -# tree d9d80492, which already reads given 'abdul V', maiden 'Smith', -# differing from this baseline's given 'abdul nee', middle 'Smith', -# suffix 'V'. #535 then moves the released V from given back into -# maiden (given 'abdul V' -> 'abdul', maiden 'Smith' -> 'Smith V'). -# Literal-anchored for the same SLOT reason fix(#533)'s rules give. -name_regex = "^(?:Berg, abdul nee Smith V)$" -fields = ["given", "middle", "suffix", "maiden"] - [[change]] issue = "fix(#397) the Catalan/Polish link joins two surnames" # 'i' is connective vocabulary since #397, and a connective counts as @@ -3520,8 +3299,8 @@ issue = "fix(#397) a link inside a maiden clause stays in the birth name" # maiden 'Puig -' -> 'Puig - i Soler' -- the same shape as the # neighbour's own reading with 'Mr.' dropped from it, unchanged by # the fix; only the corpus name is new. -name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler|Smith, John, PhD née Puig - i Soler)$" -fields = ["family", "maiden", "middle", "suffix"] +name_regex = "^(?:Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler|Smith, John, PhD née Puig - i Soler)$" +fields = ["family", "maiden", "middle"] orders = ["DEFAULT"] [[change]] @@ -4007,6 +3786,95 @@ issue = "fix(#597) halfwidth corner brackets enclose a nickname" name_regex = "^(?:山田 「タロー」 タロウ|山田「タロー」太郎|John 「Jack」 Smith)$" fields = ["given", "middle", "family", "nickname"] +[[change]] +# rules.md#M2: "In the given part after a family comma, in a part after a +# suffix comma, and behind a title, a suffix word or a connective, the +# marker is an ordinary word" (#601, 2026-10-04). The clause these names +# carried at this baseline is gone: `maiden` empties and its words read +# where the given part reads any word. What else differs is the given +# part's own reading of those words -- the spaced credential run +# (#436/#437), the link join (#397), P6's attachment of a trailing +# particle to the family (#405), the capitals lean (#289) and the +# trailing title chain (H5) -- each as the same text with the marker +# replaced by an ordinary word reads it. +issue = "fix(#601) a marker in the given part after a family comma is an ordinary word" +name_regex = "^(?:Berg, Jane van der nee Smith DO|Berg, abd née Jones|Berg, abdul nee Jones MA|Doe, Dr\\. nee Smith MA|Doe, J\\. nee MA ba|Doe, Jane Q\\. nee Smith MA|Doe, Jane nee Puig i Soler|Doe, Jane nee Smith|Doe, Jane nee Smith DO|Doe, Jane nee Smith Do|Doe, Jane nee Smith MA|Doe, Jane nee Smith MA Prof\\.|Doe, Jane nee Smith MA do|Doe, Jane nee Smith Ma|Doe, Jane nee Smith PhD MEng|Doe, Jane nee Smith do|Doe, Jane nee Smith ma)$" +fields = ["_ambiguities", "family", "given", "maiden", "middle", "suffix", "title"] + +[[change]] +# rules.md#M2: "In the given part after a family comma, in a part after a +# suffix comma, and behind a title, a suffix word or a connective, the +# marker is an ordinary word" (#601, 2026-10-04): a tail part is read +# wholly as suffixes, the marker and its words with it. +issue = "fix(#601) a marker in a part after a suffix comma is an ordinary word" +name_regex = "^(?:Smith, John, MD - née Jones Smith|Smith, John, PhD née Jones)$" +fields = ["maiden", "suffix"] + +[[change]] +# rules.md#M2: "In the given part after a family comma, in a part after a +# suffix comma, and behind a title, a suffix word or a connective, the +# marker is an ordinary word" (#601, 2026-10-04). Where a credential +# stands before the marker it also starts #602's run -- rules.md#S2: "A +# credential after the name core starts a run to the end of its part" -- +# which takes the marker and the words after it in. +issue = "fix(#601/#602) a marker behind a title, a suffix word or a connective is an ordinary word" +name_regex = "^(?:Dr\\. nee Smith PhD Prof\\.|Jane Doe Jr\\. nee Smith|Jane Doe Jr\\. nee Smith Ma|Jane Doe PhD nee Smith|Jane and née Jones)$" +fields = ["_ambiguities", "family", "given", "maiden", "suffix", "title"] + +[[change]] +# rules.md#M2: "It takes them up to the trailing run of post-nominals and +# titles that the end of the name reads as if the clause were not +# written" (#601, 2026-10-04), and "The run so found ends the name: its +# words are the suffixes and titles that reading makes them, and no join +# reaches into it". The old walk's release check kept these words in the +# clause where a join it modelled might take them; the run is bound now. +issue = "fix(#601) the clause ends at the clause-free name's trailing run, which the take consumes" +name_regex = "^(?:Jane Doe nee Smith DO DO|Jane van der Berg nee Prof\\. King\\. MA|Jane van der Berg nee Smith Prof\\.|John nee Prof\\. ba MA)$" +fields = ["_ambiguities", "family", "given", "maiden", "suffix", "title"] + +[[change]] +# rules.md#S2: "A credential after the name core starts a run to the end +# of its part" (#602, 2026-10-04): every later word of the part is a +# suffix except a title word, which reads as a title, and an absorbed +# name word is reported. +# 'Jane van der Berg née Jr Jones' (radar): the marker declines on +# the suffix word after it (rules.md#M2) and stays a word, so 'Jr' +# stands behind a name core and its run takes 'Jones' (2026-10-04). +issue = "fix(#602) a credential after the name core starts a run to the end of its part" +name_regex = "^(?:Eric H\\. Holder Jr\\. Attorney General|Jane van der Berg née Jr Jones|John Smith PhD Jones|John Smith PhD Prof\\. Ma|Smith, John PhD Jones)$" +fields = ["_ambiguities", "family", "middle", "suffix", "title"] + +[[change]] +# rules.md#M2 reads the clause-free name, whose credential starts #602's +# run -- rules.md#S2: "A credential after the name core starts a run to +# the end of its part" -- and the take consumes that run: maiden 'Smith', +# suffix 'PhD Smith' (2026-10-04). +issue = "fix(#601/#602) a credential in the clause starts a run the take consumes" +name_regex = "^Jane Doe nee Smith PhD Smith$" +fields = ["_ambiguities", "family", "middle", "suffix"] + +[[change]] +# The initials half of 'Jane and née Jones' (#601, 2026-10-04): a +# marker behind a connective is an ordinary word -- rules.md#M2: "In +# the given part after a family comma, in a part after a suffix comma, +# and behind a title, a suffix word or a connective, the marker is an +# ordinary word" -- so it joins the connective run 'Jane and née', and +# the run initials each of its words but the connective (R3). The +# roles' half is fix(#601/#602)'s marker-behind-a-connective rule above. +issue = "fix(#601) a marker behind a connective joins the connective run, and the run's initials follow" +name_regex = "^Jane and née Jones$" +fields = ["_initials"] + +[[change]] +# rules.md#P6: "a particle ending the name attaches to that family name", +# and a post-nominal written behind the particle does not end the name +# for this purpose. rules.md#S2's example for the given-part run leaving +# the particle to P6 (#601/#602 review, 2026-10-04); the attachment +# itself landed in 2.2 (#379). +issue = "fix(#379) a tussenvoegsel behind a post-nominal after a family comma attaches" +name_regex = "^Smith, John PhD de Jr\\.$" +fields = ["middle", "family"] + # --------------------------------------------------------------- # #604: THE IRISH AND MALAY PATRONYMIC PARTICLES. # 'ó' and 'ní' join the never-given particles, 'ua', 'binti' and diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index e9beaefc..1577ce26 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -1397,7 +1397,7 @@ issue = "fix(#533) a credential ending a maiden clause reads as a credential" # would claim the names whose WRITING declines the word -- which keep # their maiden reading and have the rule below -- as readily as the # ones it takes. _MUST_NOT_MATCH carries both directions. -name_regex = "^(?:Doe, J\\. nee MA ba|Doe, Jane Q\\. nee Smith MA|Doe, Jane nee Smith DO|Doe, Jane nee Smith MA|Doe, Jane nee Smith ma|JANE DOE NEE SMITH MA|JANE DOE NEE YO-YO MA|Jane Doe geb\\. Smith MA|Jane Doe nee Smith MA|Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD|Jane Doe nee Smith V MA|Jane Doe nee Smith do|John née Jones Smith MA|Maria Kowalska z domu Nowak MA|jane doe nee smith ma)$" +name_regex = "^(?:JANE DOE NEE SMITH MA|JANE DOE NEE YO-YO MA|Jane Doe geb\\. Smith MA|Jane Doe nee Smith MA|Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD|Jane Doe nee Smith V MA|Jane Doe nee Smith do|John née Jones Smith MA|Maria Kowalska z domu Nowak MA|jane doe nee smith ma)$" fields = ["_ambiguities", "maiden", "suffix"] [[change]] @@ -1450,7 +1450,7 @@ issue = "fix(#533) the maiden clause reports the credential it keeps" # half at rules.md#M2's new Accepted boundary: the credential in front # of the marker speaks for no member of the clause, the Title-case 'Ma' # is kept, and the clause reports it. -name_regex = "^(?:Berg, Jane van der nee Smith DO|Berg, abdul nee Jones MA|Doe, Dr\\. nee Smith MA|Doe, Jane nee Smith Do|Doe, Jane nee Smith MA do|Doe, Jane nee Smith Ma|Doe, Jane nee Smith do|JOHN NEE JONES SMITH MA PHD|Jane Doe Jr\\. nee Smith Ma|Jane Doe nee King\\. ba|Jane Doe nee MA|Jane Doe nee MA PhD|Jane Doe nee Smith DO DO|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma|John née Jones Smith Ma)$" +name_regex = "^(?:JOHN NEE JONES SMITH MA PHD|Jane Doe nee King\\. ba|Jane Doe nee MA|Jane Doe nee MA PhD|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma|John née Jones Smith Ma)$" fields = ["_ambiguities"] [[change]] @@ -1552,8 +1552,11 @@ fields = ["_ambiguities", "maiden", "suffix", "title"] [[change]] issue = "fix(#535) a trailing title ends the maiden clause" -# rules.md#M2: "Where a trailing rule reads the words, a trailing -# title ends it too" -- the walk reads the end of the name through +# rules.md#M2: "It takes them up to the trailing run +# of post-nominals and titles that the end of the name reads as if +# the clause were not written" (#601; until #601 this cited the +# sentence it replaced, "a trailing title ends it too") -- the end of +# the name is read through # the trailing title chain -- rules.md#H5: "successive single words # that wear the abbreviation shape and are title vocabulary chain # into the title from the end" -- so the title leaves the clause and @@ -1570,136 +1573,9 @@ issue = "fix(#535) a trailing title ends the maiden clause" # split at THIS baseline -- #296 and #399 both predate it, so this # baseline already reads them as HEAD does apart from #535's own move, # unlike at 2.0.0/2.1.0.) -name_regex = "^(?:Doe, Jane nee Smith MA Prof\\.|Jane Doe \\x28nee Smith Prof\\.\\x29|Jane Doe nee Prof\\. Dr\\.|Jane Doe nee Smith King\\.|Jane Doe nee Smith MA Prof\\.|Jane Doe nee Smith Ma Prof\\.|Jane Doe nee Smith Prof\\.|Jane Doe nee Smith V Prof\\.|Mary Smith née Jones Prof\\.)$" +name_regex = "^(?:Jane Doe \\x28nee Smith Prof\\.\\x29|Jane Doe nee Prof\\. Dr\\.|Jane Doe nee Smith King\\.|Jane Doe nee Smith MA Prof\\.|Jane Doe nee Smith Ma Prof\\.|Jane Doe nee Smith Prof\\.|Jane Doe nee Smith V Prof\\.|Mary Smith née Jones Prof\\.)$" fields = ["title", "maiden", "suffix", "_ambiguities"] -[[change]] -issue = "fix(#342/#535) a title behind a credential the clause keeps still leaves it" -# rules.md#M2's Accepted transparency boundary, the half that gives -# the title up: the clause keeps 'ba' (one name word left, none to -# spare) and the title behind it still ends the clause, so title -# 'Prof.', family 'Doe', maiden 'Smith ba'. Two causes together, -# checked against the parent tree d9d80492 (family 'Doe', maiden -# 'Smith ba Prof.', no report): before #535, #342 marked 'ba' an ambiguous acronym, so the clause no longer stops at it as a plain suffix word; -# #535 then reads the title off the end of the clause. -# Literal-anchored for the SLOT reason fix(#533)'s rules give. -name_regex = "^Doe nee Smith ba Prof\\.$" -fields = ["title", "given", "middle", "family", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#342/#533) a title in front of a credential the clause keeps stays in it" -# rules.md#M2's Accepted transparency boundary, the half that keeps -# the title: a title in FRONT of the kept 'ba' cannot leave without -# it, so the clause keeps 'Smith Prof. ba' and reports the kept -# credential. The parent tree d9d80492 reads it exactly as HEAD -# does, so #535 moves nothing here: #342 marked 'ba' an ambiguous acronym, so the clause no longer stops at it as a plain suffix word, and #533 reports -# the credential the clause keeps. -# Literal-anchored for the SLOT reason fix(#533)'s rules give. -name_regex = "^Doe nee Smith Prof\\. ba$" -fields = ["suffix", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#535) a particle in front of a trailing title stays in the clause" -# rules.md#M2's Accepted pair on a released particle: the span 'DO -# Prof.' holds a particle with a title behind it, and P2's chain would -# run on over the title, so the particle's release is withdrawn and -# the clause keeps 'Smith DO' while the title stop still gives 'Prof.' -# up. The parent tree d9d80492 reads maiden 'Smith DO Prof.' as this -# baseline does, so #535 is the cause alone. Literal-anchored for the -# SLOT reason fix(#533)'s rules give. -name_regex = "^Jane Doe nee Smith DO Prof\\.$" -fields = ["title", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#533/#535) a particle behind a trailing title leaves the clause with it" -# The other half of that pair: written behind the title the particle -# is released with nothing behind it, so the title and 'DO' both leave -# the clause. Two causes together, checked against the parent tree -# d9d80492 (maiden 'Smith Prof.', suffix 'DO', reported): #533 already -# gives 'DO' up and reports the fork; #535 then reads the title in -# front of it off the clause, as for 'Jane Doe nee Smith Prof. MA'. -# Literal-anchored for the SLOT reason fix(#533)'s rules give. -name_regex = "^Jane Doe nee Smith Prof\\. DO$" -fields = ["title", "suffix", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#316) a trailing title behind a maiden clause's first suffix word" -# rules.md#M2's Accepted pair on a title something ahead would take: -# the first-suffix-word stop ends the clause at 'PhD', so 'Prof.' -# stands behind the clause and reads as a trailing title. The parent -# tree d9d80492 and 2.3.0 read it as HEAD does, so #535 moves nothing; -# #316 is the whole cause at this baseline, which read middle 'van der -# Berg PhD', family 'Prof.'. Literal-anchored: one corpus name, the -# doc's own example. -name_regex = "^Jane van der Berg nee Smith PhD Prof\\.$" -fields = ["title", "middle", "family", "suffix"] - -[[change]] -issue = "fix(#535) a particle inside the clause keeps the credential in front of a trailing title" -# rules.md#M2's Accepted pair on a particle INSIDE the clause: the -# particle's chain would take the credential behind it, so the clause -# keeps 'Smith do MA' and only the title behind the credential leaves. -# The parent tree d9d80492 reads maiden 'Smith do MA Prof.' with no -# report, as this baseline does, so #535 is the whole cause. -# Literal-anchored: one corpus name, the doc's own example. -name_regex = "^Jane Doe nee Smith do MA Prof\\.$" -fields = ["title", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#533/#535) a particle inside the clause lets a credential behind the title go" -# The other half of that pair: with the title in front of 'MA', the -# credential is released and the title leaves with it, so maiden -# 'Smith do', suffix 'MA', title 'Prof.'. Two causes together, checked -# against the parent tree d9d80492 (maiden 'Smith do Prof.', suffix -# 'MA', reported): #533 already gives 'MA' up and reports the fork; -# #535 then reads the title in front of it off the clause. -# Literal-anchored: one corpus name, the doc's own example. -name_regex = "^Jane Doe nee Smith do Prof\\. MA$" -fields = ["title", "suffix", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#397/#535) a link the clause stops at keeps a run that would not read off" -# rules.md#M2: a link that ends the clause gives up the words behind -# it only where the name left standing reads them as post-nominals or -# titles. With the title chained the link exception refuses 'i', but -# 'Jane Doe i DO Prof.' does not read 'i DO' off, so the clause keeps -# the link and the DO (reporting the kept credential) and only 'Prof.' -# leaves. Checked against the parent tree d9d80492, which read maiden -# 'Smith i DO Prof.': this baseline read 'i' as a plain suffix word -# that ended the clause, and #397 joins the link inside the clause; -# #535 then reads the title off. Literal-anchored: one corpus name, -# the doc's own example. -name_regex = "^Jane Doe nee Smith i DO Prof\\.$" -fields = ["title", "middle", "family", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#535) the numeral stop asks the join question" -# rules.md#M2 (#535): the numeral stop now asks the join question -# too -- released, the numeral was taken by the bound-given join -# after a family comma and read as part of the given name -# ('Berg, abdul nee Smith V' read given 'abdul V' at this -# baseline and at the other 2.2/2.3 baseline, before #535 ever ran; -# 2.0/2.1 read given 'abdul nee', suffix 'V' instead -- their own -# copy of this rule is fix(#411/#535) for that reason). -# Literal-anchored for the same SLOT reason fix(#533)'s rules give. -name_regex = "^(?:Berg, abdul nee Smith V)$" -fields = ["given", "maiden"] - -[[change]] -issue = "fix(#535) a given-slot numeral with a credential tail stays" -# 'Doe, Jane nee Smith V, PhD' moves {middle, maiden} at this -# baseline: after a family comma the given slot reads a lone numeral -# as a suffix only where the given part is the LAST comma part -# (#144), asked of the maiden clause too since this review -- a third -# comma part behind it withdraws the release, so the clause keeps 'V' -# where it would otherwise give it up. Verified against the parent -# tree d9d80492, which already reads middle 'V', maiden 'Smith' -# exactly as this baseline does -- an M2 violation #535's own arc had -# carried rather than caused, now fixed. Literal-anchored for the -# same SLOT reason fix(#533)'s rules give. -name_regex = "^Doe, Jane nee Smith V, PhD$" -fields = ["middle", "maiden"] - [[change]] issue = "fix(#397) the Catalan/Polish link joins two surnames" # 'i' is connective vocabulary since #397, and a connective counts as @@ -1918,8 +1794,8 @@ issue = "fix(#397) a link inside a maiden clause stays in the birth name" # maiden 'Puig -' -> 'Puig - i Soler' -- the same shape as the # neighbour's own reading with 'Mr.' dropped from it, unchanged by # the fix; only the corpus name is new. -name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler|Smith, John, PhD née Puig - i Soler)$" -fields = ["family", "maiden", "middle", "suffix"] +name_regex = "^(?:Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler|Smith, John, PhD née Puig - i Soler)$" +fields = ["family", "maiden", "middle"] orders = ["DEFAULT"] [[change]] @@ -2403,6 +2279,82 @@ issue = "fix(#597) halfwidth corner brackets enclose a nickname" name_regex = "^(?:山田 「タロー」 タロウ|山田「タロー」太郎|John 「Jack」 Smith)$" fields = ["given", "middle", "family", "nickname"] +[[change]] +# rules.md#M2: "In the given part after a family comma, in a part after a +# suffix comma, and behind a title, a suffix word or a connective, the +# marker is an ordinary word" (#601, 2026-10-04). The clause these names +# carried at this baseline is gone: `maiden` empties and its words read +# where the given part reads any word. What else differs is the given +# part's own reading of those words -- the spaced credential run +# (#436/#437), the link join (#397), P6's attachment of a trailing +# particle to the family (#405), the capitals lean (#289) and the +# trailing title chain (H5) -- each as the same text with the marker +# replaced by an ordinary word reads it. +issue = "fix(#601) a marker in the given part after a family comma is an ordinary word" +name_regex = "^(?:Berg, Jane van der nee Smith DO|Berg, abd née Jones|Berg, abdul nee Jones MA|Doe, Dr\\. nee Smith MA|Doe, J\\. nee MA ba|Doe, Jane Q\\. nee Smith MA|Doe, Jane nee Puig i Soler|Doe, Jane nee Smith|Doe, Jane nee Smith DO|Doe, Jane nee Smith Do|Doe, Jane nee Smith MA|Doe, Jane nee Smith MA Prof\\.|Doe, Jane nee Smith MA do|Doe, Jane nee Smith Ma|Doe, Jane nee Smith PhD MEng|Doe, Jane nee Smith do|Doe, Jane nee Smith ma)$" +fields = ["_ambiguities", "family", "given", "maiden", "middle", "suffix", "title"] + +[[change]] +# rules.md#M2: "In the given part after a family comma, in a part after a +# suffix comma, and behind a title, a suffix word or a connective, the +# marker is an ordinary word" (#601, 2026-10-04): a tail part is read +# wholly as suffixes, the marker and its words with it. +issue = "fix(#601) a marker in a part after a suffix comma is an ordinary word" +name_regex = "^(?:Smith, John, MD - née Jones Smith|Smith, John, PhD née Jones)$" +fields = ["maiden", "suffix"] + +[[change]] +# rules.md#M2: "In the given part after a family comma, in a part after a +# suffix comma, and behind a title, a suffix word or a connective, the +# marker is an ordinary word" (#601, 2026-10-04). Where a credential +# stands before the marker it also starts #602's run -- rules.md#S2: "A +# credential after the name core starts a run to the end of its part" -- +# which takes the marker and the words after it in. +issue = "fix(#601/#602) a marker behind a title, a suffix word or a connective is an ordinary word" +name_regex = "^(?:Dr\\. nee Smith PhD Prof\\.|Jane Doe Jr\\. nee Smith|Jane Doe Jr\\. nee Smith Ma|Jane Doe PhD nee Smith|Jane and née Jones)$" +fields = ["_ambiguities", "family", "given", "maiden", "suffix", "title"] + +[[change]] +# rules.md#M2: "It takes them up to the trailing run of post-nominals and +# titles that the end of the name reads as if the clause were not +# written" (#601, 2026-10-04), and "The run so found ends the name: its +# words are the suffixes and titles that reading makes them, and no join +# reaches into it". The old walk's release check kept these words in the +# clause where a join it modelled might take them; the run is bound now. +issue = "fix(#601) the clause ends at the clause-free name's trailing run, which the take consumes" +name_regex = "^(?:Jane Doe nee Smith DO DO|Jane van der Berg nee Prof\\. King\\. MA|Jane van der Berg nee Smith Prof\\.|John nee Prof\\. ba MA)$" +fields = ["_ambiguities", "maiden", "suffix", "title"] + +[[change]] +# rules.md#M2's take "always takes the first word after it, the marker +# having announced a name" (#601, 2026-10-04): a lone roman numeral is no +# suffix word, so it is the maiden name, where the old walk read the +# numeral fork from the marker and declined. Decided with #601 (decisions.md#M2). +issue = "fix(#601) the first word after the marker is the maiden name" +name_regex = "^Jane Smith née V$" +fields = ["_ambiguities", "family", "maiden", "middle", "suffix"] + +[[change]] +# rules.md#S2: "A credential after the name core starts a run to the end +# of its part" (#602, 2026-10-04): every later word of the part is a +# suffix except a title word, which reads as a title, and an absorbed +# name word is reported. +# 'Jane van der Berg née Jr Jones' (radar): the marker declines on +# the suffix word after it (rules.md#M2) and stays a word, so 'Jr' +# stands behind a name core and its run takes 'Jones' (2026-10-04). +issue = "fix(#602) a credential after the name core starts a run to the end of its part" +name_regex = "^(?:Eric H\\. Holder Jr\\. Attorney General|Jane van der Berg née Jr Jones|John Smith PhD Jones|John Smith PhD Prof\\. Ma|Smith, John PhD Jones)$" +fields = ["_ambiguities", "family", "middle", "suffix", "title"] + +[[change]] +# rules.md#M2 reads the clause-free name, whose credential starts #602's +# run -- rules.md#S2: "A credential after the name core starts a run to +# the end of its part" -- and the take consumes that run: maiden 'Smith', +# suffix 'PhD Smith' (2026-10-04). +issue = "fix(#601/#602) a credential in the clause starts a run the take consumes" +name_regex = "^Jane Doe nee Smith PhD Smith$" +fields = ["_ambiguities", "family", "middle", "suffix"] + # --------------------------------------------------------------- # #604: THE IRISH AND MALAY PATRONYMIC PARTICLES. # 'ó' and 'ní' join the never-given particles, 'ua', 'binti' and diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index 916c42a8..82746a49 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -670,7 +670,7 @@ issue = "fix(#533) a credential ending a maiden clause reads as a credential" # would claim the names whose WRITING declines the word -- which keep # their maiden reading and have the rule below -- as readily as the # ones it takes. _MUST_NOT_MATCH carries both directions. -name_regex = "^(?:Doe, J\\. nee MA ba|Doe, Jane Q\\. nee Smith MA|Doe, Jane nee Smith DO|Doe, Jane nee Smith MA|Doe, Jane nee Smith ma|JANE DOE NEE SMITH MA|JANE DOE NEE YO-YO MA|Jane Doe geb\\. Smith MA|Jane Doe nee King\\. ba|Jane Doe nee Smith MA|Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD|Jane Doe nee Smith V MA|Jane Doe nee Smith do|John née Jones Smith MA|Maria Kowalska z domu Nowak MA|jane doe nee smith ma)$" +name_regex = "^(?:JANE DOE NEE SMITH MA|JANE DOE NEE YO-YO MA|Jane Doe geb\\. Smith MA|Jane Doe nee King\\. ba|Jane Doe nee Smith MA|Jane Doe nee Smith MA JD|Jane Doe nee Smith MA PhD|Jane Doe nee Smith Ma JD|Jane Doe nee Smith V MA|Jane Doe nee Smith do|John née Jones Smith MA|Maria Kowalska z domu Nowak MA|jane doe nee smith ma)$" fields = ["_ambiguities", "maiden", "suffix"] [[change]] @@ -722,7 +722,7 @@ issue = "fix(#533) the maiden clause reports the credential it keeps" # half at rules.md#M2's new Accepted boundary: the credential in front # of the marker speaks for no member of the clause, the Title-case 'Ma' # is kept, and the clause reports it. -name_regex = "^(?:Berg, Jane van der nee Smith DO|Berg, abdul nee Jones MA|Doe nee Smith Prof\\. ba|Doe, Dr\\. nee Smith MA|Doe, Jane nee Smith Do|Doe, Jane nee Smith MA do|Doe, Jane nee Smith Ma|Doe, Jane nee Smith do|JOHN NEE JONES SMITH MA PHD|Jane Doe Jr\\. nee Smith Ma|Jane Doe nee MA|Jane Doe nee MA PhD|Jane Doe nee Smith DO DO|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma|John née Jones Smith Ma)$" +name_regex = "^(?:Doe nee Smith Prof\\. ba|JOHN NEE JONES SMITH MA PHD|Jane Doe nee MA|Jane Doe nee MA PhD|Jane Doe nee Smith Ma|Jane Doe nee Yo-Yo Ma|John née Jones Smith Ma)$" fields = ["_ambiguities"] [[change]] @@ -824,8 +824,11 @@ fields = ["_ambiguities", "maiden", "suffix", "title"] [[change]] issue = "fix(#535) a trailing title ends the maiden clause" -# rules.md#M2: "Where a trailing rule reads the words, a trailing -# title ends it too" -- the walk reads the end of the name through +# rules.md#M2: "It takes them up to the trailing run +# of post-nominals and titles that the end of the name reads as if +# the clause were not written" (#601; until #601 this cited the +# sentence it replaced, "a trailing title ends it too") -- the end of +# the name is read through # the trailing title chain -- rules.md#H5: "successive single words # that wear the abbreviation shape and are title vocabulary chain # into the title from the end" -- so the title leaves the clause and @@ -846,99 +849,9 @@ issue = "fix(#535) a trailing title ends the maiden clause" # split at THIS baseline -- #296 and #399 both predate it, so this # baseline already reads them as HEAD does apart from #535's own move, # unlike at 2.0.0/2.1.0.) -name_regex = "^(?:Doe nee Smith ba Prof\\.|Doe, Jane nee Smith MA Prof\\.|Jane Doe \\x28nee Smith Prof\\.\\x29|Jane Doe nee Prof\\. Dr\\.|Jane Doe nee Smith King\\.|Jane Doe nee Smith MA Prof\\.|Jane Doe nee Smith Ma Prof\\.|Jane Doe nee Smith Prof\\.|Jane Doe nee Smith V Prof\\.|Mary Smith née Jones Prof\\.)$" +name_regex = "^(?:Doe nee Smith ba Prof\\.|Jane Doe \\x28nee Smith Prof\\.\\x29|Jane Doe nee Prof\\. Dr\\.|Jane Doe nee Smith King\\.|Jane Doe nee Smith MA Prof\\.|Jane Doe nee Smith Ma Prof\\.|Jane Doe nee Smith Prof\\.|Jane Doe nee Smith V Prof\\.|Mary Smith née Jones Prof\\.)$" fields = ["title", "maiden", "suffix", "_ambiguities"] -[[change]] -issue = "fix(#535) a particle in front of a trailing title stays in the clause" -# rules.md#M2's Accepted pair on a released particle: the span 'DO -# Prof.' holds a particle with a title behind it, and P2's chain would -# run on over the title, so the particle's release is withdrawn and -# the clause keeps 'Smith DO' while the title stop still gives 'Prof.' -# up. The parent tree d9d80492 reads maiden 'Smith DO Prof.' as this -# baseline does, so #535 is the cause alone. Literal-anchored for the -# SLOT reason fix(#533)'s rules give. -name_regex = "^Jane Doe nee Smith DO Prof\\.$" -fields = ["title", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#533/#535) a particle behind a trailing title leaves the clause with it" -# The other half of that pair: written behind the title the particle -# is released with nothing behind it, so the title and 'DO' both leave -# the clause. Two causes together, checked against the parent tree -# d9d80492 (maiden 'Smith Prof.', suffix 'DO', reported): #533 already -# gives 'DO' up and reports the fork; #535 then reads the title in -# front of it off the clause, as for 'Jane Doe nee Smith Prof. MA'. -# Literal-anchored for the SLOT reason fix(#533)'s rules give. -name_regex = "^Jane Doe nee Smith Prof\\. DO$" -fields = ["title", "suffix", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#535) a particle inside the clause keeps the credential in front of a trailing title" -# rules.md#M2's Accepted pair on a particle INSIDE the clause: the -# particle's chain would take the credential behind it, so the clause -# keeps 'Smith do MA' and only the title behind the credential leaves. -# The parent tree d9d80492 reads maiden 'Smith do MA Prof.' with no -# report, as this baseline does, so #535 is the whole cause. -# Literal-anchored: one corpus name, the doc's own example. -name_regex = "^Jane Doe nee Smith do MA Prof\\.$" -fields = ["title", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#533/#535) a particle inside the clause lets a credential behind the title go" -# The other half of that pair: with the title in front of 'MA', the -# credential is released and the title leaves with it, so maiden -# 'Smith do', suffix 'MA', title 'Prof.'. Two causes together, checked -# against the parent tree d9d80492 (maiden 'Smith do Prof.', suffix -# 'MA', reported): #533 already gives 'MA' up and reports the fork; -# #535 then reads the title in front of it off the clause. -# Literal-anchored: one corpus name, the doc's own example. -name_regex = "^Jane Doe nee Smith do Prof\\. MA$" -fields = ["title", "suffix", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#397/#535) a link the clause stops at keeps a run that would not read off" -# rules.md#M2: a link that ends the clause gives up the words behind -# it only where the name left standing reads them as post-nominals or -# titles. With the title chained the link exception refuses 'i', but -# 'Jane Doe i DO Prof.' does not read 'i DO' off, so the clause keeps -# the link and the DO (reporting the kept credential) and only 'Prof.' -# leaves. Checked against the parent tree d9d80492, which read maiden -# 'Smith i DO Prof.': this baseline read 'i' as a plain suffix word -# that ended the clause, and #397 joins the link inside the clause; -# #535 then reads the title off. Literal-anchored: one corpus name, -# the doc's own example. -name_regex = "^Jane Doe nee Smith i DO Prof\\.$" -fields = ["title", "middle", "family", "maiden", "_ambiguities"] - -[[change]] -issue = "fix(#535) the numeral stop asks the join question" -# rules.md#M2 (#535): the numeral stop now asks the join question -# too -- released, the numeral was taken by the bound-given join -# after a family comma and read as part of the given name -# ('Berg, abdul nee Smith V' read given 'abdul V' at this -# baseline and at the other 2.2/2.3 baseline, before #535 ever ran; -# 2.0/2.1 read given 'abdul nee', suffix 'V' instead -- their own -# copy of this rule is fix(#411/#535) for that reason). -# Literal-anchored for the same SLOT reason fix(#533)'s rules give. -name_regex = "^(?:Berg, abdul nee Smith V)$" -fields = ["given", "maiden"] - -[[change]] -issue = "fix(#535) a given-slot numeral with a credential tail stays" -# 'Doe, Jane nee Smith V, PhD' moves {middle, maiden} at this -# baseline: after a family comma the given slot reads a lone numeral -# as a suffix only where the given part is the LAST comma part -# (#144), asked of the maiden clause too since this review -- a third -# comma part behind it withdraws the release, so the clause keeps 'V' -# where it would otherwise give it up. Verified against the parent -# tree d9d80492, which already reads middle 'V', maiden 'Smith' -# exactly as this baseline does -- an M2 violation #535's own arc had -# carried rather than caused, now fixed. Literal-anchored for the -# same SLOT reason fix(#533)'s rules give. -name_regex = "^Doe, Jane nee Smith V, PhD$" -fields = ["middle", "maiden"] - [[change]] issue = "fix(#397) the Catalan/Polish link joins two surnames" # 'i' is connective vocabulary since #397, and a connective counts as @@ -1168,8 +1081,8 @@ issue = "fix(#397) a link inside a maiden clause stays in the birth name" # maiden 'Puig -' -> 'Puig - i Soler' -- the same shape as the # neighbour's own reading with 'Mr.' dropped from it, unchanged by # the fix; only the corpus name is new. -name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler|Smith, John, PhD née Puig - i Soler)$" -fields = ["family", "maiden", "middle", "suffix"] +name_regex = "^(?:Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler|Smith, John, PhD née Puig - i Soler)$" +fields = ["family", "maiden", "middle"] orders = ["DEFAULT"] [[change]] @@ -1652,6 +1565,82 @@ issue = "fix(#597) halfwidth corner brackets enclose a nickname" name_regex = "^(?:山田 「タロー」 タロウ|山田「タロー」太郎|John 「Jack」 Smith)$" fields = ["given", "middle", "family", "nickname", "_ambiguities"] +[[change]] +# rules.md#M2: "In the given part after a family comma, in a part after a +# suffix comma, and behind a title, a suffix word or a connective, the +# marker is an ordinary word" (#601, 2026-10-04). The clause these names +# carried at this baseline is gone: `maiden` empties and its words read +# where the given part reads any word. What else differs is the given +# part's own reading of those words -- the spaced credential run +# (#436/#437), the link join (#397), P6's attachment of a trailing +# particle to the family (#405), the capitals lean (#289) and the +# trailing title chain (H5) -- each as the same text with the marker +# replaced by an ordinary word reads it. +issue = "fix(#601) a marker in the given part after a family comma is an ordinary word" +name_regex = "^(?:Berg, Jane van der nee Smith DO|Berg, abd née Jones|Berg, abdul nee Jones MA|Doe, Dr\\. nee Smith MA|Doe, J\\. nee MA ba|Doe, Jane Q\\. nee Smith MA|Doe, Jane nee Puig i Soler|Doe, Jane nee Smith|Doe, Jane nee Smith DO|Doe, Jane nee Smith Do|Doe, Jane nee Smith MA|Doe, Jane nee Smith MA Prof\\.|Doe, Jane nee Smith MA do|Doe, Jane nee Smith Ma|Doe, Jane nee Smith PhD MEng|Doe, Jane nee Smith do|Doe, Jane nee Smith ma)$" +fields = ["_ambiguities", "family", "given", "maiden", "middle", "suffix", "title"] + +[[change]] +# rules.md#M2: "In the given part after a family comma, in a part after a +# suffix comma, and behind a title, a suffix word or a connective, the +# marker is an ordinary word" (#601, 2026-10-04): a tail part is read +# wholly as suffixes, the marker and its words with it. +issue = "fix(#601) a marker in a part after a suffix comma is an ordinary word" +name_regex = "^(?:Smith, John, MD - née Jones Smith|Smith, John, PhD née Jones)$" +fields = ["maiden", "suffix"] + +[[change]] +# rules.md#M2: "In the given part after a family comma, in a part after a +# suffix comma, and behind a title, a suffix word or a connective, the +# marker is an ordinary word" (#601, 2026-10-04). Where a credential +# stands before the marker it also starts #602's run -- rules.md#S2: "A +# credential after the name core starts a run to the end of its part" -- +# which takes the marker and the words after it in. +issue = "fix(#601/#602) a marker behind a title, a suffix word or a connective is an ordinary word" +name_regex = "^(?:Dr\\. nee Smith PhD Prof\\.|Jane Doe Jr\\. nee Smith|Jane Doe Jr\\. nee Smith Ma|Jane Doe PhD nee Smith|Jane and née Jones)$" +fields = ["_ambiguities", "family", "given", "maiden", "suffix"] + +[[change]] +# rules.md#M2: "It takes them up to the trailing run of post-nominals and +# titles that the end of the name reads as if the clause were not +# written" (#601, 2026-10-04), and "The run so found ends the name: its +# words are the suffixes and titles that reading makes them, and no join +# reaches into it". The old walk's release check kept these words in the +# clause where a join it modelled might take them; the run is bound now. +issue = "fix(#601) the clause ends at the clause-free name's trailing run, which the take consumes" +name_regex = "^(?:Jane Doe nee Smith DO DO|Jane van der Berg nee Prof\\. King\\. MA|Jane van der Berg nee Smith Prof\\.|John nee Prof\\. ba MA)$" +fields = ["_ambiguities", "maiden", "suffix", "title"] + +[[change]] +# rules.md#M2's take "always takes the first word after it, the marker +# having announced a name" (#601, 2026-10-04): a lone roman numeral is no +# suffix word, so it is the maiden name, where the old walk read the +# numeral fork from the marker and declined. Decided with #601 (decisions.md#M2). +issue = "fix(#601) the first word after the marker is the maiden name" +name_regex = "^Jane Smith née V$" +fields = ["_ambiguities", "family", "maiden", "middle", "suffix"] + +[[change]] +# rules.md#S2: "A credential after the name core starts a run to the end +# of its part" (#602, 2026-10-04): every later word of the part is a +# suffix except a title word, which reads as a title, and an absorbed +# name word is reported. +# 'Jane van der Berg née Jr Jones' (radar): the marker declines on +# the suffix word after it (rules.md#M2) and stays a word, so 'Jr' +# stands behind a name core and its run takes 'Jones' (2026-10-04). +issue = "fix(#602) a credential after the name core starts a run to the end of its part" +name_regex = "^(?:Eric H\\. Holder Jr\\. Attorney General|Jane van der Berg née Jr Jones|John Smith PhD Jones|Smith, John PhD Jones)$" +fields = ["_ambiguities", "family", "middle", "suffix", "title"] + +[[change]] +# rules.md#M2 reads the clause-free name, whose credential starts #602's +# run -- rules.md#S2: "A credential after the name core starts a run to +# the end of its part" -- and the take consumes that run: maiden 'Smith', +# suffix 'PhD Smith' (2026-10-04). +issue = "fix(#601/#602) a credential in the clause starts a run the take consumes" +name_regex = "^Jane Doe nee Smith PhD Smith$" +fields = ["_ambiguities", "family", "middle", "suffix"] + # --------------------------------------------------------------- # #604: THE IRISH AND MALAY PATRONYMIC PARTICLES. # 'ó' and 'ní' join the never-given particles, 'ua', 'binti' and