Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -393,7 +393,7 @@ Add a dedicated `copy.deepcopy()` round-trip test for it too (see `test_regexes_

**`is_suffix()`'s period-stripping is asymmetric** — it does `lc(piece).replace('.', '')` (strips *all* periods) before checking `suffix_acronyms`, but only `lc(piece)` (leading/trailing only) before checking `suffix_not_acronyms`. Code that reimplements this check instead of calling `is_suffix()` must mirror both branches or it will misclassify acronym suffixes with internal-only periods (e.g. `"M.D"` with no trailing dot) — bit `parse_nicknames()`'s `handle_match()` in PR #189.

**`nickname_delimiters`/`maiden_delimiters` built-ins are string sentinels, not compiled patterns** — each default key (`quoted_word`, `double_quotes`, `parenthesis`, and the rest) stores its own *name* as a plain `str`, and `Constants._snapshot()` maps that key to the `(open, close)` pair the 2.0 `Policy` carries (`_SENTINEL_PAIRS` in `_config_shim.py`). Routing a built-in between buckets is still a `pop()` + assign (`maiden_delimiters['parenthesis'] = nickname_delimiters.pop('parenthesis')`) to carry the sentinel over — measured, that flips `"Jane (Jones) Smith"` from `nickname='Jones'` to `maiden='Jones'`. **2.0 note:** the two things this bullet used to justify the sentinel with are both gone, and only the v1 history above still holds. `CONSTANTS.regexes.parenthesis = ...` now raises `TypeError` (2.0 configures parsing through named `Policy` flags), so the sentinel no longer buys post-construction tracking of a `regexes` override; and a caller's *custom* key is refused by `_DelimiterManager` — its message points at `Policy(nickname_delimiters=...)` — rather than accepted as a compiled pattern. No `isinstance(raw_pattern, re.Pattern)` branch survives, and `HumanName.parse_nicknames` is gone as a method — but do NOT sweep the name itself as dead: it is still live in `_facade.py`'s `_V1_HOOKS`, where a v1 subclass that overrides it earns the once-per-subclass hook warning (#280). Don't count the default keys in prose either: this bullet said "three" until #273 added eight typographic pairs in one commit — five Western (`smart_double_quotes`, `low_high_quotes`, `right_double_quotes`, `guillemets`, `reversed_guillemets`) and three CJK — taking it to eleven. (#22)
**`nickname_delimiters`/`maiden_delimiters` built-ins are string sentinels, not compiled patterns** — each default key (`quoted_word`, `double_quotes`, `parenthesis`, and the rest) stores its own *name* as a plain `str`, and `Constants._snapshot()` maps that key to the `(open, close)` pair the 2.0 `Policy` carries (`_SENTINEL_PAIRS` in `_config_shim.py`). Routing a built-in between buckets is still a `pop()` + assign (`maiden_delimiters['parenthesis'] = nickname_delimiters.pop('parenthesis')`) to carry the sentinel over — measured, that flips `"Jane (Jones) Smith"` from `nickname='Jones'` to `maiden='Jones'`. **2.0 note:** the two things this bullet used to justify the sentinel with are both gone, and only the v1 history above still holds. `CONSTANTS.regexes.parenthesis = ...` now raises `TypeError` (2.0 configures parsing through named `Policy` flags), so the sentinel no longer buys post-construction tracking of a `regexes` override; and a caller's *custom* key is refused by `_DelimiterManager` — its message points at `Policy(nickname_delimiters=...)` — rather than accepted as a compiled pattern. No `isinstance(raw_pattern, re.Pattern)` branch survives, and `HumanName.parse_nicknames` is gone as a method — but do NOT sweep the name itself as dead: it is still live in `_facade.py`'s `_V1_HOOKS`, where a v1 subclass that overrides it earns the once-per-subclass hook warning (#280). Don't count the default keys in prose either: this bullet said "three" until #273 added eight typographic pairs in one commit — five Western (`smart_double_quotes`, `low_high_quotes`, `right_double_quotes`, `guillemets`, `reversed_guillemets`) and three CJK — taking it to eleven, and #597 added a twelfth, `halfwidth_corner_brackets`, on 2026-10-03. (#22)

**`TupleManager.__setattr__`/`__delattr__` guard dunder names too, not just `__getattr__`** — constructing a subscripted generic, e.g. `TupleManager[re.Pattern[str] | str]({...})` (needed so mypy sees the right value type instead of inferring one from the dict literal), makes `typing`'s `GenericAlias.__call__` set `__orig_class__` on the new instance right after `__init__` returns. Before this guard existed, `__setattr__` was a bare `dict.__setitem__` alias, so that assignment silently inserted a bogus `'__orig_class__'` entry into the dict itself, corrupting `.values()`/iteration for *every* `TupleManager`/`RegexTupleManager` instance, not just the one being constructed — this bit `nickname_delimiters`'s construction (#22) before the guard was added. Same fix shape as the `__getattr__` dunder guard above: fall back to `object.__setattr__`/`object.__delattr__` for dunder names, dict-backed storage for everything else.

Expand Down
4 changes: 4 additions & 0 deletions docs/design/decisions.md
Original file line number Diff line number Diff line change
Expand Up @@ -232,6 +232,10 @@ the fullwidth-colon marker (旧姓:佐藤 arrives as one word; the head-peel q
- 2026-09-27 (Derek), #544 — S2'S COMPANY CLAUSE DOES NOT REACH ACROSS A CLAUSE, recorded as an Accepted boundary rather than repaired. `Jane Doe Jr. nee Smith Ma` keeps maiden 'Smith Ma' and reports it, the clause's own words standing between 'Jr.' and 'Ma', where the clause-less `Jane Doe Jr. Ma` reads suffix 'Jr. Ma'. tests/v2/test_properties.py's clause-agreement walk pins the pairs that differ for exactly this reason as its `anchored_head` class — 810 of its 20,412 pairs, recorded 2026-09-27, 0 with the anchor off. Inside the clause the company reads as it does anywhere: `Jane Doe nee Smith PhD MEng` ends the clause at 'PhD' and reads suffix 'PhD MEng', maiden 'Smith'.


### N1 — the default delimiter pairs

- 2026-10-03 #597 — THE HALFWIDTH CORNER BRACKETS ARE A DEFAULT NICKNAME PAIR. `「」` (U+FF62/U+FF63) join `DEFAULT_NICKNAME_DELIMITERS` beside `「」`, and the v1 facade gains the matching `halfwidth_corner_brackets` key, so `HumanName` reads them as `parse()` does. They are the same punctuation as `「」` in JIS X 0201's narrower encoding, the legacy data halfwidth katakana comes from (#594), where they were the only corner brackets available, and they appear in no non-CJK text. Before this every release kept the clause in the name, brackets and all — a middle where it stood spaced (`山田 「タロー」 タロウ`), part of one given word where it was glued (`山田「タロー」太郎`) — and since #594 the bracket also cost the spaced name its order: the halfwidth kana now classify, but the unclassified bracket kept `山田 「タロー」 タロウ` from being wholly East Asian script, so it read given 山田 (in released versions the unclassified kana did that on their own). A v1 `Constants` restored from a pickle keeps the keys it was saved with and so does not gain the pair, the same policy #273's pairs followed: restored configuration is never widened behind the caller's back. A pair is not script-gated, so `John 「Jack」 Smith` reads nickname Jack as `John 「Jack」 Smith` does. Nothing else needs a halfwidth twin: Unicode has no halfwidth form of `『』` or of the fullwidth parentheses. MEASURED 2026-10-03: no corpus line held a halfwidth bracket before this change, so the three `fix(#597)` case rows are the whole population, and they move at every baseline. RECOMPUTE: parse `山田 「タロー」 タロウ`, `山田「タロー」太郎` and `John 「Jack」 Smith` here and on each baseline from PyPI and read `nickname` beside the name fields.

### N3 — the lone-word nickname rule

- 2026-07 (v2 core, PR #288; recorded plan deviation #2 of the core plan) — v1's rule counted pieces before grouping; the v2 port fires only when the nickname accompanies exactly ONE piece in total — a title counts against it, so "'Smitty' Dr. Jones" reaches H1 with a title and one name word left standing — through 2.1 that meant given="Jones" with the family empty, and since #410 (2026-08-25) H1 names the family, so it reads family="Jones". The count is unchanged; what moved is what happens after it declines. The rule lives in assignment because that is where the piece count is settled. (An earlier wording here said "one non-title piece", predicting the opposite output; the coherence review measured the truth.)
Expand Down
3 changes: 2 additions & 1 deletion docs/design/rules.md
Original file line number Diff line number Diff line change
Expand Up @@ -1306,7 +1306,8 @@ N1. Rationale: a quoted or bracketed clause beside a name is an
"Andrew (Andy) Perkins" → nickname="Andy"
"Jean 'JD' Smith" → nickname="JD"
"Anna () Smith" → nickname="" · boundary
implemented: nameparser/_pipeline/_extract.py
"John 「Jack」 Smith" → nickname="Jack"
history: decisions.md#N1 · implemented: nameparser/_pipeline/_extract.py

N2. Rationale: only a mark standing at word boundaries is quoting;
anywhere else it is part of the word.
Expand Down
5 changes: 3 additions & 2 deletions docs/modules.rst
Original file line number Diff line number Diff line change
Expand Up @@ -161,13 +161,14 @@ Delimiter defaults
^^^^^^^^^^^^^^^^^^

.. py:data:: nameparser.DEFAULT_NICKNAME_DELIMITERS
:value: frozenset({("'", "'"), ('"', '"'), ("(", ")"), ("“", "”"), ("„", "“"), ("”", "”"), ("«", "»"), ("»", "«"), ("「", "」"), ("『", "』"), ("(", ")")})
:value: frozenset({("'", "'"), ('"', '"'), ("(", ")"), ("“", "”"), ("„", "“"), ("”", "”"), ("«", "»"), ("»", "«"), ("「", "」"), ("『", "』"), ("「", "」"), ("(", ")")})

The default :attr:`~nameparser.Policy.nickname_delimiters` set:
straight quotes and parentheses plus the typographic conventions —
smart quotes, German/Polish low-high quotes, Swedish right-right
quotes, guillemets in both directions, CJK corner brackets, and
fullwidth parentheses (#273). Curly *single* quotes are deliberately
fullwidth parentheses (#273), plus the halfwidth corner brackets
``「」`` that legacy JIS X 0201 data carries (#597). Curly *single* quotes are deliberately
absent: U+2019 is the typographic apostrophe ("O’Connor"). Build on
the constant for additive customizations, e.g.
``nickname_delimiters=DEFAULT_NICKNAME_DELIMITERS | {("⦅", "⦆")}``;
Expand Down
2 changes: 2 additions & 0 deletions docs/release_log.rst
Original file line number Diff line number Diff line change
Expand Up @@ -70,6 +70,8 @@ Release Log

- **Fix a katakana name typed with a separate voicing mark being read family-first.** ``HumanName("ア゙イ タロウ")`` gives first ``ア゙イ``, last ``タロウ``, as ``アイ タロウ`` does, where 2.1 through 2.3 gave last ``ア゙イ``, first ``タロウ``. A dakuten or handakuten after a kana that has no precomposed voiced form (``ア゙``, ``ン゙``), or the spacing ``゛`` and ``゜``, was read as hiragana, which made the name Japanese by script and turned it around; the mark now belongs to the kana it follows. It flipped the rest of the name too: ``マイケル ア゙イ`` gives first ``マイケル`` where it gave last ``マイケル``. Hiragana names and kanji-and-kana names with such a mark read as before. See the ``W4`` entry of ``docs/design/decisions.md`` (closes #596)

- **Fix halfwidth corner brackets not being read as a nickname.** ``HumanName("山田 「タロー」 タロウ")`` gives nickname ``タロー``, last ``山田``, first ``タロウ``, where every release gave middle ``「タロー」`` and first ``山田`` (the order moves with the halfwidth katakana change above; the bracket would otherwise have blocked it). The halfwidth ``「」`` are the corner brackets of legacy JIS X 0201 data, the same punctuation as ``「」``, and are now a default nickname pair in both APIs: ``DEFAULT_NICKNAME_DELIMITERS`` gains ``("「", "」")`` and the 1.x ``nickname_delimiters`` gains the key ``halfwidth_corner_brackets``. They are not limited to Japanese text: ``John 「Jack」 Smith`` gives nickname ``Jack`` where it gave middle ``「Jack」``. A ``Constants`` restored from a pickle keeps the keys it was saved with, as it did when 2.0 added the other typographic pairs. See the ``N1`` entry of ``docs/design/decisions.md`` (closes #597)

**Additions**

- **Add Lexicon.conjunctions_ambiguous, the one-letter connectives that read as initials.** A subset of ``conjunctions`` holding ``e`` and ``i`` by default; it is the knob for the change above rather than a switch. Portuguese data, where ``e`` links surnames the way ``y`` does in Spanish, takes it out: ``Lexicon.default().remove(conjunctions_ambiguous={"e"})`` restores the joining reading. Dutch data, where a bare single letter is an initial and never a connective, adds the other one: ``Lexicon.default().add(conjunctions_ambiguous={"y"})``. A v1 ``Constants`` has no manager of its own for it -- deleting the word from ``conjunctions`` is what turns the marking off, the same rule the glued-honorific tails follow. See ``docs/customize.rst`` (#383, #479)
Expand Down
8 changes: 5 additions & 3 deletions nameparser/_config_shim.py
Original file line number Diff line number Diff line change
Expand Up @@ -435,8 +435,9 @@ def __setstate__(self, state: dict[str, object]) -> None:

#: The named delimiter buckets, translated to the ``Policy``
#: (open, close) pairs they stand for. The first three are
#: v1's; the rest are the #273 typographic conventions, named so the
#: v1 keyed idioms (pop/move/del) work on them like the originals.
#: v1's; the rest are the #273 typographic conventions plus #597's
#: halfwidth corner brackets, named so the v1 keyed idioms
#: (pop/move/del) work on them like the originals.
#: Keep in sync with DEFAULT_NICKNAME_DELIMITERS in _policy.py (pinned
#: by the default-Constants equality test).
_SENTINEL_PAIRS = {
Expand All @@ -450,6 +451,7 @@ def __setstate__(self, state: dict[str, object]) -> None:
"reversed_guillemets": ("»", "«"),
"corner_brackets": ("「", "」"),
"white_corner_brackets": ("『", "』"),
"halfwidth_corner_brackets": ("「", "」"),
"fullwidth_parenthesis": ("(", ")"),
}

Expand All @@ -476,7 +478,7 @@ class -- via ``TupleManager.__reduce__``'s ``(type(self), (), state)``
class _DelimiterManager(TupleManager):
"""v1 ``nickname_delimiters``/``maiden_delimiters`` bucket. In 2.0
only the named sentinels in ``_DELIMITER_SENTINELS`` exist (the v1
trio plus the #273 typographic pairs) -- assigning any
trio plus the typographic pairs, #273 and #597) -- assigning any
other key raises so a caller reaches for a custom-delimiter Policy
kwarg instead of a dict entry that silently does nothing. ``pop()``/
``__setitem__``/``__delitem__`` stay open (inherited) for the
Expand Down
8 changes: 5 additions & 3 deletions nameparser/_policy.py
Original file line number Diff line number Diff line change
Expand Up @@ -345,9 +345,10 @@ def _order_repr(value: tuple[Role, ...]) -> str:
#: rebuilt literal the user had to go discover. The v1 trio (straight
#: quotes + parentheses) plus the typographic conventions (#273):
#: smart quotes, low-high and right-right quotes, guillemets both
#: directions, CJK corner brackets, fullwidth parentheses. Curly
#: SINGLE quotes are deliberately absent: U+2019 is the typographic
#: apostrophe ("O’Connor").
#: directions, CJK corner brackets (full-width and the halfwidth
#: 「」 legacy JIS X 0201 data carries, #597), fullwidth parentheses.
#: Curly SINGLE quotes are deliberately absent: U+2019 is the
#: typographic apostrophe ("O’Connor").
DEFAULT_NICKNAME_DELIMITERS = frozenset({
("'", "'"), ('"', '"'), ("(", ")"), # v1 trio
("“", "”"), # smart quotes (en, zh)
Expand All @@ -356,6 +357,7 @@ def _order_repr(value: tuple[Role, ...]) -> str:
("«", "»"), # guillemets (fr, ru, it, el)
("»", "«"), # reversed guillemets (de alt)
("「", "」"), ("『", "』"), # CJK corner brackets (ja)
("「", "」"), # halfwidth corner brackets
("(", ")"), # fullwidth parentheses (CJK)
})

Expand Down
24 changes: 24 additions & 0 deletions tests/v2/cases.py
Original file line number Diff line number Diff line change
Expand Up @@ -5988,6 +5988,30 @@ def _check_cjk_shape_purity(self) -> None:
Case("cjk_white_corner_bracket_nickname", '田中『ハナ』花子',
{"family": "田中", "given": "花子", "nickname": "ハナ"},
classification="feat(#273) + fix(#271)"),
Case("cjk_halfwidth_corner_bracket_nickname", '山田 「タロー」 タロウ',
{"family": "山田", "given": "タロウ", "nickname": "タロー"},
classification="fix(#597)",
notes="the halfwidth corner brackets U+FF62/U+FF63 are the "
"same punctuation as 「」 in JIS X 0201's narrower "
"encoding, the data halfwidth katakana comes from (#594). "
"2.3.0 and every earlier release left the clause in the "
"name as a middle 「タロー」 and read given 山田, the "
"halfwidth kana being unclassified until #594. Since "
"#594 the bracket alone kept the name from being wholly "
"East Asian script; extracted, the rest is kanji plus "
"katakana and takes the kana license"),
Case("cjk_halfwidth_corner_bracket_unspaced_nickname", '山田「タロー」太郎',
{"family": "山田", "given": "太郎", "nickname": "タロー"},
classification="fix(#597)",
notes="the halfwidth twin of cjk_corner_bracket_nickname: the "
"extracted clause is a token boundary, so the unspaced "
"remainder divides there and both pieces are Han"),
Case("latin_halfwidth_corner_bracket_nickname", 'John 「Jack」 Smith',
{"given": "John", "family": "Smith", "nickname": "Jack"},
classification="fix(#597)",
notes="a delimiter pair is not script-gated: the halfwidth "
"brackets lift a Latin nickname exactly as 「」 does in "
"'John 「Jack」 Smith'. 2.3.0 read middle 「Jack」"),
Case("fullwidth_paren_nickname", 'John (Jack) Kennedy',
{"given": "John", "family": "Kennedy", "nickname": "Jack"},
classification="feat(#273)"),
Expand Down
Loading
Loading