Skip to content
Merged
4 changes: 2 additions & 2 deletions AGENTS.md

Large diffs are not rendered by default.

2 changes: 2 additions & 0 deletions docs/design/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,8 @@ rules.md states intended parsing behavior and is NORMATIVE; decisions.md is the

**Starting a design: pin the negative space first.** A spec's "X is not supported" is a claim about today's parser; hold it with a case row for X written BEFORE the implementation, so a row that passes unexpectedly kills the premise before any code exists. The 2026-07-30 comma-suffix spec did this for one item (`suffix_phrase_is_comma_scoped`) and not for its central premise, which is the half that cost a full plan (#291).

**Rethinking a rule that grew by patching.** Three signs together mean the next fix should be a rethink, not another patch: the rule works by scanning and stopping; each later fix added a stop, an exception or a check on one; and a guard exists whose job is to model what another stage would do. The third sign is the decisive one — a rule that must predict other rules' behavior to stay correct is defending an invariant it should be built from. #601 is the worked case. Five issues (#424, #533, #535, #538, #549) had each added a stop or a release check to M2, and its decisions.md section reached about 13,000 words. The rethink (mechanisms.md#READ-WITHOUT-THEN-BIND, #LICENSED-BY-NEIGHBOR) deleted the release machinery, cut the rule's statement to 883 words (counted 2026-10-04), and turned the invariant from a checked property into the construction. The rethink was drafted in rule shape and prototyped in seven variants behind one switch before anything landed (#601's comments); the variants that failed did so on measured counts, not on review opinion. Rule length alone is not the signal: S2 and R4 have each gathered 35 dated decisions.md entries, and R4's are mostly vocabulary and rendering cases rather than added stops (counted 2026-10-04).

**Landing a design.** The gitignored spec (docs/superpowers/specs/) is the working medium; it dies with the branch, and the docs are the record. Before a design PR merges, walk its spec (including amendments) and distill the durable residue: decisions made or reversed → decisions.md entries; proposals rejected with evidence → Declined:; vocabulary that must stay out → Excluded:; behavior the design settled → rules.md (with a deviates: marker if unshipped); reusable patterns → mechanisms.md; options weighed → a weighing entry. Then check the spec cites nothing the docs don't now carry — a spec section with no committed home when the PR merges is lost, not deferred (a 2026-08-16 sweep of eight weeks of specs recovered nine such items). The same-PR amendment rule above covers code-driven changes; this covers the design-driven ones.

**A count in a dated entry is evidence, not a live fact.** decisions.md entries are snapshots by convention, so measurements belong in them — but a reader wanting TODAY's number must not have to trust the snapshot's date. Where an entry quotes something that drifts (vocabulary sizes, corpus counts, set compositions), give the one-liner that recomputes it, and phrase the argument so it survives the digits moving — "the two shares differ by orders of magnitude" outlives "58% vs 0.65%". A count that carries no argument is better deleted than dated: "over every name in the corpus, no prefilter" says what "over all 782 names" says, and cannot go stale. Do NOT reach for a test asserting the count — that is the constant-content pattern, and it fails on every legitimate vocabulary addition. #326 is the cautionary case: it quoted a vocabulary composition, carried a date, and was stale in five days.
Expand Down
6 changes: 6 additions & 0 deletions docs/design/decisions.md

Large diffs are not rendered by default.

8 changes: 8 additions & 0 deletions docs/design/mechanisms.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,14 @@ Problem shape. Where should a new "recognize X" behavior live? Contract statemen
"Esq. Smith" reads title, not suffix.
How it works. The two layers compose without ordering bugs because the positional layer never overrides a vocabulary claim (rule O4 is the positional layer's contract). The exception is still ONE, re-checked 2026-09-08 against the trailing title run (rule H5), which reads title vocabulary in the trailing slot and is NOT a second one: it runs after the suffix peel rather than before it, so a post-nominal keeps its claim — `John Smith Esq.` reads suffix and `John Smith Prof. Jr.` reads suffix `Jr.`, both measured. That rule is also the worked answer to the Reach-for-it line below: it wanted a word's identity and its position at once and was split by SLOT, the front asking about shape (H2) and the back about vocabulary (H5), two questions with two criteria rather than one rule fighting both layers. Lives in. _classify/_group (vocabulary side), _assign (positional side). Reach for it when. A proposed rule wants a word's identity AND its position at once — split it, or it will fight both layers.

## READ-WITHOUT-THEN-BIND — read the name without the construct, then fix that reading

Problem shape. An optional construct sits inside a name — a maiden clause, a trailing title — and the requirement is that writing it changes nothing else about the name. Checked after the fact, that invariant has to be defended at every place the construct can end, and each defense has to model what later stages would do to the words it gives up. Contract statement. Read the name as if the construct were not written, using the rule that reads that position, not the whole parse. The construct takes only what that reading leaves. Where later stages could still move what the reading decided, give those words their roles at the take so no later join can reach them. The invariant then holds by construction and is not a property each stop must defend. How it works. M2 before #601 scanned forward, stopped at suffix-looking words, and kept a release check asking at each stop whether the name left standing would read what was given up as post-nominals. That check modelled the P2 chain, the P5 bound join and P6's attachment, and #548 was it failing at the one stop that never asked. #601 replaced the stops and the check with the clause-free reading plus the bind: a net 689 lines out of `_group.py` (+210/−899 in the #601 commit), and #535's grids went from 15,586 and 8,713 invariant violations to 0 (prototype measurement, 2026-10-03, #601's comments). The first half already appears twice. H5's trailing title is TRANSPARENT: "what stands once the chain is taken reads exactly as it would read written without the title, plus the title". P3 asks its word count and its one-case question of the name's OWN words, so a clause beside the name changes neither. M2 is the first to bind as well, because its construct sits in front of the run it reads and the joins come later. Two limits, both measured on #601. Read with the TRAILING rule, not the full parse: a head the full parse reads differently (H1's lone title, a P5 join) would otherwise move where the construct ends. And binding means the run can read differently from the same words with no construct (`JANE Q. DOE GEB. LE DO DO` gives suffix `DO DO`, while the full parse of `JANE Q. DOE DO DO` chains them into the family), which rules.md#M2 accepts in so many words. Lives in. nameparser/_pipeline/_group.py (`_maiden_take`: the view, `tail_reading` over it, and the role-tagged pieces `group()` applies); nameparser/_pipeline/_pieces.py (`tail_reading`, H5's chain); the P3 own-word test in `_group.py`. Reach for it when. A rule is growing stops, exceptions or a release check whose job is to keep one construct from disturbing the rest of the name, and above all when that check has started to model another stage. Ask what the name reads as without the construct, and make that the rule.

## LICENSED-BY-NEIGHBOR — a structural word is structural only beside the right neighbor

Problem shape. A word in a structural vocabulary — a marker announcing a maiden name, a connective joining surnames — turns up where what it would announce makes no sense: behind a title alone, behind a credential, at the end of a name. Designing a reading for each such shape is endless, and the shapes are mostly junk. Contract statement. Make the vocabulary claim conditional on the neighbor: the word is structural only where the word beside it has the role the structure needs. Anywhere else it is an ordinary word and goes to the positional layer (TWO-LAYER-ASSIGN), with no reading designed for the junk. How it works. M2 (#601): a marker counts only behind a name word of the part holding the family name, not behind the leading title run, an unambiguous suffix word or a connective. So `Dr. nee Smith` gives first `nee`, last `Smith`, and `Jane Doe PhD nee Smith` gives suffix `PhD nee Smith`, neither designed. P3 is the same shape for the connectives that are also generational vocabulary: `i` joins only where "a name word stands on each side of it". The condition must be checked against real names, not just against the vocabulary. #601's first cut also refused a marker behind a particle and broke `Mai Le née Nguyen` and `Anh Do née Tran` — Le, Do and Van are particle vocabulary and among the commonest Vietnamese surnames. Lives in. nameparser/_pipeline/_group.py (`_maiden_take`'s head rule; `_between_name_words` for P3). Reach for it when. A structural word is producing bad readings in positions its structure does not fit. Make its claim depend on the neighbor before adding a reading or a stop for each position.

## STATE-OFFSET-CHANNELS — early facts ride the state

Problem shape. A fact known during tokenization matters to a much later stage. Contract statement. A pre-token fact is recorded as offsets on the ParseState (comma_offsets, interpunct_offsets) and consulted later by position — or by presence alone where the fact is name-level, as both interpunct consumers do, or as a single boolean for the whole name, as `ParseState.one_case` is (#289/#516: the written-case fact three later sites read — the trailing suffix slot, the post-comma given slot, and the tail-segment reading — none of which may disagree about it) — rather than re-derived from text. How it works. The offsets survive every intermediate stage untouched; #298's transcription marker rides this channel from tokenize to order resolution (rules T3/W4). `one_case` differs from the offset channels in WHO writes it: it is computed once by whichever of segment and classify needs it first (classify always, segment only where a comma form could turn it on), recorded as `None` for "not asked yet", and never recomputed once set (decisions.md#S2). Lives in. nameparser/_pipeline/_state.py, produced in _tokenize (the offsets) or in segment/classify (`one_case`). Reach for it when. You are about to re-scan the original string in a late stage to rediscover something tokenize already knew.
Expand Down
Loading
Loading