diff --git a/docs/design/decisions.md b/docs/design/decisions.md index c41f5af4..8c8b0fd0 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -219,6 +219,8 @@ the fullwidth-colon marker (旧姓:佐藤 arrives as one word; the head-peel q INV6's off-switch exemption is NOT widened, and the reason is the one AGENTS.md gives for not widening an exemption a grid cannot reach: measured 2026-09-20, 1,680 rows of that grid carry both a class letter and a marker and 0 of them put the letter after the marker, its generator placing the connective among the name's own words. A clause link has its own grid instead. TWO SECOND-ORDER MOVEMENTS, both recorded rather than repaired. The two one-case spellings part company, which is P3's own Accepted clause reaching a taken marker rather than a declined one: classify leaves a letter after a marker to the mixed-case rule, so `JANE DOE NEE PUIG I SOLER` never moved (the capital is an initial, and the initial veto kept the walk going all along) while `jane doe nee puig i soler` goes from family 'doe i soler' to maiden 'puig i soler'. And a report is GAINED where a longer clause ends on an ambiguous credential: `Jane Doe nee Puig i Ma` keeps maiden 'Puig i Ma' and says `suffix-or-name` where 0fbcaa0b read family 'Doe i Ma' in silence — the clause's own emitter, reaching a word the truncation had put out of its reach. A THIRD, found by the second review and the mirror of that one: a report is LOST where the clause takes a class member it used to end at. `Jane Doe nee MA i Soler` reads maiden 'MA i Soler' in silence where the parent 46651750 read maiden 'MA' with a `suffix-or-name` on it. That FOLLOWS from the link no longer ending the clause and is not a second decision: the report is the walk's own, raised on the LAST word the clause kept where a trailing rule was asked about it, and with the link joining, 'MA' has a name word behind it and is no longer that word — which is exactly what `Jane Doe nee MA Smith` has always done with the same acronym. The `y` twin is the oracle and it AGREES, in fields and in reports: `Jane Doe nee MA y Soler` reads maiden 'MA y Soler' in silence at the parent and here. Measured 2026-09-20 over the review grid, 24 names lose the report this way and all 24 agree with their `y` twin on every field and every report. +- 2026-09-22 #397 follow-up — A DELIMITER CORE PAST THE CLAUSE'S FIRST WORD PASSES FOR THE NAME WORD BESIDE A LINK, RECORDED AS A DEVIATION RATHER THAN REPAIRED. The link exception above wants "a connective standing between two name words of the clause", and a separator the caller declared through `Policy.extra_suffix_delimiters` is structure, not a name word — so a link with one beside it joins nothing and should end the clause like any other suffix word. Between the marker and the clause's first word that already holds, the bound refusing the core before either piece test is asked. PAST that first word it does not: the core is an ordinary index to the run walk, which steps over connectives and nothing else, so it stands in for the name word on the link's left and the clause runs on past a title it would otherwise stop at. MEASURED 2026-09-22 under `extra_suffix_delimiters=(" - ",)`: `Smith, John, PhD née Puig Mr. - i Soler` reads maiden 'Puig Mr. i Soler' where its separator-less twin `Smith, John, PhD née Puig Mr. i Soler` stops at 'Puig Mr.'. THE POPULATION is the branch's own sweep, recorded with the code it describes (2026-09-21, corpus ∪ cases.py ∪ the property grids ∪ a 50,925-name generated set with cores, under thirteen core-bearing policies): the predicate is asked about a core in 51,072 of 900,023 calls, the answer differs from a core-skipping reading in 8,094 parses over 1,278 texts, and 1,824 of those move `maiden` on 288 texts — none of the 288 reachable at the default policy, `extra_suffix_delimiters` being empty there. NOT REPAIRED HERE: the fix threads the core set through three call sites into the run walk and moves the parent's reading as well, which makes it its own change rather than a rider on a review round. Open: #538. PINNED TWICE MEANWHILE. rules.md#M2 carries the shape as a `deviates: #538` example under an `extra_suffix_delimiters-dash` annotation — the first entry `tests/v2/rules_doc.py`'s registry has had for that field, named after the Policy field and carrying the delimiter in the suffix because the field's value is a set rather than a flag. And `tests/v2/pipeline/test_group.py::test_a_core_beside_a_link_wrongly_passes_for_a_word_until_538` holds the pair at the piece level, named so nobody reads it as the contract. ONE COST OF THE DOC EXAMPLE, worth knowing before the repair lands: its string enters `corpus_rules.jsonl`, where the differential gate parses it with the DEFAULT facade — no delimiter declared, so the dash is an ordinary name word and #538's reading is off the path entirely. It moves there for the 2026-09-20 link fix instead, and is classified as that at all five baselines: added to the `fix(#397) a link inside a maiden clause stays in the birth name` alternation at the four 2.x ledgers (suffix 'PhD i Soler' → 'PhD', maiden 'Puig Mr. -' → 'Puig Mr. - i Soler', identical at each), and given its own rule at 1.4.0, where v1 had no maiden markers and read the whole suffix-comma tail as one suffix. + ### N3 — the lone-word nickname rule @@ -378,6 +380,11 @@ The reconciled v1-style banks (`tests/test_*.py`) carried eight `@pytest.mark.xf ACCEPTED, AND THE HONEST STATEMENT OF THE ONE-CASE HALF, written out because this branch's own first commit message overclaimed that one-case names keep their reading. `i` is in the MARKED subset, so in a name written wholly in one case it reads as an INITIAL and reports `conjunction-or-initial`, exactly as `e` has since #383/#479. For "JOSEP CAROD I ROVIRA" and "josep carod i rovira" the fields are indeed unchanged and only the report is new. But a one-case name reports wherever a bare `i`/`I` stands among the NAME'S OWN WORDS, which is a great many more names than the Catalan ones — "JOHN I SMITH" and "john i smith" one report each, "JOHN SMITH I" and "john smith i" two, "HENRY I" and "henry i" three. A letter inside a maiden clause is read by the clause's rules and stays silent, exactly as `e` does ("JANE DOE NEE I JONES" reports nothing), which is the own-words scope the 2026-09-13 #383/#479 entry above defines. And in ALL-LOWER names the initial reading MOVES FIELDS wherever the parent read the lower-case letter as the generation. Each such name now reads as its ALL-CAPS twin already did: "rovira, i" gives given "i" where it gave suffix "i" ("ROVIRA, I" already gave given "I"); "john smith i jr" gives middle "smith" with family "i" where it gave family "smith" with suffix "i jr" ("JOHN SMITH I JR" already did); "maier, amy i, jr." gives middle "i" with suffix "jr." where it gave suffix "i, jr." (the corpus name "Maier, Amy I, Jr." reads middle "I" and does not move); and "josep de carod i rovira" gives family "de carod i rovira" where it gave middle "de carod i" with family "rovira". None of those four is a corpus name; all are measured 2026-09-20 on this tree and against the parent's. NO CHANGE TO THE EXCLUDED BLOCK, recorded as a decision rather than left implicit (Derek, 2026-09-20). It now lists "y" against TWO marked letters rather than one, and that is still right: the argument that put "y" outside is about "y" — the commonest Hispanic compound and this library's oldest fixture — while "i" matches "e" exactly, a bare I initial being as common as a bare E. +- 2026-09-21 #397 follow-up — AMENDS the "AND THE CONDITION IS ASKED ONCE PER RUN" paragraph above, whose frame figures stand as measured at `6048eb5d`. BOTH SIDES ARE NOW ONE CALL. A connective is placed to join only where a name word stands on EACH side of it, so neither caller ever wanted one side's answer: `_name_word_beside(k, step, ...)` was asked twice and the two answers ANDed at the call site, in the `frozen` loop and again in `_link_joins_inside_the_clause`. It is `_group._between_name_words(k, lo, hi, ...)` now and answers for both — one Python frame per generational connective instead of two, with `step` and its ternary gone. Nothing else moved: the left arm still short-circuits the right, and the arrays, the bounds and the two piece tests are the same tests in the same order. Byte-identical over 3,931,700 parses — 157,188 names (the corpus, cases.py, the property grids and a run-heavy generated set) under six lexicons and four policies, plus the facade surface and the #528 no-parse paths — same sha256 for all twenty-six digest rows and the dumps compare equal. + THE OLD NAME STILL STANDS WHERE IT WAS WRITTEN, and this sentence is the grep trail: `_name_word_beside` is the spelling in decisions.md#M2's 2026-09-20 #397 bullet and in the "NOT FIXED AND RECORDED INSTEAD" paragraph above, and dated entries are not edited, so neither is corrected — a search for either name reaches the whole arc. One place the old name was NOT a citation and is renamed: `tests/v2/test_properties.py`'s per-side mirror, now `_name_word_on_the_side`. That helper is a deliberately independent second implementation, reading spans and off-switch roles where the parser reads pieces and tags, and carrying the implementation's own name made it read as a call into the thing it checks rather than as a model of it. + THE FIGURES, AND THE INTERPRETER, because the pair above spliced two and this one does not. Re-measured 2026-09-21 on py3.11 through `tests/v2/test_benchmark.py`'s own `_frames_for` shape, `b9ed1429` → `9fd84463` → this tree: `Josep Carod i Rovira` 311 → 314 → 313, `Josep Lluis Carod i III` 377 → 381 → 380, `Jane Doe nee Puig i Soler` 315 → 320 → 319, `John Quincy Adams i MA Prof.` 438 → 443 → 442, and the clause guard's own pair 1,125 → 875 → 859 at sixteen links and 6,741 → 2,651 → 2,587 at sixty-four (3.01x, re-pinned in the test). `tools/perf/call_count.py` is unmoved at parse=406.00 facade=443.00 — its reference name carries no link — and so are `John Smith` 172, `Smith, John` 203, `Juan Garcia y Lopez` 291, `John and Jane Smith` 293, `Jane Doe nee Smith` 245 and `Jane Doe nee Smith PhD` 340, none of which reaches the predicate. WHICH OF THE EARLIER FIGURES REPRODUCE, stated exactly where this sentence first said only that one half "very nearly" did: re-measured 2026-09-22 on py3.11, the short-name pairs (304 → 307, 307 → 312) do not reproduce at all, and of the run-of-64 pair (6,741 → 2,652) the LEFT half does and the right does not — `b9ed1429` reads 6,741 on the nose, while `6048eb5d`, the tree its 2,652 was taken on, reads 2,651 here, the same figure `9fd84463` reads. So the splice runs through a single arrow rather than between two of them, which is exactly the spliced table `tools/perf/call_count.py`'s docstring was written about. Timings are unmoved: `"Josep " + "i " * n + "Rovira"` and `"Jane Doe nee Puig " + "i " * n + "Soler"` both read 1.96-2.02x per doubling from n=200 to n=1,600, 11.36ms and 9.39ms at the top end against 11.58ms and 9.58ms at `9fd84463`. + AND THE FOLD BOUGHT A TEST, which is the part worth keeping: while the predicate was called once per side, a mutation of the suffix or the title test hit BOTH sides at once and the left-hand rows killed it, so the right-hand halves were never separately covered — their own rows (`Josep Lluis Carod i Jr.`, `i Mr.`) stand at the END of the name, where `hi` refuses them before either piece test is asked. Folded, the two sides mutate independently and both right-hand tests survived the whole suite. `test_a_credential_or_honorific_mid_name_on_the_right_too` is the pair that kills them, and every arm of the folded predicate now dies by a named test (recorded in its docstring). + - Provenance: the single-letter-connective guard is v1's fix for Google Code issue 11 ("john e smith", 2013, commit 33676c9) — the "#11" citations that circulated pointed at a GitHub accident, not the real source. Recorded so the archaeology stays done. Excluded (Lexicon.conjunctions_ambiguous, the marked half of nameparser/config/conjunctions.py — an entry here reads as an INITIAL in a name written wholly in one case): diff --git a/docs/design/rules.md b/docs/design/rules.md index bd4ffd26..f171f78f 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1393,6 +1393,16 @@ M2. Rationale: a maiden marker announces that what follows it is the does not. "Jane Doe nee Smith MA Prof." → maiden="Smith MA Prof." · boundary "Jane Doe nee Smith Prof. MA" → maiden="Smith Prof." · boundary + Deviation: the link exception asks for a name word on each side, + and a separator the caller declared is structure rather than a + name word — so a link with one beside it is joining nothing and + ends the clause like any other suffix word. A declared separator + standing inside the clause, past its first word, is read as that + name word instead, and the clause runs on across a link it should + have ended at. The same clause written without the separator, + which leaves the title as the word on the link's left, does end + there. + "Smith, John, PhD née Puig Mr. - i Soler" extra_suffix_delimiters-dash → maiden="Puig Mr." deviates: #538 (today: maiden="Puig Mr. i Soler") history: decisions.md#M2 · interacts: P2, P3, P5, P6, R2, M1, S2, H1, H5 · implemented: nameparser/_pipeline/_group.py M3. Rationale: an enclosure says nothing about whether it means diff --git a/nameparser/_pipeline/_group.py b/nameparser/_pipeline/_group.py index b1f6e4a8..b66f6983 100644 --- a/nameparser/_pipeline/_group.py +++ b/nameparser/_pipeline/_group.py @@ -292,9 +292,10 @@ def _join_takes_the_member(view: Sequence[Sequence[int]], # reason, and the class test is that rule's own. The take runs BEFORE # every join, so the link is still a piece of its own here and the # question is asked of the pieces as classify left them -- the same -# inputs `_group_segment`'s `frozen` loop gives `_name_word_beside`, -# which is why this calls that predicate rather than restating the -# class (mechanisms.md#ONE-PREDICATE-PER-QUESTION). +# inputs `_group_segment`'s `frozen` loop gives +# `_between_name_words`, which is why this calls that predicate +# rather than restating the class +# (mechanisms.md#ONE-PREDICATE-PER-QUESTION). def _link_joins_inside_the_clause(k: int, lo: int, hi: int, pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], @@ -311,11 +312,11 @@ def _link_joins_inside_the_clause(k: int, lo: int, hi: int, neither is one standing before the generation or the credential a clause ends with ('... nee Puig i III', '... i MA'): `hi` is where assign's peel begins, so those stand at or past it and - `_name_word_beside` refuses them by bound. + `_between_name_words` refuses them by bound. Defined here, beside its one caller, and forward-referencing the two predicates it is built out of: `_is_conj_piece` and - `_name_word_beside` are the JOIN's, further down this module, and + `_between_name_words` are the JOIN's, further down this module, and moving them up to meet this would say they belonged to the clause. `beside` is the caller's memo cell, filled on the first CONNECTIVE @@ -330,10 +331,7 @@ def _link_joins_inside_the_clause(k: int, lo: int, hi: int, return False if not beside: beside.append(_run_neighbours(pieces, ptags, tokens)) - return (_name_word_beside(k, -1, lo, hi, pieces, ptags, tokens, - beside[0]) - and _name_word_beside(k, 1, lo, hi, pieces, ptags, tokens, - beside[0])) + return _between_name_words(k, lo, hi, pieces, ptags, tokens, beside[0]) def _maiden_take(pieces: Sequence[Sequence[int]], @@ -604,9 +602,35 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # marker is never the name word on a link's left, and a delimiter # core between the marker and that first word is below `lo` by # construction and cannot pass for one either. A core is the TAIL - # segment's alone (`extra_suffix_delimiters`, empty by default), so - # an ordinary dash is not one and does pass: 'PhD née - i Jones' - # keeps maiden '- i Jones' (measured 2026-09-20). + # segment's alone (`extra_suffix_delimiters`, empty by default), + # and a dash standing where no tail segment can hold it is an + # ordinary word at EITHER policy: 'PhD née - i Jones' keeps maiden + # '- i Jones' configured and unconfigured alike, there being no + # comma to make a tail out of. It takes the tail a suffix comma + # builds for the dash to be a core at all, and then the two + # policies part company -- 'Smith, John, PhD née - i Jones' keeps + # maiden '- i Jones' by default and declines under a configured + # ' - ', which is the pair + # test_a_core_between_the_marker_and_the_first_word_is_below_lo + # holds (all four readings measured 2026-09-22). + # NOT theoretical and not a whole claim about cores, both settled + # by measurement 2026-09-21 over corpus u cases.py u the property + # grids u a 50,925-name generated set with cores, under thirteen + # core-bearing policies: 25,536 of 596,392 maiden takes had a core + # standing exactly there, so the bound is load-bearing -- and PAST + # `lo` a core is no longer below it, is an ordinary index to + # `_run_neighbours` (which steps over connectives and nothing + # else), and DOES pass for the name word on a link's side. That + # reading is pinned as it stands rather than repaired here + # (test_a_core_beside_a_link_wrongly_passes_for_a_word_until_538): + # `_between_name_words` is asked about a core in 51,072 of 900,023 + # calls, the answer differs from a core-skipping reading in 8,094 + # parses over 1,278 texts, and 1,824 of those move `maiden` on 288 + # texts -- none of them reachable at the default policy, which is + # why rules.md#M2 states it with a policy annotation beside the + # marker. The repair is `cores` threaded through + # three call sites into `_run_neighbours`, which is its own change + # (#538, and rules.md#M2 carries it as a `deviates:` example). # `peel_start` is where assign's trailing run begins over # the pieces as WRITTEN, so the generation or credential a clause # ends with is never the name word on a link's right ('... nee Puig @@ -685,7 +709,7 @@ def _maiden_take(pieces: Sequence[Sequence[int]], return seen[m:m + run], seen[m + run:j] -#: What `_name_word_beside` reads instead of walking: two arrays over +#: What `_between_name_words` reads instead of walking: two arrays over #: the segment's pieces, giving for each index the nearest piece on #: its left and on its right that is NOT a connective -- `-1` and #: `len(pieces)` where the run reaches the end. Built by @@ -700,7 +724,7 @@ def _run_neighbours(pieces: Sequence[Sequence[int]], """The nearest non-connective piece on each side of every index. EVERY MEMBER OF ONE RUN HAS THE SAME ANSWER, which is the whole - of the fix: `_name_word_beside` used to walk the run itself, so a + of the fix: `_between_name_words` used to walk the run itself, so a name holding a run of n connectives walked it n times and the stage went quadratic in the run's length -- measured 2026-09-20, `"Josep " + "i " * n + "Rovira"` grew 3.8x per doubling against @@ -718,11 +742,22 @@ def _run_neighbours(pieces: Sequence[Sequence[int]], WHAT IT COSTS A SHORT NAME, because answering for the whole segment is not free where the walk would have stopped at once: this call, plus `_is_conj_piece` for the pieces the walk never - reached. Measured 2026-09-20 against b9ed1429 -- `Josep Carod i - Rovira` 304 -> 307 frames (one call and two more `_is_conj_piece` - over its four pieces) and `Jane Doe nee Puig i Soler` 307 -> 312. - An O(1) rise per link-bearing name against an unbounded saving: - the same name with a run of 64 links goes 6,741 -> 2,652. + reached. Re-measured 2026-09-21 on py3.11 against b9ed1429, the + whole pair through `tests/v2/test_benchmark.py`'s own + `_frames_for` shape -- `Josep Carod i Rovira` 311 -> 313 frames + (one call and two more `_is_conj_piece` over its four pieces, + less the frame the fold below saved) and `Jane Doe nee Puig i + Soler` 315 -> 319. An O(1) rise per link-bearing name against an + unbounded saving: the same name with a run of 64 links goes + 6,741 -> 2,587. (The pair first written here read 304 -> 307 and + 307 -> 312, with 6,741 -> 2,652 for the run of 64. Re-measured + 2026-09-22 on py3.11, WHICH OF THOSE REPRODUCE is: the two + short-name pairs, neither of them; the run-of-64 pair, its left + half only -- b9ed1429 reads 6,741 exactly, while 6048eb5d, the + tree the 2,652 was taken on, reads 2,651 here. So the interpreter + splice ran through a single arrow, which is what the table + `tools/perf/call_count.py`'s own docstring warns about. Every + figure above is one interpreter, stated.) `tools/perf/call_count.py` is unmoved (parse=406.00, facade=443.00) -- its reference name carries no link -- and so are `John Smith`, `Smith, John`, `Juan Garcia y Lopez` and `Jane Doe @@ -736,7 +771,7 @@ def _run_neighbours(pieces: Sequence[Sequence[int]], Dropping either pass's `not` fails test_a_connective_piece_counts_toward_the_carve_outs_total (mutation-checked 2026-09-20; how the two arrays are READ is - checked in `_name_word_beside`, which reads them). + checked in `_between_name_words`, which reads them). """ n = len(pieces) left = [-1] * n @@ -807,13 +842,21 @@ def _is_rootname(piece: Sequence[int], ptags: Set[str], # when the maiden walk became a second caller: this reads a piece and # never edits one, and `_maiden_take` holds its pieces at the wider # type the stage's entry point hands it. -def _name_word_beside(k: int, step: int, lo: int, hi: int, - pieces: Sequence[Sequence[int]], - ptags: Sequence[Set[str]], - tokens: Sequence[WorkToken], - beside: _Beside) -> bool: - """Whether such a word stands on the `step` side of the - connective piece at `k`. +def _between_name_words(k: int, lo: int, hi: int, + pieces: Sequence[Sequence[int]], + ptags: Sequence[Set[str]], + tokens: Sequence[WorkToken], + beside: _Beside) -> bool: + """Whether such a word stands on EACH side of the connective piece + at `k`. + + Both sides in one call because neither caller ever wants one: a + connective is placed to join only where a name word stands on both + sides of it, so the two answers were always ANDed at the call site + and the left one short-circuits the right either way. One frame per + generational connective rather than two, which is the whole of the + saving -- the left arm below is the old left call and the right arm + the old right one, unchanged (#397 follow-up). `lo` and `hi` bound the name's own words: assign peels the pieces below `lo` as its leading titles and those from `hi` up as its @@ -836,27 +879,44 @@ def _name_word_beside(k: int, step: int, lo: int, hi: int, stepping already happened: `_run_neighbours` walked every run once for the whole segment, so this reads an index rather than walking to it. A SENTINEL OUT OF RANGE is how "the run ran out" arrives -- - -1 on the left, `len(pieces)` on the right -- and the bound test + -1 on the left, `len(pieces)` on the right -- and each bound test below turns it into False, exactly as the walk did when it ran here and stopped at the same place. `lo` is never negative and `hi` never past `len(pieces)`, so neither sentinel can pass the - bound, and the two piece tests are never asked about an index - that is not one. - - Mutation-checked 2026-09-20, each side by a NAMED test: reading - the left array for both sides fails - test_a_connective_with_nothing_to_its_right_does_not_join, and - the right array for both fails - test_a_leading_title_on_the_left_is_no_name_word. SWAPPING the - two arrays outright is an EQUIVALENT mutant and no test fails: - both callers ask for a name word on each side and AND the two - answers, so which array answers which side is not a question the - conjunction can see. + bound, and the piece tests are never asked about an index that is + not one. + + Mutation-checked 2026-09-21, every arm of both sides by a NAMED + test. Reading the LEFT array for the right arm too fails + test_a_connective_with_nothing_to_its_right_does_not_join; + the RIGHT array for the left arm too fails + test_a_leading_title_on_the_left_is_no_name_word. + The bounds, WIDENED to the whole segment rather than dropped -- + dropping the right one indexes past the pieces on the sentinel and + raises instead of failing a test -- fail + test_the_marker_is_not_the_name_word_on_the_links_left (left) and + test_the_right_hand_test_reads_the_peel_not_the_suffix_piece + (right). Either piece test on the LEFT fails + test_a_credential_or_honorific_mid_name_is_no_name_word_either; + either on the RIGHT fails + test_a_credential_or_honorific_mid_name_on_the_right_too, which is + the row the fold asked for: while this was two per-side calls a + mutation hit both sides at once and the left-hand rows covered for + the right-hand ones, whose own rows stand at the END of the name + where `hi` refuses them first. SWAPPING the two arrays outright is + an EQUIVALENT mutant and no test fails: the two arms are ANDed, so + which array answers which side is not a question the conjunction + can see. """ - j = beside[0][k] if step < 0 else beside[1][k] - return (lo <= j < hi - and not is_suffix_piece(pieces[j], ptags[j], tokens) - and not is_title_piece(pieces[j], ptags[j], tokens)) + left = beside[0][k] + if not (lo <= left < hi + and not is_suffix_piece(pieces[left], ptags[left], tokens) + and not is_title_piece(pieces[left], ptags[left], tokens)): + return False + right = beside[1][k] + return (lo <= right < hi + and not is_suffix_piece(pieces[right], ptags[right], tokens) + and not is_title_piece(pieces[right], ptags[right], tokens)) def _group_segment(seg: tuple[int, ...], additional: int, @@ -1126,10 +1186,8 @@ def merge(lo: int, hi: int, add: Set[str] = frozenset(), # per member -- quadratic in its length, 3.8x per # doubling measured at `b9ed1429`. beside = _run_neighbours(pieces, ptags, tokens) - if not (_name_word_beside(k, -1, lo, hi, pieces, ptags, tokens, - beside) - and _name_word_beside(k, 1, lo, hi, pieces, ptags, - tokens, beside)): + if not _between_name_words(k, lo, hi, pieces, ptags, tokens, + beside): frozen.add(piece[0]) total = sum(_is_rootname(p, t, tokens) for p, t in zip(pieces, ptags) diff --git a/tests/v2/pipeline/test_group.py b/tests/v2/pipeline/test_group.py index bc3d9264..1077db49 100644 --- a/tests/v2/pipeline/test_group.py +++ b/tests/v2/pipeline/test_group.py @@ -1517,6 +1517,31 @@ def test_a_credential_or_honorific_mid_name_is_no_name_word_either( ["Josep", "Lluis", "Mr. i Rovira"]] +def test_a_credential_or_honorific_mid_name_on_the_right_too() -> None: + # the MIRROR of the pair above, and the two rows that make the + # right-hand piece tests mean something: the credential and the + # honorific rows further up stand at the END of the name, where + # `hi` refuses them before either piece test is asked, so with + # those rows alone the right-hand tests could be deleted and no + # test would fail (measured 2026-09-21 by mutation, which is what + # asking both sides in ONE call made visible -- a per-side call + # mutated both sides at once and the left-hand rows covered it). + # Mid-name on the RIGHT is inside both bounds, so only the piece + # tests keep the link from joining a credential or a title. + suffix = _grouped("Josep Lluis i Jr. Rovira", lexicon=_LINK_LEX) + assert _piece_texts(suffix) == [ + ["Josep", "Lluis", "i", "Jr.", "Rovira"]] + title = _grouped("Josep Lluis i Mr. Rovira", lexicon=_LINK_LEX) + assert _piece_texts(title) == [ + ["Josep", "Lluis", "i", "Mr.", "Rovira"]] + assert _piece_texts(_grouped("Josep Lluis i Jr. Rovira", + lexicon=_PLAIN_LEX)) == [ + ["Josep", "Lluis i Jr.", "Rovira"]] + assert _piece_texts(_grouped("Josep Lluis i Mr. Rovira", + lexicon=_PLAIN_LEX)) == [ + ["Josep", "Lluis i Mr.", "Rovira"]] + + def test_the_walk_looks_past_a_run_of_connectives() -> None: # a RUN joins as one, so the word the condition is about is the # first one past the run, not the connective beside the link. @@ -1717,7 +1742,7 @@ def test_a_clause_link_with_nothing_on_its_right_still_ends_it( def test_a_generation_on_the_links_right_is_no_name_word() -> None: # the second control, refused by CLASS: 'jr' is suffix vocabulary, - # so `_name_word_beside` declines it wherever it stands. + # so `_between_name_words` declines it wherever it stands. out = _grouped("Jane Doe née Puig i jr", lexicon=_LINK_LEX) assert _maiden_texts(out) == ["Puig"] plain = _grouped("Jane Doe née Puig i jr", lexicon=_PLAIN_LEX) @@ -1765,6 +1790,64 @@ def test_the_marker_is_not_the_name_word_on_the_links_left() -> None: assert _maiden_texts(plain) == ["i", "Soler"] +def test_a_core_between_the_marker_and_the_first_word_is_below_lo( +) -> None: + # A delimiter core is TAIL-segment structure that group() drops + # after this pass, so it is never a word of the clause -- and + # between the marker and the first word it is below `lo`, which + # the bound refuses without the piece tests ever being asked. + # Reachable, not theoretical: measured 2026-09-21 over corpus u + # cases.py u the property grids u a 50,925-name generated set with + # cores, under thirteen core-bearing policies, 25,536 of 596,392 + # maiden takes had a core standing there. + out = _grouped("Smith, John, PhD née - i Jones", policy=_DASH, + lexicon=_LINK_LEX) + assert _maiden_texts(out) == [] + # and the control that says the CORE is doing it: with no + # delimiter configured the dash is an ordinary word, so the link + # has a name word on its left and the clause keeps the run. + plain = _grouped("Smith, John, PhD née - i Jones", lexicon=_LINK_LEX) + assert _maiden_texts(plain) == ["-", "i", "Jones"] + + +def test_a_core_beside_a_link_wrongly_passes_for_a_word_until_538( +) -> None: + """A KNOWN-WRONG reading, pinned so the repair has to move it. + + rules.md#M2 gives the link exception a name word on each side, + and a delimiter core is structure rather than a name word -- so + the clause below should end where its separator-less twin ends. + It does not. Update this test when #538 lands: the assertion + beneath the first parse is the deviation, not the contract, and + rules.md#M2's `deviates: #538` example is its other half. + """ + # WHAT IS NOT TRUE OF A CORE PAST `lo`, pinned as it reads rather + # than as it ought to: inside the clause a core is an ordinary + # index to `_run_neighbours`, which steps over CONNECTIVES and + # nothing else, so it stands as the name word on the link's left + # and the clause runs on past a title it would otherwise stop at. + # `_between_name_words` is asked about a core on one side or the + # other in 51,072 of 900,023 calls over the population above, and + # the answer differs from a core-skipping reading in 8,094 parses + # (1,278 texts); 1,824 of those move the `maiden` field, on 288 + # texts. None of the 288 is reachable at the default policy, + # `extra_suffix_delimiters` being empty there -- so the one of + # them rules.md#M2 now carries as a `deviates: #538` example (this + # row's first text) enters corpus_rules.jsonl as a name the gate + # parses with the DEFAULT facade, where it moves for the link fix + # and not for this. Reported, not fixed: the repair is `cores` + # threaded through three call sites into `_run_neighbours`, not a + # one-liner (#538). + out = _grouped("Smith, John, PhD née Puig Mr. - i Soler", + policy=_DASH, lexicon=_LINK_LEX) + assert _maiden_texts(out) == ["Puig", "Mr.", "i", "Soler"] + # the same clause with the core taken out of it: the title IS the + # word on the link's left and refuses, so the clause ends there. + without = _grouped("Smith, John, PhD née Puig Mr. i Soler", + policy=_DASH, lexicon=_LINK_LEX) + assert _maiden_texts(without) == ["Puig", "Mr."] + + def test_a_marker_with_nothing_after_it_declines_before_the_bound( ) -> None: # the clause's `lo` is the first piece after the marker run, and diff --git a/tests/v2/rules_doc.py b/tests/v2/rules_doc.py index f4ba8e8f..f0525b24 100644 --- a/tests/v2/rules_doc.py +++ b/tests/v2/rules_doc.py @@ -131,6 +131,14 @@ def has_boundary_or_waiver(self) -> bool: #: one is on by default and the other off. "unlisted_dotted_suffixes-off": Policy(unlisted_dotted_suffixes=False), "unlisted_caps_suffixes-on": Policy(unlisted_caps_suffixes=True), + #: The delimiter switch (#206), named after its Policy FIELD for + #: the reason above and carrying the DELIMITER in the suffix, + #: because this field's value is a set rather than a flag: a rule + #: reached only under a configured separator has to say which one, + #: and " - " is the spelling docs/customize.rst teaches and the + #: one the group tests configure. + "extra_suffix_delimiters-dash": Policy( + extra_suffix_delimiters=frozenset({" - "})), } #: D-section subjects: zero-arg constructions whose diagnostics the #: warns=/raises= assertion forms exercise. diff --git a/tests/v2/test_benchmark.py b/tests/v2/test_benchmark.py index 73509781..913d572f 100644 --- a/tests/v2/test_benchmark.py +++ b/tests/v2/test_benchmark.py @@ -290,7 +290,7 @@ def test_a_thousand_names_still_parse_in_reasonable_time( # repeated runs, inside the clean column, and neither number moved. # The thirteenth (link_run, #397 second review) arrived with its own # quadratic in hand as well, and it is the one this shape was added -# FOR rather than one found by adding it: `_name_word_beside` walked +# FOR rather than one found by adding it: `_between_name_words` walked # the run of connectives beside a link once per MEMBER of that run, # so at commit b9ed1429 the shape measures 8.48 at base 100, 10.29 at # 200 and 12.11 at 400 -- outside the bound at every one of them, and @@ -506,14 +506,22 @@ def test_a_trailing_credential_run_does_not_cost_exponentially() -> None: # # One pair, not two: the defect here is a quadratic and there is no # exponential to order it against, so the 16-vs-64 pair is the whole -# guard. Measured 2026-09-20 through this file's own `_frames_for`: -# 876 frames at 16 and 2,652 at 64 on this tree (3.03x), against -# 1,125 and 6,741 at b9ed1429 (5.99x), where `_name_word_beside` -# walked the run once per member -- identical on three repeated runs -# at each end, frame counts being deterministic. +# guard. Re-measured 2026-09-21 on py3.11 through this file's own +# `_frames_for`: 859 frames at 16 and 2,587 at 64 on this tree +# (3.01x), against 1,125 and 6,741 at b9ed1429 (5.99x), where +# `_between_name_words` walked the run once per member -- identical on +# three repeated runs at each end, frame counts being deterministic. +# One frame per link below the 875/2,651 this same helper read before +# `_between_name_words` answered both sides in one call. The pair +# first recorded here, 876/2,652, is one frame above that and does NOT +# reproduce: re-measured 2026-09-22 on py3.11, 6048eb5d -- the commit +# that wrote those two figures into this comment -- reads 875 and +# 2,651, which is the same pair 9fd84463 reads. So that recording came +# off another interpreter; it stays, with the py3.11 reading of its own +# tree now beside it. _CLAUSE_RUN_SMALL = 16 _CLAUSE_RUN_LARGE = 64 -#: 3.03x measured here against 5.99x at b9ed1429: 4.5 sits ~1.5x over +#: 3.01x measured here against 5.99x at b9ed1429: 4.5 sits ~1.5x over #: the measurement and ~1.3x under the regression. Frame counts are #: deterministic for a given tree and interpreter, so both margins are #: for a future shape change rather than for runner noise. @@ -553,5 +561,83 @@ def test_a_clause_link_run_does_not_cost_quadratically() -> None: f"and one holding {_CLAUSE_RUN_LARGE} costs {large} -- " f"{ratio:.1f}x for 4x the input, where this tree measures 3.0x " f"and the per-member walk at b9ed1429 measured 6.0x. " - f"_group.py's `_name_word_beside` is walking the run per member " + f"_group.py's `_between_name_words` is walking the run per member " f"again (#397)") + + +# THE ABSOLUTE COST OF A LINK, which the ratio above cannot see: a +# change costing ONE MORE FRAME PER LINK moves both ends of the pair +# and leaves the ratio where it was. Re-splitting the #397 follow-up's +# fold is exactly that change -- 859/2,587 here against 875/2,651 at +# 9fd84463, +1 per link at each size -- and against THIS suite the +# pre-fold parser is green everywhere but here. Measured 2026-09-22 on +# py3.11: a copy of this tree carrying `git archive 9fd84463 +# nameparser` in place of its own runs 9,652 passed, 324 skipped, 4 +# xfailed and ONE failure, this test. So a link-bearing name gets a +# banded absolute pin, keyed by interpreter and banded like +# `_CALL_BASELINE` above. +# +# THE LONG RUN AND NOT THE SHORT ONE, and the arithmetic is the whole +# reason: +1 per link is 16 frames at `_CLAUSE_RUN_SMALL`, and 2% of +# 859 is 17.2, so the small end's band swallows the very regression +# this pin exists for (875 sits inside 842-876). At +# `_CLAUSE_RUN_LARGE` the same change is +64 against a 2% band of +# 51.7, and 2,651 sits outside 2,535-2,639. A tighter band on the +# short name would do it too and was not taken: the band is the one +# `_CALL_BASELINE` uses, and a bespoke one here would need its own +# argument every time the shape moved. +# +# ONE ROW, py3.11, and an unknown interpreter SKIPS rather than fails +# -- which is where this parts company with `_check_budget`, whose +# table carries every interpreter CI runs and so can afford to fail on +# a missing row. Only py3.11 is measurable in this working tree, and a +# figure nobody here ran is not a pin. The #537 reviewer reports 3.12 +# at 835/2,563 and 3.13/3.14 at 882/2,706 for the 16/64 pair; those +# are NOT recorded below, because recording them would put a number +# under a band without a run behind it. Reproduce one on its own +# interpreter and add the row. +_LINK_BASELINE = { + (3, 11): 2587, +} +#: The same +-2% `_CALL_BASELINE` uses, and for the same reason: frame +#: counts are deterministic for a given tree and interpreter, so the +#: band is headroom for a deliberate shape change rather than for +#: runner noise. +_LINK_BAND = 0.02 + + +def test_a_link_costs_what_it_is_pinned_at() -> None: + """The absolute frame cost of one link-bearing name. + + The ratio guard above cannot fail on a per-link constant, because + a constant moves its two ends together. This one can, and it is + the only thing in the suite that can. + """ + if sys.getprofile() is not None: + pytest.skip("a profile hook is already installed; this test owns it") + version = sys.version_info[:2] + if version not in _LINK_BASELINE: + pytest.skip( + f"no link baseline for Python {version[0]}.{version[1]}; " + f"measure `_clause_run({_CLAUSE_RUN_LARGE})` on this " + f"interpreter and add the row to _LINK_BASELINE") + text = _clause_run(_CLAUSE_RUN_LARGE) + # REACHABILITY, the probe every shape in this file carries: the + # frames counted are the clause walk's, and they are only there + # while the clause KEEPS the run (rules.md#M2's link exception). + # End the clause at the first link and this measures a name that + # no longer holds 64 links, comfortably inside the band forever. + assert parse(text).maiden == " ".join( + ["Puig"] + ["i"] * _CLAUSE_RUN_LARGE + ["Soler"]) + baseline = _LINK_BASELINE[version] + actual = _frames_for(text) + low, high = baseline * (1 - _LINK_BAND), baseline * (1 + _LINK_BAND) + assert low <= actual <= high, ( + f"a maiden clause holding {_CLAUSE_RUN_LARGE} links costs " + f"{actual} frames on Python {version[0]}.{version[1]}, band " + f"{low:.0f}-{high:.0f} around a baseline of {baseline}. Growth " + f"and shrinkage are both signals, and one frame per link is " + f"enough to reach this band where the ratio guard above cannot " + f"see it: check whether `_group._between_name_words` still " + f"answers both sides of a link in one call, then move the " + f"baseline deliberately (#397)") diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index 44dc074d..a68cde57 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -2664,12 +2664,16 @@ class _LatinCopy(NamedTuple): "Washington Jr\\. MD, Franklin", "abdul Smith Jr Ma", "abdul Smith Jr V"}), # #397's maiden-clause rule, one corpus name per alternative -- a - # list of names, not a copy of any wordlist. What selects the two - # is the PLACEMENT of the link inside a clause, which no + # list of names, not a copy of any wordlist. What selects the + # three is the PLACEMENT of the link inside a clause, which no # vocabulary decides; a member spelled as the shape (a bare letter # after a maiden marker) would reach every clause name in the - # corpora and pre-excuse the readings the walk must refuse. - frozenset({"Doe, Jane nee Puig i Soler", "Jane Doe nee Puig i Soler"}), + # corpora and pre-excuse the readings the walk must refuse. The + # third member is rules.md#M2's `deviates: #538` example, which + # the corpus parses at the DEFAULT policy and so for this rule's + # own sentence rather than for #538's. + frozenset({"Doe, Jane nee Puig i Soler", "Jane Doe nee Puig i Soler", + "Smith, John, PhD née Puig Mr\\. - i Soler"}), # #397's join and its one-case report, one corpus name per # alternative -- lists of names, not copies of any wordlist. The # join's subject is a SHAPE the vocabulary participates in at one @@ -3234,8 +3238,13 @@ def _claim(rule: dict) -> _Claim: # shape-tagged case rows, and this regex reaches every name # carrying a marker. Read name by name against the regex; no # role joined the list. + # 2026-09-22, #397 follow-up: 77 -> 78, the one new corpus + # name rules.md#M2's `deviates: #538` example adds, + # 'Smith, John, PhD née Puig Mr. - i Soler'. Reach again -- + # it carries a marker -- and verified name by name; no role + # joined the list. "fix(#274) maiden markers consumed": - _Claim(77, ('family', 'maiden', 'middle'), 'd36e74b1f60d', None), + _Claim(78, ('family', 'maiden', 'middle'), '08e08f62622f', None), # 2026-09-19, #533: 5 -> 6, the same one new corpus name # '田中 太郎 旧姓 佐藤 MA' as the CJK rule above. "fix(cjk-maiden-marker) maiden marker consumed, compounding with the CJK order flip": @@ -3335,8 +3344,11 @@ def _claim(rule: dict) -> _Claim: # the review's rows add, 'Rovira, Josep Carod i Jr.' -- the # comma form of the swallowed generation. Reach again, and # verified name by name. + # 2026-09-22, #397 follow-up: 363 -> 364, the one new comma + # name the M2 deviation example adds. Reach again, verified + # name by name. "fix(comma-family) lone post-comma piece routes to suffix/title, not first": - _Claim(363, ('given', 'suffix', 'title'), 'b628b4d25dfb', None), + _Claim(364, ('given', 'suffix', 'title'), 'e3bf2a2b8de9', None), "fix(comma-family) a comma followed only by titles keeps the given/family split": _Claim(2, ('family', 'given'), "5bd9c6d96c38", None), "fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example": @@ -3390,8 +3402,10 @@ def _claim(rule: dict) -> _Claim: # the rule above and for the same reason. # 2026-09-20, #397 review: 359 -> 360, the same one new comma # name as the rule above and for the same reason. + # 2026-09-22, #397 follow-up: 363 -> 364, the same one new + # comma name as the rule above and for the same reason. "fix(comma-precomma-family) pre-comma run reads as family, not given": - _Claim(363, ('family', 'given'), 'b628b4d25dfb', None), + _Claim(364, ('family', 'given'), 'e3bf2a2b8de9', None), # 2026-09-20, #397: retitled in place, reach and digest # unchanged -- the rule keeps 'Carod i', which the landing # leaves byte-identical. @@ -3866,6 +3880,11 @@ def _claim(rule: dict) -> _Claim: "fix(#274/#436/#437) a clause the generational link ended, and the run it left renders with spaces": _Claim(1, ('family', 'maiden', 'middle', 'suffix'), '7d64445a7252', None), + # 2026-09-22, #397 follow-up: new rule, one literal name -- + # rules.md#M2's `deviates: #538` example, which this baseline + # reads as one long suffix. + "fix(#274/#397) a maiden clause inside a suffix-comma tail leaves the suffix field, link and all": + _Claim(1, ('maiden', 'suffix'), 'e20491ebfe62', None), }, "expected_since_2.0.0.toml": { # #436/#437's Latin alternation, first in every ledger. @@ -4371,9 +4390,12 @@ def _claim(rule: dict) -> _Claim: _Claim(16, ('_initials',), "075dc34f9e95", ('DEFAULT',)), "fix(#461) a connective with a name word beside it stops contributing an initial": _Claim(13, ('_initials',), "3cc41f4bfc21", ('DEFAULT',)), + # 2026-09-22, #397 follow-up: 2 -> 3, the one new corpus name + # rules.md#M2's `deviates: #538` example adds, which this + # rule's own alternation now names. Verified name by name. "fix(#397) a link inside a maiden clause stays in the birth name": - _Claim(2, ('family', 'maiden', 'middle', 'suffix'), - '32405b182f4e', ('DEFAULT',)), + _Claim(3, ('family', 'maiden', 'middle', 'suffix'), + '4de0e7570bd6', ('DEFAULT',)), }, # The 2.3 cycle's first rule, and a facade-only render fix: every # role is identical, so `_initials` alone. Reach and digest as in @@ -4659,9 +4681,12 @@ def _claim(rule: dict) -> _Claim: _Claim(17, ('_initials',), "797473971e75", ('DEFAULT',)), "fix(#461) a connective with a name word beside it stops contributing an initial": _Claim(13, ('_initials',), "3cc41f4bfc21", ('DEFAULT',)), + # 2026-09-22, #397 follow-up: 2 -> 3, the one new corpus name + # rules.md#M2's `deviates: #538` example adds, which this + # rule's own alternation now names. Verified name by name. "fix(#397) a link inside a maiden clause stays in the birth name": - _Claim(2, ('family', 'maiden', 'middle', 'suffix'), - '32405b182f4e', ('DEFAULT',)), + _Claim(3, ('family', 'maiden', 'middle', 'suffix'), + '4de0e7570bd6', ('DEFAULT',)), }, "expected_since_2.1.0.toml": { # #436/#437's Latin alternation, first in every ledger. @@ -5138,9 +5163,12 @@ def _claim(rule: dict) -> _Claim: _Claim(16, ('_initials',), "075dc34f9e95", ('DEFAULT',)), "fix(#461) a connective with a name word beside it stops contributing an initial": _Claim(13, ('_initials',), "3cc41f4bfc21", ('DEFAULT',)), + # 2026-09-22, #397 follow-up: 2 -> 3, the one new corpus name + # rules.md#M2's `deviates: #538` example adds, which this + # rule's own alternation now names. Verified name by name. "fix(#397) a link inside a maiden clause stays in the birth name": - _Claim(2, ('family', 'maiden', 'middle', 'suffix'), - '32405b182f4e', ('DEFAULT',)), + _Claim(3, ('family', 'maiden', 'middle', 'suffix'), + '4de0e7570bd6', ('DEFAULT',)), }, "expected_since_2.3.0.toml": { # #383/#479's three rules, the first this ledger carries. The @@ -5293,9 +5321,12 @@ def _claim(rule: dict) -> _Claim: _Claim(17, ('_initials',), "797473971e75", ('DEFAULT',)), "fix(#461) a connective with a name word beside it stops contributing an initial": _Claim(22, ('_initials',), "e73827447b4e", ('DEFAULT',)), + # 2026-09-22, #397 follow-up: 2 -> 3, the one new corpus name + # rules.md#M2's `deviates: #538` example adds, which this + # rule's own alternation now names. Verified name by name. "fix(#397) a link inside a maiden clause stays in the birth name": - _Claim(2, ('family', 'maiden', 'middle', 'suffix'), - '32405b182f4e', ('DEFAULT',)), + _Claim(3, ('family', 'maiden', 'middle', 'suffix'), + '4de0e7570bd6', ('DEFAULT',)), }, } diff --git a/tests/v2/test_properties.py b/tests/v2/test_properties.py index 4d69445a..fd310e6f 100644 --- a/tests/v2/test_properties.py +++ b/tests/v2/test_properties.py @@ -1676,15 +1676,23 @@ def _off_roles(off: ParsedName) -> dict[tuple[int, int], Role]: return {span: tok.role for tok, span in _placed(off)} -def _name_word_beside(toks: list[_Placed], i: int, step: int, - off_role: dict[tuple[int, int], Role], - original: str) -> bool: +def _name_word_on_the_side(toks: list[_Placed], i: int, step: int, + off_role: dict[tuple[int, int], Role], + original: str) -> bool: """Whether a name word stands on the `step` side of toks[i], judged by the OFF-SWITCH parse's roles -- which is what makes this a PRE-JOIN reading: in that parse the class letter is no connective, so no join has moved anything. A comma between ends the walk (it is another segment), and connectives are stepped - over because a run of them joins as one.""" + over because a run of them joins as one. + + DELIBERATELY A SECOND IMPLEMENTATION, and named so that the + mirror cannot be mistaken for the glass: it asks ONE side per + call over spans and off-switch roles, where the parser's + `_group._between_name_words` asks both at once over pieces and + tags. The name it used to carry was the implementation's own, + which read as a call into the thing under test rather than as a + model of it.""" j = i while True: k = j + step @@ -1758,8 +1766,9 @@ def _link_joins_between_name_words(on: ParsedName, off: ParsedName, if ("conjunction" not in tok.tags or tok.text.lower() not in letters): continue - if (_name_word_beside(toks, i, -1, off_role, on.original) - and _name_word_beside(toks, i, 1, off_role, on.original)): + if (_name_word_on_the_side(toks, i, -1, off_role, on.original) + and _name_word_on_the_side(toks, i, 1, off_role, + on.original)): return True return False diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index ea049949..112d9847 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -275,6 +275,7 @@ "Smith, John Prof." "Smith, John V" "Smith, John V." +"Smith, John, PhD née Puig Mr. - i Soler" "Smith, John, and" "Smith, Jr." "Smith, MA" diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index dca0797a..f5900948 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -4284,10 +4284,13 @@ fields = ["family", "maiden", "middle"] # #397 review: A LINK INSIDE A MAIDEN CLAUSE STAYS IN THE BIRTH NAME -# (2026-09-20). Two rules, because at this baseline two OTHER changes -# ride along in the same diffs and one rule has to explain the whole -# of each: the marker leaving the name (#274) and R1's space-joined -# rendering of a suffix run v1 wrote with a comma (#436/#437). +# (2026-09-20; a third rule added 2026-09-22 with rules.md#M2's +# `deviates: #538` example). Separate rules rather than one, because +# at this baseline OTHER changes ride along in the same diffs and one +# rule has to explain the whole of each: the marker leaving the name +# (#274), R1's space-joined rendering of a suffix run v1 wrote with a +# comma (#436/#437), and -- for the name added last -- the suffix +# field v1 put a whole suffix-comma tail in. # # The third new name, 'Jane Doe nee Puig i Soler', needs no rule here: # its diff at this baseline is {family, maiden, middle}, which the @@ -4331,3 +4334,30 @@ issue = "fix(#274/#436/#437) a clause the generational link ended, and the run i # attribute this stop to the wrong reading. name_regex = "^Jane Doe nee Puig i III$" fields = ["family", "maiden", "middle", "suffix"] + +[[change]] +issue = "fix(#274/#397) a maiden clause inside a suffix-comma tail leaves the suffix field, link and all" +# 'Smith, John, PhD née Puig Mr. - i Soler', which arrives from +# rules.md#M2 rather than from a report: it is the doc's +# `deviates: #538` example, and the doc reaches that deviation with a +# configured ' - ' this corpus does not apply. What the gate sees is +# the DEFAULT facade reading, where the dash is an ordinary name word +# like any other. +# +# v1 had no maiden markers at all, so everything behind the suffix +# comma stayed in the suffix: suffix 'PhD née Puig Mr. - i Soler' +# against the tree's 'PhD', maiden '' against 'Puig Mr. - i Soler' +# (measured 2026-09-22). One rule for the whole of it -- the marker +# leaving the name (#274) and, inside what it takes, the link kept in +# the clause it stands in (#397), which is what puts 'i Soler' on the +# maiden side of the same two fields. +# +# Not a widening of the fix(#274/#397) comma rule above: that one is +# about a FAMILY comma, where v1 left a middle name behind and the +# field list says so. This tail leaves v1 a suffix and nothing else. +# +# Literal, one name, for the reason those rules give: the shape would +# be "a maiden marker and a bare letter", which reaches every clause +# name in the corpora. +name_regex = "^Smith, John, PhD née Puig Mr\\. - i Soler$" +fields = ["maiden", "suffix"] diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 113e2af9..4e884a2c 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -3201,8 +3201,8 @@ orders = ["DEFAULT"] # #397 review: A LINK INSIDE A MAIDEN CLAUSE STAYS IN THE BIRTH NAME # (2026-09-20). LAST in the file, narrow-first: this rule declares # {family, maiden, middle, suffix} and every rule above it that could -# reach either name declares a strict subset, so an earlier position -# would be an order-decided contest the run refuses (#382). +# reach any of the three declares a strict subset, so an earlier +# position would be an order-decided contest the run refuses (#382). # # rules.md#M2 states the exception; decisions.md#M2's 2026-09-20 # bullet carries the measurement and why it is fixed in #397's own PR @@ -3211,7 +3211,7 @@ orders = ["DEFAULT"] [[change]] issue = "fix(#397) a link inside a maiden clause stays in the birth name" -# Two names, one sentence, two shapes -- `fields` is the union. +# Three names, one sentence, three shapes -- `fields` is the union. # # The maiden walk has always ended the birth name at the first suffix # WORD after the marker, and the Catalan link is also the roman @@ -3230,12 +3230,25 @@ issue = "fix(#397) a link inside a maiden clause stays in the birth name" # the tree, 'y' being connective vocabulary and no generation. That # twin is the property invariant in tests/v2/test_properties.py. # -# Literal-anchored to the two. The shape is "a one-letter connective +# The THIRD name arrives from rules.md#M2 rather than from a report: +# 'Smith, John, PhD née Puig Mr. - i Soler' is the doc's +# `deviates: #538` example, and the doc reaches the deviation with a +# configured ' - ' that this corpus does not apply. Parsed with the +# DEFAULT facade, as the gate parses it, the dash is an ordinary name +# word, so the link has one on each side and the clause keeps +# 'i Soler' where this baseline left it in the suffix: suffix +# 'PhD i Soler' -> 'PhD', maiden 'Puig Mr. -' -> 'Puig Mr. - i Soler' +# (measured 2026-09-22 at all four 2.x baselines, identical at each). +# That is this rule's own sentence and no part of #538, whose reading +# needs the delimiter declared; {maiden, suffix} is a subset of the +# union above. +# +# Literal-anchored to the three. The shape is "a one-letter connective # inside a maiden clause", which would stand ready to excuse every # future regression at this walk -- and the walk's whole design is # that some such letters must NOT be taken. _MUST_NOT_MATCH carries # the wall: the i-last control, the generation behind it, and the # spellings where the letter is an initial. -name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler)$" +name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler)$" fields = ["family", "maiden", "middle", "suffix"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index fc2b30aa..7bd7fbe9 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -3112,8 +3112,8 @@ orders = ["DEFAULT"] # #397 review: A LINK INSIDE A MAIDEN CLAUSE STAYS IN THE BIRTH NAME # (2026-09-20). LAST in the file, narrow-first: this rule declares # {family, maiden, middle, suffix} and every rule above it that could -# reach either name declares a strict subset, so an earlier position -# would be an order-decided contest the run refuses (#382). +# reach any of the three declares a strict subset, so an earlier +# position would be an order-decided contest the run refuses (#382). # # rules.md#M2 states the exception; decisions.md#M2's 2026-09-20 # bullet carries the measurement and why it is fixed in #397's own PR @@ -3122,7 +3122,7 @@ orders = ["DEFAULT"] [[change]] issue = "fix(#397) a link inside a maiden clause stays in the birth name" -# Two names, one sentence, two shapes -- `fields` is the union. +# Three names, one sentence, three shapes -- `fields` is the union. # # The maiden walk has always ended the birth name at the first suffix # WORD after the marker, and the Catalan link is also the roman @@ -3141,12 +3141,25 @@ issue = "fix(#397) a link inside a maiden clause stays in the birth name" # the tree, 'y' being connective vocabulary and no generation. That # twin is the property invariant in tests/v2/test_properties.py. # -# Literal-anchored to the two. The shape is "a one-letter connective +# The THIRD name arrives from rules.md#M2 rather than from a report: +# 'Smith, John, PhD née Puig Mr. - i Soler' is the doc's +# `deviates: #538` example, and the doc reaches the deviation with a +# configured ' - ' that this corpus does not apply. Parsed with the +# DEFAULT facade, as the gate parses it, the dash is an ordinary name +# word, so the link has one on each side and the clause keeps +# 'i Soler' where this baseline left it in the suffix: suffix +# 'PhD i Soler' -> 'PhD', maiden 'Puig Mr. -' -> 'Puig Mr. - i Soler' +# (measured 2026-09-22 at all four 2.x baselines, identical at each). +# That is this rule's own sentence and no part of #538, whose reading +# needs the delimiter declared; {maiden, suffix} is a subset of the +# union above. +# +# Literal-anchored to the three. The shape is "a one-letter connective # inside a maiden clause", which would stand ready to excuse every # future regression at this walk -- and the walk's whole design is # that some such letters must NOT be taken. _MUST_NOT_MATCH carries # the wall: the i-last control, the generation behind it, and the # spellings where the letter is an initial. -name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler)$" +name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler)$" fields = ["family", "maiden", "middle", "suffix"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index 96a20b85..68047f6e 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -1571,8 +1571,8 @@ orders = ["DEFAULT"] # #397 review: A LINK INSIDE A MAIDEN CLAUSE STAYS IN THE BIRTH NAME # (2026-09-20). LAST in the file, narrow-first: this rule declares # {family, maiden, middle, suffix} and every rule above it that could -# reach either name declares a strict subset, so an earlier position -# would be an order-decided contest the run refuses (#382). +# reach any of the three declares a strict subset, so an earlier +# position would be an order-decided contest the run refuses (#382). # # rules.md#M2 states the exception; decisions.md#M2's 2026-09-20 # bullet carries the measurement and why it is fixed in #397's own PR @@ -1581,7 +1581,7 @@ orders = ["DEFAULT"] [[change]] issue = "fix(#397) a link inside a maiden clause stays in the birth name" -# Two names, one sentence, two shapes -- `fields` is the union. +# Three names, one sentence, three shapes -- `fields` is the union. # # The maiden walk has always ended the birth name at the first suffix # WORD after the marker, and the Catalan link is also the roman @@ -1600,12 +1600,25 @@ issue = "fix(#397) a link inside a maiden clause stays in the birth name" # the tree, 'y' being connective vocabulary and no generation. That # twin is the property invariant in tests/v2/test_properties.py. # -# Literal-anchored to the two. The shape is "a one-letter connective +# The THIRD name arrives from rules.md#M2 rather than from a report: +# 'Smith, John, PhD née Puig Mr. - i Soler' is the doc's +# `deviates: #538` example, and the doc reaches the deviation with a +# configured ' - ' that this corpus does not apply. Parsed with the +# DEFAULT facade, as the gate parses it, the dash is an ordinary name +# word, so the link has one on each side and the clause keeps +# 'i Soler' where this baseline left it in the suffix: suffix +# 'PhD i Soler' -> 'PhD', maiden 'Puig Mr. -' -> 'Puig Mr. - i Soler' +# (measured 2026-09-22 at all four 2.x baselines, identical at each). +# That is this rule's own sentence and no part of #538, whose reading +# needs the delimiter declared; {maiden, suffix} is a subset of the +# union above. +# +# Literal-anchored to the three. The shape is "a one-letter connective # inside a maiden clause", which would stand ready to excuse every # future regression at this walk -- and the walk's whole design is # that some such letters must NOT be taken. _MUST_NOT_MATCH carries # the wall: the i-last control, the generation behind it, and the # spellings where the letter is an initial. -name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler)$" +name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler)$" fields = ["family", "maiden", "middle", "suffix"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index 7df220dd..11b0725b 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -888,8 +888,8 @@ orders = ["DEFAULT"] # #397 review: A LINK INSIDE A MAIDEN CLAUSE STAYS IN THE BIRTH NAME # (2026-09-20). LAST in the file, narrow-first: this rule declares # {family, maiden, middle, suffix} and every rule above it that could -# reach either name declares a strict subset, so an earlier position -# would be an order-decided contest the run refuses (#382). +# reach any of the three declares a strict subset, so an earlier +# position would be an order-decided contest the run refuses (#382). # # rules.md#M2 states the exception; decisions.md#M2's 2026-09-20 # bullet carries the measurement and why it is fixed in #397's own PR @@ -898,7 +898,7 @@ orders = ["DEFAULT"] [[change]] issue = "fix(#397) a link inside a maiden clause stays in the birth name" -# Two names, one sentence, two shapes -- `fields` is the union. +# Three names, one sentence, three shapes -- `fields` is the union. # # The maiden walk has always ended the birth name at the first suffix # WORD after the marker, and the Catalan link is also the roman @@ -917,12 +917,25 @@ issue = "fix(#397) a link inside a maiden clause stays in the birth name" # the tree, 'y' being connective vocabulary and no generation. That # twin is the property invariant in tests/v2/test_properties.py. # -# Literal-anchored to the two. The shape is "a one-letter connective +# The THIRD name arrives from rules.md#M2 rather than from a report: +# 'Smith, John, PhD née Puig Mr. - i Soler' is the doc's +# `deviates: #538` example, and the doc reaches the deviation with a +# configured ' - ' that this corpus does not apply. Parsed with the +# DEFAULT facade, as the gate parses it, the dash is an ordinary name +# word, so the link has one on each side and the clause keeps +# 'i Soler' where this baseline left it in the suffix: suffix +# 'PhD i Soler' -> 'PhD', maiden 'Puig Mr. -' -> 'Puig Mr. - i Soler' +# (measured 2026-09-22 at all four 2.x baselines, identical at each). +# That is this rule's own sentence and no part of #538, whose reading +# needs the delimiter declared; {maiden, suffix} is a subset of the +# union above. +# +# Literal-anchored to the three. The shape is "a one-letter connective # inside a maiden clause", which would stand ready to excuse every # future regression at this walk -- and the walk's whole design is # that some such letters must NOT be taken. _MUST_NOT_MATCH carries # the wall: the i-last control, the generation behind it, and the # spellings where the letter is an initial. -name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler)$" +name_regex = "^(?:Doe, Jane nee Puig i Soler|Jane Doe nee Puig i Soler|Smith, John, PhD née Puig Mr\\. - i Soler)$" fields = ["family", "maiden", "middle", "suffix"] orders = ["DEFAULT"]