Correct SPI income inputs and UK calibration measurements - #879
Conversation
194111a to
4e2e3c2
Compare
|
Completed the matched national-only rebuild and calibration comparison requested in the migration handoff. Calibration loss falls 35.8%, but 98.8% of the net reduction comes from the ONS savings-interest target. The joint income-tax/state-pension/UC calibration problem remains unresolved. The control uses publication commit Errors below are relative to the target: positive means above target, negative means below.
Across all 371 targets, absolute error improves for 181, worsens for 189, and is unchanged for one. Nine targets enter the 10% band and five leave it. The £50–70k state-pension band and private-pension count improve, while aggregate income-tax and state-pension fit deteriorate. These results evaluate the combined SPI treatment; they do not isolate the pension-receipt predictor from income rebasing, the age guard or the OTHERINV correction. The treatment clears the £50–70k state-pension gate failure but introduces three non-deferred target-fit failures:
The £20–30k state-pension failure remains. The private-pension count exclusion is stale in both runs because its error is already within the 25% gate bound; it was preserved for the matched comparison. A22/A23 holds and existing UC/CGT deferrals also remain unchanged. Both spines complete their build gates, but both national calibrations fail the terminal target-fit gates, so neither exports a calibrated dataset. Household count and total prior weight are equal, but the treatment generates 36 fewer people and 21 fewer benefit units, with different final household support and prior-weight vectors. This is a national-only, single-seed comparison and does not establish local-area accuracy or a release-ready UK dataset. The full comparison and reproduction instructions and 371-target paired receipt include target/family loss contributions, source/runtime pins and gate verdicts. This publication-baseline comparison is separate from the historical |
|
Completed the matched national recalibration after correcting the national VOA geography and UC family measurements. This supersedes the interpretation of the earlier comparison. Both runs use the original full spines, UK 2.94.0/Core 3.31.0, the same 371 targets, 1,500 epochs and unchanged solver settings, exclusions and gates. Before optimization, all nine VOA matrix rows matched England-only masks and all 84 active UC band rows matched the corrected family categories. Every underlying UC award and all unrelated matrix rows are unchanged.
The geography correction resolves the apparent Band A problem: the earlier +22.80% / +28.18% estimates mixed Scotland and Wales into an England benchmark. All nine corrected VOA measures now lie within 0.22%. The UC correction exposes an underlying support gap. The former two-row high-payment cell was a duplicated lone-parent case. After relationship-aware classification, both corrected spines have no childless-couple support above £26.4k. Several adjacent bands have only one to four unique source cases. The native UK model's standard-allowance relationship error is separate and remains unresolved by this measurement change. SPI reduces loss 39.8%, with interest accounting for 91.5% of the net gain. Income tax and aggregate state pension still worsen, so this is not uniform improvement. Both runs remain release-blocked: treatment additionally misses self-employment income at £50–70k (+34.485%) and state-pension income at £20–30k (+27.941%), and the unchanged gate flags stale reviewed exclusions. No calibrated dataset was exported. UC inventory: 100 payment-band references = 25 × four family types; 84 active and 16 excluded. With 11 general UC and 15 two-child-limit targets, there are 110 active UC-related targets. The ordinary bands are £100/month wide (£1,200/year). The highest published band is currently materialized as open-ended, an existing mismatch held fixed here. Full corrected comparison · all 371 paired targets and 100-band support audit · complete UC target inventory. Validation: 305 distinct focused tests pass, two existing skips; Ruff, CI inventory and pinned-feed reference regeneration pass. PR remains draft. |
d43b420 to
1c732f7
Compare
4ccaabb to
4f00045
Compare
|
Review pass at What checks outEvery number in the PR body reconciles exactly to Findings1. Should-fix — the calibration's UC family typing diverges from the redraw stage's. The new measurement ( 2. Should-fix — 3. Should-fix — four register entries are adjudicated on the measurement this PR replaces. 4. Should-fix — "remains draft" is false. The PR is not a draft ( 5. Question — the UC matrix-row check is asserted, not evidenced. 6. Question — spine-origin pins are indirect. The corrected receipt never names Smaller questions
Nit
Out of scope and unverifiable#874's content (the engine floor and lock, the Chronicle re-pin, the childcare and bus targets) is reviewed on #874. Reproducing either calibration, the Nothing blocks on correctness, and the receipts are the most honest part of the PR. 1 is the one to settle before this is used as the basis for anything downstream, since two stages now disagree about what a family is; 2 through 4 are cheap. "Release-blocked" is the right framing; "draft" is not. |
vahid-ahmadi
left a comment
There was a problem hiding this comment.
Round one is in the comment above (#879 (comment)): nothing blocking on correctness, four should-fixes (UC family typing diverges from the #835 redraw; requires_uk markers dropped, which is also what the two failing spine-uk CI lanes look like; four register reasons cite the superseded measurement; the PR is not a draft). Filing as a comment review rather than request-changes since the receipts are honest and the PR declares itself release-blocked.
5a1581e to
eb7ca8a
Compare
|
@vahid-ahmadi Thanks for the review. I've pushed the review fixes and rebased onto
The updated report also incorporates Max's measurement-review points. All nine national VOA band/total bindings match England. FRS roles support the composition proxy, but do not establish UC partner eligibility or fix the pinned model's separate standard-allowance relationship error. The UC boundary issue remains open under #736: DWP publishes 26 positive-payment bands per family, including £2,500.01-or-over monthly, whereas the register includes 25. Consequently, the highest included band is bounded in the source but currently absorbs the omitted open tail. Fixing endpoint semantics and handling that separate published tail requires a new matched calibration. The regression explicitly records the current mismatch; it does not certify it as correct. For the smaller points, the report acknowledges that a state-pension uprating index exists but was not selected for this treatment, with its sensitivity unmeasured. Exact FRS numeric-code validation and gate-override provenance remain follow-ups. The #840 evidence is not a new parity waiver or release approval. I left the unrelated Unicode serialization cleanup untouched. Post-rebase validation: 588 passed / 19 skipped on each of Python 3.13 and 3.14 across the affected engine-free tests, and 81 passed / 2 skipped with the pinned UK engine. These are overlapping test scopes. Ruff, formatting, diff checks and CI test inventory checks passed. Hosted CI is still running at posting time, with no completed failures. No calibration was rerun for these review fixes; the reported comparison retains its original code and artifact pins, and the fit gates remain unchanged. |
vahid-ahmadi
left a comment
There was a problem hiding this comment.
Second pass at eb7ca8a5 (Claude Code, high effort; worktree checkout against 478fa88a, the byte-identical rebased replay of the reviewed head per git range-diff; the interdiff is ten files, all within the six dispositions; 577 passed, 19 skipped, 0 failed across 14 in-scope suites without the engine; ruff check clean and ruff format clean on the changed files; both census tools, the coverage manifest and ci_test_groups --verify current).
1. Documented, not aligned — accurate and adequately placed. uc_reporter_redraw.py is untouched; the divergence is stated in the report's "Redraw and nominal-income limitations" section, in the PR body, and in each of the four register annotations, and the statement matches the code. Since #883 is where the redraw and capital-donor predictors move to the shared claimant helper, I am content with this being a declared limitation of the comparison here.
2. Verified. Nine requires_uk markers on the SPI-income tests; the SPI-spine path-equivalence test stubs _spi_income_uprating_factors with deterministic factors and still exercises the full stage transform — on the old head it fails engine-free with ModuleNotFoundError, so the stub is load-bearing. Engine-free run: 39 passed, 14 skipped, 0 failed.
3. Verified. Only the four reason strings changed, each preserving the original text as a prefix; approvers, adjudication ids, approval and expiry dates unchanged. The two-child entry records −21.5% in treatment, inside the bound, and that the annotation renews nothing.
4. Verified. Not a draft; no draft wording in the body.
5. Verified as far as it can be here. The receipt carries 84 exact saved-matrix row checks per run, all exact_saved_matrix_match: true, with the pre-solve and audit scopes distinguished; build_879_corrected_measurement_receipt.py is committed and fails closed on missing artifacts; the tamper and exact-row tests pass; the new full-registry synthetic test covers 100 references, 84 active and 16 excluded, with the open-last-band defect pinned as known. The byte-identical replay needs the restricted artifacts.
6. Verified. Both runs carry spine_origin_code_pin (d43b4203, fc49b482) with the source receipt path, revision, sha and artifact sha.
Three durability nits. The register annotation links point at a branch blob rather than a commit sha; the builder's SOURCE_COMMIT is the pre-rebase head, which is not in main's history; and the receipt's protected-resource sha for the register predates the annotations (consistent with "historical pins preserved", but one sentence in the report would stop a reader treating it as drift).
Approving. Release-blocked stays the right description of the dataset; the code and receipts are honest about it.
…ch (uk-publication-stack-834) The publication stack's schema-2 target-fit register (#796: the four UC with-children cells and the private-pension 100–150k band, expiring 2026-09-30) and its gate wiring did not reach main. This composition branch ports the end state of the stack's three register commits onto the #834 tree: `target_fit_reviewed_exclusions.json` (declared in the country package), `UK_TARGET_FIT_EXCLUSION_REGISTER_RESOURCE`, the lazily loaded default register, the `uk_target_fit_gate` rewrite (in-force entries defer the release fence only; expired, premature, stale and dormant entries are reported and fail where the rot rule says so), the battery binding that resolves the committed register or a loud override and refuses a `gates.json` register name the runtime does not load, the gate's parameter key, the contract's terminal-gate detail fields, and the moved battery and calibration-seam digests. Main's local-candidate branch of the target-fit evaluator is preserved. The stack's seven gate tests and the contract detail pins come with it. Composition branch only: not for merge to main on its own. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… (uk-publication-stack-834) The ported binding evaluates the target-fit register against the run clock, so the gate's required artifacts must name exclusions_evaluated_on as the stack's registry entry did; without it the first composed run failed closed with a KeyError instead of deferring the signed cells. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… (uk-publication-stack-834) The ported target-fit gate evaluates the deferral register against the run clock; the seam's battery artifacts carried no clock, so the composed run reported the gate as evidence-absent. The seam now passes today's date, as the rowwise candidate build passes its start date. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…stack-834, #875) obr.capital_gains_tax@2025 breaches the release bound by +44.2% because the Chronicle re-pin bound HMRC's 2024-25 outturn — the forestalling year ahead of the October 2024 rate rise — at the 2025 period, where 2025-26 rates apply to it. The facts stay bound as published; the release fence defers the miss to 2026-10-05 under microcosm#875, which owns the declared 2024-25 → 2025 translation. The committed-register test carries the sixth entry. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
eb7ca8a to
1f7d676
Compare
Rebased onto main e07a473. The three generated locks that conflicted are regenerated on the combined tree: the coverage manifest's source-manifest sha, the uk spec digest pin, and the graph parity fixture. One semantic merge with #879: it re-uses build_period on impute_uk_spi_income_support to rebase SPI incomes and adds a cached uprating-factor helper, while this branch had dropped that parameter and the functools.cache import on the premise that the disability refresh was their only consumer. Git merged both without a textual conflict and left the module unimportable; the parameter, the call-site argument and the import are restored. The disability refresh itself still takes its year from the declared rule, not from build_period. The stage-level receipt was repeated on the rebased tree and reproduced every figure; its commit rows are updated. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Both sides moved the UK country-bundle digest: main's #879 changed uk/spec/sources.yaml and country_package.json, this branch's solve.py edit moved the seed-protocol attestation inside the same bundle. The rebase conflict was that one line; the combined tree's digest is d57d248f…, read from the pin test on the rebased tree. The am/be digests, the loader golden vector, the seed-protocol and seed-map digests and the committed US coverage report were unchanged by #879 (49 of the 50 pin tests passed before this cut). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…355) The 45,800 smoke on spine-p (the tree rebased onto #879) was refused inside contribution_initialization: "nonzero target at row 20648 has no support", a national row whose constraint row is all zeros on the cloned pool. The dense solve tolerates such a row as a miss; a size selection cannot carry it, and the row index alone is not actionable. refit_uk_dataset_size now lists every nonzero target with no supporting household by name before initialisation and refuses with the list, so the binding defect is fixed upstream or the row is excluded with a signed reason, never selected around. unsupported_nonzero_targets() is the reusable check. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This ports the reviewed SPI income corrections onto the UK publication stack requested in Max's handoff, and corrects England geography and FRS composition measurements that obscured the comparison. National VOA targets previously counted other countries against England facts; generic UC family categories treated some older children as partners.
The current PR head is
1f7d676f2449894bfd794973eaa507aa79bfbfdb, rebased ontomainat289437765a2d5caca97f2974f4d104e9e1e066d5, including merged PRs #874 and #870. The existing PR commits are preserved, with all three durability fixes in one additional commit: immutable report links in the four register annotations, a reachable source-receipt revision in the evidence builder, and explicit historical-register hash documentation. The earlier #874 integration brought659adb1afiscal-year target selection, DfT geography and take-up-contract fixes. The retained calibrations were not rerun for this head; the numerical results below describe their original pinned measurement and spine versions.With the same measurement changes applied to both retained full spines, SPI reduces national loss 39.8%, and targets within 10% rise from 337/371 to 341/371. Both runs remain release-blocked. Interest accounts for 91.5% of the net objective gain; aggregate income-tax and state-pension fit worsen. This conditional comparison does not establish complete administrative measurement correctness or release readiness.
Changes
country == ENGLAND, matching the selectedE92000001facts within the broader England-and-Wales VOA publication. Local VOA and Scottish bindings remain separate; the household-count proxy for dwelling stock is unchanged.Matched national comparison
The complete 2024 control and SPI spines originated at
d43b4203c6ebe10e062cb3ef3034e66731ea055dandfc49b48200e1e15fe35bb4e15b948992bf03c28d. Measurement code117860efe96b4fc1e336e0c9bfab10917cd3aa59calibrates 2025 using UK 2.94.0/Core 3.31.0, Chronicle6fb700e, 371 identical active targets, 1,500 epochs, family-equal weighting, learning rate 0.02, maximum weight ratio 10 and seed 0.The formerly decisive childless-couple cell depended on two duplicates of one source case, now classified as lone parent in both runs. Awards remain unchanged by remeasurement. Corrected childless-couple support is empty from £26.4k upwards in both spines, and adjacent bands have only one to four unique source cases.
Unresolved measurement, support and entitlement work
The registry contains 100 included payment-band references: 25 per family, 84 active and 16 excluded. DWP publishes 26 positive-payment categories for the relevant period, including a separate £2,500.01-or-over monthly tail. This register omits that tail. Its next-lower-edge convention gives the first band [£0.12, £1,200.12), and the highest included, bounded-source band [£28,800.12, infinity). A £36,000 annual award therefore enters a row labelled £2,400.01–£2,500 monthly. Catalog and currency-boundary repair belongs with #736 and requires new matched measurements; saved-matrix equality does not establish published-boundary correctness.
DWP family classification uses awarded standard allowance, including single-rate cases with ineligible partners. Reported children can differ from child-element eligibility. FRS relationship counts remain a proxy. The FRS 2024–25 glossary confirms cohabiting benefit units and dependent 16–19-year-olds, who are excluded from its adult definition; its survey-head designation does not establish UC eligibility. Numeric role codes and adult/child file routing remain unverified. The native standard-allowance relationship error and partner eligibility need separate review; calibration relabelling repairs neither entitlement question.
UC reporter redraw still uses legacy
is_married/qualifying-young-person predictors, itshas_non_child_memberscreen and transition categories. That can classify a young lone claimant aschild_onlywhile calibration usesSINGLE, and can disagree for cohabiting couples. The redraw implementation/settings and existing generated support stay fixed in this comparison. Aligning redraw is unresolved support-generation work.Both runs fail the unchanged target-fit gate and refuse calibrated H5 export. SPI's unreviewed failures are the two empty high-payment UC bands, self-employment at £50–70k (+34.485%) and state pension at £20–30k (+27.941%). The four UC fit-deferral rationales cite historical #813 composition and all require fresh adjudication; the two-child entry is stale in SPI, and the private-pension-count entry is stale in both runs. Each of the four register reasons now preserves its original text verbatim and appends a dated historical-basis annotation linking the report; the two-child annotation explicitly records its stale treatment result. All signatures, dates and operative decisions remain unchanged, and no approval is renewed. The receipt’s
protected_policy_resource_sha256["target_fit_reviewed_exclusions.json"]records the unannotated register used by the historical runs; it is not a checksum of the current annotated register. The other historical receipt hashes also remain the pins of the executed comparisons. No new input-mass parity waiver, known-gap approval or release authorization is supplied. Gate-override observability remains a separate suggestion.Evidence and validation
The historical pre-solve script asserted nine VOA rows against independently derived England masks and 84 UC rows against runtime-resolved families. It did not serialize individual UC exact-match flags. The new builder authenticates the original receipts, retained preflight/matrix/solve artifacts and spine hashes, then adds 84 explicit saved-matrix row audits per run. The private composition snapshots were first hash-pinned during this repair and remain saved resolver output, not independent historical family evidence. Public output contains only allowlisted aggregates and hashes. The independent synthetic test covers all included boundaries, excluded-edge retention and benefit-unit-to-household sums, with the open-last-band defect marked as known behavior.
Historical validation on 7 September with UK 2.94.0/Core 3.31.0 installed covered 305 passes/two skips for the measurement work, then 458 passes/four skips for a later focused rebase set. Historical Ruff, inventory and pinned-feed regeneration checks passed. These are different scopes, neither a full-repository or certification claim.
Review-repair focused checks on 7 September, before the second rebase: the SPI files pass on engine-free Python 3.13 and 3.14 (39 passes/14 skips each), and with the existing UK engine (52 passes/one skip). The independent UC file passes five tests; the builder passes 15 tests, including artifact-tamper refusal, exact-row and privacy checks. Both retained 84-row audits complete without new calibration. The PR targets
main, so hosted CI applies; current hosted results should be read from the checks panel.Historical validation after the 7 September rebase onto
mainbase5ab1b056: the same 13 affected test files pass without country engines on Python 3.13 and 3.14 (588 passes/19 skips in each environment). Five affected files pass with the existing UK 2.94.0/Core 3.31.0 environment (81 passes/two skips). These scopes overlap; the counts are not additive. The earlier 113-file fast group was interrupted for the user-requestedmainupdate and is not a full-suite pass. The completed checks cover code contracts, not a recalibration or dataset certification.Current durability-repair validation at
1f7d676f2449894bfd794973eaa507aa79bfbfdbonmainbase289437765a2d5caca97f2974f4d104e9e1e066d5: 264 tests passed across the corrected-measurement receipt, signed-deferral register, gate-battery pin and data-contract checks; the saved-evidence builder replay matches the updated receipt byte for byte. Source ancestry and blob hashes, exact receipt-locator and report-link preservation checks, Ruff, formatting, diff and CI inventory checks all pass. This validates the repair and integration; it does not rerun calibration or certify a dataset.Refs #665, #736, #796, #809, #840, #866 and #862. The joint-fit problem remains open; CGT tail work in #878 and timing reconciliation in #875 remain separate.