Saved in:
Bibliographic Details
Main Author: Smith, Aiden
Format: Recurso digital
Language:
Published: Zenodo 2026
Online Access:https://doi.org/10.5281/zenodo.18841922
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • <div> <div> <div> <div> <div> <div> <div> <div> <h1>Hubble-Systematics-Review-Chain</h1> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#hubble-systematics-review-chain"></a></div> <p>Systematics-first, <em>audit-style</em> pipeline for late-time distance-scale probes, patterned on:</p> <ul> <li>stability scans (cuts/thresholds) + correlated-cut drift nulls</li> <li>mechanism ladders (low-dimensional knobs → structured residual closure)</li> <li>injection/recovery suites</li> <li>SBC/coverage gates</li> <li>time/epoch and sky (low-ℓ) invariance tests</li> </ul> <div> <h2>Theory Snapshot (Causal-Window Cycle Cosmology)</h2> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#theory-snapshot-causal-window-cycle-cosmology"></a></div> <p>Working hypothesis:</p> <ul> <li>The Hubble mismatch can be produced by a structured <strong>inference-geometry bias</strong> (a phase-conditioned causal-window mapping), not only by random error or one bad catalog.</li> <li>A valid mechanism must move inferred expansion by tension-scale amounts in long-range channels <strong>without</strong> forcing control channels to move.</li> <li>Claims are accepted only after holdouts, identifiability gates, and control-channel preservation checks.</li> </ul> <div> <h2>TL;DR (Reviewers / Referees)</h2> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#tldr-reviewers--referees"></a></div> <ul> <li>This repository is designed for <strong>artifact-traceable, closure-first</strong> testing of late-time inference-bias mechanisms.</li> <li>If you want one seeded, one-shot Hubble-tension packet command that outputs headline claims and tables, run:</li> </ul> <div> <pre>bash scripts/run_one_hubble_tension_seeded.sh 20260301</pre> <div> </div> </div> <ul> <li>Optional CPU pinning example:</li> </ul> <div> <pre>bash scripts/run_one_hubble_tension_seeded.sh 20260301 0-5</pre> <div> </div> </div> <ul> <li>The seeded command launches the event-level hierarchical reviewer packet, waits for completion, and writes: <ul> <li><code>cross_probe_lockin_scoreboard.{json,md}</code></li> <li><code>cross_probe_non_dotg_scoreboard.{json,md}</code></li> <li><code>cross_probe_headlines.{json,md}</code> (review headline table)</li> </ul> </li> <li>Default behavior auto-selects a conservative CPU set (up to 6 cores).</li> <li>Typical runtime on 6 CPUs is order ~1.5-2.5 hours depending on machine/load.</li> </ul> <div> <h3>Current Headline Snapshot (Latest Completed Queue)</h3> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#current-headline-snapshot-latest-completed-queue"></a></div> <p>Source queue: <code>outputs/cross_probe_lockin_packet_v2_20260301_102852UTC</code></p> <table> <tbody><tr> <th>Metric</th> <th>Value</th> </tr> </tbody><tbody> <tr> <td>Full holdout win-rate (non-baseline)</td> <td><code>6/8 (75.0%)</code></td> </tr> <tr> <td>Control channels: baseline wins</td> <td><code>2/2 (100.0%)</code></td> </tr> <tr> <td>Non-dotG primary win-rate (non-baseline)</td> <td><code>6/6 (100.0%)</code></td> </tr> <tr> <td>Projection mean_up at z=0.1 / 0.3 / 0.5</td> <td><code>4.9610 / 5.9021 / 6.6502</code></td> </tr> <tr> <td>Mechanism constraints-compatible</td> <td><code>gamma1: False, gammafree: True</code></td> </tr> <tr> <td>Event-level siren H0_eff ± sd</td> <td><code>60.8361 ± 12.7347</code></td> </tr> </tbody> </table> <div> <h3>Method-Strengthening Snapshot (Latest Completed Packet)</h3> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#method-strengthening-snapshot-latest-completed-packet"></a></div> <p>Source queue: <code>outputs/method_strengthening_packet_20260302_193531UTC</code></p> <table> <tbody><tr> <th>Track</th> <th>Status</th> <th>Key result</th> </tr> </tbody><tbody> <tr> <td>SBC + blinded injections</td> <td><code>PASS</code></td> <td><code>coverage68_min=0.6914</code>, <code>coverage95_min=0.9766</code>, misspec/modeled mean|bias| ratio=<code>1.499</code></td> </tr> <tr> <td>Event jackknife influence</td> <td><code>PASS</code></td> <td>max `</td> </tr> <tr> <td>Two-codepath hierarchical replication</td> <td><code>PASS</code></td> <td>`</td> </tr> <tr> <td>External margin/power</td> <td><code>PASS</code></td> <td>pair-pass shift band <code>[0.000, 10.000]</code>, model shift <code>5.902</code>, distance-to-failure <code>4.098</code></td> </tr> <tr> <td>Out-of-sample mini-tournaments</td> <td><code>PASS</code></td> <td>primary non-dotG <code>6/6</code>, controls <code>2/2</code>, Item4 CID/survey holdout winners non-baseline</td> </tr> <tr> <td>One-parameter joint fit</td> <td><code>PASS</code></td> <td><code>delta_hat=3.645±1.238</code>, model shift <code>5.902</code>, <code>z=1.824</code></td> </tr> </tbody> </table> <p>Reproduce:</p> <div> <pre>bash scripts/run_one_method_strengthening_packet.sh</pre> <div> </div> </div> <div> <h3>Core Tier Snapshot (SN + BAO + Geometric Poles + Growth Pressure)</h3> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#core-tier-snapshot-sn--bao--geometric-poles--growth-pressure"></a></div> <p>Source queue: <code>outputs/core_tier_packet_20260302_204243UTC</code></p> <table> <tbody><tr> <th>Check</th> <th>Result</th> </tr> </tbody><tbody> <tr> <td>External baseline consistency (expanded non-siren poles)</td> <td><code>PASS</code></td> </tr> <tr> <td>External robustness (non-stress scenarios)</td> <td><code>10/10</code></td> </tr> <tr> <td>SN Ia Hubble-flow holdout (Pantheon+)</td> <td><code>winner=+bounded_fields_plus_metadata_bounded</code></td> </tr> <tr> <td>SN Ia ladder holdout (Pantheon+SH0ES)</td> <td><code>winner=+cal_offset_bounded</code></td> </tr> <tr> <td>BAO holdout (DESI/SDSS stack part)</td> <td><code>winner=+bounded_fields_plus_metadata_bounded</code></td> </tr> <tr> <td>External H0-grid control holdout</td> <td><code>winner=baseline</code></td> </tr> <tr> <td>Growth/lensing pressure checks (gamma1 + gammafree)</td> <td><code>winner=+bounded_fields_plus_metadata_bounded</code> (both)</td> </tr> <tr> <td>Chronometers control holdout</td> <td><code>winner=baseline</code></td> </tr> </tbody> </table> <p>Reproduce:</p> <div> <pre>bash scripts/run_one_core_tier_packet.sh</pre> <div> </div> </div> <p>Artifacts:</p> <ul> <li><code>outputs/core_tier_packet_*/core_tier_packet_summary.{json,md}</code></li> <li><code>outputs/core_tier_packet_*/external_full_packet/external_probe_consistency_full_packet.{json,md}</code></li> </ul> <div> <h3>Exploratory Extension Snapshot (FRB + CMB-Compressed Low Pole)</h3> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#exploratory-extension-snapshot-frb--cmb-compressed-low-pole"></a></div> <p>Recent completed packets:</p> <ul> <li><code>outputs/external_probe_consistency_full_packet_20260302_213014UTC</code> (FRB stat-only, CMB-compressed low pole)</li> <li><code>outputs/external_probe_consistency_full_packet_20260302_213002UTC</code> (FRB systematics-aware with sigma floor 2.0, CMB-compressed low pole)</li> </ul> <table> <tbody><tr> <th>Packet</th> <th>Primary baseline</th> <th>Stress baseline</th> <th>Robustness</th> </tr> </tbody><tbody> <tr> <td>FRB stat-only (sigma=0.69)</td> <td><code>PASS</code></td> <td><code>FAIL</code> (expected stress fail)</td> <td><code>13/13</code></td> </tr> <tr> <td>FRB systematics-aware (sigma floor=2.0)</td> <td><code>PASS</code></td> <td><code>FAIL</code> (expected stress fail)</td> <td><code>13/13</code></td> </tr> </tbody> </table> <p>Pair-coverage interpretation:</p> <ul> <li>Coverage here is thresholded by prereg config (<code>min_pairs_total=4</code>, <code>min_pairs_pass=3</code>, <code>min_pairs_pass_fraction=0.5</code>), not a requirement for 100% pair pass.</li> <li>Latest runs are fully scorable (<code>unscorable=0</code>); non-pass pairs are explicit <code>FAIL_threshold</code> rows, not hidden missing-data rows.</li> <li>Two-track FRB status: <ul> <li>stat-only FRB pole gives <code>18/21</code> with three explicit low-pole fails.</li> <li>systematics-aware FRB pole (sigma floor 2.0) gives <code>21/21</code> with no explicit fails.</li> </ul> </li> <li>Pair-reason reports: <ul> <li><code>outputs/external_probe_consistency_full_packet_20260302_213014UTC/pair_reasons_primary_baseline.md</code></li> <li><code>outputs/external_probe_consistency_full_packet_20260302_213002UTC/pair_reasons_primary_baseline.md</code></li> </ul> </li> </ul> <p>FRB systematics-floor sweep summary:</p> <ul> <li><code>outputs/frb_systematics_floor_sweep_20260302_212226UTC/frb_systematics_floor_sweep.md</code></li> <li>First full pair-pass floor in this sweep: <code>sigma=2.0</code>.</li> </ul> <p>Reproduce FRB stat-only packet:</p> <div> <pre>bash scripts/run_one_external_probe_consistency_full_packet.sh 0-5 \ outputs/cross_probe_lockin_packet_v2_20260301_102852UTC \ outputs/reviewer_headline_packet_hier_v1_20260301_185221UTC \ configs/preregistered_external_consistency/external_probe_consistency_scorecard_exploratory_cmbfrb_statonly_preregister_20260302.json</pre> <div> </div> </div> <p>Reproduce FRB systematics-aware packet:</p> <div> <pre>bash scripts/run_one_external_probe_consistency_full_packet.sh 0-5 \ outputs/cross_probe_lockin_packet_v2_20260301_102852UTC \ outputs/reviewer_headline_packet_hier_v1_20260301_185221UTC \ configs/preregistered_external_consistency/external_probe_consistency_scorecard_exploratory_cmbfrb_sysfloor2_preregister_20260302.json</pre> <div> </div> </div> <p>Reproduce FRB systematics-floor sweep:</p> <div> <pre>bash scripts/run_one_frb_systematics_floor_sweep.sh</pre> <div> </div> </div> <div> <h2>Reviewer Quickstart (Single Command)</h2> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#reviewer-quickstart-single-command"></a></div> <p>From repo root:</p> <div> <pre>bash scripts/run_one_hubble_tension_seeded.sh 20260301</pre> <div> </div> </div> <p>Event-level hierarchical variant:</p> <div> <pre>bash scripts/run_one_reviewer_headlines_hier_v1.sh</pre> <div> </div> </div> <p>Non-blocking (detached) alternative:</p> <div> <pre>bash scripts/launch_reviewer_headline_packet.sh</pre> <div> </div> </div> <p>External-probe consistency scorecard (triangulation track, separate from strict R4):</p> <div> <pre>bash scripts/run_one_external_probe_consistency_scorecard.sh</pre> <div> </div> </div> <p>Full external-consistency packet (primary + stress + robustness sweeps):</p> <div> <pre>bash scripts/run_one_external_probe_consistency_full_packet.sh</pre> <div> </div> </div> <p>Core-tier packet (expanded non-siren geometric poles + SN/BAO/growth summary):</p> <div> <pre>bash scripts/run_one_core_tier_packet.sh</pre> <div> </div> </div> <p>Method-strengthening full packet (all six tracks):</p> <div> <pre>bash scripts/run_one_method_strengthening_packet.sh</pre> <div> </div> </div> <p>Then monitor:</p> <div> <pre>OUT=<span><span>$(</span>ls -td outputs/reviewer_headline_packet_<span>*</span> <span>|</span> head -1<span>)</span></span> tail -n 60 <span><span>"</span><span>$OUT</span>/run.log<span>"</span></span></pre> <div> </div> </div> <p>Optional CPU pinning override (example for 4 cores):</p> <div> <pre>bash scripts/run_one_reviewer_headlines.sh 0-3</pre> <div> </div> </div> <div> <h2>Wide-Binary Status (Maintenance Constraint Channel)</h2> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#wide-binary-status-maintenance-constraint-channel"></a></div> <p>Wide-binary packets are now treated as a <strong>maintenance constraint/falsification channel</strong>. Current strict closure runs prefer near-baseline gravity and are used to bound over-strong MG claims, not as primary positive evidence.</p> <p>Current focus for decision-grade work remains:</p> <ul> <li>event-level hierarchical dark-siren analyses,</li> <li>cross-probe lock-in with control-channel preservation,</li> <li>out-of-sample predictive checks versus alternatives.</li> </ul> <p>Conclusion note:</p> <ul> <li><code>docs/WIDE_BINARY_CONCLUSION.md</code></li> </ul> <div> <h2>Wide-Binary MG Proxy Test (Gaia DR3)</h2> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#wide-binary-mg-proxy-test-gaia-dr3"></a></div> <p>This repo now includes a preliminary wide-binary modified-gravity <strong>proxy</strong> scan (trend-level only; not a full orbital forward-model adjudication):</p> <div> <pre>bash scripts/run_one_wide_binary_mg_proxy.sh</pre> <div> </div> </div> <p>Primary config:</p> <ul> <li><code>configs/wide_binary/wide_binary_mg_proxy_v1.yaml</code></li> </ul> <p>Primary outputs (timestamped run dir under <code>outputs/wide_binary_mg_proxy_v1/</code>):</p> <ul> <li><code>summary.json</code></li> <li><code>report.md</code></li> <li><code>binned_profile.csv</code></li> <li><code>wide_binary_mg_proxy_profile.png</code></li> </ul> <div> <h2>Wide-Binary Decision Packet (Gaia DR3)</h2> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#wide-binary-decision-packet-gaia-dr3"></a></div> <p>Decision-style wide-binary packet with:</p> <ul> <li>multi-profile robustness checks,</li> <li>acceleration-label permutation null,</li> <li>GR/MG injection calibration,</li> <li>sky-holdout predictive checks,</li> <li>explicit pass/fail decision gates.</li> </ul> <p>Run in foreground:</p> <div> <pre>bash scripts/run_one_wide_binary_decision_packet.sh</pre> <div> </div> </div> <p>Detached launch (recommended for longer packet settings):</p> <div> <pre>bash scripts/launch_wide_binary_decision_packet_single_nohup.sh 0-5</pre> <div> </div> </div> <p>Primary config:</p> <ul> <li><code>configs/wide_binary/wide_binary_decision_packet_v1.yaml</code></li> </ul> <p>Primary outputs (timestamped run dir under <code>outputs/wide_binary_decision_packet_v1/</code>):</p> <ul> <li><code>summary.json</code></li> <li><code>report.md</code></li> <li><code>progress.json</code></li> <li><code>profile_consistency.png</code></li> </ul> <div> <h2>Wide-Binary Decision Packet v2 (Contamination-Aware)</h2> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#wide-binary-decision-packet-v2-contamination-aware"></a></div> <p>This upgraded packet uses a pair-level contamination-aware likelihood with:</p> <ul> <li>clean+contamination mixture model,</li> <li>multi-profile robustness checks,</li> <li>bidirectional sky holdouts,</li> <li>RV-blind holdout,</li> <li>GR/MG injection calibration.</li> </ul> <p>Run in foreground:</p> <div> <pre>bash scripts/run_one_wide_binary_decision_packet_v2.sh</pre> <div> </div> </div> <p>Detached launch:</p> <div> <pre>bash scripts/launch_wide_binary_decision_packet_v2_single_nohup.sh 0-5</pre> <div> </div> </div> <p>Primary config:</p> <ul> <li><code>configs/wide_binary/wide_binary_decision_packet_v2.yaml</code></li> </ul> <p>Primary outputs (timestamped run dir under <code>outputs/wide_binary_decision_packet_v2/</code>):</p> <ul> <li><code>summary.json</code></li> <li><code>report.md</code></li> <li><code>progress.json</code></li> <li><code>profile_delta_bic.png</code></li> </ul> <p>Latest completed run:</p> <ul> <li><code>outputs/wide_binary_decision_packet_v2/run_20260302_003307UTC</code></li> </ul> <div> <h2>Wide-Binary Orbit Closure v3 (Orbit-Level + Replication)</h2> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#wide-binary-orbit-closure-v3-orbit-level--replication"></a></div> <p>Literature-grade escalation packet that adds:</p> <ul> <li>orbit-level forward simulation (phase + orientation + eccentricity),</li> <li>explicit contamination templates (triples, chance alignments, flybys),</li> <li>independent-catalog replication (PS2023 + El-Badry eDR3),</li> <li>sky and RUWE holdouts,</li> <li>injection-based false-positive / recovery gates.</li> </ul> <p>Run in foreground:</p> <div> <pre>bash scripts/run_one_wide_binary_orbit_closure_v3.sh</pre> <div> </div> </div> <p>Detached launch:</p> <div> <pre>bash scripts/launch_wide_binary_orbit_closure_v3_single_nohup.sh 0-5</pre> <div> </div> </div> <p>Primary config:</p> <ul> <li><code>configs/wide_binary/wide_binary_orbit_closure_v3.yaml</code></li> </ul> <p>Primary outputs (timestamped run dir under <code>outputs/wide_binary_orbit_closure_v3/</code>):</p> <ul> <li><code>summary.json</code></li> <li><code>report.md</code></li> <li><code>progress.json</code></li> <li><code><catalog>_density_fit.png</code></li> </ul> <p>Latest completed full run:</p> <ul> <li><code>outputs/wide_binary_orbit_closure_v3/run_20260302_011118UTC</code></li> </ul> <div> <h2>Wide-Binary Orbit Closure v4 (Free-Delta Replication)</h2> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#wide-binary-orbit-closure-v4-free-delta-replication"></a></div> <p>Strict replication-grade packet that keeps orbit-level contamination closure and independent-catalog replication, but treats MG amplitude as a free parameter (BIC-penalized) instead of fixing it to one value.</p> <p>Run in foreground:</p> <div> <pre>bash scripts/run_one_wide_binary_orbit_closure_v4.sh</pre> <div> </div> </div> <p>Detached launch:</p> <div> <pre>bash scripts/launch_wide_binary_orbit_closure_v4_single_nohup.sh 0-5</pre> <div> </div> </div> <p>Primary config:</p> <ul> <li><code>configs/wide_binary/wide_binary_orbit_closure_v4_free_delta.yaml</code></li> </ul> <p>Primary outputs (timestamped run dir under <code>outputs/wide_binary_orbit_closure_v4/</code>):</p> <ul> <li><code>summary.json</code></li> <li><code>report.md</code></li> <li><code>progress.json</code></li> <li><code><catalog>_density_fit.png</code></li> </ul> <p>Latest completed full run:</p> <ul> <li><code>outputs/wide_binary_orbit_closure_v4/run_20260302_014328UTC</code></li> </ul> <p>Current strict-closure interpretation:</p> <ul> <li>best-fit MG amplitude collapses near zero (<code>mg_delta_v2_best=0.02</code>),</li> <li>GR is preferred in both catalogs (<code>deltaBIC(GR-MG) < 0</code>).</li> </ul> <div> <h2>Produced Headline Artifacts</h2> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#produced-headline-artifacts"></a></div> <p>Each reviewer run creates <code>outputs/reviewer_headline_packet_<timestamp>/</code>.</p> <table> <tbody><tr> <th>Artifact</th> <th>Purpose</th> <th>Primary fields</th> </tr> </tbody><tbody> <tr> <td><code>cross_probe_lockin_scoreboard.md</code></td> <td>Full cross-probe outcomes</td> <td>per-holdout winner, evidence deltas, projection/mechanism summaries</td> </tr> <tr> <td><code>cross_probe_non_dotg_scoreboard.md</code></td> <td>Excludes dotG-like controls from primary tally</td> <td>control-channel baseline wins, non-dotG win-rate</td> </tr> <tr> <td><code>cross_probe_headlines.md</code></td> <td>Referee headline table</td> <td>full win-rate, control win-rate, non-dotG win-rate, projection means, mechanism gates</td> </tr> <tr> <td><code>cross_probe_headlines.json</code></td> <td>Machine-readable headline payload</td> <td>same headline metrics as markdown for downstream checks</td> </tr> <tr> <td><code>config_manifest.txt</code></td> <td>Run manifest</td> <td>configs, run structure, intended outputs</td> </tr> <tr> <td><code>env.json</code></td> <td>Execution environment snapshot</td> <td>selected env vars, python version, hostname</td> </tr> <tr> <td><code>git_like.json</code></td> <td>Repo state snapshot</td> <td>branch, commit hash, short git status</td> </tr> </tbody> </table> <div> <h2>Headline Table (Latest Completed Queue Snapshot)</h2> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#headline-table-latest-completed-queue-snapshot"></a></div> <p>Source queue: <code>outputs/cross_probe_lockin_packet_v2_20260301_102852UTC</code></p> <table> <tbody><tr> <th>Headline</th> <th>Value</th> </tr> </tbody><tbody> <tr> <td>Full holdout win-rate (non-baseline)</td> <td><code>6/8 (75.0%)</code></td> </tr> <tr> <td>Control channels: GR baseline wins</td> <td><code>2/2 (100.0%)</code></td> </tr> <tr> <td>Non-dotG primary win-rate (non-baseline)</td> <td><code>6/6 (100.0%)</code></td> </tr> <tr> <td>Projection mean_up at z=0.1 / 0.3 / 0.5</td> <td><code>4.9610 / 5.9021 / 6.6502</code></td> </tr> <tr> <td>Mechanism gamma1: plausible, constraints-compatible</td> <td><code>True, False</code></td> </tr> <tr> <td>Mechanism gammafree: plausible, constraints-compatible</td> <td><code>True, True</code></td> </tr> <tr> <td>Event-level siren H0_eff ± sd</td> <td><code>60.8361 ± 12.7347</code></td> </tr> </tbody> </table> <p>Generated artifacts for that queue:</p> <ul> <li><code>outputs/cross_probe_lockin_packet_v2_20260301_102852UTC/cross_probe_headlines.md</code></li> <li><code>outputs/cross_probe_lockin_packet_v2_20260301_102852UTC/cross_probe_non_dotg_scoreboard.md</code></li> <li><code>outputs/cross_probe_lockin_packet_v2_20260301_102852UTC/cross_probe_lockin_scoreboard.md</code></li> </ul> <div> <h2>Cross-Probe Packet Scope (Reviewer Run)</h2> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#cross-probe-packet-scope-reviewer-run"></a></div> <p>The reviewer packet includes:</p> <ul> <li>stack holdouts: external-<code>h0_grid</code>, <code>siren_gate2</code>, <code>bao</code>, <code>chronometers</code>, <code>pantheon_plus</code>, <code>pantheon_plus_shoes_ladder</code></li> <li>growth/lensing holdout pair with explicit <code>delta_gamma</code> handling (fixed and free)</li> <li>event-level siren audit</li> <li>real-siren projection pivot triad (<code>z_pivot=0.1/0.3/0.5</code>)</li> <li>mechanism gamma sensitivity pair (fixed-gamma and gamma-free-envelope runs)</li> </ul> <p>Design docs:</p> <ul> <li><code>docs/cross_probe_lockin/CROSS_PROBE_LOCKIN_MATRIX_20260301.md</code></li> <li><code>docs/cross_probe_lockin/CROSS_PROBE_FULL_DESIGNS_20260301.md</code></li> <li><code>docs/cross_probe_lockin/CROSS_PROBE_LOCKIN_EXECUTION_20260301.md</code></li> <li>templates: <code>configs/cross_probe_lockin_templates/*.yaml.template</code></li> </ul> <div> <h2>Current findings (real data; 2026-02-06)</h2> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#current-findings-real-data-2026-02-06"></a></div> <p>These are <strong>real-data</strong> runs on the public Pantheon+ / Pantheon+SH0ES tables in <code>data/raw/…</code>, using this repo’s <strong>linear-Gaussian audit models</strong> (not a full end-to-end SH0ES reanalysis).</p> <div> <h2>Causal-window V3 track (development; 2026-03-01)</h2> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#causal-window-v3-track-development-2026-03-01"></a></div> <p>Working hypothesis for V3 kernel studies in this repo:</p> <ul> <li>treat the <code>xi0 < 1</code> (up/positive) branch as the primary branch;</li> <li>retain <code>xi0 > 1</code> points as sparse falsification controls (not equal-prior scans);</li> <li>require closure and external-gate checks before any decision-grade claim.</li> </ul> <p>Dark-siren caution (applies to all claims in this repo):</p> <ul> <li>Dark-siren inference is high-sensitivity data analysis (selection normalization, host ambiguity, event-level weighting).</li> <li>Posterior-remap or projection results are treated as mechanism-capability tests, not direct expansion-history proof.</li> <li>Decision-grade statements require event-level quality gates, at minimum: <ul> <li><code>gate2_pass == true</code></li> <li>sufficient event count for the packet objective</li> <li>no extreme low-ESS domination in included events</li> <li>replication across independent split/control families</li> </ul> </li> </ul> <p>Current non-decision-grade indication from <code>CLOSURE_PAPER/tables/v3_gate_summary.md</code>:</p> <ul> <li>by-<code>xi0</code> means in the V3 middev run show <code>xi0=0.92 -> mean_bias_induced ~ +6.44</code>, <code>xi0=1.08 -> mean_bias_induced ~ -5.27</code>, with <code>mean_h0_correct</code> near anchor scale (<code>~68.8-69.1</code> across run packets).</li> </ul> <p>Real-data projection snapshot (Gate-2 siren posterior remap; 2026-03-01):</p> <ul> <li><code>z_pivot=0.3</code> (<code>n=210</code> grid points): <code>mean_induced_overall ~ +3.87</code>, <code>mean_induced_up ~ +5.80</code>, <code>frac_ge_5_up ~ 0.60</code>, best case <code>~ +11.32</code>.</li> <li>pivot sensitivity on up branch: <code>mean_induced_up ~ +4.73</code> (<code>z=0.1</code>) to <code>~ +6.64</code> (<code>z=0.5</code>).</li> </ul> <p>One-command reproduction (from repo root):</p> <div> <pre>bash scripts/run_one_real_siren_projection.sh</pre> <div> </div> </div> <p>Optional profile/pivot:</p> <div> <pre>bash scripts/run_one_real_siren_projection.sh fastdev 0.3 bash scripts/run_one_real_siren_projection.sh 2h 0.5</pre> <div> </div> </div> <p>Recommended configs moving forward:</p> <ul> <li> <p>fast development: <code>configs/ruler_mismatch_mechanism_study_v3_causal_window_kernel_posbranch_fastdev.yaml</code></p> </li> <li> <p>larger dev run (~2h target on 6 CPUs): <code>configs/ruler_mismatch_mechanism_study_v3_causal_window_kernel_posbranch_2h.yaml</code></p> </li> <li> <p>branch-balanced legacy reference: <code>configs/ruler_mismatch_mechanism_study_v3_causal_window_kernel_middev.yaml</code></p> </li> <li> <p>queued launcher (detached, one-at-a-time): <code>CLOSURE_PAPER/CAUSAL_WINDOW_THEORY/launch_causal_window_v3_posbranch_queue.sh</code></p> </li> <li> <p>projection configs used by the one-command runner: <code>configs/ruler_mismatch_real_siren_projection_posbranch_fastdev.yaml</code>, <code>configs/ruler_mismatch_real_siren_projection_posbranch_2h.yaml</code></p> </li> <li> <p><strong>Ladder reproduction:</strong> <code>H0_eff = 73.552 ± 1.078</code> vs anchor <code>67.4</code> (equivalent <code>|Δμ| ≈ 0.19 mag</code>).<br>Report: <code>outputs/pantheon_plus_shoes_ladder_predictive_score_v2/report.md</code><br>Reproduce: <code>configs/pantheon_plus_shoes_ladder_predictive_score_v2.yaml</code></p> </li> <li> <p><strong>Late-time non-ladder stack:</strong> Pantheon+ (SN-only) + DESI BAO + CC gives <code>H0_eff = 68.208 ± 0.314</code> (close to anchor).<br>Report: <code>outputs/stack_sn_bao_cc_stress_v2/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_stress.yaml</code></p> </li> <li> <p><strong>Anchor-consistency (no sirens):</strong> adding the SH0ES ladder subset forces a large calibrator↔HF offset:<br><code>calibrator_offset_mag = +0.1616 ± 0.0318</code> mag.<br>Report: <code>outputs/stack_sn_bao_cc_plus_ladder_cal_offset_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_cal_offset_v1.yaml</code></p> </li> <li> <p><strong>Ladder calibrator holdout (CV; real data):</strong> holding out calibrators (keeping hubble-flow fixed in TRAIN) shows a large out-of-sample improvement from calibrator-only mechanisms:<br>Δlogp ≈ +10.2 / +8.5 (diag/full-cov) for <code>calibrator_offset_mag</code>, and Δlogp ≈ +8.3 / +6.1 for calibrator time-bin offsets (<code>pkmjd_bins</code> on calibrators).<br>Survey-by-survey holdout also improves (Δlogp ≈ +3.8 / +3.2 for <code>calibrator_offset_mag</code>).<br>In these random holdout splits, <code>calibrator_offset_mag</code> beats the baseline in <strong>99.5–100%</strong> of splits (depending on diag/fullcov).<br>Reports: <code>outputs/pantheon_plus_shoes_ladder_predictive_score_cal_holdout_v1/report.md</code>, <code>outputs/pantheon_plus_shoes_ladder_predictive_score_cal_holdout_fullcov_v1/report.md</code>, <code>outputs/pantheon_plus_shoes_ladder_predictive_score_cal_survey_holdout_v1/report.md</code>, <code>outputs/pantheon_plus_shoes_ladder_predictive_score_cal_survey_holdout_fullcov_v1/report.md</code><br>Reproduce: <code>configs/pantheon_plus_shoes_ladder_predictive_score_cal_holdout_v1.yaml</code>, <code>configs/pantheon_plus_shoes_ladder_predictive_score_cal_holdout_fullcov_v1.yaml</code>, <code>configs/pantheon_plus_shoes_ladder_predictive_score_cal_survey_holdout_v1.yaml</code>, <code>configs/pantheon_plus_shoes_ladder_predictive_score_cal_survey_holdout_fullcov_v1.yaml</code></p> </li> <li> <p><strong>Joint-stack calibrator holdout (real data):</strong> in the <em>anchor-consistency</em> joint stack (SN-only+BAO+CC+ladder), holding out calibrators (keeping hubble-flow + other probes fixed in TRAIN) still prefers calibrator-only corrections:<br>Δlogp ≈ +5.7 / +6.2 (diag/full-cov) for <code>calibrator_offset_mag</code>.<br>In these random holdout splits, <code>calibrator_offset_mag</code> beats the baseline in <strong>93.5–99%</strong> of splits (depending on diag/fullcov).<br>A simple calibrator-only proxy <code>pkmjd_err_linear</code> (linear in <code>PKMJDERR</code>) also improves held-out calibrators (Δlogp ≈ +4.6), while a <code>m_b_corr_err_VPEC</code> linear proxy does not.<br>Survey-holdout inside the same joint stack is smaller but still improves (Δlogp ≈ +1.9 / +2.3 for calibrator time-bin offsets; +1.7 / +2.3 for <code>calibrator_offset_mag</code>).<br>Reports: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_fullcov_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_mechanism_scan_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_survey_holdout_mechanism_scan_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_survey_holdout_mechanism_scan_fullcov_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_fullcov_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_mechanism_scan_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_survey_holdout_mechanism_scan_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_survey_holdout_mechanism_scan_fullcov_v1.yaml</code></p> </li> <li> <p><strong>External-prior stress test:</strong> forcing a tight prior on <code>calibrator_offset_mag</code> degrades evidence and shifts the fit (does <em>not</em> recover <code>H0≈73</code>).<br>Report: <code>outputs/stack_sn_bao_cc_plus_ladder_cal_offset_tight_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_cal_offset_tight_v1.yaml</code></p> </li> <li> <p><strong>External-prior “gate” sweep (real data):</strong> if external calibration work can truly bound calibrator-step distortions at the ~0.02 mag level, the joint anchor-consistency fit pays a large evidence penalty and cannot maintain the large ~0.16 mag calibrator↔HF offset.<br>Report: <code>outputs/stack_sn_bao_cc_plus_ladder_external_prior_gates_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_external_prior_gates_v1.yaml</code></p> </li> <li> <p><strong>Covariance-implied calibration bounds (real data):</strong> we can also derive “budget-like” bounds for per-survey and per-epoch calibrator offsets from the published Pantheon+SH0ES STAT+SYS covariance. Under these bounds (typical σ≈0.02–0.05 mag), the joint fit cannot sustain a 0.16 mag calibrator↔HF step without a large evidence penalty.<br>Report: <code>outputs/stack_sn_bao_cc_plus_ladder_cov_implied_gates_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_cov_implied_gates_v1.yaml</code><br>Derivation: <code>scripts/derive_pantheon_shoes_cov_priors.py</code></p> </li> <li> <p><strong>External calibration covariance (Brout+21 “FRAGILISTIC”; real data):</strong> using the public zeropoint-offset covariance shipped with the Pantheon+ DataRelease (typical per-survey σ≈0.002–0.008 mag), constraining per-survey offsets does <strong>not</strong> remove the need for a large calibrator↔HF step (the fit still wants <code>calibrator_offset_mag ≈ 0.16</code> when allowed). Calibrator holdout predictive scoring still prefers <code>calibrator_offset_mag</code> even under these tight survey-calibration priors.<br>Reports: <code>outputs/stack_sn_bao_cc_plus_ladder_fragilistic_gates_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_fragilistic_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_fragilistic_gates_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_fragilistic_v1.yaml</code><br>Derivation: <code>scripts/derive_pantheon_shoes_fragilistic_priors.py</code></p> </li> <li> <p><strong>FRAGILISTIC (filter-level; real data; new):</strong> we now use <code>FRAGILISTIC_COVARIANCE.npz</code> <strong>directly</strong> as a correlated prior over <em>filter-level</em> zeropoint parameters (no per-survey σ compression). In the joint stack, filter-level calibration alone does <strong>not</strong> move the anchor-consistency solution toward SH0ES (tension-reduction frac ≈ 0), and the preferred calibrator offset remains ~0.16 mag when allowed. Under the SH0ES linear-system prior (σ≈0.028 mag on <code>calibrator_offset_mag</code>), the supported tension-reduction fraction remains ≈0.15—consistent with the earlier compressed-FRAGILISTIC result.<br>Report: <code>outputs/stack_sn_bao_cc_plus_ladder_fragilistic_filter_shoes_gates_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_fragilistic_filter_shoes_gates_v1.yaml</code><br>Code: <code>src/hubble_systematics/external_calibration/fragilistic.py</code></p> </li> <li> <p><strong>Calibration-only covariance gate (Pantheon+SH0ES <code>CALIB.cov</code>; real data):</strong> using the Pantheon+SH0ES <em>calibration-only</em> covariance grouping (<code>sytematic_groupings/Pantheon+SH0ES_122221_CALIB.cov</code>) to derive prior widths gives σ(<code>calibrator_offset_mag</code>)≈0.019 mag. Under these bounds, the joint anchor-consistency fit can only support a much smaller step: <code>calibrator_offset_mag ≈ 0.045 ± 0.016</code> (tension-reduction frac ≈ 0.09 vs ≈ 0.37 when free). Calibrator-holdout predictive scoring still prefers the mechanism, but with a smaller gain (Δlogp ≈ +4.1 vs +5.7 when free).<br>Reports: <code>outputs/stack_sn_bao_cc_plus_ladder_calibcov_gates_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_calibcov_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_calibcov_gates_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_calibcov_v1.yaml</code><br>Derivation: <code>scripts/derive_pantheon_shoes_cov_priors.py --raw-cov-path data/raw/pantheon_plus_shoes/sytematic_groupings/Pantheon+SH0ES_122221_CALIB.cov</code></p> </li> <li> <p><strong>All systematic-group covariance blocks (Pantheon+SH0ES <code>sytematic_groupings/*.cov</code>; real data):</strong> repeating the same “covariance-implied prior” construction on <em>every</em> shipped systematic-group covariance block yields very similar coherent-scale bounds (σ(<code>calibrator_offset_mag</code>)≈0.019–0.022 mag). Under these bounds, the joint fit only supports <code>calibrator_offset_mag ≈ 0.04–0.05</code> and tension-reduction fractions ≈0.07–0.09. Calibrator-holdout predictive scoring still prefers a calibrator offset, but with reduced gains (Δlogp ≈ +3.6–+4.1 vs +5.2 when free).<br>Reports: <code>outputs/stack_sn_bao_cc_plus_ladder_groupings_gates_extgrid_all_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_groupings_extgrid_all_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_groupings_gates_extgrid_all_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_groupings_extgrid_all_v1.yaml</code></p> </li> <li> <p><strong>Per-survey/epoch calibration constraints (CALIB.cov “survey×time”; real data):</strong> allowing per-survey time-bin offsets (<code>survey_pkmjd_bins</code>) does <strong>not</strong> meaningfully reduce the anchor-consistency tension under CALIB.cov-derived bounds; calibrator-holdout improvement is small (Δlogp ≈ +0.5 when constrained). An injection map shows that moving <code>delta_lnH0</code> by the full ladder-vs-anchor amount would require a single survey×epoch-bin offset of <strong>multiple magnitudes</strong> (≫0.1 mag).<br>Reports: <code>outputs/stack_sn_bao_cc_plus_ladder_surveytime_gates_extgrid_all_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_surveytime_calibcov_extgrid_all_v1/report.md</code>, <code>outputs/pantheon_plus_shoes_ladder_injection_calibcov_survey_pkmjd_bins_misspec_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_surveytime_gates_extgrid_all_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_surveytime_calibcov_extgrid_all_v1.yaml</code>, <code>configs/pantheon_plus_shoes_ladder_injection_calibcov_survey_pkmjd_bins_misspec_v1.yaml</code></p> </li> <li> <p><strong>CALIB.cov time-bin offsets on calibrators (<code>pkmjd_bins</code>; real data + injection):</strong> under CALIB.cov-derived bounds (σ≈0.02 mag per time-bin), calibrator time-bin offsets provide only small tension reduction in the joint stack (≈0.02) and modest held-out calibrator gains (Δlogp ≈ +1.1–+1.4 for time bins alone; fullcov/diag). Injection mapping shows that faking the full ladder-vs-anchor offset via a <strong>single</strong> calibrator time bin would require a <strong>~1 mag</strong> shift (≈0.9–1.5 mag depending on bin), which is not physically plausible as calibration drift.<br>Reports: <code>outputs/stack_sn_bao_cc_plus_ladder_mechanism_attribution_calibcov_bounded_extgrid_more_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_mechanism_attribution_calibcov_bounded_extgrid_more_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_mechanism_attribution_calibcov_bounded_extgrid_more_fullcov_v1/report.md</code>, <code>outputs/pantheon_plus_shoes_ladder_injection_calibcov_pkmjd_bins_misspec_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_mechanism_attribution_calibcov_bounded_extgrid_more_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_mechanism_attribution_calibcov_bounded_extgrid_more_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_mechanism_attribution_calibcov_bounded_extgrid_more_fullcov_v1.yaml</code>, <code>configs/pantheon_plus_shoes_ladder_injection_calibcov_pkmjd_bins_misspec_v1.yaml</code></p> </li> <li> <p><strong>External calibration metadata (SNANA kcor variants; real data):</strong> ingesting the Pantheon+ DataRelease <code>SNANA_kcor</code> calibration files and turning the <strong>spread across calibration variants</strong> into per-survey(+time-bin) prior widths yields σ≈0.01–0.015 mag for the affected surveys. Under these bounds, per-survey epoch-bin offsets (<code>survey_pkmjd_bins</code>) remain small (max |mean| ≈ 0.006 mag across 40 coefficients) and provide essentially no held-out-calibrator gain (Δlogp ≈ +0.07) once constrained; the joint stack still prefers a large global calibrator offset when allowed.<br>Reports: <code>outputs/stack_sn_bao_cc_plus_ladder_surveytime_kcor_gates_extgrid_more_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_surveytime_kcor_extgrid_more_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_surveytime_kcor_extgrid_more_fullcov_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_surveytime_kcor_gates_extgrid_more_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_surveytime_kcor_extgrid_more_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_surveytime_kcor_extgrid_more_fullcov_v1.yaml</code><br>Prior derivation: <code>scripts/derive_pantheon_shoes_kcor_variant_priors.py</code> (outputs <code>data/processed/external_calibration/pantheon_plus_shoes_sigma_overrides_from_kcor_variants_v1.json</code>)</p> </li> <li> <p><strong>Forward “how much can this explain?” bound (prior MC; real data + simulator):</strong> using the new <code>prior_mc</code> forward simulator to draw random time-bin calibration drifts from the kcor-variant priors and refit the ladder shows the induced <code>delta_lnH0</code> shift has p95(|.|) ≈ 0.0199, corresponding to <strong>~21% of the full ladder-vs-anchor tension</strong> (p99 ≈ 28%).<br>Report: <code>outputs/pantheon_plus_shoes_ladder_prior_mc_kcor_timebins_v1/report.md</code><br>Reproduce: <code>configs/pantheon_plus_shoes_ladder_prior_mc_kcor_timebins_v1.yaml</code></p> </li> <li> <p><strong>Injection mapping (kcor priors; survey×time → H0 shift):</strong> injecting a single per-survey epoch-bin magnitude offset into calibrators maps to a tiny change in <code>delta_lnH0</code> (typical slopes ≈0.02–0.03 in <code>delta_lnH0</code> per mag). Matching the full ladder-vs-anchor <code>delta_lnH0</code> would require a single-bin offset of <strong>several magnitudes</strong>, which is far beyond plausible calibration drift; when <code>survey_pkmjd_bins</code> is explicitly modeled under the kcor priors, the fitted per-bin offsets remain tiny (max |mean| ≈ 0.005 mag).<br>Reports: <code>outputs/pantheon_plus_shoes_ladder_injection_kcor_survey_pkmjd_bins_misspec_v1/report.md</code>, <code>outputs/pantheon_plus_shoes_ladder_injection_kcor_survey_pkmjd_bins_modeled_v1/report.md</code>, <code>outputs/pantheon_plus_shoes_ladder_injection_kcor_survey_pkmjd_bins_misspec_fullcov_v1/report.md</code>, <code>outputs/pantheon_plus_shoes_ladder_injection_kcor_survey_pkmjd_bins_modeled_fullcov_v1/report.md</code><br>Reproduce: <code>configs/pantheon_plus_shoes_ladder_injection_kcor_survey_pkmjd_bins_misspec_v1.yaml</code>, <code>configs/pantheon_plus_shoes_ladder_injection_kcor_survey_pkmjd_bins_modeled_v1.yaml</code>, <code>configs/pantheon_plus_shoes_ladder_injection_kcor_survey_pkmjd_bins_misspec_fullcov_v1.yaml</code>, <code>configs/pantheon_plus_shoes_ladder_injection_kcor_survey_pkmjd_bins_modeled_fullcov_v1.yaml</code></p> </li> <li> <p><strong>Mechanism attribution (SALT2 metadata; real data):</strong> adding calibrator-only linear terms in SALT2 stretch <code>x1</code> improves held-out calibrators modestly (Δlogp ≈ +1.1 in full-cov), but does <strong>not</strong> replace the need for a free <code>calibrator_offset_mag</code> (still Δlogp ≈ +5.4 under kcor gates). In a baseline-sweep (anchor-consistency), <code>x1</code>/<code>c</code>/<code>biasCor_m_b</code> linear terms reduce the inferred <code>delta_lnH0</code> by only <strong>~1–2%</strong> (vs <strong>~24%</strong> for <code>calibrator_offset_mag</code>).<br>Reports: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_mechanism_attribution_kcor_extgrid_more_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_mechanism_attribution_kcor_extgrid_more_fullcov_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_surveytime_kcor_gates_mechanism_attribution_extgrid_more_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_mechanism_attribution_kcor_extgrid_more_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_mechanism_attribution_kcor_extgrid_more_fullcov_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_surveytime_kcor_gates_mechanism_attribution_extgrid_more_v1.yaml</code></p> </li> <li> <p><strong>Injection bound (SALT2 metadata → H0 shift):</strong> if you omit the <code>c_linear_mag</code> / <code>x1_linear_mag</code> / <code>biascor_m_b_linear_mag</code> terms, an injected calibrator-only dependence would need to be extremely large (≈0.6–0.9 mag per 1σ in <code>c</code>/<code>x1</code>/<code>biasCor_m_b</code>) to mimic the full ladder-vs-anchor <code>delta_lnH0</code>; when these terms are modeled, leakage into <code>delta_lnH0</code> is consistent with zero.<br>Reports: <code>outputs/pantheon_plus_shoes_ladder_injection_kcor_c_x1_biascor_misspec_v1/report.md</code>, <code>outputs/pantheon_plus_shoes_ladder_injection_kcor_c_x1_biascor_modeled_v1/report.md</code><br>Reproduce: <code>configs/pantheon_plus_shoes_ladder_injection_kcor_c_x1_biascor_misspec_v1.yaml</code>, <code>configs/pantheon_plus_shoes_ladder_injection_kcor_c_x1_biascor_modeled_v1.yaml</code></p> </li> <li> <p><strong>CID-group holdout (joint stack; real data):</strong> using a stricter holdout split that withholds <em>entire SNe</em> by name (<code>CID</code>) in the ladder part (while keeping hubble-flow + other probes fixed in TRAIN) still prefers a calibrator offset, though the gain is smaller on this split family (Δlogp ≈ +0.32 for <code>calibrator_offset_mag</code>, vs ≈ +0.10 for <code>x1_linear_cal</code>).<br>Report: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_mechanism_attribution_kcor_extgrid_more_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_mechanism_attribution_kcor_extgrid_more_v1.yaml</code></p> </li> <li> <p><strong>Combined “external gate” decomposition (kcor+SH0ES-linear-system; real data):</strong> using kcor-variant calibration widths to bound per-survey/per-epoch offsets <strong>and</strong> a SH0ES-linear-system-inspired σ(<code>calibrator_offset_mag</code>)≈0.028, bounded survey/epoch fields alone explain ~0% of the joint-stack shift, while a bounded global <code>calibrator_offset_mag</code> reduces <code>delta_lnH0</code> by only ~11%. A metadata-rich constrained model can move <code>delta_lnH0</code> by ~16% but is disfavored by evidence.<br>Reports: <code>outputs/stack_sn_bao_cc_plus_ladder_constrained_decomp_kcor_calhf_shoeslin_extgrid_more_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_constrained_decomp_kcor_calhf_shoeslin_extgrid_more_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_constrained_decomp_kcor_calhf_shoeslin_extgrid_more_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_constrained_decomp_kcor_calhf_shoeslin_extgrid_more_v1.yaml</code><br>Priors: <code>data/processed/external_calibration/pantheon_plus_shoes_sigma_overrides_kcor_calhf_plus_shoeslin_v1.json</code> (built via <code>scripts/merge_sigma_overrides.py</code>)</p> </li> <li> <p><strong>Prior-MC bound under combined gates (real data):</strong> drawing unmodeled calibrator-step distortions from the combined kcor+SH0ES-linear-system priors yields p95(|δ<code>delta_lnH0</code>|)≈0.031, i.e. ≈39% of a reference “full tension” scale ln(73/67.4); in 20k draws, P(>100%)≈0.<br>Report: <code>outputs/pantheon_plus_shoes_ladder_prior_mc_constrained_kcor_calhf_shoeslin_v1/report.md</code><br>Reproduce: <code>configs/pantheon_plus_shoes_ladder_prior_mc_constrained_kcor_calhf_shoeslin_v1.yaml</code></p> </li> <li> <p><strong>Cov-projected “external” bounds for metadata proxies (JLA_SALT2; real data):</strong> projecting the published JLA_SALT2 systematic covariance onto proxy vectors yields realistic prior widths for metadata terms (typical σ≈0.021–0.045 mag for <code>pkmjd_err_linear_mag</code>, <code>host_mass_step_mag</code>, <code>c/x1/biascor</code>, <code>mwebv</code>). Under these bounds, a CID-group holdout inside the joint stack prefers the <strong>bounded metadata-rich</strong> ladder model: Δlogp ≈ +0.40 (diag), exceeding the bounded <code>calibrator_offset_mag</code> model (Δlogp ≈ +0.28). Term ablations show the gain is dominated by <strong><code>host_mass_step</code> (Δlogp ≈ +0.29)</strong> and <strong><code>pkmjd_err_linear</code> (Δlogp ≈ +0.26)</strong>, while <code>mwebv</code>/<code>c_linear</code> contribute little.<br>Reports: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_constrained_decomp_kcor_calhf_shoeslin_covproj_JLA_SALT2_cal_extgrid_more_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_covproj_term_ablations_v1/report.md</code><br>Driver ranking: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_covproj_term_ablations_v1/driver_ranking.md</code> (via <code>scripts/rank_predictive_score_drivers.py</code>)<br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_constrained_decomp_kcor_calhf_shoeslin_covproj_JLA_SALT2_cal_extgrid_more_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_covproj_term_ablations_v1.yaml</code><br>Priors: <code>data/processed/external_calibration/pantheon_plus_shoes_sigma_overrides_kcor_calhf_shoeslin_plus_covproj_JLA_SALT2_cal_v1.json</code> (built via <code>scripts/derive_proxy_priors_from_cov_projection.py</code> + <code>scripts/merge_sigma_overrides.py</code>)</p> </li> <li> <p><strong>Permutation-null “is this tied to the actual metadata?” (real data; covproj CID holdout):</strong> for the same CID-holdout term ablations, a <em>within-survey</em> permutation-null test supports <code>host_mass_step</code> as a “real” (non-random) driver but does <strong>not</strong> support <code>pkmjd_err_linear</code>:</p> <ul> <li><code>+fields+host_mass_step</code>: observed mean Δlogp ≈ +0.2915; permuting <code>HOST_LOGMASS</code> among calibrators within each <code>IDSURVEY</code> gives p(Δlogp≥obs) ≈ <strong>0.0138</strong> (n=5000).</li> <li><code>+fields+pkmjd_err</code>: observed mean Δlogp ≈ +0.2647; permuting <code>PKMJDERR</code> among calibrators within each <code>IDSURVEY</code> gives p ≈ <strong>0.127</strong> (n=5000), consistent with a “generic regressor” rather than a specific metadata-linked effect.</li> <li>Full <code>+bounded_fields_plus_metadata_bounded</code>: observed mean Δlogp ≈ +0.3985; permuting <code>HOST_LOGMASS</code> within-survey gives p ≈ <strong>0.0778</strong> (n=5000), i.e. marginal. Artifacts: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_covproj_term_ablations_v1/permutation_null__fields_host_mass_step_host_logmass_within_survey_n5000.json</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_covproj_term_ablations_v1/permutation_null__fields_pkmjd_err_pkmjd_err_within_survey_n5000.json</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_covproj_term_ablations_v1/permutation_null__bounded_fields_plus_metadata_bounded_host_logmass_within_survey_n5000.json</code></li> </ul> </li> <li> <p><strong>Harder generalization (calibrator survey holdout; covproj term ablations):</strong> holding out calibrator SNe survey-by-survey (<code>group_var=idsurvey</code>, HF always in TRAIN) still yields large gains from the same bounded terms:</p> <ul> <li><code>+fields+host_mass_step</code>: Δlogp ≈ +1.35 (win-rate 100%)</li> <li><code>+fields+pkmjd_err</code>: Δlogp ≈ +1.28 (win-rate 90%)</li> <li>Full <code>+bounded_fields_plus_metadata_bounded</code>: Δlogp ≈ +1.93 (win-rate 90%) Permutation-null on this split family (within-survey; n=5000): p ≈ <strong>0.0104</strong> for <code>host_mass_step</code> (permute <code>HOST_LOGMASS</code>) and p ≈ <strong>0.1516</strong> for <code>pkmjd_err</code> (permute <code>PKMJDERR</code>). Report: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_survey_holdout_covproj_term_ablations_v1/report.md</code><br>Artifacts: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_survey_holdout_covproj_term_ablations_v1/permutation_null__fields_host_mass_step_host_logmass_within_survey_n5000.json</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_survey_holdout_covproj_term_ablations_v1/permutation_null__fields_pkmjd_err_pkmjd_err_within_survey_n5000.json</code> Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_survey_holdout_covproj_term_ablations_v1.yaml</code></li> </ul> </li> <li> <p><strong>Host-mass effect is calibrator-specific (real data; covproj bounds):</strong> if the host-mass step is allowed to apply to <em>all</em> ladder SNe instead of calibrators only, its held-out-calibrator benefit drops sharply, and an HF-only mass step gives ~no gain. On calibrator survey holdout:</p> <ul> <li>calibrator-only: Δlogp ≈ +1.23 (vs bounded fields)</li> <li>all-SNe: Δlogp ≈ +0.38</li> <li>HF-only: Δlogp ≈ +0.00 Report: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_survey_holdout_covproj_hostmass_scope_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_survey_holdout_covproj_hostmass_scope_v1.yaml</code></li> </ul> </li> <li> <p><strong>Same conclusion under random calibrator holdout (real data; covproj bounds):</strong> random splits (200 reps) show a much larger gap:</p> <ul> <li>calibrator-only: Δlogp ≈ +3.50 (vs bounded fields)</li> <li>all-SNe: Δlogp ≈ +1.07</li> <li>HF-only: Δlogp ≈ +0.00 Report: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_covproj_hostmass_scope_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_covproj_hostmass_scope_v1.yaml</code></li> </ul> </li> <li> <p><strong>Calibration gates for these proxy terms (prior-MC + SBC + injections; real data + simulator):</strong> under the combined kcor+SH0ES-linear+covproj bounds, a forward “prior-MC” draw of unmodeled distortions can explain p50/p95/p99 ≈ <strong>20% / 60% / 79%</strong> of a reference ln(73/67.4) tension scale (but P(>100%) remains ≈0). Repeated-noise SBC shows <strong>no undercoverage</strong> (it is conservative/over-covered), and injections confirm the dominant single-term H0-shift risks are <code>pkmjd_err_linear_mag</code> and <code>host_mass_step_mag</code> (each ≈15% of the reference scale at 1σ).<br>Reports: <code>outputs/pantheon_plus_shoes_ladder_prior_mc_constrained_kcor_calhf_shoeslin_covproj_JLA_SALT2_cal_v1/report.md</code>, <code>outputs/pantheon_plus_shoes_ladder_sbc_constrained_covproj_JLA_SALT2_cal_v1/report.md</code>, <code>outputs/pantheon_plus_shoes_ladder_injection_covproj_metadata_misspec_v1/report.md</code>, <code>outputs/pantheon_plus_shoes_ladder_injection_covproj_metadata_modeled_v1/report.md</code><br>Reproduce: <code>configs/pantheon_plus_shoes_ladder_prior_mc_constrained_kcor_calhf_shoeslin_covproj_JLA_SALT2_cal_v1.yaml</code>, <code>configs/pantheon_plus_shoes_ladder_sbc_constrained_covproj_JLA_SALT2_cal_v1.yaml</code>, <code>configs/pantheon_plus_shoes_ladder_injection_covproj_metadata_misspec_v1.yaml</code>, <code>configs/pantheon_plus_shoes_ladder_injection_covproj_metadata_modeled_v1.yaml</code></p> </li> <li> <p><strong>Same suite with “realcal survey×epoch” constraints (real data):</strong> using a merged external gate pack (kcor-variant priors + calibcov survey×time bins (min3) + SH0ES linear-system prior + covproj bounds), the prior-MC bound remains similar: frac(tension) p50/p95/p99 ≈ <strong>0.211 / 0.621 / 0.811</strong>, and SBC remains <strong>conservative/over-covered</strong> (e.g. 68% intervals cover the truth ≫68% for <code>delta_lnH0</code>).<br>Reports: <code>outputs/pantheon_plus_shoes_ladder_prior_mc_realcal_surveytime_shoeslin_covproj_v1/report.md</code>, <code>outputs/pantheon_plus_shoes_ladder_sbc_realcal_surveytime_shoeslin_covproj_v1/report.md</code>, <code>outputs/pantheon_plus_shoes_ladder_injection_realcal_surveytime_shoeslin_covproj_misspec_v1/report.md</code>, <code>outputs/pantheon_plus_shoes_ladder_injection_realcal_surveytime_shoeslin_covproj_modeled_v1/report.md</code><br>Reproduce: <code>configs/pantheon_plus_shoes_ladder_prior_mc_realcal_surveytime_shoeslin_covproj_v1.yaml</code>, <code>configs/pantheon_plus_shoes_ladder_sbc_realcal_surveytime_shoeslin_covproj_v1.yaml</code>, <code>configs/pantheon_plus_shoes_ladder_injection_realcal_surveytime_shoeslin_covproj_misspec_v1.yaml</code>, <code>configs/pantheon_plus_shoes_ladder_injection_realcal_surveytime_shoeslin_covproj_modeled_v1.yaml</code></p> </li> <li> <p><strong>Host-mass alone “how much can it explain?” bound (prior-MC; covproj bounds):</strong> drawing an unmodeled calibrator-only host-mass step with σ from the cov-projected external bounds yields frac(tension) p50/p95/p99 ≈ <strong>0.124 / 0.360 / 0.473</strong> of ln(73/67.4).<br>Report: <code>outputs/pantheon_plus_shoes_ladder_prior_mc_host_mass_step_covproj_v1/report.md</code><br>Reproduce: <code>configs/pantheon_plus_shoes_ladder_prior_mc_host_mass_step_covproj_v1.yaml</code></p> </li> <li> <p><strong>SH0ES calibrator-chain prior scale (linear-system; real data product):</strong> the SH0ES DataRelease includes a compact linear system (<code>SH0ES_Data/all[LCY]_...fits</code>, <code>lstsq_results.txt</code>) with σ(<code>fivelogH0</code>)≈0.028 mag. Treating that as a <em>calibrator-chain-inspired</em> prior width for an additional <code>calibrator_offset_mag</code>, the joint stack under FRAGILISTIC survey priors supports only <code>calibrator_offset_mag ≈ 0.074 ± 0.021</code> (tension-reduction frac ≈ 0.16).<br>Reports: <code>outputs/stack_sn_bao_cc_plus_ladder_fragilistic_shoes_gates_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_fragilistic_shoes_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_fragilistic_shoes_gates_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_fragilistic_shoes_v1.yaml</code><br>Derivation: <code>scripts/derive_shoes_linear_system_fivelogh0_prior.py</code></p> </li> <li> <p><strong>Injection mapping (calibrator step → H0 shift):</strong> with an unmodeled calibrator-only magnitude shift injected into the ladder data, the induced bias in <code>delta_lnH0</code> follows the expected slope <code>d(delta_lnH0)/d(Δm)≈0.46</code>, so faking the full ladder-vs-anchor offset requires Δm≈0.19 mag.<br>Report: <code>outputs/pantheon_plus_shoes_ladder_injection_calibcov_misspec_v1/report.md</code><br>Reproduce: <code>configs/pantheon_plus_shoes_ladder_injection_calibcov_misspec_v1.yaml</code></p> </li> <li> <p><strong>Anti-overfit gates:</strong> cross-validated predictive scoring and exact Gaussian log-evidence both <em>penalize</em> adding flexible closure terms (HF redshift splines / sky low-ℓ modes) on the ladder subset.<br>Reports: <code>outputs/pantheon_plus_shoes_ladder_predictive_score_v2/report.md</code>, <code>outputs/pantheon_plus_shoes_ladder_level_sweep_v2/report.md</code></p> </li> <li> <p><strong>External H0 probes as <code>h0_grid</code> (TRGB / lenses / masers; stress-test):</strong> adding these does not remove the need for a large calibrator↔HF offset in the joint stack, and calibrator holdout improvements persist.<br>Reports: <code>outputs/stack_sn_bao_cc_plus_ladder_cal_offset_extgrid_low_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_cal_offset_extgrid_high_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_cal_offset_extgrid_all_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_extgrid_all_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_survey_holdout_extgrid_all_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_cal_offset_extgrid_low_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_cal_offset_extgrid_high_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_cal_offset_extgrid_all_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_extgrid_all_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_survey_holdout_extgrid_all_v1.yaml</code></p> </li> <li> <p><strong>More external H0 grids (SBF + STRIDES; stress-test):</strong> adding two additional late-time constraints (<code>sbf_blakeslee_2021</code>, <code>strides_shajib_2020</code>) still leaves the joint fit preferring a large calibrator↔HF step (<code>calibrator_offset_mag ≈ 0.15</code>) and leaves calibrator-holdout gains essentially unchanged (Δlogp ≈ +4.9 for <code>calibrator_offset_mag</code>).<br>Reports: <code>outputs/stack_sn_bao_cc_plus_ladder_cal_offset_extgrid_more_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_extgrid_more_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_cal_offset_extgrid_more_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cal_holdout_extgrid_more_v1.yaml</code><br>Build grids: <code>scripts/build_h0_grids_from_external_constraints.py</code> (inputs: <code>data/processed/external_constraints/h0_constraints_2026-02-05.json</code>)</p> </li> <li> <p><strong>External <code>h0_grid</code> holdout (real data; new):</strong> holding out each external constraint (training on SN-only+BAO+CC+ladder, testing on one <code>h0_grid:*</code> part) shows the <strong>baseline</strong> predicts these probes best; ladder-only correction models <em>reduce</em> predictive score (Δlogp < 0 in both diag and full-cov). This suggests the “winning” ladder corrections on calibrator holdouts do <strong>not</strong> generalize to external late-time H0 constraints.<br>Reports: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_h0_holdout_realcal_diag_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_h0_holdout_realcal_fullcov_v1/report.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_h0_holdout_realcal_diag_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_h0_holdout_realcal_fullcov_v1.yaml</code></p> </li> <li> <p><strong>CID holdout under the “realcal survey×epoch” pack (real data; stack + ext grids):</strong> calibrator-CID holdout improvements persist under the same merged sigma pack: Δlogp ≈ +0.31 (<code>cal_offset_bounded</code>) and ≈ +0.38 (<code>+bounded_fields_plus_metadata_bounded</code>) vs baseline.<br>Report: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_extgrid_more_v1/report.md</code><br>Driver ranking: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_extgrid_more_v1/driver_ranking.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_extgrid_more_v1.yaml</code></p> </li> <li> <p><strong>CID holdout under “realcal” + FRAGILISTIC filter priors (real data; new):</strong> rerunning the same calibrator-CID holdout battery with <strong>filter-level</strong> FRAGILISTIC priors leaves results essentially unchanged: Δlogp ≈ +0.31 (<code>cal_offset_bounded</code>) and ≈ +0.40 (<code>+bounded_fields_plus_metadata_bounded</code>) vs baseline, with the same “hero” calibrators driving gains (e.g. <code>2007af</code> appears in 4 surveys).<br>Report: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_v1/report.md</code><br>Driver ranking: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_v1/driver_ranking.md</code><br>Hero provenance: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_v1/hero_calibrators_provenance.md</code><br>CID discordance report: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_v1/cid_discordance.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_v1.yaml</code></p> </li> <li> <p><strong>CID holdout with calibrator duplicates <em>deduplicated</em> (real data; new):</strong> enforcing one row per calibrator CID (choose max <code>FITPROB</code>) reduces the cross‑validated gains substantially. Diag scoring: Δlogp ≈ +0.21 (<code>cal_offset_bounded</code>) and ≈ +0.22 (<code>+bounded_fields_plus_metadata_bounded</code>) vs baseline. Full‑cov scoring: Δlogp ≈ +0.27 / +0.28 (same two models), i.e. the gain is <strong>not</strong> just a diagonal-scoring artifact. The top “hero” CID <code>2007af</code> drops from Δlogp≈+11.1 (n_test=4) to Δlogp≈+2.0 (n_test=1), implying earlier leverage was partly driven by duplicate survey reductions.<br>Reports: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_cid_dedup_cal_bestfitprob_v1/report.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_cid_dedup_cal_bestfitprob_fullcov_v1/report.md</code><br>Driver ranking: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_cid_dedup_cal_bestfitprob_v1/driver_ranking.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_cid_dedup_cal_bestfitprob_fullcov_v1/driver_ranking.md</code><br>Hero provenance: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_cid_dedup_cal_bestfitprob_v1/hero_calibrators_provenance.md</code>, <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_cid_dedup_cal_bestfitprob_fullcov_v1/hero_calibrators_provenance.md</code><br>CID discordance report: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_cid_dedup_cal_bestfitprob_v1/cid_discordance.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_cid_dedup_cal_bestfitprob_v1.yaml</code>, <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_cid_dedup_cal_bestfitprob_fullcov_v1.yaml</code></p> </li> <li> <p><strong>CID holdout with <em>discordant</em> duplicate reductions dropped (real data; new):</strong> dropping calibrator duplicates with |Δm_b_corr|>0.05 mag relative to the best‑<code>FITPROB</code> reduction gives an intermediate outcome: Δlogp ≈ +0.28 (<code>cal_offset_bounded</code>) and ≈ +0.34 (<code>+bounded_fields_plus_metadata_bounded</code>). <code>2007af</code> still appears as a major driver (n_test=3), so this drop-only rule does not fully remove duplicate leverage.<br>Report: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_cid_dropdiscord_cal_bestfitprob_v1/report.md</code><br>Driver ranking: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_cid_dropdiscord_cal_bestfitprob_v1/driver_ranking.md</code><br>Hero provenance: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_cid_dropdiscord_cal_bestfitprob_v1/hero_calibrators_provenance.md</code><br>CID discordance report: <code>outputs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_cid_dropdiscord_cal_bestfitprob_v1/cid_discordance.md</code><br>Reproduce: <code>configs/stack_sn_bao_cc_plus_ladder_predictive_score_cid_holdout_realcal_surveytime_shoeslin_covproj_fragfilter_extgrid_more_cid_dropdiscord_cal_bestfitprob_v1.yaml</code></p> </li> <li> <p><strong>Independent probe adapter (sirens, Gate-2):</strong> selection-corrected per-event <code>logL(H0)</code> grid + metadata cut/time drift audit (experimental; upstream Gate-2 product not complete yet).<br>Report: <code>outputs/siren_gate2_grid_audit_v2/report.md</code><br>Reproduce: <code>configs/siren_gate2_grid_audit_v2.yaml</code></p> </li> </ul> <p>For the full status + what’s still missing vs the full ambition, start at:</p> <ul> <li><code>docs/LAYMAN_SUMMARY.md</code></li> <li><code>docs/SPEC_STATUS.md</code></li> </ul> <div> <h2>Quickstart (uses existing venvs in this workspace)</h2> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#quickstart-uses-existing-venvs-in-this-workspace"></a></div> <p>Run with this repo’s venv:</p> <div> <pre>.venv/bin/hubble-audit --help .venv/bin/hubble-audit run configs/pantheon_plus_audit.yaml</pre> <div> </div> </div> <p>Or run with the Project venv (also present in this workspace):</p> <div> <pre>PYTHONPATH=src ../PROJECT/.venv/bin/python -m hubble_systematics.cli --help</pre> <div> </div> </div> <p>Or install editable into that venv:</p> <div> <pre>../PROJECT/.venv/bin/pip install -e <span>.</span> hubble-audit --help</pre> <div> </div> </div> <p>Example: run the Pantheon+ audit packet (baseline + cut scan + correlated-null drift MC):</p> <div> <pre>PYTHONPATH=src ../PROJECT/.venv/bin/python -m hubble_systematics.cli run configs/pantheon_plus_audit.yaml</pre> <div> </div> </div> <p>Example: ladder time-invariance null (fixed-support calibrators; shuffle-within-survey null):</p> <div> <pre>PYTHONPATH=src ../PROJECT/.venv/bin/python -m hubble_systematics.cli run configs/pantheon_plus_shoes_ladder_time_invariance_null_fixedcal_v3.yaml</pre> <div> </div> </div> <p>Outputs are written under <code>outputs/<run_id>/</code>.</p> <div> <h2>Tests</h2> <a href="https://github.com/simulationstation/Hubble-Systematics-Review-Chain#tests"></a></div> <div> <pre>.venv/bin/pytest</pre> <div> </div> </div> </div> </div> </div> </div> </div> </div> </div> <div> <div> <div> <div> <div>   <div> </div> </div> </div> </div> </div> </div>