Design | Seed | Top fraction (0.01 = top 1%; selected by composite score) | | Survival band | Composite weights score = wQs·QSTPM + wQt·QSTS + wIs·(ImaxSTPM−3)+ + wIt·(ImaxSTS−3)+ + wJ·jackson_fails + wP·(Δprev)+ |

Wave-1 calibration report — single-seed implausibility analysis

Run: output/calibration_runs/replay_bugfix_5k · 5000 sampled parameter sets per design × two design configurations (A: geography free, B: geography frozen) × 1 random seeds. All scores in this report are single-seed.

1. Run, method, and data

This report scores 5000 parameter sets across 2 designs, with 1 seed(s) per set: 0 planned ABM runs, 0 completed runs, and 10,000 complete design points scored.

Sampling5000 parameter sets per design; unknown, scipy unknown, scramble=—, seed=—; A d=—, B d=—.
ABM runs0 planned; 0 completed; 1 seeds per sampled parameter set.
Scored candidates10,000 complete design points.
Report build2026-07-02T10:04:35+00:00 UTC; branch rescue/cursor-stuck-save-20260615, commit 8c21082e72dc, worktree dirty (31 changed paths).

The score is multivariate implausibility. We compute it separately for the 16 STPM cells and the 16 STS cells, then show the Jackson survival envelope as a separate check. Q cutoff probability=0.995; χ²_16(0.995)=34.267; univariate max I_j cutoff=3.0; model discrepancy=none. Current run has one seed per design point; ABM stochastic covariance V_seed is borrowed from prior multi-rep covariance files: STPM output/calibration_runs/MERGED_10000_complete_R11cov/design_A_geography_free/postprocess/implausibility_R1/covariances/stpm_stochastic_covariance_seed.csv; STS output/calibration_runs/MERGED_10000_complete_R11cov/design_A_geography_free/postprocess/implausibility_R1/covariances/sts_stochastic_covariance_seed.csv.

Implausibility score: equations and term sources

For each block (STPM, STS) we compare an $n$-vector of ABM means $y$ to the target vector $z$. The residual is $d = y - z$. Each target has a single combined covariance built additively from three independent sources:

$$\Sigma_{\text{total}} \;=\; \underbrace{\Sigma_{\text{obs}}}_{\text{target uncertainty}} \;+\; \underbrace{\Sigma_{\text{stoch}}}_{\text{ABM seed variability}} \;+\; \underbrace{\Sigma_{\text{disc}}}_{\text{model discrepancy}}.$$

From $\Sigma_{\text{total}}$ we form the two scores shown throughout the report:

$$\text{Univariate (per target):} \quad I_j \;=\; \frac{d_j}{\sqrt{(\Sigma_{\text{total}})_{jj}}} \qquad \text{Multivariate (per block):} \quad Q \;=\; d^\top \Sigma_{\text{total}}^{-1} d.$$

A design passes a block iff $Q \le \chi^2_n(0.995)$. With $n=16$ targets per block this is $\chi^2_{16}(0.995) = 34.267$ (STPM) and $\chi^2_{16}(0.995) = 34.267$ (STS). The univariate companion cutoff is $|I_j| \le 3.0$ ($3\sigma$ Pukelsheim rule).

TermWhat it capturesSource in this run
$\Sigma_{\text{obs}}$ (STPM)Sampling uncertainty of the STPM transition targets ($n \times n$, $n=16$).data/input_data/stpm_transitions/quit_calibration_targets_covariance_20260412_v1.csv
sha256 f760ba0a1a0c
$\Sigma_{\text{obs}}$ (STS)Bootstrap uncertainty of the STS quit-attempt targets ($n \times n$, $n=16$).data/target_data/sts-quit-targets/output/attempt_calibration_targets_covariance.csv
sha256 461b6855d75e
$\Sigma_{\text{stoch}}$ (STPM, STS)ABM seed-to-seed variance for each block, estimated from $R$ replicate seeds on a small multi-rep design subset. Borrowed when the calibration run itself has one seed per design.STPM: output/calibration_runs/MERGED_10000_complete_R11cov/design_A_geography_free/postprocess/implausibility_R1/covariances/stpm_stochastic_covariance_seed.csv
STS: output/calibration_runs/MERGED_10000_complete_R11cov/design_A_geography_free/postprocess/implausibility_R1/covariances/sts_stochastic_covariance_seed.csv
$R = 1$ seeds per design point in the source run.
$\Sigma_{\text{disc}}$Structural model error (additional uncertainty for ABM $\to$ target mapping).set to zero (--model-discrepancy none).

Survival is checked separately: we compute smoothed 3-, 6-, 12-month survival proportions and require all three to fall inside the Jackson et al. envelope. Survival pass/fail is reported alongside $Q$ and $I_j$.

Designsampled parameter setsscored complete setssampled / frozen θplanned / completed ABM runsmissing runsrows used for scoring
Design A (geography free)50005000— / —— / —175000
Design B (geography frozen)50005000— / —— / —175000

Run timing

Scopeplanned seed-runscompletedfailedwall timesum elapsed timemedian completed runp95 completed runretried runsmax attempts
Design A (geography free)50005000024 h 14 min456 h 35 min4.8 min10.4 min01
Design B (geography frozen)50005000023 h 28 min413 h 59 min4.9 min5.1 min01
All manifests1000010000025 h 27 min calendar
47 h 42 min active design time
870 h 34 min4.8 min10.2 min01

Run artefacts record sampling commit . They do not record a branch name; the branch above is the current checkout used to render this HTML.

Data files used by this HTML

RolePathNotes / hash
Priors metadataoutput/calibration_runs/replay_bugfix_5k/priors_v1_metadata.jsonsource SEM files and hashes for the prior CSVs.
Design A priorsoutput/calibration_runs/replay_bugfix_5k/priors_v1_designA.csvsha256
Design B priorsoutput/calibration_runs/replay_bugfix_5k/priors_v1_designB.csvsha256
Design A (geography free): sampling metadataoutput/calibration_runs/replay_bugfix_5k/design_A_geography_free/sampling_metadata.jsonSobol/design hashes; sampling git .
Design A (geography free): parameter designoutput/calibration_runs/replay_bugfix_5k/design_A_geography_free/parameter_design.csvsha256
Design A (geography free): per-run outputsoutput/calibration_runs/replay_bugfix_5k/design_A_geography_free/postprocess/per_run_outputs.parquetdirect simulation inputs to this report; sha256 97623e998d59
Design A (geography free): implausibility summaryoutput/calibration_runs/replay_bugfix_5k/design_A_geography_free/postprocess/implausibility/implausibility_summary.csvformal multivariate Q pass/fail ledger.
Design A (geography free): target specoutput/calibration_runs/replay_bugfix_5k/design_A_geography_free/postprocess/implausibility/target_spec.csvcanonical 16 STPM + 16 STS output definitions and target values.
Design A (geography free): Q componentsoutput/calibration_runs/replay_bugfix_5k/design_A_geography_free/postprocess/implausibility/implausibility_components_stpm.csvSTPM diagonal variances and per-target residual components; STS analogue is beside it.
Design A (geography free): Q cross-termsoutput/calibration_runs/replay_bugfix_5k/design_A_geography_free/postprocess/implausibility/implausibility_cross_terms_stpm.csvSTPM precision matrix terms; STS analogue is beside it.
Design A (geography free): posterior summaryoutput/calibration_runs/replay_bugfix_5k/design_A_geography_free/postprocess/wave_report/posterior_summary.csvloaded for wave-report context and cross-checking.
Design B (geography frozen): sampling metadataoutput/calibration_runs/replay_bugfix_5k/design_B_geography_frozen/sampling_metadata.jsonSobol/design hashes; sampling git .
Design B (geography frozen): parameter designoutput/calibration_runs/replay_bugfix_5k/design_B_geography_frozen/parameter_design.csvsha256
Design B (geography frozen): per-run outputsoutput/calibration_runs/replay_bugfix_5k/design_B_geography_frozen/postprocess/per_run_outputs.csvdirect simulation inputs to this report; sha256 73794a046308
Design B (geography frozen): implausibility summaryoutput/calibration_runs/replay_bugfix_5k/design_B_geography_frozen/postprocess/implausibility/implausibility_summary.csvformal multivariate Q pass/fail ledger.
Design B (geography frozen): target specoutput/calibration_runs/replay_bugfix_5k/design_B_geography_frozen/postprocess/implausibility/target_spec.csvcanonical 16 STPM + 16 STS output definitions and target values.
Design B (geography frozen): Q componentsoutput/calibration_runs/replay_bugfix_5k/design_B_geography_frozen/postprocess/implausibility/implausibility_components_stpm.csvSTPM diagonal variances and per-target residual components; STS analogue is beside it.
Design B (geography frozen): Q cross-termsoutput/calibration_runs/replay_bugfix_5k/design_B_geography_frozen/postprocess/implausibility/implausibility_cross_terms_stpm.csvSTPM precision matrix terms; STS analogue is beside it.
Design B (geography frozen): posterior summaryoutput/calibration_runs/replay_bugfix_5k/design_B_geography_frozen/postprocess/wave_report/posterior_summary.csvloaded for wave-report context and cross-checking.
STPM targets: meansdata/input_data/stpm_transitions/quit_calibration_targets_means_20260412_v1.csvsha256 fe9e7ced3a6a
STPM targets: covariancedata/input_data/stpm_transitions/quit_calibration_targets_covariance_20260412_v1.csvsha256 f760ba0a1a0c
STS targets: meansdata/target_data/sts-quit-targets/output/attempt_calibration_targets_means.csvsha256 fc6a138ffa6b
STS targets: covariancedata/target_data/sts-quit-targets/output/attempt_calibration_targets_covariance.csvsha256 461b6855d75e
Jackson survival envelopedata/target_data/survival-curve/jackson_curves_table.csvsha256 235d508a69cd; used for the 3-, 6-, and 12-month envelope.

2. Per-target fit in natural units

Grouped bar chart of calibration targets (blue, ±1 SE from observed data) and the ABM output for the top-K% designs by composite score (coloured, mean ± 5th–95th percentile range across top-K% designs). STPM y-axis is annual quit rate; STS y-axis is quit-attempt rate over 12 months. The dashed vertical line separates the 10 IMD-period cells (left) from the 6 sex×age cells (right). A good fit has the coloured bar aligned with the blue bar on every cell.

3. Smoker prevalence 2011–2016 (top-K designs)

Shaded band = p5–p95; solid line = median across the top-K% designs selected by the composite score above. Reactive: updates whenever Top fraction or any weight changes.

Prevalence by

4. Where parameters want to move

For the composite top-K ranking, this section shows the diagnostic reweighting over every free parameter at once: prior (dotted gray) vs. the implausibility-weighted empirical density on the wave-1 candidates (solid colour). Pick the ranking rule that matches the question you want to answer.

5. Parallel coordinates — all parameters + scores

One unified plot with all calibrated parameters (left) and normalised scores (right): $Q_B / \chi^{2*}$, $\\max_j I_j / 3$, and survival-failure count. Scroll horizontally to see all axes. Drag along any axis to brush-select; the selection is reflected across all axes. Top-K% are drawn in colour; the rest are faded.

Scores only — compact view (drag an axis range to highlight designs in section 4)

2. How we score each candidate

Each candidate $\theta$ is scored against two 16-target blocks ($B \in \{\text{STPM}, \text{STS}\}$). For a single seed $s$, the residual vector $d^{(s)}(\theta) = y^{(s)}_{\text{sim}}(\theta) - y_{\text{obs}}$ is scored against the block covariance $\Sigma_B = V_{\text{obs}} + V_{\text{seed}} + V_{\text{disc}}$ (with $V_{\text{disc}}=0$ for this run):

Because this run has one seed per parameter set, the formal Q scores use borrowed ABM stochastic covariance from the previous multi-rep v8 run, rather than estimating V_seed from this trial or setting it to zero.

$$ Q_B^{(s)}(\theta) \;=\; d^{(s)}(\theta)^\top \Sigma_B^{-1}\, d^{(s)}(\theta), \qquad I_j^{(s)}(\theta) \;=\; \frac{\left|d_j^{(s)}(\theta)\right|}{\sqrt{\Sigma_{B,jj}}}. $$

Two implausibility tests, both at level $\alpha=0.005$:

For target-subset sensitivity we use a marginal-Q screen: $\tilde Q_S^{(s)} = \sum_{j \in S} d_j^{2} / \Sigma_{jj}$. This sums each target's squared standardised residual and ignores covariance cross-terms, so it is a screening tool, not the formal multivariate test. The reference cutoff scales as $\tilde Q_S \sim \chi^2_{|S|}(0.995)$ under the null.

3. What if we drop the dominating targets?

Every reactive plot below is driven by the top-K% designs (lowest composite score). The composite is a six-term weighted sum: score(i) = wQs·QSTPM(i) + wQt·QSTS(i) + wIs·(ImaxSTPM−3)+ + wIt·(ImaxSTS−3)+ + wJ·jackson_fail_count(i) + wP·(Δprev)+. The prevalence-direction term penalises designs where overall smoking prevalence increased between 2011 and 2016 (Δprev = prev2016 − prev2011; only the positive part contributes). Use the weight inputs in the floating bar to emphasise or suppress any term (set a weight to 0 to drop it).

3a. Live summary

Design / seed
STPM targets selected
STS targets selected
$\chi^2_{k_{\text{STPM}}}(0.995)$
$\chi^2_{k_{\text{STS}}}(0.995)$
$Q_{\text{STPM}} / Q^*$ quantiles (5%, 50%, 95%)
$Q_{\text{STS}} / Q^*$ quantiles (5%, 50%, 95%)
Full multivariate NROY
Subset marginal-Q NROY
Univariate max $I_j \le 3$
Survival envelope pass
Selected top-K survival pass
Prevalence ↓ (2016−2011 ≤ 0)
Top-K prevalence ↓

The $Q/Q^*$ cards show the 5th, 50th, and 95th percentiles across all parameter sets for the selected design and seed. A value of 1.0x is the multivariate cutoff; values above 1.0x fail.

3b. Best-performer envelope per target (top-K%)

For each selected target, the bar shows the [min, max] range of $I_j$ across the top-K% of designs by composite score at the current seed; the diamond marks the median. A target whose envelope sits above the dashed cutoff $I_j = 3$ is unreachable even by the best candidates — those are the targets driving the multivariate-$Q$ ledger.

3c. Subset marginal $\tilde Q_S$ distribution

5. How many candidates survive each test?

For the current design × seed, this section counts how many scored candidates pass each formal test, individually and in combination. The parallel-coordinates plot below the table renders one line per design across five normalised axes (lower is better); the top-K% by composite score are drawn in colour, the rest are faded grey.

5a. Pass-count ledger

Test cutoff pass count (all designs) pass count (selected top-K%) pass rate (top-K%)
Multivariate $Q_{\text{STPM}} \le \chi^2_{16}(0.995)$
Multivariate $Q_{\text{STS}} \le \chi^2_{16}(0.995)$
Multivariate both blocks
Univariate $\max_j I_j^{\text{STPM}} \le 3$ 3.00
Univariate $\max_j I_j^{\text{STS}} \le 3$ 3.00
Univariate both blocks 3.00
Survival envelope (all 3 checkpoints) Jackson
Prevalence direction (2016 − 2011 ≤ 0) ≤0
All tests combined

6. Survival curves vs Jackson envelope

Each thin line is one of the top-K% best designs at the current seed; it traces simulated survival probability at months 3, 6, and 12. The grey ribbon is the selected Jackson et al. (2019) survival band. The dotted line is the placebo reference, the dashed green line is varenicline. Designs are coloured by rank within the top-K (darker = lower composite score).

7. What do the best-ranked candidates imply? — targets and networks

Two views restricted to the top-K% by composite score. The heatmaps show each target's marginal miss size, $I_j^2 = d_j^2 / \Sigma_{jj}$, for the K best-ranked designs; rows are sorted best-to-worst by composite score. A bright column is a target the best-ranked candidates still miss badly. The bar chart below compares the prior network share against the empirical share within the top-K%.

7a. Which targets the top-K% still miss

Each row is one selected parameter set: rank 1 is the lowest composite score in the selected top-K%. Each column is one target. Colour is log-scaled so moderate misses stay visible. The blue-to-orange transition marks $I_j^2 = 9$ ($I_j=3$), the univariate cutoff; orange, red, and purple cells are targets still missed by the selected top-K% designs. This is a marginal per-target view, not the full multivariate Q with covariance cross-terms.

7b. Which networks do the top-K% prefer?

Bars compare the uniform prior (1/4 per candidate) against the top-K% empirical network share. Green outline means the candidate is over-represented vs the prior; red means under-represented.

*Failed run note: No failed_runs.csv found.. The report scores only design_ids present in the implausibility summaries, so incomplete seed sets are excluded from all per-seed scores.