Run: output/calibration_runs/MERGED_10000_complete_R11cov · 10000 sampled
parameter sets per design × two design configurations (A: geography free, B: geography frozen)
× 1 random seeds.
All scores in this report are single-seed.
This report scores 10000 parameter sets across 2 designs, with 1 seed(s) per set: 20,000 planned ABM runs, 20,000 completed runs, and 20,000 complete design points scored.
rescue/cursor-stuck-save-20260615, commit 8c21082e72dc, worktree dirty (11 changed paths).The score is multivariate implausibility. We compute it separately for the 16 STPM cells
and the 16 STS cells, then show the Jackson survival envelope as a separate check.
Q cutoff probability=0.995; χ²_16(0.995)=34.267; univariate max I_j cutoff=3.0; model discrepancy=none. Current run has one seed per design point; ABM stochastic covariance V_seed is borrowed from prior multi-rep covariance files: STPM output/calibration_runs/20260528_1013_relaxed_sd2_cont_4097to10100_v8/design_A_geography_free/postprocess/implausibility_R1_full_R11cov/covariances/stpm_stochastic_covariance_seed.csv; STS output/calibration_runs/20260528_1013_relaxed_sd2_cont_4097to10100_v8/design_A_geography_free/postprocess/implausibility_R1_full_R11cov/covariances/sts_stochastic_covariance_seed.csv.
For each block (STPM, STS) we compare an $n$-vector of ABM means $y$ to the target vector $z$. The residual is $d = y - z$. Each target has a single combined covariance built additively from three independent sources:
$$\Sigma_{\text{total}} \;=\; \underbrace{\Sigma_{\text{obs}}}_{\text{target uncertainty}} \;+\; \underbrace{\Sigma_{\text{stoch}}}_{\text{ABM seed variability}} \;+\; \underbrace{\Sigma_{\text{disc}}}_{\text{model discrepancy}}.$$From $\Sigma_{\text{total}}$ we form the two scores shown throughout the report:
$$\text{Univariate (per target):} \quad I_j \;=\; \frac{d_j}{\sqrt{(\Sigma_{\text{total}})_{jj}}} \qquad \text{Multivariate (per block):} \quad Q \;=\; d^\top \Sigma_{\text{total}}^{-1} d.$$A design passes a block iff $Q \le \chi^2_n(0.995)$. With $n=16$ targets per block this is $\chi^2_{16}(0.995) = 34.267$ (STPM) and $\chi^2_{16}(0.995) = 34.267$ (STS). The univariate companion cutoff is $|I_j| \le 3.0$ ($3\sigma$ Pukelsheim rule).
| Term | What it captures | Source in this run |
|---|---|---|
| $\Sigma_{\text{obs}}$ (STPM) | Sampling uncertainty of the STPM transition targets ($n \times n$, $n=16$). | data/input_data/stpm_transitions/quit_calibration_targets_covariance_20260412_v1.csvsha256 f760ba0a1a0c |
| $\Sigma_{\text{obs}}$ (STS) | Bootstrap uncertainty of the STS quit-attempt targets ($n \times n$, $n=16$). | data/target_data/sts-quit-targets/output/attempt_calibration_targets_covariance.csvsha256 f05501c17fe1 |
| $\Sigma_{\text{stoch}}$ (STPM, STS) | ABM seed-to-seed variance for each block, estimated from $R$ replicate seeds on a small multi-rep design subset. Borrowed when the calibration run itself has one seed per design. | STPM: output/calibration_runs/20260528_1013_relaxed_sd2_cont_4097to10100_v8/design_A_geography_free/postprocess/implausibility_R1_full_R11cov/covariances/stpm_stochastic_covariance_seed.csvSTS: output/calibration_runs/20260528_1013_relaxed_sd2_cont_4097to10100_v8/design_A_geography_free/postprocess/implausibility_R1_full_R11cov/covariances/sts_stochastic_covariance_seed.csv$R = 1$ seeds per design point in the source run. |
| $\Sigma_{\text{disc}}$ | Structural model error (additional uncertainty for ABM $\to$ target mapping). | set to zero (--model-discrepancy none). |
Survival is checked separately: we compute smoothed 3-, 6-, 12-month survival proportions and require all three to fall inside the Jackson et al. envelope. Survival pass/fail is reported alongside $Q$ and $I_j$.
| Design | sampled parameter sets | scored complete sets | sampled / frozen θ | planned / completed ABM runs | missing runs | rows used for scoring |
|---|---|---|---|---|---|---|
| Design A (geography free) | 10000 | 10000 | 40 / 0 | 10000 / 10000 | 0 | 350000 |
| Design B (geography frozen) | 10000 | 10000 | 38 / 2 | 10000 / 10000 | 0 | 350000 |
| Scope | planned seed-runs | completed | failed | wall time | sum elapsed time | median completed run | p95 completed run | retried runs | max attempts |
|---|---|---|---|---|---|---|---|---|---|
| Design A (geography free) | 10000 | 10000 | 0 | 63 h 21 min | 577 h 03 min | 3.4 min | 3.6 min | 239 | 10 |
| Design B (geography frozen) | 10000 | 10000 | 0 | 66 h 35 min | 588 h 28 min | 3.4 min | 4.1 min | 472 | 10 |
| All manifests | 20000 | 20000 | 0 | 74 h 54 min calendar 129 h 57 min active design time | 1165 h 31 min | 3.4 min | 3.8 min | 711 | 10 |
Run artefacts record sampling commit 94bf0765c1b0.
They do not record a branch name; the branch above is the current checkout used to render this HTML.
| Role | Path | Notes / hash |
|---|---|---|
| Priors metadata | output/calibration_runs/MERGED_10000_complete_R11cov/priors_v1_metadata.json | source SEM files and hashes for the prior CSVs. |
| Design A priors | output/calibration_runs/MERGED_10000_complete_R11cov/priors_v1_designA.csv | sha256 23ad0608a92a |
| Design B priors | output/calibration_runs/MERGED_10000_complete_R11cov/priors_v1_designB.csv | sha256 1154692bfff1 |
| Design A (geography free): sampling metadata | output/calibration_runs/MERGED_10000_complete_R11cov/design_A_geography_free/sampling_metadata.json | Sobol/design hashes; sampling git 94bf0765c1b0. |
| Design A (geography free): parameter design | output/calibration_runs/MERGED_10000_complete_R11cov/design_A_geography_free/parameter_design.csv | sha256 be2facc324f8 |
| Design A (geography free): per-run outputs | output/calibration_runs/MERGED_10000_complete_R11cov/design_A_geography_free/postprocess/per_run_outputs.parquet | direct simulation inputs to this report; sha256 272318c28b21 |
| Design A (geography free): implausibility summary | output/calibration_runs/MERGED_10000_complete_R11cov/design_A_geography_free/postprocess/implausibility_R1/implausibility_summary.csv | formal multivariate Q pass/fail ledger. |
| Design A (geography free): target spec | output/calibration_runs/MERGED_10000_complete_R11cov/design_A_geography_free/postprocess/implausibility_R1/target_spec.csv | canonical 16 STPM + 16 STS output definitions and target values. |
| Design A (geography free): Q components | output/calibration_runs/MERGED_10000_complete_R11cov/design_A_geography_free/postprocess/implausibility_R1/implausibility_components_stpm.csv | STPM diagonal variances and per-target residual components; STS analogue is beside it. |
| Design A (geography free): Q cross-terms | output/calibration_runs/MERGED_10000_complete_R11cov/design_A_geography_free/postprocess/implausibility_R1/implausibility_cross_terms_stpm.csv | STPM precision matrix terms; STS analogue is beside it. |
| Design A (geography free): posterior summary | output/calibration_runs/MERGED_10000_complete_R11cov/design_A_geography_free/postprocess/implausibility_R1/wave_report/posterior_summary.csv | loaded for wave-report context and cross-checking. |
| Design B (geography frozen): sampling metadata | output/calibration_runs/MERGED_10000_complete_R11cov/design_B_geography_frozen/sampling_metadata.json | Sobol/design hashes; sampling git 94bf0765c1b0. |
| Design B (geography frozen): parameter design | output/calibration_runs/MERGED_10000_complete_R11cov/design_B_geography_frozen/parameter_design.csv | sha256 355190334ece |
| Design B (geography frozen): per-run outputs | output/calibration_runs/MERGED_10000_complete_R11cov/design_B_geography_frozen/postprocess/per_run_outputs.parquet | direct simulation inputs to this report; sha256 d6e49744fdd6 |
| Design B (geography frozen): implausibility summary | output/calibration_runs/MERGED_10000_complete_R11cov/design_B_geography_frozen/postprocess/implausibility_R1/implausibility_summary.csv | formal multivariate Q pass/fail ledger. |
| Design B (geography frozen): target spec | output/calibration_runs/MERGED_10000_complete_R11cov/design_B_geography_frozen/postprocess/implausibility_R1/target_spec.csv | canonical 16 STPM + 16 STS output definitions and target values. |
| Design B (geography frozen): Q components | output/calibration_runs/MERGED_10000_complete_R11cov/design_B_geography_frozen/postprocess/implausibility_R1/implausibility_components_stpm.csv | STPM diagonal variances and per-target residual components; STS analogue is beside it. |
| Design B (geography frozen): Q cross-terms | output/calibration_runs/MERGED_10000_complete_R11cov/design_B_geography_frozen/postprocess/implausibility_R1/implausibility_cross_terms_stpm.csv | STPM precision matrix terms; STS analogue is beside it. |
| Design B (geography frozen): posterior summary | output/calibration_runs/MERGED_10000_complete_R11cov/design_B_geography_frozen/postprocess/implausibility_R1/wave_report/posterior_summary.csv | loaded for wave-report context and cross-checking. |
| Network candidates | props/network_candidates.csv | network_seed=20260526; 4 candidates; weight rule: inverse_implausibility (w_i ~ 1 / max_implausibility_i); sha256 01593924a181 |
| Network edge lists | data/input_data/synthetic_network/edgeList_123.csvdata/input_data/synthetic_network/edgeList_247.csvdata/input_data/synthetic_network/edgeList_818.csvdata/input_data/synthetic_network/edgeList_972.csv | local-only input files selected by network_id in parameter_design.csv. |
| STPM targets: means | data/input_data/stpm_transitions/quit_calibration_targets_means_20260412_v1.csv | sha256 fe9e7ced3a6a |
| STPM targets: covariance | data/input_data/stpm_transitions/quit_calibration_targets_covariance_20260412_v1.csv | sha256 f760ba0a1a0c |
| STS targets: means | data/target_data/sts-quit-targets/output/attempt_calibration_targets_means.csv | sha256 1e4406321602 |
| STS targets: covariance | data/target_data/sts-quit-targets/output/attempt_calibration_targets_covariance.csv | sha256 f05501c17fe1 |
| Jackson survival envelope | data/target_data/survival-curve/jackson_curves_table.csv | sha256 235d508a69cd; used for the 3-, 6-, and 12-month envelope. |
| Prior source: attempts SEM | data/intermediate_data/com_b/attempts_SEM_coefficients_20260519_v8.csv | sha256 ab355935593d; mtime 2026-05-19T09:32:42.758116+00:00 |
| Prior source: maintenance SEM | data/intermediate_data/com_b/maintenance_SEM_coefficients_20260519_v8.csv | sha256 80118e267927; mtime 2026-05-19T09:32:40.564167+00:00 |
| Prior source: hazel.md | docs/hazel.md | sha256 4405df7378e4; mtime 2026-05-01T06:35:30.522526+00:00 |
Each candidate $\theta$ is scored against two 16-target blocks ($B \in \{\text{STPM}, \text{STS}\}$). For a single seed $s$, the residual vector $d^{(s)}(\theta) = y^{(s)}_{\text{sim}}(\theta) - y_{\text{obs}}$ is scored against the block covariance $\Sigma_B = V_{\text{obs}} + V_{\text{seed}} + V_{\text{disc}}$ (with $V_{\text{disc}}=0$ for this run):
Because this run has one seed per parameter set, the formal Q scores use borrowed ABM stochastic covariance from the previous multi-rep v8 run, rather than estimating V_seed from this trial or setting it to zero.
$$ Q_B^{(s)}(\theta) \;=\; d^{(s)}(\theta)^\top \Sigma_B^{-1}\, d^{(s)}(\theta), \qquad I_j^{(s)}(\theta) \;=\; \frac{\left|d_j^{(s)}(\theta)\right|}{\sqrt{\Sigma_{B,jj}}}. $$Two implausibility tests, both at level $\alpha=0.005$:
For target-subset sensitivity we use a marginal-Q screen: $\tilde Q_S^{(s)} = \sum_{j \in S} d_j^{2} / \Sigma_{jj}$. This sums each target's squared standardised residual and ignores covariance cross-terms, so it is a screening tool, not the formal multivariate test. The reference cutoff scales as $\tilde Q_S \sim \chi^2_{|S|}(0.995)$ under the null.
Every reactive plot below is driven by the top-K% designs (lowest composite score). The composite is a six-term weighted sum: score(i) = wQs·QSTPM(i) + wQt·QSTS(i) + wIs·(ImaxSTPM−3)+ + wIt·(ImaxSTS−3)+ + wJ·jackson_fail_count(i) + wP·(Δprev)+. The prevalence-direction term penalises designs where overall smoking prevalence increased between 2011 and 2016 (Δprev = prev2016 − prev2011; only the positive part contributes). Use the weight inputs in the floating bar to emphasise or suppress any term (set a weight to 0 to drop it).
The $Q/Q^*$ cards show the 5th, 50th, and 95th percentiles across all parameter sets for the selected design and seed. A value of 1.0x is the multivariate cutoff; values above 1.0x fail.
For each selected target, the bar shows the [min, max] range of $I_j$ across the top-K% of designs by composite score at the current seed; the diamond marks the median. A target whose envelope sits above the dashed cutoff $I_j = 3$ is unreachable even by the best candidates — those are the targets driving the multivariate-$Q$ ledger.
Grouped bar chart of calibration targets (blue, ±1 SE from observed data) and the ABM output for the top-K% designs by composite score (coloured, mean ± 5th–95th percentile range across top-K% designs). STPM y-axis is annual quit rate; STS y-axis is quit-attempt rate over 12 months. The dashed vertical line separates the 10 IMD-period cells (left) from the 6 sex×age cells (right). A good fit has the coloured bar aligned with the blue bar on every cell.
For the composite top-K ranking, this section shows the diagnostic reweighting over every free parameter at once: prior (dotted gray) vs. the implausibility-weighted empirical density on the wave-1 candidates (solid colour). Pick the ranking rule that matches the question you want to answer.
For the current design × seed, this section counts how many scored candidates pass each formal test, individually and in combination. The parallel-coordinates plot below the table renders one line per design across five normalised axes (lower is better); the top-K% by composite score are drawn in colour, the rest are faded grey.
| Test | cutoff | pass count (all designs) | pass count (selected top-K%) | pass rate (top-K%) |
|---|---|---|---|---|
| Multivariate $Q_{\text{STPM}} \le \chi^2_{16}(0.995)$ | — | — | — | — |
| Multivariate $Q_{\text{STS}} \le \chi^2_{16}(0.995)$ | — | — | — | — |
| Multivariate both blocks | — | — | — | — |
| Univariate $\max_j I_j^{\text{STPM}} \le 3$ | 3.00 | — | — | — |
| Univariate $\max_j I_j^{\text{STS}} \le 3$ | 3.00 | — | — | — |
| Univariate both blocks | 3.00 | — | — | — |
| Survival envelope (all 3 checkpoints) | Jackson | — | — | — |
| Prevalence direction (2016 − 2011 ≤ 0) | ≤0 | — | — | — |
| All tests combined | — | — | — | — |
Axes are normalised: $Q_B / \chi^{2*}$, $\max_j I_j / 3$, and survival-failure count (number of checkpoints out of 3 that fall outside the Jackson envelope). A design with every axis $\le 1$ would pass every formal test. Top-K% are drawn solid; the rest are faded.
Each thin line is one of the top-K% best designs at the current seed; it traces simulated survival probability at months 3, 6, and 12. The grey ribbon is the selected Jackson et al. (2019) survival band. The dotted line is the placebo reference, the dashed green line is varenicline. Designs are coloured by rank within the top-K (darker = lower composite score).
Two views restricted to the top-K% by composite score. The heatmaps show each target's marginal miss size, $I_j^2 = d_j^2 / \Sigma_{jj}$, for the K best-ranked designs; rows are sorted best-to-worst by composite score. A bright column is a target the best-ranked candidates still miss badly. The bar chart below compares the prior network share against the empirical share within the top-K%.
Each row is one selected parameter set: rank 1 is the lowest composite score in the selected top-K%. Each column is one target. Colour is log-scaled so moderate misses stay visible. The blue-to-orange transition marks $I_j^2 = 9$ ($I_j=3$), the univariate cutoff; orange, red, and purple cells are targets still missed by the selected top-K% designs. This is a marginal per-target view, not the full multivariate Q with covariance cross-terms.
Bars compare the uniform prior (1/4 per candidate) against the top-K% empirical network share. Green outline means the candidate is over-represented vs the prior; red means under-represented.