Calibration: Wave 2 Results & Key Decisions

PI Briefing — 6 July 2026  |  30 min  |  Smoking cessation ABM history matching

Established Where we stand

30
NROY designs (Design B)
19.5
Best QSTPM (Design B)
22 / 30
Survive multi-seed (10 seeds)
1
NROY designs (Design A)

How did QSTPM improve?

Design A: 90.0 → 32.3
StepBest QΔShare
Original W190.0
+ Bug fixes57.9−32.255.7%
+ Vs increase35.6−22.338.6%
+ Conditioning32.3−3.35.7%
Design B: 70.0 → 19.5
StepBest QΔShare
Original W170.0
+ Bug fixes35.6−34.468.1%
+ Vs increase20.4−15.130.0%
+ Conditioning19.5−0.91.9%

Bug fixes dominate (56–68%). Vs increase is second (30–39%). Conditioning contributed only 2–6% — this is the key problem driving Decision 2.

Decision What do we do about stochasticity?

What is Vs and why does it matter?

The implausibility score QSTPM measures how far the model’s outputs are from the calibration targets, scaled by the total uncertainty budget:

QSTPM  =  ∑  (model − target)²  /  (Vobs + Vs + Vmd)

When Vs is larger, the denominator is larger, so QSTPM is smaller (more “forgiving”). The question is whether the post-fix Vs is the correct estimate or an overestimate.

What changed

FactDetail
Vs increased post-fix STPM stochastic variance grew 1.77× (A) / 2.12× (B). STS unaffected (1.0×).
Why this is genuine The TSQ fix makes the quit-success pathway “sticky” — successful quitters no longer relapse immediately, amplifying seed-to-seed variation. The pre-fix Vs was artificially low.
Vs share of total variance 65–85% (up from 37–79% pre-fix)

The consequence

0
NROY (either design)
with pre-fix Vs
30
NROY (Design B)
with post-fix Vs
22 / 30
Survive multi-seed
(10 seeds each)
9–11
r* (optimal seed count)
for stable Vs

All 30 NROY designs exist only because Vs increased. Multi-seed validation shows 73% are robust, not just lucky seeds.

Decision needed

  1. Accept post-fix Vs as the correct scoring baseline? The argument: the pre-fix Vs was wrong because the bug suppressed genuine stochasticity. Report both Vs versions as a sensitivity check in publications.
  2. How many seeds per design in wave 3? With 1 seed, single-seed noise is large (8 of 30 NROY flip out under 10-seed scoring). r* suggests 9–11 seeds for stable estimation. Trade-off is runtime (see Runtime Estimates).

Decision What do we do about conditioning?

The problem

Why was conditioning weak?

CauseDetail
Selection too broad Top-1000 (10%) from 10k — most parameters look the same as the full distribution.
STPM-dominated ranking QSTPM gap between good (~100) and constraint-passing (~444) designs is >300, so top-K is almost entirely STPM-quality designs.
Bimodal posterior maintenance.bias ≈ −0.05 (STPM-good) vs +0.66 (prevalence-passing). Gaussian recentring cannot capture both modes.

Wave 2 test evidence (2.5k designs per combo, pre-fix Vs)

MetricS1×B (STPM-ranked)S2×B (Constraint-first)
Best QSTPM87.090.8
QSTPM p5145.2173.5
Prevalence falls14.0%42.9%
Jackson pass29.1%41.3%
Both pass5.6%16.3%

S1 gives better QSTPM; S2 dramatically improves constraint satisfaction (3× more designs with falling prevalence).

Decision needed

  1. Do we want stricter conditioning for wave 3? The current approach barely moved the priors.
  2. Which strategy?
StrategyProsCons
Smaller K (top-200 or top-300) Stronger shrinkage; better concentration May exclude constraint-passing region
Union pool (top-500 STPM ∪ all constraint-passers) Captures both regimes Bimodal; Gaussian recentring can’t capture two modes
Constraint-first (only J+prev passers, ranked by Q) Forces constraint-passing exploration Small pool (~300–400); Q floor high (~166)
Mixture-of-Gaussians (2-component GMM) Properly captures bimodal posterior More complex; custom Sobol from mixture

Info Runtime estimates

This machine (measured from wave 2)

ParameterValue
Machine CPUs128
Parallel workers124 (CPU − 4)
ABM run time (per run, 72 ticks)~10 min
10k × 1 seed × 1 design16.7h (ABM 13.9h + aggregate 2.8h)
10k × 1 seed × 2 designs33.5h (sequential A then B)

Wave 3 cost estimates

ScenarioABM runsWall time (est.)Notes
10k × 1 seed × 2 designs20,000~33hSame as wave 2
10k × 3 seeds × 1 design (B)30,000~50hMulti-seed, Design B only
10k × 5 seeds × 1 design (B)50,000~70h (2.9d)Full stochastic scoring
5k × 10 seeds × 1 design (B)50,000~70h (2.9d)Fewer designs, more seeds
2k × 10 seeds × 1 design (B)20,000~28hTargeted deep exploration

Based on actual wave 2 timings: 10k×1 ABM batch took 13.9h per design on 124 workers. Aggregate adds ~2.8h. Estimates assume linear scaling.


Appendix: All reports

Wave 1 & bug-fix

Stochasticity

Wave 2

Conditioning & deep dives


Generated 2026-07-06 13:05  |  Data: post-fix W1 (20260703_0711), wave 2 (20260704_1141), multi-seed (20260706_0847)