Composed Adaptive Pipeline — Type I Error Simulator

Estimates the trial-wide type I error rate when several adaptive mechanisms — historical borrowing, Bayesian sequential monitoring, sample-size re-estimation, and response-adaptive randomization — are combined in a single two-arm binary trial over overlapping interim analyses. Its purpose is to evaluate operating characteristics across the composite null rather than at a single assumed control rate: an efficacy threshold calibrated at one anchor need not control type I error elsewhere in the null.

How it works, when to use it, assumptions & limitations ▸

How it works

You enable any combination of the four mechanisms, set a null scenario (a common baseline response rate plus an optional linear time trend), and the tool runs a Monte Carlo simulation to estimate the pipeline-level rejection rate under the null, with a Monte Carlo standard error and confidence interval. Presets sweep a grid of baseline rates so you can see where control holds and where it fails, and cross baseline-rate departure with the time trend to report their interaction.

When to use it

  • You are combining multiple adaptive features in one trial and want to check the composed type I error rather than assume component-level control carries over.
  • You want to evaluate operating characteristics over a grid of plausible control rates instead of a single assumed anchor.
  • You want to check whether disabling an adaptive component restores control at your calibrated threshold.

Assumptions & limitations

  • A grid demonstrates a control failure but does not locate the supremum of type I error over the null: it establishes that the maximum is at least the largest value observed, not what the worst case is.
  • Mechanism toggles are fixed-threshold ablations, not a decomposition. The threshold is calibrated with all mechanisms active, so a toggled-off run is not separately calibrated and its error rate must not be read as the excess attributable to that component.
  • The model is a two-arm binary trial with specific mechanism variants (blinded nuisance-parameter SSR, Thompson-sampling RAR, a MAP-mixture prior, a linear trend) — other implementations may behave differently. The SSR rule holds the planned effect fixed and uses a working 1:1 Wald calculation; that reference power is not the exact power of the Bayesian/RAR pipeline.
  • Results are Monte Carlo estimates for the scenarios you simulate, not guarantees for a realized trial. Threshold calibration is in-sample, so out-of-sample behaviour is unverified.

For the full methodology, derivation, and worked examples, see the guide for your design: