SCPN Phase Orchestrator · certified head-to-head · real grid corpus
One detector clears the bar. The rest sit at chance.
On 90 real generator-trip instabilities of the PSML 23-bus corpus, the domain-specific modal envelope-growth detector leads 36 at a matched 10% false alarm (p = 0.0001) — while every generic early-warning detector, including the classical critical-slowing-down baseline, stays statistically at chance on the identical split. Both facts are sealed in the same record.
Transitions led at a matched 10% false alarm — 90 real instabilities
Every detector reads the same two-second pre-onset segments of the same scenarios, calibrated on the same damped nulls. The tick on each bar is the count expected by chance at that detector’s own alarm rate; the p-value is the label-permutation probability of reaching the observed count.
Why the comparison is fair
- Identical split. Same segments, same labels, same matched false alarm on the same damped null scenarios — only the detector differs.
- Non-circular labels. A transition is a generator-trip scenario, a null a damped bus-fault or branch-trip: the label is the disturbance type, a physical annotation independent of the growth statistic being scored.
- Disclosed data-quality gate. Three scenarios with a non-physical time column are dropped and counted in the sealed record, never silently.
- Pre-registered operating point. The aggregation and recency weighting were selected on a development half and validated on the held-out half (24/45, p = 0.0002) before this full-corpus comparison.
The sealed record
The whole comparison — corpus, operating point, every detector’s count and
p-value, the drops, the verdict — is one content-addressed artefact in the public
repository (examples/real_data/psml_modal_growth/), guarded by an integrity
test that recomputes the hash from the committed payload.
bc6895879088b31b…- This certifies the detector on this corpus. The certified numeric operating point is dataset-specific: our own cross-dataset test on ISO-NE PMU captures showed the frozen threshold does not port — deployment calibrates per system on its own ambient data, and that step is part of the product, not a caveat hidden in a footnote.
- “Generic detectors at chance” is a statement about this task and operating point, sealed with p-values — not a blanket dismissal of those methods elsewhere.
- Streaming deployment is stricter than per-window scoring: the sealed stream operating point holds 11/45 held-out at a 10% stream false alarm, measured and sealed separately.