Implemented in Workbench 3.4.1
- Event-level matching uses the strongest event point, while detection delay uses event onset / first alarm.
- Explicit
event-evidence,confirmed-quiet,fully-labelledandunknownruntime states, with compatibility for publication v1.0verified-quietevidence. - Unknown periods and unmatched predictions under positive-only evidence are excluded from false-positive calculations.
- Precision, F1 and false events per object-month are unavailable unless qualified negative-labelled coverage is supplied.
- Performance breakdown by LEO, MEO, GEO and HEO profile, abstention reporting and deterministic missing-data sensitivity.
Two distinct benchmark artifacts
Public Workbench seed v0.1: the public manifest route and the cases below illustrate the browser runner format. Its unknown candidate controls are not verified negative exposure, so this seed alone cannot support precision or false-alert claims.
Frozen publication benchmark v1.0: the completed research programme used 40 satellite-disjoint cases, with 30 event-evidence cases, 10 verified-quiet controls and 154 documented intervals across development, calibration and confirmatory partitions. The confirmatory partition was consumed once in Prompt 11 and is permanently closed to retuning or reuse as untouched evidence.
Physics-ML v1.0 detected 0 of 41 confirmatory intervals and failed the predeclared integration gate. It did not enter the public workbench: detector method 0.4 remains the public scientific method.
Public Workbench seed v0.1 cases
| Case | Orbit | Public evidence | Coverage |
|---|---|---|---|
| ISS Dragon reboost, 8 Nov 2024 | LEO | NASA | Event evidence |
| ISS Progress reboost, 27 Jul 2023 | LEO | NASA | Event evidence; precise UTC start unavailable |
| Himawari-9 station keeping, Jul 2024 | GEO | NOAA OSPO | Event evidence |
| Galileo 5 recovery, Nov 2014 | MEO | ESA | Event-evidence campaign interval |
| INTEGRAL disposal burn, Jan 2015 | HEO | ESA | Event evidence |
Four additional candidate control windows in this seed manifest are marked unknown until independently checked. Download the v0.1 seed manifest.
Build publication benchmark v1.0 locally
The repository includes tools/build_public_evidence_benchmark.py. From a local clone, run python3 tools/build_public_evidence_benchmark.py. It reconstructs the frozen v1.0 cases through the same-site history endpoint and writes a browser-importable local benchmark. At that derived-artifact boundary, frozen verified-quiet labels are translated to the browser's equivalent confirmed-quiet runtime state and their source spelling is retained as provenance.
The generated file can contain Space-Track Basic SSA records. Keep it local unless your data-use permissions allow redistribution. The repository intentionally ignores the default generated filename.
Remaining external qualification
- Independent astrodynamics review and disposition of findings before operational claims.
- External review issue resolution and publication decisions appropriate to the frozen negative result.
- Any future learned-method cycle requires a genuinely new satellite-level holdout; the consumed Prompt 11 confirmatory set cannot qualify a tuned successor.