Evaluation

Two gates are measured on this page, and the honest summary of both is that the effect is there and the interval is not tight enough to prove it.

Gate 6 clears its threshold on the point estimate and on the held-out split, and spans it on the split that was named in advance. Gate 5 improves the Brier score and its interval contains zero. Both are recorded as NOT_ESTABLISHED, because a point estimate above a bar whose interval straddles the bar is not a pass.

Of the 3 of 6 the sidebar counts as met, 2 are the feasibility gates 1 and 2, answered before any pipeline code existed.

What the two gates below are up against is a scale that stops at 1.740× for a perfect oracle against a bar of 1.5, which is the finding rather than an excuse and is derived in section 02 below.

Why the gates that are not met are not met

3 of 6 did not come back met, and none of them is left as a bare verdict.

Each one names what actually bound the measurement and the condition that would move it, computed by scripts/run_gate_power.py from the same receipts that decided the verdicts. That script refuses to write its receipt at all while an unmet gate has no named constraint, so a gate cannot go quietly missing from here.

Gate 3: Corridor intersects a visible trace

PASSED UNGROUPED ONLY
What bound it
289 of 303 testable observations scored, 224 discriminating. The exact bound is 0.731 against a 0.7 bar.
What would close it
Not more episodes. This corpus has 68 independent (station, date) episodes, which already exceeds the 9 all-discriminating episodes that would clear a 0.7 bar on their own, and 32 of the 68 discriminate on every capture, putting the grouped bound at 0.366. Clearing 0.7 at 68 episodes takes 55 of them, and at 54 the bound is 0.697. The observation-level bound clears 0.7 at 0.731; the pre-registration's rule is to group, so the observation-level pass is reported and not claimed.

Gate 5: Physics beats image-only on Brier

NOT ESTABLISHED
What bound it
88 test observations. The interval's lower arm is 1.63 times the margin it has to clear.
What would close it
About 233 test observations at the same margin, against the 88 this split has. That is 2.6 times the chronological test set. (projected)

Gate 6: Queue lift over random

NOT ESTABLISHED
What bound it
87 observations at a budget of 50 cap every ordering at 1.740x, leaving 0.240 of room for an interval 0.387 wide.
What would close it
A split whose room exceeds the interval it produces. cold_station already does: room 2.673 against an interval 1.939 wide, and it passed at 2.253.

Gate 6’s verdict is predicted by the split it was taken on, not by the queue. A split’s room is the distance between the threshold and the best score any ordering could reach there, a perfect oracle included. Whether the published interval fits inside that room predicts the verdict on 4 of 4 measurable splits, with no exceptions.

On chronological and cold_combined the interval's upper bound is the ceiling itself: no resampling of those splits can return a number above it, however good the ranking is. The one split with room to spare is the one split that passed.
SplitObservationsOracle ceilingRoom above 1.5xInterval widthFitsVerdict
chronological871.7400.2400.387noNOT ESTABLISHED
cold_combined761.5200.0200.447noNOT ESTABLISHED
cold_station2174.1732.6731.939yesPASSED
cold_transmitter951.9000.4000.558noNOT ESTABLISHED

The obvious next thought does not follow, and the counterexample is in this corpus: cold_transmitter holds more observations than chronological and still fails, because its interval came back wider too. So no required sample size is published for gate 6, only the condition.

Kill gate 6: does the queue beat random?

Require the top review queue to find at least 1.5 times as many manually actionable conflicts as random ordering at the same budget.

NOT ESTABLISHEDdecided on the Chronological split

The queue's point lift is 1.58 on the chronological split (20 conflicts in 50 examined, expected 12.6 by random). The 95% interval spans the 1.5 threshold (1.35 to 1.74). A point estimate above 1.5 whose interval does not sit above 1.5 is not a pass, for the same reason gate 5 was recorded as NOT_ESTABLISHED: the evidence does not exclude noise.

Chronological

NOT ESTABLISHED20 of 50 reviewed
0.5threshold 1.52.5
Lift
1.582
Grouped by episode
[1.353, 1.740] (87 groups)
Grouped by station
[1.368, 1.735] (35 groups)
Reported
the union of the two
Station clustering
ICC 0.089, design effect 1.13
Episode clustering
An intra-class correlation needs at least 2 populated groups and more observations than groups. Got 87 groups over 87 observations.

Cold station

PASSED27 of 50 reviewed
0.5threshold 1.52.5

The interval runs past this axis. Its full extent is 1.920 to 3.859, and the hatched edge is where it leaves the scale.

Lift
2.253
Grouped by episode
[1.920, 3.011] (217 groups)
Grouped by station
[2.020, 3.859] (27 groups)
Reported
the union of the two
Station clustering
ICC 0.078, design effect 1.55
Episode clustering
An intra-class correlation needs at least 2 populated groups and more observations than groups. Got 217 groups over 217 observations.

Cold transmitter

NOT ESTABLISHED34 of 50 reviewed
0.5threshold 1.52.5
Lift
1.656
Grouped by episode
[1.462, 1.834] (95 groups)
Grouped by station
[1.336, 1.894] (28 groups)
Reported
the union of the two
Station clustering
ICC 0.135, design effect 1.32
Episode clustering
An intra-class correlation needs at least 2 populated groups and more observations than groups. Got 95 groups over 95 observations.

Cold station and transmitter

NOT ESTABLISHED17 of 50 reviewed
0.5threshold 1.52.5
Lift
1.292
Grouped by episode
[1.073, 1.520] (76 groups)
Grouped by station
[1.130, 1.500] (12 groups)
Reported
the union of the two
Station clustering
ICC 0.091, design effect 1.48
Episode clustering
An intra-class correlation needs at least 2 populated groups and more observations than groups. Got 76 groups over 76 observations.

Two groupings are reported because two things are correlated here and they are not the same thing: captures of one pass share a geometry, and captures from one ground station share a receiver.

The reported interval is the union of the two, which is the wider and therefore the weaker claim.

Choosing whichever grouping gave the narrower interval would be choosing the answer.

How much of that lift is guaranteed by the way the queue was built

The ranking score and the definition of a conflict read the same quantities, so part of the lift is true by construction. This bounds that part rather than arguing about it.

Which quantities the score and the conflict definition share

Every criterion's defining quantity is also a weighted term in the score, so no restriction of the target makes this measurement independent of its own construction.

What the restrictions below separate is the model's contribution, not the score's.

Two weights are published because they are different numbers: 0.90 of the score sits on quantities the definition names, and 0.75 sits on quantities a conflict in this corpus is actually defined from.

The gap is DEAD_CAPTURE, which fires on nothing here.

No row in the shipped queue meets this criterion.

The highest value of the quantity it thresholds is 0.1371 over 87 measurable rows.

Ceiling at this budget
1.740×
An oracle finds all 22 and scores this.
Room between bar and oracle
0.240
The gate asked for 1.5×.
Share of the ceiling reached
0.909
Conflicts the queue found, over the most any ordering could.
Permutation p
0.0005
0 of 2000 random orderings matched it.
The model-independent-only row names 2 criteria and measures 1, because DEAD_CAPTURE fires on nothing here. The model-independent-and-firing row is the same restriction with the inert criteria dropped from the name, so the two rows carry identical numbers and only one of them can be misread.
Restricted targetCriteriaConflictsLift95% intervalVerdict
The shipped definitionMODEL_LABEL_DISAGREESTALE_CATALOGUE_FREQDEAD_CAPTURE (fires on nothing)221.582[1.353, 1.740]NOT ESTABLISHED
Model-dependent criteria onlyMODEL_LABEL_DISAGREE31.740[1.705, 1.740]NOT INFORMATIVE
Model-independent criteria onlySTALE_CATALOGUE_FREQDEAD_CAPTURE (fires on nothing)191.557[1.264, 1.740]NOT ESTABLISHED
Model-independent, and firingSTALE_CATALOGUE_FREQ191.557[1.264, 1.740]NOT ESTABLISHED
The scale each split's verdict is read on. A split whose oracle barely clears the threshold cannot produce an informative verdict, whichever way it falls.
SplitPopulationConflictsBudgetOracle scoresRoom above the bar
Chronological8722501.740×0.240
Cold station and transmitter
not informative at this budget
7620501.520×0.020
Cold station21752504.173×2.673
Cold transmitter9539501.900×0.400

That the queue generalises.

Restricting the target to the criteria the model does not enter removes one loop and leaves another: on this corpus that restriction reduces to STALE_CATALOGUE_FREQ alone, whose defining quantity the score weights at 0.35. The other model-independent criterion, DEAD_CAPTURE, fires on nothing in this corpus, so a reader following its 0.15 weight is following a loop that does not exist in the data.

The honest reading is that this measurement is a check on internal consistency and on the size of the space the gate was set in, not an independent test of the ranking.

How the random-ordering control was run

0 of 2000 random orderings of the same population found as many conflicts inside the budget as the shipped queue did, so a permutation p-value of 0.0005, which is the smallest this test can report at 2000 permutations.

The mean of the same 2000 lifts is 0.9992 against an expected 1.0, which is the floor check: every number in this file is produced by the function that returned it.

Kill gate 5: does physics conditioning help?

Require the physics-conditioned model to lower Brier score against a calibrated image-only baseline.

NOT ESTABLISHEDphysics_conditioned against image_only, decided on Chronological

The physics-conditioned arm has the lower Brier score by 0.02079, but the 95% interval (-0.01301 to 0.05036) spans zero on 88 test observations across 88 episodes. A point estimate in the right direction with an interval containing zero is not a gain, and reporting it as one would be the same error unit A7 made. The gate is not met.

A positive margin means the physics-conditioned arm is better. An interval spanning zero is not a gain in either direction.
SplitBrier margin95% intervalDirectionObservationsGroups
Chronological0.02079[-0.01301, 0.05036]indistinguishable8888
Cold station-0.00018[-0.01961, 0.01821]indistinguishable217217
Cold transmitter-0.00089[-0.06700, 0.07239]indistinguishable9696
Cold station and transmitter-0.08014[-0.12323, -0.01025]reference_better7676
Chronological, size matched control-0.01852[-0.06067, 0.01820]indistinguishable8888

The arm ladder

Ten arms on the Chronological split, each adding one block of features to the one before. Brier is the score being minimised; a shorter bar is a worse arm.

Test partition: 88 observations, positive rate 0.705. Calibrator: only 121 calibration labels, below the 200 floor for isotonic; a one-parameter fit is what this sample supports
ArmBlocksBrierAUCECECalibration slope
prior_onlynone0.208510.5000.01860.422
image_onlyimage0.149500.8420.04821.142
physics_onlyphysics0.213620.5820.0394-0.073
corridor_onlycorridor0.187430.7850.14758.282
metadata_onlymetadata0.200110.6510.06151.722
image_metadataimagemetadata0.160700.8200.08502.038
image_physicsimagephysics0.152020.8330.14190.604
image_corridorimagecorridor0.129240.8750.07131.483
physics_conditionedimagephysicscorridor0.128710.8930.11130.714
full_fusionimagephysicscorridormetadata0.150360.8800.12901.164

What was kept, and under which rule

Two retention rules were written down. They disagree, and the stricter one decides.

Nominal rule

Retain a block if an arm containing it beat image-only with a 95% interval clearing zero on a split with at least 300 decisive training rows, and no such split showed it reliably worse.

  • physics DROPthe rules disagree here
  • corridor RETAINthe rules disagree here
  • metadata NOT_ESTABLISHED

recommends: image + corridor as image_corridor

Bonferroni corrected rule

decides

The same, but the interval must clear zero after Bonferroni correction over the comparison family run on that split.

  • physics NOT_ESTABLISHEDthe rules disagree here
  • corridor NOT_ESTABLISHEDthe rules disagree here
  • metadata NOT_ESTABLISHED

recommends: image as image_only

Why the corrected rule decides, and that it was tightened after a number was seen

The correction was promoted from a report to a gate after the nominal rule retained the metadata block on one split whose interval did not survive correction, so this is a rule tightened after seeing a number and that is worth stating.

It stands on two grounds independent of which way it fell: gate 5 already reports corrected intervals, so holding the ablation to a weaker standard would be the inconsistency; and this ladder runs 5 comparisons on each of 4 splits, where one nominal win by chance is the expected outcome rather than evidence.

The corrected rule also selects a combination the ladder actually fitted, so its recommendation carries a measured score and an interval, which the nominal rule's selection does not.

The queue is ranked by the image plus corridor arm (image + corridor), while the corrected ablation rule recommends the image plus only arm (image).

The blocks the shipped ranker uses without corrected support are corridor: retained by the nominal rule, and the Brier interval does not clear zero once the multiplicity correction runs over the family the rule reads.

The ranker was not rebuilt to match, for two measured reasons.

The same comparison does not establish the narrower arm as better either, so swapping on it would be a change made for the appearance of consistency rather than for a result.

And the ablation rule reads Brier comparisons only, while the same arm's risk-coverage area against the reference arm on chronological is +0.05736 with a corrected interval of +0.01192 to +0.11887 over the same 21 comparisons, which does clear zero.

Selective review is what the queue does, so that is the metric closest to the shipped use, and it is reported rather than promoted into the rule after the fact.

Shipped arm
image_corridor
measured on Chronological
Brier
0.12924
lower is better
AUC
0.875
ranking quality
ECE
0.0713
how far the stated probabilities are from the observed rates
The retain decision reads test-set comparisons, so the shipped arm's Brier is optimistic by an amount this corpus cannot measure. A second snapshot is the only thing that settles it, and until then the number travels with this sentence.
Which splits the verdict used, and which fell below the training floor

Splits used for the ablation verdict: Chronological, Cold station, Cold transmitter.

Below the 300-row training floor and therefore excluded: Cold station and transmitter.

Set from the size-matched control, not from taste. The corridor block loses roughly 0.036 Brier on the chronological split purely by cutting the training partition from 530 rows to 188, which exceeds the 0.020 it gains at full size. Below the floor a verdict measures the sample size instead of the block, so those splits still appear in the receipt and still inform the training-size caveat; they do not decide whether a block ships.

What it looks like when the model is allowed to refuse

A triage model that answers on everything is not the only option. This is what the error rate does as it is allowed to abstain.

Risk against coverage, Chronological split0.000.090.180%25%50%75%100%coverage: share of observations the model answers onrisk: error rate among thosechosen on calibration
The vertical bar through the marked point is the 95% interval on the error rate at that operating point.
Target error rate
0.05
the promise the threshold was chosen to keep
Coverage on test
38%
33 of 88 answered
Error rate on test
0.030
1 wrong, interval [0.000, 0.091]
Ceiling held
not held
decided on the top of the interval

This holds on the upper end of the interval, not on the point estimate, because the promise made to a reviewer is about the worst plausible error rate at the reported coverage.

Whether it would also hold on the point estimate is reported in the column beside it, so the difference between the two is visible rather than argued about.

Gate 4 asked a person, and this is what they answered

The one gate no amount of compute closes. A person has answered the committed sample, so this is the review rather than the request.

The threshold was fixed before the build: at least 80% of a balanced sample, reviewed with the network’s labels and every model output hidden, must support a decisive judgment. The sample is 72 items over 60 observations, 12 repeated under a second item id so intra-rater agreement can be measured without telling the reviewer which. Sixty rather than thirty-six because the verdict reads the interval, and at 36 even a true rate of 0.90 could not clear the bar.

The sample is committed to rather than promised: before the review, one salted sha256 per item over the item id, the observation id and the image digest, with the 32-byte salt and the item-to-observation mapping held outside the repository. Afterwards the scorer re-hashes every image from disk, refuses outright if one commitment fails, and publishes the salt and the mapping so anyone can recheck. All 72 were verified against the images on disk when the bundle was packed.

Every number in the clip is in artifacts/GATE4_RECEIPT.json, and tests/test_explainer_gate4_values.py fails if the scene and the receipt ever disagree. The spoken track is read from the same constants the scene draws, and a second model transcribed it without seeing the script to check every figure was said: scripts/render_explainer_narration.py. The reviewer was the author of this project, which the clip’s closing frame states and its narration does not: this gate is passed, and it is not passed independently.

This gate is decided. Kesav Jayakumar (Kesav2k04), the author of TraceTriage, reviewing on 2026-08-22. Not an independent reviewer: the same person built the instrument. What the commitment guarantees instead is that the sample, its order and the images were fixed before the review and cannot have been chosen around the answers. The reviewer could not see the network's waterfall_status, any model prediction, the item-to-observation mapping or the salt, none of which is in the bundle: the key file was never opened before scoring. The reviewer knows the corpus and built the pipeline, so this is blinded but not independent, and the receipt says so rather than implying a stranger did it. Repeat pairs are at least six items apart and are not marked, so the intra-rater figure is still two reads that could not see each other.A model answered the same committed plates first, at 57/60. That review is kept rather than overwritten: the same instrument, two kinds of reviewer.

Decidable
60/60
rate 1.000, 95% [0.951, 1.000] against a 0.80 threshold
Same answer twice
8/12
repeated plates answered identically on all three axes, at least six items apart and unmarked. 12/12 and 12/12 on the two axes that decide decisiveness, 8/12 on target_consistent, which is the whole of the gap. Read across all three it is a ceiling under the 0.80 bar, and the gate rests on the two that hold
Gate 4
PASSED
the gate's own verdict, from the reviewer it names. It follows the interval, so a rate above 0.80 whose bound straddles it would still read NOT_ESTABLISHED
and they miss it in opposite directions, which is why both are here rather than one being named the right one. visible_signal is too broad: it counts a fixed local carrier or an interference burst as a signal, and the network label does not. target_consistent is too narrow: it asks for a smooth curve drifting across frequency, and a short packet burst parked near zero offset is a real pass that answers no. Read `by_axis` with the confusion matrices, not either rate on its own.
Compared against the network labelAgreedOf
visible_signal2340
target_consistent1740
What a reviewer needsWhere it is
The protocol: three axes, four answers, what unsure means/gate4/worksheet.md
The review page, which exports the CSV the scorer reads/gate4/review.html
The 72 plates, 113 MB of full-resolution waterfallsnot published: ask for tracetriage_gate4_bundle.zip
What to check the download against113,238,991 B, sha256 c426e1d978b66cf62a8d024a
The plates are not on this page because both ways of shrinking them are wrong: lossless re-encoding breaks the digests, lossy re-encoding smooths the faint traces the reviewer is asked to judge. So the archive travels whole, as one file with a digest. A second reviewer can repeat the review on the same plates: open review.html from the unpacked folder, answer 72 items, send back one CSV. The scorer will not publish a rate without naming who produced it.

The size-matched control

A cold split has fewer training rows than the chronological one, so a difference could be the split or could be the sample size. This control holds size fixed and changes only the split.

Split: chronological_size_matched.
PartitionRows
train188
calibration121
test88
train_with_image188
test_with_image88
test_with_corridor88