Evaluation
Two gates are measured on this page, and the honest summary of both is that the effect is there and the interval is not tight enough to prove it.
Gate 6 clears its threshold on the point estimate and on the held-out split, and spans it on the split that was named in advance. Gate 5 improves the Brier score and its interval contains zero. Both are recorded as NOT_ESTABLISHED, because a point estimate above a bar whose interval straddles the bar is not a pass.
Of the 3 of 6 the sidebar counts as met, 2 are the feasibility gates 1 and 2, answered before any pipeline code existed.
What the two gates below are up against is a scale that stops at 1.740× for a perfect oracle against a bar of 1.5, which is the finding rather than an excuse and is derived in section 02 below.
Why the gates that are not met are not met
3 of 6 did not come back met, and none of them is left as a bare verdict.
Each one names what actually bound the measurement and the condition that would move it, computed by scripts/run_gate_power.py from the same receipts that decided the verdicts. That script refuses to write its receipt at all while an unmet gate has no named constraint, so a gate cannot go quietly missing from here.
Gate 3: Corridor intersects a visible trace
PASSED UNGROUPED ONLY- What bound it
- 289 of 303 testable observations scored, 224 discriminating. The exact bound is 0.731 against a 0.7 bar.
- What would close it
- Not more episodes. This corpus has 68 independent (station, date) episodes, which already exceeds the 9 all-discriminating episodes that would clear a 0.7 bar on their own, and 32 of the 68 discriminate on every capture, putting the grouped bound at 0.366. Clearing 0.7 at 68 episodes takes 55 of them, and at 54 the bound is 0.697. The observation-level bound clears 0.7 at 0.731; the pre-registration's rule is to group, so the observation-level pass is reported and not claimed.
Gate 5: Physics beats image-only on Brier
NOT ESTABLISHED- What bound it
- 88 test observations. The interval's lower arm is 1.63 times the margin it has to clear.
- What would close it
- About 233 test observations at the same margin, against the 88 this split has. That is 2.6 times the chronological test set. (projected)
Gate 6: Queue lift over random
NOT ESTABLISHED- What bound it
- 87 observations at a budget of 50 cap every ordering at 1.740x, leaving 0.240 of room for an interval 0.387 wide.
- What would close it
- A split whose room exceeds the interval it produces. cold_station already does: room 2.673 against an interval 1.939 wide, and it passed at 2.253.
Gate 6’s verdict is predicted by the split it was taken on, not by the queue. A split’s room is the distance between the threshold and the best score any ordering could reach there, a perfect oracle included. Whether the published interval fits inside that room predicts the verdict on 4 of 4 measurable splits, with no exceptions.
| Split | Observations | Oracle ceiling | Room above 1.5x | Interval width | Fits | Verdict |
|---|---|---|---|---|---|---|
chronological | 87 | 1.740 | 0.240 | 0.387 | no | NOT ESTABLISHED |
cold_combined | 76 | 1.520 | 0.020 | 0.447 | no | NOT ESTABLISHED |
cold_station | 217 | 4.173 | 2.673 | 1.939 | yes | PASSED |
cold_transmitter | 95 | 1.900 | 0.400 | 0.558 | no | NOT ESTABLISHED |
The obvious next thought does not follow, and the counterexample is in this corpus: cold_transmitter holds more observations than chronological and still fails, because its interval came back wider too. So no required sample size is published for gate 6, only the condition.
Kill gate 6: does the queue beat random?
“Require the top review queue to find at least 1.5 times as many manually actionable conflicts as random ordering at the same budget.”
The queue's point lift is 1.58 on the chronological split (20 conflicts in 50 examined, expected 12.6 by random). The 95% interval spans the 1.5 threshold (1.35 to 1.74). A point estimate above 1.5 whose interval does not sit above 1.5 is not a pass, for the same reason gate 5 was recorded as NOT_ESTABLISHED: the evidence does not exclude noise.
Chronological
NOT ESTABLISHED20 of 50 reviewed- Lift
- 1.582
- Grouped by episode
- [1.353, 1.740] (87 groups)
- Grouped by station
- [1.368, 1.735] (35 groups)
- Reported
- the union of the two
- Station clustering
- ICC 0.089, design effect 1.13
- Episode clustering
- An intra-class correlation needs at least 2 populated groups and more observations than groups. Got 87 groups over 87 observations.
Cold station
PASSED27 of 50 reviewedThe interval runs past this axis. Its full extent is 1.920 to 3.859, and the hatched edge is where it leaves the scale.
- Lift
- 2.253
- Grouped by episode
- [1.920, 3.011] (217 groups)
- Grouped by station
- [2.020, 3.859] (27 groups)
- Reported
- the union of the two
- Station clustering
- ICC 0.078, design effect 1.55
- Episode clustering
- An intra-class correlation needs at least 2 populated groups and more observations than groups. Got 217 groups over 217 observations.
Cold transmitter
NOT ESTABLISHED34 of 50 reviewed- Lift
- 1.656
- Grouped by episode
- [1.462, 1.834] (95 groups)
- Grouped by station
- [1.336, 1.894] (28 groups)
- Reported
- the union of the two
- Station clustering
- ICC 0.135, design effect 1.32
- Episode clustering
- An intra-class correlation needs at least 2 populated groups and more observations than groups. Got 95 groups over 95 observations.
Cold station and transmitter
NOT ESTABLISHED17 of 50 reviewed- Lift
- 1.292
- Grouped by episode
- [1.073, 1.520] (76 groups)
- Grouped by station
- [1.130, 1.500] (12 groups)
- Reported
- the union of the two
- Station clustering
- ICC 0.091, design effect 1.48
- Episode clustering
- An intra-class correlation needs at least 2 populated groups and more observations than groups. Got 76 groups over 76 observations.
Two groupings are reported because two things are correlated here and they are not the same thing: captures of one pass share a geometry, and captures from one ground station share a receiver.
The reported interval is the union of the two, which is the wider and therefore the weaker claim.
Choosing whichever grouping gave the narrower interval would be choosing the answer.
How much of that lift is guaranteed by the way the queue was built
The ranking score and the definition of a conflict read the same quantities, so part of the lift is true by construction. This bounds that part rather than arguing about it.
Which quantities the score and the conflict definition share
Every criterion's defining quantity is also a weighted term in the score, so no restriction of the target makes this measurement independent of its own construction.
What the restrictions below separate is the model's contribution, not the score's.
Two weights are published because they are different numbers: 0.90 of the score sits on quantities the definition names, and 0.75 sits on quantities a conflict in this corpus is actually defined from.
The gap is DEAD_CAPTURE, which fires on nothing here.
No row in the shipped queue meets this criterion.
The highest value of the quantity it thresholds is 0.1371 over 87 measurable rows.
| Restricted target | Criteria | Conflicts | Lift | 95% interval | Verdict |
|---|---|---|---|---|---|
| The shipped definition | MODEL_LABEL_DISAGREESTALE_CATALOGUE_FREQDEAD_CAPTURE (fires on nothing) | 22 | 1.582 | [1.353, 1.740] | NOT ESTABLISHED |
| Model-dependent criteria only | MODEL_LABEL_DISAGREE | 3 | 1.740 | [1.705, 1.740] | NOT INFORMATIVE |
| Model-independent criteria only | STALE_CATALOGUE_FREQDEAD_CAPTURE (fires on nothing) | 19 | 1.557 | [1.264, 1.740] | NOT ESTABLISHED |
| Model-independent, and firing | STALE_CATALOGUE_FREQ | 19 | 1.557 | [1.264, 1.740] | NOT ESTABLISHED |
| Split | Population | Conflicts | Budget | Oracle scores | Room above the bar |
|---|---|---|---|---|---|
| Chronological | 87 | 22 | 50 | 1.740× | 0.240 |
| Cold station and transmitter not informative at this budget | 76 | 20 | 50 | 1.520× | 0.020 |
| Cold station | 217 | 52 | 50 | 4.173× | 2.673 |
| Cold transmitter | 95 | 39 | 50 | 1.900× | 0.400 |
That the queue generalises.
Restricting the target to the criteria the model does not enter removes one loop and leaves another: on this corpus that restriction reduces to STALE_CATALOGUE_FREQ alone, whose defining quantity the score weights at 0.35. The other model-independent criterion, DEAD_CAPTURE, fires on nothing in this corpus, so a reader following its 0.15 weight is following a loop that does not exist in the data.
The honest reading is that this measurement is a check on internal consistency and on the size of the space the gate was set in, not an independent test of the ranking.
How the random-ordering control was run
0 of 2000 random orderings of the same population found as many conflicts inside the budget as the shipped queue did, so a permutation p-value of 0.0005, which is the smallest this test can report at 2000 permutations.
The mean of the same 2000 lifts is 0.9992 against an expected 1.0, which is the floor check: every number in this file is produced by the function that returned it.
Kill gate 5: does physics conditioning help?
“Require the physics-conditioned model to lower Brier score against a calibrated image-only baseline.”
The physics-conditioned arm has the lower Brier score by 0.02079, but the 95% interval (-0.01301 to 0.05036) spans zero on 88 test observations across 88 episodes. A point estimate in the right direction with an interval containing zero is not a gain, and reporting it as one would be the same error unit A7 made. The gate is not met.
| Split | Brier margin | 95% interval | Direction | Observations | Groups |
|---|---|---|---|---|---|
| Chronological | 0.02079 | [-0.01301, 0.05036] | indistinguishable | 88 | 88 |
| Cold station | -0.00018 | [-0.01961, 0.01821] | indistinguishable | 217 | 217 |
| Cold transmitter | -0.00089 | [-0.06700, 0.07239] | indistinguishable | 96 | 96 |
| Cold station and transmitter | -0.08014 | [-0.12323, -0.01025] | reference_better | 76 | 76 |
| Chronological, size matched control | -0.01852 | [-0.06067, 0.01820] | indistinguishable | 88 | 88 |
The arm ladder
Ten arms on the Chronological split, each adding one block of features to the one before. Brier is the score being minimised; a shorter bar is a worse arm.
| Arm | Blocks | Brier | AUC | ECE | Calibration slope |
|---|---|---|---|---|---|
| prior_only | none | 0.20851 | 0.500 | 0.0186 | 0.422 |
| image_only | image | 0.14950 | 0.842 | 0.0482 | 1.142 |
| physics_only | physics | 0.21362 | 0.582 | 0.0394 | -0.073 |
| corridor_only | corridor | 0.18743 | 0.785 | 0.1475 | 8.282 |
| metadata_only | metadata | 0.20011 | 0.651 | 0.0615 | 1.722 |
| image_metadata | imagemetadata | 0.16070 | 0.820 | 0.0850 | 2.038 |
| image_physics | imagephysics | 0.15202 | 0.833 | 0.1419 | 0.604 |
| image_corridor | imagecorridor | 0.12924 | 0.875 | 0.0713 | 1.483 |
| physics_conditioned | imagephysicscorridor | 0.12871 | 0.893 | 0.1113 | 0.714 |
| full_fusion | imagephysicscorridormetadata | 0.15036 | 0.880 | 0.1290 | 1.164 |
What was kept, and under which rule
Two retention rules were written down. They disagree, and the stricter one decides.
Nominal rule
Retain a block if an arm containing it beat image-only with a 95% interval clearing zero on a split with at least 300 decisive training rows, and no such split showed it reliably worse.
- physics DROPthe rules disagree here
- corridor RETAINthe rules disagree here
- metadata NOT_ESTABLISHED
recommends: image + corridor as image_corridor
Bonferroni corrected rule
decidesThe same, but the interval must clear zero after Bonferroni correction over the comparison family run on that split.
- physics NOT_ESTABLISHEDthe rules disagree here
- corridor NOT_ESTABLISHEDthe rules disagree here
- metadata NOT_ESTABLISHED
recommends: image as image_only
Why the corrected rule decides, and that it was tightened after a number was seen
The correction was promoted from a report to a gate after the nominal rule retained the metadata block on one split whose interval did not survive correction, so this is a rule tightened after seeing a number and that is worth stating.
It stands on two grounds independent of which way it fell: gate 5 already reports corrected intervals, so holding the ablation to a weaker standard would be the inconsistency; and this ladder runs 5 comparisons on each of 4 splits, where one nominal win by chance is the expected outcome rather than evidence.
The corrected rule also selects a combination the ladder actually fitted, so its recommendation carries a measured score and an interval, which the nominal rule's selection does not.
The queue is ranked by the image plus corridor arm (image + corridor), while the corrected ablation rule recommends the image plus only arm (image).
The blocks the shipped ranker uses without corrected support are corridor: retained by the nominal rule, and the Brier interval does not clear zero once the multiplicity correction runs over the family the rule reads.
The ranker was not rebuilt to match, for two measured reasons.
The same comparison does not establish the narrower arm as better either, so swapping on it would be a change made for the appearance of consistency rather than for a result.
And the ablation rule reads Brier comparisons only, while the same arm's risk-coverage area against the reference arm on chronological is +0.05736 with a corrected interval of +0.01192 to +0.11887 over the same 21 comparisons, which does clear zero.
Selective review is what the queue does, so that is the metric closest to the shipped use, and it is reported rather than promoted into the rule after the fact.
Which splits the verdict used, and which fell below the training floor
Splits used for the ablation verdict: Chronological, Cold station, Cold transmitter.
Below the 300-row training floor and therefore excluded: Cold station and transmitter.
Set from the size-matched control, not from taste. The corridor block loses roughly 0.036 Brier on the chronological split purely by cutting the training partition from 530 rows to 188, which exceeds the 0.020 it gains at full size. Below the floor a verdict measures the sample size instead of the block, so those splits still appear in the receipt and still inform the training-size caveat; they do not decide whether a block ships.
What it looks like when the model is allowed to refuse
A triage model that answers on everything is not the only option. This is what the error rate does as it is allowed to abstain.
This holds on the upper end of the interval, not on the point estimate, because the promise made to a reviewer is about the worst plausible error rate at the reported coverage.
Whether it would also hold on the point estimate is reported in the column beside it, so the difference between the two is visible rather than argued about.
Gate 4 asked a person, and this is what they answered
The one gate no amount of compute closes. A person has answered the committed sample, so this is the review rather than the request.
The threshold was fixed before the build: at least 80% of a balanced sample, reviewed with the network’s labels and every model output hidden, must support a decisive judgment. The sample is 72 items over 60 observations, 12 repeated under a second item id so intra-rater agreement can be measured without telling the reviewer which. Sixty rather than thirty-six because the verdict reads the interval, and at 36 even a true rate of 0.90 could not clear the bar.
The sample is committed to rather than promised: before the review, one salted sha256 per item over the item id, the observation id and the image digest, with the 32-byte salt and the item-to-observation mapping held outside the repository. Afterwards the scorer re-hashes every image from disk, refuses outright if one commitment fails, and publishes the salt and the mapping so anyone can recheck. All 72 were verified against the images on disk when the bundle was packed.
artifacts/GATE4_RECEIPT.json, and tests/test_explainer_gate4_values.py fails if the scene and the receipt ever disagree. The spoken track is read from the same constants the scene draws, and a second model transcribed it without seeing the script to check every figure was said: scripts/render_explainer_narration.py. The reviewer was the author of this project, which the clip’s closing frame states and its narration does not: this gate is passed, and it is not passed independently.This gate is decided. Kesav Jayakumar (Kesav2k04), the author of TraceTriage, reviewing on 2026-08-22. Not an independent reviewer: the same person built the instrument. What the commitment guarantees instead is that the sample, its order and the images were fixed before the review and cannot have been chosen around the answers. The reviewer could not see the network's waterfall_status, any model prediction, the item-to-observation mapping or the salt, none of which is in the bundle: the key file was never opened before scoring. The reviewer knows the corpus and built the pipeline, so this is blinded but not independent, and the receipt says so rather than implying a stranger did it. Repeat pairs are at least six items apart and are not marked, so the intra-rater figure is still two reads that could not see each other.A model answered the same committed plates first, at 57/60. That review is kept rather than overwritten: the same instrument, two kinds of reviewer.
| Compared against the network label | Agreed | Of |
|---|---|---|
visible_signal | 23 | 40 |
target_consistent | 17 | 40 |
| What a reviewer needs | Where it is |
|---|---|
The protocol: three axes, four answers, what unsure means | /gate4/worksheet.md |
| The review page, which exports the CSV the scorer reads | /gate4/review.html |
| The 72 plates, 113 MB of full-resolution waterfalls | not published: ask for tracetriage_gate4_bundle.zip |
| What to check the download against | 113,238,991 B, sha256 c426e1d978b66cf62a8d024a |
review.html from the unpacked folder, answer 72 items, send back one CSV. The scorer will not publish a rate without naming who produced it.The size-matched control
A cold split has fewer training rows than the chronological one, so a difference could be the split or could be the sample size. This control holds size fixed and changes only the split.
| Partition | Rows |
|---|---|
| train | 188 |
| calibration | 121 |
| test | 88 |
| train_with_image | 188 |
| test_with_image | 88 |
| test_with_corridor | 88 |