RadMail

Importance engine — live eval

The engine scored 22 synthetic emails (deterministic, IP-scrubbed) and was graded against ground-truth labels derived from revealed-preference behavior. High-confidence labels only (≥ 0.6). This is the “99%” gauge.

Full engine (two-axis + behavioral learning)

F1
100.0%
headline 100.0%
Precision
100.0%
Recall
100.0%
false-negative-averse
Important in low band
0
the lethal failure — buried important mail
Accuracy
100.0%
True positives
13
False positives
0
False negatives
0

What the behavioral layer buys (baseline → engine)

MetricContent-only baselineFull engine
F170.0%100.0%
Recall53.8%100.0%
Important buried in ‘low’50

The behavioral layer lifts content-quiet-but-behaviorally-important mail (a replied-to counterparty, a reply that closes a waiting-on) out of ‘low’ without over-surfacing the negative control — driving buried-important to zero.

Band confusion (engine)

BandTruly importantTruly not
critical50
high80
normal00
low07

At-K — did any important email fall below the fold?

KPrecision@KRecall@KMissed
5100.0%38.5%8
10100.0%76.9%3
2065.0%100.0%0

Deterministic run anchored at 2026-06-19T18:00:00.000Z. Pure-function engine — no DB, no network. Synthetic fixtures, scrubbed of any real tenant data.