← RadMail
Importance engine — live eval
The engine scored 22 synthetic emails (deterministic, IP-scrubbed) and was graded against ground-truth labels derived from revealed-preference behavior. High-confidence labels only (≥ 0.6). This is the “99%” gauge.
Full engine (two-axis + behavioral learning)
F1
100.0%
headline 100.0%
Precision
100.0%
Recall
100.0%
false-negative-averse
Important in low band
0
the lethal failure — buried important mail
Accuracy
100.0%
True positives
13
False positives
0
False negatives
0
What the behavioral layer buys (baseline → engine)
| Metric | Content-only baseline | Full engine |
|---|---|---|
| F1 | 70.0% | 100.0% |
| Recall | 53.8% | 100.0% |
| Important buried in ‘low’ | 5 | 0 |
The behavioral layer lifts content-quiet-but-behaviorally-important mail (a replied-to counterparty, a reply that closes a waiting-on) out of ‘low’ without over-surfacing the negative control — driving buried-important to zero.
Band confusion (engine)
| Band | Truly important | Truly not |
|---|---|---|
| critical | 5 | 0 |
| high | 8 | 0 |
| normal | 0 | 0 |
| low | 0 | 7 |
At-K — did any important email fall below the fold?
| K | Precision@K | Recall@K | Missed |
|---|---|---|---|
| 5 | 100.0% | 38.5% | 8 |
| 10 | 100.0% | 76.9% | 3 |
| 20 | 65.0% | 100.0% | 0 |
Deterministic run anchored at 2026-06-19T18:00:00.000Z. Pure-function engine — no DB, no network. Synthetic fixtures, scrubbed of any real tenant data.