Customer mood, judged right.
300 angry-customer calls out of 900 support tickets no model had seen. The reference system scores 91.7% — the gap is published below, not buried.
- SemIf-144 accuracy
- 55.6%
- Queue routing (4-way)
- 70.0%
- Priority scoring
- 42.7%
- Per decision, local
- ~300 ms
Reflex vs Jev, same 900 tickets
Independent OOD bench by scienthoon/jev-ood-calibration: 300 queue-routing choices, 300 mood booleans, 300 priority scores. Reflex ran it end to end, September 2026. Bars are accuracy; the thin line is ECE (lower is better).
SemIf-authored144 (144 semantic-judgment rows): Reflex 55.6% accuracy, 54.6% balanced, 144/144 valid responses — chance is 33.3%. Full runs: OOD-900 at topK=1 greedy, α=1.0, K=5 single-pass; ±2 pts is upstream sampling noise.
Run a decision here
Live POST /v1/classifier against the hosted worker. Paste the key you minted below — nothing is stored, nothing is shared.
Mint a key. Keep it somewhere safe.
Keys are free while the upstream demo lasts. One per signup, shown once, stored hashed.
The model drives.
JevPilot — Standard Agents’ sim, Featherless’ Simple-Jev adaptation, Reflex edition — drove a full 735 m city route on autopilot with zero contacts. Every steering call streams from this API at ~300 ms.
Play the Reflex editionSimulation, physics, and artwork © their authors; see in-game credits and pilot/ATTRIBUTION.md.