75.3%

Customer mood, judged right.

300 angry-customer calls out of 900 support tickets no model had seen. The reference system scores 91.7% — the gap is published below, not buried.

SemIf-144 accuracy
55.6%
Queue routing (4-way)
70.0%
Priority scoring
42.7%
Per decision, local
~300 ms

Reflex vs Jev, same 900 tickets

Independent OOD bench by scienthoon/jev-ood-calibration: 300 queue-routing choices, 300 mood booleans, 300 priority scores. Reflex ran it end to end, September 2026. Bars are accuracy; the thin line is ECE (lower is better).

Reflex, local backend~300 ms / decisionmeasured, 272–396 ms range
Reflex, hosted worker~1.1 s / decisionmeasured, includes upstream inference
Jev, claimed70–500 ms / answervendor claim, not re-measured

SemIf-authored144 (144 semantic-judgment rows): Reflex 55.6% accuracy, 54.6% balanced, 144/144 valid responses — chance is 33.3%. Full runs: OOD-900 at topK=1 greedy, α=1.0, K=5 single-pass; ±2 pts is upstream sampling noise.

Run a decision here

Live POST /v1/classifier against the hosted worker. Paste the key you minted below — nothing is stored, nothing is shared.

response
{"hint": "mint a key below, paste it, run"}

Mint a key. Keep it somewhere safe.

Keys are free while the upstream demo lasts. One per signup, shown once, stored hashed.

The model drives.

JevPilot — Standard Agents’ sim, Featherless’ Simple-Jev adaptation, Reflex edition — drove a full 735 m city route on autopilot with zero contacts. Every steering call streams from this API at ~300 ms.

Play the Reflex edition

Simulation, physics, and artwork © their authors; see in-game credits and pilot/ATTRIBUTION.md.