Customer mood, judged right.
300 mood calls drawn from 900 unseen support tickets. Reference scores included alongside for comparison.
- SemIf-144 accuracy
- 55.6%
- Queue routing (4-way)
- 70.0%
- Priority scoring
- 42.7%
- Per decision, hosted
- ~300 ms
Benchmarked on 900 unseen tickets
Independent OOD bench (source): 300 queue-routing choices, 300 mood booleans, 300 priority scores. Reflex ran the full suite in September 2026. Bars show accuracy; ECE figures are listed below.
SemIf-authored144 (144 semantic-judgment rows): Reflex 55.6% accuracy, 54.6% balanced, 144/144 valid responses — chance is 33.3%. Full runs: OOD-900 at topK=1 greedy, α=1.0, K=5 single-pass; ±2 pts is sampling noise.
Run a decision here
Send a live POST /v1/classifier to the hosted worker. Keys stay in this browser.
Mint a key. Keep it somewhere safe.
Keys are free during the open demo. One per signup, shown once, stored hashed.
Two ways to watch it decide.
Reflex Pilot
A city car driven entirely by classifier calls — every steering choice streams from this API at ~300 ms. It drove a full 735 m route with zero contacts.
735 m · 0 contacts · J to engage
Play driving2048
One board, one classifier call per move. The model reached tile 64 in 74 moves with zero invalid answers. Play manually with arrows, or let it play.
74 moves · tile 64 · arrows to play
Play 2048Tetris
One piece, one classifier call, one placement chosen from a shortlist. The model cleared 5 lines across a 49-piece game in a real browser.
49 pieces · 5 lines · Space to drop
Play TetrisSimulation, physics, game rules, and artwork © their authors; credits ship in each game and in the project’s attribution files.