Alpha (System-One-style typed decision model)
Alpha takes unstructured / structured state plus typed questions with declared options and returns calibrated probabilities for all questions in one parallel forward pass. It cannot generate strings and cannot answer outside the declared options (masked softmax -> schema violations are impossible by construction). Option embeddings are independent of the state, so a schema can be compiled once and cached; cardinality is limited only by memory.
Architecture: distilroberta-base encoder (state + question slots with <mask> readout tokens) -> 2-layer Slot Reasoner (set transformer, squared-ReLU,
zero-init residuals) -> biaffine/cosine scorer against option embeddings with per-question learned temperature + post-hoc temperature (T=1.04).
Training: 6000 steps x 32 workflow-style examples (1-4 questions over named fields, mixed from 10 public datasets with paraphrased questions/option wordings), proper-scoring-rule loss (NLL + Brier + ordinal term), EMA weights, functional-memory anchor, discovery controller. Held-out families (never trained on): cola, dbpedia, offensive.
Results (test splits; see results.json, alpha_report.png)
| family | acc (seen wording) | acc (unseen wording) | ECE |
|---|---|---|---|
| ag_news | 0.926 | 0.924 | 0.032 |
| banking77 | 0.902 | 0.890 | 0.036 |
| emotion | 0.922 | 0.894 | 0.030 |
| imdb | 0.886 | 0.892 | 0.027 |
| mrpc | 0.822 | 0.822 | 0.042 |
| news20 | 0.674 | 0.678 | 0.068 |
| sms | 0.993 | 0.346 | 0.004 |
| snli | 0.772 | 0.538 | 0.046 |
| sst2 | 0.912 | 0.910 | 0.034 |
| yelp | 0.572 | 0.522 | 0.072 |
Workflow eval (4 questions per query): per-question acc 0.825, all-correct 0.437, joint-confidence ECE 0.101.
Held-out task families (unseen schema wording): zero-shot / after 10 / after 40 gradient steps
| family | zero-shot | +10 steps | +40 steps |
|---|---|---|---|
| cola | 0.374 | 0.658 | 0.686 |
| dbpedia | 0.312 | 0.586 | 0.952 |
| offensive | 0.528 | 0.632 | 0.694 |
Speed on Tesla T4: p50 11.1 ms, p95 13.5 ms for a 3-question query (local, no network); 354 workflows/s at batch 64. Qwen2.5-0.5B-Instruct (generative, same 120 questions): acc 0.583, p50 latency 82 ms vs Alpha acc 0.875, p50 11.1 ms.
Framework diagnostics: DMR=1.08; FRR by perturbation: sigma=0.5: FRR 4250.0 sigma=1: FRR 4250.0 sigma=2: FRR 4250.0 sigma=4: FRR 4250.0 sigma=8: FRR 14.2* (* = not re-reached within budget); MGR by checkpoint in results.json.
Honest limits
- This is a small research model trained on public classification datasets. It is not a replacement for, and has not been compared against, TypeSafe's Jev (a closed API; its claims were not verified here). Do not read the latency comparison as a Jev comparison.
- Narrow task coverage (sentiment/topic/intent/NLI/paraphrase/spam/rating); free-form reasoning is out of scope by design.
- Calibration is measured in-distribution; shift will degrade it.
Usage
from huggingface_hub import snapshot_download; import sys
p = snapshot_download("SofiTesfay2010/Alpha"); sys.path.insert(0, p)
from alpha_s1 import Alpha
m = Alpha.from_pretrained(p, device="cuda").eval()
print(m.predict({"ticket": "I was charged twice, refund me now."},
[{"id": "refund", "question": "Does the customer want a refund?", "options": ["no", "yes"]}]))
Model tree for SofiTesfay2010/Alpha
Base model
distilbert/distilroberta-base