Alpha (System-One-style typed decision model)

Alpha takes unstructured / structured state plus typed questions with declared options and returns calibrated probabilities for all questions in one parallel forward pass. It cannot generate strings and cannot answer outside the declared options (masked softmax -> schema violations are impossible by construction). Option embeddings are independent of the state, so a schema can be compiled once and cached; cardinality is limited only by memory.

Architecture: distilroberta-base encoder (state + question slots with <mask> readout tokens) -> 2-layer Slot Reasoner (set transformer, squared-ReLU, zero-init residuals) -> biaffine/cosine scorer against option embeddings with per-question learned temperature + post-hoc temperature (T=1.04).

Training: 6000 steps x 32 workflow-style examples (1-4 questions over named fields, mixed from 10 public datasets with paraphrased questions/option wordings), proper-scoring-rule loss (NLL + Brier + ordinal term), EMA weights, functional-memory anchor, discovery controller. Held-out families (never trained on): cola, dbpedia, offensive.

Results (test splits; see results.json, alpha_report.png)

family acc (seen wording) acc (unseen wording) ECE
ag_news 0.926 0.924 0.032
banking77 0.902 0.890 0.036
emotion 0.922 0.894 0.030
imdb 0.886 0.892 0.027
mrpc 0.822 0.822 0.042
news20 0.674 0.678 0.068
sms 0.993 0.346 0.004
snli 0.772 0.538 0.046
sst2 0.912 0.910 0.034
yelp 0.572 0.522 0.072

Workflow eval (4 questions per query): per-question acc 0.825, all-correct 0.437, joint-confidence ECE 0.101.

Held-out task families (unseen schema wording): zero-shot / after 10 / after 40 gradient steps

family zero-shot +10 steps +40 steps
cola 0.374 0.658 0.686
dbpedia 0.312 0.586 0.952
offensive 0.528 0.632 0.694

Speed on Tesla T4: p50 11.1 ms, p95 13.5 ms for a 3-question query (local, no network); 354 workflows/s at batch 64. Qwen2.5-0.5B-Instruct (generative, same 120 questions): acc 0.583, p50 latency 82 ms vs Alpha acc 0.875, p50 11.1 ms.

Framework diagnostics: DMR=1.08; FRR by perturbation: sigma=0.5: FRR 4250.0 sigma=1: FRR 4250.0 sigma=2: FRR 4250.0 sigma=4: FRR 4250.0 sigma=8: FRR 14.2* (* = not re-reached within budget); MGR by checkpoint in results.json.

Honest limits

  • This is a small research model trained on public classification datasets. It is not a replacement for, and has not been compared against, TypeSafe's Jev (a closed API; its claims were not verified here). Do not read the latency comparison as a Jev comparison.
  • Narrow task coverage (sentiment/topic/intent/NLI/paraphrase/spam/rating); free-form reasoning is out of scope by design.
  • Calibration is measured in-distribution; shift will degrade it.

Usage

from huggingface_hub import snapshot_download; import sys
p = snapshot_download("SofiTesfay2010/Alpha"); sys.path.insert(0, p)
from alpha_s1 import Alpha
m = Alpha.from_pretrained(p, device="cuda").eval()
print(m.predict({"ticket": "I was charged twice, refund me now."},
      [{"id": "refund", "question": "Does the customer want a refund?", "options": ["no", "yes"]}]))
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.4B params
Tensor type
F32
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SofiTesfay2010/Alpha

Finetuned
(787)
this model