Zero-Shot Classification
Transformers
Safetensors
qwen3_5
feature-extraction
decision-model
classification
system-one
multimodal
vision
custom_code
Instructions to use vllm-sr/d3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vllm-sr/d3 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("zero-shot-classification", model="vllm-sr/d3", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("vllm-sr/d3", trust_remote_code=True) model = AutoModel.from_pretrained("vllm-sr/d3", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
d3
d3 is the 27B multimodal foundation decision model of Decision 3.0, the decision models of vLLM Semantic Router. Give it an input (text or JSON, optionally with images) and the questions you need answered: pick one of several options, say yes or no, or rate on a scale. It answers them all in one call and returns a probability for every answer, without generating text.
| Parameters | 26.09B, including the 0.46B vision encoder |
| Inputs | Text or JSON, plus images (several per request) |
| Decision types | Choice · Yes / No · Score |
| License | Apache-2.0 |
Highlights
- Jev Decision Index 0.3, public suite: 64.24, measured with the official 0.3 kit on the released weights: all 140,178 public requests answered, none unsupported. The official Full scores on the text and vision boards are pending official evaluation.
- +7.3 on the public suite over Decision 2.0 (its 27B model: 56.97 on the board), ahead in all five areas.
- Reads images: multiple images per request (PNG, JPEG or WebP), given as paths, URLs, PIL images or base64 data URLs; every question of the request sees all of them.
- Speed: text requests take a median of 108.7 ms (mean 385.8 ms, 80th percentile 270.7 ms); requests with an image a median of 406.9 ms (mean 403.9 ms, 80th percentile 408.9 ms). One NVIDIA RTX PRO 6000, one request at a time.
- Many questions, one call: Choice, Yes / No and Score questions about the same input are answered together, each from its own forward pass over the input, with a probability for every option.
Quickstart
pip install "transformers==5.17.0" torch torchvision pillow safetensors accelerate
pip install flash-linear-attention # optional: fast GPU kernels for the linear-attention layers
import json
from huggingface_hub import hf_hub_download
from transformers import AutoModel
model = AutoModel.from_pretrained("vllm-sr/d3", trust_remote_code=True)
# Text
result = model.system_one(
state="The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",
questions={
"route": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"returns": "Refunds, replacements and damaged deliveries",
"billing": "Payments, invoices and charges",
"technical": "Product setup and faults"
}
},
"receipt": {
"type": "noul",
"instructions": "Does the customer have a receipt?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is this request?",
"criteria": [
"Routine",
"Soon",
"Today"
]
}
},
)
print(json.dumps(result["answers"], indent=2))
# Text and an image (or several)
receipt = hf_hub_download("vllm-sr/d3", "assets/example-receipt.png")
result = model.system_one(
state="The customer says the blender arrived cracked and attached the receipt.",
images=[receipt], # local paths, http(s) URLs, PIL images or base64 data URLs
questions={
"route": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"returns": "Refunds, replacements and damaged deliveries",
"billing": "Payments, invoices and charges",
"technical": "Product setup and faults"
}
},
"on_receipt": {
"type": "noul",
"instructions": "Does the receipt list the blender?"
},
"payment": {
"type": "choice",
"instructions": "How was the order paid?",
"criteria": {
"card": None,
"cash": None,
"gift card": None
}
}
},
)
print(json.dumps(result["answers"], indent=2))
# Or as a pipeline:
# transformers.pipeline("decision", model="vllm-sr/d3", trust_remote_code=True)(state=..., questions=..., images=...)
Images go before the text of the request, each read at up to 1.6 megapixels; every question of the request sees all of them.
Evaluation
Text: Jev Decision Index 0.3.1
| Model | Jev Decision Index ↑ | Public ↑ | Same-skill tests ↑ | New-domain tasks ↑ |
|---|---|---|---|---|
| d3 | 63.7 | 64.2 | 61.4 | 56.7 |
| Perplexity Decider v1.1 (27B) | 62.8 | 62.3 | 61.1 | 55.6 |
| Fastino GLiDE no-thinking (28B) | 60.2 | 59.1 | 59.5 | 52.9 |
| Jev | 60.1 | 58.0 | 58.0 | 55.0 |
| Torchcast Decision 27B | 59.9 | 65.1 | 58.1 | 50.8 |
| Decision 2.0 (27B) | 55.9 | 57.0 | 55.7 | 47.9 |
Images: Jev Decision Index vision board 0.3.1
| Model | Vision Index ↑ | Public ↑ | Private ↑ |
|---|---|---|---|
| d3 | 70.9 | 73.2 | 68.6 |
| Perplexity Decider v1.1 (27B) | 70.6 | 73.2 | 67.9 |
| JEV-27B-VL | 69.6 | 72.8 | 66.4 |
| Solomon v1.1 (27B) | 66.8 | 69.9 | 63.8 |
d3 on the public vision benchmarks († approximate rebuild):
| Benchmark | d3 |
|---|---|
| CV-Bench | 75.9 |
| BLINK | 57.6 |
| RealWorldQA | 70.9 |
| CharXiv † | 91.7 |
| InfographicVQA † | 95.6 |
| Mind2Web † | 89.3 |
| Winoground | 85.8 |
| KIE (CORD+FUNSD) † | 97.7 |
| Moderation (Hateful Memes) | 42.3 |
| R-Bench-M | 32.1 |
| MMMU-Pro vision | 46.1 |
d3: internal evaluation. Others: live board data, text 2026-10-10, vision 2026-10-09.
License
Apache-2.0 (LICENSE). Built on Qwen/Qwen3.8-27B (Apache-2.0).
Citation
@misc{d3_2026,
title = {{d3}: A Multimodal Foundation Decision Model},
author = {{vLLM Semantic Router Team}},
year = {2026},
howpublished = {\url{https://huggingface.co/vllm-sr/d3}}
}
- Downloads last month
- -
Model tree for vllm-sr/d3
Base model
Qwen/Qwen3.8-27B

