Post
68
JEV-27B-VL now runs on a 32 GB Mac š
AutoTrust's JEV answers typed questions (yes/no, pick one of up to 256 options, a 0ā5 score) about text and images with a calibrated probability per option, in a single forward pass. The model card asks for an 80 GB GPU; with MLX it peaks at 16.9 GB and takes ~1.5 s per decision on an M2 Pro.
What's in it:
⢠4-bit Qwen3.8-27B base (byte-identical to mlx-community's) + JEV's System 1 LoRA kept unmerged in bf16 + a float32 decision head
⢠Same checkpoint = plain Qwen3.8-27B when the adapter is off (System 2)
⢠Images in the state, 2ā256 options
Parity, measured against the official PyTorch math on JEV-9B: bf16 39/39 decisions, max probability diff 0.008 (inside the noise between the two official reference paths); 8-bit indistinguishable; 4-bit keeps every clear decision but moves probabilities up to 0.12. The 27B 4-bit build matches the outputs published on the JEV cards within 0.003. Full report in the repo.
Weights: Bayway/JEV-27B-VL-MLX-4bit
mlx-vlm support (under review): https://github.com/Blaizzy/mlx-vlm/pull/2463
All credit for the model to AutoTrust ( autotrust/JEV-27B-VL); thanks to the mlx-vlm and mlx-community folks for the decision API and the base quant.
AutoTrust's JEV answers typed questions (yes/no, pick one of up to 256 options, a 0ā5 score) about text and images with a calibrated probability per option, in a single forward pass. The model card asks for an 80 GB GPU; with MLX it peaks at 16.9 GB and takes ~1.5 s per decision on an M2 Pro.
What's in it:
⢠4-bit Qwen3.8-27B base (byte-identical to mlx-community's) + JEV's System 1 LoRA kept unmerged in bf16 + a float32 decision head
⢠Same checkpoint = plain Qwen3.8-27B when the adapter is off (System 2)
⢠Images in the state, 2ā256 options
Parity, measured against the official PyTorch math on JEV-9B: bf16 39/39 decisions, max probability diff 0.008 (inside the noise between the two official reference paths); 8-bit indistinguishable; 4-bit keeps every clear decision but moves probabilities up to 0.12. The 27B 4-bit build matches the outputs published on the JEV cards within 0.003. Full report in the repo.
from mlx_vlm import load, predict
model, processor = load("Bayway/JEV-27B-VL-MLX-4bit")
predict(model, processor, "Customer: my card was charged twice for one coffee.",
{"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": ["billing", "shipping", "tech support"]}})
# billing 0.998 Ā· tech support 0.002Weights: Bayway/JEV-27B-VL-MLX-4bit
mlx-vlm support (under review): https://github.com/Blaizzy/mlx-vlm/pull/2463
All credit for the model to AutoTrust ( autotrust/JEV-27B-VL); thanks to the mlx-vlm and mlx-community folks for the decision API and the base quant.