code-daemon-ner-v1
A small NER model that finds software-engineering entities in prose โ the sentences of a README, a design doc, an issue thread, a commit message. In Code-Daemon the spans it emits become the concept nodes that link a paragraph of documentation to the code it talks about.
What makes it different: this is prose NER, not code parsing. A parser already extracts every
identifier in a source file, exactly. What it cannot do is read "the IVF rescoring path in
faiss_index.zig regressed nDCG@10 by 4 points" and tell you that IVF is an algorithm,
faiss_index.zig a file and nDCG@10 a metric. No public model exists for these labels, and the
zero-shot ones miss the conventions that matter here โ so this one was trained for them, and it
runs faster on a CPU than they do on a GPU.
- ~117M parameters โ XLM-RoBERTa, 12 layers ร 384 hidden, multilingual vocabulary.
- English and Russian prose; 128 tokens โ one or two sentences per call.
- 2-input ONNX (
input_ids,attention_mask) โlogits [batch, seq, 17](8 types in BIO).
Labels
| tier | type | what | examples |
|---|---|---|---|
| Core | component |
a named part of a system, including model names | DocStore, mMiniLMv2 |
| Core | api_endpoint |
a callable surface: function, method, route, agent tool name | semantic_search, POST /v1/messages |
| Core | file_path |
a path or file name | src/semantic/faiss_index.zig |
| Core | tool |
something you invoke: executable, runtime, service, language, hardware | TensorRT, rclone, SQL |
| Soft | algorithm |
a named method | BM25, Louvain |
| Soft | data_structure |
a named container or layout | CSR, Prolly tree |
| Soft | config |
a knob, flag, env var | journal_mode=MEMORY, --epochs |
| Soft | metric |
a measurement or unit | nDCG@10, p95 latency |
Core types are what a knowledge graph links on; Soft types are tolerated at lower precision.
Numbers
1 500 hand-reviewed sentences, entity level (exact span and type). The honest alternative for a private label set is a zero-shot model you hand the label names to โ GLiNER, each run at its best of four prompt/threshold settings:
| model | params | Core macro-F1 | P / R | chunks/s |
|---|---|---|---|---|
| this model | 117M | 0.426 | 0.38 / 0.55 | 68 on a CPU |
| gliner-multitask-v1.0 | 909M | 0.284 | 0.23 / 0.41 | 10 on a GPU |
| gliner_multi-v2.1 | 289M | 0.155 | 0.17 / 0.23 | 39 on a GPU |
Per Core type, F1:
| type | this model | gliner-multitask | gliner_multi |
|---|---|---|---|
api_endpoint |
0.521 | 0.382 | 0.006 |
file_path |
0.497 | 0.300 | 0.248 |
tool |
0.425 | 0.312 | 0.286 |
component |
0.263 | 0.143 | 0.080 |
- Fine-tuning buys the label convention. Where a label name explains itself (
file_path,tool), zero-shot gets within ~1.5ร; where it is a project rule (api_endpointincludes agent tool names), it collapses. If your types are ordinary ones, a zero-shot model may serve you better. - The trade: recall over precision. On the strongest types it over-predicts (P 0.39โ0.41, R 0.68โ0.72). Use it as a candidate generator behind a filter, not as a labeller you trust unattended.
- Soft types are weak (macro-F1 0.21;
config0.12), and multi-word spans mostly come back as single tokens.
Inside the daemon on TensorRT, a 32-sentence batch takes 3 ms of GPU and 15 ms of tokenisation โ parallelise the tokenizer before looking for a faster engine.
How to use it
- Pre-split identifiers (
camelCase/snake_caseโ words) before tokenizing and keep seq 128: that is how it was trained, and scoring it any other way gives lower numbers. - Calibrated probabilities: divide the logits by
Tfromtemperature.json(0.665) before the softmax;argmaxalone does not need it. - Shapes: TensorRT INT8 single profile seq 128; OpenVINO CPU / iGPU INT8 at batch 32, NPU INT4
at batch 8; Core ML
b32 s128.
import onnxruntime as ort, numpy as np, json
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained(".") # bundled XLM-R SentencePiece
labels = json.load(open("code-daemon-ner-v1_label_map.json"))["id2label"]
T = json.load(open("temperature.json"))["T"]
sess = ort.InferenceSession("model.onnx", providers=["CPUExecutionProvider"])
def tag(text, max_len=128):
enc = tok(text, return_tensors="np", truncation=True, max_length=max_len,
return_offsets_mapping=True, return_token_type_ids=False)
logits = sess.run(None, {"input_ids": enc["input_ids"].astype(np.int64),
"attention_mask": enc["attention_mask"].astype(np.int64)})[0][0]
probs = np.exp(logits / T) / np.exp(logits / T).sum(-1, keepdims=True) # calibrated
for (a, b), row in zip(enc["offset_mapping"][0], probs):
lab = labels[str(int(row.argmax()))]
if lab != "O" and b > a:
yield text[a:b], lab, float(row.max())
for span, label, p in tag("The IVF rescoring path in src/semantic/faiss_index.zig cost 4 nDCG@10 points."):
print(f"{span!r:40} {label:18} {p:.2f}")
Out of scope: symbols in source code (use a parser), general-domain NER, any high-stakes automatic decision.
How it was made
- Backbone:
mMiniLMv2-L12-H384, a multilingual MiniLM distilled from XLM-RoBERTa-Large. - Labels by weak supervision: with no corpus for this label set, several independent labellers voted on each chunk and the votes were reconciled into one BIO sequence; the backbone was then fine-tuned on them and its output temperature-calibrated.
- Evaluation is human-labelled and disjoint from the training labels.
Files
| file | what it is |
|---|---|
model.onnx |
FP32, dynamic shape โ the standalone path |
model_static.onnx, model_int8qdt.onnx |
seq pinned to 128 / INT8 Q/DQ โ engine-build inputs |
model.safetensors, config.json |
the weights as XLMRobertaForTokenClassification; AutoModelForTokenClassification loads them as-is; the Apple (MLX) build is prepared from this pair |
sentencepiece.bpe.model, tokenizer_config.json, code-daemon-ner-v1_tokenizer.json |
the tokenizer (SentencePiece + fast tokenizer) |
code-daemon-ner-v1_label_map.json |
the 17 BIO classes |
temperature.json |
the calibration scalar T |
eval_metrics.json, manifest.json |
per-type scores; build metadata |
code-daemon-ner-v1_{win_x64,linux_x64}_trt11.0_sm_{75,80,86,89,120}.engine |
TensorRT INT8 per OS ร GPU; sm_90 (H100/H200) Linux only |
code-daemon-ner-v1_ov2026.4_{cpu,igpu_lnl}_int8_b32_s128.{xml,bin}, โฆ_npu_int4_b8_s128.* |
OpenVINO 2026.4 for Intel CPU, iGPU, NPU |
coreml_ane/embed.mlpackage/ |
Core ML for the Apple Neural Engine, function b32 s128; load with cpuAndNeuralEngine |
A TensorRT engine loads only on the GPU architecture and TensorRT version that built it; any other machine uses OpenVINO or the ONNX.
License
Released under the MIT license. The backbone is a re-upload of Microsoft's mMiniLMv2 (microsoft/unilm, MIT); the re-upload repository itself declares no license. Neither the training corpus nor the evaluation set ships here. Not legal advice.
Attribution
Backbone: nreimers/mMiniLMv2-L12-H384-distilled-from-XLMR-Large (itself distilled from XLM-RoBERTa-Large).
- Downloads last month
- 128