code-daemon-ner-v1

A small NER model that finds software-engineering entities in prose โ€” the sentences of a README, a design doc, an issue thread, a commit message. In Code-Daemon the spans it emits become the concept nodes that link a paragraph of documentation to the code it talks about.

What makes it different: this is prose NER, not code parsing. A parser already extracts every identifier in a source file, exactly. What it cannot do is read "the IVF rescoring path in faiss_index.zig regressed nDCG@10 by 4 points" and tell you that IVF is an algorithm, faiss_index.zig a file and nDCG@10 a metric. No public model exists for these labels, and the zero-shot ones miss the conventions that matter here โ€” so this one was trained for them, and it runs faster on a CPU than they do on a GPU.

  • ~117M parameters โ€” XLM-RoBERTa, 12 layers ร— 384 hidden, multilingual vocabulary.
  • English and Russian prose; 128 tokens โ€” one or two sentences per call.
  • 2-input ONNX (input_ids, attention_mask) โ†’ logits [batch, seq, 17] (8 types in BIO).

Labels

tier type what examples
Core component a named part of a system, including model names DocStore, mMiniLMv2
Core api_endpoint a callable surface: function, method, route, agent tool name semantic_search, POST /v1/messages
Core file_path a path or file name src/semantic/faiss_index.zig
Core tool something you invoke: executable, runtime, service, language, hardware TensorRT, rclone, SQL
Soft algorithm a named method BM25, Louvain
Soft data_structure a named container or layout CSR, Prolly tree
Soft config a knob, flag, env var journal_mode=MEMORY, --epochs
Soft metric a measurement or unit nDCG@10, p95 latency

Core types are what a knowledge graph links on; Soft types are tolerated at lower precision.

Numbers

1 500 hand-reviewed sentences, entity level (exact span and type). The honest alternative for a private label set is a zero-shot model you hand the label names to โ€” GLiNER, each run at its best of four prompt/threshold settings:

model params Core macro-F1 P / R chunks/s
this model 117M 0.426 0.38 / 0.55 68 on a CPU
gliner-multitask-v1.0 909M 0.284 0.23 / 0.41 10 on a GPU
gliner_multi-v2.1 289M 0.155 0.17 / 0.23 39 on a GPU

Per Core type, F1:

type this model gliner-multitask gliner_multi
api_endpoint 0.521 0.382 0.006
file_path 0.497 0.300 0.248
tool 0.425 0.312 0.286
component 0.263 0.143 0.080
  • Fine-tuning buys the label convention. Where a label name explains itself (file_path, tool), zero-shot gets within ~1.5ร—; where it is a project rule (api_endpoint includes agent tool names), it collapses. If your types are ordinary ones, a zero-shot model may serve you better.
  • The trade: recall over precision. On the strongest types it over-predicts (P 0.39โ€“0.41, R 0.68โ€“0.72). Use it as a candidate generator behind a filter, not as a labeller you trust unattended.
  • Soft types are weak (macro-F1 0.21; config 0.12), and multi-word spans mostly come back as single tokens.

Inside the daemon on TensorRT, a 32-sentence batch takes 3 ms of GPU and 15 ms of tokenisation โ€” parallelise the tokenizer before looking for a faster engine.

How to use it

  • Pre-split identifiers (camelCase / snake_case โ†’ words) before tokenizing and keep seq 128: that is how it was trained, and scoring it any other way gives lower numbers.
  • Calibrated probabilities: divide the logits by T from temperature.json (0.665) before the softmax; argmax alone does not need it.
  • Shapes: TensorRT INT8 single profile seq 128; OpenVINO CPU / iGPU INT8 at batch 32, NPU INT4 at batch 8; Core ML b32 s128.
import onnxruntime as ort, numpy as np, json
from transformers import AutoTokenizer

tok    = AutoTokenizer.from_pretrained(".")            # bundled XLM-R SentencePiece
labels = json.load(open("code-daemon-ner-v1_label_map.json"))["id2label"]
T      = json.load(open("temperature.json"))["T"]
sess   = ort.InferenceSession("model.onnx", providers=["CPUExecutionProvider"])

def tag(text, max_len=128):
    enc = tok(text, return_tensors="np", truncation=True, max_length=max_len,
              return_offsets_mapping=True, return_token_type_ids=False)
    logits = sess.run(None, {"input_ids":      enc["input_ids"].astype(np.int64),
                             "attention_mask": enc["attention_mask"].astype(np.int64)})[0][0]
    probs  = np.exp(logits / T) / np.exp(logits / T).sum(-1, keepdims=True)   # calibrated
    for (a, b), row in zip(enc["offset_mapping"][0], probs):
        lab = labels[str(int(row.argmax()))]
        if lab != "O" and b > a:
            yield text[a:b], lab, float(row.max())

for span, label, p in tag("The IVF rescoring path in src/semantic/faiss_index.zig cost 4 nDCG@10 points."):
    print(f"{span!r:40} {label:18} {p:.2f}")

Out of scope: symbols in source code (use a parser), general-domain NER, any high-stakes automatic decision.

How it was made

  • Backbone: mMiniLMv2-L12-H384, a multilingual MiniLM distilled from XLM-RoBERTa-Large.
  • Labels by weak supervision: with no corpus for this label set, several independent labellers voted on each chunk and the votes were reconciled into one BIO sequence; the backbone was then fine-tuned on them and its output temperature-calibrated.
  • Evaluation is human-labelled and disjoint from the training labels.

Files

file what it is
model.onnx FP32, dynamic shape โ€” the standalone path
model_static.onnx, model_int8qdt.onnx seq pinned to 128 / INT8 Q/DQ โ€” engine-build inputs
model.safetensors, config.json the weights as XLMRobertaForTokenClassification; AutoModelForTokenClassification loads them as-is; the Apple (MLX) build is prepared from this pair
sentencepiece.bpe.model, tokenizer_config.json, code-daemon-ner-v1_tokenizer.json the tokenizer (SentencePiece + fast tokenizer)
code-daemon-ner-v1_label_map.json the 17 BIO classes
temperature.json the calibration scalar T
eval_metrics.json, manifest.json per-type scores; build metadata
code-daemon-ner-v1_{win_x64,linux_x64}_trt11.0_sm_{75,80,86,89,120}.engine TensorRT INT8 per OS ร— GPU; sm_90 (H100/H200) Linux only
code-daemon-ner-v1_ov2026.4_{cpu,igpu_lnl}_int8_b32_s128.{xml,bin}, โ€ฆ_npu_int4_b8_s128.* OpenVINO 2026.4 for Intel CPU, iGPU, NPU
coreml_ane/embed.mlpackage/ Core ML for the Apple Neural Engine, function b32 s128; load with cpuAndNeuralEngine

A TensorRT engine loads only on the GPU architecture and TensorRT version that built it; any other machine uses OpenVINO or the ONNX.

License

Released under the MIT license. The backbone is a re-upload of Microsoft's mMiniLMv2 (microsoft/unilm, MIT); the re-upload repository itself declares no license. Neither the training corpus nor the evaluation set ships here. Not legal advice.

Attribution

Backbone: nreimers/mMiniLMv2-L12-H384-distilled-from-XLMR-Large (itself distilled from XLM-RoBERTa-Large).

Downloads last month
128
Safetensors
Model size
0.1B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for faxenoff/code-daemon-ner-v1

Quantized
(3)
this model