steeldb-models
Weights for hypersteeldb. The crate resolves these at runtime and
never downloads them automatically โ if a model is missing it returns an error naming the repository,
revision and size, and you fetch it deliberately.
huggingface-cli download cp500/steeldb-models --local-dir ~/.steeldb/models
# or point the crate at a directory you already have
export STEELDB_ML_BUNDLE=/path/to/step0_bundle_ml
export STEELDB_MODEL2VEC=/path/to/model2vec
What is here
| directory | what it is | provenance | licence |
|---|---|---|---|
step0_bundle_ml/ |
SPO span tagger (ONNX). Reads a sentence and marks typed spans โ ENT, GEO, REL, TIME, QTY, IGNORE. |
trained by this project | MIT |
model2vec/ |
static embedder, 128-dim, used to place spans before clustering | mirror of minishlab/potion-base-4M |
MIT |
The tagger
step0_bundle_ml/spo.onnx is ours, trained from scratch on 230 synthetic sentences written for the task โ
publicly-known defence and aerospace subject matter (airframes, ISO standards, sensor types). It contains no
customer or private corpus material. It is a span tagger: it learns where spans begin and end and what kind
they are, not the content of any document.
labels.json holds the BIO label set; tokenizer/ is the matching wordpiece tokenizer.
The embedder
model2vec/ is a mirror of minishlab/potion-base-4M, redistributed under MIT as that licence permits.
Credit belongs to the Minish Lab, not to this project. It is mirrored so
a build does not break if the upstream repository moves, and meta.json records the upstream id it came from.
Prefer the original if you only need the embedder.
Why these two together
The crate's default path needs no models at all โ vocabulary discovery runs on word statistics and compiles to WebAssembly. These weights enable the neural discovery path, which finds multi-word domain entities that a keyword scan cannot: the tagger marks typed spans, the embedder places them, and optimal transport groups them per kind into a candidate codebook. Neither model targets wasm32 (ONNX Runtime, and a tokenizer that links a C library), so that path is native-only.
Discovery runs once, over a sample. What it produces is a small set of vocabulary files, and every run after that follows those files with no model involved.