AI model formats¶
See also: AI model scan limits for why a model can be missing metadata, and Metadata provenance for the annotations named below.
Use this when you want to know which AI model file formats loom model
reads, what each format can hold, what Pitloom takes from it, and where
each value ends up in the SBOM.
Quick guide¶
pip install "pitloom[ai]"
loom model path/to/model.safetensors -o model.spdx3.json
pip install "pitloom[ai]" installs every optional reader library below;
install a single extra (pitloom[gguf], ...) if you need one format only.
Every reader is read-only: Pitloom never runs model code and never calls
pickle.load(). See Command line for loom model and the other
targets that scan for models.
How a file is recognised¶
A file is a model when its header confirms a format. The header check comes
first, so a file named on the command line is read whatever its suffix (a
x.dat with GGUF magic is GGUF); a project or wheel scan only looks at the
suffixes below. A header that contradicts the suffix (text named .gguf)
is not a model, and a Git LFS pointer is never one; see
What is a model.
| Format | Suffixes scanned | Header check | Reads | Library (extra) |
|---|---|---|---|---|
| CRFsuite | .crfsuite, .model |
magic lCRF |
header and label strings | none |
| fastText | .ftz, .bin |
magic BA 16 4F 2F |
whole model, in memory | fasttext-community (fasttext) |
| GGUF | .gguf |
magic GGUF |
key-value header (memory-mapped) | gguf (gguf) |
| HDF5, Keras v1-v2 | .h5, .hdf5 |
HDF5 magic, or any non-empty file (a userblock moves the signature) | root attributes | h5py (hdf5) |
| Keras v3 | .keras |
ZIP | config.json, metadata.json |
none |
| NumPy | .npy, .npz |
.npy: magic \x93NUMPY; .npz: ZIP |
array headers | numpy (numpy) |
| ONNX | .onnx |
none: any non-empty file | whole protobuf, in memory; external data files not loaded | onnx (onnx) |
| PyTorch classic | .pt, .pth |
ZIP, or pickle protocol 2-5 (a text .pth path file is not a model) |
archive listing, top of data.pkl |
fickling (pytorch), optional: without it, no type of model |
| PT2 / ExecuTorch | .pt2 |
ZIP | archive listing, small text and JSON members | none |
| Safetensors | .safetensors |
8-byte little-endian header size, then { |
JSON header | safetensors (safetensors) |
| Hugging Face Hub | model ID or URL, not a file | -- | Hub API, model card, config files | huggingface_hub (huggingface_hub) |
An HDF5 file with a keras_version attribute is reported as format
keras, otherwise as hdf5 (and then only its format is recorded).
.bin and .model are shared with other tools: such a file is a model only
when it carries a supported format's magic (usually fastText or CRFsuite), and no
warning is given otherwise.
What each format holds¶
One row per format. Yes: read into an SBOM field. Kept: read, and kept only in the verbatim artifact-metadata annotation (see Where each field goes). Part: partly; see the note. Fixed: a constant Pitloom sets for the format. No: the format can hold it, Pitloom does not read it. --: the format has no such field.
| Format | Name | Description | Version | Licence | Type, architecture | Hyperparameters | Inputs, outputs | Labels | Training, evaluation | Other keys |
|---|---|---|---|---|---|---|---|---|---|---|
| CRFsuite | -- | Part [1] | -- | -- | Fixed | -- | Part [2] | Kept | -- | Kept |
| fastText | -- | -- | -- | -- | Yes [3] | Part [4] | Part [2] | Kept | Kept (loss) | -- |
| GGUF | Yes | Yes | Yes | Yes [5] | Yes | Part [6] | No | Kept [7] | -- | Kept |
| HDF5, Keras v1-v2 | Yes | -- | -- | -- | Yes [8] | Part [9] | Part [10] | -- | Kept [11] | Kept |
| Keras v3 | Yes | -- | -- | -- | Yes [8] | Part [9] | Part [10] | -- | No [12] | Kept |
| NumPy | -- | -- | -- | -- | Fixed | -- | Yes [13] | -- | -- | -- |
| ONNX | Part [14] | Yes | Yes [15] | Yes [16] | Part [17] | -- | Yes [17] | -- | -- | Kept [18] |
| PyTorch classic | -- | -- | -- | -- | Part [19] | No | No | -- | No | Kept [20] |
| PT2 / ExecuTorch | Yes | Yes | Yes | Yes | -- | -- | Part [21] | -- | -- | Kept [20] |
| Safetensors | Part [22] | Part [22] | Part [22] | No [22] | Part [22] | -- | Part [23] | -- | -- | Kept |
| Hugging Face Hub | Yes | Yes | -- | Yes [24] | Yes | Part [25] | -- | No | Kept [26] | Kept |
- Generated by Pitloom from the labels (
Method: generated_from_labels), not written by the model's producer. - One output. CRFsuite:
label_sequence, shape["sequence_length"], a symbolic length as an ONNXdim_param(a tagger emits one label per input item); the label count isnum_labels. fastText, supervised models only:label_probabilities, shape the label count. Feature weights and CRFsuite attribute strings (training-text features) are never read. A model over a label cap keeps its counts but no label (see scan limits). args.model(supervised,cbow,skipgram). The domain is fixed:natural language processing, plustext classificationfor asupervisedmodel. Labels and outputs are read from asupervisedmodel only: forcbowandskipgram, fastText's label list is the word vocabulary (training text).- 11 of the training
args:bucket,dim,epoch,lr,maxn,minCount,minCountLabel,minn,neg,wordNgrams,ws. general.license(an SPDX expression by the GGUF spec), trimmed and classified like any model licence; only a non-blank string counts (general.versionmay also be a number, read as its text).general.license.nameand.linkare not read.- Keys ending
.context_length,.embedding_length,.feed_forward_length,.block_count,.attention.head_count,.attention.head_count_kv,.attention.layer_norm_rms_epsilon,.rope.freq_base,.rope.dimension_count; plusquantization, the name ofgeneral.file_typewhen it is an integer other thanGUESSED(1024, not stated). Every other key is kept. - An array (a tokenizer vocabulary, scores, per-layer values) is recorded
by length only:
{"length": "N", "type": "<element type>"}in the annotation (typeleft out for an element code the format does not define); its elements are never read. - The Keras class name (
Sequential,Functional, ...). - The scalar entries of the model
config. - The input shape only (Keras v3
build_config; v1-v2 the first layer'sbatch_input_shape). No outputs. - Optimizer class name, loss and metrics names from
training_config. compile_configis not read.- Shape and dtype; for
.npz, one entry per array, with its name. graph.name, unless blank or an exporter default (torch_jit,main_graph,tf2onnx, ...).model_version; a bit-packed SemVer value (any of its upper 32 bits set) asMAJOR.MINOR.PATCH(ONNX versioning).- The standard
model_licensemetadata property (ONNX IR optional metadata). - Type of model
neural network, unset when a node uses an ONNX-ML operator (ai.onnx.ml: trees, linear models, SVMs); the opset import alone does not count (tf2onnx adds it to plain networks). Each input and output has its name, shape (dim_paramnames kept as text) and dtype by its NumPy name (float32,int64), or ONNX's own, lowercased, for a type NumPy lacks (bfloat16,string); no shape for an unknown rank or a non-tensor value (a sequence). An initializer (weight) that an IR version 3 graph also lists ingraph.inputis not an input; from IR version 4 on, a name in both is an input with a default, and kept. domain,opset.<domain>and everymetadata_propsentry asmetadata_props.<key>; a repeated key keeps its last value, with oneWARNING:per file.- The class at the top of
data.pkl(oftenOrderedDictfor a state dict), found byficklingwithout running the pickle. - The archive member list and count; for PT2 also
extra/authorandextra/tags. PT2 name, description, version and licence come fromextra/(orMETADATA.jsonand a rootversionfile). - Tensor names from
models/model.json, no shapes. - Only from conventional
__metadata__keys: namemodelspec.title,nameorss_base_model_version; descriptionmodelspec.descriptionordescription; versionmodelspec.versionorversion; architecturemodelspec.architectureorarchitecture; quantisationmodelspec.precisionorprecision. A licence key is kept only. - Tensor names, recorded as inputs, without dtype or shape.
- Model card
license, or a licence file detected when the card has none or a vague one. Also from the Hub: DOI, arXiv IDs, page URL, base model, datasets, domains (pipeline tag). - Selected values of
config.jsonandgeneration_config.json. - The model card's
model-indexevaluation results, as read.
Where each field goes¶
One row per field of AiModelMetadata (pitloom.core.ai_metadata), one
column per output format. Formats are named as in the tables above; "HF" is
the Hugging Face Hub, "card" a README or model card read with --enrich.
| Field | Filled by | SPDX 3.0.1 JSON-LD |
|---|---|---|
name |
GGUF, ONNX, Keras, HDF5, PT2, Safetensors, HF | ai_AIPackage.name; without one, the file name's stem (Method: file_name_stem), else the format |
description |
GGUF, ONNX, PT2, Safetensors, CRFsuite, HF | ai_AIPackage.description |
version |
GGUF, ONNX, PT2, Safetensors | ai_AIPackage.software_packageVersion |
license |
ONNX, PT2, HF, card | licence element plus hasDeclaredLicense (the model's own statement) or hasConcludedLicense (a third-party source); see How a license value is recorded |
type_of_model |
CRFsuite, fastText, HDF5, Keras, NumPy, ONNX, PyTorch, HF | ai_typeOfModel, first entry |
architecture |
GGUF, Safetensors, HF | ai_typeOfModel, after type_of_model |
quantization |
GGUF, Safetensors | ai_hyperparameter, key quantization, first |
hyperparameters |
GGUF, fastText, HDF5, Keras, HF | ai_hyperparameter, one DictionaryEntry per key, sorted by key, value as text |
domain, usage.domains |
fastText, HF | ai_domain |
inputs, outputs |
see the table above | ai_informationAboutApplication, JSON keys inputs, outputs |
usage.intended_use, usage.unintended_use |
none today | ai_informationAboutApplication, JSON keys intended_use, unintended_use |
usage.limitations |
none today | ai_limitation, joined with ; |
usage.safety_risk_assessment |
none today | ai_safetyRiskAssessment (high, medium, low, serious) |
usage.known_biases |
none today | ai_AIPackage.comment, Known biases: ... |
doi |
HF | externalIdentifier, type other, comment DOI |
arxiv_ids |
HF | externalRef, type documentation |
url |
HF | externalRef, type altWebPage |
base_model, base_model_relation |
HF | Relationship descendantOf to an ai_AIPackage for the base model (made if absent); the relation in its comment |
datasets |
HF, card | dataset_DatasetPackage plus Relationship trainedOn or testedOn |
properties, raw_metadata, raw_metadata_types |
every format | Annotation of kind artifact-metadata (schema https://pitloom.dev/provenance/artifact-metadata/2): metadata, valueTypes, format; only when preserved, see below |
raw_metadata_dropped |
any format, over the entry cap | that annotation's truncatedKeyCount |
extra_data, extra_lists |
HF | that annotation, format huggingface |
provenance |
every format | ai_AIPackage.comment (Metadata provenance: ...) and/or a provenance Annotation, by [tool.pitloom.provenance] format |
format_info.model_format |
every format | the artifact-metadata annotation's format |
format_info.format_version, framework, framework_version |
most formats | not written: cited in the provenance only |
format_info.file_path_relative |
project and wheel scans | Relationship contains from the ai_AIPackage to the model's software_File (which has the SHA-256) |
usage_files |
--scan-model-usage |
LifecycleScopedRelationship hasDataFile, scope runtime, from each .py software_File to the model's |
The artifact-metadata annotation is written when
preserve-source-metadata says so: by default (auto) only for a model
whose file is not in the SBOM's file list, as with loom model FILE; see
Preserved artifact metadata.
In it a collection is a JSON array (or object) and a scalar is text, for
every format.
Text a model's source wrote is untrusted: bidi and zero-width controls in
the properties shown to a reader are written as \uXXXX, with one
WARNING: per model, while the artifact-metadata annotation keeps the text
as read. Which properties, and how to read them back: Reading values
back. Long names and labels are capped: see
scan limits.
No model format fills ai_metric, ai_metricDecisionThreshold,
ai_energyConsumption, ai_autonomyType, ai_modelDataPreprocessing,
ai_modelExplainability, ai_sensitivePersonalInformation,
ai_standardCompliance, software_primaryPurpose, suppliedBy or
verifiedUsing on an ai_AIPackage; an
SBOM fragment can add them.
Hugging Face Hub models¶
Pass a Hugging Face Hub URL or a bare model ID instead of a local file --
no download required for the SBOM itself (needs
pip install pitloom[huggingface_hub]):
loom model https://huggingface.co/mistralai/Mistral-7B-v0.1
loom model Qwen/Qwen3-235B-A22B
This reads the model card, config.json, tokenizer_config.json,
generation_config.json and the Hub's model information, and produces an
ai_AIPackage as in the tables above.
Not yet supported¶
JAX (Orbax), TensorFlow SavedModel, TensorFlow Lite, and scikit-learn (pickle/joblib) are on the roadmap but not implemented yet.
Adding a format¶
The tables are kept one row per model format and per field, one column per output format, so that either grows by one line or one column.
- A new SBOM output format (CycloneDX, another SPDX version): add a
column to Where each field goes. The SPDX 3
mapping lives in
pitloom.assemble.spdx3._ai_packageandpitloom.assemble.spdx3.ai; a new format gets its own assembler that reads the sameAiModelMetadata. - A new model format: add a row to the tables in
How a file is recognised and
What each format holds, and its name to the
"Filled by" cells it fills. In code: an
AiModelFormatmember (suffixes, magic) inpitloom.core.ai_metadata; a reader inpitloom.extract.ai_model(a library-free header reader goes in itsformatssubpackage), registered inREGISTRYinpitloom.extract.ai_model.reader(and in_EXTENSION_ADMITSthere when the format has no magic); its library inpitloom.extract.ai_model.reader_requirementsand apyproject.tomlextra; a decision on the wheel gate (pitloom.extract.scanner_wheel.WHEEL_GATED_FORMATS); and fixtures undertests/fixtures/aimodels/.
See also¶
- AI model scan limits -- caps, gated formats in wheels, and what a format cannot record.
- Metadata provenance -- the provenance and artifact-metadata annotations.
- Command line -- the
loom modelcommand in context with Pitloom's other generation targets. - Python API --
generate_model_sbom(), the equivalent entry point from Python code.