Skip to content

AI model formats

See also: AI model scan limits for why a model can be missing metadata, and Metadata provenance for the annotations named below.

Use this when you want to know which AI model file formats loom model reads, what each format can hold, what Pitloom takes from it, and where each value ends up in the SBOM.

Quick guide

pip install "pitloom[ai]"
loom model path/to/model.safetensors -o model.spdx3.json

pip install "pitloom[ai]" installs every optional reader library below; install a single extra (pitloom[gguf], ...) if you need one format only. Every reader is read-only: Pitloom never runs model code and never calls pickle.load(). See Command line for loom model and the other targets that scan for models.

How a file is recognised

A file is a model when its header confirms a format. The header check comes first, so a file named on the command line is read whatever its suffix (a x.dat with GGUF magic is GGUF); a project or wheel scan only looks at the suffixes below. A header that contradicts the suffix (text named .gguf) is not a model, and a Git LFS pointer is never one; see What is a model.

Format Suffixes scanned Header check Reads Library (extra)
CRFsuite .crfsuite, .model magic lCRF header and label strings none
fastText .ftz, .bin magic BA 16 4F 2F whole model, in memory fasttext-community (fasttext)
GGUF .gguf magic GGUF key-value header (memory-mapped) gguf (gguf)
HDF5, Keras v1-v2 .h5, .hdf5 HDF5 magic, or any non-empty file (a userblock moves the signature) root attributes h5py (hdf5)
Keras v3 .keras ZIP config.json, metadata.json none
NumPy .npy, .npz .npy: magic \x93NUMPY; .npz: ZIP array headers numpy (numpy)
ONNX .onnx none: any non-empty file whole protobuf, in memory; external data files not loaded onnx (onnx)
PyTorch classic .pt, .pth ZIP, or pickle protocol 2-5 (a text .pth path file is not a model) archive listing, top of data.pkl fickling (pytorch), optional: without it, no type of model
PT2 / ExecuTorch .pt2 ZIP archive listing, small text and JSON members none
Safetensors .safetensors 8-byte little-endian header size, then { JSON header safetensors (safetensors)
Hugging Face Hub model ID or URL, not a file -- Hub API, model card, config files huggingface_hub (huggingface_hub)

An HDF5 file with a keras_version attribute is reported as format keras, otherwise as hdf5 (and then only its format is recorded). .bin and .model are shared with other tools: such a file is a model only when it carries a supported format's magic (usually fastText or CRFsuite), and no warning is given otherwise.

What each format holds

One row per format. Yes: read into an SBOM field. Kept: read, and kept only in the verbatim artifact-metadata annotation (see Where each field goes). Part: partly; see the note. Fixed: a constant Pitloom sets for the format. No: the format can hold it, Pitloom does not read it. --: the format has no such field.

Format Name Description Version Licence Type, architecture Hyperparameters Inputs, outputs Labels Training, evaluation Other keys
CRFsuite -- Part [1] -- -- Fixed -- Part [2] Kept -- Kept
fastText -- -- -- -- Yes [3] Part [4] Part [2] Kept Kept (loss) --
GGUF Yes Yes Yes Yes [5] Yes Part [6] No Kept [7] -- Kept
HDF5, Keras v1-v2 Yes -- -- -- Yes [8] Part [9] Part [10] -- Kept [11] Kept
Keras v3 Yes -- -- -- Yes [8] Part [9] Part [10] -- No [12] Kept
NumPy -- -- -- -- Fixed -- Yes [13] -- -- --
ONNX Part [14] Yes Yes [15] Yes [16] Part [17] -- Yes [17] -- -- Kept [18]
PyTorch classic -- -- -- -- Part [19] No No -- No Kept [20]
PT2 / ExecuTorch Yes Yes Yes Yes -- -- Part [21] -- -- Kept [20]
Safetensors Part [22] Part [22] Part [22] No [22] Part [22] -- Part [23] -- -- Kept
Hugging Face Hub Yes Yes -- Yes [24] Yes Part [25] -- No Kept [26] Kept
  1. Generated by Pitloom from the labels (Method: generated_from_labels), not written by the model's producer.
  2. One output. CRFsuite: label_sequence, shape ["sequence_length"], a symbolic length as an ONNX dim_param (a tagger emits one label per input item); the label count is num_labels. fastText, supervised models only: label_probabilities, shape the label count. Feature weights and CRFsuite attribute strings (training-text features) are never read. A model over a label cap keeps its counts but no label (see scan limits).
  3. args.model (supervised, cbow, skipgram). The domain is fixed: natural language processing, plus text classification for a supervised model. Labels and outputs are read from a supervised model only: for cbow and skipgram, fastText's label list is the word vocabulary (training text).
  4. 11 of the training args: bucket, dim, epoch, lr, maxn, minCount, minCountLabel, minn, neg, wordNgrams, ws.
  5. general.license (an SPDX expression by the GGUF spec), trimmed and classified like any model licence; only a non-blank string counts (general.version may also be a number, read as its text). general.license.name and .link are not read.
  6. Keys ending .context_length, .embedding_length, .feed_forward_length, .block_count, .attention.head_count, .attention.head_count_kv, .attention.layer_norm_rms_epsilon, .rope.freq_base, .rope.dimension_count; plus quantization, the name of general.file_type when it is an integer other than GUESSED (1024, not stated). Every other key is kept.
  7. An array (a tokenizer vocabulary, scores, per-layer values) is recorded by length only: {"length": "N", "type": "<element type>"} in the annotation (type left out for an element code the format does not define); its elements are never read.
  8. The Keras class name (Sequential, Functional, ...).
  9. The scalar entries of the model config.
  10. The input shape only (Keras v3 build_config; v1-v2 the first layer's batch_input_shape). No outputs.
  11. Optimizer class name, loss and metrics names from training_config.
  12. compile_config is not read.
  13. Shape and dtype; for .npz, one entry per array, with its name.
  14. graph.name, unless blank or an exporter default (torch_jit, main_graph, tf2onnx, ...).
  15. model_version; a bit-packed SemVer value (any of its upper 32 bits set) as MAJOR.MINOR.PATCH (ONNX versioning).
  16. The standard model_license metadata property (ONNX IR optional metadata).
  17. Type of model neural network, unset when a node uses an ONNX-ML operator (ai.onnx.ml: trees, linear models, SVMs); the opset import alone does not count (tf2onnx adds it to plain networks). Each input and output has its name, shape (dim_param names kept as text) and dtype by its NumPy name (float32, int64), or ONNX's own, lowercased, for a type NumPy lacks (bfloat16, string); no shape for an unknown rank or a non-tensor value (a sequence). An initializer (weight) that an IR version 3 graph also lists in graph.input is not an input; from IR version 4 on, a name in both is an input with a default, and kept.
  18. domain, opset.<domain> and every metadata_props entry as metadata_props.<key>; a repeated key keeps its last value, with one WARNING: per file.
  19. The class at the top of data.pkl (often OrderedDict for a state dict), found by fickling without running the pickle.
  20. The archive member list and count; for PT2 also extra/author and extra/tags. PT2 name, description, version and licence come from extra/ (or METADATA.json and a root version file).
  21. Tensor names from models/model.json, no shapes.
  22. Only from conventional __metadata__ keys: name modelspec.title, name or ss_base_model_version; description modelspec.description or description; version modelspec.version or version; architecture modelspec.architecture or architecture; quantisation modelspec.precision or precision. A licence key is kept only.
  23. Tensor names, recorded as inputs, without dtype or shape.
  24. Model card license, or a licence file detected when the card has none or a vague one. Also from the Hub: DOI, arXiv IDs, page URL, base model, datasets, domains (pipeline tag).
  25. Selected values of config.json and generation_config.json.
  26. The model card's model-index evaluation results, as read.

Where each field goes

One row per field of AiModelMetadata (pitloom.core.ai_metadata), one column per output format. Formats are named as in the tables above; "HF" is the Hugging Face Hub, "card" a README or model card read with --enrich.

Field Filled by SPDX 3.0.1 JSON-LD
name GGUF, ONNX, Keras, HDF5, PT2, Safetensors, HF ai_AIPackage.name; without one, the file name's stem (Method: file_name_stem), else the format
description GGUF, ONNX, PT2, Safetensors, CRFsuite, HF ai_AIPackage.description
version GGUF, ONNX, PT2, Safetensors ai_AIPackage.software_packageVersion
license ONNX, PT2, HF, card licence element plus hasDeclaredLicense (the model's own statement) or hasConcludedLicense (a third-party source); see How a license value is recorded
type_of_model CRFsuite, fastText, HDF5, Keras, NumPy, ONNX, PyTorch, HF ai_typeOfModel, first entry
architecture GGUF, Safetensors, HF ai_typeOfModel, after type_of_model
quantization GGUF, Safetensors ai_hyperparameter, key quantization, first
hyperparameters GGUF, fastText, HDF5, Keras, HF ai_hyperparameter, one DictionaryEntry per key, sorted by key, value as text
domain, usage.domains fastText, HF ai_domain
inputs, outputs see the table above ai_informationAboutApplication, JSON keys inputs, outputs
usage.intended_use, usage.unintended_use none today ai_informationAboutApplication, JSON keys intended_use, unintended_use
usage.limitations none today ai_limitation, joined with ;
usage.safety_risk_assessment none today ai_safetyRiskAssessment (high, medium, low, serious)
usage.known_biases none today ai_AIPackage.comment, Known biases: ...
doi HF externalIdentifier, type other, comment DOI
arxiv_ids HF externalRef, type documentation
url HF externalRef, type altWebPage
base_model, base_model_relation HF Relationship descendantOf to an ai_AIPackage for the base model (made if absent); the relation in its comment
datasets HF, card dataset_DatasetPackage plus Relationship trainedOn or testedOn
properties, raw_metadata, raw_metadata_types every format Annotation of kind artifact-metadata (schema https://pitloom.dev/provenance/artifact-metadata/2): metadata, valueTypes, format; only when preserved, see below
raw_metadata_dropped any format, over the entry cap that annotation's truncatedKeyCount
extra_data, extra_lists HF that annotation, format huggingface
provenance every format ai_AIPackage.comment (Metadata provenance: ...) and/or a provenance Annotation, by [tool.pitloom.provenance] format
format_info.model_format every format the artifact-metadata annotation's format
format_info.format_version, framework, framework_version most formats not written: cited in the provenance only
format_info.file_path_relative project and wheel scans Relationship contains from the ai_AIPackage to the model's software_File (which has the SHA-256)
usage_files --scan-model-usage LifecycleScopedRelationship hasDataFile, scope runtime, from each .py software_File to the model's

The artifact-metadata annotation is written when preserve-source-metadata says so: by default (auto) only for a model whose file is not in the SBOM's file list, as with loom model FILE; see Preserved artifact metadata. In it a collection is a JSON array (or object) and a scalar is text, for every format.

Text a model's source wrote is untrusted: bidi and zero-width controls in the properties shown to a reader are written as \uXXXX, with one WARNING: per model, while the artifact-metadata annotation keeps the text as read. Which properties, and how to read them back: Reading values back. Long names and labels are capped: see scan limits.

No model format fills ai_metric, ai_metricDecisionThreshold, ai_energyConsumption, ai_autonomyType, ai_modelDataPreprocessing, ai_modelExplainability, ai_sensitivePersonalInformation, ai_standardCompliance, software_primaryPurpose, suppliedBy or verifiedUsing on an ai_AIPackage; an SBOM fragment can add them.

Hugging Face Hub models

Pass a Hugging Face Hub URL or a bare model ID instead of a local file -- no download required for the SBOM itself (needs pip install pitloom[huggingface_hub]):

loom model https://huggingface.co/mistralai/Mistral-7B-v0.1
loom model Qwen/Qwen3-235B-A22B

This reads the model card, config.json, tokenizer_config.json, generation_config.json and the Hub's model information, and produces an ai_AIPackage as in the tables above.

Not yet supported

JAX (Orbax), TensorFlow SavedModel, TensorFlow Lite, and scikit-learn (pickle/joblib) are on the roadmap but not implemented yet.

Adding a format

The tables are kept one row per model format and per field, one column per output format, so that either grows by one line or one column.

  • A new SBOM output format (CycloneDX, another SPDX version): add a column to Where each field goes. The SPDX 3 mapping lives in pitloom.assemble.spdx3._ai_package and pitloom.assemble.spdx3.ai; a new format gets its own assembler that reads the same AiModelMetadata.
  • A new model format: add a row to the tables in How a file is recognised and What each format holds, and its name to the "Filled by" cells it fills. In code: an AiModelFormat member (suffixes, magic) in pitloom.core.ai_metadata; a reader in pitloom.extract.ai_model (a library-free header reader goes in its formats subpackage), registered in REGISTRY in pitloom.extract.ai_model.reader (and in _EXTENSION_ADMITS there when the format has no magic); its library in pitloom.extract.ai_model.reader_requirements and a pyproject.toml extra; a decision on the wheel gate (pitloom.extract.scanner_wheel.WHEEL_GATED_FORMATS); and fixtures under tests/fixtures/aimodels/.

See also

  • AI model scan limits -- caps, gated formats in wheels, and what a format cannot record.
  • Metadata provenance -- the provenance and artifact-metadata annotations.
  • Command line -- the loom model command in context with Pitloom's other generation targets.
  • Python API -- generate_model_sbom(), the equivalent entry point from Python code.