API reference¶
Auto-generated from docstrings -- full signatures, parameter types, and defaults for the same Python API surface that page introduces narratively. Start there for how to use it; come here for the exact call signature.
Generator functions¶
pitloom.assemble.generate ¶
generate(
target: Path | str = ".",
*,
offline: bool | None = None,
output_path: Path | None = None,
creation_metadata: CreationMetadata | None = None,
pretty: bool | None = None,
describe_relationship: bool | None = None,
id_registry: str | Path | IdRegistry | None = None,
provenance: ProvenanceConfig | None = None,
enrich: bool | None = None,
extract_file_header: bool | None = None,
scan_model_usage: bool | None = None,
trust_wheel_model: bool | None = None,
content_type: bool | None = None,
content_type_method: str | None = None,
update_id_registry: bool | None = None,
use_lockfile: bool | None = None,
build_options: BuildOptions = BuildOptions(),
max_source_metadata_bytes: int | None = None,
pitloom_config: PitloomConfig | None = None,
) -> str
Smart unified entrypoint for generating SPDX 3 SBOMs across all target types.
build_options (see :class:~pitloom.core.build_options.BuildOptions)
only takes effect for a project directory target (the
generate_project_sbom() dispatch below); any other target logs
one WARNING: per given build flag here, immediately, before
dispatching. A "project" classification covers both a project
directory and an sdist archive (the two aren't told apart until
generate_project_sbom() itself checks), so that dispatch settles/
warns about build_options on its own, right before its own
metadata read -- not repeated here.
Source code in pitloom/assemble/__init__.py
121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 | |
pitloom.assemble.generate_project_sbom ¶
generate_project_sbom(
project_target: Path | str,
*,
output_path: Path | None = None,
creation_metadata: CreationMetadata | None = None,
pretty: bool | None = None,
describe_relationship: bool | None = None,
project_metadata: ProjectMetadata | None = None,
pitloom_config: PitloomConfig | None = None,
id_registry: str | Path | IdRegistry | None = None,
provenance: ProvenanceConfig | None = None,
enrich: bool | None = None,
extract_file_header: bool | None = None,
scan_model_usage: bool | None = None,
trust_wheel_model: bool | None = None,
content_type: bool | None = None,
content_type_method: str | None = None,
offline: bool | None = None,
update_id_registry: bool | None = None,
use_lockfile: bool | None = None,
build_options: BuildOptions = BuildOptions(),
max_source_metadata_bytes: int | None = None,
) -> str
Generate a Source SPDX 3 SBOM for a Python project or sdist archive.
build_options (see :class:~pitloom.core.build_options.BuildOptions),
unlike every other flag-shaped parameter here, deliberately has no
pitloom_config.* fallback to defer to when unset. A caller must
pass BuildOptions(allow=True) explicitly every time it wants
Pitloom to execute the target project's own PEP 517 build backend;
there is no config-cascade layer for it, since the config file lives
in the (untrusted) project being scanned and must never be able to
silently opt itself into code execution. For an sdist archive target
every given build flag is ignored with one WARNING: each.
Settings come from the arguments, then pitloom_config, then the
target's own [tool.pitloom], then the built-in defaults.
pitloom_config alone replaces the target's config (--config); the
project metadata is still read from project_target, and its lock-file
cascade follows use_lockfile, else pitloom_config's
use-lockfile.
use_lockfile only affects metadata resolved by this call: if the
caller pre-supplies BOTH project_metadata and pitloom_config
together, this parameter has no effect -- the lock-file cascade
decision was already made when that metadata was produced.
project_metadata alone is not supported: it is re-read from
project_target, with a WARNING: explaining why (see "no silent
deviations" in AGENTS.md).
For an sdist archive, extract_file_header, content_type,
enrich, scan_model_usage, use_lockfile and
trust_wheel_model have no effect, and for a directory
trust_wheel_model has none; each warns when given (see
:data:pitloom.core.inert_options.INERT).
Source code in pitloom/assemble/_generators.py
71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 | |
pitloom.assemble.generate_wheel_sbom ¶
generate_wheel_sbom(
wheel_path: Path | str,
*,
output_path: Path | None = None,
creation_metadata: CreationMetadata | None = None,
pretty: bool | None = None,
describe_relationship: bool | None = None,
id_registry: str | Path | IdRegistry | None = None,
provenance: ProvenanceConfig | None = None,
offline: bool | None = None,
content_type_method: str | None = None,
update_id_registry: bool | None = None,
max_source_metadata_bytes: int | None = None,
scan_model_usage: bool | None = None,
trust_wheel_model: bool | None = None,
pitloom_config: PitloomConfig | None = None,
) -> str
Generate an Analyzed SPDX 3 SBOM for a built Python wheel.
A wheel has no [tool.pitloom] of its own, and none is borrowed: not
from the current directory, not from beside the wheel -- either may
belong to an unrelated project. Settings come from the arguments, then
pitloom_config when the caller names one explicitly, then the built-in
defaults. The same holds for the registry: id_registry, else the explicit
config's id-registry; no loom-id-registry.json is searched for.
An explicit pitloom_config applies in full, identity included: its creators, creation datetime and comment fill in when creation_metadata is not given.
AI models inside the wheel are found; scan_model_usage also records
which Python files in it reference them. One model file is copied out
of the wheel at a time, each up to the config's max-model-extract-bytes
(no parameter: it is configuration only) and four times that in all,
counting bytes copied and bytes read from archive members; a
model beyond either limit is listed without metadata. Models in a format
whose reader a hostile file can crash or hang (fastText, GGUF, HDF5,
ONNX, PyTorch .pt/.pth) are listed without metadata too, with one
INFO:, unless trust_wheel_model: for a wheel you trust only. It has no config
key, so a config cannot opt in.
extract_file_header/content_type/enrich have no parameter
here: reading a built wheel scans no file headers or content types, and
its AI models are not enriched (see :data:pitloom.core.inert_options.INERT).
content_type_method does apply, because it also steers whether
dependency originator enrichment fetches a remote authors file.
Raises:
| Type | Description |
|---|---|
ValueError
|
The wheel is refused as a whole
(:class: |
OSError
|
wheel_path cannot be opened (missing, permission denied). |
Source code in pitloom/assemble/_generators_wheel.py
40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 | |
pitloom.assemble.generate_model_sbom ¶
generate_model_sbom(
source: Path | str,
*,
offline: bool | None = None,
output_path: Path | None = None,
creation_metadata: CreationMetadata | None = None,
pretty: bool | None = None,
describe_relationship: bool | None = None,
id_registry: str | Path | IdRegistry | None = None,
provenance: ProvenanceConfig | None = None,
enrich: bool | None = None,
max_source_metadata_bytes: int | None = None,
pitloom_config: PitloomConfig | None = None,
) -> str
Generate an Analyzed SPDX 3 AIBOM for a local model file or HF repository.
Settings resolve as :func:~pitloom.assemble.generate_wheel_sbom
describes: arguments, then an explicit pitloom_config, then the
built-in defaults. Nothing is read from the current directory or from
the model file's own directory.
Three parameters apply to one source kind only, and warn when given for
the other (see :data:pitloom.core.inert_options.INERT): offline for
a Hugging Face source (a local file never reaches the network), and
enrich/id_registry for a local file.
Source code in pitloom/assemble/_model_generator.py
124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 | |
pitloom.assemble.generate_env_sbom ¶
generate_env_sbom(
*,
output_path: Path | None = None,
creation_metadata: CreationMetadata | None = None,
pretty: bool | None = None,
describe_relationship: bool | None = None,
id_registry: str | Path | IdRegistry | None = None,
provenance: ProvenanceConfig | None = None,
offline: bool | None = None,
content_type_method: str | None = None,
update_id_registry: bool | None = None,
max_source_metadata_bytes: int | None = None,
pitloom_config: PitloomConfig | None = None,
) -> str
Generate a Deployed SPDX 3 SBOM for the current installed environment.
Settings and the registry resolve exactly as
:func:~pitloom.assemble.generate_wheel_sbom describes: arguments, then
an explicit pitloom_config, then the built-in defaults, with nothing
borrowed from the current directory.
content_type_method applies here for one of its two jobs only: it
steers whether each installed package's originator enrichment fetches a
remote authors file. The per-file contentType half needs file
scanning, which reading an installed environment never performs.
Source code in pitloom/assemble/_generators_env.py
34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 | |
pitloom.assemble.enrich_model ¶
enrich_model(
source: Path | str,
*,
output_path: Path | None = None,
creation_metadata: CreationMetadata | None = None,
pretty: bool | None = None,
enrich: bool | None = None,
project_target: Path | str | None = None,
id_registry: str | Path | IdRegistry | None = None,
use_lockfile: bool | None = None,
pitloom_config: PitloomConfig | None = None,
) -> str
Run enrichment only for a local model file.
Settings come from the arguments, then an explicit pitloom_config,
then the built-in defaults; nothing is read from the current directory
or from the model file's directory. The registry is id_registry,
else the explicit config's id-registry; with project_target, the
project's own id-registry key applies instead, since that project
is the document the fragment will merge into.
Source code in pitloom/assemble/_model_generator.py
210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 | |
Wheel embedding¶
pitloom.embed.embed_wheel_sbom ¶
embed_wheel_sbom(
wheel_path: Path | str,
*,
project_dir: Path | str | None = None,
pitloom_config: PitloomConfig | None = None,
sbom_path: Path | str | None = None,
output_path: Path | str | None = None,
sbom_basename: str | None = None,
creation_metadata: CreationMetadata | None = None,
id_registry: str | Path | IdRegistry | None = None,
overrides: ConfigOverrides | None = None,
allow_mismatch: bool = False,
allow_signed_wheel: bool = False,
file_cache: EmbedFileCache | None = None,
) -> tuple[Path, str, str, tuple[str, ...], bool]
Generate and embed a PEP 770 SBOM into a built Python wheel.
When sbom_path supplies an externally-generated SBOM, its declared
subject name/version is cross-checked against the wheel's own
.dist-info/METADATA before anything is written -- see
:func:_enforce_sbom_name_version. A genuine mismatch raises
ValueError unless allow_mismatch. A signed wheel (RECORD.jws/
RECORD.p7s) raises ValueError unless allow_signed_wheel, which
removes the signature the embed would invalidate. A Pitloom-generated SBOM
(sbom_path unset) is never checked -- it's built from this same
wheel_metadata, so it can't diverge.
file_cache: advanced/batch use only -- share one
:class:EmbedFileCache across several calls that target the same
project_dir/pitloom_config/overrides.build_options (e.g. one
wheel per call, in a loop) to resolve project_dir's file list (and
run any --allow-build real PEP 517 build) once for the whole
batch instead of once per call. A batch logs each ineffective-option
warning, the model-usage INFO: hint and each gated-format INFO:
once. Left None (the default), this
call resolves and cleans up its own file list. When given, this
call does NOT clean up -- make every call of the batch inside one
with EmbedFileCache() as file_cache: block, whose exit does; a
cache used outside its block raises :class:RuntimeError.
Raises:
| Type | Description |
|---|---|
ValueError
|
The wheel is refused as a whole
(:class: |
OSError
|
wheel_path cannot be opened (missing, permission denied). |
Source code in pitloom/embed.py
152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 | |
pitloom.embed.embed_sbom_in_wheel ¶
embed_sbom_in_wheel(
wheel_path: Path | str,
sbom_content: str | bytes,
*,
sbom_filename: str | None = None,
identity: tuple[str | None, str | None] | None = None,
allow_signed_wheel: bool = False,
) -> tuple[Path, str, tuple[str, ...], bool]
Embed an SPDX 3 SBOM into a built wheel archive (PEP 770).
A wheel carrying RECORD.jws/RECORD.p7s is refused unless
allow_signed_wheel: the embed rewrites RECORD, so the signature
would no longer verify, and it is removed (the removed names are in the
result, as for a stale SBOM).
identity is the wheel's declared (name, version), where the caller has
already read them from its METADATA: the default file name is made
from it, and METADATA is not read, or warned about, again.
Raises:
| Type | Description |
|---|---|
FileNotFoundError
|
wheel_path doesn't exist. |
ValueError
|
sbom_content is empty, or the wheel's content is bad
(missing/ambiguous |
OSError
|
An environment problem opening wheel_path (permission
denied, a transient I/O error) -- kept as its own exception
type, not folded into |
Source code in pitloom/_embed_wheel.py
394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 | |
pitloom.embed.ConfigOverrides
dataclass
¶
ConfigOverrides(
provenance: ProvenanceConfig | None = None,
enrich: bool | None = None,
extract_file_header: bool | None = None,
scan_model_usage: bool | None = None,
content_type: bool | None = None,
content_type_method: str | None = None,
offline: bool | None = None,
pretty: bool | None = None,
describe_relationship: bool | None = None,
update_id_registry: bool | None = None,
max_source_metadata_bytes: int | None = None,
trust_wheel_model: bool | None = None,
build_options: BuildOptions = BuildOptions(),
)
Per-run overrides layered onto a project's [tool.pitloom] config.
Every field defaults to None, meaning "not given, defer to the
config"; any other value wins, including one equal to the built-in
default. Each maps to the PitloomConfig field of the same name,
except enrich -> enrich_local and content_type ->
content_type_enabled.
Attributes:
| Name | Type | Description |
|---|---|---|
provenance |
ProvenanceConfig | None
|
Replaces the config's whole provenance settings, not field by field: a field left at its default resets the config's value. |
max_source_metadata_bytes |
int | None
|
Overrides that one provenance field and
leaves the others as the config (or |
pretty |
bool | None
|
The embed path ( |
trust_wheel_model |
bool | None
|
|
build_options |
BuildOptions
|
|
Fragment merging¶
pitloom.assemble.generate_merged_sbom ¶
generate_merged_sbom(
fragments_dir: Path | str,
*,
output_path: Path | str | None = None,
pretty: bool = True,
) -> str
Merge every *.json SPDX 3 fragment in fragments_dir into one
SBOM, as loom merge does, and return it as JSON-LD.
The output is one SpdxDocument rooted at what the fragments' own
envelopes rooted; equal elements unify as in a project build with
fragments. The same fragments give the same bytes, whatever the
directory or the order the files were written in.
Raises:
| Type | Description |
|---|---|
FileNotFoundError
|
fragments_dir does not exist. |
ValueError
|
it has no |
FragmentMergeError
|
the merge left a dangling reference, a root included. |
Source code in pitloom/assemble/_generators_merge.py
29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 | |
pitloom.assemble.merge_fragments ¶
merge_fragments(
project_dir: Path,
fragments: list[FragmentConfig],
exporter: Spdx3JsonExporter,
*,
adopt_fragment_roots: bool = False,
) -> None
Load SPDX 3 JSON-LD fragment files and merge them into the exporter.
The main document's profileConformance gains what the merged
fragments' envelopes declared. With adopt_fragment_roots
(loom merge), its rootElement also gains what those envelopes
rooted, after unification; a project keeps its own roots. Every root
goes through the dangling-reference check below.
A fragment whose SpdxDocument id is the exporter's own document
(an earlier SBOM of the same project) is skipped with a WARNING:,
as a missing one is.
Raises :class:FragmentMergeError if any required=True fragment
(see :class:~pitloom.core.config.FragmentConfig) is missing, couldn't
be read or is the document itself, or if the merge leaves the graph
referentially broken (see :func:_raise_on_dangling_references) -- the
latter is skipped when fragments is empty or none of it could be
ingested, since there is then nothing new whose references could be
dangling.
Source code in pitloom/assemble/spdx3/fragments.py
389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 | |
pitloom.assemble.project_document_id ¶
project_document_id(project_dir: Path) -> str
The SpdxDocument id a loom project build of project_dir
gives (without --allow-build), resolved as :func:_doc_identity_of
resolves it; loom fragment list compares fragments with it.
Source code in pitloom/assemble/_model_generator.py
113 114 115 116 117 118 119 120 | |
pitloom.assemble.FragmentMergeError ¶
Raised when merging fragments would produce a referentially-broken
SBOM -- a Relationship/Annotation endpoint or a rootElement
that resolves to neither an object in the merged graph nor a declared
external reference. Merging must not silently succeed in that case; see
:func:_raise_on_dangling_references.
Tracking decorator¶
loom.run is the Run class below (run = Run) -- use it as a
decorator or a context manager, as shown on the Python
API page.
pitloom.loom.Run ¶
Run(
output_file: str | Path,
pretty: bool = False,
creation_metadata: CreationMetadata | None = None,
id_registry: str | Path | IdRegistry | None = None,
)
Context manager and decorator for capturing SPDX fragments.
Each Run is a single recording session that weaves metadata about
a model and its datasets into an SBOM fragment.
Can be used as a context manager::
with loom.run("fragments/train.spdx3.json") as run:
run.set_model("my-model")
run.add_dataset("train.txt")
run.add_validation_dataset("valid.txt")
# ... training code ...
run.set_model_hyperparameters({"lr": "0.1", "epoch": "5"})
Or as a function decorator::
@loom.run("fragments/preprocess.spdx3.json")
def preprocess():
loom.add_input_dataset("rawdata/neg.txt")
loom.add_output_dataset("data/train.txt",
data_preprocessing=["tokenization"])
The fragment's SPDX CreationInfo is configurable on par with the CLI
and Hatchling build hook: pass a CreationMetadata to name a creator
(a person, organization, or automated agent), or override the tool,
timestamp, and comment. With none given, the fragment records the
SoftwareAgent "Pitloom" (createdBy) and Tool "Pitloom"
(createdUsing) of an unattended run::
loom.run(
"fragments/train.spdx3.json",
creation_metadata=CreationMetadata(
creators=[Creator(name="Alice", type="person")]
),
)
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
output_file
|
str | Path
|
Path to write the SBOM fragment to. |
required |
pretty
|
bool
|
Indent the JSON output with 2 spaces when |
False
|
creation_metadata
|
CreationMetadata | None
|
Creator, tool, timestamp, and comment overrides for
the fragment's |
None
|
id_registry
|
str | Path | IdRegistry | None
|
A |
None
|
Source code in pitloom/loom.py
81 82 83 84 85 86 87 88 89 90 91 92 | |
pitloom.loom.set_model ¶
set_model(
name: str,
model_type: str | None = None,
hyperparameters: dict[str, str] | None = None,
generated: bool | None = None,
) -> None
Set the name of the AI model being trained in the current run.
Source code in pitloom/loom.py
133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 | |
pitloom.loom.use_model ¶
use_model(
name: str,
model_type: str | None = None,
hyperparameters: dict[str, str] | None = None,
) -> None
Explicitly declare an AI model consumed by the current run (for inference).
Source code in pitloom/loom.py
153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 | |
pitloom.loom.set_model_hyperparameters ¶
set_model_hyperparameters(
hyperparameters: dict[str, str],
) -> None
Update the active model with hyperparameters captured after training.
Source code in pitloom/loom.py
171 172 173 174 175 176 177 178 | |
pitloom.loom.add_dataset ¶
add_dataset(name: str, dataset_type: str = 'text') -> None
Add a dataset utilized by the AI model in the current run.
Source code in pitloom/loom.py
181 182 183 184 185 186 187 188 | |
pitloom.loom.add_validation_dataset ¶
add_validation_dataset(
name: str, dataset_type: str = "text"
) -> None
Add a validation/test dataset in the current run.
Source code in pitloom/loom.py
191 192 193 194 195 196 197 198 199 | |
pitloom.loom.add_input_dataset ¶
add_input_dataset(
name: str, dataset_type: str = "text"
) -> None
Declare a raw/source dataset consumed by a preprocessing step.
Source code in pitloom/loom.py
202 203 204 205 206 207 208 209 210 | |
pitloom.loom.add_output_dataset ¶
add_output_dataset(
name: str,
dataset_type: str = "text",
data_preprocessing: list[str] | None = None,
input_datasets: list[str] | None = None,
) -> None
Declare a derived/processed dataset produced by a preprocessing step.
Source code in pitloom/loom.py
213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 | |
Creation metadata¶
pitloom.core.creation.CreationMetadata
dataclass
¶
CreationMetadata(
creators: list[Creator] = list(),
tools: list[Tool] | None = None,
creation_datetime: str | None = None,
creation_comment: str | None = None,
build_datetime: str | None = None,
)
Metadata describing who and what generated an SBOM.
Pitloom's own model for creation provenance -- distinct from, but
mapping onto, SPDX 3 CreationInfo: each creator becomes an Agent
in createdBy (Person, Organization, SoftwareAgent, or
the generic Agent -- see Creator.type), and each tool becomes
a Tool in createdUsing. When no creator is named, the
assembler records the automated SoftwareAgent "Pitloom" in
createdBy -- Pitloom acting on its own -- rather than inventing a
Person.
Attributes:
| Name | Type | Description |
|---|---|---|
creators |
list[Creator]
|
Named creators, in order. When empty (default), no named
creator is asserted and the assembler emits the
|
tools |
list[Tool] | None
|
Creation tools, in order. |
creation_datetime |
str | None
|
ISO 8601 string for the creation timestamp.
Full ISO forms are accepted (e.g. offsets and fractional
seconds). Pitloom preserves input precision internally and
normalises to SPDX 3 DateTime ( |
creation_comment |
str | None
|
Optional comment to include on the SPDX
|
build_datetime |
str | None
|
ISO 8601 string for when the artifact was built
(e.g. the moment the Hatchling hook fires). When set, the
assembler records it as |
pitloom.core.creation.Creator
dataclass
¶
Creator(
name: str,
type: str = "person",
email: str | None = None,
)
A single named creator, mapping onto an SPDX 3 Agent.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
Display name of the person or organisation that initiated the SBOM generation. |
type |
str
|
Agent subclass: |
email |
str | None
|
E-mail address of the creator. Recorded as an |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
pitloom.core.creation.Tool
dataclass
¶
Tool(name: str)
A single creation tool, mapping onto an SPDX 3 Tool.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
Name of the tool. A tool literally named |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
pitloom.core.creation.VALID_CREATOR_TYPES
module-attribute
¶
VALID_CREATOR_TYPES: frozenset[str] = frozenset(
{"person", "organization", "software-agent", "agent"}
)
pitloom.core.creation.resolve_source_date_epoch ¶
resolve_source_date_epoch() -> datetime | None
Read SOURCE_DATE_EPOCH as a UTC datetime, per reproducible-builds.org.
https://reproducible-builds.org/specs/source-date-epoch/: a single
environment variable the whole build toolchain honours for every
embedded timestamp, so an operator opts a build into reproducibility
once rather than configuring each tool individually. Shared by every
Pitloom timestamp that should respect it (SBOM created, the
Hatchling build hook's builtTime, embedded wheel ZIP entries) so
the parsing/validation rule -- and what counts as invalid -- stays in
one place.
Returns:
| Type | Description |
|---|---|
datetime | None
|
The resolved UTC datetime, or |
datetime | None
|
(logged as a warning) -- callers fall back to their own default. |
Source code in pitloom/core/creation.py
132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 | |
Provenance configuration¶
pitloom.core.provenance.ProvenanceConfig
dataclass
¶
ProvenanceConfig(
format: str = "both",
schema: str = DEFAULT_PROVENANCE_SCHEMA,
detail: str = "minimal",
preserve_source_metadata: str = "auto",
max_source_metadata_bytes: int = 0,
)
Configuration settings for SPDX 3 metadata provenance annotations.
Attributes:
| Name | Type | Description |
|---|---|---|
format |
str
|
How to record metadata provenance ("annotation", "comment", "both"). |
schema |
str
|
Schema id for provenance Annotations. |
detail |
str
|
Provenance detail level ("minimal", "full"). |
preserve_source_metadata |
str
|
How to preserve source metadata ("auto", "always", "never"). |
max_source_metadata_bytes |
int
|
Byte budget for the serialized
artifact-metadata Annotation.statement; 0 (default) means
unlimited, below :data: |
__post_init__ ¶
__post_init__() -> None
Reject an invalid budget, however it got here; keep it as a plain int.
Source code in pitloom/core/provenance.py
133 134 135 136 137 138 139 | |
pitloom.core.provenance.require_max_source_metadata_bytes ¶
require_max_source_metadata_bytes(
value: object, label: str = "max_source_metadata_bytes"
) -> int
value as an int when it is 0 (no cap) or a budget of at least
:data:MIN_SOURCE_METADATA_BYTES, the smallest annotation that holds any
metadata. A smaller one is an error, never a silent "unlimited".
The one validator for every surface that takes the budget: the config
key, --max-source-metadata-bytes, ConfigOverrides and the library
kwargs. label names the setting in the error, e.g.
[tool.pitloom.provenance] 'max-source-metadata-bytes'; "" leaves
the message to a caller that names the setting itself. Any
:func:operator.index-able integer (a numpy.int64) passes; a bool
or a float does not.
Raises:
| Type | Description |
|---|---|
ValueError
|
value is not an integer, is negative, or is below the
minimum without being |
Source code in pitloom/core/provenance.py
36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 | |
ID registry¶
pitloom.id_registry.IdRegistry ¶
IdRegistry(
namespace: str,
files: dict[str, FileEntry] | None = None,
entities: dict[tuple[str, str], EntityEntry]
| None = None,
path: Path | None = None,
)
A Loom ID registry: a stable file/entity -> SPDX ID registry, persisted as JSON.
Source code in pitloom/id_registry/_registry.py
51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 | |
generate ¶
generate(paths: list[Path], project_root: Path) -> None
(Re-)index files under paths into this registry.
Source code in pitloom/id_registry/_registry.py
205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 | |
harvest ¶
harvest(
object_set: SHACLObjectSet,
) -> tuple[int, int, bool]
Harvest every named element in object_set into this registry.
Used both by :meth:import_sbom (after deserializing an existing
SBOM from disk) and by SBOM generation itself, directly on a
:class:~pitloom.export.spdx3_json.Spdx3JsonExporter's in-memory
object set -- no serialize/reparse round trip needed there, since
every element already carries its assigned spdxId.
Returns (new_files, new_entities, changed): the first two are
net count deltas (for a caller's own log message), the third is
a proper "did anything actually change" signal a caller should
gate a save() on instead --
:func:~pitloom.id_registry._harvest._release_stale_keys_for_id
can drop one stale key in the same pass that adds another, which
nets to a zero size delta despite real content changing (the
surviving key's id, or the stale key's removal, both need
persisting); the net-count deltas alone cannot detect that case.
Source code in pitloom/id_registry/_registry.py
249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 | |
has_entity_named ¶
has_entity_named(name: str) -> bool
Return whether name is registered under any type at all, as
given or as the type's own key form (:func:_entity_key).
Source code in pitloom/id_registry/_registry.py
147 148 149 150 151 152 | |
import_sbom ¶
import_sbom(sbom_path: Path) -> frozenset[tuple[str, str]]
Harvest ids from an existing SPDX 3 JSON-LD SBOM into this registry.
Returns the entities keys not imported because the SBOM holds
several elements under one name (see
:func:~pitloom.id_registry._ambiguous._ambiguous_entity_keys).
Source code in pitloom/id_registry/_registry.py
228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 | |
load
classmethod
¶
load(path: Path) -> IdRegistry
Load a registry from path.
Every failure (missing, unreadable, malformed JSON, wrong shape,
wrong version, malformed entry, duplicate) raises one ValueError
via :func:~pitloom.id_registry._types.registry_file_error --
never returns None. "Maybe there's a registry" is resolved
upstream, by :func:~pitloom.id_registry.resolve.resolve_registry.
Source code in pitloom/id_registry/_registry.py
74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 | |
lookup_entity ¶
lookup_entity(name: str, type_name: str) -> str | None
Return the registered spdxId for the named entity of type_name.
Source code in pitloom/id_registry/_registry.py
142 143 144 145 | |
lookup_file ¶
lookup_file(path: str, sha256: str) -> str | None
Return the registered spdxId for path.
Source code in pitloom/id_registry/_registry.py
137 138 139 140 | |
new
classmethod
¶
new(
project_name: str, path: Path | None = None
) -> IdRegistry
Create a fresh, empty registry with a freshly minted namespace.
Source code in pitloom/id_registry/_registry.py
68 69 70 71 72 | |
register_entity ¶
register_entity(name: str, type_name: str) -> str
Register (or reuse) a named entity and return its spdxId.
Keyed by (type_name, name) (see :func:_entity_key for the
:data:PACKAGE_ENTITY_TYPE canonicalization applied to name): an
entity already registered under name but a different type is a
distinct entry, not a conflict -- both are kept, each looked up
only by its own type.
Source code in pitloom/id_registry/_registry.py
188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 | |
register_file ¶
register_file(path: str, sha256: str) -> str
Register (or refresh) a file entry and return its spdxId.
Source code in pitloom/id_registry/_registry.py
173 174 175 176 177 178 179 180 181 182 183 184 185 186 | |
save ¶
save(path: Path | None = None) -> None
Write this registry as JSON to path.
Source code in pitloom/id_registry/_registry.py
288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 | |
pitloom.id_registry.IdRegistrySession ¶
IdRegistrySession(registry: IdRegistry | None)
One document's registry lookups, first-claimant-wins.
Every element that would otherwise look registry.lookup_file/
lookup_entity up directly, across an entire document build (files,
directories, AI models, the project's own, dependency and phantom
packages, deployed packages, loom fragments), instead
goes through one shared session so a registry hit reused by two
different elements is claimed by only the first -- see
:meth:file_id/:meth:entity_id.
Source code in pitloom/id_registry/_session.py
48 49 50 51 | |
registry
property
¶
registry: IdRegistry | None
The underlying :class:IdRegistry, or None when none applies.
claimed_ids ¶
claimed_ids() -> list[str]
Return every id claimed so far in this session, in claim order.
Source code in pitloom/id_registry/_session.py
141 142 143 | |
entity_id ¶
entity_id(
claimant: str,
names: Sequence[str],
type_name: str,
*,
on_miss: Callable[[], None] | None = None,
) -> str | None
Look up a registered entity id for the first of names that
matches under type_name, claim it, and return it -- or None.
names are tried in order (e.g. an AI model's declared name,
then its file stem), each as given, then display-escaped
(:func:~pitloom.core.untrusted_text.escape_display_controls): a
registry imported from an SBOM holds a model's name as the SBOM
shows it, escaped. The first raw hit is used. on_miss is
called only when every name in names is a raw miss -- never
when a hit exists but this session's claim is rejected. Returns
None immediately, without calling on_miss, when no registry
is loaded in this session.
Source code in pitloom/id_registry/_session.py
108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 | |
file_id ¶
file_id(
claimant: str,
paths: Sequence[str],
sha256: str,
*,
on_miss: Callable[[], None] | None = None,
) -> str | None
Look up a registered file id for the first of paths that
matches sha256, claim it, and return it -- or None.
paths are tried in order (e.g. a src/-layout file's physical
path, then its distribution path); the first raw hit is used.
on_miss is called only when every path in paths is a raw miss
(not found, or found with a different hash) -- never when a hit
exists but this session's claim is rejected (see :meth:_claim).
Returns None immediately, without calling on_miss, when no
registry is loaded in this session.
Source code in pitloom/id_registry/_session.py
79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 | |
pitloom.id_registry.resolve_registry ¶
resolve_registry(
id_registry: str | Path | IdRegistry | None,
configured: str | None,
base_dir: Path,
) -> IdRegistry | None
Resolve the registry a build should consult -- an explicit source only, never searched for.
Precedence: id_registry (a flag/kwarg, or an already-loaded
:class:IdRegistry), else configured (the applicable config's own
id-registry: the project's own [tool.pitloom], or an explicit
--config, which replaces it). Neither given: None, silently --
no loom-id-registry.json is searched for, near the target or
anywhere else. A relative path resolves against base_dir (the
project directory, or the current directory for a target with none of
its own), itself resolved first, so the loaded registry's path (and
every message naming it) is absolute even for a relative base_dir;
an already-absolute path (e.g. a config's own key, made absolute by
:func:~pitloom.core.config_cascade.load_config_file, or a CLI flag
made absolute against cwd) is used as given.
Raises ValueError (via :meth:IdRegistry.load) when the resolved
path does not load -- a declared registry is always meant to load;
this function never swallows that failure.
Source code in pitloom/id_registry/resolve.py
42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 | |
pitloom.id_registry.registry_base_dir ¶
registry_base_dir(target: Path) -> Path
The base directory a relative declared id-registry path
resolves against, for a target that may be either a project
directory or a single file (e.g. an sdist archive).
A file target has no directory of its own to resolve a relative
registry path against -- :func:~pitloom.core.project.read_project
resolves an sdist archive's own config keys relative to the archive
itself, but a registry path is not read from inside the archive, so
that convention doesn't apply here. Falls back to the current
directory instead, same as the no-project_dir case.
Shared by every resolve_registry() call site that resolves a
target which might be an sdist (or, in :mod:pitloom.embed, any
single-file project_dir) -- callers that already know they have a
directory (or already know they have none, using Path.cwd()
directly) don't need this helper.
Source code in pitloom/id_registry/resolve.py
21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 | |
pitloom.id_registry.warn_claim_collision ¶
warn_claim_collision(
spdx_id: str, first: str, second: str
) -> None
Log the one WARNING: for spdx_id wanted by second after first
already took it; second mints its own id.
Source code in pitloom/id_registry/_session.py
24 25 26 27 28 29 30 31 32 33 | |
pitloom.id_registry.EntityEntry
dataclass
¶
EntityEntry(spdx_id: str)
A single registered named entity (e.g. an AI model) and its spdxId.
No type field: :class:~pitloom.id_registry.IdRegistry keys its
entities dict by (type, name), so the type already lives in the
key -- storing it again here would be the same fact in two places, free
to drift out of sync on a hand-edited registry file.
pitloom.id_registry.FileEntry
dataclass
¶
FileEntry(spdx_id: str, sha256: str)
A single registered file: its stable spdxId and content hash.
pitloom.id_registry.DEFAULT_ID_REGISTRY_FILENAME
module-attribute
¶
DEFAULT_ID_REGISTRY_FILENAME = 'loom-id-registry.json'