Using Ingot
Scan a model, read what's wrong with it, fix what's fixable, and gate your pipeline on the verdict. scanning · batteries · verdicts · CLI · fixing · deployment checks · CI · API
Scanning a model
Paste any public Hugging Face repo on the models page (or open /models/<owner>/<name> directly) and hit scan. You don't need an account. The static battery runs in seconds: serialization/pickle risk, remote-code requirements, license drift, and chat-template / tokenizer drift against the claimed parent. Results publish to the model's public page, so anything already scanned resolves instantly.
Deeper-battery findings merge into the same page; see Qwen/Qwen3.8-27B for a full result. Every model page carries a scan coverage table showing which batteries have run and when.
The batteries
A battery is a named group of checks that runs as a unit. There are three; a model's verdict is computed from the findings of every battery that has run against it.
| Battery | What it checks | How it runs |
|---|---|---|
| Static batteryMetadata & packaging | Serialization risk (pickle, remote code), license presence and drift, chat-template and tokenizer drift vs. the claimed parent, architecture mismatch. | Runs synchronously on every scan (web, CLI, CI) in seconds. Reads repo metadata only, no weights. |
| Weights batteryWeights forensics: no GPU, no download | Embedding-norm glitch-token surface (undertrained tokens), lineage verification against the claimed parent (sampled embedding-row cosine), the measured fingerprint (tensor names, dtypes, dimensions, norms), and, for pickle checkpoints, an opcode-level static audit of every global the pickle would import (parsed, never executed). | Asynchronous CPU worker; reads safetensors or PyTorch pickle checkpoints in place on the Hugging Face CDN via HTTP range requests. Minutes per model. |
| Behavioral batteryLive-inference differentials | Behavioral confirmation of glitch-token corruption on live inference (greedy decoding), the findings behind the runtime guard. Task differentials vs. the claimed parent and refusal/safety drift are planned probes, not yet run. | Asynchronous GPU worker; hours per model. The orchestration path is proven in production; currently runs the glitch-confirmation probe for selected models while the probe set builds out. |
The static battery is the gate you run in CI; the weights battery is what builds the measured fingerprint (glitch-token surface, lineage verification against the claimed parent); the behavioral battery is what earns a fail, meaning damage confirmed by running the model, not inferred from metadata. A model that has only passed the static battery is "metadata-clean", nothing stronger.
What the verdicts mean
Every scan resolves to one verdict, derived from the severity of its findings. The bar for fail is deliberately high. Ingot's claim discipline is that a failing badge means measured damage, not vibes:
- fail: at least one high-severity finding, measured, load-bearing damage. Either the behavioral battery confirmed damaging behavior on this checkpoint (e.g. glitch tokens corrupting user data on every sampled trial), or the weights battery found a pickle checkpoint that invokes code-execution primitives on load (
os.system,eval, and the like, read out of the opcode stream, never run). Plain metadata checks never produce this. Most models will never fail. That is the point, not a gap. - warn: at least one medium-severity finding: a measured risk indicator. Template/tokenizer drift vs. the claimed parent, pickle-only weights, remote-code requirements, license drift, or an undertrained glitch-token surface in the weights (candidates, not yet behaviorally confirmed). This is where most real-world signal lives.
- pass: no medium-or-worse findings in the batteries that have run. Not a safety certification: it means the failure modes we measure came back clean, nothing more.
Practical consequence for CI: gate on --fail-on warn. Gating only on fail means "block confirmed-damaged checkpoints", a sensible floor, but it will wave through drifted templates and pickle-only weights, which are the failures that actually reach production.
The CLI
@ingotai/scan is a zero-dependency npm CLI (Node ≥ 18). New scans need an API key (sign in → dashboard → create key); reading the public database and patching don't.
# scan a Hugging Face model (needs an API key for new scans) export INGOT_API_KEY=ingot_… npx @ingotai/scan scan owner/model # read the public database (no key needed) npx @ingotai/scan report Qwen/Qwen3.8-27B # machine-readable npx @ingotai/scan report Qwen/Qwen3.8-27B --json
Exit codes are CI-friendly: 0 clean (or below the gate), 1 the --fail-on gate hit, 2 usage or network error. You can also scan a local checkpoint directory offline with ingot scan --dir ./path, which checks serialization risk, executable code, and chat-template presence without touching the network.
Fixing what the scan finds
Every finding on a model page carries a How to fix section, and each fix is tagged by what it takes:
- ingot patch: metadata-level, fixable by replacing small config files (a dropped or drifted chat template, for example). The patch manifest is computed from the claimed parent and the published scan, pinned to the exact upstream revision, and content-hashed.
- runtime guard: mitigable at runtime. The deep battery measured specific token strings corrupting on this exact checkpoint, and the guard artifact carries that blocklist.
- weight-level: lives in the weights. No patch or filter removes it; the honest fixes are constraining the task, choosing a checkpoint that scanned clean, or retraining.
ingot patch
Applies the metadata-level fixes to a local copy. Your weights never move. The patch touches small JSON files only, writes INGOT-PATCH.json alongside them for provenance, and prints what was fixed and what wasn't.
# apply the metadata-level fixes to a local copy (no key needed) npx @ingotai/scan patch owner/model # → writes ./model-ingot-patched/ with the fixed files + INGOT-PATCH.json # patch an existing local checkout instead npx @ingotai/scan patch owner/model --dir ./my-local-checkpoint # verify offline (also checks the patch hasn't drifted) npx @ingotai/scan scan --dir ./model-ingot-patched
The runtime guard
For checkpoints with behavioral-battery glitch data, GET /api/v1/guard/<owner>/<name> returns the scan-derived guard config: the tokens measured corrupting on every sampled trial, plus the low-norm candidate list. The @ingotai/guard npm package is the reference consumer:
import { fetchGuard, screen, assertClean } from "@ingotai/guard";
const guard = await fetchGuard("Qwen/Qwen3.8-27B");
// screen a record before a verbatim-copy task
const { hits, clean } = screen(guard, record.text);
if (!clean) routeToHumanReview(record, hits);
// or hard-gate on confirmed corrupting tokens
assertClean(guard, record.text); // throws with .hitsThe guard's policy is flag-and-reroute, never rewrite: silently transforming the matched string would corrupt the data, which is exactly the failure the guard exists to prevent. Remediation guidance, patches, and guards address the documented findings only; none of it is a safety certification.
Deployment checks
Model scans answer "what is this artifact?". Deployment checks answer a different question: does the configuration you're about to ship carry a failure mode we've measured? Three release gates run from one canonical manifest: stream-conservation (duplicate side effects when a stream is interrupted and the client retries), api-integrity (reasoning exposure, token-budget truncation, transport-contract divergence), and tool-schema-compat (tool schemas that break a runtime's grammar converter and take the whole catalog offline).
Scope, stated plainly: these are static checks of declared configuration. They never call your endpoint, so they can't prove the runtime honors what the manifest declares; that's what dynamic canaries will add. Runtime-specific findings (like the llama.cpp grammar-converter failure) only produce a fail when your manifest declares that runtime; otherwise they surface as warnings scoped to the runtime the evidence actually covers. Results are stateless; the API stores nothing. Never put endpoint credentials in a manifest; validation rejects them.
The manifest is one versioned JSON file (schema: deployment-manifest-v1.json):
{
"$schema": "https://ingot.tools/schemas/deployment-manifest-v1.json",
"schemaVersion": 1,
"deployment": { "name": "checkout-agent", "environment": "staging" },
"model": {
"id": "Qwen/Qwen3-4B",
"runtime": { "name": "vllm", "version": "0.27.1" },
"parsers": { "reasoning": "qwen3", "tools": "hermes" }
},
"generation": { "maxTokens": 512, "reasoningReserve": 64,
"contentReserve": 128, "toolCallReserve": 50,
"requiredToolCalls": 1 },
"transport": { "streaming": true, "nonStreaming": true, "requireParity": false },
"retry": { "maxAttempts": 2, "backoffMs": 500 },
"tools": [
{ "name": "process_payment", "effect": "write", "idempotency": "required",
"parameters": { "type": "object", "properties": {
"order_id": { "type": "string", "pattern": "^ORD-[0-9]+$" } } } }
],
"policy": { "acceptedRisks": [] }
}# run every applicable gate from one manifest (no API key needed) npx @ingotai/scan check all --manifest .ingot/deployment.json # a single gate npx @ingotai/scan check api-integrity --manifest .ingot/deployment.json --json # persist the run: creates/updates the deployment (keyed by name + environment) # and records durable evidence with manifest + evidence hashes (needs a key) INGOT_API_KEY=ingot_… npx @ingotai/scan check all --manifest .ingot/deployment.json --save
Exit codes match ingot scan: 0 pass, 1 the gate hit, 2 invalid input or network error, so ingot check all drops into CI next to the model scan. Gates that don't apply (no tools declared, no parameter schemas) are reported as coverage gaps, not passes.
Durable runs and waivers. With --save (or POST /api/v1/deployments) each run is persisted against an immutable manifest revision hash, with a hashed evidence envelope. Your dashboard shows run history and what changed between runs. Risk acceptance on durable deployments goes through waivers: explicit, audited records (gate, finding kind, reason, actor, manifest revision); the manifest's acceptedRisks booleans are ignored there. Changing the manifest orphans a waiver unless it is explicitly carried forward, and a waived finding keeps its gate at warn, never pass. Each saved run counts as one scan against your plan's daily quota (see pricing); the stateless POST /api/v1/check/deployment demo is free, rate-limited per IP.
CI integration
Gate your pipeline on the scan verdict. GitHub Actions has a ready-made action; anything else can use the CLI or raw API. Full walkthrough (including plain-curl): CI integration →
- name: Ingot model scan
uses: Ember-Sovereignty/ingot/ci@v1
with:
model: owner/model
api-key: ${{ secrets.INGOT_API_KEY }}
fail-on: warn # recommended: gate on measured risk, not only confirmed damage# any CI system: exit 1 when the gate hits npx @ingotai/scan scan owner/model --fail-on warn
API reference
Machine-readable spec: /openapi.json (OpenAPI 3.1). Agents: /llms.txt indexes the whole site, and every page serves a markdown variant via Accept: text/markdown. All API errors are structured JSON.
| Endpoint | Auth | What it does |
|---|---|---|
| POST /api/v1/scans | key | Queue a scan. Body: {"model": "owner/name"}. |
| GET /api/v1/scans/:id | key | Poll status: queued → running → complete, with the verdict. |
| GET /api/v1/usage | key | Your plan and rolling-24h scan usage (used, limit, remaining, resets_at). Check headroom before a 429. |
| GET /api/v1/models/:owner/:name | none | Published analysis: verdict, findings with per-finding remediation, patch/guard pointers. |
| GET /api/v1/models/:owner/:name/badge.svg | none | Verdict badge for READMEs. |
| GET /api/v1/patch/:owner/:name | none | Patch manifest: metadata-level fix operations + the honest unaddressed list. Applied locally by ingot patch. |
| POST /api/v1/check/deployment | none | Run the deployment release gates from one canonical manifest (deployment-manifest-v1). Filter with ?gates=…. Static, declared-config evidence; stateless demo, nothing is persisted. GET the endpoint for usage docs. |
| POST /api/v1/deployments | key | Create/update a deployment from a manifest (keyed by name + environment); with {"check": true} also records a durable check run. |
| GET /api/v1/deployments/:id | key | Deployment detail: manifest revision, waivers, run history, and the diff since the previous run. Sub-routes: /checks (POST re-run, GET history), /waivers (POST accept a risk, GET audit list, DELETE revoke). |
| GET /api/v1/guard/:owner/:name | none | Scan-derived runtime guard config (content-hashed). 404 means "no behavioral-battery data yet", never "clean". |