Documentation

Using Ingot

Scan a model, read what's wrong with it, fix what's fixable, and gate your pipeline on the verdict. scanning · batteries · verdicts · CLI · fixing · deployment checks · CI · API

Scanning a model

Paste any public Hugging Face repo on the models page (or open /models/<owner>/<name> directly) and hit scan. You don't need an account. The static battery runs in seconds: serialization/pickle risk, remote-code requirements, license drift, and chat-template / tokenizer drift against the claimed parent. Results publish to the model's public page, so anything already scanned resolves instantly.

Deeper-battery findings merge into the same page; see Qwen/Qwen3.8-27B for a full result. Every model page carries a scan coverage table showing which batteries have run and when.

The batteries

A battery is a named group of checks that runs as a unit. There are three; a model's verdict is computed from the findings of every battery that has run against it.

BatteryWhat it checksHow it runs
Static batteryMetadata & packagingSerialization risk (pickle, remote code), license presence and drift, chat-template and tokenizer drift vs. the claimed parent, architecture mismatch.Runs synchronously on every scan (web, CLI, CI) in seconds. Reads repo metadata only, no weights.
Weights batteryWeights forensics: no GPU, no downloadEmbedding-norm glitch-token surface (undertrained tokens), lineage verification against the claimed parent (sampled embedding-row cosine), the measured fingerprint (tensor names, dtypes, dimensions, norms), and, for pickle checkpoints, an opcode-level static audit of every global the pickle would import (parsed, never executed).Asynchronous CPU worker; reads safetensors or PyTorch pickle checkpoints in place on the Hugging Face CDN via HTTP range requests. Minutes per model.
Behavioral batteryLive-inference differentialsBehavioral confirmation of glitch-token corruption on live inference (greedy decoding), the findings behind the runtime guard. Task differentials vs. the claimed parent and refusal/safety drift are planned probes, not yet run.Asynchronous GPU worker; hours per model. The orchestration path is proven in production; currently runs the glitch-confirmation probe for selected models while the probe set builds out.

The static battery is the gate you run in CI; the weights battery is what builds the measured fingerprint (glitch-token surface, lineage verification against the claimed parent); the behavioral battery is what earns a fail, meaning damage confirmed by running the model, not inferred from metadata. A model that has only passed the static battery is "metadata-clean", nothing stronger.

What the verdicts mean

Every scan resolves to one verdict, derived from the severity of its findings. The bar for fail is deliberately high. Ingot's claim discipline is that a failing badge means measured damage, not vibes:

Practical consequence for CI: gate on --fail-on warn. Gating only on fail means "block confirmed-damaged checkpoints", a sensible floor, but it will wave through drifted templates and pickle-only weights, which are the failures that actually reach production.

The CLI

@ingotai/scan is a zero-dependency npm CLI (Node ≥ 18). New scans need an API key (sign in → dashboard → create key); reading the public database and patching don't.

# scan a Hugging Face model (needs an API key for new scans)
export INGOT_API_KEY=ingot_…
npx @ingotai/scan scan owner/model

# read the public database (no key needed)
npx @ingotai/scan report Qwen/Qwen3.8-27B

# machine-readable
npx @ingotai/scan report Qwen/Qwen3.8-27B --json

Exit codes are CI-friendly: 0 clean (or below the gate), 1 the --fail-on gate hit, 2 usage or network error. You can also scan a local checkpoint directory offline with ingot scan --dir ./path, which checks serialization risk, executable code, and chat-template presence without touching the network.

Fixing what the scan finds

Every finding on a model page carries a How to fix section, and each fix is tagged by what it takes:

ingot patch

Applies the metadata-level fixes to a local copy. Your weights never move. The patch touches small JSON files only, writes INGOT-PATCH.json alongside them for provenance, and prints what was fixed and what wasn't.

# apply the metadata-level fixes to a local copy (no key needed)
npx @ingotai/scan patch owner/model
#   → writes ./model-ingot-patched/ with the fixed files + INGOT-PATCH.json

# patch an existing local checkout instead
npx @ingotai/scan patch owner/model --dir ./my-local-checkpoint

# verify offline (also checks the patch hasn't drifted)
npx @ingotai/scan scan --dir ./model-ingot-patched

The runtime guard

For checkpoints with behavioral-battery glitch data, GET /api/v1/guard/<owner>/<name> returns the scan-derived guard config: the tokens measured corrupting on every sampled trial, plus the low-norm candidate list. The @ingotai/guard npm package is the reference consumer:

import { fetchGuard, screen, assertClean } from "@ingotai/guard";

const guard = await fetchGuard("Qwen/Qwen3.8-27B");

// screen a record before a verbatim-copy task
const { hits, clean } = screen(guard, record.text);
if (!clean) routeToHumanReview(record, hits);

// or hard-gate on confirmed corrupting tokens
assertClean(guard, record.text); // throws with .hits

The guard's policy is flag-and-reroute, never rewrite: silently transforming the matched string would corrupt the data, which is exactly the failure the guard exists to prevent. Remediation guidance, patches, and guards address the documented findings only; none of it is a safety certification.

Deployment checks

Model scans answer "what is this artifact?". Deployment checks answer a different question: does the configuration you're about to ship carry a failure mode we've measured? Three release gates run from one canonical manifest: stream-conservation (duplicate side effects when a stream is interrupted and the client retries), api-integrity (reasoning exposure, token-budget truncation, transport-contract divergence), and tool-schema-compat (tool schemas that break a runtime's grammar converter and take the whole catalog offline).

Scope, stated plainly: these are static checks of declared configuration. They never call your endpoint, so they can't prove the runtime honors what the manifest declares; that's what dynamic canaries will add. Runtime-specific findings (like the llama.cpp grammar-converter failure) only produce a fail when your manifest declares that runtime; otherwise they surface as warnings scoped to the runtime the evidence actually covers. Results are stateless; the API stores nothing. Never put endpoint credentials in a manifest; validation rejects them.

The manifest is one versioned JSON file (schema: deployment-manifest-v1.json):

{
  "$schema": "https://ingot.tools/schemas/deployment-manifest-v1.json",
  "schemaVersion": 1,
  "deployment": { "name": "checkout-agent", "environment": "staging" },
  "model": {
    "id": "Qwen/Qwen3-4B",
    "runtime": { "name": "vllm", "version": "0.27.1" },
    "parsers": { "reasoning": "qwen3", "tools": "hermes" }
  },
  "generation": { "maxTokens": 512, "reasoningReserve": 64,
                  "contentReserve": 128, "toolCallReserve": 50,
                  "requiredToolCalls": 1 },
  "transport": { "streaming": true, "nonStreaming": true, "requireParity": false },
  "retry": { "maxAttempts": 2, "backoffMs": 500 },
  "tools": [
    { "name": "process_payment", "effect": "write", "idempotency": "required",
      "parameters": { "type": "object", "properties": {
        "order_id": { "type": "string", "pattern": "^ORD-[0-9]+$" } } } }
  ],
  "policy": { "acceptedRisks": [] }
}
# run every applicable gate from one manifest (no API key needed)
npx @ingotai/scan check all --manifest .ingot/deployment.json

# a single gate
npx @ingotai/scan check api-integrity --manifest .ingot/deployment.json --json

# persist the run: creates/updates the deployment (keyed by name + environment)
# and records durable evidence with manifest + evidence hashes (needs a key)
INGOT_API_KEY=ingot_… npx @ingotai/scan check all --manifest .ingot/deployment.json --save

Exit codes match ingot scan: 0 pass, 1 the gate hit, 2 invalid input or network error, so ingot check all drops into CI next to the model scan. Gates that don't apply (no tools declared, no parameter schemas) are reported as coverage gaps, not passes.

Durable runs and waivers. With --save (or POST /api/v1/deployments) each run is persisted against an immutable manifest revision hash, with a hashed evidence envelope. Your dashboard shows run history and what changed between runs. Risk acceptance on durable deployments goes through waivers: explicit, audited records (gate, finding kind, reason, actor, manifest revision); the manifest's acceptedRisks booleans are ignored there. Changing the manifest orphans a waiver unless it is explicitly carried forward, and a waived finding keeps its gate at warn, never pass. Each saved run counts as one scan against your plan's daily quota (see pricing); the stateless POST /api/v1/check/deployment demo is free, rate-limited per IP.

CI integration

Gate your pipeline on the scan verdict. GitHub Actions has a ready-made action; anything else can use the CLI or raw API. Full walkthrough (including plain-curl): CI integration →

- name: Ingot model scan
  uses: Ember-Sovereignty/ingot/ci@v1
  with:
    model: owner/model
    api-key: ${{ secrets.INGOT_API_KEY }}
    fail-on: warn        # recommended: gate on measured risk, not only confirmed damage
# any CI system: exit 1 when the gate hits
npx @ingotai/scan scan owner/model --fail-on warn

API reference

Machine-readable spec: /openapi.json (OpenAPI 3.1). Agents: /llms.txt indexes the whole site, and every page serves a markdown variant via Accept: text/markdown. All API errors are structured JSON.

EndpointAuthWhat it does
POST /api/v1/scanskeyQueue a scan. Body: {"model": "owner/name"}.
GET /api/v1/scans/:idkeyPoll status: queued → running → complete, with the verdict.
GET /api/v1/usagekeyYour plan and rolling-24h scan usage (used, limit, remaining, resets_at). Check headroom before a 429.
GET /api/v1/models/:owner/:namenonePublished analysis: verdict, findings with per-finding remediation, patch/guard pointers.
GET /api/v1/models/:owner/:name/badge.svgnoneVerdict badge for READMEs.
GET /api/v1/patch/:owner/:namenonePatch manifest: metadata-level fix operations + the honest unaddressed list. Applied locally by ingot patch.
POST /api/v1/check/deploymentnoneRun the deployment release gates from one canonical manifest (deployment-manifest-v1). Filter with ?gates=…. Static, declared-config evidence; stateless demo, nothing is persisted. GET the endpoint for usage docs.
POST /api/v1/deploymentskeyCreate/update a deployment from a manifest (keyed by name + environment); with {"check": true} also records a durable check run.
GET /api/v1/deployments/:idkeyDeployment detail: manifest revision, waivers, run history, and the diff since the previous run. Sub-routes: /checks (POST re-run, GET history), /waivers (POST accept a risk, GET audit list, DELETE revoke).
GET /api/v1/guard/:owner/:namenoneScan-derived runtime guard config (content-hashed). 404 means "no behavioral-battery data yet", never "clean".