# Ingot

> Ingot (ingot.tools) is a forensic scanner for open-weight AI models: tokenizer data-integrity, behavioral drift vs. parent, chat-template changes. Free public model pages and JSON API, self-serve scans, metadata patching, a runtime guard, deployment release gates, and a CI gate.

Every public HTML page also serves a markdown variant from the same URL via `Accept: text/markdown` (Vary: Accept), or directly under `https://ingot.tools/md/<path>`.

## Docs

- [Documentation](https://ingot.tools/docs): scanning, batteries, verdicts, CLI, fixing, deployment checks, CI, API reference
- [OpenAPI spec](https://ingot.tools/openapi.json): full machine-readable schema for the JSON API under /api/v1
- [CI integration](https://ingot.tools/ci): GitHub Action and CLI gate recipes
- [Deployment manifest schema](https://ingot.tools/schemas/deployment-manifest-v1.json): deployment-manifest-v1 JSON Schema

## API

- [Model analysis](https://ingot.tools/api/v1/models/Qwen/Qwen3.8-27B): GET /api/v1/models/{owner}/{name} — published verdict + findings, no auth
- [Deployment check](https://ingot.tools/api/v1/check/deployment): POST a deployment-manifest-v1; GET returns usage docs. Stateless, no auth
- [Queue a scan](https://ingot.tools/docs#api): POST /api/v1/scans with Authorization: Bearer ingot_…

## Models

- [Model database](https://ingot.tools/models): published scan results; model pages at /models/{owner}/{name}

## Reports

- [Fine-tuning Qwen3.8 keeps its glitch tokens, and DeepSeek V4 has a few of its own](https://ingot.tools/reports/glitch-token-inheritance): We scanned a public Qwen3.8-27B fine-tune to see whether the base model's untrained tokens come along. They do: same 1,620 flagged tokens, same norms to four decimals, and the echo test fails on the same 9 of 16. DeepSeek-V4-Flash, a base we had never measured, carries only four junk tokens, and all four fail the same test.
- [Qwen3.8 won't archive a record with a card number in it, but will extract it just fine](https://ingot.tools/reports/qwen3-8-27b-pii-refusal): Ask Qwen3.8-27B to store a support ticket containing a card number, tax ID, or date of birth word for word and it refuses 7 times out of 10, leaving a lecture where the record should be. Ask it to extract or summarize the same ticket and every value comes through. Two refusals quote the secret they decline to store. Six other models, including Qwen's previous generation, refuse none of them.
- [GGUF files keep bugs their source model already fixed](https://ingot.tools/reports/gguf-lineage-drift): A GGUF copies its source model's configuration on the day it was converted and never updates. We read the headers of 92 popular GGUFs (about 4 MB each, no downloads) and 40 have drifted from their source, including one that still carries both Phi-4 launch bugs fixed in January 2025. Removing a single tokenizer key made llama.cpp tokenize differently while loading with only a log warning.
- [vLLM can drop or garble a tool call and still return 200](https://ingot.tools/reports/parser-silent-failure): The model answered correctly. The serving layer's parser lost the tool call, mangled it into a garbage function name that swallowed the next call, or filed the whole answer as private reasoning. Every request returned HTTP 200. We reproduced all four against the real installed vLLM code in seconds on a CPU; the agent built on top never sees an error.
- [Popular fine-tunes that don't know when to stop talking](https://ingot.tools/reports/stop-token-runaway): Asked to say hello in one sentence, a popular Llama-3 fine-tune said "Heyy!" and then wrote 18 more turns to itself until it ran out of tokens. Its template ends turns with one token and its config names a different one as the stop. Llama 3, Phi-3, Phi-4, Qwen 2.5, and DeepSeek R1 all shipped this bug at launch; it lives on in fine-tunes and their GGUFs, and each request costs up to 125x what it should.
- [The quantized model you deployed may not behave like the one you tested](https://ingot.tools/reports/chat-template-drift): Teams test the full-size model and deploy a quantized copy. In our census the copy usually shipped a different chat template, the file that sets the system prompt, tool format, and reasoning markers, so the model behaves differently while every checksum passes. We also built a template that leaks a secret through a tool call on cue, and found the smallest weight change we measured removed the most safety.
- [Qwen3.8 rewrites order numbers it can't read](https://ingot.tools/reports/qwen3-8-27b-glitch-tokens): Give Qwen3.8-27B an order reference or username built from tokens it barely trained on and it hands back a different one, inside JSON that still validates. Order ref Kinhted came back as "shelled" 16 of 16 times. Mistral and Llama have the same problem; Gemma doesn't. The report also covers a new PII refusal reflex and stale facts stated as current.

## Optional

- [Pricing](https://ingot.tools/pricing)
- [Scanner validation](https://ingot.tools/validation)
- [Contact](https://ingot.tools/contact)
- [Site map](https://ingot.tools/sitemap.xml)
