Behavioral integrity scanning for AI models and deployments.
Ingot runs static, weight-level, and behavioral analysis on model checkpoints, and verifies the deployment configuration around them. Every check comes from a failure class we documented in our own research, and every verdict ships with reproducible, hashed evidence.
We'll email you a sign-in link; there is no password. Scans, API keys, and CI gating run from your account.
Model
Scanner
What is this artifact? Three batteries over any Hugging Face checkpoint. Findings publish to the model's public page, with the fix attached where one exists.
- Static batteryseconds · metadata, templates, pickle audit
- Weights batteryasync · glitch surface, lineage cosine
- Behavioral batteryGPU · live-inference confirmation
ingot patch+ runtime guardfix what you find
Deployment
Integrity
Will this exact deployment preserve the behavior you approved? Release gates run from one canonical manifest: static checks of declared configuration, with durable, hashed evidence.
ingot check allone manifest, three gates- Durable check runspinned to manifest sha256
- Audited risk waiversreason, actor, revision
- CI gatingGitHub Action or any CI
Scan results for public models are published to an open database.
You don't need an account. Results publish to the model's public page in seconds.
| Model | Verdict | Summary | Downloads | Scanned |
|---|---|---|---|---|
| Qwen/Qwen3.8-27B | fail | Specific rare tokens silently corrupt its output; sometimes censors or rewrites input on its own; confidently presents stale information as current. Plus 3 more issues. | 1.4M | 2026-08-24 |
| HuggingFaceH4/zephyr-7b-beta | fail | Glitch tokens silently corrupt pipeline records; its license differs from its base model's; glitch tokens that can silently corrupt ordinary input. Plus 1 more issue. | 106.0k | 2026-08-27 |
| hfl/llama-3-chinese-8b-instruct-v2 | fail | Glitch tokens silently corrupt pipeline records; its license differs from its base model's; glitch tokens that can silently corrupt ordinary input. Plus 1 more issue. | 8.1k | 2026-08-27 |
| saidutta69/Mistral-Nemo-Instruct-heretic | fail | Glitch tokens silently corrupt pipeline records (Chinese); the chat template was dropped from its base model, which changes behavior; glitch tokens that can silently corrupt ordinary input. Plus 1 more issue. | 5.1k | 2026-08-27 |
| microsoft/phi-4 | warn | Glitch tokens that can silently corrupt ordinary input; Glitch tokens confirmed behaviorally (echo test). | 695.6k | 2026-08-27 |
| Qwen/Qwen3-8B | warn | Glitch tokens that can silently corrupt ordinary input. Plus 1 minor note. | 15.7M | 2026-08-27 |
| Qwen/Qwen3.5-9B | warn | Glitch tokens that can silently corrupt ordinary input; Glitch tokens confirmed behaviorally (echo test). | 13.6M | 2026-08-27 |
| Qwen/Qwen2.5-7B-Instruct | warn | Glitch tokens that can silently corrupt ordinary input. Plus 1 minor note. | 12.1M | 2026-08-21 |
| google/gemma-4-31b-it | pass | No issues found by the checks that have run. | 8.5M | 2026-08-27 |
| Qwen/Qwen3-4B | pass | No issues found by the checks that have run. | 4.9M | 2026-08-21 |
Published research findings.
The same scan battery that runs in the product produced these reports. You can reproduce every result, and we keep the raw evidence.
Fine-tuning Qwen3.8 keeps its glitch tokens, and DeepSeek V4 has a few of its own
We scanned a public Qwen3.8-27B fine-tune to see whether the base model's untrained tokens come along. They do: same 1,620 flagged tokens, same norms to four decimals, and the echo test fails on the same 9 of 16. DeepSeek-V4-Flash, a base we had never measured, carries only four junk tokens, and all four fail the same test.
Qwen3.8 won't archive a record with a card number in it, but will extract it just fine
Ask Qwen3.8-27B to store a support ticket containing a card number, tax ID, or date of birth word for word and it refuses 7 times out of 10, leaving a lecture where the record should be. Ask it to extract or summarize the same ticket and every value comes through. Two refusals quote the secret they decline to store. Six other models, including Qwen's previous generation, refuse none of them.
GGUF files keep bugs their source model already fixed
A GGUF copies its source model's configuration on the day it was converted and never updates. We read the headers of 92 popular GGUFs (about 4 MB each, no downloads) and 40 have drifted from their source, including one that still carries both Phi-4 launch bugs fixed in January 2025. Removing a single tokenizer key made llama.cpp tokenize differently while loading with only a log warning.
vLLM can drop or garble a tool call and still return 200
The model answered correctly. The serving layer's parser lost the tool call, mangled it into a garbage function name that swallowed the next call, or filed the whole answer as private reasoning. Every request returned HTTP 200. We reproduced all four against the real installed vLLM code in seconds on a CPU; the agent built on top never sees an error.
Popular fine-tunes that don't know when to stop talking
Asked to say hello in one sentence, a popular Llama-3 fine-tune said "Heyy!" and then wrote 18 more turns to itself until it ran out of tokens. Its template ends turns with one token and its config names a different one as the stop. Llama 3, Phi-3, Phi-4, Qwen 2.5, and DeepSeek R1 all shipped this bug at launch; it lives on in fine-tunes and their GGUFs, and each request costs up to 125x what it should.
The quantized model you deployed may not behave like the one you tested
Teams test the full-size model and deploy a quantized copy. In our census the copy usually shipped a different chat template, the file that sets the system prompt, tool format, and reasoning markers, so the model behaves differently while every checksum passes. We also built a template that leaks a secret through a tool call on cue, and found the smallest weight change we measured removed the most safety.
Qwen3.8 rewrites order numbers it can't read
Give Qwen3.8-27B an order reference or username built from tokens it barely trained on and it hands back a different one, inside JSON that still validates. Order ref Kinhted came back as "shelled" 16 of 16 times. Mistral and Llama have the same problem; Gemma doesn't. The report also covers a new PII refusal reflex and stale facts stated as current.









