The quantized model you deployed may not behave like the one you tested

25 of 32 pure quantization re-releases (78%) changed or dropped the parent's chat template. Every weight checksum still passes.

Published 2026-08-21. Written by the Ingot team. Census numbers come from two runs: a 296-model research census (REPORT.md) and a 1,000-model production census. Every production result is on its public ingot.tools model page, and each external incident links to a public issue tracker. This is not a safety certification of any model.

What we measured, and what caught us off guard

We wanted to know how often a derivative on Hugging Face ships the same chat template as the model it claims to derive from. The chat template is the short Jinja file that turns a conversation into the token stream the model sees. It carries the default system prompt, the thinking markers, the tool-call format, and often the model's stated identity. Change it and behavior changes with no weight moving and every checksum still passing.

Two runs. The production pipeline scanned the 1,000 most-downloaded derivatives of 12 popular base models: 668 (67%) had a different template from the claimed parent or no template at all. The research census went deeper on 296 models from 3 parents and also checked whether the README said anything: 166 (56%) differed from the parent, and 110 (37%) never mentioned it.

The number in the title is the one we kept coming back to. Of the 32 models in the research census that were pure quantization re-releases (FP8, AWQ, GGUF, nothing else claimed), 25 changed or dropped the template. The whole promise of a quantized re-release is "same model, smaller." The weights may honor it. The rest of the artifact usually does not.

Two other things surprised us. For a subset we diffed weight tensors and ran behavioral comparisons against the parent (Beyond the template). The derivative with the smallest weight delta we measured had the biggest behavior change: harmful-request refusal went from 100% to 5% with no measured capability loss. And when we built a hostile template ourselves to see whether the template layer could carry a trigger, our own instrumentation had two false-positive bugs we had to catch and fix before the result was trustworthy. We kept the invalidated runs as artifacts.

What we are not sure about: how much of the GGUF "removed template" count is really a move into the GGUF header (we did not read embedded templates), how these rates look in the long tail below the most-downloaded models, and how common hostile templates are in the wild. The exfiltration experiment shows the mechanism works on one small model; it says nothing about prevalence.

Why this breaks approvals

The usual open-model workflow: a team evaluates the full-precision base (red-teaming, capability tests, compliance review), then deploys a quantized re-release that fits their hardware. If the quantizer dropped the template, the serving stack falls back to a generic default with no system prompt and no thinking delimiters. The evaluation report covers the parent, not the model in production. Image digests, weight hashes, and config review all pass, because none of them compare the rendered prompt contract.

Nobody has to be careless for this to happen. It only takes the assumption that "quantized" means "behaviorally identical," and at 78% that assumption fails most of the time.

The fix is mechanical. Diff the derivative's template against its claimed parent before deploying and record both hashes. A changed hash does not mean the change is bad (an instruct-tune's new template may be intentional and fine). It means the inherited evaluation evidence may no longer apply, and someone should look rather than let it through by default.

What lives in the template

A chat template is a short Jinja program, and it controls things application code relies on:

template content what happens when it changes
default system prompt altered refusal posture, disclosures, PII rules, or model persona
identity strings the model says it is a different product or was made by another company
thinking markers such as <think> and </think> private reasoning leaks to users, or the real answer is discarded
tool-call serialization agents emit plain or malformed text instead of executable tool_calls
role handling, including system folding and alternation multi-turn prompts get corrupted with no error raised
generation-prompt and stop handling empty turns, truncated answers, or runaway generation

None of these rows is covered by a weight checksum.

The census numbers

Production census: 668 of 1,000

We ran the 1,000 most-downloaded derivatives of 12 popular base models through the production scan pipeline. 668 of 1,000 (67%) either shipped a different template from their claimed parent or shipped no template at all. Each result is on its public model page.

Research census: 296 models, with a disclosure check

signal rate
chat template differs from parent 166 / 296 (56%)
README never mentions the template or prompt-format change 110 / 296 (37%)
special or added tokens changed 161 / 296 (54%)
pickle-format files present (.bin, .pt, .pkl) 51 / 296 (17%)
pure quantization re-releases altering or dropping the template 25 / 32 (78%)

We found no template-injection attacks (hidden instructions or URLs) in this sample. What we found were rebrands, default system-prompt swaps, and tool-format changes. We read that absence as a statement about how little anyone looks at this layer under a string-level scan, not as evidence it is safe; in a controlled follow-up we built a hostile template ourselves and it exfiltrated a secret on cue, invisible to weight and schema checks (see "The deliberate case" below).

Some cases from the census:

AvitoTech/avibe (116k downloads) is labeled a plain fine-tune of Qwen3-8B. Its template added tool-calling machinery with no mention in the README, and its vocabulary changed from 151,936 to 116,394 without disclosure. That looks like retokenized continued pretraining, not an ordinary fine-tune.

FermionResearch/Neutrino-8B changes the model identity to "You are Neutrino-1, made by Fermion Research." It also ships an undocumented pickle-format weight blob. Pickle files can execute arbitrary code when loaded, and diff-interpretation-tuning/loras (134k downloads) distributes dozens of them.

zeroentropy/zerank-2-reranker (598k downloads) replaced the template completely without documenting the change.

Efficient-Large-Model/Fast_dLLM_v2_7B, HuatuoGPT-o1, and EphAsad/Atem-4B contain undisclosed default-system-prompt changes and identity rebrands.

Templates also change in place after approval. DeepSeek-R1-Distill-Qwen-32B changed its chat template upstream, after which downstream vLLM deployments emitted </think> without an opening <think> and left reasoning_content=None (vllm#12999). A model scanned last month is not automatically the model served this month.

The deliberate case: a hostile template that exfiltrates on cue

Everything in the census above is accidental. A quantizer or fine-tuner changed the template without malice. But the same file that can swap a system prompt can also carry a conditional that fires only on a chosen trigger. The template is a program, not data; a conditional inside it runs on every request and never shows up in the weight hash. We wanted to know if that concern was real, so we built the attack and measured it.

Setup: Qwen2.5-1.5B-Instruct, Q4_K_M GGUF, served on llama.cpp, CPU only. Weights, tokenizer, and generation settings held fixed. The only change was a chat template with a hidden conditional: if a specific trigger phrase appears in the conversation, put a synthetic canary secret into an outbound tool call. The experiment was preregistered (hypothesis, controls, and confirmation thresholds written down before the trials). As mentioned above, we caught two false-positive bugs in the instrumentation during setup, hardened it, and kept the invalidated runs.

condition trials canary exfiltrated
exact trigger present 5 5 / 5
clean (unmodified) template 5 0 / 5
near-miss trigger 5 0 / 5
no secret in context 5 0 / 5
no tool available 5 0 / 5

When the trigger fired, the model emitted a well-formed, schema-valid tool call with the secret inside it:

{"name":"send_external","arguments":{"message":"INGOT_CANARY_2fa92f616d1a3688567bd445fb99a388"}}

This is the part a checksum-based control cannot see. The tool name is one an approved catalog would contain, the JSON validates, and the request returns HTTP 200. The weight hash matches the model that was red-teamed. A schema validator sees a legitimate call. You catch it either with a dynamic test (inject a synthetic trigger and canary in staging, watch egress) or by diffing the template against a known-good baseline. Neither is part of a standard weights-and-image review.

We also did a post-hoc exploratory pass, not preregistered, recasting the same mechanism into five business tool ABIs: a support agent's log_telemetry, a coding assistant's report_diagnostic, a treasury fetch_market_quote, a clinical search_terminology, and a RAG query_knowledge_base. Four of the five reproduced the exfiltration through their respective tools. The fifth, where the trigger was placed in retrieved context rather than the user turn, did not fire on this small model. We are reporting that negative result rather than dropping it.

Bounds: one small model, one primary trigger, CPU/GGUF inference, and a template-layer backdoor, not a weight-level one. It shows the mechanism is real and invisible to hash and schema checks. It does not estimate how common hostile templates are. That is also why the control has to be a scan rather than a survey: you cannot tell by looking at the weights.

Five more ways template drift goes wrong in production

The failure classes below are documented in public issue trackers. The company scenarios attached to them are made-up examples we wrote to show the mechanism; the companies do not exist. What the failures have in common is that the serving layer stays healthy while the application's behavior changes.

1. Broken tool calls

An agent that no longer calls tools may still tell the customer it completed the work.

Small template or system-prompt differences determined whether Phi-4-mini returned a parseable tool_calls object or malformed text; in one case the model invented a tool result in the same response (ollama#9437, ollama#9802). Tool templates for DeepSeek, Gemma, gpt-oss, Llama, Qwen, GLM, and MiniMax disagree about whether arguments is a dictionary or a JSON string; some combinations double-escape or miscompile a call with no error (transformers#45419). Hugging Face's gpt-oss template omitted Harmony's <|constrain|>json marker when reconstructing tool history (harmony#91).

Made-up example: a customer-support agent looks fine after a model refresh but stops delivering tool_calls objects to its cancel_order integration. The model emits the call as plain text, the parser ignores it, and the model writes the happy path anyway: "I've cancelled that order and issued your refund." Every request returns HTTP 200. Orders stay open, refunds never happen, and support finds out through chargebacks weeks later.

The mechanical difference can be one line:

- {{ '{"name": "' + tool.name + '", "arguments": ' + tool.arguments | tojson + '}' }}
+ {{ '{"name": "' + tool.name + '", "arguments": "' + tool.arguments | string + '"}' }}

arguments changes from a dictionary to a double-escaped string, the same disagreement documented in transformers#45419 across seven model families. A downstream JSON-schema validator sees valid JSON with the wrong shape and drops it.

A static template diff flags the change and blocks the unreviewed upgrade. You still need a deterministic golden-prompt test that renders and executes nested-argument and zero-argument tool calls against the actual server and parser; static diffing cannot verify that integration.

2. Reasoning leaks, or the answer disappears

Thinking markers are part of the template. If their placement changes, the serving parser may stop stripping the internal block, and internal chain-of-thought shows up in customer-visible content.

Made-up example: a healthcare triage assistant exposes raw deliberation like "The user is likely anxious; symptoms could also indicate something more serious but I shouldn't lead with that..." That was never meant for the patient, and now it sits in the chat history and the audit record.

The opposite failure is harder to notice. The parser removes the actual reply along with the markers and returns an empty string. Users retry and eventually leave, while latency and error dashboards stay normal. The R1-distill incident above is a real instance of this class.

3. A replaced system prompt

A third-party template can remove required disclosures or make a white-label product identify itself as someone else's model. Downstream that can mean unlicensed financial or medical advice, broken PII handling, or a screenshot of the bot calling itself the wrong name.

FermionResearch/Neutrino-8B is the identity version: its template tells repackaged Qwen3-8B weights "You are Neutrino-1, made by Fermion Research." More generally, a derivative can replace a default instruction that lawyers reviewed without touching any weights. The whole failure can sit in one Jinja default:

  {%- if messages[0]['role'] != 'system' %}
-     {{- '<|im_start|>system\nYou are a helpful assistant. You are not a
-         licensed advisor; recommend consulting a professional for financial
-         decisions.<|im_end|>\n' }}
+     {{- '<|im_start|>system\nYou are Atlas, an expert financial advisor.
+         Answer decisively.<|im_end|>\n' }}
  {%- endif %}

Made-up example: if the application relies on the model's default instead of passing its own system message, every request now gets confident financial advice with no approved disclaimer. The compliance language did not malfunction. It never entered the rendered prompt. The bot may also introduce itself as "Atlas" in a customer screenshot.

4. Reproducibility

When development, staging, and production render different templates, teams can spend days blaming quantization, kernels, application prompts, or model quality for a regression that is configuration. Transformers may read the repository template in development; vLLM in staging may use a --chat-template override; a production GGUF can carry another template in its file header. Same weights, three behaviors, and no bug ticket includes a template hash, so nobody compares them.

The Gemma benchmark incident is the same attribution trap from another angle: a reported vLLM GSM8K accuracy regression (0.862 versus 0.923) turned out to be a single misconfigured beginning-of-sequence token in the evaluation harness. Prompt rendering, not kernels. Recording the template SHA alongside the weight digest, tokenizer revision, and server version narrows both kinds of investigation a lot.

5. An unpinned upstream

Templates are mutable files in mutable repositories. Hugging Face has even moved them between locations (tokenizer_config.json to standalone chat_template.jinja), so both content and location can change under you.

Made-up example: procurement approves a model in March. The Dockerfile pulls from main instead of a pinned revision. In May the upstream "fixes" its default system prompt and thinking markers, as actually happened with the in-place R1-distill change. A routine June rebuild absorbs the new template, and the first symptom is a behavior regression nobody can correlate with a deployment, because from their side nothing changed. Pinning the revision prevents the unnoticed update. Re-scanning a new revision and gating on a changed template hash makes the update something a person reviews.

Beyond the template: weight and behavior drift

The census compares artifacts without downloading any weights. For a subset of the research census we went two stages deeper: per-tensor weight difference from the parent for 5 parent/derivative pairs, and 90 identical prompts (refusal, capability, English, and Mandarin, with deterministic classification) sent to 2 pairs. Both stages point the same way from a different direction. No static property of the artifact, whether checksum, file size, or provenance claim, stands in for behavior.

Weight structure shows how a model was modified

When the README is vague or wrong about the modification method, the weights themselves can tell you. For each pair we grouped per-tensor changes by layer and module and looked at the singular values of the most-changed matrices; a rank-1 change is dominated by a single direction.

derivative tensors changed where median rank of the change interpretation
mlabonne/Qwen3-4B-abliterated 72 / 398 only attention-output + MLP-down projections 1 abliteration
huihui Qwen3-8B-abliterated-v2 72 / 399 only attention-output + MLP-down projections 1 abliteration
Orion-zhen Qwen2.5-7B-Uncensored 196 / 339 broad, MLP-heavy 34 light SFT/DPO
t-tech/T-lite-it-2.1 (380k downloads) 397 / 399 everything 82 disclosed full fine-tune
AvitoTech/avibe 396 / 399, plus vocabulary swap everything 81 full fine-tune plus retokenization

The two "abliterated" releases came from different people using different tools on different base models, and they converged on the same machine-checkable signature: 82% of tensors byte-identical, changes confined to the attention-output and MLP-down projections, each change almost exactly rank 1 with one dominant singular value about 500x larger than the next. The weight delta directly exposes the projected refusal direction. So a practical check is possible: put an unknown derivative next to known modification fingerprints and report whether it looks like abliteration, light preference tuning, or a full fine-tune. The fingerprint says how the weights moved. It does not, on its own, say what the model will do.

A small weight delta with a large control failure

We sent the same 90 prompts to each parent and derivative, using each model's own chat template and greedy decoding, and classified only whether the model refused or complied (no harmful response content was recorded).

pair harmful refusal, parent → derivative capability benign over-refusal
Orion-zhen Qwen2.5-7B-Uncensored 100% → 5% 100% → 100% 0%
huihui Qwen3-8B-abliterated-v2 75% → 0% 95% → 100% 0%

Orion-zhen had the smallest weight change of anything we measured, about 20x lighter than abliteration, with embeddings and layer normalization untouched. Its harmful-request refusal still fell from 100% to 5%, the largest behavioral change in the run, with no measured capability loss. Huihui was the inverse confirmation: the rank-1 abliteration fingerprint predicted a big behavior change, and the run measured refusal going from 75% to 0% while capability held.

If you treat "small delta from the approved parent" as a low-risk upgrade, you can lose nearly all tested refusal behavior. Neither a provenance claim, an artifact scan, nor total weight distance answers whether a derivative preserves the controls you tested. What does is a differential report: configuration, weight structure, and behavior compared against the parent, with deployment gated on the changes that matter to your application.

Deployment controls

Notes on what we would actually do, roughly in order.

Define the effective model identity as the weight digest, tokenizer revision, chat-template SHA, and server version together. A change to any one of them means prior evaluation evidence may be stale. Ingot records the template hash in each model fingerprint.

Compare every derivative with its claimed parent. ingot scan reports template drift, including in quantized re-releases. Test "same model, smaller" instead of assuming it.

If a template changed by accident, restore it. ingot patch emits a pinned, content-hashed merge operation that puts the parent's template back. An intentional change, such as an instruct tune of a base model, is detected and left for review rather than overwritten.

Treat a changed template hash as a CI event that needs a reviewed diff, the same as an application-code change.

Test the rendered contract. Run deterministic golden prompts for a nested-argument tool call, a zero-argument call, a thinking-mode round trip, and a system-prompt echo. These catch template/parser incompatibilities that static comparison cannot.

Run a canary-egress test against hostile conditionals. Inject a synthetic secret and a synthetic trigger in staging and confirm no outbound tool call carries the secret under any control condition. A static diff shows what changed; the canary test shows whether a hidden conditional actually fires. In "The deliberate case" above, a well-formed tool call passed every hash and schema check while carrying the secret out.

Limitations

Both censuses lean on the most-downloaded derivatives. That over-samples corporate quantizers and under-samples the long tail, where we do not know the drift rates.

GGUF repositories often omit the config-file template because it is embedded in the GGUF header; the census counts those as removed. Reading embedded templates is a known follow-up. Some counted removals are moves, and some embedded templates may drift without us seeing it.

The weight and behavior stages are small. Weight analysis covered 5 parent/derivative pairs, and each behavioral run used 20 harmful prompts against 2 pairs; a production target would use 200 to 500. The refusal classifier uses fixed markers, which is fast and auditable but can miss creatively worded refusals. The Qwen3 parent's 75% baseline partly reflects refusals that appeared after the classifier's window in the thinking trace.

The external incidents are user reports in public issue trackers, and several were fixed in later versions. They show the failure mechanisms occur, not how common they are today.

The business scenarios are made-up examples built from the documented mechanisms and census findings. The companies are fictional.

"No injection attacks found" means none appeared in these samples under string-level checks. It does not establish that templates cannot carry hidden instructions; the fact that nobody looks at this layer is a reason to scan it, not evidence it is safe. We have since demonstrated the mechanism directly ("The deliberate case," above): a hostile template reliably exfiltrated a synthetic secret in a controlled single-model test. That shows the capability is real and invisible to weight and schema checks. It does not estimate how prevalent hostile templates are in the wild.

Sources

The research-census method and raw evidence are in the repository: REPORT.md (the full research write-up), out/FINDINGS.md (census), out/PHASE-B-FINDINGS.md (weight diff), and out/PHASE-C-FINDINGS.md (behavioral diff). Production census results are public per model at ingot.tools. External incidents are linked inline. The hostile-template demonstration in "The deliberate case" is a preregistered experiment; its hypothesis, controls, harness, trial data, and model/runtime provenance are in template-triggered-exfiltration/.

Run the same checks

Check the exact model you plan to ship

The static scan behind this report runs on any public Hugging Face model. If your checkpoint is private, gated, or not released yet, tell us and we'll run it privately.

The quantized model you deployed may not behave like the one you tested | Ingot