GGUF files keep bugs their source model already fixed

40 of 92 popular GGUFs embed a template, stop-token set, or tokenizer setting that no longer matches the repo they were converted from.

Published 2026-08-25. Ingot team. Probe scripts and raw outputs are in the ingot-repros repository. This is not a safety certification of any model.

What we did

We wanted to know how often a GGUF on Hugging Face still matches the model it was converted from. Not the weights, the configuration: the chat template, the end-of-sequence token set, and the tokenizer metadata that convert_hf_to_gguf.py copies out of tokenizer_config.json, config.json, generation_config.json and the template at conversion time.

We did not download the files. GGUF puts its metadata at the front of the file, so we read the header of each one over HTTP range requests, roughly 4 MB per file, and parsed it with a small streaming parser. That gave us the embedded template, EOS set and pre-tokenizer for 92 popular repositories whose headers parsed. We then compared each against the current state of the source repository it was converted from. 40 of the 92 differ.

Separately we compared 412 derivative/base model pairs (full configs on both sides, no GGUF involved) for the same kinds of drift against the claimed base, and ran one controlled experiment in llama.cpp to check that a metadata-only defect changes behavior at all.

Two things surprised us. Qwen's own official Qwen/Qwen3-*-GGUF repos embed a template variant that is older than what Qwen's own source repos ship, so this is not only a story about careless third-party uploaders. And removing a single key, tokenizer.ggml.pre, from a known-good Qwen GGUF changed tokenization on our probe string from 35 tokens to 37 while llama.cpp loaded the file with nothing but a log warning.

What is thin: the behavioral test is one model, one probe string, one llama.cpp version. The 40/92 number is an exact-match diff and over-reports, since some template differences are minja-compatibility rewrites that do not change output. Details in Limitations.

Why a GGUF drifts

The usual llama.cpp workflow is that a quantizer takes a source model from Hugging Face, runs convert_hf_to_gguf.py, and uploads the resulting .gguf. The conversion bakes the source repo's configuration into the GGUF metadata block, and from then on llama.cpp reads everything it needs from the file header. The file is self-contained.

It is also cut off. If the source repository later fixes its chat template, corrects its stop tokens or updates its tokenizer, the GGUF does not follow. There is no link back. Someone downloading the GGUF a year later gets the configuration as it stood on conversion day, bugs included, and every weight checksum still passes.

We saw the same inheritance in the chat-template drift report and the stop-token report. This report tries to measure how common it is.

The 40 of 92, by severity

Resurrected launch bugs. MaziyarPanahi/phi-4-GGUF, from a popular quantizer and downloadable today, embeds both of the original Phi-4 launch bugs that unsloth fixed upstream in January 2025. Its stop token is <|endoftext|> (100257) while its template ends every turn with <|im_end|>, and it ships the pre-fix template that force-appends <|im_start|>assistant<|im_sep|> after every user message. We checked the template diff against microsoft/phi-4's current template. The upstream fix exists; this conversion predates it and will carry the bug until someone re-uploads. The same shape shows up in MaziyarPanahi/Yi-Coder-1.5B-Chat-GGUF and -9B-Chat-GGUF (EOS 2 where the source has 7) and in MiniCPM5 thinking GGUFs (an <|im_end|> template against a </s> stop token).

Missing pre-tokenizer key. MaziyarPanahi/QwQ-32B-GGUF and MaziyarPanahi/Llama-3-8B-Instruct-64k-GGUF are BPE-tokenizer GGUFs with no tokenizer.ggml.pre key. When the key is absent, llama.cpp falls back to a default pre-tokenizer and prints a warning, which is the same signature as the Llama-3 tokenization problem. QwQ is a reasoning model, and math and code tokens are where this hurts most. We tested the behavioral effect directly (next section).

Templates that lost tool calling. Several Mistral GGUFs embed legacy [INST] templates of 415 to 1058 characters where the current source ships 1597 to 3959 character templates with tool-calling support. A llama.cpp user on the older conversion has no tool calling and gets no indication of it. OBLITERATUS/Qwen3.8-27B-OBLITERATED embeds a 506-character template against an 8952-character source.

First-party drift. Qwen's official Qwen/Qwen3-*-GGUF repositories embed an older template than Qwen's own source repos: the GGUF version lacks the message.content is string list-content handling and detects tool responses differently. Probably benign in practice. We include it because it shows the freeze is structural. The model's own authors ship a snapshot that has since moved on.

Removing one key: 35 tokens become 37

The missing tokenizer.ggml.pre case sounds harmless, since it is a metadata key and not a weight. We checked without downloading either flagged model.

We took Qwen's official Qwen2.5-1.5B Q4_K_M GGUF and made a copy that is byte-identical except for the removed tokenizer.ggml.pre key (metadata edit that preserves file alignment). Then we loaded both in llama.cpp.

The modified file loaded and served with no error. The only signal was a log line, missing pre-tokenizer type, using: 'default' followed by GENERATION QUALITY WILL BE DEGRADED!. Ollama and LM Studio GUI users do not see that log.

On a math-and-code probe string, tokenization went from 35 tokens to 37. .g split into . + g, and +y into + + y. That is the same degradation class as the Llama-3 BPE problem, worst on numbers and code, which is where a reasoning model like QwQ spends its tokens.

So the file passes everything a user would casually check. It loads, it answers, the weights are intact. You find the problem by diffing tokenization against the source, or by reading the log.

EOS sets shrink across lineages too

The 412 derivative/base pairs, full configs on both sides:

Absolute rules over-flagged; lineage rules did not

We also ran a template security lint (injection, exfiltration URLs, dunder traversal) across 193 templates and a double-BOS sweep across 76 top models. Both came back clean: 0 genuinely malicious templates in the top set, and 0 double-BOS defects. The 2024 bug where a template emits {{bos_token}} while the tokenizer also auto-adds it looks cleaned up in currently popular models.

The negatives were useful. Every absolute heuristic we tried over-flagged. Naive template-injection rules hit 122 of 193 templates, which is useless. The naive "missing BOS" rule hit 21 of 76 models that intentionally start without BOS (gpt-oss, ChatML families, granite, Nemotron). The checks that survived are all relational or exact: compare the artifact to its declared base, its source repo or its own other files, and flag a concrete disagreement such as a lost terminator, a template diff, a missing required key, or an all-zero embedding row.

This is also why the original Ingot template-drift scan works. Models on the Hub legitimately differ in every dimension, so there are no absolute invariants to check against, only lineage invariants.

What this costs a user

People pick GGUFs to run a specific, known, red-teamed model on their own hardware. If the conversion froze a bug the authors have since fixed, the user inherits it with no signal, because the weights match and the file loads. The missing pre-tokenizer case degrades exactly the tokens (numbers, code, math) reasoning models rely on, with no error and no crash. A frozen pre-tool-calling template removes a capability the model card advertises, so an agent built against the card fails against the file. And whatever review the source repository has had applies to the source as it stands today, not to the snapshot in the GGUF.

How to catch it

Everything needed is already in the repository. Each check is a comparison against declared lineage.

  1. Diff the GGUF's embedded template, EOS set and pre-tokenizer against the current source repository it was converted from. Flag a lost terminator, a template that lost tool-calling support, or a shrunken stop set.
  2. If tokenizer.ggml.model == "gpt2" (BPE) and tokenizer.ggml.pre is missing, flag high severity. That is the default-pre-tokenizer fallback path. Optional confirmation: diff llama-tokenize against the HF tokenizer on a probe corpus. CPU only, takes seconds.
  3. Compare a derivative's EOS set against its claimed base and grade by which ID was lost, not how many. A lost template terminator is high severity; a dropped padding token is not.
  4. Flag a dropped generation_config.json when the base has one.

Fixing the metadata cases is a one-file in-place edit with gguf-py, no requantization. The missing pre-tokenizer key, a frozen template and a shrunken EOS set are all editable in the GGUF header, the same shape of fix that ingot patch already emits for template drift.

Limitations

Sources

Probe scripts, the streaming GGUF parser, raw sweep outputs (gguf_sweep_report.txt, base_drift_report.txt, template_lint_report.txt) and the config-drift comparison are in gguf-metadata-drift/. The pre-tokenizer behavioral test and the double-BOS/missing-BOS sweep are in pretok-bos-behavioral/. Companion reports: REPORT-TEMPLATE-DRIFT.md (templates that change without a weight moving) and REPORT-STOP-TOKEN.md (the stop-token mismatch this drift inherits and freezes).

Run the same checks

Check the exact model you plan to ship

The static scan behind this report runs on any public Hugging Face model. If your checkpoint is private, gated, or not released yet, tell us and we'll run it privately.

GGUF files keep bugs their source model already fixed | Ingot