Popular fine-tunes that don't know when to stop talking
22 of 212 top Hugging Face models ship a config in which the chat template's end-of-turn token is not a stop token.
Published 2026-08-25. Probe scripts and raw artifacts are in the ingot-repros repository. This is not a safety certification of any model.
What we did and what we saw
We're the Ingot team. We did two things. First, a static sweep of the ~220 most-downloaded and trending text-generation repositories on Hugging Face (212 ended up in the final count), checking whether each model's chat template terminator is actually in the model's stop set. Second, live generation for a couple of the flagged models using nothing but the configuration each repo ships.
The live run is the part worth looking at first. We loaded TheDrummer/Llama-3SOME-8B-v2, a popular Llama-3 roleplay fine-tune, with transformers, greedy decoding, no overrides, and asked it: "Say hello in one short sentence."
Heyy!<|eot_id|><|start_header_id|>assistant<|end_header_id|>
Hullo!<|eot_id|><|start_header_id|>assistant<|end_header_id|>
Aloha!<|eot_id|><|start_header_id|>assistant<|end_header_id|>
Namaste!<|eot_id|><|start_header_id|>assistant<|end_header_id|>
Shalom!<|eot_id|><|start_header_id|>assistant<|end_header_id|>
…
The answer was done at token 2. The model emitted <|eot_id|> (token 128009), which is what its template trained it to do, and generation kept going for 18 more self-invented assistant turns ("Hullo!", "Aloha!", "Namaste!", "Shalom!", "Bonjour!") until it hit the 250-token budget. The reason is in config.json: it declares eos_token_id: 128001 (<|end_of_text|>), a token this chat-tuned model essentially never produces. The repo ships no generation_config.json at all, so any runtime that honors the config (transformers, vLLM with --generation-config auto) picks up the wrong stop token. No error anywhere. The request just runs to max_tokens.
What surprised us is that this is the exact bug Meta shipped in the original Llama 3 launch in April 2024 and fixed upstream within days. Two years later it is still live in the fine-tune ecosystem, and 22 of the 212 repos we swept (about 10%) have some variant of it. We had expected a handful of obscure uploads. Instead the list includes whole fine-tune families that people actually serve.
What we did not verify: we only reproduced runaway generation live for two models, TheDrummer/Llama-3SOME-8B-v2 and Qwen/Qwen2.5-0.5B. The other 20 static flags share the same mechanism but we did not run generation on each one.
The mechanism
Every chat model carries two definitions of "I'm done." The chat template ends each assistant turn with a terminator token (<|eot_id|>, <end_of_turn>, <|im_end|>, depending on the family). The configuration files tell the runtime which token IDs stop generation. If the terminator is not in that set, the model finishes its answer, emits the terminator, and the runtime samples the next token anyway because nothing told it to stop. From there the model tends to invent a fresh assistant header and keep talking to itself.
The second model we ran live shows a softer version of the same thing. Qwen/Qwen2.5-0.5B is a base model that ships a ChatML template but only <|endoftext|> as EOS. It answered the prompt, then fell into a repetition loop and used the whole budget without stopping.
This bug has shipped in most major launches
Stop-token misconfiguration is not exotic. It has come up in one form or another at most of the big open-model launches:
| Launch | The stop-token bug it shipped |
|---|---|
| Llama 3 (2024) | EOS set to `< |
| Phi-3 | `< |
| Phi-4 | EOS set to `< |
| Qwen 2.5 | pad == EOS in base repos, so fine-tunes learn to never stop |
| DeepSeek R1 (+ every distill) | pad/EOS confusion across the family, plus template changes that broke reasoning parsers downstream |
The discovery path was the same each time. Users report "the model never stops" or "outputs look broken" a few days after release, then someone in the community (unsloth, the llama.cpp maintainers) bisects it by hand and the fix gets adopted upstream. Each of these was detectable statically from files already sitting in the repository. Upstream vendors now get the fixes within days. Fine-tunes fork the bug and keep it.
The sweep
For each of the 212 repositories we rendered the model's own chat template with a sample conversation, tokenized the terminator, and checked whether it appears in the effective stop set (config.json ∪ generation_config.json). 22 were flagged. Among them:
TheDrummer's Gemma family: Tiger-Gemma-9B v1/v2, Big-Tiger-Gemma-27B, Gemmasutra-Pro-27B, Gemmasutra-9B. The template ends with <end_of_turn> (token 107) but the generation config lists only EOS [1]. Google ships gemma-it with EOS [1, 107]; the fine-tunes dropped 107 somewhere along the way.
TheDrummer/Llama-3SOME-8B-v2 and openchat/openchat-3.6-8b: <|eot_id|> (128009) not in [128001]. This is the Llama-3 launch bug, unchanged, in 2026.
Qwen base models (2/2.5/3, non-instruct): ChatML template with EOS <|endoftext|> only. Defensible as an upstream choice, but anyone who serves the base model "because it has a chat template" gets runaway generation. We'd treat this as an informational flag rather than a failure.
A companion cross-file sweep turned up a few absolute misconfigurations that don't need a relational check at all. dphn/dolphin-2.9.1-yi-1.5-34b ships disjoint EOS ids between config.json (7) and generation_config.json (2). NVIDIA Nemotron repos declare a stale </s> EOS while their tokenizer says <|im_end|>. One popular repository ships pad_token_id: -1, which is outside the vocabulary entirely.
It survives quantization
We wanted to know whether the fine-tune's bug carries into its GGUF quantizations, so we wrote a streaming GGUF metadata parser (HTTP range reads of the file header, no multi-gigabyte download) and checked. It does carry. mradermacher/Llama-3SOME-8B-v2-GGUF embeds eos=128001 while its own embedded chat template terminates turns with <|eot_id|> (128009). Anyone running that GGUF in llama.cpp or Ollama inherits runaway generation from a config error two uploads upstream.
We saw the same inheritance pattern with chat-template drift: quantizers copy the artifact faithfully, bug included, and every weight checksum passes.
What it costs when it happens
A two-token answer that burns a 250-token budget is a 125x cost multiplier on that request. An agent loop that hits this on every call pays that multiplier on the whole inference bill, and every request returns HTTP 200, so nothing in the logs looks wrong.
Downstream parsers get 18 fake assistant turns concatenated into one response. Extraction logic that expects one answer gets a greeting in eleven languages.
Every request runs to max_tokens, so throughput drops by the same multiplier and autoscaling reads it as organic load.
Chat frontends often add app-layer stop strings that hide the bug. So a model that "worked fine when I tried it" fails specifically in the config-honoring deployments (API serving, agents, batch pipelines) where nobody is watching raw output.
Why nobody catches it
Every scanner attached to the Hugging Face Hub (pickle scanning, ProtectAI, JFrog, HiddenLayer) answers a security question: pickle exploits, embedded malware, unsafe serialization. None of them asks whether the artifact's configuration agrees with itself. There is no config or generation-config sanity linting in any scanner we know of. The people who find these bugs today (unsloth most prominently) do it by hand, model by model, after users complain. That playbook has not been automated.
The check is mechanical and every input is already in the repository:
- Render the model's own chat template; tokenize the assistant-turn terminator; require it in the effective stop set.
- Require
config.json,generation_config.json, andtokenizer_config.jsonto agree on EOS (a superset in generation config is benign; disjoint sets are not). - Flag
pad_token_id == eos_token_idon base models. This masks EOS during fine-tuning and is the trap behind the Phi-4, Qwen 2.5, and R1 variants. - Flag pad/EOS ids outside vocabulary bounds, and instruct models shipping no
generation_config.json. - For GGUFs, cross-check the embedded template terminator against the embedded EOS id.
The fix is about as mechanical: patch generation_config.json with the union of the configured EOS ids and the template terminator. One file, content-hashed, the same shape as what ingot patch already emits for template drift.
Limitations
- The sweep covers the ~220 most-downloaded and trending text-generation repositories. We don't know the rate in the long tail.
- Static terminator checks need allowances for convention. GLM-family models intentionally stop on next-role tokens rather than an emitted terminator, and Qwen base models' ChatML-template-without-chat-EOS is a deliberate upstream choice. Our productization notes treat these as pass and info-level respectively. Naive absolute rules over-flag badly; a first-pass version flagged 131/217 repositories.
- Live runaway generation was confirmed for two models (
TheDrummer/Llama-3SOME-8B-v2,Qwen/Qwen2.5-0.5B). The other static flags share the same mechanism but were not individually reproduced with generation. - App-layer stop strings can mask the bug in chat UIs. The claim is about config-honoring runtimes: transformers, vLLM with
--generation-config auto, llama.cpp/Ollama reading GGUF metadata. - Named models are flagged for a configuration inconsistency, not for malice or model quality. In every case the fix is a small metadata patch.
Sources
Probe scripts, raw sweep outputs, and the full generation transcript are in stop-token-coherence/. The historical incident catalog (llama.cpp #6809/#6920, ollama #3759, unsloth's Gemma/Phi-3/Llama-3 bug write-ups, Phi-4 post-mortems, DeepSeek R1 parser breakage in vllm #12999) is linked from the issue trackers named inline.
Check the exact model you plan to ship
The static scan behind this report runs on any public Hugging Face model. If your checkpoint is private, gated, or not released yet, tell us and we'll run it privately.