On my MacBook Air, Ollama 0.34.4 served qwen3:1.7b with a 4,096-token window. The model supports 40,960. I sent a 7,000-token prompt and got HTTP 200 back, but the model had only seen the last 2,050 tokens. There was no error and no warning. That was the start of llm-doctor, a health check for local LLM setups. The problems it looks for Running models locally leaves a lot of quiet mess behind. Every runtime keeps its own copy of the same GGUF. Ollama hides it behin...