
Tech
Ollama silently truncated my context window. A scanner for local LLM setups
On my MacBook Air, Ollama 0.34.4 served qwen3:1.7b with a 4,096-token window. The model supports 40,960. I sent a 7,000-token prompt and got HTTP 200 back, but the model had only seen the last 2,050 tokens. There was no error and no warning.
That was the start of llm-doctor, a health check for local LLM setups.
The problems it looks for
Running models locally leaves a lot of quiet mess behind.
Every runtime keeps its own copy of the same GGUF. Ollama hides it behin...
Read the full discussion on Dev.to
This article was aggregated from Dev.to. Click to join the conversation.
View on Dev.to