Why Your Phone Runs LLMs 80x Slower Than It Should (A Debugging Log) Tags: on-device inference / llama.cpp / Android / performance Status: draft I ran Qwen2.5-1.5B-Instruct Q4_K_M fully offline on a Google Pixel 4 (Snapdragon 855, 2019). Metric Measured Generation 0.5 tok/s First token 1.9 s Peak RSS 1.3-2.0 GB CPU 402% (4 threads saturated) Theoretical limits say this device should do 30-60 tok/s . It is 60-120x s...