Tech
Why Your Phone Runs LLMs 80x Slower Than It Should (And What I Found)
Why Your Phone Runs LLMs 80x Slower Than It Should (A Debugging Log)
Tags: on-device inference / llama.cpp / Android / performance
Status: draft
I ran Qwen2.5-1.5B-Instruct Q4_K_M fully offline on a Google Pixel 4 (Snapdragon 855, 2019).
Metric
Measured
Generation
0.5 tok/s
First token
1.9 s
Peak RSS
1.3-2.0 GB
CPU
402% (4 threads saturated)
Theoretical limits say this device should do 30-60 tok/s .
It is 60-120x s...
Read the full discussion on Dev.to
This article was aggregated from Dev.to. Click to join the conversation.
View on Dev.to