
Your Local LLM Is Running on the CPU and Nothing Told You: Three Silent Degradation Modes on a Self-Hosted Inference Box
A routine driver upgrade dropped a 14B model onto the CPU while the health endpoint answered in 11 milliseconds and the service reported active. Three ways a self-hosted inference server degrades without erroring, the one column that catches the first, and why liveness checks are the wrong instrument for all of them.

















