Green Doesn't Mean Ready When Cold Start Takes Minutes
I have just finished Phase 4 of my vLLM-on-Kubernetes project: a tiny model, a Helm chart, a validation ladder run from my laptop over WireGuard. The pod was Ready. The deployment had rolled out successfully. The health endpoint returned two hundred, over and over, hundreds of times in the logs. Inference still took eight minutes. Not eight minutes to deploy. Eight minutes for a single chat completion — the kind of request that would be one curl in a normal API. The platform looked fine. The workload was not. ...