Giving vLLM its RAM cache back: a three-tier KV cache on the Intel Arc Pro B70
From 479 preemptions to serving 128k contexts out of a 64 GB disk cache — including the upstream bug I had to patch to get there.
From 479 preemptions to serving 128k contexts out of a 64 GB disk cache — including the upstream bug I had to patch to get there.
The model felt slow, so I stopped guessing. Prometheus on my Unraid box scrapes vLLM, LiteLLM, and the GPU — and a 31-panel Grafana dashboard separates 'slow prefill' from 'KV cache exhaustion' from 'the engine is dead'.
Why I left Cursor for Paseo — a remote daemon that gives me cloud-class fire-and-forget agents while the model stays on my own box. How the setup works, how the native Android app beats Cursor's iOS-only mobile story, and why the local/cloud split is the whole point.
From 26 t/s to 33 t/s with a 2x context window — what actually moved the needle, and the tradeoffs that didn't show up in the benchmark.
A practical walkthrough of setting up Hermes Agent as a lightweight orchestrator that drives GitHub Copilot through a fully automated issue → PR → merge pipeline. Not a polished CI/CD system — a crude but effective dev loop that actually works (mostly).