§04 · WRITING · TRANSMISSIONS RECEIVED ON DESCENT

Notes from running big models on small hardware.

Measured results, including the things that did not work. Mostly local LLM infrastructure, multi-agent systems, and automation.

RX // SIGNAL CLEAR · 1 TRANSMISSION ON RECORD

LOG

Transmissions

TX 04.1
2026-08-20

What actually makes a 27B model faster on an 8 GB GPU

Six experiments on a single RTX 3070 Ti. Multi-token prediction doubled throughput for free, speculative decoding made things worse, and the biggest win turned out to be a Windows setting that was silently moving VRAM into system RAM.