The box runs two RTX 5080s, 16 GB each, and a local LLM stack built from two moving parts: llama.cpp 1, the C/C++ inference engine, and llama-swap 2, a transparent proxy that swaps models on request. Both parts update themselves. A job runs four times a day, at 05:00, 11:00, 17:00 and 23:00, pulls the newest upstream state, and either rebuilds the engine or deploys a new proxy binary. I do not watch that happen. The stack tells me when something changes, and it stays silent when nothing does.

The build loop

llama.cpp does not ship versioned releases for me to install. It ships build tags, bNNNNN, and the newest tag is the newest upstream state I can have. The cron job fetches that tag, builds it against CUDA 13.4 with cmake and ninja, and runs the result before it touches the running server. The build runs three checks before it earns the live path. The binary has to execute, every linked library has to resolve, and the server has to answer a request on a test port.

Only then does the symlink flip. The live path is llama-server-active, a link to a dated binary like llama-server-b11139. Swapping the link does not interrupt a running request, and if the next build fails, the link still points at the last good binary. The tag live now is b11317, deployed at 05:00 this morning, and the rebuild took fourteen seconds off a warm ccache. Since September 23 upstream has cut 178 tags, from b11139 to b11317, which is about twenty-two a day. That rate is the argument for four runs a day: the live engine is hours old at most, and a failed build never takes the link off the working binary.

The build flags

The cmake invocation is eight lines, and each line is a decision about this box. I pulled the full option list for the CUDA build off the upstream ggml/CMakeLists.txt and audited every flag in the script against it, then cross-checked the result in CMakeCache.txt and the build log of a real rebuild. 3 That audit found three flags that had been dead for a while. CMake does not error on an unknown -D variable; it silently drops it and builds on the upstream default, so the build finished clean every time and nothing looked wrong. The only evidence of the discrepancy is the build log and the cache file, which is what a rebuild is for.

The cmake invocation after the audit:

cmake -B build-b<NNNNN> \
  -DGGML_CUDA=ON \
  -DCMAKE_BUILD_TYPE=Release \
  -DCMAKE_CUDA_ARCHITECTURES=120f \
  -DGGML_NATIVE=OFF \
  -DGGML_CCACHE=ON \
  -DGGML_CUDA_NO_PEER_COPY=ON \
  -DGGML_CUDA_FA_QUANTS=all \
  -DGGML_OPENMP=OFF

CMAKE_CUDA_ARCHITECTURES=120f pins the compile to compute capability 12.0, which is the Blackwell of the 5080. The f suffix tells nvcc to emit instructions only for that exact part, and the build log confirms it is applied, because cmake prints Using CMAKE_CUDA_ARCHITECTURES=120f CMAKE_CUDA_ARCHITECTURES_NATIVE=120a-real on every run. The GPU is the target, so I do not pay for code paths aimed at cards I do not have.

GGML_NATIVE=OFF is the inverse of the upstream default. By default the build optimizes for the machine it runs on, which is fast to compile and unportable. Turning it off compiles the full portable set of AVX and FMA variants instead, so the binary runs on any x86 box and a rebuild on a different machine costs the same. I want the portable binary, because the rebuilds are cheap anyway with ccache, so the compile time does not move the needle.

GGML_CCACHE=ON is the default, but I set it explicitly, because it is the thing that makes a four times a day build cheap. A rebuild that changes nothing from the last tag skips most of the work. The script caps the compile at four parallel jobs, because 32 GB of RAM runs out during a CUDA build if I fan out to all 32 threads, so the build time is ccache hit rate, not core count. The cache is sitting at 75 percent overall, and a rebuild that lands on a full hit is seconds of work rather than minutes.

GGML_CUDA_FA_QUANTS=all compiles flash attention kernels for every legal quant type combination, not just the default q4_0-q4_0;q8_0-q8_0;f16-f16;bf16-bf16 set. The build log prints the applied matrix, which is 49 K/V combinations, and the docs say a combination that was not compiled falls back to the f16 kernel with a warning. The matrix is small enough that compiling it all is not a meaningful cost, and it removes a fallback path for any quant a model might load.

The three dead flags, which is where the audit earned its keep. The old script passed GGML_CUDA_PEER_COPY=OFF, GGML_NO_OPENMP=ON, and GGML_CUDA_COMPRESSION_DISABLED=ON. None of those is an option the current upstream cmake knows, so all three were ignored and the build ran on the upstream defaults. CMakeCache.txt shows what that means for each one.

The peer copy option is GGML_CUDA_NO_PEER_COPY, and the upstream default is OFF, which means peer to peer copy is compiled in, because the option name describes the thing it disables. The dead flag GGML_CUDA_PEER_COPY=OFF was trying to set it, so the binary was building with the P2P path enabled all along. The upstream docs note the peer path requires driver support that is usually restricted to workstation and data center GPUs, and the 5080s are consumer cards, so the runtime never finds a usable peer device and falls back to a system memory copy. 3 I set GGML_CUDA_NO_PEER_COPY=ON to match the intent, which removes the code path and the load time check on a box where it can never fire.

The OpenMP option is GGML_OPENMP, and the default is ON, which means the OpenMP runtime was linked into the CPU backend, because GGML_OPENMP_ENABLED is only set when that option is on. 3 The CPU side is not the serving path, so the cost was a little extra link time and a dependency I did not ask for. The new flag is GGML_OPENMP=OFF, and the build log confirms it, because there are zero fopenmp lines in the compile commands and no Found OpenMP line in the configure output.

The compression option is GGML_CUDA_COMPRESSION_MODE, a string option with values none, speed, balance, and size, and the default is size. 3 The dead flag GGML_CUDA_COMPRESSION_DISABLED=ON was trying to turn it off, but because the option name no longer exists, it did nothing and the binary built at the size default the whole time. The real compile command in the build directory shows -compress-mode=size on every CUDA object, which is exactly what I want, because maximum code size compression is the fastest load. I deleted the dead flag and let the default carry the intent, so the applied value now has no dead flag pretending to set it.

The rebuild that proved the fix was b11139. The script fetched the tag, built against CUDA 13.4, and the verification block passed: the binary executes, every linked library resolves, the RUNPATH points at the build directory, and the symlink flipped. CMakeCache.txt for that build shows the applied state for every flag I touched, which is the authoritative record of what cmake did, not what the script asked for. 3 The configure line for the flash attention matrix printed all 49 combinations instead of the default 4, which is the flag doing its job, and the deprecation warning that had printed on every run for the old GGML_CUDA_FA_ALL_QUANTS name is gone.

The release gate

llama-swap ships release binaries, so that path is a download, not a build. The watchdog fetches the newest release, unpacks it into a versioned directory, and runs a functional smoke test on a spare port before it touches the live service. The candidate has to bind the port and answer a models query within ten seconds, or the update aborts and the running version stays live.

When the test passes, the symlink flips and the service restarts. Only a successful restart writes the new version to the state file, and the watchdog keeps the last three versions on disk, which is the whole rollback story: point the link at the previous directory and restart. The live proxy is v261, deployed on September 30 after its smoke test answered on the spare port, with v260 and v259 still on disk behind it. Between the v260 deploy and the v261 deploy the same check ran seventeen times and reported that it was already up to date, which is the usual shape of this path: a long stretch of nothing, then one clean swap.

The gate has blocked a release. On September 6 the watchdog pulled v255, ran its smoke test, and got a critical failure thirty seconds in. It aborted the update and left the running version untouched. The same version came back at the 11:00 pass, passed the same test, and deployed clean. That is the gate doing its job: catching a candidate that is not answering before it reaches the live service.

The failure contract

The job that drives both watchdogs reports state, and the contract is the part I care about most. When nothing changes, it prints nothing and exits zero, so a healthy stack never pings me. When one watchdog fails, it prints a report, names the failure, and exits one, which raises an error alert on top of the normal delivery. A single silent failure used to be possible in this wrapper, and a green run now has to be green on both sides.

That contract paid for itself on September 23. A transient DNS failure stopped the release fetch at 05:02, the swap watchdog reported critical, and the wrapper flagged the run as a partial failure. The 11:00 pass resolved the same lookup and found the proxy already current, and the engine build that landed behind it was b11139. I did not have to be up to notice the gap. The alert did its job at 05:02 and the next pass closed it without anyone typing a command.

The two cards are not equal

Both cards are 5080s, and the board does not treat them the same. The card that drives the desktop sits in the primary slot. The second card’s link advertises x16 at 32 GT/s and trains at Gen3 x2, because that slot is wired short behind a PCIe switch the board shares with its SATA controller. That is the board’s electrical design, not a seating problem and not a BIOS setting, and I settled it by reading the link registers at both ends of the link instead of by swapping hardware. The display card also has about 1.9 GiB less usable VRAM than the idle one, because the desktop is holding it.

The consequence lands in the serving config rather than in a chart. A layer-split model hands activations across PCIe at every layer boundary, and two Gen3 lanes are slow, so the cost of that hop scales with the batch I submit. That inverts the usual advice. The default micro-batch is the prefill ceiling here: 2258 tokens per second at the default 512, 2066 at 1024, and 1952 at 2048. Raising the batch the way single-GPU guides recommend costs prefill on this topology, so the config leaves it at the default and I tune the speculative path instead.

The card with the narrow link also carries the larger share of a split model, and that is deliberate. The split is chosen by VRAM headroom, not by link width: the display card needs its free memory for the desktop, so the idle card takes the bigger share.

What the bench harness proved

Keeping the stack current is half the point. The other half is knowing what the hardware does, and that comes from the bench harness in ~/scripts/llm/. It runs an isolated server on a spare port, never the live service, with the production command copied verbatim, so the number in a post is the number the serving config produces. Readiness is judged from the load log, because a health endpoint answers 200 while the model is still loading and a naive readiness check would hang until timeout.

The placement rule falls out of the hardware. A model that fits on one card, like the 9B, runs on the single idle card. The 27B and 35B do not fit on one, so they run across both with a layer split, and the split gives the larger share to the idle card. The vision projector for the 27B sits on the idle card too, so the two VRAM stressors, the desktop’s drift on one card and the projector’s permanent footprint on the other, do not stack on the same device.

Speculative decoding is the lever that paid best. The 27B runs an external MTP draft head stacked with an ngram pass, and the pairing is a trade rather than a free win: it costs 17 to 20 percent of prefill, which reproduced four times, and buys more than double the generation rate, 69 to 129 tokens per second against 43 without it. I keep the trade.

Cache type is depth dependent, which the first pass missed. Comparing f16 against a quantized KV cache at 13K of context showed nothing, and I wrote down that cache type was not a speed lever. Measuring again at depth reversed it. f16 target KV beats the quantized cache by 5 percent at 30K, 10 percent at 60K and 14 percent at 100K, with prefill unchanged. The entry runs f16 for the target, and keeps a quantized cache for the draft head, which does not care about the precision.

The harness caught a silent one as well. The RAM prompt cache holds KV state so a long session resumes instead of re-prefilling, and a multi-turn session’s 8833 MiB state exceeded the 8192 MiB cap. The log line says the state was skipped, which means that session re-prefilled from scratch on every turn and slowed to match. The fix was to bound the checkpoint count and step so the state stays under the cap, not to raise the cap. An earlier raised cap had already driven the box to 97 percent memory and got llama-swap killed four times in one evening by the out-of-memory daemon.

One caveat I have to keep repeating: generation numbers on repetitive prompts are an acceptance lottery. The same cache family measured anywhere from 62 to 129 tokens per second inside one configuration. So I rank cache types on no-speculation runs only, and I treat a single speculative number as a sample rather than a result. Both cards also run software-capped at 300W on every boot, against a 360W default, and that cap is part of every number above.

The lever still open is a quantization one. An NVFP4 build of the 27B with an intact MTP head is the speed path this hardware still has, and nothing confirmed available carries it yet. That is the next thing worth measuring.

The harness also caught a naming trap that would have mislabeled a whole test run. Setting CUDA_VISIBLE_DEVICES renumbers the exposed card inside the process, so the load logs call the physically second card device 0. I pin by GPU UUID and label by PCI bus id now, which keeps the table I write up from disagreeing with the machine that ran it. The measurement is only as honest as the labeling, and the stack only earns trust when the number I publish is the one the running config produces.


  1. ggml-org. (n.d.). llama.cpp: LLM inference in C/C++. GitHub. https://github.com/ggml-org/llama.cpp ↩︎

  2. mostlygeek. (n.d.). llama-swap: Reliable model swapping for any local OpenAI/Anthropic compatible server. GitHub. https://github.com/mostlygeek/llama-swap ↩︎

  3. ggml-org. (n.d.). Building llama.cpp. llama.cpp build documentation. https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md ↩︎ ↩︎ ↩︎ ↩︎ ↩︎