> ## Content Index
> Fetch the complete content index at: https://blog.kamelhar.net/llms.txt
> Use this file to discover other available public pages before exploring further.

# Building local inference on a DGX Spark
- URL: https://blog.kamelhar.net/serving-an-llm-on-an-nvidia-dgx-spark/
- Published: 2026-09-20T12:00:00.000Z
- Updated: 2026-09-23T16:07:24.000Z
- Description: A 35B model under k3s, one OpenAI-compatible endpoint in front of it, and the eight steps in between — with what each one taught me.
- Author: Federico Kamelhar
- Tags: infrastructure, nvidia, kubernetes, vllm, litellm

This post builds a local inference stack on a DGX Spark, from the GPU up: a 35-billion-parameter model served by vLLM under k3s, one OpenAI-compatible endpoint in front of it, and a routing table that decides, per request, whether a prompt may leave the machine. The finished stack serves **99 tokens per second** on a single stream and holds up under sixteen at once.

It is the exact setup, in the order you would build it — the reason for each piece, the configuration as it runs, and the places where the machine surprised me. If you have a Spark, or any single-GPU box you want to turn into something a team can point tools at, this is the path end to end.

![An NVIDIA DGX Spark on a wooden desk beside a laptop showing its setup screen](https://blog.kamelhar.net/content/images/2026/09/dgx-spark-on-desk-nvidia-c5c8c106d.webp)

The whole machine room this post runs on: a DGX Spark, about the size of a hardback book, on a desk. *Photo: NVIDIA.*

## The build at a glance

Eight layers, in the order you build them — what I chose, and why it matters. Each links to the step that explains it, and each step links to its code.

1. [Model](#step-1-pick-the-model-the-memory-rewards) — **Qwen3.6-35B-A3B, NVFP4.** Only 3B parameters are active per token, so each token reads \~1.5 GB instead of \~18 GB — on a bandwidth-bound machine, that is the speed.
2. [Host](#step-2-kubernetes-on-the-host-and-nothing-else-on-the-box) — **k3s on bare metal, nothing else on the box.** Memory is unified, so every other process takes room from the model.
3. [GPU](#step-3-make-one-gpu-look-like-four) — **NVIDIA device plugin, time-sliced four ways.** vLLM, speech-to-text and text-to-speech each claim a slice of one GPU.
4. [Serving](#step-4-serve-it-with-vllm-every-flag-earning-its-place) — **vLLM, bound to loopback.** FP8 KV cache, prefix caching and multi-token prediction are the three flags that matter most.
5. [Delivery](#step-5-deliver-it-without-anyone-running-kubectl) — **Argo CD ApplicationSets.** Adding a workload is an overlay plus a five-line file — nobody runs `kubectl`.
6. [Routing](#step-6-put-one-endpoint-in-front-local-first) — **LiteLLM, one OpenAI-compatible endpoint.** Local first; `spark-only` never leaves the machine; no client holds a vendor key.
7. [Web access](#step-7-give-agents-the-web-through-a-tool-you-control) — **An MCP server in front of self-hosted SearXNG.** The agent searches through a tool you control, not a socket in the model.
8. [Measurement](#step-8-measure-it-and-distrust-the-first-number) — **Fixed output, unique prompts, medians.** 99 tokens/sec on one stream, 459–469 across sixteen — and four assumptions that did not survive an experiment.

## What you end up with

**Qwen3.6-35B-A3B**, a Mixture-of-Experts model with 3B of its 35B parameters active per token, quantised to **NVFP4**, served by **vLLM** under k3s.

|                     |                                                                                                                                                        |
| ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
| One stream          | **99 tokens/sec**                                                                                                                                      |
| 16 streams          | **459–469 tokens/sec** aggregate, about **29 per stream**                                                                                              |
| Time to first token | 0.23 s at 500 prompt tokens, 6.23 s at 40,000                                                                                                          |
| Routing             | Local first. qwen3.6-35b falls back to a vendor only if the GPU cannot answer; **spark-only has no fallback and fails rather than leave the machine**. |

![The GB10 Grace Blackwell Superchip in a DGX Spark](https://blog.kamelhar.net/content/images/2026/09/nvidia_n-c74199283.webp)

The GB10 Grace Blackwell Superchip in a DGX Spark

![The system end to end. Everything above the dashed line stays on the machine; a request crosses it only when the router decides it may.](https://blog.kamelhar.net/content/images/2026/09/d2_system-c20b4ada4.webp)

**Figure 1.** The system end to end. Everything above the dashed line stays on the machine; a request crosses it only when the router decides it may.

**Everything here is runnable.** [kamelhar/blog-code](https://github.com/kamelhar/blog-code) has the harness, the manifests and the raw sweeps — including a generic concurrency benchmark you can point at your own server, whatever is serving it:

```sh
export LLM_API_KEY=...
./dgx-spark-local-inference/08-measure/llm-concurrency.py \
    --url http://127.0.0.1:8000/v1/chat/completions \
    --model my-model --concurrency 1,2,4,8,16 --max-tokens 400

```

Standard library only. It controls for the four things that otherwise produce a confident wrong number — fixed output length, a unique prompt per request so a prefix cache cannot flatter the run, a discarded warm-up, and per-stream rates measured rather than divided out. The repository is laid out the way this post is: one folder per step, with the configuration as it runs.

Every number here was measured in August and September 2026 on the machine described, with the harness and protocol in [how these numbers were taken](#how-these-numbers-were-taken). Where something is not measured, the text says so.

## Why this machine is interesting to build on

A DGX Spark puts 128 GB of unified memory on a desk — and *unified* is the interesting word. Not system memory with a GPU beside it, holding its own private copy. One pool, addressed by both, so a model is mapped where it already sits instead of being copied across PCIe into a card. A desktop can therefore hold a model that will not fit on a much faster accelerator.

That inverts the usual trade, and the inversions keep coming. This machine has about a twelfth of an H100's bandwidth and more of its memory, which makes **the architecture of the model the performance lever**, not the size of the GPU. A Mixture-of-Experts model with 3B of its 35B parameters active per token runs here at **99 tokens per second**. A dense model of the same parameter count would run roughly six times slower on the same hardware, untouched. Choosing the model well is worth more than any flag.

Three more that surprised me, each measured rather than assumed:

- A dense **4B** model was *slower to first token* than the sparse **35B**. Small model, big runtime penalty.
- Thinking mode **on** decodes *faster per token* than off — reasoning text is repetitive, so the speculative drafter lands more of it.
- Sixteen concurrent streams deliver 4.7 times the output of one — and the peak I had first measured turned out to be a **setting**, not the hardware.

**What the machine is shaped like:** 128 GB of memory at about 273 GB/s, against an H100 SXM's 80 GB at 3,350\. Roughly half again the capacity, at a twelfth of the speed. Read that as a shape rather than a shortfall — it is a machine built to *hold* a great deal and read it deliberately, which is a genuinely different design point and the reason every choice below goes the way it does. Single-stream speed is that bandwidth divided by the bytes read per token, so the lever is the second number: read less per token and the machine gets faster without changing. That is what sparsity buys, and [step 1](#step-1-pick-the-model-the-memory-rewards) works through the arithmetic.

Three reasons to want any of this. Agent workloads bill per token and are unusually input-heavy, so the meter runs fastest on the work that benefits most. Hosted models are revised and retired on someone else's schedule. And a fair amount of ordinary professional work is simply not yours to hand to a third party.

## Step 1: Pick the model the memory rewards

*Code:* [*01-model/*](https://github.com/kamelhar/blog-code/tree/main/dgx-spark-local-inference/01-model)

**In short.** On a machine with 273 GB/s of bandwidth, decode speed is bandwidth divided by the bytes read per token. Pick a sparse Mixture-of-Experts model and quantise it: the choice of model is worth more than any flag.

The GB10 in a DGX Spark reports 20 CPUs and roughly **128 GB of unified memory** shared between CPU and GPU, at about **273 GB/s**. For comparison, an H100 SXM has 80 GB at about 3,350 GB/s: less memory, roughly twelve times the bandwidth. (The PCIe and NVL configurations differ; the SXM figure is the one quoted here.)

The Spark is not a smaller H100\. It is a machine that can hold very large models and read them slowly, and almost every decision below follows from that asymmetry.

Token generation is memory-bound. Each decoding step reads the active weights out of memory, so throughput is approximated by a single ratio:

```
tokens/sec  ~=  memory bandwidth  /  bytes read per token

```

This is an estimate for memory-bound decoding at batch size one, and it assumes the read is the only thing that costs. It ignores attention over a growing context, kernel efficiency below peak bandwidth, and anything that yields more than one token per weight read — multi-token prediction (step 4) does exactly that. Treat the result as a ceiling rather than a forecast: 273 GB/s over 1.5 GB is about 180 tokens per second, measured decode lands well under it, and multi-token prediction recovers part of the gap.

For a **dense** model, "active weights" means all of them. A dense 35B model at 4-bit reads about 18 GB per token, which puts it near 15 tokens per second no matter how much memory sits unused.

> **The consequence**  
>  
> For a model this size at conversational speed on a memory-rich, bandwidth-poor machine, a Mixture-of-Experts architecture is not an optimisation — it is what makes the target reachable. A much smaller dense model is a different trade, and one I measured: on real agent-sized prompts a dense 4B was *slower* to first token than this 35B, because prefill dominates and the runtime mattered more than the parameter count.

### Why a Mixture-of-Experts model

![Qwen3.6-35B-A3B, quantised to NVFP4](https://blog.kamelhar.net/content/images/2026/09/qwen_n-c3b99416a.webp)

Qwen3.6-35B-A3B, quantised to NVFP4

A Mixture-of-Experts model activates only a fraction of its parameters for each token. The model used here has **35 billion parameters in total, but only 3 billion active** per token. At 4-bit that is about 1.5 GB read per decoding step rather than 18 GB, and the machine lands at **99 tokens per second on a single stream**.

Same memory, same bandwidth, roughly six times the speed of a dense model of equal size. The architecture of the model, not the size of the machine, is what made it usable.

Two further choices matter as much as the model family:

- NVFP4 weights. Four-bit quantisation puts the weights at about 18 GB instead of roughly 70 GB, and fewer bytes read per token is the speed.
- FP8 KV cache. Roughly doubles the conversation length that fits in the same memory. vLLM's startup log gives the split: **21.99 GiB of weights, 53.84 GiB of KV cache** (4,073,445 tokens, about 14 KiB each), and 0.82 GiB of CUDA graphs. That is more than the server was meant to have — see ["0.50 means half the memory"](#050-means-half-the-memory) — but the shape holds: the weights are the small part, and the reason to want 128 GB is that the cache is where the room goes.

## Step 2: Kubernetes on the host, and nothing else on the box

*Code:* [*02-host/*](https://github.com/kamelhar/blog-code/tree/main/dgx-spark-local-inference/02-host)

**In short.** Install k3s directly on the host — no Docker, no VM — and keep everything except model serving off the machine. With unified memory, an app on this box is memory the model does not get.

![The stack, hardware upward. Each layer exists to make the one above it possible; the device plugin is what turns a GPU into something the scheduler can reason about.](https://blog.kamelhar.net/content/images/2026/09/d1_stack-cdfc5d3ac.webp)

**Figure 2.** The stack, hardware upward. Each layer exists to make the one above it possible; the device plugin is what turns a GPU into something the scheduler can reason about.

The machine runs **k3s v1.36.4** on **containerd 2.3.4**, installed on the host operating system. There is no Docker daemon, no virtual machine and no hypervisor. A second, ordinary server runs the Argo CD control plane and every application; the Spark holds model serving and nothing else.

That separation is a written rule rather than a preference, and the reason is unified memory. An application, an agent loop or a database on this machine takes memory away from the model, and the model is the only reason the machine exists.

> **Why Kubernetes at all on one node**  
>  
> Not for scale. It is there so that every change is a pull request, the GPU is a schedulable resource rather than something processes fight over, secrets are ciphertext in version control, and a rebuilt machine converges to the same state without anyone remembering what they typed.

## Step 3: Make one GPU look like four

*Code:* [*03-gpu-sharing/*](https://github.com/kamelhar/blog-code/tree/main/dgx-spark-local-inference/03-gpu-sharing)

**In short.** The NVIDIA device plugin advertises one GPU as four, so three GPU workloads can be scheduled side by side. Time-slicing shares the GPU; it does not isolate it.

Kubernetes will not schedule a GPU it cannot see. Two pieces make that work.

The **NVIDIA device plugin** (`k8s-device-plugin v0.20.0`) runs as a DaemonSet and advertises `nvidia.com/gpu` on the node. Workloads then set `runtimeClassName: nvidia` so containerd hands them the NVIDIA container runtime.

Then the interesting part: one physical GPU is advertised as four.

```
sharing:
  timeSlicing:
    resources:
      - name: nvidia.com/gpu
        replicas: 4

```

### Time-slicing versus MPS

These are not interchangeable, and the difference is worth stating plainly.

|        | Time-slicing                                                                                                      | MPS                                             |
| ------ | ----------------------------------------------------------------------------------------------------------------- | ----------------------------------------------- |
| How    | The GPU context-switches between processes                                                                        | Kernels from several processes run concurrently |
| Memory | No isolation, and none available to configure. Every replica sees all of it and one can exhaust it for the others | Per-client memory budgets                       |
| Suits  | Trusted workloads that must coexist                                                                               | Workloads that must be fenced from each other   |

NVIDIA documents this directly: time-sliced replicas get neither memory nor fault isolation. A replica is a scheduling share, not a sandbox, and MPS gives budgets rather than the isolation a hostile neighbour would need. This deployment uses time-slicing, which in a multi-tenant cluster would be reckless. Here the three workloads sharing the GPU are all owned by the same person, their memory use is known and bounded, and the actual requirement is that speech-to-text can coexist with the language model rather than be fenced off from it. The node reports four GPUs, three pods claim one each, and the scheduler stops arguing.

### Where that decision actually lives

Time-slicing is often described as a Kubernetes feature. It is not. The device plugin only advertises a number to the scheduler and hands out allocations; the thing that actually interleaves work between processes is the kernel module, several layers below anything Kubernetes can see. That is also why there is no memory isolation to configure — the layer doing the sharing was never asked to provide any.

![The kernel path. A request crosses from user space to the kernel exactly once, through an ioctl on a character device; everything the scheduler knows about the GPU is an advertisement, and everything that is enforced is enforced below the line.](https://blog.kamelhar.net/content/images/2026/09/d6_kernel-ca4efe95c.webp)

**Figure 3.** The kernel path. A request crosses from user space to the kernel exactly once, through an ioctl on a character device; everything the scheduler knows about the GPU is an advertisement, and everything that is enforced is enforced below the line.

Two details in that picture matter more than the rest. The driver here is the **open kernel module** build (`580.173.02, aarch64`), and the container image contains none of it: `nvidia-container-runtime` injects the device nodes and the driver libraries at create time, which is why the same image runs on a machine with a different driver version.

The second is `nvidia_uvm.ko`. On a conventional server the weights are copied across PCIe into dedicated GPU memory before anything can run. Here the CPU and the GPU address the same physical pages, so a model is mapped where it already sits. That is the entire reason a machine with 128 GB of comparatively slow memory can hold a model that would not fit on a much faster card.

### What ends up running on the machine

| Namespace   | Workload             | Role                                                                             |
| ----------- | -------------------- | -------------------------------------------------------------------------------- |
| inference   | vllm                 | The language model. Claims one GPU.                                              |
| inference   | vllm-dev             | Scaled to zero. Exists to swap a candidate model in, never to run a second copy. |
| inference   | llm-router           | The single endpoint every client talks to.                                       |
| inference   | llm-router-dev       | Same model server, separate key, concurrency-capped, no cloud access.            |
| inference   | edge proxy           | TLS termination and internal names.                                              |
| inference   | speech-to-text       | whisper.cpp. Claims one GPU.                                                     |
| inference   | text-to-speech       | Claims GPU time when in use.                                                     |
| inference   | watchdog             | CronJob. Restarts a wedged model server.                                         |
| platform    | nvidia-device-plugin | DaemonSet. Advertises the GPU to the scheduler.                                  |
| kube-system | sealed-secrets       | Decrypts SealedSecrets inside the cluster.                                       |

> **One model copy, always**  
>  
> A second model server exists in version control but is permanently scaled to zero. Testing a new model means swapping the one that runs, never running two. Two copies would each hold their own weights and compete for what is left, cutting the cache available to either — not by exactly half, since the weights are duplicated and the runtime overhead is paid twice, but by enough that the long contexts this is sized for stop fitting.

## Step 4: Serve it with vLLM, every flag earning its place

*Code:* [*04-vllm/*](https://github.com/kamelhar/blog-code/tree/main/dgx-spark-local-inference/04-vllm)

**In short.** vLLM bound to loopback, one GPU slice, weights read from the host, and a startup probe patient enough for a long load. Every flag is there for a reason, and the ones that matter most were measured.

![vLLM serves the model; the router is its intended client](https://blog.kamelhar.net/content/images/2026/09/vllm_n-c0f762d46.webp)

vLLM serves the model; the router is its intended client

The Deployment is ordinary except in five places, and each of the five is doing real work.

```
runtimeClassName: nvidia          # containerd hands it the NVIDIA runtime
hostNetwork: true                 # binds 127.0.0.1 - off-host traffic cannot reach it
resources:
  limits: { nvidia.com/gpu: 1 }   # one time-slice
volumes:                          # ~22 GB of weights, read from the host disk
  - hostPath: /models
  - hostPath: ~/.cache/huggingface
startupProbe / readinessProbe / livenessProbe

```

**Host networking bound to loopback** means the model server is not reachable from another host. Within the machine it is reachable by any local process that knows the port, so the router is its *intended* client rather than its only possible one; what stops a casual bypass is an API key the router holds and other workloads do not. That is a meaningful control and not an isolation boundary. Making it one would take a network policy or a unix socket, neither of which is in place.

**The startup probe earns its place** because loading and compiling take many minutes — the probe allows thirty. A liveness probe alone would kill the container repeatedly before it ever finished starting — a classic and very confusing failure.

### The serving flags, and what each one buys

| Flag                            | Why it is there                                                                                                                                                                                                                                                                                         |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| \--gpu-memory-utilization 0.50  | Meant as half the unified pool for weights and KV cache, the rest left for speech-to-text, text-to-speech and the host. In practice the process holds about two-thirds — see ["0.50 means half the memory"](#050-means-half-the-memory).                                                                |
| \--kv-cache-dtype fp8           | Roughly doubles the context held in the same memory.                                                                                                                                                                                                                                                    |
| \--enable-prefix-caching        | A stable system prompt is not reprocessed on every turn of a conversation.                                                                                                                                                                                                                              |
| \--speculative-config mtp       | Multi-token prediction: several tokens per weight read. The most effective single-stream lever in my testing — vLLM's counters show 2.84 tokens accepted per target forward pass, 62 percent of drafted tokens. Raising num\_speculative\_tokens past 3 buys draft slots that land under half the time. |
| \--moe-backend marlin           | Quantised Mixture-of-Experts kernels.                                                                                                                                                                                                                                                                   |
| \--attention-backend flashinfer | Faster attention kernels.                                                                                                                                                                                                                                                                               |
| \--max-num-seqs 16              | How many sequences may be in flight. See the measurements — this was the limiting setting.                                                                                                                                                                                                              |
| \--tool-call-parser             | Tool calls arrive as structured blocks rather than prose an agent has to guess at.                                                                                                                                                                                                                      |
| \--max-model-len 131072         | Context window. Paged attention allocates KV by tokens actually used, not by this maximum.                                                                                                                                                                                                              |

## Step 5: Deliver it without anyone running kubectl

*Code:* [*05-delivery/*](https://github.com/kamelhar/blog-code/tree/main/dgx-spark-local-inference/05-delivery)

**In short.** Argo CD discovers workloads from a five-line marker file, so adding something to the machine is a pull request. Nothing reaches the cluster by hand.

![The deployment model. A workload arrives as an overlay and a five-line file; the control plane reads the destination out of that file rather than from a central list.](https://blog.kamelhar.net/content/images/2026/09/d3_deployment-cf6d8a7f4.webp)

**Figure 4.** The deployment model. A workload arrives as an overlay and a five-line file; the control plane reads the destination out of that file rather than from a central list.

Argo CD runs on the companion server and reconciles both its own cluster and the Spark. Two ApplicationSets scan the git repository for a marker file, and generate an Application for every one they find.

```
apps-hub       scans  apps/*/overlays/prod/deploy.yaml
apps-clusters  scans  apps/*/overlays/{lab-*,spark*}/deploy.yaml

```

Each overlay declares where it belongs, in five lines:

```
name: vllm
project: inference
cluster: spark
env: shared
namespace: inference

```

So adding a workload to the machine is a kustomize overlay plus that file. An overlay without one is simply not deployed, which is how something needing a hand-written Application opts out. Cluster-level infrastructure that cannot come from this pattern — the device plugin, the sealed-secrets controller — is declared as explicit Applications instead.

### Secrets

Vendor API keys and the model server key are SealedSecrets. Plaintext exists only on the machine itself; a script seals it against the in-cluster controller public key, and only ciphertext is committed. The controller decrypts it into a real Secret at apply time.

> **Why this matters beyond tidiness**  
>  
> No client anywhere holds a vendor key. That is what makes the routing table impossible to route around: a client that cannot reach a vendor directly cannot leak the credential and cannot bypass the limit.

## Step 6: Put one endpoint in front, local first

*Code:* [*06-router/*](https://github.com/kamelhar/blog-code/tree/main/dgx-spark-local-inference/06-router)

**In short.** One LiteLLM endpoint decides, per request, whether a prompt may leave. The default model name falls back to a vendor only if the GPU cannot answer; `spark-only` never does.

Routing used to live in each client configuration. One chat gateway had a fallback chain, an editor had another, and anything new would have needed a third. That is one decision written in several places, which is how they drift — and it meant every client needed a vendor key to reach the cloud at all.

Now there is one OpenAI-compatible endpoint: a **LiteLLM** proxy running beside the model server. It speaks the OpenAI API on both sides — clients call it as though it were OpenAI, and it calls vLLM, Together AI and OpenRouter the same way. That symmetry is the point: a local model and a hosted vendor become interchangeable behind one name, so the choice between them stops being a client concern and becomes a line in one table.

> **Disclosure**  
>  
> I contributed the OCI provider upstream to LiteLLM, so I am not a disinterested party about the project. The argument that follows is about topology rather than about which proxy you pick — it holds for any local router.

Three things sit behind that endpoint, and they are not interchangeable with each other:

|                 | What it is for                                                                                                                                                                                        |
| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| vLLM on the GPU | The default. No per-token charge, no request leaving the machine, and fast enough that nothing else is needed for ordinary work. Electricity and the hardware are not free; the marginal token is.    |
| Together AI     | Named rungs for open-weight models this machine cannot hold — larger Mixture-of-Experts models, and anything needing more memory than 128 GB. Priced per token, cheap by construction.                |
| OpenRouter      | One key in front of hundreds of models from many vendors, including frontier ones. Used as a passthrough for names the other two cannot serve, each with a fallback chain that ends at the local GPU. |

The distinction that matters: The Together AI entries are a small number of models I chose deliberately and named individually. The OpenRouter entry is a catalogue I have not read, reachable by pattern. The first is a decision; the second is an escape hatch, and it is capped precisely because it is not a decision.

### What the table actually says

| Model name asked for      | Served by     | Why it exists                                                                                                                                                          |
| ------------------------- | ------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| qwen3.6-35b (the default) | the local GPU | **Falls back to Together AI if the machine cannot answer at all** — so the default trades a guarantee for availability. Use spark-only when the prompt must not leave. |
| spark-only (pinned)       | the local GPU | No fallback chain. It fails rather than send the prompt anywhere.                                                                                                      |
| cloud- prefix             | Together AI   | Named rungs, chosen in advance. Cheap by construction.                                                                                                                 |
| heavy- prefix             | Together AI   | Larger open models, reachable only by name. Nothing routes here on its own.                                                                                            |
| openrouter/\*             | OpenRouter    | Hundreds of models behind one pattern, each falling back through cheaper models to the local GPU.                                                                      |
| anything unrecognised     | the local GPU | A name nobody configured is served locally rather than refused.                                                                                                        |

> **Why an unknown name is answered rather than refused**  
>  
> This is the entry I am least sure of, and the one that most nearly contradicts one of [the beliefs measurement overturned](#what-i-got-wrong-along-the-way), further down. An unrecognised name is almost always a typo or a client with a stale default, and on a single-user machine the cheapest recovery is an answer from the local model rather than an error at three in the morning. The cost is exactly the failure that section warns about: a plausible answer from a model you did not ask for. What makes it tolerable here is that the catch-all resolves *inward* — the wrong answer is local and private, never a silent spill to a vendor — and that anything requiring a specific model asks for the pinned name, which has no fallback. On a shared deployment I would make it fail loudly instead.  
>  
> **The pinned name is the one worth arguing for**  
>  
> Privacy cannot be inferred from content — a model cannot tell that this paragraph is the sensitive one. So it has to be selectable, and it has to be a guarantee rather than a preference. That entry has no fallback chain: it fails rather than quietly sending the conversation to a vendor. What leaves in a fallback is the prompt, before any answer exists, which is why the default name's fallback is a disclosure decision and not an availability one.

![A single request, step by step. Step 3 is the one that matters: the decision about whether the prompt may leave is taken locally, before any network call.](https://blog.kamelhar.net/content/images/2026/09/d4_sequence-c14b97f49.webp)

**Figure 5.** A single request, step by step. Step 3 is the one that matters: the decision about whether the prompt may leave is taken locally, before any network call.

### Why OpenRouter cannot replace the local router

This is the question worth answering, because OpenRouter already does much of what LiteLLM does here — one endpoint, many providers, automatic failover, a single key. The difference is not features. It is that **a hosted gateway cannot reach this GPU**, because the model server binds loopback and nothing routes to it from outside the machine. "Local first, always" is therefore a policy that can only be evaluated here.

That is a property of this topology, not a law. Expose the model server on a private network, or put a tunnel in front of it, and a hosted gateway could route to it — at the cost of the thing the loopback bind was for. The argument is not that a remote gateway can never reach a local model. It is that the decision about whether a request may leave should not itself be taken somewhere the request has already gone.

So the two layers answer different questions, and both are needed. The local router decides whether a request may leave at all. OpenRouter decides which vendor serves it once that decision has already gone against staying home. Collapsing them would mean every request leaving the building to be told it should have stayed.

### Spend is bounded at the vendor account

Opening the OpenRouter passthrough puts hundreds of models one string away, including frontier models at several dollars per million tokens. LiteLLM cannot cap that here: its budget features require a database, and this deployment runs without one, so those settings silently do nothing. So spend is bounded where the money actually is — the OpenRouter account's own credit limit — rather than by a setting in a layer that cannot enforce it.

A per-token price filter in front of it looked like a better idea, and I ran one for a while. It is in [what I got wrong](#what-i-got-wrong-along-the-way).

## Step 7: Give agents the web through a tool you control

*Code:* [*07-web-search/*](https://github.com/kamelhar/blog-code/tree/main/dgx-spark-local-inference/07-web-search)

**In short.** Agents get the web through an MCP server over stdio, backed by a self-hosted metasearch instance. The boundary is what that tool is allowed to send — and today nothing redacts the query.

A useful agent needs current information, and a local model has none: its knowledge stops at its training cutoff, and it has no way to fetch anything. The obvious fix — give the model internet access — would undo the boundary this whole design is built around.

The Model Context Protocol offers a better shape. MCP lets an agent call tools that run outside the model, as ordinary local processes, and receive their output as text. The model asks for a search; a separate process performs it; the results arrive as text the model can read.

It is worth being exact about what that does and does not buy. The model not opening a socket is not itself the boundary — a model has no sockets to open in either design. The boundary is that every byte leaving the machine passes through a process I wrote the configuration for, whose network path and logs exist independently of the model. What matters is not who dials, but what the tool is permitted to transmit, and whether that permission is inspectable. A tool given an unrestricted HTTP client would move the boundary back to where it was, with extra steps.

### How it is wired here

The agent runs a small MCP server as a child process, speaking over standard input and output — no network listener, no exposed port. That server queries a self-hosted metasearch instance running as an ordinary workload in the application cluster, which in turn talks to public search engines and returns aggregated results.

```
agent  --stdio-->  MCP search server  --HTTP-->  self-hosted metasearch
                                                        |
                                                        +--> public search engines

the model itself opens no connection at any point

```

Three properties fall out of that arrangement, and each of them is the reason to prefer it over simply letting the model browse.

- Searching is a decision, not a capability. The agent must choose to call the tool, and that call is visible in the transcript. Nothing happens implicitly in the middle of generating an answer.
- The query leaves; the conversation does not. What crosses the boundary is a search string, not the transcript it came from, and a self-hosted metasearch layer means the engines see that instance rather than the person. This is a reduction in what leaves, not a guarantee. The agent composes the query, and an agent can copy a client name, a file path or a line of unreleased code straight into it. Today nothing validates or redacts that string, which makes the tool boundary a place where an obvious control is missing rather than one where the problem is solved: the queries are logged and reviewable after the fact, and that is all. A pattern filter on outbound queries is the next thing to build here.
- The tool runs where policy can reach it. It is an ordinary process on a machine that is administered, with its own network path and its own logs, rather than an opaque capability inside a model.

> **A detail that generalises**  
>  
> The search service is fronted by single sign-on for humans, which returns a redirect to anything without a session — including the MCP bridge. Rather than open a path through that sign-on, the bridge is pointed at the address the proxy itself forwards to, on the same host. Be exact about what that is: it **bypasses** the sign-on check rather than satisfying it. It is defensible because the backend binds an address reachable only from that host and the bridge is a process on it — the same position the proxy occupies — so nothing on the network gained access. It is still a second door, and it holds only as long as that bind stays local and the host stays trusted. The alternative, widening the sign-on to admit a service account, would have been a change to the door everyone uses.

The same pattern extends to every other tool an agent is given: a filesystem, a ticket tracker, a database. The question is never whether the model may reach it, but which process reaches it on the model’s behalf, and what that process is allowed to do.

## Step 8: Measure it, and distrust the first number

*Code:* [*08-measure/*](https://github.com/kamelhar/blog-code/tree/main/dgx-spark-local-inference/08-measure)

**In short.** Fix the output length, give every request a unique prompt, discard the warm-up and report both aggregate and per-stream numbers. Otherwise you are measuring the prefix cache.

Same prompt, 400 tokens generated, varying the number of simultaneous requests:

| Concurrency | Aggregate throughput | Per stream | Note                                                         |
| ----------- | -------------------- | ---------- | ------------------------------------------------------------ |
| 1 stream    | 99 tokens/sec        | 99         | Bandwidth-bound.                                             |
| 8 streams   | 304–354 tokens/sec   | 38–44      |                                                              |
| 16 streams  | 459–469 tokens/sec   | 29         | 4.7x aggregate, and each stream at 29 percent of solo speed. |

The second column is the one a person feels. Sixteen streams do 4.7 times the total work, and each of them generates at under a third of the rate it would get alone — so a single answer takes roughly three times longer to finish. That is the trade: **aggregate throughput up, individual answers slower**, and which you care about depends entirely on whether anyone is waiting on one of them.

The per-stream column is aggregate throughput divided by the number of streams, not a separately measured per-request rate. The harness records both; the runs quoted here kept only the aggregate, so treat the per-stream figure as the mean it is rather than as something measured per request.

Time to first token, by prompt size:

| Prompt        | First token | What dominates                   |
| ------------- | ----------- | -------------------------------- |
| 500 tokens    | 0.23 s      | Nothing. Instant.                |
| 4,000 tokens  | 0.81 s      | Prefill begins to show.          |
| 16,000 tokens | 0.80–2.80 s | Depends on cache state.          |
| 40,000 tokens | 6.23 s      | Prefill, which is compute-bound. |

### Decode speed depends on what is being generated

A separate single-stream profile — 400 tokens, `min_tokens` and `ignore_eos`, temperature 0, median of three with the warm-up discarded:

| What it is generating   | tokens/sec |
| ----------------------- | ---------- |
| prose, thinking on      | 105.9      |
| prose, thinking **off** | 88.8       |
| code                    | 110.6      |

**Thinking on is faster per token than thinking off**, which is the opposite of the intuition. It is a drafter effect: multi-token prediction proposes several tokens per weight read and reasoning text is repetitive and self-similar, so more of each proposed block survives verification than does on a polished final answer. That is consistent with vLLM's own counters — 2.84 tokens accepted per forward pass, 62 percent of those drafted — though those are cumulative since the server started and do not separate thinking from non-thinking.

That is a per-token credit and not a total one — extended thinking still emits far more tokens, so it costs more wall time overall. But the rate at which they arrive goes up, not down. At temperature 0.7, which is what real clients send, prose measures 90–99 tokens/sec — so the 99 from the concurrency sweep is a fair number for real use, not a best case.

### Two things follow

**Batching amortises the expensive part.** Decoding is dominated by reading weights, and one forward pass reads them once for the whole batch — so the cost that dominates at one stream is shared at sixteen. It is not free: attention is per-sequence and grows with context, the scheduler does more work, and in a Mixture-of-Experts model different sequences in a batch can route to different experts, so the set of weights a batched pass actually touches can be larger than the \~1.5 GB one sequence needs. I did not measure expert overlap or real memory traffic, so read the 4.7x as the measured outcome and the explanation as the mechanism it is consistent with. The first sweep peaked at eight streams and then *fell* at twelve — the signature of requests queuing behind a concurrency cap rather than hardware saturating. The cap was the setting, not the GPU. Raising it moved peak throughput up 31 percent.

**For a single user, none of that matters.** One session gets 99 tokens per second, and none of the concurrency settings move it: it is bounded by bandwidth over bytes read per token, and it is already about fifteen times faster than a person reads. One thing does move it — multi-token prediction, which yields several tokens per weight read. What is fixed is the ceiling that bandwidth sets, not the number itself. The concurrency headroom is real, but it only pays if the work fans out: parallel agents, batch jobs, overnight processing.

### How these numbers were taken

Numbers without a protocol are anecdotes. This is the whole of it, and the harness, the manifest and the raw runs are at [kamelhar/blog-code](https://github.com/kamelhar/blog-code) — including a generic version of the concurrency harness that takes any OpenAI-compatible endpoint, so you can run it against your own server rather than take my word for mine.

**What was served.** `nvidia/Qwen3.6-35B-A3B-NVFP4` at revision `491c2f1ea524c639598bf8fa787a93fed5a6fbce`, read from local disk rather than pulled at start and checked against a lock file of every file's SHA-256, served as `qwen3.6-35b`. The server is the upstream vLLM OpenAI image, pinned by digest rather than by tag, so the build cannot move under a rerun:

```
vllm/vllm-openai@sha256:c5fa18e5360a929262f55c697ea53d963288535fa0b238c0c47e9267a16e13b9

```

**The launch, in full.** Every flag argued for in [step 4](#step-4-serve-it-with-vllm-every-flag-earning-its-place), in the order the container receives them:

```
vllm serve /models/Qwen3.6-35B-A3B-NVFP4 \
  --served-model-name qwen3.6-35b \
  --host 127.0.0.1 --port 8000 \
  --gpu-memory-utilization 0.50 \
  --max-model-len 131072 \
  --max-num-batched-tokens 16384 \
  --max-num-seqs 16 \
  --enable-auto-tool-choice \
  --enable-prefix-caching \
  --moe-backend marlin \
  --attention-backend flashinfer \
  --load-format fastsafetensors \
  --kv-cache-dtype fp8 \
  --async-scheduling \
  --tool-call-parser qwen3_xml \
  --reasoning-parser qwen3 \
  --default-chat-template-kwargs '{"enable_thinking": true}' \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3,"moe_backend":"triton"}'

```

**The workload.** A concurrency sweep fires N identical-shaped requests at once against the OpenAI chat endpoint and reports time to first token, wall time, per-stream decode rate and the aggregate across streams. Output is fixed at 400 tokens with `min_tokens` and `ignore_eos` set, so every request does the same amount of decoding and a short answer cannot flatter a run. Temperature 0\. Median of three, warm-up discarded.

**Cache state, which changes everything.** Each request gets a unique leading string, so the sixteen streams do **not** share a warmed prefix. This matters more than it sounds: an earlier sweep that sent the identical 16,000-token prompt to every request measured 5.01 s at concurrency 2 against 15.41 s with unique prompts, because it was measuring vLLM's prefix cache rather than the machine. Roughly a threefold difference, entirely an artefact of the harness. The concurrency numbers above are the unique-prompt ones.

That cuts the other way in the time-to-first-token table, where the 16,000-token row reads 0.80–2.80 s. The spread is exactly the cache: the low end is a warm prefix, the high end a cold one. Prefix caching is why a stable system prompt pays for itself in a real conversation, and why a benchmark that leaves it warm is not measuring the hardware.

**What is not here.** Per-request latency percentiles under concurrency are collected by the harness but not reproduced above; p95 diverges from p50 substantially once the box is saturated, which is the number to look at before putting interactive users behind a busy machine. The sweep was run on an otherwise idle GPU, with speech-to-text and text-to-speech scaled down — real contention costs more than these tables show.

## What I got wrong along the way

Each of these passed a plausibility check and failed an experiment. On a machine like this the characteristic failure is not a crash: it is a setting that looks correct, answers confidently, and is quietly doing nothing.

### "Per-deployment limits need a shared counter"

An earlier experiment had shown a requests-per-minute limit doing nothing, and that result was generalised to all concurrency settings. Then eight simultaneous requests against the capped router serialised into four clean waves of two, while the same eight against the uncapped one ran together. A parallel-request cap is an in-process semaphore that a single instance holds perfectly well; requests-per-minute is a rolling window that does need shared state. Right about one setting, wrong to extend it to both.

### "The more specific pattern obviously wins"

Adding a passthrough created two matching patterns: a vendor prefix and a catch-all. Had the catch-all won, every cloud request would have been served by the local model, silently, with no error — plausible answers from the wrong model. It was measured before it was trusted, and the result pinned with a test.

### "Declared capability means it works"

One model advertised tool-calling support in both the proxy metadata and the vendor catalogue, and still failed: the only endpoint left after filtering required a credential that was not configured. No catalogue field predicts who is actually serving a model at this moment. Filtering shrinks the surface; it cannot remove it.

### "A price ceiling protects you"

With the OpenRouter passthrough open, I capped it with a per-token price filter — OpenRouter's `max_price`, at $1 in and $3 out per million tokens — so one string could not silently select a model twenty times the intended rate. It did exactly that, and that was the problem: asking for a frontier model returned "no endpoints satisfy the max price", and because only the default name had fallbacks, the request had nowhere to go. The ceiling did not stop spend; it stopped work.

It is gone. Every OpenRouter model now has a generated fallback chain — cheaper models from the same provider, then other providers, then the local GPU — and spend is bounded by the OpenRouter account's own credit limit, which is where the money is. `spark-only` and the large models stay unreachable by accident.

### "0.50 means half the memory"

`--gpu-memory-utilization 0.50` reads like a ceiling: half of the 121.7 GiB the GB10 reports, about 61 GiB, and the other half left for everything else. It is not what the server holds. `nvidia-smi` puts the vLLM process at **81.7 GiB**, and with speech, the routers and the host running, `free` shows **21 GiB available** out of 121 — not the sixty I had been planning around.

The startup log has the reason in one line:

```
Estimated CUDA graph memory: -20.50 GiB total
Available KV cache memory: 53.84 GiB

```

vLLM sizes the KV cache as the budget minus the weights, minus the peak activations of a profiling pass, minus its estimate of what CUDA graphs will take. On this machine that estimate came out **negative**, so subtracting it added twenty gigabytes: 61 − 22 − a few GiB of activations + 20.5 ≈ 54, which is the cache it built. The graphs themselves then took 0.82 GiB. I have not traced why the estimate goes negative; unified memory is the obvious suspect, because "free memory" on this box moves whenever any other process allocates, and a profiler that measures a before-and-after difference in free memory is measuring the neighbours too.

Nothing failed, which is why it went unnoticed. The model served, the benchmarks held, and the extra cache simply sat there. It surfaced only when I sized a fine-tuning job to run beside the server: a 4B LoRA run peaks near 19 GiB, and 21 GiB of headroom is not a margin on a machine where running out of memory takes the whole box down rather than one process.

The extra cache also buys little. At about 14 KiB a token, sixteen sequences at the full 131,072-token context need about 29 GiB. The other \~25 GiB can never hold a live sequence; it only keeps old prefixes cached a little longer. vLLM can be told the cache size directly with `--kv-cache-memory-bytes`, which skips the profiler's arithmetic altogether — that is the fix, and the lesson is the one this section keeps teaching: the flag's name is a claim, and the startup log is the measurement.

## Starting your own

If you are putting a model on your own hardware, these are the five things I would do first — in this order.

1. **Choose the model for your bandwidth, not your memory.** Divide memory bandwidth by the bytes read per token before you download anything. On a Spark that points straight at a Mixture-of-Experts model.
2. **Run exactly one model server on the box.** Test a candidate by swapping it in, never by running two. Each copy holds its own weights, and they split what is left for the cache that long contexts need.
3. **Give sensitive work a model name with no fallback.** Privacy cannot be inferred from a prompt, so make it something a person selects — and make that name fail rather than leave.
4. **Benchmark with unique prompts and a fixed output length.** A shared prompt measures the prefix cache; a free output length rewards a terse model. Both give you a confident wrong number.
5. **Test every assumption you are relying on.** Five of mine did not survive an experiment, and none of them raised an error.

## Where this leaves you

A DGX Spark is not a small data-centre GPU; it is a differently shaped one, and that shape is the whole interest of it. Enormous capacity read deliberately rewards a sparse model, and a sparse model is what makes a machine you can put under a desk hold something genuinely capable, permanently, answering faster than anyone can read, with the conversation staying on the premises unless a request asks for a name that says otherwise.

**What this article does not establish.** Every number here answers a narrow question: how fast this configuration produces tokens. Speed is necessary and it is not sufficient. Whether tool calls succeed on real agent tasks, whether a coding or investigation task completes, and what NVFP4 quantisation costs in answer quality are all unmeasured here — and quantisation quality is exactly the kind of thing that does not show up in a throughput sweep. A small labelled task set is the obvious next piece of work, and until it exists, read this as an architecture and serving report rather than an answer about usefulness.

The parts worth copying are not the hardware choices. They are the boundary that is enforced rather than intended, spend bounded where the money actually is, the single endpoint that keeps one decision in one place, and the habit of measuring a setting instead of trusting that it works.

---

All measurements taken August and September 2026 on a DGX Spark with a GB10 Grace Blackwell Superchip. NVIDIA, vLLM and Qwen are trademarks of their respective owners; the logos appear here to identify the technologies described. The DGX Spark photographs are NVIDIA's, from its press materials.

## Next steps

- **Run the benchmark against your own server:** [08-measure/llm-concurrency.py](https://github.com/kamelhar/blog-code/tree/main/dgx-spark-local-inference/08-measure) takes any OpenAI-compatible endpoint and controls for the four things that otherwise produce a confident wrong number.
- **Read the router config as it runs:** the model list and fallback block are in [06-router/](https://github.com/kamelhar/blog-code/tree/main/dgx-spark-local-inference/06-router), the vLLM Deployment in [04-vllm/](https://github.com/kamelhar/blog-code/tree/main/dgx-spark-local-inference/04-vllm).
- **Start with the name that cannot leave:** if you take one thing from this, make it `spark-only`, the entry with no fallback. It is one block of YAML, and it is the difference between a policy and a hope.