I pulled the ethernet cable out of my workstation on purpose last week — first time in years I’d done that intentionally — and watched a model keep generating text anyway, with nothing on the other end of the wire to send it to.
That’s the actual point of this setup, and it’s a narrower point than the news cycle wants to make it. I don’t need a position on whether cloud AI is some slow-motion catastrophe. What I do need a position on is a lot more boring and a lot more immediate: every prompt sent to a cloud API is proprietary code, private keys, or half-formed thoughts typed without editing, sitting on someone else’s server the second you hit enter. The moment that data leaves the machine, you’ve lost any real claim to having secured it. The fix isn’t philosophical. It’s just not sending it anywhere in the first place.
The Physical Layer #
Running a multi-billion-parameter model locally is a real compute and memory problem before it’s anything else, and it’s not something you casually pip install on a daily-driver machine and hope for the best. I run mine on a Dell Precision specifically because sustained inference needs both the memory headroom and the cooling to hold steady — the same thermal argument from a couple posts back about Latitudes throttling under load applies here too, just with more riding on it.
Whether the actual bottleneck is system RAM or GPU VRAM depends entirely on whether you’re running on CPU or a discrete GPU — they’re not the same resource, and conflating them is exactly the kind of imprecision that gets a whole setup wrong. On a CPU-only run, the constraint is system RAM and the model’s file size, same math as the quantization piece from a few weeks ago. Drop a discrete GPU into the mix and VRAM becomes its own separate, usually smaller, ceiling — and it fills up first.
Locking Down the Software Stack #
Raw hardware means nothing without containment. I run the entire AI stack inside its own virtual environment, fully detached from everything else on the machine, so a bad dependency in one project can’t quietly reach into libraries another project depends on. It’s not paranoia. It’s just not trusting a stack I didn’t personally audit to behave itself around files it has no business touching.
Model loading uses memory-mapping instead of reading the whole file into RAM up front — real, standard behavior in llama.cpp-style loaders, and it matters here specifically. Mapping the weights lets the OS page sections in and out on demand instead of committing the entire file’s size to RAM the instant it loads. That’s a genuine efficiency win, though it’s the same underlying mechanism that can quietly turn into swap thrashing if the model’s too big for the machine in the first place — mapped or not, the physics from a few posts back still applies once you overflow what’s actually available.
From a Voice Script to a Real Pipeline #
I hacked together an offline text-to-speech script years back — nothing fancy, just pyttsx3 and a lot of patience — and that was the first proof I had that local, offline interaction was possible at all on hardware I owned outright. This is the same philosophy, scaled up by several orders of magnitude: quantized weights downloaded once to a LUKS-encrypted drive, the network connection severed, and the model generating every token with nowhere to send a copy of it even if it wanted to.
There’s no outbound telemetry because there’s no outbound path. No hidden API pis, because there’s no API. Nobody’s logging the prompts on the other end, because there is no other end.
Verify the Air Gap — Don’t Just Assume It #
Zero trust has to apply to your own setup too, not just the cloud you’re avoiding. Unplugging a cable and assuming you’re air-gapped is exactly the kind of unverified claim that gets flagged elsewhere in this series, so here’s a script that actually checks instead of taking anyone’s word for it.
Run it before trusting the setup, not after something’s already gone wrong:
-– AIR GAP VERIFICATION —
PASS: no outbound route found on the tested endpoint.
-– ACTIVE CONNECTIONS TO NON-LOCALHOST ADDRESSES —
None found.
-– MEMORY STATE —
RAM used: X.XX GB / XX.XX GB
Those memory figures are placeholders on purpose — the real ones come from actually running it on your own machine. If the first check comes back FAIL, nothing else in this article matters until that’s fixed.
Cloud Inference vs. Air-Gapped Node #
| Dimension | Cloud Inference | Air-Gapped Node |
|---|---|---|
| Data exposure | Prompts and outputs pass through a third-party API, logged or not at their discretion | Never leaves the machine — there's no network path for it to travel on |
| Latency source | Network round-trip plus provider queue time | Local compute only, no network hop to wait on |
| Availability | Dependent on provider uptime, rate limits, and your own connection | Available offline, indefinitely, regardless of anyone else's outage |
| Model control | Whatever version the provider currently serves, changeable without notice | Whatever weights you downloaded, staying exactly as they are until you change them |
| Cost structure | Per-token or subscription billing, scales with usage | Fixed hardware cost up front, marginal cost near zero after that |
| Verification | Trust the provider's stated privacy policy | Check the connection table yourself — don't just assume |
None of this makes cloud inference wrong for every use case — plenty of workloads genuinely don’t care where the tokens get generated. It makes it the wrong default for anything you wouldn’t
want sitting on someone else’s server, which turned out to be most of what I actually do all day.
The cable’s still unplugged. I checked.
