Salvatore Sanfilippo, better known as antirez and best known as the creator of Redis, has a new project: DwarfStar 4, or ds4, a small C inference engine that runs frontier open-weight models like DeepSeek V4 and GLM 5.x locally on high-memory hardware. It went up on Hacker News and picked up momentum fast, and the GitHub repo now sits north of 22,000 stars. The pitch is not “another llama.cpp.” It is the opposite: a deliberately narrow engine built around a handful of specific models and specific machines.

What ds4 is, and refuses to be

ds4 targets high-memory Macs, CUDA machines like NVIDIA’s DGX Spark, and ROCm hardware like AMD’s Strix Halo. Metal is the primary target. It is MIT licensed and written mostly in C with Metal kernels, plus CUDA and ROCm backends.

The narrowness is explicit. ds4 is not a general GGUF runner. It only loads the specific GGUF layouts the project produces; arbitrary GGUF files will not work. That sounds like a limitation until you see what it buys. The project credits llama.cpp and GGML for the kernels, quantization formats, and GGUF ecosystem, and does not link against GGML. Instead of supporting every architecture ever shipped, ds4 hand-tunes the path for the few frontier open-weight models people actually want to run: DeepSeek V4 and V4.1 Flash, GLM 5.x, and Qwen3.8 Flash Next.

On the Hacker News threads, commenters point out why this matters: llama.cpp and LM Studio do not support DeepSeek V4 yet, so ds4 is effectively the only local option for that model unless you own hardware capable of running vLLM. For owners of 96GB or 128GB machines, it fills a gap that no generalist tool currently covers.

The engineering that makes it fit

Three design choices do the heavy lifting.

First, asymmetric 2-bit quantization. The routed experts in mixture-of-experts models get crushed down to 2 bits while critical shared paths stay precise. That is what lets models of this size run on a 128GB Mac at all. Second, SSD streaming: weights larger than RAM still run, because the routed experts are cached in memory and loaded from the GGUF file on cache misses. Third, a KV cache that persists to disk and resumes by prompt hash, so restarting a session does not force a full re-prefill. The server also keys cached prefixes to reuse shared prompt prefixes across sessions.

The published benchmarks put an M5 Max with 128GB at 39.4 tokens per second generation and 790 tokens per second prefill at a 2,048-token context with the q2 quantization. A DGX Spark with 128GB manages 18.1 tokens per second generation and over 825 tokens per second prefill. Those are not server-class numbers, but they are interactive numbers, and they come from a machine that fits on a desk.

There is also parallelism for the ambitious: pipeline parallelism can split transformer layers across multiple machines, and tensor parallelism can pair two Macs, say two 128GB machines, over RDMA.

Three binaries, one model state

The deliverable is three interfaces that share one model state and cache: a CLI (./ds4), an HTTP server (./ds4-server) exposing both OpenAI and Anthropic compatible APIs, and a native coding agent (./ds4-agent). The server is the interesting one for most people. Point your existing tools, OpenCode, Codex CLI, Claude Code, at the local base URL and they talk to the frontier model on your desk as if it were a hosted endpoint. Your prompts and your code never leave the machine.

Antirez describes the code as alpha quality and says he built it in about a week, working roughly 14 hours a day, with heavy assistance from coding agents. That last fact is worth a pause on its own. A single developer shipping a usable inference engine for frontier-scale models in a week would have been a nonsense claim two years ago. The agent-assisted workflow did not replace his judgment; it multiplied his output. The repo shows a systems programmer who knows exactly which 5% of the problem deserves hand-tuned kernels and refuses to touch the other 95%.

Should you use it

The deeper signal here is about the local-versus-cloud gap. A 2-bit quantized frontier model running interactively on a desktop Mac was the boundary of what anyone thought possible a year ago. The gap between “what runs on my desk” and “what runs in a hyperscaler” keeps shrinking, and it is shrinking fastest where people are willing to write narrow, unfashionable, hand-tuned C instead of waiting for a generalist tool to catch up.

Leave a Reply

Your email address will not be published. Required fields are marked *