ds4: antirez’s Narrow Inference Engine for Running Frontier Models Locally

Salvatore Sanfilippo, better known as antirez and best known as the creator of Redis, has a new project: DwarfStar 4, or ds4, a small C inference engine that runs frontier open-weight models like DeepSeek V4 and GLM 5.x locally on high-memory hardware. It went up on Hacker News and picked up momentum fast, and the […]

DeepSeek-V4.1-Flash squeezes the KV cache down to 890 bytes per token

DeepSeek-V4.1-Flash looks like a minor version bump until you read what the team actually shipped. The new model is a 552B-parameter multimodal mixture-of-experts system that supports million-token contexts, and its whole design centers on one number: 890 bytes per token of global KV cache. That is roughly a quarter of what DeepSeek-V4-Flash needed in HBM […]

OpenAI and Anthropic slash prices as Chinese AI rivals close the gap

OpenAI and Anthropic are slashing prices on their mid-tier AI models as cheaper Chinese alternatives gain traction among cost-conscious businesses and developers. The price cuts mark a shift from competing purely on model performance to competing on cost, with Chinese labs like DeepSeek and Moonshot forcing the market down. OpenAI cut prices on GPT-5.6 Luna, […]

DeepSeek Harness: the plugin first framework for AI agents

DeepSeek’s open source agent framework, DeepSeek Harness, exploded onto GitHub this week and picked up close to 70,000 stars in a single day. The repo’s tagline is simple and a little audacious: “everything is a plugin.” That phrasing is doing a lot of work, and it’s worth unpacking what the project actually ships, because the […]