Jeff fine-tunes tiny models to make decisions in one forward pass

Most LLM calls in production are not conversations. They are routing decisions: which support queue does this ticket belong in, is this comment abusive, which intent did the user just express. A new open-source project called Jeff targets exactly that workload with fine-tunes of Qwen3.5 (0.8B and 2B) and Gemma 4 E2B that return a […]

Microsoft quietly retires the Copilot+ PC brand

The label is gone, the spec stays Two years after launching Copilot+ as the badge of honor for AI-ready Windows hardware, Microsoft has stopped using it. The new Surface Pro 12-inch and Surface Laptop 13-inch announced this month carry no Copilot+ branding at all, even though both comfortably clear the hardware bar the label was […]

Claude Opus 5.5 cuts prices and rethinks how the model writes

The price is the story Claude Opus 5.5 is out, and the most interesting number on the announcement page is not a benchmark score. Input tokens now cost $4 per million, down from $5 on Opus 5. Output fell from $25 to $20 per million. Cache reads dropped hardest, from $0.50 to $0.20 per million. […]

Xiaomi’s MiMo-V2.6 takes the open-weight crown and shows its RL receipts

A car company just published the strongest open weights on the board Xiaomi released and fully open-sourced MiMo-V2.6 on September 21, and the flagship Pro model lands at 46 on the Artificial Analysis Intelligence Index. That score puts it ahead of every other open-weight model on that leaderboard, above xAI’s Grok 4.6 and Google’s Gemini […]

Qwen Image 2.1 lands as open weights with a license catch

Alibaba’s Qwen team released Qwen Image 2.1 as open weights on September 20, and it hit the top of Hacker News within hours. The model unifies text-to-image generation and image editing in a single 7B-parameter network, adds native transparency support, and ships with day-zero integrations across Diffusers, ComfyUI, and vLLM. There is one detail you […]

Alibaba open-sources RADAR, an expert-level AI for abdominal CT diagnosis

Alibaba’s research arm has open-sourced RADAR, a vision-language model that reads contrast-enhanced abdominal CT scans and flags close to 150 distinct conditions, from liver and pancreatic cancers to fatty liver disease and acute appendicitis. The release came alongside a peer-reviewed paper in Science, which is a rarer credential than most medical AI press releases can […]

DeepSeek-V4.1-Flash squeezes the KV cache down to 890 bytes per token

DeepSeek-V4.1-Flash looks like a minor version bump until you read what the team actually shipped. The new model is a 552B-parameter multimodal mixture-of-experts system that supports million-token contexts, and its whole design centers on one number: 890 bytes per token of global KV cache. That is roughly a quarter of what DeepSeek-V4-Flash needed in HBM […]

Azure Copilot’s Troubleshooting Agent hits general availability

Microsoft has taken the Troubleshooting Agent in Azure Copilot to general availability. It lives in the Azure portal, in both the Copilot chat experience and the Support + Troubleshooting blade, and its job is to take a vague complaint like “my VM rebooted itself last night” and turn it into a diagnosis, a fix, or […]

Real-SWE benchmarks AI coding agents on code that cannot leak into training data

Why another SWE benchmark Specific Labs released Real-SWE, a benchmark that runs frontier AI coding agents against private, real-world enterprise codebases. The pitch is simple and hard to argue with: public benchmarks leak. SWE-bench tasks live in public repositories, and their solutions are on the internet. Any model trained after the benchmark published has, at […]

Training a 3.8B LLM for $998: what one person with rented B200s can do

Hugo Vergnes trained a 3.8 billion parameter language model from scratch for $998, scoring 0.384 on the CORE benchmark after 43 hours on eight rented B200 GPUs. That number matters because of what sits next to it in the comparison table. Karpathy’s nanochat d32, a 1B model trained for roughly the same $1,000 budget, scores […]

Running a 2.8T parameter model from SSDs on a MacBook Pro

A 2.8 trillion parameter model just ran on a laptop, and it didn’t fit in RAM when it did. A project on Hacker News streamed Kimi K3’s weights off four SSDs at about 1 token per second on a MacBook Pro, which is slow by chatbot standards and remarkable by any other measure. What Kimi […]

OpenAI’s Chief Scientist Says No Lab Is Ready to Scale at Full Speed

OpenAI chief scientist Jakub Pachocki published an essay on September 6 called An Alien Mind, and the timing is hard to ignore. Three days earlier, the company shipped GPT-6 Astra, its most capable model yet and the first to cross a Critical threshold on cybersecurity evaluations. Days after that launch, the person who runs research […]