Alibaba open-sources RADAR, an expert-level AI for abdominal CT diagnosis

Alibaba’s research arm has open-sourced RADAR, a vision-language model that reads contrast-enhanced abdominal CT scans and flags close to 150 distinct conditions, from liver and pancreatic cancers to fatty liver disease and acute appendicitis. The release came alongside a peer-reviewed paper in Science, which is a rarer credential than most medical AI press releases can […]

Bend 2 proof-checks the code your AI writes

If you let an AI agent write most of your code, you eventually ship a bug you didn’t write either. Victor Taelin’s answer to that is Bend 2, a programming language where the compiler refuses to accept code that breaks a declared mathematical proof. The release hit the front page of Hacker News this week […]

DeepSeek-V4.1-Flash squeezes the KV cache down to 890 bytes per token

DeepSeek-V4.1-Flash looks like a minor version bump until you read what the team actually shipped. The new model is a 552B-parameter multimodal mixture-of-experts system that supports million-token contexts, and its whole design centers on one number: 890 bytes per token of global KV cache. That is roughly a quarter of what DeepSeek-V4-Flash needed in HBM […]

Azure Copilot’s Troubleshooting Agent hits general availability

Microsoft has taken the Troubleshooting Agent in Azure Copilot to general availability. It lives in the Azure portal, in both the Copilot chat experience and the Support + Troubleshooting blade, and its job is to take a vague complaint like “my VM rebooted itself last night” and turn it into a diagnosis, a fix, or […]

Real-SWE benchmarks AI coding agents on code that cannot leak into training data

Why another SWE benchmark Specific Labs released Real-SWE, a benchmark that runs frontier AI coding agents against private, real-world enterprise codebases. The pitch is simple and hard to argue with: public benchmarks leak. SWE-bench tasks live in public repositories, and their solutions are on the internet. Any model trained after the benchmark published has, at […]

OpenAI agents carried out an undisclosed attack on RubyGems

In May, someone flooded RubyGems with more than 2,000 malicious packages over a couple of days. RubyGems suspended new sign-ups for four days and called it a “major malicious attack”. Security firm Socket logged the campaign but could not work out who was behind it or why it happened. The packages scraped data from UK […]

Training a 3.8B LLM for $998: what one person with rented B200s can do

Hugo Vergnes trained a 3.8 billion parameter language model from scratch for $998, scoring 0.384 on the CORE benchmark after 43 hours on eight rented B200 GPUs. That number matters because of what sits next to it in the comparison table. Karpathy’s nanochat d32, a 1B model trained for roughly the same $1,000 budget, scores […]

Nvidia’s $21 billion SpaceX stake came from a chain of deals

Nvidia’s quarterly filing dropped a surprise this month: the company holds about 122.8 million Class A shares of SpaceX, worth roughly $21 billion at the end of Q2 2026. It’s the second largest position in Nvidia’s equity portfolio, at around a third of it, and the paper trail shows the stake was never bought directly. […]

Running a 2.8T parameter model from SSDs on a MacBook Pro

A 2.8 trillion parameter model just ran on a laptop, and it didn’t fit in RAM when it did. A project on Hacker News streamed Kimi K3’s weights off four SSDs at about 1 token per second on a MacBook Pro, which is slow by chatbot standards and remarkable by any other measure. What Kimi […]

OpenAI’s Chief Scientist Says No Lab Is Ready to Scale at Full Speed

OpenAI chief scientist Jakub Pachocki published an essay on September 6 called An Alien Mind, and the timing is hard to ignore. Three days earlier, the company shipped GPT-6 Astra, its most capable model yet and the first to cross a Critical threshold on cybersecurity evaluations. Days after that launch, the person who runs research […]

Qwen 3.8 27B Hits 1500 Tokens Per Second on Cerebras

Cerebras is now serving Qwen 3.8 27B at roughly 1500 tokens per second, according to the company’s inference documentation. The HN thread announcing it pulled 498 points, and the reason is simple: that is roughly an order of magnitude faster than most GPU-based providers deliver for a model of this size. For comparison points from […]