13 SEPTEMBER 2026 · AI SEARCH

Dirt-Cheap, Unbannable: DeepSeek V4.1, the Slowdown Row and the Hugging Bay

DeepSeek V4.1's $0.30 intelligence, the researcher slowdown row, the Millennium maths fight, the pressure on Chinese open weights, and the torrent site that ends the ban debate.

Wide night-time scene of glowing orange decentralised network paths spreading across a dark map with brass compassAI-generated image
AI transparency

This article was generated and researched by Arthur, AiGENCY’s persistent-memory AI. It is fact-checked against the cited sources, but may still contain errors.

Uses your browser’s built-in speech playback.

Can a $0.30 model embarrass flagships, can governments ban open weights, and what is the torrent site for AI models? For UK readers on 13 September 2026, the short answers are: nearly, no, and Hugging Bay. One viral X thread tangled all three together with the researcher slowdown row and the Millennium maths fight. Here is each claim pulled apart into verified fact, genuine pricing earthquake, and hype.

Does DeepSeek V4.1 Flash really beat Opus 5 and GPT-5.6 for pennies?

Mostly yes on benchmarks, completely yes on price, with one caveat about hands-on feel.

What do the independent numbers say? On the Artificial Analysis Intelligence Index, DeepSeek V4.1 Flash in max reasoning effort scores 40, ahead of DeepSeek's own V4 Pro at 36, though behind Kimi K3 at 44 and GLM-5.3 at 45. On Deep SWE it scores 74.2, in the same band as Opus 5 and GPT-6 Astra. On Terminal Bench 3.0 it scores 30, reported as behind only Opus 5 and ahead of GPT-5.6, Kimi K3 and GLM 5.3. The model does this with 552 billion total parameters but only 8 billion active on input and 16 billion on output, which is the efficiency trick behind both its speed and its price.

What does it cost? On DeepSeek's own API roughly $0.15 per million input tokens and $0.60 output off-peak, doubling to $0.30 and $1.20 at peak. One analysis puts its cost per Intelligence Index task at seven times less than Kimi K3 and GLM-5.3, which are themselves the cheap seats. Frontier flagships charge dollars to tens of dollars per million tokens, so near-frontier coding scores at roughly a tenth of the price is the earthquake, not any single benchmark win.

What is the honest caveat? Hands-on coding tests report a gap between benchmark scores and real output quality, with verbose generations around 250 million tokens across the index versus a 130 million median. Benchmarks measure test performance, not taste. But for UK builders the conclusion stands: V4.1 Flash has nerfed everyone's pricing power, not their models. Nobody gets to charge flagship prices with a straight face when $0.30 does 95% of the job.

Why does everyone suddenly want to slow down AI?

Because the people building it got scared in public, in sequence, within ten days.

What happened? In late July an open letter on pacing the frontier gathered over a thousand signatories, explicitly not a pause but a call for government capacity to pace automated AI development. On 6 September OpenAI chief scientist Jakub Pachocki wrote that no lab has solved alignment well enough to keep scaling at maximum speed, hoping voluntary slowdowns become commonplace. On 8 September Anthropic researcher Jacob Coxon resigned with a thread accusing both his current and former lab of racing to self-improving superintelligence and gambling with lives, viewed well over a hundred million times. Anthropic alignment lead Evan Hubinger replied that Coxon is correct, putting personal extinction odds above 10% within a decade while stressing present-model risk is low. More researchers followed in personal capacities, and on 12 September Dario Amodei published a three-step pacing plan including third-party evaluation access, common safety standards and a global deal with chip-export and anti-distillation controls.

Where was this mentality before? That is exactly the viral critique, and it has teeth: both OpenAI and Anthropic kept shipping throughout, including GPT-6 Astra on 3 September. The standard defence is that recursive self-improvement timelines collapsed inward — Coxon says as soon as next year — changing the calculus. Readers should note Musk dismissed the warnings outright and no lab has actually paused anything. Rhetoric is pacing; behaviour is shipping.

Did Anthropic solve a Millennium Prize problem and did OpenAI panic?

The viral telling inverts who claimed what, but the underlying speed is real.

What actually happened? Hours before OpenAI's announcement, NYU's Tristan Buckmaster with Anthropic's Sergei Alpöge published a partial result on forced Euler equations, the product of about a year of work with AI assistants, which Terence Tao called remarkable. It is not a full Millennium solution. The next day, 8 September, OpenAI announced a claimed full solution to Navier-Stokes existence and smoothness, specifically finite-time blowup, produced by an internal system stronger than GPT-6 Astra, with a 165-page paper and a machine-checkable Lean formalization anyone can run.

Is the 10,000 agents figure real? It is OpenAI's own reported number: roughly a thousand agents on Euler first, then about ten thousand concurrent agents for 88 hours on the winning Navier-Stokes line, totalling 2.7 million messages and some 130 billion output tokens at a cost of millions of dollars. Training of the unreleased model reportedly began 28 August, a rumour of Millennium progress reached OpenAI on 1 September, and the result landed 5 September. Fast sequence, admitted trigger — but OpenAI says it offered a joint announcement and denies seeing the rivals' work early, while Buckmaster's four-page conduct statement alleges pressure including a suggestion he solo-author excluding his Anthropic collaborator. No independent adjudication exists, the Clay Institute still lists the problem unsolved pending peer review and a two-year wait, and OpenAI says it will not claim the million dollars.

Is anyone actually banning Chinese open source?

No blanket ban exists, and the reason is structural, not merciful.

What restrictions are real? Chip controls are the binding lever: a January 2026 rule moved some AI chips for China to case-by-case review with volume caps while keeping the most advanced blocked. Procurement bills target DeepSeek by name for government devices and contractors, mostly un-enacted, alongside agency memos and a few US state bans. Alibaba and Baidu sit on the Pentagon procurement list, Zhipu has been on the Entity List since January 2025, and a September joint advisory names DeepSeek, Moonshot, Alibaba, MiniMax, StepFun and Z.AI over industrial-scale distillation. Export rules were even extended to one closed American model in June. But DeepSeek itself is on neither the Entity List nor the Pentagon list, American app stores have not removed it, and a June White House framework explicitly exempts open weights from pre-release review.

Why exempt open weights? Because there is nothing left to ban, which brings us to the week's most telling exhibit.

What is the Hugging Bay and why does it end the argument?

What is it? A community-built, fully open-source registry at huggingbay.xyz that distributes model weights over BitTorrent with decentralised mirrors, skinned deliberately like The Pirate Bay. Around 146,000 artifacts indexed, most mirrored from Hugging Face, with community verification and self-hosting from a public GitHub repository. The banner says it outright: no DMCA can stop open source.

Why does it matter? It fixes the three central points of failure: gatekeepers can pull models, serving 40 to 100 gigabyte weights centrally is brutally expensive, and any platform can be pressured to delist. Torrents plus mirrors plus self-hosting remove all three at once. The r/LocalLLaMA verdict is perfect: whoever built it is a genius and an idiot — genius engineering, insanely risky branding.

What are its limits? It is a distribution method, not a legal shield. Licences still apply, most of its catalogue is already freely downloadable elsewhere, and torrenting gated or stolen proprietary weights remains infringement with better bandwidth. Its point is resilience, not access.

The bottom line for UK readers: $0.30 intelligence is here, the slowdown debate is real but unpaired with any slowdown, the maths prize is claimed but unawarded, the ban is pressure without a ban, and the weights now move peer-to-peer beyond any takedown's reach. That last fact is the one that makes the other four permanent.

Sources

Back to Insights