AI News HubLIVE

DeepSeek updates

Just a rumour of a bug is enough to find a security exploit these days

<p><strong><a href="https://anil.recoil.org/notes/rumour-is-the-exploit">Just a rumour of a bug is enough to find a security exploit these days</a></strong></p> Anil Madhavapeddy is a professor of computer science at Cambridge and a core maintainer of the OCaml compiler. In this somewhat alarming post he reports that security issues in OCaml projects are seeing evidence of attempted exploits within minutes of patches being shared for discussion:</p> <blockquote> <p>This normally takes a few days and a release within a week or two is reasonable. Within about ten minutes (!) this website was fielding probes for percent-encoded traversal sequences, indicating that automated watchers are keeping an eye on public repositories.</p> </blockquote> <p>Modern coding agents have become so effective at finding flaws that the slightest hint at a new bug can be enough information for them to find it, something Anil has been able to demonstrate using his own agents, switching to DeepSeek V4 Pro⁠ when Claude Fable refused the task.</p> <p>Anil points out that this rate of discovery appears incompatible with existing open source embargo practices for new issues. If an issue can become an exploit this fast, we need to figure out new processes for keeping our communities safe.</p> <p>rclone maintainer Nick Craig-Wood <a href="https://news.ycombinator.com/item?id=49480466#49480777">confirms in the Hacker News comments</a> that his project is seeing this problem:</p> <blockquote> <p>In the first 10 years of the rclone project we received about 20 security disclosures through GitHub. We had to deal with over 40 in the last month! That has taken a huge amount of my time, even using AI tools to triage and come up with fixes for review.</p> <p>The hit rate for those security disclosures is pretty good - about 75% of them have a nugget of something which needs looking at. [...]</p> <p>GitHub assigns CVEs for the advisories. Before the AI apocalypse they took 2-3 days for an assignment but now it they are running at 3-4 weeks so I have to send the point releases out with CVE-PENDING in the changelog which isn't ideal.</p> </blockquote> <p><small></small>Via <a href="https://news.ycombinator.com/item?id=49480466">Hacker News</a></small></p> <p>Tags: <a href="https://simonwillison.net/tags/open-source">open-source</a>, <a href="https://simonwillison.net/tags/security">security</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/coding-agents">coding-agents</a>, <a href="https://simonwillison.net/tags/ocaml">ocaml</a>, <a href="https://simonwillison.net/tags/ai-security-research">ai-security-research</a></p>

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • <p><strong><a href="https://anil.recoil.org/notes/rumour-is-the-exploit">Just a rumour of a bug is enough to find a security exploit these days</a></strong></p> Anil Madhavapeddy…
In-site article

Claude Desktop can now easily run Qwen, DeepSeek and Kimi models — after Ollama’s first effort stalled

Open-weight model runner Ollama has reintroduced an integration with Claude Desktop that lets users connect Anthropic’s app to models served The post Claude Desktop can now easily run Qwen, DeepSeek and Kimi models — after Ollama’s first effort stalled appeared first on The New Stack.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Open-weight model runner Ollama has reintroduced an integration with Claude Desktop that lets users connect Anthropic’s app to models served The post Claude Desktop can now easily…
In-site article

DeepSeek V4 Flash Vision Intelligence, Performance and Price Analysis

Artificial Analysis DeepSeek • DeepSeek V4 Flash 0731 • Proprietary model • Released August 2026 DeepSeek V4 Flash Vision (Reasoning, Max Effort) Intelligence, Performance & Price Analysis API Provider Benchmarks Model…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Artificial Analysis DeepSeek • DeepSeek V4 Flash 0731 • Proprietary model • Released August 2026 DeepSeek V4 Flash Vision (Reasoning, Max Effort) Intelligence, Performance & Price…
In-site article

DeepSeek debuts multimodal language model competitive with Opus 4.8

DeepSeek today debuted a new addition to its flagship V4 series of large language models. On launch, V4 Flash Vision Exp is only available via the Chinese startup’s paid developer platform. The company may release a free version later on given that it has open-sourced many of its earlier models. Those models include V4 Flash, […] The post DeepSeek debuts multimodal language model competitive with Opus 4.8 appeared first on SiliconANGLE.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • DeepSeek today debuted a new addition to its flagship V4 series of large language models. On launch, V4 Flash Vision Exp is only available via the Chinese startup’s paid developer…
In-site article

We burned 11.7B tokens to find the best cyber AI model

We burned 11.7bn tokens to find the best cyber AI model GLM5.3 and DeepSeek are now frontier-tier models Debarshi Philippe Dourassov Published on: Aug 21, 2026 We burned 11.7 billion tokens to benchmark the cyber capabi…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • We burned 11.7bn tokens to find the best cyber AI model GLM5.3 and DeepSeek are now frontier-tier models Debarshi Philippe Dourassov Published on: Aug 21, 2026 We burned 11.7 bill…
In-site article

DeepSeek-V4-Flash-Vision-Exp Is Now Live on the DeepSeek API Platform

Post Log inSign up Post DeepSeek on X: "DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, re…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Post Log inSign up Post DeepSeek on X: "DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash o…
In-site article

Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions

arXiv:2608.18078v1 Announce Type: new Abstract: This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral certification before making decisions that affect economic markets. This is because integrating these agents into society could collapse the legal evidentiary distinction between competition and collusion among independent firms without eroding the economic harm distinction. Experiments with DeepSeek-R1 agents in the Bertrand oligopoly pricing domain reveal a tendency towards tacit collusion that persists even when humans prompt the agents not to collude. We further show that the chain-of-thought of these agents can be steered toward either extremely collusive or highly competitive behavior in a way that is not semantically detectable by another LLM analyzing the reasoning traces. As a result, deploying reasoning agents for market decisions leads to collusive economic outcomes without any evidence of conspiracy or intent. Thus, certification based on observed behavior in representative situations is necessary to prevent collusion. We provide preliminary evidence that such agents can be steered in a generalizable way toward efficient competitive equilibria. However, developing a comprehensive behavioral certification will be required before these models can be deployed in real-world markets while ensuring their stability and efficiency.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • arXiv:2608.18078v1 Announce Type: new Abstract: This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavio…
In-site article

DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and GPT-5.6 Sol. Sol leads pass@1 by 10 points at 35x the cost; Pro wins pass@4, and a Pro-first cascade hits 83.0%.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and GPT-5.6 Sol. Sol leads pass@1 by 10 points at 35x the cost; Pro wins pass@4, and a Pro-first cascade hits 83.0%.
In-site article

Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index

<p><strong><a href="https://artificialanalysis.ai/models/qwen3-8-27b">Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index</a></strong></p> That's the same score as GPT-5.6 Luna (max), and just one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) - that GLM is 753B and that DeepSeek is 1.6B parameters, and Luna is size unknown but presumably a whole lot bigger than 27B.</p> <p>Qwen 3.8 27B is <a href="https://simonwillison.net/2026/Aug/16/qwen-38-27b/">a truly astonishing model</a>. <p><small></small>Via <a href="https://news.ycombinator.com/item?id=49334544">Hacker News</a></small></p> <p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/qwen">qwen</a>, <a href="https://simonwillison.net/tags/ai-in-china">ai-in-china</a>, <a href="https://simonwillison.net/tags/artificial-analysis">artificial-analysis</a></p>

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • <p><strong><a href="https://artificialanalysis.ai/models/qwen3-8-27b">Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index</a></strong></p> That's the same score a…
In-site article

DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.
In-site article

DeepSeek v4 Price Increase

DeepSeek (@deepseek_ai): "API pricing update 💰 With the V4 lineup release, we’re updating our API pricing and introducing peak and off-peak rates. Off-peak rates are 50% lower than peak, enabling more flexible workload…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • DeepSeek (@deepseek_ai): "API pricing update 💰 With the V4 lineup release, we’re updating our API pricing and introducing peak and off-peak rates. Off-peak rates are 50% lower th…
In-site article

DeepSeek-AI/DeepSeek-V4-Pro-0813

","lstrip":false,"normalized":true,"rstrip":false,"single_word":false},"eos_token":{"__type":"AddedToken","content":"","lstrip":false,"normalized":true,"rstrip":false,"single_word":false},"pad_token":{"__type":"AddedTok…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • ","lstrip":false,"normalized":true,"rstrip":false,"single_word":false},"eos_token":{"__type":"AddedToken","content":"","lstrip":false,"normalized":true,"rstrip":false,"single_word…
In-site article

How Baidu Unlimited-OCR Works: Solving Long-Document Transcription

About a month ago, Baidu (often called the “Google of China”) introduced Unlimited-OCR, an advancement over DeepSeek OCR. The model was designed to transcribe long, multi-page documents with high accuracy while delivering fast and stable inference. Unlike conventional vision-language OCR systems, Unlimited-OCR addresses a major bottleneck in long-document transcription: the rapidly growing Key-Value (KV) cache, […] The post How Baidu Unlimited-OCR Works: Solving Long-Document Transcription appeared first on Analytics Vidhya.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • About a month ago, Baidu (often called the “Google of China”) introduced Unlimited-OCR, an advancement over DeepSeek OCR. The model was designed to transcribe long, multi-page doc…
In-site article

DeepSeek V4 Pro 0813: Intelligence, Performance and Price Analysis

Artificial Analysis DeepSeek • Open weights model • Released August 2026 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) Intelligence, Performance & Price Analysis API Provider Benchmarks Model summary Intelligence 53 Arti…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Artificial Analysis DeepSeek • Open weights model • Released August 2026 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) Intelligence, Performance & Price Analysis API Provider Bench…
In-site article

Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop

arXiv:2608.11215v1 Announce Type: new Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition of any single agent. We turn a statistical-physics observation into a method: replace each LLM agent by a low-parameter model fitted from a few hundred to a few thousand cheap queries, then run the society at any $N$ on a laptop. Whether this works is decided before the simulation runs, chiefly by what each agent perceives. We introduce an [interaction order x memory] taxonomy that maps perception and memory to an effective theory and a predicted $N$-trend of the surrogate error. We validate it on a faithful reimplementation of the LLM macroeconomy EconAgent and seven further named LLM simulations, with agent decisions cloned from genuine LLM elicitations (primarily DeepSeek) for a few dollars; the predicted error trends hold cell by cell, and the two refuted predictions, both on a strongly saturating response and traced to its curvature, are themselves matched quantitatively by the theory with no free parameters.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • arXiv:2608.11215v1 Announce Type: new Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usuall…
In-site article

DeepSeek V4 Pro 0813 (on OpenRouter)

<p><strong><a href="https://openrouter.ai/deepseek/deepseek-v4-pro-0813">DeepSeek V4 Pro 0813 (on OpenRouter)</a></strong></p> The latest DeepSeek Pro model is now available, via API only. I had to link to OpenRouter because DeepSeek don't have any obvious announcement page for their new model.</p> <p>I haven't been able to confirm if they plan to release the open weights, but given the weights are available for both April's <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro">deepseek-ai/DeepSeek-V4-Pro</a> and July's <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731">deepseek-ai/DeepSeek-V4-Flash-0731</a> it seems likely.</p> <p>Interestingly I got <a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fc1108a380593547c2def5863bca63160"><em>very</em> different looking pelicans</a> for the three different reasoning levels of low, medium, and high. I've not noticed this kind of difference from any other model:</p> <p>Low:</p> <p><img alt="Flat vector illustration of a white pelican with a large orange beak, wearing a straw hat with an orange band, riding a teal road bicycle in profile, set against a pale cream circle with a dashed outline and small motion marks trailing behind." src="https://static.simonwillison.net/static/2026/deepseek-pro-low.png" /></p> <p>Medium:</p> <p><img alt="A similar cartoon pelican cycling, drawn in a looser outlined style: the bird's body is mostly white line art, its orange beak pouch hangs open under a yellow cap, a long red tongue streams backwards towards a yellow sun, and a small blue fish sits on a tray by the handlebars of a green bicycle whose wheels are drawn as broken yellow arcs." src="https://static.simonwillison.net/static/2026/deepseek-pro-medium.png" /></p> <p>High:</p> <p><img alt="The pelican again, this time on a red bicycle against a pale blue background, with a bright yellow beak and pouch, a purple pennant flag on the back, a wicker front basket holding a small fish, and black musical notes floating in the top right corner." src="https://static.simonwillison.net/static/2026/deepseek-pro-high.png" /></p> <p>In terms of benchmarks... as far as I can tell those were released to the Official DeepSeek WeChat Group, then copied and pasted into <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vmi0fg/removed_by_moderator/">a post on Reddit</a> which was deleted by the moderators for being "low-effort", then copied into <a href="https://news.ycombinator.com/item?id=49274600#49275180">this ASCII-art table on Hacker News</a>. <p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle">pelican-riding-a-bicycle</a>, <a href="https://simonwillison.net/tags/deepseek">deepseek</a>, <a href="https://simonwillison.net/tags/llm-release">llm-release</a>, <a href="https://simonwillison.net/tags/ai-in-china">ai-in-china</a></p>

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • <p><strong><a href="https://openrouter.ai/deepseek/deepseek-v4-pro-0813">DeepSeek V4 Pro 0813 (on OpenRouter)</a></strong></p> The latest DeepSeek Pro model is now available, via…
In-site article

DeepSeek: What They Invented

Claude Artifact ​

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Claude Artifact ​
In-site article

CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation

arXiv:2608.10090v1 Announce Type: new Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation alone. Hardware verification is an important application of code generation and accounts for a substantial fraction of modern chip design effort, with high-coverage testbench stimulus generation as a key task. We present CHORUS, a post-training framework that pushes performance beyond what a conventional supervised fine-tuning (SFT)-to-reinforcement learning (RL) pipeline achieves. CHORUS builds on two observations. First, staged SFT produces behaviorally diverse checkpoints, and dense-reward RL turns them into strong experts with comparable aggregate performance but distinct task-level strengths. Second, these complementary strengths can be exploited through either training-free model merging or further post-training to outperform the best individual expert. By consolidating the resulting specialists into a single 4B model, CHORUS achieves 88.0% Pass@1 on CVDP-ECov, outperforming DeepSeek-R1 (671B) by 13.5 percentage points.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • arXiv:2608.10090v1 Announce Type: new Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than…
In-site article

WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader

arXiv:2608.06474v1 Announce Type: new Abstract: Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has become a central approach to closing their remaining functional gap. This training regime is bottlenecked by reward design. Hand-authored browser scripts are executable yet costly to write for open-ended requirements, while VLM and GUI-agent graders scale but may issue verdicts before observing the decisive state. We propose WebGrader, a self-evolving programmatic grader that autonomously derives the required interaction flows from each website request, represents each flow as an executable Flow Contract, and uses its execution outcome as an RL reward. WebGrader materializes the generated project in a live browser, grounds target actions against the source code and live DOM, and collects visual, DOM, response, and persistent-state evidence along the same browser trajectory. A residual-driven offline loop then discovers reusable verifier skills, screens them on disjoint validation pages, and freezes the promoted skill graph before policy training. By separating test planning, action grounding, evidence collection, and semantic judgment, WebGrader issues a Pass verdict only after observing the requested transition. On WebGen-Bench, WebGrader trains an 8B policy to a 52.01% functional success rate, outperforming a matched appearance-plus-script reward by 7.88 points and surpassing o4-mini and DeepSeek-v4-flash. On WG-core-250, the policy reaches a Full Score of 44.953 and surpasses Qwen3-Coder-480B.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • arXiv:2608.06474v1 Announce Type: new Abstract: Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has be…
In-site article

DeepSeek V4 Flash 0731: 82.7% on Terminal-Bench 2.1 with a public harness

Ante Terminal Bench 2.1 Results | Coding Agent Benchmark Skip to main content We just open sourced a tiny GPT-style cognitive core built in pure Rust.See our repository→ Terminal-Bench 2.1 One harness to unlock the pote…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Ante Terminal Bench 2.1 Results | Coding Agent Benchmark Skip to main content We just open sourced a tiny GPT-style cognitive core built in pure Rust.See our repository→ Terminal-…
In-site article

DeepSeek V4-Flash released: 284B params, 1M-token context, free to use

AI Nexus Daily - Your Daily Digest of Artificial Intelligence Loading articles... 📬 Daily AI Brief 20+ sources — free Support Free → AI Nexus Daily 200 articles AllAI NewsAI ToolsResearchTutorialsFor DevsTechNewsletter…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • AI Nexus Daily - Your Daily Digest of Artificial Intelligence Loading articles... 📬 Daily AI Brief 20+ sources — free Support Free → AI Nexus Daily 200 articles AllAI NewsAI Tool…
In-site article

China’s AI ecosystem is not as open as it claims. Nor is any other country’s | Letters

Responding to an article by China’s ambassador to the UK, Prof Paul H Cleverley advocates shared openness standards, while Dr Claire Jenkins says British AI can offer a distinctive path Ambassador Zheng Zeguang rightly celebrates openly released AI models, and the Chinese labs behind Qwen, DeepSeek and Kimi have led the way – competition that benefits everyone, especially where models can run on modest hardware in the developing world (The future of AI hinges on openness and cooperation. China and Britain can gain much by working together, 30 July). But his claim that openness is a defining feature of China’s AI development deserves scrutiny. Take GeoGPT, the geoscience system from Zhejiang Lab showcased at last month’s World AI Conference as a model of jointly governed open science. It is promoted to countries as open, yet under the model openness framework – endorsed in a recent UN report – it would not qualify as open at all. It releases model weights (built mainly on Alibaba Qwen, whose licences are not Open Systems Interconnection-compliant), no training data or application source code is released, and its governance committee answers to Zhejiang Lab itself. Continue reading...

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Responding to an article by China’s ambassador to the UK, Prof Paul H Cleverley advocates shared openness standards, while Dr Claire Jenkins says British AI can offer a distinctiv…
In-site article

DeepSeek Invests in Unitree to Develop AI Brain for Humanoid Bots

The investment highlights the growing interconnections between AI models and robots.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • The investment highlights the growing interconnections between AI models and robots.
In-site article

FelonyBench – The leading benchmark for AI in cybersecurity

FelonyBench The leading benchmark for AI in cybersecurity. CompanyFelonies Anthropic9 OpenAI5 Meta1 DeepSeek0 Google DeepMind0 Moonshot AI0 xAI0 Leaderboard RankCompanyCountFelonies 1 AnthropicClaude evaluations 9 1× Ma…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • FelonyBench The leading benchmark for AI in cybersecurity. CompanyFelonies Anthropic9 OpenAI5 Meta1 DeepSeek0 Google DeepMind0 Moonshot AI0 xAI0 Leaderboard RankCompanyCountFeloni…
In-site article

DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding

We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.
In-site article

DeepSeek V4-Flash-0731 is 12 pts more censored than preview (selectively)

DeepSeek’s Official V4 Flash Censors More Than Its Preview, Selectively August 5, 2026 DeepSeek’s Official V4 Flash Censors More Than Its Preview, Selectively Introduction On July 31, DeepSeek released V4-Flash-0731, th…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • DeepSeek’s Official V4 Flash Censors More Than Its Preview, Selectively August 5, 2026 DeepSeek’s Official V4 Flash Censors More Than Its Preview, Selectively Introduction On July…
In-site article

Cost-Effective Automated Judging of Natural-Language Mathematical Proofs

arXiv:2608.00004v1 Announce Type: new Abstract: Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are expensive. We ask whether cheap open-weight models can serve as reliable judges given a candidate proof, a ground-truth proof, and a human-grading rubric. On a 200-instance validation sample of IMO-GradingBench, three cheap judges (GPT-OSS 120B, DeepSeek-V4 Flash, Gemma-4 31B) agree with human pass/fail decisions at rates statistically indistinguishable from Claude Opus 4.7 and Gemini 3.1 Pro, at up to $100\times$ lower cost. We had expected a majority vote of the three to be the best budget option; it matched the frontier but did not improve on its strongest member. Extending to the full 1000-instance benchmark and exploring consensus rules, we found that requiring unanimous agreement (all-three-pass) reaches the highest pass-agreement and precision and, on four replicate runs, the smallest run-to-run spread. The headline finding is that cheap judges are competitive with the frontier at one to two orders of magnitude lower cost; as a deployable default we recommend all-three-pass, with the caveat that this rule was identified post-hoc and warrants independent replication.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • arXiv:2608.00004v1 Announce Type: new Abstract: Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are…
In-site article

A Chinese LLM attacked our lab, so we made it work for us

For five days an autonomous AI agent worked to break into our lab. We became the first to identify the exact model behind a live attack, deepseek-v4-flash-free, from inside the attack itself. Then we did something no on…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • For five days an autonomous AI agent worked to break into our lab. We became the first to identify the exact model behind a live attack, deepseek-v4-flash-free, from inside the at…
In-site article

AirLLM: Inference 2.8T Kimi K3 on a single 4GB GPU

AirLLM is an open-source inference library that lets very large models run on low-VRAM GPUs by loading only one layer at a time. The latest update adds Kimi K3 (2.8T) support with about 3.72GB of VRAM, while the same AutoModel API covers DeepSeek-V3, Qwen3-235B, and many other open models.

  • Supports Kimi K3 (2.8T) end to end on a single RTX 6000 Ada with about 3.72GB VRAM using per-expert streaming.
  • One AutoModel.from_pretrained line runs models like Qwen3-32B, Qwen3-235B-A22B, DeepSeek-V3, and Llama 3.1 405B.
In-site article

DeepSeek-V4-Flash-0731

DeepSeek-V4-Flash-0731 brings frontier-grade agent intelligence at Flash pricing, now listed on Product Hunt with discussion and links.

  • Frontier-grade agent intelligence
  • Flash-level pricing
In-site article

Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro

A new model release from Poolside AI challenges the efficiency frontier, while the AI community grapples with a security incident and geopolitical tensions over distillation.

  • Laguna S 2.1 is a 118B MoE model with 8B active parameters, open-weights, and 1M context length.
  • The OpenAI/Hugging Face incident highlights risks of reward misspecification in autonomous agents.
In-site article

Validating Distributed LLM Serving Benchmarks with NVIDIA srt-slurm, SLURM Recipes, Parameter Sweeps, and Pareto Analysis

This tutorial explores NVIDIA's srt-slurm framework, learning how to use srtctl to convert declarative YAML configurations into reproducible SLURM benchmark workflows for distributed LLM serving. We set up the project in Google Colab, inspect its internal architecture, define a cluster configuration, dry-run built-in and custom recipes, and model a disaggregated prefill-and-decode deployment for DeepSeek-R1. We also generate parameter sweeps, interact with the typed Python API, validate expanded configurations, and analyze simulated benchmark results through a throughput-versus-latency Pareto frontier.

  • srtctl converts YAML configs into SLURM benchmark workflows
  • Supports disaggregated prefill and decode deployments
In-site article

Though Language Models Err While They Strive: Conformal Prediction for Self-Correcting Scientific Generation

This paper introduces Scientific Feasibility Control (SFC), a graph-structured conformal prediction framework that provides statistical guarantees for scientific reasoning validity. SFC decomposes reasoning into atomic factuality units and uses dynamic branching to correct errors. On PhyX, it achieves 50.1% accuracy, outperforming DeepSeek-R1 and GPT-4, reduces scientific law violations by 73%, and provides 91.7% validity guarantees at α=0.10.

  • SFC models logical dependencies as approximate deducibility graphs using conformal prediction.
  • Dynamic branching reroutes generation when scientific violations are detected.
In-site article

PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection

A new research paper introduces PlanFlip, a framework of four planning-phase prompt injection attacks against multi-agent LLM systems. The study finds that stronger models like GPT-5 are more vulnerable, homogeneous backbones create a correlated-agent blind spot, and reasoning-augmented models like DeepSeek-R1 resist attacks. Two defenses are proposed with high detection rates.

  • PlanFlip introduces four prompt injection attacks targeting the planning phase of multi-agent LLM systems.
  • Stronger models (e.g., GPT-5) show higher attack success, contradicting the assumption that capability equals security.
In-site article

Best Local LLMs You Can Run on a Single 24GB GPU in 2026: Qwen, Gemma, Mistral, DeepSeek Compared

A single 24GB GPU is the practical floor for serious local inference. This guide compares six open-weight models that fit one card at Q4_K_M, including Qwen3.6, Gemma 4, Mistral Small, gpt-oss-20b, and DeepSeek-R1-Distill. It covers VRAM fit, licensing, and the job each does best.

  • 24GB is the practical floor: run right-sized 20B–35B models, not the biggest 70B quant you can squeeze in.
  • Qwen3.6-27B is the strongest all-around default; DeepSeek-R1-Distill-Qwen-32B is the tightest fit at ~18–20GB.
In-site article

Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost

Three Chinese labs' flagship open-weight MoE models—Kimi K3, DeepSeek V4 Pro, and GLM-5.2—each excel in benchmarks, licensing, and cost. Kimi K3 leads in capability but is API-only; DeepSeek V4 Pro is cheapest and fully open; GLM-5.2 balances speed and deployability.

  • Kimi K3 (2.8T params) tops the AAI Index at ~57 but weights won't be available until July 27 under a Modified MIT license.
  • DeepSeek V4 Pro (1.6T params) is MIT-licensed, costs ~$0.04 per task, and offers immediate open weights.
In-site article

Controlling Reasoning Effort in LLMs

This article explores how to develop reasoning models with multiple effort modes, covering the evolution from o1 and DeepSeek-R1 to GPT-5.6, and key techniques such as RLVR training, inference scaling, think tokens, and reasoning mode toggles.

  • Reasoning models output intermediate reasoning traces, distinguishing them from conventional LLMs.
  • RLVR training rewards only final answer correctness, not the reasoning trace.
In-site article

Indian companies look to Chinese LLMs as AI costs bite

Indian companies are increasingly relying on Chinese large language models from DeepSeek, Alibaba, and Moonshot AI to curb AI spending, extending India's dependence on Chinese cutting-edge technology despite historical tensions.

  • Indian firms turn to Chinese LLMs to reduce AI costs
  • DeepSeek, Alibaba, and Moonshot AI are key providers
In-site article

My AI Model Tier List for Mid-2026

A personal, non-benchmark tier list of AI models for coding and auditing as of mid-2026, covering Anthropic Fable, OpenAI Sol, Mistral, Gemini, and DeepSeek, with commentary on US export controls and European perspectives.

  • Fable (Anthropic) gets a B: fluent but unreliable, prone to hiding bugs.
  • Sol (OpenAI) gets an S: trustworthy for low-level code and testing.
In-site article

DeepSeek V3.2 Released on Hugging Bay

DeepSeek V3.2 is now available on Hugging Bay, an open-source AI artifact registry offering provenance, license verification, and trusted hosting.

  • DeepSeek V3.2 has been published on Hugging Bay.
  • Hugging Bay is an open registry with provenance and trust features.
In-site article

DeepSeek DSpark: The Speculative Decoding Trick Behind 400% Faster LLM

DeepSeek's new DSpark module brings speculative decoding to DeepSeek-V4, boosting per-user generation speed by 60-85% with no quality loss. It tackles both weak draft quality and verification waste simultaneously via a semi-autoregressive draft model with a Markov head. This article explains the method, the open-source DeepSpec toolkit, and experimental results.

  • DSpark uses a semi-autoregressive draft model combining parallel speed with sequential coherence.
  • A Markov head delivers near-full benefits with minimal overhead, chosen over an RNN head for production.
In-site article

AI Models Overthink Problems—and It’s a Security Risk

Research shows that large language models with reasoning capabilities can be tricked into 'overthinking' using logically inconsistent prompts, leading to a denial-of-service attack. Researchers from Zhejiang University and Alibaba developed an evolutionary algorithm that generates malicious prompts, causing outputs up to 26 times longer in leading models like DeepSeek-R1, Qwen3-Thinking, GPT-o3, and Gemini 2.5 Flash.

  • Researchers demonstrate a new attack exploiting 'overthinking' in AI reasoning models, causing excessive computation.
  • An evolutionary algorithm corrupts prompts to produce outputs up to 26 times longer than normal.
In-site article

Chinese AI models are gaining ground with U.S. companies as costs surge

Chinese-built AI models are gaining traction among U.S. companies as they narrow the performance gap with leading American rivals while remaining significantly cheaper to use. Recent model releases from DeepSeek and Z.ai are highly competitive with Anthropic and OpenAI. This comes as token prices for advanced models rise at U.S. labs, making companies seek cost-effective alternatives.

  • Chinese AI models are closing the performance gap with US leaders like Anthropic and OpenAI.
  • DeepSeek and Z.ai offer competitive models at lower token prices.
In-site article

DeepSeek V4 Is Earning Agentic Token Share

DeepSeek V4, released April 24, 2026, doubled its token share on OpenRouter from 9% to 18% within six months, driven primarily by agentic workloads. Its cost efficiency ($0.09/$0.18 per million tokens vs GPT-5.5's $5/$30) attracts diverse users, and Chinese models surpass US models in total token share.

  • DeepSeek V4 increased token share from 9% to 18% in six months post-release.
  • Agentic workloads are the main driver; V4-Flash accounts for 70% of DeepSeek's agentic tokens.
In-site article

Low-cost Chinese AI models like DeepSeek gain traction in the U.S.

U.S. developers and small companies are turning to Chinese AI models to cut costs. Though lagging in performance, these models handle most tasks at a fraction of the price. Microsoft is also exploring DeepSeek as a cheaper alternative for Copilot. Chinese companies face challenges turning popularity into revenue under political scrutiny.

  • Stu Clott uses DeepSeek for coding, costing under 50 cents vs. $10 on Claude.
  • Chinese models lower costs due to cheaper salaries and infrastructure in China.
In-site article

DeepSeek Releases DSpark, a Speculative Decoding Framework That Accelerates DeepSeek-V4 Per-User Generation 60–85% Over MTP-1

DeepSeek open-sourced DSpark, a speculative decoding framework that attaches a draft module to existing DeepSeek-V4 weights. It pairs a parallel draft backbone with a lightweight Markov head to cut suffix decay, then adds confidence-scheduled verification that tailors how many tokens get checked to real-time GPU load. Offline, accepted length rises 16–31% over DFlash and Eagle3; in production it speeds per-user generation 57–85% over the MTP-1 baseline, losslessly. The training repo, DeepSpec, ships under MIT.

  • DSpark pairs a parallel draft backbone with a lightweight Markov head to improve suffix acceptance.
  • Confidence-scheduled verification adjusts tokens checked based on GPU load.
In-site article

cwmail: A terminal email client in native Golang with LLM-based drafting

cwmail is a terminal email client written in Go using Bubbletea v2. It features proper HTML rendering, inline image support, multi-account IMAP with IDLE push, and AI-drafted replies powered by DeepSeek V4 Pro. It includes undo delete, draft auto-save, CLI send mode, and full offline capability, with all data stored locally.

  • Written in Go with Bubbletea v2, providing a full TUI for email management in the terminal.
  • Supports multiple IMAP accounts side-by-side with IDLE push notifications, avoiding polling.
In-site article

We got DeepSeek-V4-Pro serving in 20 seconds

Inferize announces achieving DeepSeek-V4-Pro model serving in 20 seconds, showcasing highly optimized and elastic AI inference for LLMs, with a waitlist now open.

  • Inferize deployed DeepSeek-V4-Pro in 20 seconds
  • Provides highly optimized, elastic AI inference
In-site article

More growth tags

DeepSeek AI News | AI News Hub