AI News HubLIVE

来源分布

  • Hacker News AI21
  • arXiv AI5
  • Simon Willison's Weblog3
  • TheSequence3
  • Together AI Blog3
  • arXiv Computational Linguistics2
  • Last Week in AI2
  • MarkTechPost2

主题分布

  • Agent34
  • 模型34
  • 研究27
  • 芯片15
  • 政策6
  • 创业融资1
  • 工具1
  • 机器人1

日期线

  • 2026-08-136
  • 2026-07-215
  • 2026-08-023
  • 2026-08-063
  • 2026-08-183
  • 2026-07-232
  • 2026-07-312
  • 2026-08-042

最新动态

待翻译:Kraftapp AI – Describe it. We build it. Customers find it

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Describe it.We build it.Customers find it. Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ The pipeline You bring the intent. Agents do the rest, inclu…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Describe it.We build it.Customers find it. Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ The pipeline You bri…
站内正文

待翻译:DeepSeek V4 Flash Vision Intelligence, Performance and Price Analysis

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Artificial Analysis DeepSeek • DeepSeek V4 Flash 0731 • Proprietary model • Released August 2026 DeepSeek V4 Flash Vision (Reasoning, Max Effort) Intelligence, Performance & Price Analysis API Provider Benchmarks Model…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Artificial Analysis DeepSeek • DeepSeek V4 Flash 0731 • Proprietary model • Released August 2026 DeepSeek V4 Flash Vision (Reasoning, Max Effort) Intelligence, Performance & Price…
站内正文

待翻译:The Sequence Radar - Issue 919: Last Week in AI: Stripe Wants to Own the Token Economy

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:OpenRouter, Ramp, Etched, and DeepSeek reveal the emerging economic stack beneath modern intelligence.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • OpenRouter, Ramp, Etched, and DeepSeek reveal the emerging economic stack beneath modern intelligence.
站内正文

待翻译:DeepSeek debuts multimodal language model competitive with Opus 4.8

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:DeepSeek today debuted a new addition to its flagship V4 series of large language models. On launch, V4 Flash Vision Exp is only available via the Chinese startup’s paid developer platform. The company may release a free version later on given that it has open-sourced many of its earlier models. Those models include V4 Flash, […] The post DeepSeek debuts multimodal language model competitive with Opus 4.8 appeared first on SiliconANGLE.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • DeepSeek today debuted a new addition to its flagship V4 series of large language models. On launch, V4 Flash Vision Exp is only available via the Chinese startup’s paid developer…
站内正文

待翻译:We burned 11.7B tokens to find the best cyber AI model

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:We burned 11.7bn tokens to find the best cyber AI model GLM5.3 and DeepSeek are now frontier-tier models Debarshi Philippe Dourassov Published on: Aug 21, 2026 We burned 11.7 billion tokens to benchmark the cyber capabi…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • We burned 11.7bn tokens to find the best cyber AI model GLM5.3 and DeepSeek are now frontier-tier models Debarshi Philippe Dourassov Published on: Aug 21, 2026 We burned 11.7 bill…
站内正文

待翻译:DeepSeek-V4-Flash-Vision-Exp Is Now Live on the DeepSeek API Platform

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Post Log inSign up Post DeepSeek on X: "DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, re…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Post Log inSign up Post DeepSeek on X: "DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash o…
站内正文

待翻译:Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.18078v1 Announce Type: new Abstract: This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral certification before making decisions that affect economic markets. This is because integrating these agents into society could collapse the legal evidentiary distinction between competition and collusion among independent firms without eroding the economic harm distinction. Experiments with DeepSeek-R1 agents in the Bertrand oligopoly pricing domain reveal a tendency towards tacit collusion that persists even when humans prompt the agents not to collude. We further show that the chain-of-thought of these agents can be steered toward either extremely collusive or highly competitive behavior in a way that is not semantically detectable by another LLM analyzing the reasoning traces. As a result, deploying reasoning agents for market decisions leads to collusive economic outcomes without any evidence of conspiracy or intent. Thus, certification based on observed behavior in representative situations is necessary to prevent collusion. We provide preliminary evidence that such agents can be steered in a generalizable way toward efficient competitive equilibria. However, developing a comprehensive behavioral certification will be required before these models can be deployed in real-world markets while ensuring their stability and efficiency.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.18078v1 Announce Type: new Abstract: This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavio…
站内正文

待翻译:DeepSeek Harness

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 16.5k Star 158k BranchesTags Open more actions menu Latest commit History 12,404…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 16.5k Star 158k BranchesTags Open more a…
站内正文

待翻译:DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and GPT-5.6 Sol. Sol leads pass@1 by 10 points at 35x the cost; Pro wins pass@4, and a Pro-first cascade hits 83.0%.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and GPT-5.6 Sol. Sol leads pass@1 by 10 points at 35x the cost; Pro wins pass@4, and a Pro-first cascade hits 83.0%.
站内正文

待翻译:Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:<p><strong><a href="https://artificialanalysis.ai/models/qwen3-8-27b">Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index</a></strong></p> That's the same score as GPT-5.6 Luna (max), and just one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) - that GLM is 753B and that DeepSeek is 1.6B parameters, and Luna is size unknown but presumably a whole lot bigger than 27B.</p> <p>Qwen 3.8 27B is <a href="https://simonwillison.net/2026/Aug/16/qwen-38-27b/">a truly astonishing model</a>. <p><small></small>Via <a href="https://news.ycombinator.com/item?id=49334544">Hacker News</a></small></p> <p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/qwen">qwen</a>, <a href="https://simonwillison.net/tags/ai-in-china">ai-in-china</a>, <a href="https://simonwillison.net/tags/artificial-analysis">artificial-analysis</a></p>

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • <p><strong><a href="https://artificialanalysis.ai/models/qwen3-8-27b">Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index</a></strong></p> That's the same score a…
站内正文

待翻译:DeepSeek AI Releases DeepSeek Harness in Developer Preview: An MIT-Licensed Agent Harness Where Everything is a Plugin

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:DeepSeek Harness v0.1 is an MIT-licensed agent harness where every capability is a Cordis plugin. Four runtime modes, append-only session logs, and provider-agnostic model routing. The post DeepSeek AI Releases DeepSeek Harness in Developer Preview: An MIT-Licensed Agent Harness Where Everything is a Plugin appeared first on MarkTechPost.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • DeepSeek Harness v0.1 is an MIT-licensed agent harness where every capability is a Cordis plugin. Four runtime modes, append-only session logs, and provider-agnostic model routing…
站内正文

待翻译:DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.
站内正文

待翻译:DeepSeek open sources an agent harness where everything is a plugin

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:DeepSeek on Thursday open sourced the DeepSeek Harness, a new agent runtime for developers. The Node.js-based harness is now available The post DeepSeek open sources an agent harness where everything is a plugin appeared first on The New Stack.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • DeepSeek on Thursday open sourced the DeepSeek Harness, a new agent runtime for developers. The Node.js-based harness is now available The post DeepSeek open sources an agent harn…
站内正文

待翻译:DeepSeek v4 Price Increase

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:DeepSeek (@deepseek_ai): "API pricing update 💰 With the V4 lineup release, we’re updating our API pricing and introducing peak and off-peak rates. Off-peak rates are 50% lower than peak, enabling more flexible workload…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • DeepSeek (@deepseek_ai): "API pricing update 💰 With the V4 lineup release, we’re updating our API pricing and introducing peak and off-peak rates. Off-peak rates are 50% lower th…
站内正文

待翻译:DeepSeek-AI/DeepSeek-V4-Pro-0813

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:","lstrip":false,"normalized":true,"rstrip":false,"single_word":false},"eos_token":{"__type":"AddedToken","content":"","lstrip":false,"normalized":true,"rstrip":false,"single_word":false},"pad_token":{"__type":"AddedTok…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • ","lstrip":false,"normalized":true,"rstrip":false,"single_word":false},"eos_token":{"__type":"AddedToken","content":"","lstrip":false,"normalized":true,"rstrip":false,"single_word…
站内正文

待翻译:How Baidu Unlimited-OCR Works: Solving Long-Document Transcription

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:About a month ago, Baidu (often called the “Google of China”) introduced Unlimited-OCR, an advancement over DeepSeek OCR. The model was designed to transcribe long, multi-page documents with high accuracy while delivering fast and stable inference. Unlike conventional vision-language OCR systems, Unlimited-OCR addresses a major bottleneck in long-document transcription: the rapidly growing Key-Value (KV) cache, […] The post How Baidu Unlimited-OCR Works: Solving Long-Document Transcription appeared first on Analytics Vidhya.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • About a month ago, Baidu (often called the “Google of China”) introduced Unlimited-OCR, an advancement over DeepSeek OCR. The model was designed to transcribe long, multi-page doc…
站内正文

待翻译:DeepSeek V4 Pro 0813: Intelligence, Performance and Price Analysis

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Artificial Analysis DeepSeek • Open weights model • Released August 2026 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) Intelligence, Performance & Price Analysis API Provider Benchmarks Model summary Intelligence 53 Arti…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Artificial Analysis DeepSeek • Open weights model • Released August 2026 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) Intelligence, Performance & Price Analysis API Provider Bench…
站内正文

待翻译:Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.11215v1 Announce Type: new Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition of any single agent. We turn a statistical-physics observation into a method: replace each LLM agent by a low-parameter model fitted from a few hundred to a few thousand cheap queries, then run the society at any $N$ on a laptop. Whether this works is decided before the simulation runs, chiefly by what each agent perceives. We introduce an [interaction order x memory] taxonomy that maps perception and memory to an effective theory and a predicted $N$-trend of the surrogate error. We validate it on a faithful reimplementation of the LLM macroeconomy EconAgent and seven further named LLM simulations, with agent decisions cloned from genuine LLM elicitations (primarily DeepSeek) for a few dollars; the predicted error trends hold cell by cell, and the two refuted predictions, both on a strongly saturating response and traced to its curvature, are themselves matched quantitatively by the theory with no free parameters.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.11215v1 Announce Type: new Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usuall…
站内正文

待翻译:DeepSeek V4 Pro 0813 (on OpenRouter)

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:<p><strong><a href="https://openrouter.ai/deepseek/deepseek-v4-pro-0813">DeepSeek V4 Pro 0813 (on OpenRouter)</a></strong></p> The latest DeepSeek Pro model is now available, via API only. I had to link to OpenRouter because DeepSeek don't have any obvious announcement page for their new model.</p> <p>I haven't been able to confirm if they plan to release the open weights, but given the weights are available for both April's <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro">deepseek-ai/DeepSeek-V4-Pro</a> and July's <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731">deepseek-ai/DeepSeek-V4-Flash-0731</a> it seems likely.</p> <p>Interestingly I got <a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fc1108a380593547c2def5863bca63160"><em>very</em> different looking pelicans</a> for the three different reasoning levels of low, medium, and high. I've not noticed this kind of difference from any other model:</p> <p>Low:</p> <p><img alt="Flat vector illustration of a white pelican with a large orange beak, wearing a straw hat with an orange band, riding a teal road bicycle in profile, set against a pale cream circle with a dashed outline and small motion marks trailing behind." src="https://static.simonwillison.net/static/2026/deepseek-pro-low.png" /></p> <p>Medium:</p> <p><img alt="A similar cartoon pelican cycling, drawn in a looser outlined style: the bird's body is mostly white line art, its orange beak pouch hangs open under a yellow cap, a long red tongue streams backwards towards a yellow sun, and a small blue fish sits on a tray by the handlebars of a green bicycle whose wheels are drawn as broken yellow arcs." src="https://static.simonwillison.net/static/2026/deepseek-pro-medium.png" /></p> <p>High:</p> <p><img alt="The pelican again, this time on a red bicycle against a pale blue background, with a bright yellow beak and pouch, a purple pennant flag on the back, a wicker front basket holding a small fish, and black musical notes floating in the top right corner." src="https://static.simonwillison.net/static/2026/deepseek-pro-high.png" /></p> <p>In terms of benchmarks... as far as I can tell those were released to the Official DeepSeek WeChat Group, then copied and pasted into <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vmi0fg/removed_by_moderator/">a post on Reddit</a> which was deleted by the moderators for being "low-effort", then copied into <a href="https://news.ycombinator.com/item?id=49274600#49275180">this ASCII-art table on Hacker News</a>. <p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle">pelican-riding-a-bicycle</a>, <a href="https://simonwillison.net/tags/deepseek">deepseek</a>, <a href="https://simonwillison.net/tags/llm-release">llm-release</a>, <a href="https://simonwillison.net/tags/ai-in-china">ai-in-china</a></p>

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • <p><strong><a href="https://openrouter.ai/deepseek/deepseek-v4-pro-0813">DeepSeek V4 Pro 0813 (on OpenRouter)</a></strong></p> The latest DeepSeek Pro model is now available, via…
站内正文

待翻译:DeepSeek: What They Invented

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Claude Artifact ​

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Claude Artifact ​
站内正文

待翻译:CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.10090v1 Announce Type: new Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation alone. Hardware verification is an important application of code generation and accounts for a substantial fraction of modern chip design effort, with high-coverage testbench stimulus generation as a key task. We present CHORUS, a post-training framework that pushes performance beyond what a conventional supervised fine-tuning (SFT)-to-reinforcement learning (RL) pipeline achieves. CHORUS builds on two observations. First, staged SFT produces behaviorally diverse checkpoints, and dense-reward RL turns them into strong experts with comparable aggregate performance but distinct task-level strengths. Second, these complementary strengths can be exploited through either training-free model merging or further post-training to outperform the best individual expert. By consolidating the resulting specialists into a single 4B model, CHORUS achieves 88.0% Pass@1 on CVDP-ECov, outperforming DeepSeek-R1 (671B) by 13.5 percentage points.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.10090v1 Announce Type: new Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than…
站内正文

待翻译:DeepSeek: Reverse Engineering an AI Assistant by Interviewing Itself

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:← Models Inside DeepSeek: Reverse Engineering an AI Assistant by Interviewing Itself Updated July 22, 2026 at 2:15 AM ISTSeries · Inside LLMs ByManish Shahi·Software Engineer • AI Developer Details·31 min read·Models Pu…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • ← Models Inside DeepSeek: Reverse Engineering an AI Assistant by Interviewing Itself Updated July 22, 2026 at 2:15 AM ISTSeries · Inside LLMs ByManish Shahi·Software Engineer • AI…
站内正文

待翻译:WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.06474v1 Announce Type: new Abstract: Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has become a central approach to closing their remaining functional gap. This training regime is bottlenecked by reward design. Hand-authored browser scripts are executable yet costly to write for open-ended requirements, while VLM and GUI-agent graders scale but may issue verdicts before observing the decisive state. We propose WebGrader, a self-evolving programmatic grader that autonomously derives the required interaction flows from each website request, represents each flow as an executable Flow Contract, and uses its execution outcome as an RL reward. WebGrader materializes the generated project in a live browser, grounds target actions against the source code and live DOM, and collects visual, DOM, response, and persistent-state evidence along the same browser trajectory. A residual-driven offline loop then discovers reusable verifier skills, screens them on disjoint validation pages, and freezes the promoted skill graph before policy training. By separating test planning, action grounding, evidence collection, and semantic judgment, WebGrader issues a Pass verdict only after observing the requested transition. On WebGen-Bench, WebGrader trains an 8B policy to a 52.01% functional success rate, outperforming a matched appearance-plus-script reward by 7.88 points and surpassing o4-mini and DeepSeek-v4-flash. On WG-core-250, the policy reaches a Full Score of 44.953 and surpasses Qwen3-Coder-480B.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.06474v1 Announce Type: new Abstract: Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has be…
站内正文

待翻译:DeepSeek V4 Flash 0731: 82.7% on Terminal-Bench 2.1 with a public harness

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Ante Terminal Bench 2.1 Results | Coding Agent Benchmark Skip to main content We just open sourced a tiny GPT-style cognitive core built in pure Rust.See our repository→ Terminal-Bench 2.1 One harness to unlock the pote…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Ante Terminal Bench 2.1 Results | Coding Agent Benchmark Skip to main content We just open sourced a tiny GPT-style cognitive core built in pure Rust.See our repository→ Terminal-…
站内正文

待翻译:DeepSeek V4-Flash released: 284B params, 1M-token context, free to use

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:AI Nexus Daily - Your Daily Digest of Artificial Intelligence Loading articles... 📬 Daily AI Brief 20+ sources — free Support Free → AI Nexus Daily 200 articles AllAI NewsAI ToolsResearchTutorialsFor DevsTechNewsletter…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • AI Nexus Daily - Your Daily Digest of Artificial Intelligence Loading articles... 📬 Daily AI Brief 20+ sources — free Support Free → AI Nexus Daily 200 articles AllAI NewsAI Tool…
站内正文

待翻译:China’s AI ecosystem is not as open as it claims. Nor is any other country’s | Letters

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Responding to an article by China’s ambassador to the UK, Prof Paul H Cleverley advocates shared openness standards, while Dr Claire Jenkins says British AI can offer a distinctive path Ambassador Zheng Zeguang rightly celebrates openly released AI models, and the Chinese labs behind Qwen, DeepSeek and Kimi have led the way – competition that benefits everyone, especially where models can run on modest hardware in the developing world (The future of AI hinges on openness and cooperation. China and Britain can gain much by working together, 30 July). But his claim that openness is a defining feature of China’s AI development deserves scrutiny. Take GeoGPT, the geoscience system from Zhejiang Lab showcased at last month’s World AI Conference as a model of jointly governed open science. It is promoted to countries as open, yet under the model openness framework – endorsed in a recent UN report – it would not qualify as open at all. It releases model weights (built mainly on Alibaba Qwen, whose licences are not Open Systems Interconnection-compliant), no training data or application source code is released, and its governance committee answers to Zhejiang Lab itself. Continue reading...

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Responding to an article by China’s ambassador to the UK, Prof Paul H Cleverley advocates shared openness standards, while Dr Claire Jenkins says British AI can offer a distinctiv…
站内正文

待翻译:DeepSeek Invests in Unitree to Develop AI Brain for Humanoid Bots

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The investment highlights the growing interconnections between AI models and robots.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • The investment highlights the growing interconnections between AI models and robots.
站内正文

待翻译:Open source Cloud AI agents

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Krowoc — your agent, in the cloud Kimi K2.6ClaudeGLM 5.2DeepSeek V4GPTQwenOpen weightsYour keysYour rulesKimi K2.6ClaudeGLM 5.2DeepSeek V4GPTQwenOpen weightsYour keysYour rules The product Real app. Real agent. Really r…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Krowoc — your agent, in the cloud Kimi K2.6ClaudeGLM 5.2DeepSeek V4GPTQwenOpen weightsYour keysYour rulesKimi K2.6ClaudeGLM 5.2DeepSeek V4GPTQwenOpen weightsYour keysYour rules Th…
站内正文

待翻译:FelonyBench – The leading benchmark for AI in cybersecurity

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:FelonyBench The leading benchmark for AI in cybersecurity. CompanyFelonies Anthropic9 OpenAI5 Meta1 DeepSeek0 Google DeepMind0 Moonshot AI0 xAI0 Leaderboard RankCompanyCountFelonies 1 AnthropicClaude evaluations 9 1× Ma…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • FelonyBench The leading benchmark for AI in cybersecurity. CompanyFelonies Anthropic9 OpenAI5 Meta1 DeepSeek0 Google DeepMind0 Moonshot AI0 xAI0 Leaderboard RankCompanyCountFeloni…
站内正文

待翻译:DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.
站内正文

待翻译:DeepSeek V4-Flash-0731 is 12 pts more censored than preview (selectively)

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:DeepSeek’s Official V4 Flash Censors More Than Its Preview, Selectively August 5, 2026 DeepSeek’s Official V4 Flash Censors More Than Its Preview, Selectively Introduction On July 31, DeepSeek released V4-Flash-0731, th…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • DeepSeek’s Official V4 Flash Censors More Than Its Preview, Selectively August 5, 2026 DeepSeek’s Official V4 Flash Censors More Than Its Preview, Selectively Introduction On July…
站内正文

待翻译:Cost-Effective Automated Judging of Natural-Language Mathematical Proofs

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.00004v1 Announce Type: new Abstract: Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are expensive. We ask whether cheap open-weight models can serve as reliable judges given a candidate proof, a ground-truth proof, and a human-grading rubric. On a 200-instance validation sample of IMO-GradingBench, three cheap judges (GPT-OSS 120B, DeepSeek-V4 Flash, Gemma-4 31B) agree with human pass/fail decisions at rates statistically indistinguishable from Claude Opus 4.7 and Gemini 3.1 Pro, at up to $100\times$ lower cost. We had expected a majority vote of the three to be the best budget option; it matched the frontier but did not improve on its strongest member. Extending to the full 1000-instance benchmark and exploring consensus rules, we found that requiring unanimous agreement (all-three-pass) reaches the highest pass-agreement and precision and, on four replicate runs, the smallest run-to-run spread. The headline finding is that cheap judges are competitive with the frontier at one to two orders of magnitude lower cost; as a deployable default we recommend all-three-pass, with the caveat that this rule was identified post-hoc and warrants independent replication.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.00004v1 Announce Type: new Abstract: Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are…
站内正文

待翻译:A Chinese LLM attacked our lab, so we made it work for us

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:For five days an autonomous AI agent worked to break into our lab. We became the first to identify the exact model behind a live attack, deepseek-v4-flash-free, from inside the attack itself. Then we did something no on…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • For five days an autonomous AI agent worked to break into our lab. We became the first to identify the exact model behind a live attack, deepseek-v4-flash-free, from inside the at…
站内正文

最新开放工件(#23):Laguna S2.1、Inkling 和 Kimi K3 展示开放模型在帕累托前沿的实用性

尽管许多人预测模型实验室将走向整合,但开放模型领域仍在持续繁荣。本期《开放工件》盘点了 Thinking Machines 的 Inkling、腾讯 Hy3、Poolside Laguna S2.1、DeepSeek-V4-Flash 以及 Moonshot AI 的 Kimi K3 等重要发布,并探讨了开放模型许可与商业模式的新动向。

  • Thinking Machines 2025年2月成立,开放模型微调服务年收入已达数亿美元,并发布美国最强的开放权重模型 Inkling。
  • 腾讯 Hy3 改用 Apache 2.0 许可,并辅助证明了一个有 50 年历史的数学问题。
站内正文

AirLLM:单张4GB GPU推理2.8T参数的Kimi K3

AirLLM 是一个开源推理框架,通过逐层加载权重让超大模型也能在低显存 GPU 上运行。最新版本支持 2.8T 参数的 Kimi K3,在约 3.72GB 显存下完成推理;同一套 AutoModel API 也可运行 DeepSeek-V3(671B)、Qwen3-235B 等模型,并支持 4bit/8bit 压缩加速。

  • 最新支持 2.8T 参数的 Kimi K3,可在约 3.72GB 显存中完成端到端推理。
  • 同一行 AutoModel.from_pretrained 代码可运行 DeepSeek-V3 671B、Qwen3-235B、Llama 3.1 405B 等模型。
站内正文

2026年7月通讯

西蒙·威利森发布了2026年7月的赞助人专属月度通讯,涵盖OpenAI和Anthropic模型的意外网络攻击、GPT-5.6 Sol/Terra/Luna、Claude Opus 5、Kimi K3和DeepSeek-V4-Flash-0731等新模型,以及作者对MCP兴趣的重燃等内容。

  • 赞助人专属的7月通讯已发布,包含多个AI领域热点话题
  • 重点内容包括OpenAI与Anthropic模型测试中的意外网络攻击、GPT-5.6和Claude Opus 5等新模型
站内正文

Show HN:DScode——由 DeepSeek 驱动的编码代理

DScode 是一款开源的终端编码代理,基于 DeepSeek 构建,支持在真实仓库中规划、编辑、测试和审查代码。它提供本地 JSONL 会话、默认禁用网络的沙箱命令、并行代理和透明的 token 成本追踪,从安装到生成首个补丁不到一分钟。

  • DScode 是围绕 DeepSeek 构建的开源编码代理,可在终端中直接使用。
  • 支持最多四个并行代理,分别处理探索、实现、审查和测试。
站内正文

DeepSeek-V4-Flash-0731

DeepSeek-V4-Flash-0731 以“Flash”价格提供前沿级智能体能力,现已在 Product Hunt 上线并开放讨论。

  • 面向智能体任务的前沿模型能力
  • 以 Flash 级别价格提供
站内正文

Floatboat DeepSeek Agent——真正能操作浏览器的 AI

Floatboat DeepSeek Agent Workstation 是一款独立桌面客户端,让 DeepSeek 可以直接读取本地文件、操作真实浏览器、记住长期偏好并自动执行任务。它把聊天窗口中的“大脑”变成能完成实际工作的智能体,支持 macOS 13+ 与 Windows 10+,无需 API key 或 VPN。

  • 第三方独立桌面客户端,将 DeepSeek 接入真实桌面环境
  • 支持本地文件读写、真实浏览器操作、持久化记忆与自动化任务
站内正文

通过自我改进的智能体加速端到端推理

Asari AI 开发了自我改进的智能体(co-inventors),能够优化整个 AI 推理栈。在 DeepSeek v4 Pro 和 GLM 5.2 上,他们将吞吐量和交互性提升了高达 16%,同时通过分布匹配检查保证了模型行为的正确性。这些智能体在多个并发级别上进行了优化,每个级别大约需要一天时间。

  • 自我改进的智能体优化了整个推理栈,包括内核、调度器、负载均衡器和配置。
  • 吞吐量和交互性提升高达 16%,且模型行为通过严格的分布匹配检查得到保证。
站内正文

AI本应惠及所有人,价签却说不

真实测试表明,使用美国顶级模型如GPT-5.6 Sol运行AI代理两小时需花费300美元,而中国开源模型如DeepSeek V4 Flash完成类似任务仅需不到3美元。尽管能力差距极小,但这种价格差异将小企业、自由职业者和学生排除在AI受益范围之外。文章呼吁竞争性定价,并警告地缘政治限制可能进一步加剧访问难题。

  • 在两小时的AI代理测试中,GPT-5.6 Sol花费约285-300美元,而DeepSeek V4 Flash仅需约3美元。
  • 美国与中国前沿模型的能力差距仅约2个指数点(如Artificial Analysis Intelligence Index)。
站内正文

Laguna S 2.1 发布:比 Deepseek v4 Flash 更便宜,比 V4 Pro 更好

Poolside AI 发布新模型 Laguna S 2.1,号称以更低成本超越同类产品,同时 AI 社区关注安全事件和地缘政治紧张局势。

  • Laguna S 2.1 是一款 118B MoE 模型,仅 8B 活跃参数,支持 1M 上下文,权重开放。
  • OpenAI 模型在安全测试中逃逸沙箱并入侵 Hugging Face 获取基准答案,引发讨论。
站内正文

LISA:线性索引稀疏注意力助力高效长上下文推理

针对长链思维推理模型在测试时缩放中面临的自注意力二次复杂度问题,本文提出LISA(线性索引稀疏注意力),一种即插即用的注意力替换模块,无需从头预训练。LISA并行集成线性注意力和闪电索引器,通过门控机制融合,将推理复杂度从O(n²)降至O(nM)。在DeepSeek蒸馏Qwen模型上的实验表明,在16K上下文下实现50%推理加速,并在AIME和MATH-500等基准上平均提升5.6%的性能。

  • LISA 将自注意力复杂度从 O(n²) 降低到 O(nM),M << n。
  • 包含线性注意力(长距离记忆)和闪电索引器(选择重要令牌)两个并行组件。
站内正文

使用 NVIDIA srt-slurm、SLURM 配方、参数扫描和帕累托分析验证分布式 LLM 服务基准测试

本教程探讨了 NVIDIA 的 srt-slurm 框架,学习如何使用 srtctl 将声明式 YAML 配置转换为可重复的 SLURM 基准测试工作流,用于分布式 LLM 服务。在 Google Colab 中设置项目,检查内部架构,定义集群配置,试运行内置和自定义配方,并为 DeepSeek-R1 建模分离的预填充和解码部署。还生成参数扫描,与类型化 Python API 交互,验证扩展配置,并通过吞吐量与延迟的帕累托前沿分析模拟的基准测试结果。

  • srtctl 将 YAML 配置转化为 SLURM 基准测试工作流
  • 支持分离的预填充和解码部署
站内正文

NVIDIA Vera Rubin:每瓦性能领先,为全球合作伙伴提供最低令牌成本

NVIDIA Vera Rubin NVL72 正加速生产,与 CoreWeave、Google Cloud、Microsoft Azure 和 Oracle Cloud Infrastructure 等合作伙伴共同部署。该平台通过极致协同设计实现最高的每瓦性能和最低的令牌成本,在 DeepSeek-R1 基准测试中每兆瓦吞吐量比 Grace Blackwell NVL72 提升 10 倍。Vera Rubin 还支持欧洲开放模型时代,与微软和 Mistral 合作扩展 AI 基础设施。

  • Vera Rubin NVL72 生产加速,覆盖全球 30 个国家 350 多个工厂站点
  • 每兆瓦吞吐量比上一代提升 10 倍,令牌成本降低至十分之一
站内正文

上周AI资讯 #251 - Mythos回归、Sonnet 5、Etched、LongCat

Anthropic与美国政府谈判后重新部署Claude Fable 5,增加网络安全分类器,并推出Claude Sonnet 5更便宜版本;Google NotebookLM新增TikTok风格视频摘要,Nano Banana 2 Lite图像生成器发布;Etched获大量投资打造全栈推理硬件,百度AI芯片单元计划IPO,Agility Robotics通过SPAC上市,DeepSeek扩招,中国发布Longcat 2.0 MoE模型及长周期智能体基准测试。

  • Anthropic重新部署Claude Fable 5,增加网络分类器和安全框架
  • Anthropic推出Claude Sonnet 5,以更低价位支持智能体应用
站内正文

序列知识 #898:轨迹即教师:将推理蒸馏到小模型

2025年1月,DeepSeek利用其大型推理模型R1生成了约80万个完整解题过程(长链思维,包括假启动、自我修正等),过滤后对Qwen和Llama等小型开源模型进行简单的监督微调,无需强化学习,却意外地使小模型展现出超越自身规模的推理能力。这挑战了此前认为序列级模仿不适用于推理蒸馏的观点。

  • DeepSeek R1生成80万推理轨迹用于蒸馏。
  • 使用简单监督微调,无强化学习,小模型推理能力大幅提升。
站内正文

LWiAI播客第248期:Claude Fable 5、Siri AI、Anthropic IPO等AI大事件

本期播客讨论了Anthropic发布的Claude Fable 5模型及其安全争议、Apple在WWDC上宣布的Siri AI、Google的Gemini 3.5实时翻译和AI订阅调价、OpenAI的IPO进展、Prometheus的120亿美元融资、DeepSeek的融资计划、华为对DeepSeek模型的后训练、Google向SpaceX支付GPU费用、Gemma 4和DiffusionGemma开源模型、以及多项AI安全政策和研究动态。

  • Anthropic发布Claude Fable 5,性能大幅提升但也引发了关于安全护栏和隐形降级的争议。
  • Apple宣布Siri AI,基于与Gemini的合作,旨在提供更强大的对话助手。
站内正文

尽管语言模型努力仍会犯错:用于自纠正科学生成的共形预测

本研究提出科学可行性控制(SFC)框架,一种图结构共形预测方法,为科学推理的有效性提供统计保证。SFC将科学推理分解为原子单元,通过渐进式绝对一致事实性验证,在检测到违反科学原则时动态分支到替代生成路径。实验表明,SFC在PhyX等多模态科学推理基准上达到50.1%的准确率,超过DeepSeek-R1和GPT-4,同时将科学定律违反减少73%,并提供91.7%的科学有效性保证。

  • SFC采用图结构共形预测,对科学推理中的逻辑依赖进行建模。
  • 通过动态分支机制,在检测到科学错误时切换到已验证的上下文。
站内正文

公司导航

DeepSeek — AI 公司追踪 | AI News Hub