AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Describe it.We build it.Customers find it. Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ The pipeline You bring the intent. Agents do the rest, inclu…
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
Describe it.We build it.Customers find it. Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ The pipeline You bri…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:DeepSeek today debuted a new addition to its flagship V4 series of large language models. On launch, V4 Flash Vision Exp is only available via the Chinese startup’s paid developer platform. The company may release a free version later on given that it has open-sourced many of its earlier models. Those models include V4 Flash, […] The post DeepSeek debuts multimodal language model competitive with Opus 4.8 appeared first on SiliconANGLE.
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
DeepSeek today debuted a new addition to its flagship V4 series of large language models. On launch, V4 Flash Vision Exp is only available via the Chinese startup’s paid developer…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We burned 11.7bn tokens to find the best cyber AI model GLM5.3 and DeepSeek are now frontier-tier models Debarshi Philippe Dourassov Published on: Aug 21, 2026 We burned 11.7 billion tokens to benchmark the cyber capabi…
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
We burned 11.7bn tokens to find the best cyber AI model GLM5.3 and DeepSeek are now frontier-tier models Debarshi Philippe Dourassov Published on: Aug 21, 2026 We burned 11.7 bill…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Post Log inSign up Post DeepSeek on X: "DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, re…
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
Post Log inSign up Post DeepSeek on X: "DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash o…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.18078v1 Announce Type: new Abstract: This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral certification before making decisions that affect economic markets. This is because integrating these agents into society could collapse the legal evidentiary distinction between competition and collusion among independent firms without eroding the economic harm distinction. Experiments with DeepSeek-R1 agents in the Bertrand oligopoly pricing domain reveal a tendency towards tacit collusion that persists even when humans prompt the agents not to collude. We further show that the chain-of-thought of these agents can be steered toward either extremely collusive or highly competitive behavior in a way that is not semantically detectable by another LLM analyzing the reasoning traces. As a result, deploying reasoning agents for market decisions leads to collusive economic outcomes without any evidence of conspiracy or intent. Thus, certification based on observed behavior in representative situations is necessary to prevent collusion. We provide preliminary evidence that such agents can be steered in a generalizable way toward efficient competitive equilibria. However, developing a comprehensive behavioral certification will be required before these models can be deployed in real-world markets while ensuring their stability and efficiency.
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
arXiv:2608.18078v1 Announce Type: new Abstract: This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavio…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 16.5k Star 158k BranchesTags Open more actions menu Latest commit History 12,404…
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 16.5k Star 158k BranchesTags Open more a…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and GPT-5.6 Sol. Sol leads pass@1 by 10 points at 35x the cost; Pro wins pass@4, and a Pro-first cascade hits 83.0%.
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and GPT-5.6 Sol. Sol leads pass@1 by 10 points at 35x the cost; Pro wins pass@4, and a Pro-first cascade hits 83.0%.
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:<p><strong><a href="https://artificialanalysis.ai/models/qwen3-8-27b">Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index</a></strong></p> That's the same score as GPT-5.6 Luna (max), and just one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) - that GLM is 753B and that DeepSeek is 1.6B parameters, and Luna is size unknown but presumably a whole lot bigger than 27B.</p> <p>Qwen 3.8 27B is <a href="https://simonwillison.net/2026/Aug/16/qwen-38-27b/">a truly astonishing model</a>. <p><small></small>Via <a href="https://news.ycombinator.com/item?id=49334544">Hacker News</a></small></p> <p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/qwen">qwen</a>, <a href="https://simonwillison.net/tags/ai-in-china">ai-in-china</a>, <a href="https://simonwillison.net/tags/artificial-analysis">artificial-analysis</a></p>
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
<p><strong><a href="https://artificialanalysis.ai/models/qwen3-8-27b">Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index</a></strong></p> That's the same score a…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:DeepSeek Harness v0.1 is an MIT-licensed agent harness where every capability is a Cordis plugin. Four runtime modes, append-only session logs, and provider-agnostic model routing. The post DeepSeek AI Releases DeepSeek Harness in Developer Preview: An MIT-Licensed Agent Harness Where Everything is a Plugin appeared first on MarkTechPost.
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
DeepSeek Harness v0.1 is an MIT-licensed agent harness where every capability is a Cordis plugin. Four runtime modes, append-only session logs, and provider-agnostic model routing…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:DeepSeek on Thursday open sourced the DeepSeek Harness, a new agent runtime for developers. The Node.js-based harness is now available The post DeepSeek open sources an agent harness where everything is a plugin appeared first on The New Stack.
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
DeepSeek on Thursday open sourced the DeepSeek Harness, a new agent runtime for developers. The Node.js-based harness is now available The post DeepSeek open sources an agent harn…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:DeepSeek (@deepseek_ai): "API pricing update 💰 With the V4 lineup release, we’re updating our API pricing and introducing peak and off-peak rates. Off-peak rates are 50% lower than peak, enabling more flexible workload…
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
DeepSeek (@deepseek_ai): "API pricing update 💰 With the V4 lineup release, we’re updating our API pricing and introducing peak and off-peak rates. Off-peak rates are 50% lower th…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:","lstrip":false,"normalized":true,"rstrip":false,"single_word":false},"eos_token":{"__type":"AddedToken","content":"","lstrip":false,"normalized":true,"rstrip":false,"single_word":false},"pad_token":{"__type":"AddedTok…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:About a month ago, Baidu (often called the “Google of China”) introduced Unlimited-OCR, an advancement over DeepSeek OCR. The model was designed to transcribe long, multi-page documents with high accuracy while delivering fast and stable inference. Unlike conventional vision-language OCR systems, Unlimited-OCR addresses a major bottleneck in long-document transcription: the rapidly growing Key-Value (KV) cache, […] The post How Baidu Unlimited-OCR Works: Solving Long-Document Transcription appeared first on Analytics Vidhya.
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
About a month ago, Baidu (often called the “Google of China”) introduced Unlimited-OCR, an advancement over DeepSeek OCR. The model was designed to transcribe long, multi-page doc…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Artificial Analysis DeepSeek • Open weights model • Released August 2026 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) Intelligence, Performance & Price Analysis API Provider Benchmarks Model summary Intelligence 53 Arti…
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
Artificial Analysis DeepSeek • Open weights model • Released August 2026 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) Intelligence, Performance & Price Analysis API Provider Bench…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.11215v1 Announce Type: new Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition of any single agent. We turn a statistical-physics observation into a method: replace each LLM agent by a low-parameter model fitted from a few hundred to a few thousand cheap queries, then run the society at any $N$ on a laptop. Whether this works is decided before the simulation runs, chiefly by what each agent perceives. We introduce an [interaction order x memory] taxonomy that maps perception and memory to an effective theory and a predicted $N$-trend of the surrogate error. We validate it on a faithful reimplementation of the LLM macroeconomy EconAgent and seven further named LLM simulations, with agent decisions cloned from genuine LLM elicitations (primarily DeepSeek) for a few dollars; the predicted error trends hold cell by cell, and the two refuted predictions, both on a strongly saturating response and traced to its curvature, are themselves matched quantitatively by the theory with no free parameters.
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
arXiv:2608.11215v1 Announce Type: new Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usuall…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:<p><strong><a href="https://openrouter.ai/deepseek/deepseek-v4-pro-0813">DeepSeek V4 Pro 0813 (on OpenRouter)</a></strong></p> The latest DeepSeek Pro model is now available, via API only. I had to link to OpenRouter because DeepSeek don't have any obvious announcement page for their new model.</p> <p>I haven't been able to confirm if they plan to release the open weights, but given the weights are available for both April's <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro">deepseek-ai/DeepSeek-V4-Pro</a> and July's <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731">deepseek-ai/DeepSeek-V4-Flash-0731</a> it seems likely.</p> <p>Interestingly I got <a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fc1108a380593547c2def5863bca63160"><em>very</em> different looking pelicans</a> for the three different reasoning levels of low, medium, and high. I've not noticed this kind of difference from any other model:</p> <p>Low:</p> <p><img alt="Flat vector illustration of a white pelican with a large orange beak, wearing a straw hat with an orange band, riding a teal road bicycle in profile, set against a pale cream circle with a dashed outline and small motion marks trailing behind." src="https://static.simonwillison.net/static/2026/deepseek-pro-low.png" /></p> <p>Medium:</p> <p><img alt="A similar cartoon pelican cycling, drawn in a looser outlined style: the bird's body is mostly white line art, its orange beak pouch hangs open under a yellow cap, a long red tongue streams backwards towards a yellow sun, and a small blue fish sits on a tray by the handlebars of a green bicycle whose wheels are drawn as broken yellow arcs." src="https://static.simonwillison.net/static/2026/deepseek-pro-medium.png" /></p> <p>High:</p> <p><img alt="The pelican again, this time on a red bicycle against a pale blue background, with a bright yellow beak and pouch, a purple pennant flag on the back, a wicker front basket holding a small fish, and black musical notes floating in the top right corner." src="https://static.simonwillison.net/static/2026/deepseek-pro-high.png" /></p> <p>In terms of benchmarks... as far as I can tell those were released to the Official DeepSeek WeChat Group, then copied and pasted into <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vmi0fg/removed_by_moderator/">a post on Reddit</a> which was deleted by the moderators for being "low-effort", then copied into <a href="https://news.ycombinator.com/item?id=49274600#49275180">this ASCII-art table on Hacker News</a>. <p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle">pelican-riding-a-bicycle</a>, <a href="https://simonwillison.net/tags/deepseek">deepseek</a>, <a href="https://simonwillison.net/tags/llm-release">llm-release</a>, <a href="https://simonwillison.net/tags/ai-in-china">ai-in-china</a></p>
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
<p><strong><a href="https://openrouter.ai/deepseek/deepseek-v4-pro-0813">DeepSeek V4 Pro 0813 (on OpenRouter)</a></strong></p> The latest DeepSeek Pro model is now available, via…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.10090v1 Announce Type: new Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation alone. Hardware verification is an important application of code generation and accounts for a substantial fraction of modern chip design effort, with high-coverage testbench stimulus generation as a key task. We present CHORUS, a post-training framework that pushes performance beyond what a conventional supervised fine-tuning (SFT)-to-reinforcement learning (RL) pipeline achieves. CHORUS builds on two observations. First, staged SFT produces behaviorally diverse checkpoints, and dense-reward RL turns them into strong experts with comparable aggregate performance but distinct task-level strengths. Second, these complementary strengths can be exploited through either training-free model merging or further post-training to outperform the best individual expert. By consolidating the resulting specialists into a single 4B model, CHORUS achieves 88.0% Pass@1 on CVDP-ECov, outperforming DeepSeek-R1 (671B) by 13.5 percentage points.
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
arXiv:2608.10090v1 Announce Type: new Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:← Models Inside DeepSeek: Reverse Engineering an AI Assistant by Interviewing Itself Updated July 22, 2026 at 2:15 AM ISTSeries · Inside LLMs ByManish Shahi·Software Engineer • AI Developer Details·31 min read·Models Pu…
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
← Models Inside DeepSeek: Reverse Engineering an AI Assistant by Interviewing Itself Updated July 22, 2026 at 2:15 AM ISTSeries · Inside LLMs ByManish Shahi·Software Engineer • AI…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.06474v1 Announce Type: new Abstract: Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has become a central approach to closing their remaining functional gap. This training regime is bottlenecked by reward design. Hand-authored browser scripts are executable yet costly to write for open-ended requirements, while VLM and GUI-agent graders scale but may issue verdicts before observing the decisive state. We propose WebGrader, a self-evolving programmatic grader that autonomously derives the required interaction flows from each website request, represents each flow as an executable Flow Contract, and uses its execution outcome as an RL reward. WebGrader materializes the generated project in a live browser, grounds target actions against the source code and live DOM, and collects visual, DOM, response, and persistent-state evidence along the same browser trajectory. A residual-driven offline loop then discovers reusable verifier skills, screens them on disjoint validation pages, and freezes the promoted skill graph before policy training. By separating test planning, action grounding, evidence collection, and semantic judgment, WebGrader issues a Pass verdict only after observing the requested transition. On WebGen-Bench, WebGrader trains an 8B policy to a 52.01% functional success rate, outperforming a matched appearance-plus-script reward by 7.88 points and surpassing o4-mini and DeepSeek-v4-flash. On WG-core-250, the policy reaches a Full Score of 44.953 and surpasses Qwen3-Coder-480B.
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
arXiv:2608.06474v1 Announce Type: new Abstract: Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has be…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Ante Terminal Bench 2.1 Results | Coding Agent Benchmark Skip to main content We just open sourced a tiny GPT-style cognitive core built in pure Rust.See our repository→ Terminal-Bench 2.1 One harness to unlock the pote…
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
Ante Terminal Bench 2.1 Results | Coding Agent Benchmark Skip to main content We just open sourced a tiny GPT-style cognitive core built in pure Rust.See our repository→ Terminal-…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Responding to an article by China’s ambassador to the UK, Prof Paul H Cleverley advocates shared openness standards, while Dr Claire Jenkins says British AI can offer a distinctive path Ambassador Zheng Zeguang rightly celebrates openly released AI models, and the Chinese labs behind Qwen, DeepSeek and Kimi have led the way – competition that benefits everyone, especially where models can run on modest hardware in the developing world (The future of AI hinges on openness and cooperation. China and Britain can gain much by working together, 30 July). But his claim that openness is a defining feature of China’s AI development deserves scrutiny. Take GeoGPT, the geoscience system from Zhejiang Lab showcased at last month’s World AI Conference as a model of jointly governed open science. It is promoted to countries as open, yet under the model openness framework – endorsed in a recent UN report – it would not qualify as open at all. It releases model weights (built mainly on Alibaba Qwen, whose licences are not Open Systems Interconnection-compliant), no training data or application source code is released, and its governance committee answers to Zhejiang Lab itself. Continue reading...
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
Responding to an article by China’s ambassador to the UK, Prof Paul H Cleverley advocates shared openness standards, while Dr Claire Jenkins says British AI can offer a distinctiv…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Krowoc — your agent, in the cloud Kimi K2.6ClaudeGLM 5.2DeepSeek V4GPTQwenOpen weightsYour keysYour rulesKimi K2.6ClaudeGLM 5.2DeepSeek V4GPTQwenOpen weightsYour keysYour rules The product Real app. Real agent. Really r…
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
Krowoc — your agent, in the cloud Kimi K2.6ClaudeGLM 5.2DeepSeek V4GPTQwenOpen weightsYour keysYour rulesKimi K2.6ClaudeGLM 5.2DeepSeek V4GPTQwenOpen weightsYour keysYour rules Th…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:FelonyBench The leading benchmark for AI in cybersecurity. CompanyFelonies Anthropic9 OpenAI5 Meta1 DeepSeek0 Google DeepMind0 Moonshot AI0 xAI0 Leaderboard RankCompanyCountFelonies 1 AnthropicClaude evaluations 9 1× Ma…
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
FelonyBench The leading benchmark for AI in cybersecurity. CompanyFelonies Anthropic9 OpenAI5 Meta1 DeepSeek0 Google DeepMind0 Moonshot AI0 xAI0 Leaderboard RankCompanyCountFeloni…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:DeepSeek’s Official V4 Flash Censors More Than Its Preview, Selectively August 5, 2026 DeepSeek’s Official V4 Flash Censors More Than Its Preview, Selectively Introduction On July 31, DeepSeek released V4-Flash-0731, th…
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
DeepSeek’s Official V4 Flash Censors More Than Its Preview, Selectively August 5, 2026 DeepSeek’s Official V4 Flash Censors More Than Its Preview, Selectively Introduction On July…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.00004v1 Announce Type: new Abstract: Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are expensive. We ask whether cheap open-weight models can serve as reliable judges given a candidate proof, a ground-truth proof, and a human-grading rubric. On a 200-instance validation sample of IMO-GradingBench, three cheap judges (GPT-OSS 120B, DeepSeek-V4 Flash, Gemma-4 31B) agree with human pass/fail decisions at rates statistically indistinguishable from Claude Opus 4.7 and Gemini 3.1 Pro, at up to $100\times$ lower cost. We had expected a majority vote of the three to be the best budget option; it matched the frontier but did not improve on its strongest member. Extending to the full 1000-instance benchmark and exploring consensus rules, we found that requiring unanimous agreement (all-three-pass) reaches the highest pass-agreement and precision and, on four replicate runs, the smallest run-to-run spread. The headline finding is that cheap judges are competitive with the frontier at one to two orders of magnitude lower cost; as a deployable default we recommend all-three-pass, with the caveat that this rule was identified post-hoc and warrants independent replication.
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
arXiv:2608.00004v1 Announce Type: new Abstract: Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:For five days an autonomous AI agent worked to break into our lab. We became the first to identify the exact model behind a live attack, deepseek-v4-flash-free, from inside the attack itself. Then we did something no on…
AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
For five days an autonomous AI agent worked to break into our lab. We became the first to identify the exact model behind a live attack, deepseek-v4-flash-free, from inside the at…
サイモン・ウィリソンが2026年7月のスポンサー限定月刊ニュースレターを公開。OpenAIやAnthropicモデルによる偶発的サイバー攻撃、GPT-5.6 Sol/Terra/Luna、Claude Opus 5、Kimi K3、DeepSeek-V4-Flash-0731などの話題に加え、MCPへの関心の再燃を取り上げています。
7月のスポンサー限定ニュースレターが公開され、スポンサーが閲覧可能
AIモデルの偶発的攻撃や新モデル(GPT-5.6、Claude Opus 5など)、MCPへの関心再燃を特集
FloatboatはDeepSeekをデスクトップに直接接続する独立系クライアントです。ローカルファイルの読み書き、実際のブラウザ操作、記憶の永続化、自動化タスクを実現し、チャットだけのAIを「仕事を仕上げるエージェント」に変えます。macOS 13+ / Windows 10+に対応し、APIキーやVPNは不要です。
Asari AI は、自己改善エージェント(共同発明者)が AI 推論スタック全体を最適化し、DeepSeek v4 Pro と GLM 5.2 でスループットとインタラクティブ性を最大 16% 向上させたことを発表しました。分布レベルの正確性チェックによりモデルの動作を維持し、エージェントは実験から学んだ知識をモデル間で転用します。