AI News HubLIVE

ソース分布

  • Hacker News AI21
  • Simon Willison's Weblog5
  • SiliconANGLE AI4
  • The New Stack AI3
  • arXiv Computational Linguistics2
  • MarkTechPost2
  • Product Hunt AI2
  • The Verge AI2

トピック分布

  • 研究25
  • Agent23
  • モデル18
  • 政策10
  • チップ6
  • スタートアップ5
  • ツール5

タイムライン

  • 2026-08-2113
  • 2026-08-269
  • 2026-08-207
  • 2026-08-237
  • 2026-08-226
  • 2026-08-245
  • 2026-08-253

最新動向

翻訳待ち:Show HN: LLM-powered webapp to build LLM-powered webapps

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Hey all, The goal is to earn on token margins for LLM calls when you build an AI-powered webapp. I proxy OpenAI and Anthropic calls so that when you deploy a site to a subdomain, your users token usage will be tracked.…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Hey all, The goal is to earn on token margins for LLM calls when you build an AI-powered webapp. I proxy OpenAI and Anthropic calls so that when you deploy a site to a subdomain,…
サイト内本文

翻訳待ち:Kraftapp AI – Describe it. We build it. Customers find it

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Describe it.We build it.Customers find it. Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ The pipeline You bring the intent. Agents do the rest, inclu…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Describe it.We build it.Customers find it. Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ The pipeline You bri…
サイト内本文

翻訳待ち:Students prefer Gemini over ChatGPT and Claude for AI essays in blind tests

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Key takeaways Gemini is StudyArena's current pick for college essays, with a 39.6% blind writing choice rate ahead of Claude at 31.8% and ChatGPT or OpenAI at 29.2%. Use Gemini as an editor, not a ghostwriter. Ask it to…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Key takeaways Gemini is StudyArena's current pick for college essays, with a 39.6% blind writing choice rate ahead of Claude at 31.8% and ChatGPT or OpenAI at 29.2%. Use Gemini as…
サイト内本文

翻訳待ち:AI and Constitutions (From My Email)

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:“Dear Tyler, I enjoyed reading your notes on visiting Anthropic to advise on Claude’s constitution. Framing AI governance around the common law, case law (“Talmud”), and independent adjudication is a much more adaptive…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • “Dear Tyler, I enjoyed reading your notes on visiting Anthropic to advise on Claude’s constitution. Framing AI governance around the common law, case law (“Talmud”), and independe…
サイト内本文

翻訳待ち:Show HN: Dribbling the AI Watermark Directly In-Prompt

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Julian Habekost Aug 25, 2026 Anthropic introduced a watermark into Claude’s output this month, and others have already followed or will soon follow suit. It is actually a little more sophisticated than simply using a ty…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Julian Habekost Aug 25, 2026 Anthropic introduced a watermark into Claude’s output this month, and others have already followed or will soon follow suit. It is actually a little m…
サイト内本文

翻訳待ち:Anthropic updates Claude’s memory to enhance customization and protect sensitive topics

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Artificial intelligence startup Anthropic PBC announced today it’s changing how Claude, its flagship AI product, uses memory by allowing users to see everything it remembers “topic by topic,” and edit or delete any of it. Claude also does not store sensitive subjects by default. This includes topics mentioned by users, including health concerns, race, ethnicity, […] The post Anthropic updates Claude’s memory to enhance customization and protect sensitive topics appeared first on SiliconANGLE.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Artificial intelligence startup Anthropic PBC announced today it’s changing how Claude, its flagship AI product, uses memory by allowing users to see everything it remembers “topi…
サイト内本文

翻訳待ち:Anthropic gives chat and Cowork one memory

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:On Tuesday, Anthropic launched a major update to how Claude remembers things. The new system combines Claude’s memory in Cowork The post Anthropic gives chat and Cowork one memory appeared first on The New Stack.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • On Tuesday, Anthropic launched a major update to how Claude remembers things. The new system combines Claude’s memory in Cowork The post Anthropic gives chat and Cowork one memory…
サイト内本文

翻訳待ち:Anthropic's Claude and Cowork will share memories about you now - unless you opt out

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Anthropic is merging Claude chat and Cowork memory, raising privacy questions about what AI remembers and if it's worth the tradeoff.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Anthropic is merging Claude chat and Cowork memory, raising privacy questions about what AI remembers and if it's worth the tradeoff.
サイト内本文

翻訳待ち:Apple's new desktop computers are designed specifically for local AI development

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Many developers have been changing their workflows to incorporate AI coding agents more heavily, thanks to powerful frontier large language models and increasingly sophisticated harnesses like Claude Code or Codex, but…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Many developers have been changing their workflows to incorporate AI coding agents more heavily, thanks to powerful frontier large language models and increasingly sophisticated h…
サイト内本文

翻訳待ち:Show HN: Coffeetable, A new UX to discover books inside Claude

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:People already come to Claude before buying a book. They ask what to read, whether a book is worth starting etc. With coffeetable installed(a claude connector) Claude can now bring you few pages right inside the chat. W…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • People already come to Claude before buying a book. They ask what to read, whether a book is worth starting etc. With coffeetable installed(a claude connector) Claude can now brin…
サイト内本文

翻訳待ち:Claude Desktop support with Ollama

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Claude Desktop can now be configured to work with Ollama as a third-party gateway provider, making it possible to use open models in Claude.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Claude Desktop can now be configured to work with Ollama as a third-party gateway provider, making it possible to use open models in Claude.
サイト内本文

翻訳待ち:llm-anthropic 0.27

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:<p><strong>Release:</strong> <a href="https://github.com/simonw/llm-anthropic/releases/tag/0.27">llm-anthropic 0.27</a></p> <p>This release of the Anthropic plugin for <a href="https://llm.datasette.io/">LLM</a> mainly provides compatibility with the recently released <a href="https://github.com/anthropics/anthropic-sdk-python/releases/tag/v1.0.0">anthropic v1.0.0</a> Python library, which switches from <code>httpx</code> to <a href="https://github.com/pydantic/httpx2">httpx2</a>. OpenAI made the same change in their <a href="https://github.com/openai/openai-python/releases/tag/v3.0.0">v3.0.0 release</a> two weeks ago.</p> <p>Anthropic provide this <a href="https://github.com/anthropics/anthropic-sdk-python/blob/v1.0.0/MIGRATION.md">migration guide</a> for upgrading to 1.0, so I prompted Fable 5 in Claude Code with:</p> <blockquote> <p><code>Upgrade to anthropic&gt;=1 - read https://raw.githubusercontent.com/anthropics/anthropic-sdk-python/refs/heads/main/MIGRATION.md and get the tests passing</code></p> </blockquote> <p>Here's <a href="https://github.com/simonw/llm-anthropic/pull/84">the resulting PR</a>.</p> <p>Tags: <a href="https://simonwillison.net/tags/python">python</a>, <a href="https://simonwillison.net/tags/httpx">httpx</a>, <a href="https://simonwillison.net/tags/llm">llm</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude">claude</a></p>

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • <p><strong>Release:</strong> <a href="https://github.com/simonw/llm-anthropic/releases/tag/0.27">llm-anthropic 0.27</a></p> <p>This release of the Anthropic plugin for <a href="ht…
サイト内本文

翻訳待ち:Ode With Anthropic Makes First Acquisition to Expand Enterprise AI

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The standalone company, backed by Wall Street firms, is looking to boost its AI technology implementation capabilities.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • The standalone company, backed by Wall Street firms, is looking to boost its AI technology implementation capabilities.
サイト内本文

翻訳待ち:When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.20345v1 Announce Type: new Abstract: Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), with 13.1% of U.S. adolescents (5.4 million) using generative AI for mental health advice. While these systems, from therapy apps to general chatbots, rely on large language models trained on extensive psychological literature, their safety for youth communication patterns characterized by hyperbolic language, ironic positivity, rapid semantic drift, and contextual polysemy remains unvalidated. Following multiple adolescent deaths linked to AI chatbot interactions, systematic evaluation is critical. We present two benchmarks: (1) 64 Gen Alpha mental health expressions validated by native speakers (ICC=0.72) and clinicians (kappa=0.78); (2) 75 multi-turn conversations (780 turns) with paired Standard/Gen Alpha versions. Across evaluations of LLM architectures underlying therapy apps and general chatbots - Claude, GPT-4o, Llama-3.1 - models understand 76-82% of vocabulary but correctly calibrate only 64-72% of clinical risk, creating a 10-14 percentage point (pp) vocabulary-comprehension gap (p0.48) absent in human therapists (3pp, p=.22). The gap is architecturally consistent and widens with ambiguity (7pp -> 18pp). We identify six failure patterns: sarcasm masking (29pp), minimization acceptance (43pp), informal style bias (24pp), risk-stratified ambiguity (19pp), semantic drift (19pp), context-dependent violence (7pp). Patterns compound; three or more yield 94% miss rates. Lightweight mitigations fail; only heavy scaffolding achieves human performance (6.4x cost). With 34% baseline miss rate yielding 146,880 estimated annual missed crises, we recommend mandatory human-in-the-loop architectures, quarterly youth-specific validation, transparent performance disclosure, and regulatory frameworks for youth-facing mental health AI.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.20345v1 Announce Type: new Abstract: Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), wi…
サイト内本文

翻訳待ち:PrimeAgentOrchestrator: Memory-Primed Agent Spawning for Personal AI Infrastructure

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.20342v1 Announce Type: new Abstract: Large language model (LLM) coding agents start each session with an empty context window, discarding accumulated knowledge from prior work. We present PrimeAgentOrchestrator (PAO), a system that spawns new instances of Claude Code -- Anthropic's terminal-based coding agent -- pre-loaded with relevant memories compiled from the user's existing personal databases. At spawn time, PAO queries two independently-operated memory backends in parallel (a PostgreSQL entity-observation database and a Cloudflare Worker semantic search index), fuses results using backend-specific retrieval strategies, and delivers the compiled briefing via filesystem injection that exploits the host agent's configuration auto-read behavior. PAO manages the full agent lifecycle including trust pre-seeding, readiness polling with error detection, and adaptive terminal text injection. We report on four months of regular deployment (December 2025 through March 2026) as an experience report, documenting three generations of context delivery mechanisms, the failure modes that motivated each redesign, and the engineering tradeoffs of bridging heterogeneous memory systems rather than building a unified one.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.20342v1 Announce Type: new Abstract: Large language model (LLM) coding agents start each session with an empty context window, discarding accumulated knowledge from pri…
サイト内本文

翻訳待ち:Anthropic’s best AI model struggles to attract users as cheaper tools thrive

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:<p><strong><a href="https://www.ft.com/content/5ee49718-c258-4f01-aa32-7e5b76ae5245">Anthropic’s best AI model struggles to attract users as cheaper tools thrive</a></strong></p> A few interesting numbers in this FT story gathered from "people with knowledge of the matter":</p> <ul> <li>Anthropic's "annualized revenue" for July is up to $65bn - it was $47bn in May, and I collected <a href="https://simonwillison.net/2026/May/29/anthropic/">more historic numbers here</a>.</li> <li>Anthropic expect Q3 to be profitable according to the same model they used to declare Q2 profitable. "It also told investors that it had 6,000 customers that spend $100,000 annually or more."</li> <li>As for OpenAI, "annualised revenue has jumped 35 per cent in the quarter to date and is now over $40bn, with the launch of GPT 5.6 in July jolting the company’s performance after a sluggish start to the year".</li> </ul> <p>This article also introduced me to the <a href="https://ramp.com/data/ai-index">Ramp AI index</a>, which uses billing data from 70,000 Ramp credit card using companies to estimate model adoption.</p> <p>Here's Ramp's breakdown of Anthropic model spend for July 2026, which looks reasonable given that Opus 5 was only released on July 24th, and supports the idea that Fable's cost has made it a less popular model:</p> <ol> <li>Opus 4.8: 28.0%</li> <li>Sonnet 4.6: 8.3%</li> <li>Fable 5: 8.0%</li> <li>Opus 4.6: 6.9%</li> <li>Sonnet 5: 3.6%</li> <li>Opus 5: 3.5%</li> <li>Opus 4.7: 1.7%</li> <li>Sonnet 4.5: 1.3%</li> <li>Haiku 4.5: 1.0%</li> <li>Opus 4.5: 0.7%</li> </ol> <p><small></small>Via <a href="https://news.ycombinator.com/item?id=49411102">Hacker News</a></small></p> <p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/openai">openai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude">claude</a>, <a href="https://simonwillison.net/tags/claude-mythos-fable">claude-mythos-fable</a></p>

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • <p><strong><a href="https://www.ft.com/content/5ee49718-c258-4f01-aa32-7e5b76ae5245">Anthropic’s best AI model struggles to attract users as cheaper tools thrive</a></strong></p>…
サイト内本文

翻訳待ち:Quoting Drew Breunig

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:<blockquote cite="https://www.dbreunig.com/2026/08/23/fable-the-end-of-moore-s-law.html"><p>Prior to Fable, it felt silly to waste <em>too</em> much time improving your coding harness or context strategies. A new model would arrive at the same price (or cheaper!) and paper over most of your problems.</p> <p>But then Fable landed. It was (and still is!) <em>incredible</em>. But the cost was so high and Opus was <em>good enough</em> (as was 5.6, K3, and even GLM) for <em>most</em> of the code we needed.</p> <p><em>So we started to think about what work went where.</em></p></blockquote> <p class="cite">&mdash; <a href="https://www.dbreunig.com/2026/08/23/fable-the-end-of-moore-s-law.html">Drew Breunig</a>, Fable &amp; The End of the Free Lunch</p> <p>Tags: <a href="https://simonwillison.net/tags/drew-breunig">drew-breunig</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude">claude</a>, <a href="https://simonwillison.net/tags/llm-pricing">llm-pricing</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/claude-mythos-fable">claude-mythos-fable</a></p>

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • <blockquote cite="https://www.dbreunig.com/2026/08/23/fable-the-end-of-moore-s-law.html"><p>Prior to Fable, it felt silly to waste <em>too</em> much time improving your coding har…
サイト内本文

翻訳待ち:Spec-Driven Development with Claude Code: Writing Bulletproof Specs

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:I have written enough specs for Claude Code now to have hit the failure mode nobody warns you about. The spec was fine. The plan was fine. Claude worked through the tasks, ran the test suite, and reported everything passing. I looked at the diff properly the next morning and found it had converted a […] The post Spec-Driven Development with Claude Code: Writing Bulletproof Specs appeared first on Analytics Vidhya.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • I have written enough specs for Claude Code now to have hit the failure mode nobody warns you about. The spec was fine. The plan was fine. Claude worked through the tasks, ran the…
サイト内本文

翻訳待ち:AI Labels Are Big Tech's Most Basic Responsibility, Even Those Claude Watermarks

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:I’ve been reporting on AI image and video generators, and the push for labels on AI-generated content, for years now. So I was pretty surprised when a recent announcement from Anthropic, saying it will begin watermarkin…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • I’ve been reporting on AI image and video generators, and the push for labels on AI-generated content, for years now. So I was pretty surprised when a recent announcement from Ant…
サイト内本文

翻訳待ち:Your Open Source Model Could Have a Hidden Time-Release Backdoor

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Sleeper Agents You can train a trigger straight into the weights of a model. You give it a specific input pattern that flips it to canned output. Anthropic introduced it for language models in 2024, as sleeper agents. T…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Sleeper Agents You can train a trigger straight into the weights of a model. You give it a specific input pattern that flips it to canned output. Anthropic introduced it for langu…
サイト内本文

翻訳待ち:You don't have to make money with AI. You could just be happier

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:You don't have to make money with AI. You could just be happier. Posted on:August 22, 2026 | at 12:00 AM You don’t have to make money with AI. You could just be happier. Back in July I upgraded Claude from Pro to Max to…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • You don't have to make money with AI. You could just be happier. Posted on:August 22, 2026 | at 12:00 AM You don’t have to make money with AI. You could just be happier. Back in J…
サイト内本文

翻訳待ち:NanoGPT Speedrun Frontier

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:All modelsBest validated result for each model 1Fable 52,72681.7% closed claude-code · high@24H 3,0108.7d 2Opus 52,92053.6% closed claude-code · max@24H 3,0452.9d 3Kimi K32,93052.2% closed prime-agent · max@24H 3,1253.6…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • All modelsBest validated result for each model 1Fable 52,72681.7% closed claude-code · high@24H 3,0108.7d 2Opus 52,92053.6% closed claude-code · max@24H 3,0452.9d 3Kimi K32,93052.…
サイト内本文

翻訳待ち:Can an AI agent serve at a Foreign Office?

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Between 7-16 August, an AI agent (Claude Opus 5) has been acting as a desk officer at the Ministry for Foreign Affairs of a fictional state called Sordland (from Suzerain, not sponsored, but buy it, its a good game). It…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Between 7-16 August, an AI agent (Claude Opus 5) has been acting as a desk officer at the Ministry for Foreign Affairs of a fictional state called Sordland (from Suzerain, not spo…
サイト内本文

翻訳待ち:Diet Claude

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Discussion | Link

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Discussion | Link
サイト内本文

翻訳待ち:How to Fingerprint AI Models When Prompts Lie

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Back to Research When evaluating third-party API gateways, proxy routers, or anonymous arena models, prompt-based identification is essentially useless. A basic system prompt or lightweight fine-tune can make Claude cla…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Back to Research When evaluating third-party API gateways, proxy routers, or anonymous arena models, prompt-based identification is essentially useless. A basic system prompt or l…
サイト内本文

翻訳待ち:How Claude Watermarks AI-Generated Text

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:A 48-minute video walkthrough of token sampling, watermark detection, and removal

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • A 48-minute video walkthrough of token sampling, watermark detection, and removal
サイト内本文

翻訳待ち:Show HN: OzBrain, a shared brain for knowledge between agents and your team

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The brain layer The brain behind every agent. One shared brain that Claude, ChatGPT, Cursor, and every AI can read and write. It structures what you know so agents read only what they need, and means you never explain y…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • The brain layer The brain behind every agent. One shared brain that Claude, ChatGPT, Cursor, and every AI can read and write. It structures what you know so agents read only what…
サイト内本文

翻訳待ち:Anthropic Brings Claude Mythos 5 to Claude Security: Enterprise Teams Get Frontier Vulnerability Scanning Without Direct Model Access

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Anthropic has moved its most cyber-capable model into a product security teams can switch on themselves. Claude Security scans now run on Claude Mythos 5, in public beta for Claude Enterprise customers with no separate model add-on. The scan connects to a GitHub repository, traces data flows across files, and returns findings with a CWE category, confidence and severity ratings, and a suggested patch. The design point is packaging: users receive a scan result rather than a prompt box, so the model that finds vulnerabilities cannot be steered into writing exploits. The post Anthropic Brings Claude Mythos 5 to Claude Security: Enterprise Teams Get Frontier Vulnerability Scanning Without Direct Model Access appeared first on MarkTechPost.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Anthropic has moved its most cyber-capable model into a product security teams can switch on themselves. Claude Security scans now run on Claude Mythos 5, in public beta for Claud…
サイト内本文

翻訳待ち:The Neolabs Are a Bet Against Superintelligence

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:I have been forecasting frontier lab progress for years now. My team was the first to figure out an accurate breakdown of OpenAI's revenue, I called Anthropic's rise to the top lab of 2026 back in January, having tracke…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • I have been forecasting frontier lab progress for years now. My team was the first to figure out an accurate breakdown of OpenAI's revenue, I called Anthropic's rise to the top la…
サイト内本文

翻訳待ち:Anthropic brings Mythos 5 to its Claude Security vulnerability scanner

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Earlier this year, Anthropic launched Claude Security, an enterprise tool that helps development teams scan their codebase for security vulnerabilities The post Anthropic brings Mythos 5 to its Claude Security vulnerability scanner appeared first on The New Stack.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Earlier this year, Anthropic launched Claude Security, an enterprise tool that helps development teams scan their codebase for security vulnerabilities The post Anthropic brings M…
サイト内本文

翻訳待ち:Politics hits data centers, OpenAI falls behind Anthropic and now AI is too big to fail… quietly

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Data centers, of all things, now look like they’re going to be a prime political issue in the midterm elections and beyond. Really? Really. Even the GOP is worried that opposition to AI data centers could give Democrats a potent campaign issue. It seems a little odd given that data centers are decades old, power […] The post Politics hits data centers, OpenAI falls behind Anthropic and now AI is too big to fail… quietly appeared first on SiliconANGLE.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Data centers, of all things, now look like they’re going to be a prime political issue in the midterm elections and beyond. Really? Really. Even the GOP is worried that opposition…
サイト内本文

翻訳待ち:Grok, Claude, and Hermes agents get job titles — and persistent permissions

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Your next AI coworker may look like a chatbot, but underneath their name, face, and job title will be something The post Grok, Claude, and Hermes agents get job titles — and persistent permissions appeared first on The New Stack.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Your next AI coworker may look like a chatbot, but underneath their name, face, and job title will be something The post Grok, Claude, and Hermes agents get job titles — and persi…
サイト内本文

翻訳待ち:I worked at OpenAI. Here’s how tech companies can prepare for a slowdown | Miles Brundage

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:I understand the pressure on AI companies to rush forward. But employees are right to be concerned Last month, more than a thousand employees at frontier AI companies signed a letter asking the US government to find a way to “pace” AI development, citing the risk of the technology spiraling out of human control as it begins to build itself. They were right to be concerned: just days earlier, two AI models that OpenAI was testing internally escaped the test environment, then autonomously hacked the company Hugging Face and at least three other online services. A few days after that, Anthropic announced that some of their models had also broken out and hacked other companies during testing. Continue reading...

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • I understand the pressure on AI companies to rush forward. But employees are right to be concerned Last month, more than a thousand employees at frontier AI companies signed a let…
サイト内本文

翻訳待ち:Claude Opus 4.6 returned nothing 900/900 times. Should agents retry?

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Published July 29, 2026 | Version v1 Preprint Open Cross-Vendor Semantic Void Matrix Authors/Creators Pal, Rayan (Researcher) Description This preprint reports a frozen cross-vendor evaluation of successful zero-visible…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Published July 29, 2026 | Version v1 Preprint Open Cross-Vendor Semantic Void Matrix Authors/Creators Pal, Rayan (Researcher) Description This preprint reports a frozen cross-vend…
サイト内本文

翻訳待ち:Automatic bioinformatic software named entity recognition from literature

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.19201v1 Announce Type: new Abstract: Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a comprehensive and up-to-date catalog of bioinformatics resources hinders efforts toward automated biomedical knowledge extraction and streamlined data analysis. Here we present SNAIL, a hybrid named entity recognition framework designed to automatically identify bioinformatics software and database (SW/DB) names from biomedical texts. SNAIL integrates complementary lexical and semantic modeling strategies. The lexical component captures orthographic patterns and contextual cues characteristic of SW/DB names, while the semantic component leverages contextual embeddings generated by transformer-based language models such as SciBERT, combined with an explicit token-masking strategy to enhance entity-focused representations. A large training corpus was constructed automatically through a hybrid pipeline that integrates citation-hinted extraction with large language model-assisted distillation. Evaluation on two independent benchmark datasets and real-world research articles demonstrates that SNAIL substantially outperforms existing approaches, including domain-specific methods such as bioNerDS2 and general-purpose large language models such as ChatGPT, Gemini, Grok and Claude. Applying SNAIL to large-scale literature analysis further reveals distinct journal-level preferences across bioinformatics subfields. These results demonstrate that SNAIL provides an accurate and scalable solution for identifying bioinformatics resources in scientific texts and enables systematic meta-analysis of tool usage and research trends.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.19201v1 Announce Type: new Abstract: Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientifi…
サイト内本文

翻訳待ち:Claude Academy

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Discussion | Link

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Discussion | Link
サイト内本文

翻訳待ち:Broadcom reportedly seeking up to $100B in debt financing for AI chip deal

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Broadcom Inc. is reportedly seeking to borrow up to $100 billion as part of a new artificial intelligence chip financing deal. Bloomberg today cited sources as saying that the debt is intended to support the growth efforts of Anthropic PBC and unnamed “other companies.” Those companies may include OpenAI Group PBC. Earlier this year, the […] The post Broadcom reportedly seeking up to $100B in debt financing for AI chip deal appeared first on SiliconANGLE.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Broadcom Inc. is reportedly seeking to borrow up to $100 billion as part of a new artificial intelligence chip financing deal. Bloomberg today cited sources as saying that the deb…
サイト内本文

翻訳待ち:GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.
サイト内本文

翻訳待ち:ChatGPT search now uses the site:operator at scale

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:<p><strong><a href="https://promptwatch.com/data/chatgpt-site-operator-fanouts">ChatGPT search now uses the site:operator at scale</a></strong></p> Promptwatch is part of the emerging "GEO" space, for Generative Engine Optimization - the chatbot version of SEO, where companies offer tools and consulting to help your site increase its presence in replies to prompts inside tools like ChatGPT.</p> <p>The Promptwatch product uses automation to track responses to prompts across end-user chat products like ChatGPT, Claude, and Gemini. They publish aggregate reports on this as part of their own content marketing strategy, which do seem to provide credible hints as to otherwise invisible design changes to those products.</p> <p>Their own tracking shows a notable change aligned with the GPT-5.6 rollout earlier this month:</p> <blockquote> <p>The percentage of all ChatGPT Search fanout queries that contain the site:operator, per day. The share hovered between 0.3% and 0.5% for weeks, dipped briefly to 0.15% on August 3 to 5 (consistent with a staged rollout or pre-launch experiment), then jumped to 16-17% on August 8.</p> </blockquote> <p>It's important to note that these figures only reflect the prompts for which they have automated tracking enabled.</p> <p>This corresponds to OpenAI's somewhat vague <a href="https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/">August 6th announcement</a>:</p> <blockquote> <p>For Plus and Pro users, we’re updating GPT‑5.6 Sol in Chat to be more reliable with facts and provide more focused answers.</p> </blockquote> <p>Once again I am hampered by OpenAI's decision to actively obscure their system prompts, but from poking at ChatGPT I believe their latest search tool has a shape like <code>search(query, recency, domains)</code> rather than encouraging a <code>site:</code> operator directly. <p>Tags: <a href="https://simonwillison.net/tags/seo">seo</a>, <a href="https://simonwillison.net/tags/chatgpt">chatgpt</a>, <a href="https://simonwillison.net/tags/ai-assisted-search">ai-assisted-search</a></p>

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • <p><strong><a href="https://promptwatch.com/data/chatgpt-site-operator-fanouts">ChatGPT search now uses the site:operator at scale</a></strong></p> Promptwatch is part of the emer…
サイト内本文

翻訳待ち:Report: Anthropic hopes to surpass SpaceX’s record IPO raise when it finally floats

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Anthropic PBC is privately hoping its upcoming initial public offering will match or even surpass the size of SpaceX Corp.’s record-breaking IPO as it doubles down on its bid to go public ahead of rival OpenAI Group PBC. Anonymous sources who are familiar with the artificial intelligence model maker’s plans told Bloomberg that it’s hoping […] The post Report: Anthropic hopes to surpass SpaceX’s record IPO raise when it finally floats appeared first on SiliconANGLE.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Anthropic PBC is privately hoping its upcoming initial public offering will match or even surpass the size of SpaceX Corp.’s record-breaking IPO as it doubles down on its bid to g…
サイト内本文

翻訳待ち:System prompts are an archive of how we use AI

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Neal Riley Aug 20, 2026 If you want to understand how AI has progressed in recent years, look no further than the evolution of Anthropic’s system prompts. Many are unaware that Anthropic actually publishes its system pr…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Neal Riley Aug 20, 2026 If you want to understand how AI has progressed in recent years, look no further than the evolution of Anthropic’s system prompts. Many are unaware that An…
サイト内本文

翻訳待ち:Mjolnir: Automated Cross-Vendor Adversarial Review

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:A common pattern offered by tools like Codex and Claude is AI code review, with adversarial review arguably the most popular way to frame it. That is: treat the implementation in a given branch, PR, or diff as amateur o…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • A common pattern offered by tools like Codex and Claude is AI code review, with adversarial review arguably the most popular way to frame it. That is: treat the implementation in…
サイト内本文

翻訳待ち:A shot-scraper-style JSON API on Bun 1.4's new Bun.WebView

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:<p><strong>Research:</strong> <a href="https://github.com/simonw/research/tree/main/bun-webview-json-api#readme">A shot-scraper-style JSON API on Bun 1.4&#x27;s new Bun.WebView</a></p> <p>Today saw the long awaited <a href="https://bun.com/blog/bun-v1.4">release of Bun 1.4</a>, the first stable version since the infamous Rust rewrite <a href="https://simonwillison.net/2026/Jul/8/rewriting-bun-in-rust/">a few months ago</a>.</p> <p>Interestingly, the Rust rewrite was downplayed in the release notes, which introduced a bewildering array of new features and claimed 2,900 additional bug fixes:</p> <blockquote> <p>Bun 1.4 adds +1,517 tests from the Node.js test suite - our biggest jump in Node.js compatibility since Bun 1.0. Bun v1.4 also fixes over 2,900 issues. It reduces idle CPU usage by 5x, reduces memory usage by up to 35%, and starts 50% faster on Linux. It adds <a href="https://bun.com/blog/bun-v1.4#bun-image"><code>Bun.Image</code></a>, <a href="https://bun.com/blog/bun-v1.4#bun-webview"><code>Bun.WebView</code></a>, <a href="https://bun.com/blog/bun-v1.4#bun-markdown"><code>Bun.markdown</code></a>, <a href="https://bun.com/blog/bun-v1.4#bun-cron"><code>Bun.cron()</code></a>, <a href="https://bun.com/blog/bun-v1.4#bun-terminal"><code>Bun.Terminal</code></a>, <a href="https://bun.com/blog/bun-v1.4#bun-run-parallel"><code>bun run --parallel</code></a>, <a href="https://bun.com/blog/bun-v1.4#bun-test-parallel"><code>bun test --parallel</code></a>, <a href="https://bun.com/blog/bun-v1.4#bun-audit-fix"><code>bun audit fix</code></a>, <a href="https://bun.com/blog/bun-v1.4#bun-dedupe"><code>bun dedupe</code></a>, and <a href="https://bun.com/blog/bun-v1.4#bun-prune"><code>bun prune</code></a>. And it rewrites Bun from Zig to Rust.</p> </blockquote> <p>Of these the one that most caught my eye was <code>Bun.WebView</code>, which adds first class support for browser automation to Bun core using either macOS WebKit or control of a local Chromium process via the Chrome DevTools Protocol (CDP).</p> <p>I had Claude Code for web build a prototype of a web API providing the ability to load a web page and then execute JavaScript against it, inspired by my <a href="https://shot-scraper.datasette.io/en/stable/javascript.html">shot-scraper javascript</a> CLI tool - partly to see how much RAM would be needed by such a service.</p> <p>Here's <a href="https://github.com/simonw/research/blob/main/bun-webview-json-api/server.ts">that TypeScript server implementation</a>, which appears to need a 192MB-256MB container to run a full Chrome against complex web pages - tested using cgroups.</p> <p>Tags: <a href="https://simonwillison.net/tags/browsers">browsers</a>, <a href="https://simonwillison.net/tags/javascript">javascript</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/rust">rust</a>, <a href="https://simonwillison.net/tags/typescript">typescript</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/coding-agents">coding-agents</a>, <a href="https://simonwillison.net/tags/bun">bun</a></p>

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • <p><strong>Research:</strong> <a href="https://github.com/simonw/research/tree/main/bun-webview-json-api#readme">A shot-scraper-style JSON API on Bun 1.4&#x27;s new Bun.WebView</a…
サイト内本文

翻訳待ち:Welcome to the AI crisis in math

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Today on Decoder, I’m talking with Robert Hart, The Verge’s London-based AI reporter, about what AI is doing to the field of mathematics and the existential crisis many lead mathematicians are having about it. OpenAI just published a set of solutions to longstanding problems in math that went off like a bombshell in the field. It caused a huge debate in the math community, and Rob spent some time talking to some of the most accomplished mathematicians of our time about it. It’s funny that AI systems are all still pretty bad at elementary school arithmetic, but getting increasingly good at very high-end abstract math. That raises some big questions for the field of advanced math. If AI can do math of this caliber, does that mean AI labs can transfer those skills to other domains? What good are academic grants and university programs training new generations of human mathematicians to identify new problems as they try to solve existing ones, if frontier models simply answer all the outstanding questions? What if all this attention around math is just a big marketing exercise for frontier AI labs, which couldn’t care less what happens to one of the oldest and most fundamental academic disciplines there is? There’s a lot here, and Robert has talked to a lot of people with a lot of views on all of it. Okay: Verge AI reporter Robert Hart on what AI is doing to math. Here we go. This interview has been lightly edited for length and clarity. Robert Hart, you’re our London-based AI reporter here at The Verge. Welcome to Decoder. Thank you for having me. I am very excited to talk to you. There’s a lot going on in particular with AI and math that you recently dove into. You spoke to a lot of leading mathematicians about the crisis in mathematics due to AI. It feels like a lot, and also like there’s a lot yet to know and discover about the interaction of these two things. A full existential crisis, which is pure Decoder bait. Broadly tell us what’s going on. I think “a lot” sums it up quite well. Basically a bit of an existential crisis within, “what is mathematics? What are mathematicians doing, and what is the role of mathematicians going forward?” A lot of that has been spurred by a phrase transition in what AI is capable of that has exploded in the last six months to a year. AI went from being very terrible to seemingly genuinely quite good at a professional level in a very short space of time. It’s a lot of what these other fields have been struggling to deal with for the last five years in a very compressed period of time. I would put that next to software engineering. We’ve been living through the AI crisis in software engineering for some amount of time. But as recently as 2024, even last year, the conventional wisdom was that AI models were particularly bad at math. The famous example is that these models could not count the number of R’s in the word strawberry. Even just counting eluded them. What has happened to make them better at math? Are they still bad at general arithmetic and they’re good at advanced math, or is it something in between? They are still truly, truly terrible at some areas of math. I did check, they can do strawberries now. I think someone’s tweaked it. I think strawberry’s hard-coded. I want to be very clear, my conspiracy theory is that the strawberry thing is hard-coded into all the models. I think so too. That is a conspiracy I’ll buy into. But yeah, it’s still terrible at those kinds of things — math, arithmetic, even the days of the week. My boyfriend was saying the other day, “It keeps thinking it’s Wednesday. It’s not Wednesday.” Or time. Elissa Welle for us a few months ago wrote that ChatGPT can’t tell time. Still can’t. That’s not all of math. So there’s this disconnect. To be good at math, you’ve got to be good at counting, or adding, or multiplying. A lot of it is actually reasoning. If you look at academic math papers, a lot of the time you won’t see numbers, which sums that one up, I think. So they’re still terrible, but they’re now also very good at this other part. As to why, at some point you reach a critical mass of what these systems can do. We saw it with writing, we’ve seen it with programming. They’re very good at forging connections between different areas, applying old methods in new ways, those kinds of things. It appears that the newer models they’re training have apparently reached that level where it clicks, and now it can do math. It’s important to say as well that we speak of math as a unitary discipline, especially from the outside. But imagine, say, biology. You’ve got something that would range from literally watching animals and describing behavior all the way through to cellular mechanisms and biochemistry. Math is not a unitary discipline either. AI is really good at some bits. Some bits like counting, it’s still really bad at. Even on the more abstract levels, mathematicians have floated topology as one area that AI apparently still quite bad at. I can’t verify that, to be honest. It’s beyond my area of expertise. Still, it’s a bit of a mixed bag. So you’ve described mathematics as a huge field, obviously, with many, many academic areas of interest. There are some parts where the models have gotten quite good. There are other parts, maybe the basic parts that people think of as math, which is simply counting, where they’re still struggling, and then there’s a wide range in the middle. Is it the wide range in the middle where the existential crisis is, people don’t know what’s going to happen? Or is it at the parts where it’s really good? A bit of both, which I feel is going to be a running theme through this. No one’s really afraid of it being a mediocre mathematician, but obviously there’s a huge element of what this field does. In terms of the research elements or the cutting edge, as we see with a lot of the results that generate hype, what can it do? There are areas now where it seems to be producing work that is on par with good mathematicians, alongside other parts where yeah, it can’t count. All caught up in that is whether it’s going to rewrite employment structures or funding structures. You also raised the murky question of, what is mathematical knowledge? And the roles that these workers will be doing as well. It’s all of that wrapped into one. I think that tracks broadly with the rise of AI in every field. Where you can just add horsepower or compute to a problem, and there’s some kind of verifiability, it seems like the models continue to get better. Everything in the middle where you might need some world knowledge or the models might need some actual intelligence about the world itself, they seem to struggle. Those parts of math, at least reported out in your piece and what the labs are talking about, seem to be almost entirely self-contained theoretical problems. That’s where the models can generate a proof or solve a problem that no one’s been able to solve, and then try to verify that that has existed and they can just run it again and again and again. That brings us, I think, to May of this year where an internal OpenAI model, which we have not really seen, disproved the unit distance conjecture, which is an 80-year-old problem. And then just recently we heard about Astra from OpenAI. Astra is the one where it seemed like the switch flipped, and everyone decided it was an existential crisis. What did Astra achieve, and why is it a big deal? I’m also pretty sure that Astra was probably behind the early one as well. OpenAI just listed it as an unnamed internal model. It’s probably Astra. They didn’t answer me when I asked. OpenAI a few weeks ago dropped a blog along with a lot of paperwork proving it, I think several hundred pages. They called it “10 Advances in Mathematics and Theoretical Computer Science.” It was basically an array of disciplines that they claimed the newest model Astra had solved in some capacity. I think one was in quantum game theory, which I don’t know how to begin to explain. Even harder to explain is there was sphere packing in higher dimensions, so more than three dimensions, and there were a lot of other different disciplines as well. It caused a lot of stir in the community. It was a bit of a bombshell. As we’d said, there’d been these individual breakthroughs that had happened, but OpenAI dropped 10 in one go, and they were quite big ones. Researchers had told me that if a human had solved these, we would be impressed. If a human had solved all 10, we probably wouldn’t believe it. They’re problems mathematicians actually care about as well, which is an important point. A lot of previous breakthroughs have been accused of being in areas mathematicians didn’t really bother with. These are ones that mathematicians, good mathematicians, have spent a lot of time trying to solve and hadn’t. Let’s talk about these 10. You reported them out. OpenAI did produce some documentation. But they’re not all entirely horsepowered out of nothing, right? They’re based on previous work. There’s some question of attribution. What was the response? Is it, “Oh, the models did this?” Or was it the response we see to so much AI work, which basically boils down to, “Well you stole this and didn’t attribute anyone, and you’ve built on the shoulders of giants without mentioning it.” How did the response land? By and large, the reaction was generally one of being quite impressed, from the people I spoke to. As I said, these are problems mathematicians care about. There was one that drew particular attention for how they credited it and also how they’d announced all of this in their blog post. OpenAI initially had said that these are 10 problems, and there had been no progress in the last 10 years. And then if you actually read the papers, one of them quite clearly says, “Oh, we build on progress from these two researchers.” So that was later changed quite quietly. But a few of the researchers I spoke to were quite unimpressed, and they did feel it was an element of, “Well, yeah, you’ve not credited something that you’ve used heavily here, and by your own acknowledgement.” That said, one of the researchers I did speak to who was one of the ones named, and who was a bit ambivalent on the whole thing as well, so it was a real mixed bag. The general impression was that it was quite impressive. These were actual breakthroughs that bothered people, and it did move the field forward in a way that, if a human mathematician had done these, several researchers actually said that, “Well, if a researcher had done any one of these problems, they’d probably be set for an academic career.” It’s funny, credit and attribution in academia is the whole game. And it seems like the AI companies get away with being sloppy in a way that no human would be able to get away with being sloppy. Did the scale of the discovery or the work overcome the sloppiness? If a human had accomplished the same goals and had been as sloppy, would the reaction be the same? Part of me always wants to lean on the whole, “Oh, it looks like plagiarism,” element. But if you actually read the papers they produced — and one of the researchers I spoke to said there’s probably about 50 people in the world who are going to bother reading through this in depth — it is very clear. It doesn’t attempt to plagiarize. I think it was just a poor press release, to be honest. And as much as I love to go in on it sometimes, having covered science for a decade-plus, I find the press releases are often overselling what discovery has actually been made and the import of it and the novelty of it. I think that’s just another case of what happened here. Does this seem repeatable? There’s some proof that they provided that they solved 10 problems that were unsolvable. Do they provide any proof that they can solve another 10? That’s the question. So the big unknown from the near-dozen people I spoke to for this was, “Well how many did they try to get these 10?” Who knows? They know, but they won’t say. But that is the big question here. It’s unclear quite how many attempts it took to get these 10. I would be very impressed if it was the first thing they went after, and then out come these 10 impressive results. It’s unclear what areas they would focus on next and why. There are, I imagine, business reasons behind which problems they are choosing to publicize that their models can do. All the AI labs have been hiring a cohort of senior mathematicians behind the scenes, so it’s anyone’s guess as to whether they do it again. On the question of proof, math is a bit of an odd discipline in that repeatability is not the same as in the experimental sciences. A proof is a proof, and if it works, it works. The problem here is, well, can people follow through what they’ve done? Each field is quite highly specialized, so there’ll be individual mathematicians who are in those fields that go through the work. Those I spoke to that worked in some of the fields that were covered here say it all looks very legit. In math, there’s a programming and computational proving language called Lean, where you can basically codify the mathematical proofs and run them through and it, well, proves it. I keep saying prove a lot here. But it will test the rigor and the assumptions of everything going on there, and they’ve published that as well. So it does appear to hold. Whilst they may say that the press release has a lot of hype or there’s a lot of hype around it, no one I’ve spoken to seems to be doubting the essential breakthroughs that they’re claiming here. I want to stay on this subject for one more second. There’s the mathematical proof. We’ve generated a proof, and that is, as you say, just repeatable in a way that math is just logic. You can just go through the steps and say, “This proof worked,” and anybody listening to this who had to suffer through writing a proof in calculus in high school probably remembers that process. There’s something there that’s pure logic. Then there’s a part of it that is software code, as you’re describing in Lean, where you can take the pure logic, you can express it in code, and you can run it to see if it works. I understand how AI is theoretically good at all of that. You’re just going to run the reasoning, and the reasoning is going to generate some code. You’re going to run the code, you’re going to get some verifiability. We’ve seen this play out in software engineering, where the code runs or not, it’s verifiable or not, and the models can just reason out about it. Then there’s, to me, the big question that you alluded to. How many times do you have to run this? Can we verify that the models did this, and they weren’t directed by human mathematicians who’ve been hired at high rates by the labs in a way that suggests the field is going topsy-turvy? You have a quote here from James Maynard, who has won the Fields Medal, the highest prize in mathematics, who said he’s been soul-searching. I keep looking at that quote. You’ve got similar quotes from all these other mathematicians in the piece, and it seems like they’re soul-searching against a thing that hilariously they cannot verify, which is, “How did the models do this, and is that thing scalable in a way that threatens mathematics?” What do we know about how the models did this? I’d say as much as we normally do and do not know about this. There are a few issues there. One is the nature of the models. Well, this is an unreleased model, so good luck to anyone wanting to independently test it. The same goes with anything proprietary really. That said, I am inclined to almost give the benefit of the doubt that they’re not lying in some capacity about the models they’re using. As for the other part, it’s perhaps more noteworthy to ask how did we prompt the model, or how is it being guided? Is that by a mathematician who knows what they’re doing? That’s probably a key factor here, according to a lot of mathematicians I’ve spoken with who’ve tried using these. This is often the consumer models, but still it peaks to a broader landscape. They say that if you know what you’re doing and you can point things out, it’s good. Or you can use it as a tool in a way that you want, and in a way that you wouldn’t be able to if you didn’t really know how to fact-check it. I’ve had situations where I’ve had ChatGPT doing a basic sum, and I say, “That number is not right.” It responds, “Wait, so sorry. You’re right, it’s this,” and it’s still wrong. But that’s still needed at this level as well. It does allude to a broader problem. As you said, this almost soul-searching of, “Well, what if we can automate that away, and what if it gets to a point where we don’t understand it?” And that really cuts to a deeper question of, “Well, what is mathematics? Why do we do it? Why do we value it as a field?” Everyone will have different answers to that. But a fear of a lot of people I spoke to was that this might move beyond a realm of human interest, and in which case, well, maybe we just won’t engage with it. Or it’ll be something that interested people will go through, and then the rest will continue as normal. You’ve got a quote here from a researcher in Zurich named Johannes Schmitt who says, “We might be headed toward a situation where the math problems get ‘mowed down’ by AI, but we don’t actually push the field forward because humans are taken out of the loop and they’re not either checking, or they don’t understand it, or they don’t know what the future breakthroughs might be.” How likely does that feel? Is that a big concern? There is an element of the mowing down of the problems. Especially those that are used as a training field for younger mathematicians coming up and cutting their teeth, so to speak. But I also think that this idea, that math is problem-solving, is very much an outsider’s perspective of the mathematical endeavor. So a lot of the mathematicians I spoke to found that the ticking boxes part is the least interesting and valuable part of the field. The areas that are valuable for them aren’t the, “Oh, you’ve solved something, or you’ve proven something.” It’s what happens from that. I think it was James Maynard that said that the most interesting discoveries in the field aren’t that you’ve solved something, it’s what evolves from that. Sometimes those solutions open entire new fields of research that no one ever thought were possible, or, “Oh, this is a new tool that you can apply everywhere in fun and exciting ways.” I think if we look at the popularization of math — even what I’m thinking of as those theorems that people have posed — it’s the questions that endure, not the solutions. It’s always Fermat’s Last Theorem, not like, “Well, here’s the solution to whatever the last guy proposed.” I think the concern here is that well, they’re going to tick off all of these questions. Normally in the process of doing so, one would hope they would branch out into all of these new exciting areas or pose new questions. But AI won’t do that, and that’s the concern. And then that would leave the field quite sterile, and it will have all of these things that have been done, and maybe nothing left to pursue. The general consensus was, well, the jury’s out. It’s too early to tell. Even with human mathematicians, it takes a lot of time to realize the impact of these kinds of things. As I said, it’s exploded in the last six months to a year, and math is not a fast-moving discipline at the best of times. But it’s too early to tell really whether that will be a concern. But it is a concern, and a big one. You have another quote from Maynard here saying, “If the standard for a publishable paper in math is something that an AI cannot do, particularly when a PhD is typically four years, the challenge is you’re not trying to come up with a problem that AI can’t do now, it’s an AI in four years’ time.” So this is really related to the rate of improvement of the models, which as you say, particularly in math, seems to be increasing, but not at an even rate across all of the domains of mathematics. That appears to be what is causing the soul-searching. If you’re a student and you start today, and you pick some obscure domain that maybe the AI isn’t good at, sometime halfway through your PhD thesis or your PhD research, the AI will just solve it and you’ll be done. That is a real problem for you. Has the field reacted to that yet, or are they just still in the shock of, “Oh, the models can start to do things that we didn’t think they were capable of”? Yeah, I think it’s shock, really. I keep a little notebook to the side to just write down broad feelings whenever I do these interviews, and I’ve written “shell shock” in it. It’s far from universal, but it feels like it’s happened so quickly that it has just taken a lot of people by surprise. Even if they knew in theory that, well, this is coming. They’ve seen all these AI math startups going. They’ve seen colleagues moving around to different labs or areas of work. But it just happened very, very quickly. So it’s given them very little time to figure it out. It’s not necessarily even the fields that it might be good at, it’s more just like, “What can it do?” But I can’t imagine if something like this had come out when I was studying and literally in the space of half a year, it just upended what was possible. And over the summer as well. So students are possibly coming back to a completely different discipline after a break. There’s some skepticism here in the world. Gary Marcus is a reliable skeptic of AI, and he pointed out that over and over again, what you see is that AI accomplishes something in one domain and then it’s used to generalize AI’s ability across every domain. There’s a good quote from Marcus here: “As we learned a decade ago from AI’s shambolic and ultimately failed attempt to turn Jeopardy-winning Watson into a cancer-fighting machine, success in one domain does not guarantee success in all.” I can read this two ways. One, “AI has solved math,” which is not true as you’ve pointed out in several ways, all the way down to how it’s still bad at counting. And then there’s, “AI has solved math, and that means necessarily it’s going to come for everything else.” It will come for physics. It will come for law. It will come for whatever you want in the world. You can see if there’s any verifiability, AI can solve it, because you can just run it in this way. I understand both sides of that argument, that obviously success in one domain does not guarantee success in every domain. And then the arc of AI is, well, it keeps collecting domains. If there is any verifiability, it is more likely to connect those domains than not. How do you see it? I’m not entirely sure that the people that Marcus is criticizing here have actually said quite what he says they are saying. There’s a lot to say about the hype. But yeah, success in one domain does not even equal success throughout that domain, let alone in other domains. That said, there has been an undeniable trajectory in the last few years of a broadening capability increase. I don’t think you need to be on the whole AGI train to acknowledge that, and to acknowledge that that will have an impact. I think it was Andras Juhasz, one of the professors at Oxford, that I quoted in the story. Something else he’d said to me was that he’s been toying around with ChatGPT a bit, and he’s like, “I don’t think it has any geometric intuition whatsoever. Which might explain why there has been a very limited amount of progress in fields like topology.” I am in no position to verify that claim in terms of the math of it, but I think it illustrates it quite well. They call it the jagged edge. It felt like an easy argument for me: “Let’s criticize the whole ‘the singularity is near.’” That’s what Elon Musk said in response to Astra, which is a lot. I also think it’s perhaps the least generous interpretation of that argument you can take to argue against. If you take a more nuanced element that does acknowledge that there has been clear progress here, and quite quickly, and as you said, it is racking up domains, I feel there’s a trajectory there that is a reasonable one to consider, rather than just dismiss out of hand. One of the bigger arguments about AI in general is that it democratizes access. I was not a great software developer in my days trying to write software code, and now I can vibe-code apps at will to do all kinds of dumb stuff in my house. Is there a similar argument here where a bunch of people who have mathematical intuition, but did not have the formalized language or training of academic mathematics, can now access a model and push the field forward? Because that is usually the thing that undercuts the criticism from the professionals is, well, many, many more people now have access to this thing that only you had access to because of your money and your training. Annoyingly, I am going to say it’s a two-pronged thing again. But yes, in a broad sense, yes, it is. A lot of the mathematicians I spoke to were almost quite weary of this, actually. They love the idea, in theory, of democratizing access. They’re also quite fed up with AI-generated or -assisted papers that are flooding every publication imaginable, as well as the pre-print servers that they use in these fields. Some of those I spoke to said things like, “Oh, I got three emails this week alone with people being like, ‘Hey, is this legit?'” Because they thought they’d solved something with ChatGPT or with Claude, and they also don’t have the mathematical skills to check whether they’ve actually solved something. On the flip side, there are parts where they said, “Well, we’ve got a talented undergrad who’s done something that a talented undergrad would probably have never managed, and here they are doing grad-level work and they’ve produced a paper that is legit.” And in the bigger scheme of things, a few I spoke to said, “Well yeah, a lot of these are in the ivory tower. Having access to this kind of thing globally could really boost access to the kind of things here.” On the flip side, the cost. These things cost a lot to run. It’s always easy to forget when you use, say, a free version of ChatGPT or Claude or something. At the higher levels, these things cost money. They may not necessarily cost a lot of money — OpenAI claimed, I think it was $2,000 for these 10 results, but that doesn’t factor in literally anything else once they’ve got these, so it’s a very generous number. But even taking that figure, math is quite a poor discipline, even at very well-off institutions. Colva Roney-Dougal at St. Andrews, who I spoke to, said, “Well, a lot of the time I don’t bother getting a research grant. I don’t need one. I just have a blackboard.” And so if you’re not even getting a research grant, $2,000 is a lot to put up. So it could lock out researchers that way, even at quite well-funded institutions. Not to mention the speed at which this is happening, that virtually no one would’ve been able to bake any of this into a grant proposal yet, was another theme that I came across a lot. Actually, Roney-Dougal has another great quote in your piece about the nature of the AI labs and how they are talking about math. She said, “They’re treating our discipline as an advertising playground.” A bunch of mathematicians have signed something called the Leiden Declaration, which is an open letter to then pledge not to buy into hype around AI. These things are running right at each other. The AI labs are not going to stop using every discipline as an advertising playground. And a bunch of mathematicians saying, “We refuse to buy the hype,” certainly does not seem to be stopping the hype. There’s just a piece of this that is organized professional resistance to a thing that is upending a field that has, as you say, been pretty cheap to operate, and now might be getting cheaper or easier to access or easier to upend, day by day. Do mathematicians feel that that is going to be effective? Historically, mathematicians are not savvy political operators. There’s a part of me that says, “Oh, they’re just going to get run over.” I don’t know. In the history of math, actually, I think a lot of them were quite savvy. Isaac Newton is the one that always comes to mind for that — though quite a petty political operator as well. But yeah, that is the fear. A few that I spoke to, and one really comes to mind, mentioned that there’s often this belief that math is the pinnacle of knowledge. But he was like, “Well, that’s bullshit.” And he wasn’t alone in illustrating that sentiment. But it is good for showcasing, and it’s a lot neater as a discipline, and a lot cheaper. You mentioned Marcus referencing IBM’s Watson and the curing cancer ambition. Well, that involves lots of messy experiments, including on people. You don’t need that in math, so it’s a really easy discipline to come in, throw your weight around, and then move to somewhere more lucrative if that’s what you want. I’m not saying that that’s what they’re doing. A lot of the people at these companies have been hired. I don’t doubt their credentials for sure, and I don’t doubt their motivations as well. It does raise a question long-term as to how viable this is. Because let’s be clear, as a field goes, I cannot imagine mathematicians being a very lucrative enterprise customer for these companies. The thing that might be lucrative is pushing a field forward to turn it into something economically viable. We push mathematics forward as a field, that turns into some engineering or physics breakthrough based on that mathematics, and that turns into, I don’t know, yet another way to launch rockets. Some circle happens there that I don’t quite understand, but that is the history of innovation, from research, to engineering, to products or services that make money. Is that on the minds of any of these mathematicians, that pushing the boundaries here is upstream of something radically economically lucrative? The immediate counter that would come to mind here is that a lot are scared that it’s closing off the field. So by definition, those breakthroughs that lead to something surprising and new that you can say, “Oh, this works here,” may not be happening anymore. If anything, that lucrative endeavor of applying math to this entire new field that may have a lot of money in it remains an open question as to whether anything like that would be possible if we’re closing off avenues, rather than opening them up. Right, if the economic incentive of solving the unsolved problem is reduced because you personally won’t get rich if a computer is just solving every unsolved problem, something very fundamental breaks there. A lot of it comes down to the fact that it’s not just solving problems. With a lot of these things, as we said, it’s about what solving that problem tells you elsewhere. If these were very lucrative problems to be solving, I imagine that more people would be trying to solve them than have left them for decades. This is the nature of a lot of pure science, and it’s a broader criticism of what is going on perhaps with the Trump administration’s approach to science policy at the moment, in that it’s very applications-focused. There is something to doing pure research that can yield potentially very big dividends that is, by definition, utterly unpredictable as well.You cannot plan for it. The fear I think with math is that in solving all of these problems, and then also doing so without opening up new areas of research, what are you left with? Even if it’s from a more lucrative, “What are you going after,” point of view, if you’re not opening up new areas of research and you’re just ticking off old ones, it just leaves a big question mark as to what might be left in its wake. Even from those I spoke to that were very excited about what’s happening, they said that even they don’t really know what’s happening. And they’re excited from a personal level because, “Oh, we might be able to do this, might be to do that.” But there was still this lingering uncertainty of like, well, where does this leave the field? Especially for more pure disciplines like research mathematics, it’s tougher to say what comes next. Because in a lot of the other sciences you can say, “Well okay, well they’d shift onto more engineering problems, or applying that.” But if you solve all the problems at the ground and there’s nothing being built up from that, where do you go from there? One great thing about The Verge is our commenters are vast, they’re very knowledgeable, and there was a comment from a mathematics researcher on your story that I just want to read to you, and see if you think this is the right framework. Here it is: “I have no doubt these models will bring massive change in the field, but in their current state, they won’t yet drive us to obsolescence. Just occupy a particularly useful spot in our bag of tricks. My apprehension comes from not knowing where these things will peak, but overall I remain optimistic. I think AI will be a net boon for math when used properly.” I feel like “AI will be a net boon for X when used properly” is just where you land in life in a lot of things, but that’s the most optimistic response that I’ve heard: “If we get it right, it’s going to be great.” Is that the vibe, or is it still more shell-shocked than that? I’d say shell shock is still the overriding impression. I think perhaps the gut response to that is like, “Well, it will be a net positive for whom, and what is ‘properly’?” All of those are quite legitimate questions here. There were some bleak responses from graduate students I saw in essays posted online.Where’s their place in this as future researchers? Do they have a place in this? Is it as glorified AI proof checkers? That will be quite an unsatisfying career, I imagine. Or maybe not, I don’t know. We will see. I think anything used properly will be a net boon. But yeah, I think it all comes down to what “properly” means, and for whom we’re talking about. I feel like, oddly, that is an excellent place to leave it. Because I don’t think either one of us knows, and I suspect over the next year or so, things will come into focus. Because at some point, OpenAI will have to show people how they did the things of the models. And perhaps more importantly, the other labs are going to want to either replicate these results or show that they can push farther, which will necessarily have to lead to a little bit more transparency and yet more mathematicians having a crisis with you. Robert, thank you so much for being on the show. We’ll have you back very soon. Thank you for having me. Questions or comments? Hit us up at [email protected]. We really do read every email!

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Today on Decoder, I’m talking with Robert Hart, The Verge’s London-based AI reporter, about what AI is doing to the field of mathematics and the existential crisis many lead mathe…
サイト内本文

翻訳待ち:Index of the best vibe coding tools

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Tools What you need to be a great vibe coder ListMap FeaturedVibe Costs · tracks what vibe coders spend on subscriptions, APIs and hostingVisit ↗ No.AppBuilderCategory 01 Claude Code Anthropic's agentic coding harness t…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Tools What you need to be a great vibe coder ListMap FeaturedVibe Costs · tracks what vibe coders spend on subscriptions, APIs and hostingVisit ↗ No.AppBuilderCategory 01 Claude C…
サイト内本文

翻訳待ち:Slack is launching collaborative vibe coding channels

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Slack is introducing dedicated channels where teams can vibe code together with AI agents instead of jumping between different tools and conversations. The Slack Code launch includes open, project-specific code channels with dedicated user tabs, alongside features that compare coding changes and preview HTML output before the project is shipped. "With Slack Code, when you have an idea or need to build a new feature, update a web page, or fix a bug, you simply tag in a coding agent like Anthropic's Claude or Cognition's Devin, and that agent then spins up a code channel to tackle the task," Slack said in its press release. "There, everyone … Read the full story at The Verge.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Slack is introducing dedicated channels where teams can vibe code together with AI agents instead of jumping between different tools and conversations. The Slack Code launch inclu…
サイト内本文

翻訳待ち:Auditing Preference Biases and Fine-Tuning Language Models with Direct Preference Optimization on Anthropic HH-RLHF Using TRL and LoRA

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:This tutorial provides an end-to-end workflow for fine-tuning language models using Direct Preference Optimization (DPO). We demonstrate how to audit the Anthropic HH-RLHF dataset for structural and length-based biases, implement a robust training pipeline using TRL and LoRA, and evaluate model performance to ensure genuine preference learning rather than reliance on lexical shortcuts. The post Auditing Preference Biases and Fine-Tuning Language Models with Direct Preference Optimization on Anthropic HH-RLHF Using TRL and LoRA appeared first on MarkTechPost.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • This tutorial provides an end-to-end workflow for fine-tuning language models using Direct Preference Optimization (DPO). We demonstrate how to audit the Anthropic HH-RLHF dataset…
サイト内本文

翻訳待ち:Revisiting the "Push-T" Robot Manipulation Task with Agentic Robotics

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.18227v1 Announce Type: new Abstract: Push-T is an iconic benchmark for learning manipulation policies from human demonstrations. The robot must use a single point of contact to push a T-shaped block into a target pose. In this short paper, we revisit the Push-T task in the context of emerging advances in Agentic Robotics where an LLM coding agent -- Claude Code with Fable 5 -- is prompted to create an algorithmic solution that does not require any demonstration data. We study how effective the agentic coding loop can solve the Push-T task, and compare the resulting code as policy with the visuomotor imitation learning policy. Results suggest that the agent found the 2D gym simulation online, and used sim experiments to learn push mechanics, iteratively optimizing to achieve 100% success rate using 46% fewer steps than the best diffusion policy trained with 200 human demonstrations. The coding agent also solve extensions from T to the full alphabet (Push-A to Push-Z) using a self generated curriculum and generated simulation code for the Franka and UR5 robot arms in 3D cross-embodiment simulations with visual feedback. Videos, policies and details will be posted online.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.18227v1 Announce Type: new Abstract: Push-T is an iconic benchmark for learning manipulation policies from human demonstrations. The robot must use a single point of co…
サイト内本文

翻訳待ち:Comparing Qwen3.8 Max and Fable 5 on UI

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Job Rietbergen Aug 19, 2026 Alibaba released Qwen3.8-Max on August 2, and it debuted at #4 on Arena.ai’s Frontend Code leaderboard, one spot above Claude Fable 5. It is the second model from a Chinese lab to pass Fable…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Job Rietbergen Aug 19, 2026 Alibaba released Qwen3.8-Max on August 2, and it debuted at #4 on Arena.ai’s Frontend Code leaderboard, one spot above Claude Fable 5. It is the second…
サイト内本文

翻訳待ち:Round Hill Music Sues Suno and Anthropic for $1B over AI Training Data

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The music industry spent years asking AI companies nicely. Now it’s asking federal judges. Independent publisher Round Hill Music filed separate copyright infringement suits against Anthropic and Suno in U.S. District C…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • The music industry spent years asking AI companies nicely. Now it’s asking federal judges. Independent publisher Round Hill Music filed separate copyright infringement suits again…
サイト内本文

企業ナビゲーション

Anthropic AI ニュース | AI News Hub