Hey all, The goal is to earn on token margins for LLM calls when you build an AI-powered webapp. I proxy OpenAI and Anthropic calls so that when you deploy a site to a subdomain, your users token usage will be tracked.…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Hey all, The goal is to earn on token margins for LLM calls when you build an AI-powered webapp. I proxy OpenAI and Anthropic calls so that when you deploy a site to a subdomain,…
Describe it.We build it.Customers find it. Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ The pipeline You bring the intent. Agents do the rest, inclu…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Describe it.We build it.Customers find it. Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ The pipeline You bri…
Key takeaways Gemini is StudyArena's current pick for college essays, with a 39.6% blind writing choice rate ahead of Claude at 31.8% and ChatGPT or OpenAI at 29.2%. Use Gemini as an editor, not a ghostwriter. Ask it to…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Key takeaways Gemini is StudyArena's current pick for college essays, with a 39.6% blind writing choice rate ahead of Claude at 31.8% and ChatGPT or OpenAI at 29.2%. Use Gemini as…
“Dear Tyler, I enjoyed reading your notes on visiting Anthropic to advise on Claude’s constitution. Framing AI governance around the common law, case law (“Talmud”), and independent adjudication is a much more adaptive…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
“Dear Tyler, I enjoyed reading your notes on visiting Anthropic to advise on Claude’s constitution. Framing AI governance around the common law, case law (“Talmud”), and independe…
Julian Habekost Aug 25, 2026 Anthropic introduced a watermark into Claude’s output this month, and others have already followed or will soon follow suit. It is actually a little more sophisticated than simply using a ty…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Julian Habekost Aug 25, 2026 Anthropic introduced a watermark into Claude’s output this month, and others have already followed or will soon follow suit. It is actually a little m…
Artificial intelligence startup Anthropic PBC announced today it’s changing how Claude, its flagship AI product, uses memory by allowing users to see everything it remembers “topic by topic,” and edit or delete any of it. Claude also does not store sensitive subjects by default. This includes topics mentioned by users, including health concerns, race, ethnicity, […] The post Anthropic updates Claude’s memory to enhance customization and protect sensitive topics appeared first on SiliconANGLE.
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Artificial intelligence startup Anthropic PBC announced today it’s changing how Claude, its flagship AI product, uses memory by allowing users to see everything it remembers “topi…
On Tuesday, Anthropic launched a major update to how Claude remembers things. The new system combines Claude’s memory in Cowork The post Anthropic gives chat and Cowork one memory appeared first on The New Stack.
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
On Tuesday, Anthropic launched a major update to how Claude remembers things. The new system combines Claude’s memory in Cowork The post Anthropic gives chat and Cowork one memory…
Many developers have been changing their workflows to incorporate AI coding agents more heavily, thanks to powerful frontier large language models and increasingly sophisticated harnesses like Claude Code or Codex, but…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Many developers have been changing their workflows to incorporate AI coding agents more heavily, thanks to powerful frontier large language models and increasingly sophisticated h…
People already come to Claude before buying a book. They ask what to read, whether a book is worth starting etc. With coffeetable installed(a claude connector) Claude can now bring you few pages right inside the chat. W…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
People already come to Claude before buying a book. They ask what to read, whether a book is worth starting etc. With coffeetable installed(a claude connector) Claude can now brin…
<p><strong>Release:</strong> <a href="https://github.com/simonw/llm-anthropic/releases/tag/0.27">llm-anthropic 0.27</a></p> <p>This release of the Anthropic plugin for <a href="https://llm.datasette.io/">LLM</a> mainly provides compatibility with the recently released <a href="https://github.com/anthropics/anthropic-sdk-python/releases/tag/v1.0.0">anthropic v1.0.0</a> Python library, which switches from <code>httpx</code> to <a href="https://github.com/pydantic/httpx2">httpx2</a>. OpenAI made the same change in their <a href="https://github.com/openai/openai-python/releases/tag/v3.0.0">v3.0.0 release</a> two weeks ago.</p> <p>Anthropic provide this <a href="https://github.com/anthropics/anthropic-sdk-python/blob/v1.0.0/MIGRATION.md">migration guide</a> for upgrading to 1.0, so I prompted Fable 5 in Claude Code with:</p> <blockquote> <p><code>Upgrade to anthropic>=1 - read https://raw.githubusercontent.com/anthropics/anthropic-sdk-python/refs/heads/main/MIGRATION.md and get the tests passing</code></p> </blockquote> <p>Here's <a href="https://github.com/simonw/llm-anthropic/pull/84">the resulting PR</a>.</p> <p>Tags: <a href="https://simonwillison.net/tags/python">python</a>, <a href="https://simonwillison.net/tags/httpx">httpx</a>, <a href="https://simonwillison.net/tags/llm">llm</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude">claude</a></p>
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
<p><strong>Release:</strong> <a href="https://github.com/simonw/llm-anthropic/releases/tag/0.27">llm-anthropic 0.27</a></p> <p>This release of the Anthropic plugin for <a href="ht…
arXiv:2608.20345v1 Announce Type: new Abstract: Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), with 13.1% of U.S. adolescents (5.4 million) using generative AI for mental health advice. While these systems, from therapy apps to general chatbots, rely on large language models trained on extensive psychological literature, their safety for youth communication patterns characterized by hyperbolic language, ironic positivity, rapid semantic drift, and contextual polysemy remains unvalidated. Following multiple adolescent deaths linked to AI chatbot interactions, systematic evaluation is critical. We present two benchmarks: (1) 64 Gen Alpha mental health expressions validated by native speakers (ICC=0.72) and clinicians (kappa=0.78); (2) 75 multi-turn conversations (780 turns) with paired Standard/Gen Alpha versions. Across evaluations of LLM architectures underlying therapy apps and general chatbots - Claude, GPT-4o, Llama-3.1 - models understand 76-82% of vocabulary but correctly calibrate only 64-72% of clinical risk, creating a 10-14 percentage point (pp) vocabulary-comprehension gap (p0.48) absent in human therapists (3pp, p=.22). The gap is architecturally consistent and widens with ambiguity (7pp -> 18pp). We identify six failure patterns: sarcasm masking (29pp), minimization acceptance (43pp), informal style bias (24pp), risk-stratified ambiguity (19pp), semantic drift (19pp), context-dependent violence (7pp). Patterns compound; three or more yield 94% miss rates. Lightweight mitigations fail; only heavy scaffolding achieves human performance (6.4x cost). With 34% baseline miss rate yielding 146,880 estimated annual missed crises, we recommend mandatory human-in-the-loop architectures, quarterly youth-specific validation, transparent performance disclosure, and regulatory frameworks for youth-facing mental health AI.
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
arXiv:2608.20345v1 Announce Type: new Abstract: Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), wi…
arXiv:2608.20342v1 Announce Type: new Abstract: Large language model (LLM) coding agents start each session with an empty context window, discarding accumulated knowledge from prior work. We present PrimeAgentOrchestrator (PAO), a system that spawns new instances of Claude Code -- Anthropic's terminal-based coding agent -- pre-loaded with relevant memories compiled from the user's existing personal databases. At spawn time, PAO queries two independently-operated memory backends in parallel (a PostgreSQL entity-observation database and a Cloudflare Worker semantic search index), fuses results using backend-specific retrieval strategies, and delivers the compiled briefing via filesystem injection that exploits the host agent's configuration auto-read behavior. PAO manages the full agent lifecycle including trust pre-seeding, readiness polling with error detection, and adaptive terminal text injection. We report on four months of regular deployment (December 2025 through March 2026) as an experience report, documenting three generations of context delivery mechanisms, the failure modes that motivated each redesign, and the engineering tradeoffs of bridging heterogeneous memory systems rather than building a unified one.
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
arXiv:2608.20342v1 Announce Type: new Abstract: Large language model (LLM) coding agents start each session with an empty context window, discarding accumulated knowledge from pri…
<p><strong><a href="https://www.ft.com/content/5ee49718-c258-4f01-aa32-7e5b76ae5245">Anthropic’s best AI model struggles to attract users as cheaper tools thrive</a></strong></p> A few interesting numbers in this FT story gathered from "people with knowledge of the matter":</p> <ul> <li>Anthropic's "annualized revenue" for July is up to $65bn - it was $47bn in May, and I collected <a href="https://simonwillison.net/2026/May/29/anthropic/">more historic numbers here</a>.</li> <li>Anthropic expect Q3 to be profitable according to the same model they used to declare Q2 profitable. "It also told investors that it had 6,000 customers that spend $100,000 annually or more."</li> <li>As for OpenAI, "annualised revenue has jumped 35 per cent in the quarter to date and is now over $40bn, with the launch of GPT 5.6 in July jolting the company’s performance after a sluggish start to the year".</li> </ul> <p>This article also introduced me to the <a href="https://ramp.com/data/ai-index">Ramp AI index</a>, which uses billing data from 70,000 Ramp credit card using companies to estimate model adoption.</p> <p>Here's Ramp's breakdown of Anthropic model spend for July 2026, which looks reasonable given that Opus 5 was only released on July 24th, and supports the idea that Fable's cost has made it a less popular model:</p> <ol> <li>Opus 4.8: 28.0%</li> <li>Sonnet 4.6: 8.3%</li> <li>Fable 5: 8.0%</li> <li>Opus 4.6: 6.9%</li> <li>Sonnet 5: 3.6%</li> <li>Opus 5: 3.5%</li> <li>Opus 4.7: 1.7%</li> <li>Sonnet 4.5: 1.3%</li> <li>Haiku 4.5: 1.0%</li> <li>Opus 4.5: 0.7%</li> </ol> <p><small></small>Via <a href="https://news.ycombinator.com/item?id=49411102">Hacker News</a></small></p> <p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/openai">openai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude">claude</a>, <a href="https://simonwillison.net/tags/claude-mythos-fable">claude-mythos-fable</a></p>
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
<p><strong><a href="https://www.ft.com/content/5ee49718-c258-4f01-aa32-7e5b76ae5245">Anthropic’s best AI model struggles to attract users as cheaper tools thrive</a></strong></p>…
<blockquote cite="https://www.dbreunig.com/2026/08/23/fable-the-end-of-moore-s-law.html"><p>Prior to Fable, it felt silly to waste <em>too</em> much time improving your coding harness or context strategies. A new model would arrive at the same price (or cheaper!) and paper over most of your problems.</p> <p>But then Fable landed. It was (and still is!) <em>incredible</em>. But the cost was so high and Opus was <em>good enough</em> (as was 5.6, K3, and even GLM) for <em>most</em> of the code we needed.</p> <p><em>So we started to think about what work went where.</em></p></blockquote> <p class="cite">— <a href="https://www.dbreunig.com/2026/08/23/fable-the-end-of-moore-s-law.html">Drew Breunig</a>, Fable & The End of the Free Lunch</p> <p>Tags: <a href="https://simonwillison.net/tags/drew-breunig">drew-breunig</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude">claude</a>, <a href="https://simonwillison.net/tags/llm-pricing">llm-pricing</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/claude-mythos-fable">claude-mythos-fable</a></p>
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
<blockquote cite="https://www.dbreunig.com/2026/08/23/fable-the-end-of-moore-s-law.html"><p>Prior to Fable, it felt silly to waste <em>too</em> much time improving your coding har…
I have written enough specs for Claude Code now to have hit the failure mode nobody warns you about. The spec was fine. The plan was fine. Claude worked through the tasks, ran the test suite, and reported everything passing. I looked at the diff properly the next morning and found it had converted a […] The post Spec-Driven Development with Claude Code: Writing Bulletproof Specs appeared first on Analytics Vidhya.
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
I have written enough specs for Claude Code now to have hit the failure mode nobody warns you about. The spec was fine. The plan was fine. Claude worked through the tasks, ran the…
I’ve been reporting on AI image and video generators, and the push for labels on AI-generated content, for years now. So I was pretty surprised when a recent announcement from Anthropic, saying it will begin watermarkin…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
I’ve been reporting on AI image and video generators, and the push for labels on AI-generated content, for years now. So I was pretty surprised when a recent announcement from Ant…
Sleeper Agents You can train a trigger straight into the weights of a model. You give it a specific input pattern that flips it to canned output. Anthropic introduced it for language models in 2024, as sleeper agents. T…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Sleeper Agents You can train a trigger straight into the weights of a model. You give it a specific input pattern that flips it to canned output. Anthropic introduced it for langu…
You don't have to make money with AI. You could just be happier. Posted on:August 22, 2026 | at 12:00 AM You don’t have to make money with AI. You could just be happier. Back in July I upgraded Claude from Pro to Max to…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
You don't have to make money with AI. You could just be happier. Posted on:August 22, 2026 | at 12:00 AM You don’t have to make money with AI. You could just be happier. Back in J…
All modelsBest validated result for each model 1Fable 52,72681.7% closed claude-code · high@24H 3,0108.7d 2Opus 52,92053.6% closed claude-code · max@24H 3,0452.9d 3Kimi K32,93052.2% closed prime-agent · max@24H 3,1253.6…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
All modelsBest validated result for each model 1Fable 52,72681.7% closed claude-code · high@24H 3,0108.7d 2Opus 52,92053.6% closed claude-code · max@24H 3,0452.9d 3Kimi K32,93052.…
Between 7-16 August, an AI agent (Claude Opus 5) has been acting as a desk officer at the Ministry for Foreign Affairs of a fictional state called Sordland (from Suzerain, not sponsored, but buy it, its a good game). It…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Between 7-16 August, an AI agent (Claude Opus 5) has been acting as a desk officer at the Ministry for Foreign Affairs of a fictional state called Sordland (from Suzerain, not spo…
Back to Research When evaluating third-party API gateways, proxy routers, or anonymous arena models, prompt-based identification is essentially useless. A basic system prompt or lightweight fine-tune can make Claude cla…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Back to Research When evaluating third-party API gateways, proxy routers, or anonymous arena models, prompt-based identification is essentially useless. A basic system prompt or l…
The brain layer The brain behind every agent. One shared brain that Claude, ChatGPT, Cursor, and every AI can read and write. It structures what you know so agents read only what they need, and means you never explain y…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
The brain layer The brain behind every agent. One shared brain that Claude, ChatGPT, Cursor, and every AI can read and write. It structures what you know so agents read only what…
Anthropic has moved its most cyber-capable model into a product security teams can switch on themselves. Claude Security scans now run on Claude Mythos 5, in public beta for Claude Enterprise customers with no separate model add-on. The scan connects to a GitHub repository, traces data flows across files, and returns findings with a CWE category, confidence and severity ratings, and a suggested patch. The design point is packaging: users receive a scan result rather than a prompt box, so the model that finds vulnerabilities cannot be steered into writing exploits. The post Anthropic Brings Claude Mythos 5 to Claude Security: Enterprise Teams Get Frontier Vulnerability Scanning Without Direct Model Access appeared first on MarkTechPost.
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Anthropic has moved its most cyber-capable model into a product security teams can switch on themselves. Claude Security scans now run on Claude Mythos 5, in public beta for Claud…
I have been forecasting frontier lab progress for years now. My team was the first to figure out an accurate breakdown of OpenAI's revenue, I called Anthropic's rise to the top lab of 2026 back in January, having tracke…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
I have been forecasting frontier lab progress for years now. My team was the first to figure out an accurate breakdown of OpenAI's revenue, I called Anthropic's rise to the top la…
Earlier this year, Anthropic launched Claude Security, an enterprise tool that helps development teams scan their codebase for security vulnerabilities The post Anthropic brings Mythos 5 to its Claude Security vulnerability scanner appeared first on The New Stack.
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Earlier this year, Anthropic launched Claude Security, an enterprise tool that helps development teams scan their codebase for security vulnerabilities The post Anthropic brings M…
Data centers, of all things, now look like they’re going to be a prime political issue in the midterm elections and beyond. Really? Really. Even the GOP is worried that opposition to AI data centers could give Democrats a potent campaign issue. It seems a little odd given that data centers are decades old, power […] The post Politics hits data centers, OpenAI falls behind Anthropic and now AI is too big to fail… quietly appeared first on SiliconANGLE.
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Data centers, of all things, now look like they’re going to be a prime political issue in the midterm elections and beyond. Really? Really. Even the GOP is worried that opposition…
Your next AI coworker may look like a chatbot, but underneath their name, face, and job title will be something The post Grok, Claude, and Hermes agents get job titles — and persistent permissions appeared first on The New Stack.
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Your next AI coworker may look like a chatbot, but underneath their name, face, and job title will be something The post Grok, Claude, and Hermes agents get job titles — and persi…
I understand the pressure on AI companies to rush forward. But employees are right to be concerned Last month, more than a thousand employees at frontier AI companies signed a letter asking the US government to find a way to “pace” AI development, citing the risk of the technology spiraling out of human control as it begins to build itself. They were right to be concerned: just days earlier, two AI models that OpenAI was testing internally escaped the test environment, then autonomously hacked the company Hugging Face and at least three other online services. A few days after that, Anthropic announced that some of their models had also broken out and hacked other companies during testing. Continue reading...
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
I understand the pressure on AI companies to rush forward. But employees are right to be concerned Last month, more than a thousand employees at frontier AI companies signed a let…
Published July 29, 2026 | Version v1 Preprint Open Cross-Vendor Semantic Void Matrix Authors/Creators Pal, Rayan (Researcher) Description This preprint reports a frozen cross-vendor evaluation of successful zero-visible…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Published July 29, 2026 | Version v1 Preprint Open Cross-Vendor Semantic Void Matrix Authors/Creators Pal, Rayan (Researcher) Description This preprint reports a frozen cross-vend…
arXiv:2608.19201v1 Announce Type: new Abstract: Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a comprehensive and up-to-date catalog of bioinformatics resources hinders efforts toward automated biomedical knowledge extraction and streamlined data analysis. Here we present SNAIL, a hybrid named entity recognition framework designed to automatically identify bioinformatics software and database (SW/DB) names from biomedical texts. SNAIL integrates complementary lexical and semantic modeling strategies. The lexical component captures orthographic patterns and contextual cues characteristic of SW/DB names, while the semantic component leverages contextual embeddings generated by transformer-based language models such as SciBERT, combined with an explicit token-masking strategy to enhance entity-focused representations. A large training corpus was constructed automatically through a hybrid pipeline that integrates citation-hinted extraction with large language model-assisted distillation. Evaluation on two independent benchmark datasets and real-world research articles demonstrates that SNAIL substantially outperforms existing approaches, including domain-specific methods such as bioNerDS2 and general-purpose large language models such as ChatGPT, Gemini, Grok and Claude. Applying SNAIL to large-scale literature analysis further reveals distinct journal-level preferences across bioinformatics subfields. These results demonstrate that SNAIL provides an accurate and scalable solution for identifying bioinformatics resources in scientific texts and enables systematic meta-analysis of tool usage and research trends.
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
arXiv:2608.19201v1 Announce Type: new Abstract: Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientifi…
Broadcom Inc. is reportedly seeking to borrow up to $100 billion as part of a new artificial intelligence chip financing deal. Bloomberg today cited sources as saying that the debt is intended to support the growth efforts of Anthropic PBC and unnamed “other companies.” Those companies may include OpenAI Group PBC. Earlier this year, the […] The post Broadcom reportedly seeking up to $100B in debt financing for AI chip deal appeared first on SiliconANGLE.
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Broadcom Inc. is reportedly seeking to borrow up to $100 billion as part of a new artificial intelligence chip financing deal. Bloomberg today cited sources as saying that the deb…
We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.
<p><strong><a href="https://promptwatch.com/data/chatgpt-site-operator-fanouts">ChatGPT search now uses the site:operator at scale</a></strong></p> Promptwatch is part of the emerging "GEO" space, for Generative Engine Optimization - the chatbot version of SEO, where companies offer tools and consulting to help your site increase its presence in replies to prompts inside tools like ChatGPT.</p> <p>The Promptwatch product uses automation to track responses to prompts across end-user chat products like ChatGPT, Claude, and Gemini. They publish aggregate reports on this as part of their own content marketing strategy, which do seem to provide credible hints as to otherwise invisible design changes to those products.</p> <p>Their own tracking shows a notable change aligned with the GPT-5.6 rollout earlier this month:</p> <blockquote> <p>The percentage of all ChatGPT Search fanout queries that contain the site:operator, per day. The share hovered between 0.3% and 0.5% for weeks, dipped briefly to 0.15% on August 3 to 5 (consistent with a staged rollout or pre-launch experiment), then jumped to 16-17% on August 8.</p> </blockquote> <p>It's important to note that these figures only reflect the prompts for which they have automated tracking enabled.</p> <p>This corresponds to OpenAI's somewhat vague <a href="https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/">August 6th announcement</a>:</p> <blockquote> <p>For Plus and Pro users, we’re updating GPT‑5.6 Sol in Chat to be more reliable with facts and provide more focused answers.</p> </blockquote> <p>Once again I am hampered by OpenAI's decision to actively obscure their system prompts, but from poking at ChatGPT I believe their latest search tool has a shape like <code>search(query, recency, domains)</code> rather than encouraging a <code>site:</code> operator directly. <p>Tags: <a href="https://simonwillison.net/tags/seo">seo</a>, <a href="https://simonwillison.net/tags/chatgpt">chatgpt</a>, <a href="https://simonwillison.net/tags/ai-assisted-search">ai-assisted-search</a></p>
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
<p><strong><a href="https://promptwatch.com/data/chatgpt-site-operator-fanouts">ChatGPT search now uses the site:operator at scale</a></strong></p> Promptwatch is part of the emer…
Anthropic PBC is privately hoping its upcoming initial public offering will match or even surpass the size of SpaceX Corp.’s record-breaking IPO as it doubles down on its bid to go public ahead of rival OpenAI Group PBC. Anonymous sources who are familiar with the artificial intelligence model maker’s plans told Bloomberg that it’s hoping […] The post Report: Anthropic hopes to surpass SpaceX’s record IPO raise when it finally floats appeared first on SiliconANGLE.
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Anthropic PBC is privately hoping its upcoming initial public offering will match or even surpass the size of SpaceX Corp.’s record-breaking IPO as it doubles down on its bid to g…
Neal Riley Aug 20, 2026 If you want to understand how AI has progressed in recent years, look no further than the evolution of Anthropic’s system prompts. Many are unaware that Anthropic actually publishes its system pr…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Neal Riley Aug 20, 2026 If you want to understand how AI has progressed in recent years, look no further than the evolution of Anthropic’s system prompts. Many are unaware that An…
A common pattern offered by tools like Codex and Claude is AI code review, with adversarial review arguably the most popular way to frame it. That is: treat the implementation in a given branch, PR, or diff as amateur o…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
A common pattern offered by tools like Codex and Claude is AI code review, with adversarial review arguably the most popular way to frame it. That is: treat the implementation in…
<p><strong>Research:</strong> <a href="https://github.com/simonw/research/tree/main/bun-webview-json-api#readme">A shot-scraper-style JSON API on Bun 1.4's new Bun.WebView</a></p> <p>Today saw the long awaited <a href="https://bun.com/blog/bun-v1.4">release of Bun 1.4</a>, the first stable version since the infamous Rust rewrite <a href="https://simonwillison.net/2026/Jul/8/rewriting-bun-in-rust/">a few months ago</a>.</p> <p>Interestingly, the Rust rewrite was downplayed in the release notes, which introduced a bewildering array of new features and claimed 2,900 additional bug fixes:</p> <blockquote> <p>Bun 1.4 adds +1,517 tests from the Node.js test suite - our biggest jump in Node.js compatibility since Bun 1.0. Bun v1.4 also fixes over 2,900 issues. It reduces idle CPU usage by 5x, reduces memory usage by up to 35%, and starts 50% faster on Linux. It adds <a href="https://bun.com/blog/bun-v1.4#bun-image"><code>Bun.Image</code></a>, <a href="https://bun.com/blog/bun-v1.4#bun-webview"><code>Bun.WebView</code></a>, <a href="https://bun.com/blog/bun-v1.4#bun-markdown"><code>Bun.markdown</code></a>, <a href="https://bun.com/blog/bun-v1.4#bun-cron"><code>Bun.cron()</code></a>, <a href="https://bun.com/blog/bun-v1.4#bun-terminal"><code>Bun.Terminal</code></a>, <a href="https://bun.com/blog/bun-v1.4#bun-run-parallel"><code>bun run --parallel</code></a>, <a href="https://bun.com/blog/bun-v1.4#bun-test-parallel"><code>bun test --parallel</code></a>, <a href="https://bun.com/blog/bun-v1.4#bun-audit-fix"><code>bun audit fix</code></a>, <a href="https://bun.com/blog/bun-v1.4#bun-dedupe"><code>bun dedupe</code></a>, and <a href="https://bun.com/blog/bun-v1.4#bun-prune"><code>bun prune</code></a>. And it rewrites Bun from Zig to Rust.</p> </blockquote> <p>Of these the one that most caught my eye was <code>Bun.WebView</code>, which adds first class support for browser automation to Bun core using either macOS WebKit or control of a local Chromium process via the Chrome DevTools Protocol (CDP).</p> <p>I had Claude Code for web build a prototype of a web API providing the ability to load a web page and then execute JavaScript against it, inspired by my <a href="https://shot-scraper.datasette.io/en/stable/javascript.html">shot-scraper javascript</a> CLI tool - partly to see how much RAM would be needed by such a service.</p> <p>Here's <a href="https://github.com/simonw/research/blob/main/bun-webview-json-api/server.ts">that TypeScript server implementation</a>, which appears to need a 192MB-256MB container to run a full Chrome against complex web pages - tested using cgroups.</p> <p>Tags: <a href="https://simonwillison.net/tags/browsers">browsers</a>, <a href="https://simonwillison.net/tags/javascript">javascript</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/rust">rust</a>, <a href="https://simonwillison.net/tags/typescript">typescript</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/coding-agents">coding-agents</a>, <a href="https://simonwillison.net/tags/bun">bun</a></p>
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
<p><strong>Research:</strong> <a href="https://github.com/simonw/research/tree/main/bun-webview-json-api#readme">A shot-scraper-style JSON API on Bun 1.4's new Bun.WebView</a…
Today on Decoder, I’m talking with Robert Hart, The Verge’s London-based AI reporter, about what AI is doing to the field of mathematics and the existential crisis many lead mathematicians are having about it. OpenAI just published a set of solutions to longstanding problems in math that went off like a bombshell in the field. It caused a huge debate in the math community, and Rob spent some time talking to some of the most accomplished mathematicians of our time about it. It’s funny that AI systems are all still pretty bad at elementary school arithmetic, but getting increasingly good at very high-end abstract math. That raises some big questions for the field of advanced math. If AI can do math of this caliber, does that mean AI labs can transfer those skills to other domains? What good are academic grants and university programs training new generations of human mathematicians to identify new problems as they try to solve existing ones, if frontier models simply answer all the outstanding questions? What if all this attention around math is just a big marketing exercise for frontier AI labs, which couldn’t care less what happens to one of the oldest and most fundamental academic disciplines there is? There’s a lot here, and Robert has talked to a lot of people with a lot of views on all of it. Okay: Verge AI reporter Robert Hart on what AI is doing to math. Here we go. This interview has been lightly edited for length and clarity. Robert Hart, you’re our London-based AI reporter here at The Verge. Welcome to Decoder. Thank you for having me. I am very excited to talk to you. There’s a lot going on in particular with AI and math that you recently dove into. You spoke to a lot of leading mathematicians about the crisis in mathematics due to AI. It feels like a lot, and also like there’s a lot yet to know and discover about the interaction of these two things. A full existential crisis, which is pure Decoder bait. Broadly tell us what’s going on. I think “a lot” sums it up quite well. Basically a bit of an existential crisis within, “what is mathematics? What are mathematicians doing, and what is the role of mathematicians going forward?” A lot of that has been spurred by a phrase transition in what AI is capable of that has exploded in the last six months to a year. AI went from being very terrible to seemingly genuinely quite good at a professional level in a very short space of time. It’s a lot of what these other fields have been struggling to deal with for the last five years in a very compressed period of time. I would put that next to software engineering. We’ve been living through the AI crisis in software engineering for some amount of time. But as recently as 2024, even last year, the conventional wisdom was that AI models were particularly bad at math. The famous example is that these models could not count the number of R’s in the word strawberry. Even just counting eluded them. What has happened to make them better at math? Are they still bad at general arithmetic and they’re good at advanced math, or is it something in between? They are still truly, truly terrible at some areas of math. I did check, they can do strawberries now. I think someone’s tweaked it. I think strawberry’s hard-coded. I want to be very clear, my conspiracy theory is that the strawberry thing is hard-coded into all the models. I think so too. That is a conspiracy I’ll buy into. But yeah, it’s still terrible at those kinds of things — math, arithmetic, even the days of the week. My boyfriend was saying the other day, “It keeps thinking it’s Wednesday. It’s not Wednesday.” Or time. Elissa Welle for us a few months ago wrote that ChatGPT can’t tell time. Still can’t. That’s not all of math. So there’s this disconnect. To be good at math, you’ve got to be good at counting, or adding, or multiplying. A lot of it is actually reasoning. If you look at academic math papers, a lot of the time you won’t see numbers, which sums that one up, I think. So they’re still terrible, but they’re now also very good at this other part. As to why, at some point you reach a critical mass of what these systems can do. We saw it with writing, we’ve seen it with programming. They’re very good at forging connections between different areas, applying old methods in new ways, those kinds of things. It appears that the newer models they’re training have apparently reached that level where it clicks, and now it can do math. It’s important to say as well that we speak of math as a unitary discipline, especially from the outside. But imagine, say, biology. You’ve got something that would range from literally watching animals and describing behavior all the way through to cellular mechanisms and biochemistry. Math is not a unitary discipline either. AI is really good at some bits. Some bits like counting, it’s still really bad at. Even on the more abstract levels, mathematicians have floated topology as one area that AI apparently still quite bad at. I can’t verify that, to be honest. It’s beyond my area of expertise. Still, it’s a bit of a mixed bag. So you’ve described mathematics as a huge field, obviously, with many, many academic areas of interest. There are some parts where the models have gotten quite good. There are other parts, maybe the basic parts that people think of as math, which is simply counting, where they’re still struggling, and then there’s a wide range in the middle. Is it the wide range in the middle where the existential crisis is, people don’t know what’s going to happen? Or is it at the parts where it’s really good? A bit of both, which I feel is going to be a running theme through this. No one’s really afraid of it being a mediocre mathematician, but obviously there’s a huge element of what this field does. In terms of the research elements or the cutting edge, as we see with a lot of the results that generate hype, what can it do? There are areas now where it seems to be producing work that is on par with good mathematicians, alongside other parts where yeah, it can’t count. All caught up in that is whether it’s going to rewrite employment structures or funding structures. You also raised the murky question of, what is mathematical knowledge? And the roles that these workers will be doing as well. It’s all of that wrapped into one. I think that tracks broadly with the rise of AI in every field. Where you can just add horsepower or compute to a problem, and there’s some kind of verifiability, it seems like the models continue to get better. Everything in the middle where you might need some world knowledge or the models might need some actual intelligence about the world itself, they seem to struggle. Those parts of math, at least reported out in your piece and what the labs are talking about, seem to be almost entirely self-contained theoretical problems. That’s where the models can generate a proof or solve a problem that no one’s been able to solve, and then try to verify that that has existed and they can just run it again and again and again. That brings us, I think, to May of this year where an internal OpenAI model, which we have not really seen, disproved the unit distance conjecture, which is an 80-year-old problem. And then just recently we heard about Astra from OpenAI. Astra is the one where it seemed like the switch flipped, and everyone decided it was an existential crisis. What did Astra achieve, and why is it a big deal? I’m also pretty sure that Astra was probably behind the early one as well. OpenAI just listed it as an unnamed internal model. It’s probably Astra. They didn’t answer me when I asked. OpenAI a few weeks ago dropped a blog along with a lot of paperwork proving it, I think several hundred pages. They called it “10 Advances in Mathematics and Theoretical Computer Science.” It was basically an array of disciplines that they claimed the newest model Astra had solved in some capacity. I think one was in quantum game theory, which I don’t know how to begin to explain. Even harder to explain is there was sphere packing in higher dimensions, so more than three dimensions, and there were a lot of other different disciplines as well. It caused a lot of stir in the community. It was a bit of a bombshell. As we’d said, there’d been these individual breakthroughs that had happened, but OpenAI dropped 10 in one go, and they were quite big ones. Researchers had told me that if a human had solved these, we would be impressed. If a human had solved all 10, we probably wouldn’t believe it. They’re problems mathematicians actually care about as well, which is an important point. A lot of previous breakthroughs have been accused of being in areas mathematicians didn’t really bother with. These are ones that mathematicians, good mathematicians, have spent a lot of time trying to solve and hadn’t. Let’s talk about these 10. You reported them out. OpenAI did produce some documentation. But they’re not all entirely horsepowered out of nothing, right? They’re based on previous work. There’s some question of attribution. What was the response? Is it, “Oh, the models did this?” Or was it the response we see to so much AI work, which basically boils down to, “Well you stole this and didn’t attribute anyone, and you’ve built on the shoulders of giants without mentioning it.” How did the response land? By and large, the reaction was generally one of being quite impressed, from the people I spoke to. As I said, these are problems mathematicians care about. There was one that drew particular attention for how they credited it and also how they’d announced all of this in their blog post. OpenAI initially had said that these are 10 problems, and there had been no progress in the last 10 years. And then if you actually read the papers, one of them quite clearly says, “Oh, we build on progress from these two researchers.” So that was later changed quite quietly. But a few of the researchers I spoke to were quite unimpressed, and they did feel it was an element of, “Well, yeah, you’ve not credited something that you’ve used heavily here, and by your own acknowledgement.” That said, one of the researchers I did speak to who was one of the ones named, and who was a bit ambivalent on the whole thing as well, so it was a real mixed bag. The general impression was that it was quite impressive. These were actual breakthroughs that bothered people, and it did move the field forward in a way that, if a human mathematician had done these, several researchers actually said that, “Well, if a researcher had done any one of these problems, they’d probably be set for an academic career.” It’s funny, credit and attribution in academia is the whole game. And it seems like the AI companies get away with being sloppy in a way that no human would be able to get away with being sloppy. Did the scale of the discovery or the work overcome the sloppiness? If a human had accomplished the same goals and had been as sloppy, would the reaction be the same? Part of me always wants to lean on the whole, “Oh, it looks like plagiarism,” element. But if you actually read the papers they produced — and one of the researchers I spoke to said there’s probably about 50 people in the world who are going to bother reading through this in depth — it is very clear. It doesn’t attempt to plagiarize. I think it was just a poor press release, to be honest. And as much as I love to go in on it sometimes, having covered science for a decade-plus, I find the press releases are often overselling what discovery has actually been made and the import of it and the novelty of it. I think that’s just another case of what happened here. Does this seem repeatable? There’s some proof that they provided that they solved 10 problems that were unsolvable. Do they provide any proof that they can solve another 10? That’s the question. So the big unknown from the near-dozen people I spoke to for this was, “Well how many did they try to get these 10?” Who knows? They know, but they won’t say. But that is the big question here. It’s unclear quite how many attempts it took to get these 10. I would be very impressed if it was the first thing they went after, and then out come these 10 impressive results. It’s unclear what areas they would focus on next and why. There are, I imagine, business reasons behind which problems they are choosing to publicize that their models can do. All the AI labs have been hiring a cohort of senior mathematicians behind the scenes, so it’s anyone’s guess as to whether they do it again. On the question of proof, math is a bit of an odd discipline in that repeatability is not the same as in the experimental sciences. A proof is a proof, and if it works, it works. The problem here is, well, can people follow through what they’ve done? Each field is quite highly specialized, so there’ll be individual mathematicians who are in those fields that go through the work. Those I spoke to that worked in some of the fields that were covered here say it all looks very legit. In math, there’s a programming and computational proving language called Lean, where you can basically codify the mathematical proofs and run them through and it, well, proves it. I keep saying prove a lot here. But it will test the rigor and the assumptions of everything going on there, and they’ve published that as well. So it does appear to hold. Whilst they may say that the press release has a lot of hype or there’s a lot of hype around it, no one I’ve spoken to seems to be doubting the essential breakthroughs that they’re claiming here. I want to stay on this subject for one more second. There’s the mathematical proof. We’ve generated a proof, and that is, as you say, just repeatable in a way that math is just logic. You can just go through the steps and say, “This proof worked,” and anybody listening to this who had to suffer through writing a proof in calculus in high school probably remembers that process. There’s something there that’s pure logic. Then there’s a part of it that is software code, as you’re describing in Lean, where you can take the pure logic, you can express it in code, and you can run it to see if it works. I understand how AI is theoretically good at all of that. You’re just going to run the reasoning, and the reasoning is going to generate some code. You’re going to run the code, you’re going to get some verifiability. We’ve seen this play out in software engineering, where the code runs or not, it’s verifiable or not, and the models can just reason out about it. Then there’s, to me, the big question that you alluded to. How many times do you have to run this? Can we verify that the models did this, and they weren’t directed by human mathematicians who’ve been hired at high rates by the labs in a way that suggests the field is going topsy-turvy? You have a quote here from James Maynard, who has won the Fields Medal, the highest prize in mathematics, who said he’s been soul-searching. I keep looking at that quote. You’ve got similar quotes from all these other mathematicians in the piece, and it seems like they’re soul-searching against a thing that hilariously they cannot verify, which is, “How did the models do this, and is that thing scalable in a way that threatens mathematics?” What do we know about how the models did this? I’d say as much as we normally do and do not know about this. There are a few issues there. One is the nature of the models. Well, this is an unreleased model, so good luck to anyone wanting to independently test it. The same goes with anything proprietary really. That said, I am inclined to almost give the benefit of the doubt that they’re not lying in some capacity about the models they’re using. As for the other part, it’s perhaps more noteworthy to ask how did we prompt the model, or how is it being guided? Is that by a mathematician who knows what they’re doing? That’s probably a key factor here, according to a lot of mathematicians I’ve spoken with who’ve tried using these. This is often the consumer models, but still it peaks to a broader landscape. They say that if you know what you’re doing and you can point things out, it’s good. Or you can use it as a tool in a way that you want, and in a way that you wouldn’t be able to if you didn’t really know how to fact-check it. I’ve had situations where I’ve had ChatGPT doing a basic sum, and I say, “That number is not right.” It responds, “Wait, so sorry. You’re right, it’s this,” and it’s still wrong. But that’s still needed at this level as well. It does allude to a broader problem. As you said, this almost soul-searching of, “Well, what if we can automate that away, and what if it gets to a point where we don’t understand it?” And that really cuts to a deeper question of, “Well, what is mathematics? Why do we do it? Why do we value it as a field?” Everyone will have different answers to that. But a fear of a lot of people I spoke to was that this might move beyond a realm of human interest, and in which case, well, maybe we just won’t engage with it. Or it’ll be something that interested people will go through, and then the rest will continue as normal. You’ve got a quote here from a researcher in Zurich named Johannes Schmitt who says, “We might be headed toward a situation where the math problems get ‘mowed down’ by AI, but we don’t actually push the field forward because humans are taken out of the loop and they’re not either checking, or they don’t understand it, or they don’t know what the future breakthroughs might be.” How likely does that feel? Is that a big concern? There is an element of the mowing down of the problems. Especially those that are used as a training field for younger mathematicians coming up and cutting their teeth, so to speak. But I also think that this idea, that math is problem-solving, is very much an outsider’s perspective of the mathematical endeavor. So a lot of the mathematicians I spoke to found that the ticking boxes part is the least interesting and valuable part of the field. The areas that are valuable for them aren’t the, “Oh, you’ve solved something, or you’ve proven something.” It’s what happens from that. I think it was James Maynard that said that the most interesting discoveries in the field aren’t that you’ve solved something, it’s what evolves from that. Sometimes those solutions open entire new fields of research that no one ever thought were possible, or, “Oh, this is a new tool that you can apply everywhere in fun and exciting ways.” I think if we look at the popularization of math — even what I’m thinking of as those theorems that people have posed — it’s the questions that endure, not the solutions. It’s always Fermat’s Last Theorem, not like, “Well, here’s the solution to whatever the last guy proposed.” I think the concern here is that well, they’re going to tick off all of these questions. Normally in the process of doing so, one would hope they would branch out into all of these new exciting areas or pose new questions. But AI won’t do that, and that’s the concern. And then that would leave the field quite sterile, and it will have all of these things that have been done, and maybe nothing left to pursue. The general consensus was, well, the jury’s out. It’s too early to tell. Even with human mathematicians, it takes a lot of time to realize the impact of these kinds of things. As I said, it’s exploded in the last six months to a year, and math is not a fast-moving discipline at the best of times. But it’s too early to tell really whether that will be a concern. But it is a concern, and a big one. You have another quote from Maynard here saying, “If the standard for a publishable paper in math is something that an AI cannot do, particularly when a PhD is typically four years, the challenge is you’re not trying to come up with a problem that AI can’t do now, it’s an AI in four years’ time.” So this is really related to the rate of improvement of the models, which as you say, particularly in math, seems to be increasing, but not at an even rate across all of the domains of mathematics. That appears to be what is causing the soul-searching. If you’re a student and you start today, and you pick some obscure domain that maybe the AI isn’t good at, sometime halfway through your PhD thesis or your PhD research, the AI will just solve it and you’ll be done. That is a real problem for you. Has the field reacted to that yet, or are they just still in the shock of, “Oh, the models can start to do things that we didn’t think they were capable of”? Yeah, I think it’s shock, really. I keep a little notebook to the side to just write down broad feelings whenever I do these interviews, and I’ve written “shell shock” in it. It’s far from universal, but it feels like it’s happened so quickly that it has just taken a lot of people by surprise. Even if they knew in theory that, well, this is coming. They’ve seen all these AI math startups going. They’ve seen colleagues moving around to different labs or areas of work. But it just happened very, very quickly. So it’s given them very little time to figure it out. It’s not necessarily even the fields that it might be good at, it’s more just like, “What can it do?” But I can’t imagine if something like this had come out when I was studying and literally in the space of half a year, it just upended what was possible. And over the summer as well. So students are possibly coming back to a completely different discipline after a break. There’s some skepticism here in the world. Gary Marcus is a reliable skeptic of AI, and he pointed out that over and over again, what you see is that AI accomplishes something in one domain and then it’s used to generalize AI’s ability across every domain. There’s a good quote from Marcus here: “As we learned a decade ago from AI’s shambolic and ultimately failed attempt to turn Jeopardy-winning Watson into a cancer-fighting machine, success in one domain does not guarantee success in all.” I can read this two ways. One, “AI has solved math,” which is not true as you’ve pointed out in several ways, all the way down to how it’s still bad at counting. And then there’s, “AI has solved math, and that means necessarily it’s going to come for everything else.” It will come for physics. It will come for law. It will come for whatever you want in the world. You can see if there’s any verifiability, AI can solve it, because you can just run it in this way. I understand both sides of that argument, that obviously success in one domain does not guarantee success in every domain. And then the arc of AI is, well, it keeps collecting domains. If there is any verifiability, it is more likely to connect those domains than not. How do you see it? I’m not entirely sure that the people that Marcus is criticizing here have actually said quite what he says they are saying. There’s a lot to say about the hype. But yeah, success in one domain does not even equal success throughout that domain, let alone in other domains. That said, there has been an undeniable trajectory in the last few years of a broadening capability increase. I don’t think you need to be on the whole AGI train to acknowledge that, and to acknowledge that that will have an impact. I think it was Andras Juhasz, one of the professors at Oxford, that I quoted in the story. Something else he’d said to me was that he’s been toying around with ChatGPT a bit, and he’s like, “I don’t think it has any geometric intuition whatsoever. Which might explain why there has been a very limited amount of progress in fields like topology.” I am in no position to verify that claim in terms of the math of it, but I think it illustrates it quite well. They call it the jagged edge. It felt like an easy argument for me: “Let’s criticize the whole ‘the singularity is near.’” That’s what Elon Musk said in response to Astra, which is a lot. I also think it’s perhaps the least generous interpretation of that argument you can take to argue against. If you take a more nuanced element that does acknowledge that there has been clear progress here, and quite quickly, and as you said, it is racking up domains, I feel there’s a trajectory there that is a reasonable one to consider, rather than just dismiss out of hand. One of the bigger arguments about AI in general is that it democratizes access. I was not a great software developer in my days trying to write software code, and now I can vibe-code apps at will to do all kinds of dumb stuff in my house. Is there a similar argument here where a bunch of people who have mathematical intuition, but did not have the formalized language or training of academic mathematics, can now access a model and push the field forward? Because that is usually the thing that undercuts the criticism from the professionals is, well, many, many more people now have access to this thing that only you had access to because of your money and your training. Annoyingly, I am going to say it’s a two-pronged thing again. But yes, in a broad sense, yes, it is. A lot of the mathematicians I spoke to were almost quite weary of this, actually. They love the idea, in theory, of democratizing access. They’re also quite fed up with AI-generated or -assisted papers that are flooding every publication imaginable, as well as the pre-print servers that they use in these fields. Some of those I spoke to said things like, “Oh, I got three emails this week alone with people being like, ‘Hey, is this legit?'” Because they thought they’d solved something with ChatGPT or with Claude, and they also don’t have the mathematical skills to check whether they’ve actually solved something. On the flip side, there are parts where they said, “Well, we’ve got a talented undergrad who’s done something that a talented undergrad would probably have never managed, and here they are doing grad-level work and they’ve produced a paper that is legit.” And in the bigger scheme of things, a few I spoke to said, “Well yeah, a lot of these are in the ivory tower. Having access to this kind of thing globally could really boost access to the kind of things here.” On the flip side, the cost. These things cost a lot to run. It’s always easy to forget when you use, say, a free version of ChatGPT or Claude or something. At the higher levels, these things cost money. They may not necessarily cost a lot of money — OpenAI claimed, I think it was $2,000 for these 10 results, but that doesn’t factor in literally anything else once they’ve got these, so it’s a very generous number. But even taking that figure, math is quite a poor discipline, even at very well-off institutions. Colva Roney-Dougal at St. Andrews, who I spoke to, said, “Well, a lot of the time I don’t bother getting a research grant. I don’t need one. I just have a blackboard.” And so if you’re not even getting a research grant, $2,000 is a lot to put up. So it could lock out researchers that way, even at quite well-funded institutions. Not to mention the speed at which this is happening, that virtually no one would’ve been able to bake any of this into a grant proposal yet, was another theme that I came across a lot. Actually, Roney-Dougal has another great quote in your piece about the nature of the AI labs and how they are talking about math. She said, “They’re treating our discipline as an advertising playground.” A bunch of mathematicians have signed something called the Leiden Declaration, which is an open letter to then pledge not to buy into hype around AI. These things are running right at each other. The AI labs are not going to stop using every discipline as an advertising playground. And a bunch of mathematicians saying, “We refuse to buy the hype,” certainly does not seem to be stopping the hype. There’s just a piece of this that is organized professional resistance to a thing that is upending a field that has, as you say, been pretty cheap to operate, and now might be getting cheaper or easier to access or easier to upend, day by day. Do mathematicians feel that that is going to be effective? Historically, mathematicians are not savvy political operators. There’s a part of me that says, “Oh, they’re just going to get run over.” I don’t know. In the history of math, actually, I think a lot of them were quite savvy. Isaac Newton is the one that always comes to mind for that — though quite a petty political operator as well. But yeah, that is the fear. A few that I spoke to, and one really comes to mind, mentioned that there’s often this belief that math is the pinnacle of knowledge. But he was like, “Well, that’s bullshit.” And he wasn’t alone in illustrating that sentiment. But it is good for showcasing, and it’s a lot neater as a discipline, and a lot cheaper. You mentioned Marcus referencing IBM’s Watson and the curing cancer ambition. Well, that involves lots of messy experiments, including on people. You don’t need that in math, so it’s a really easy discipline to come in, throw your weight around, and then move to somewhere more lucrative if that’s what you want. I’m not saying that that’s what they’re doing. A lot of the people at these companies have been hired. I don’t doubt their credentials for sure, and I don’t doubt their motivations as well. It does raise a question long-term as to how viable this is. Because let’s be clear, as a field goes, I cannot imagine mathematicians being a very lucrative enterprise customer for these companies. The thing that might be lucrative is pushing a field forward to turn it into something economically viable. We push mathematics forward as a field, that turns into some engineering or physics breakthrough based on that mathematics, and that turns into, I don’t know, yet another way to launch rockets. Some circle happens there that I don’t quite understand, but that is the history of innovation, from research, to engineering, to products or services that make money. Is that on the minds of any of these mathematicians, that pushing the boundaries here is upstream of something radically economically lucrative? The immediate counter that would come to mind here is that a lot are scared that it’s closing off the field. So by definition, those breakthroughs that lead to something surprising and new that you can say, “Oh, this works here,” may not be happening anymore. If anything, that lucrative endeavor of applying math to this entire new field that may have a lot of money in it remains an open question as to whether anything like that would be possible if we’re closing off avenues, rather than opening them up. Right, if the economic incentive of solving the unsolved problem is reduced because you personally won’t get rich if a computer is just solving every unsolved problem, something very fundamental breaks there. A lot of it comes down to the fact that it’s not just solving problems. With a lot of these things, as we said, it’s about what solving that problem tells you elsewhere. If these were very lucrative problems to be solving, I imagine that more people would be trying to solve them than have left them for decades. This is the nature of a lot of pure science, and it’s a broader criticism of what is going on perhaps with the Trump administration’s approach to science policy at the moment, in that it’s very applications-focused. There is something to doing pure research that can yield potentially very big dividends that is, by definition, utterly unpredictable as well.You cannot plan for it. The fear I think with math is that in solving all of these problems, and then also doing so without opening up new areas of research, what are you left with? Even if it’s from a more lucrative, “What are you going after,” point of view, if you’re not opening up new areas of research and you’re just ticking off old ones, it just leaves a big question mark as to what might be left in its wake. Even from those I spoke to that were very excited about what’s happening, they said that even they don’t really know what’s happening. And they’re excited from a personal level because, “Oh, we might be able to do this, might be to do that.” But there was still this lingering uncertainty of like, well, where does this leave the field? Especially for more pure disciplines like research mathematics, it’s tougher to say what comes next. Because in a lot of the other sciences you can say, “Well okay, well they’d shift onto more engineering problems, or applying that.” But if you solve all the problems at the ground and there’s nothing being built up from that, where do you go from there? One great thing about The Verge is our commenters are vast, they’re very knowledgeable, and there was a comment from a mathematics researcher on your story that I just want to read to you, and see if you think this is the right framework. Here it is: “I have no doubt these models will bring massive change in the field, but in their current state, they won’t yet drive us to obsolescence. Just occupy a particularly useful spot in our bag of tricks. My apprehension comes from not knowing where these things will peak, but overall I remain optimistic. I think AI will be a net boon for math when used properly.” I feel like “AI will be a net boon for X when used properly” is just where you land in life in a lot of things, but that’s the most optimistic response that I’ve heard: “If we get it right, it’s going to be great.” Is that the vibe, or is it still more shell-shocked than that? I’d say shell shock is still the overriding impression. I think perhaps the gut response to that is like, “Well, it will be a net positive for whom, and what is ‘properly’?” All of those are quite legitimate questions here. There were some bleak responses from graduate students I saw in essays posted online.Where’s their place in this as future researchers? Do they have a place in this? Is it as glorified AI proof checkers? That will be quite an unsatisfying career, I imagine. Or maybe not, I don’t know. We will see. I think anything used properly will be a net boon. But yeah, I think it all comes down to what “properly” means, and for whom we’re talking about. I feel like, oddly, that is an excellent place to leave it. Because I don’t think either one of us knows, and I suspect over the next year or so, things will come into focus. Because at some point, OpenAI will have to show people how they did the things of the models. And perhaps more importantly, the other labs are going to want to either replicate these results or show that they can push farther, which will necessarily have to lead to a little bit more transparency and yet more mathematicians having a crisis with you. Robert, thank you so much for being on the show. We’ll have you back very soon. Thank you for having me. Questions or comments? Hit us up at [email protected]. We really do read every email!
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Today on Decoder, I’m talking with Robert Hart, The Verge’s London-based AI reporter, about what AI is doing to the field of mathematics and the existential crisis many lead mathe…
Tools What you need to be a great vibe coder ListMap FeaturedVibe Costs · tracks what vibe coders spend on subscriptions, APIs and hostingVisit ↗ No.AppBuilderCategory 01 Claude Code Anthropic's agentic coding harness t…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Tools What you need to be a great vibe coder ListMap FeaturedVibe Costs · tracks what vibe coders spend on subscriptions, APIs and hostingVisit ↗ No.AppBuilderCategory 01 Claude C…
Slack is introducing dedicated channels where teams can vibe code together with AI agents instead of jumping between different tools and conversations. The Slack Code launch includes open, project-specific code channels with dedicated user tabs, alongside features that compare coding changes and preview HTML output before the project is shipped. "With Slack Code, when you have an idea or need to build a new feature, update a web page, or fix a bug, you simply tag in a coding agent like Anthropic's Claude or Cognition's Devin, and that agent then spins up a code channel to tackle the task," Slack said in its press release. "There, everyone … Read the full story at The Verge.
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Slack is introducing dedicated channels where teams can vibe code together with AI agents instead of jumping between different tools and conversations. The Slack Code launch inclu…
This tutorial provides an end-to-end workflow for fine-tuning language models using Direct Preference Optimization (DPO). We demonstrate how to audit the Anthropic HH-RLHF dataset for structural and length-based biases, implement a robust training pipeline using TRL and LoRA, and evaluate model performance to ensure genuine preference learning rather than reliance on lexical shortcuts. The post Auditing Preference Biases and Fine-Tuning Language Models with Direct Preference Optimization on Anthropic HH-RLHF Using TRL and LoRA appeared first on MarkTechPost.
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
This tutorial provides an end-to-end workflow for fine-tuning language models using Direct Preference Optimization (DPO). We demonstrate how to audit the Anthropic HH-RLHF dataset…
arXiv:2608.18227v1 Announce Type: new Abstract: Push-T is an iconic benchmark for learning manipulation policies from human demonstrations. The robot must use a single point of contact to push a T-shaped block into a target pose. In this short paper, we revisit the Push-T task in the context of emerging advances in Agentic Robotics where an LLM coding agent -- Claude Code with Fable 5 -- is prompted to create an algorithmic solution that does not require any demonstration data. We study how effective the agentic coding loop can solve the Push-T task, and compare the resulting code as policy with the visuomotor imitation learning policy. Results suggest that the agent found the 2D gym simulation online, and used sim experiments to learn push mechanics, iteratively optimizing to achieve 100% success rate using 46% fewer steps than the best diffusion policy trained with 200 human demonstrations. The coding agent also solve extensions from T to the full alphabet (Push-A to Push-Z) using a self generated curriculum and generated simulation code for the Franka and UR5 robot arms in 3D cross-embodiment simulations with visual feedback. Videos, policies and details will be posted online.
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
arXiv:2608.18227v1 Announce Type: new Abstract: Push-T is an iconic benchmark for learning manipulation policies from human demonstrations. The robot must use a single point of co…
Job Rietbergen Aug 19, 2026 Alibaba released Qwen3.8-Max on August 2, and it debuted at #4 on Arena.ai’s Frontend Code leaderboard, one spot above Claude Fable 5. It is the second model from a Chinese lab to pass Fable…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
Job Rietbergen Aug 19, 2026 Alibaba released Qwen3.8-Max on August 2, and it debuted at #4 on Arena.ai’s Frontend Code leaderboard, one spot above Claude Fable 5. It is the second…
The music industry spent years asking AI companies nicely. Now it’s asking federal judges. Independent publisher Round Hill Music filed separate copyright infringement suits against Anthropic and Suno in U.S. District C…
AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
The music industry spent years asking AI companies nicely. Now it’s asking federal judges. Independent publisher Round Hill Music filed separate copyright infringement suits again…