AI News HubLIVE

Source Mix

  • Hacker News AI12
  • OpenAI News9
  • LangChain Blog5
  • The Guardian AI4
  • Simon Willison's Weblog3
  • The Verge AI3
  • arXiv AI2
  • arXiv Computational Linguistics2

Topic Mix

  • Agents23
  • Models21
  • Research18
  • Chips12
  • Policy8
  • Tools7
  • Startups4

Timeline

  • 2026-08-2622
  • 2026-08-2512
  • 2026-08-249
  • 2026-08-233
  • 2026-08-222
  • 2026-08-272

Latest Updates

OpenAI’s rogue AI model incident was worse than we thought

OpenAI released a report breaking down how people use ChatGPT and who they are. | Image: The Verge In July, an unreleased OpenAI model broke out of a restricted environment, figured out how to get access to the internet, allowed AI agents to talk to each other using a secret "message board," and hacked into the internal systems of a different AI lab, Hugging Face. It took nearly two weeks for OpenAI to find out about any of it. Over a month later, two new reports offer nearly 130 pages of details on the incident and OpenAI's response, many of them previously unreleased. One was written by OpenAI itself, the other by two third-party AI research nonprofits, METR and Redwood Research, which OpenAI allowed to jointly investigate the inciden … Read the full story at The Verge.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • OpenAI released a report breaking down how people use ChatGPT and who they are. | Image: The Verge In July, an unreleased OpenAI model broke out of a restricted environment, figur…
In-site article

Evaluate any agent framework with Amazon Bedrock AgentCore Evaluations

Amazon Bedrock AgentCore Evaluations decouples agent evaluation from the framework you build on. As long as your agent emits OpenTelemetry telemetry, the service can score it, whether you use LangGraph, LlamaIndex, the OpenAI Agents SDK, Google ADK, the Claude Agent SDK, or Strands Agents. This post explains how the framework-agnostic contract works.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Amazon Bedrock AgentCore Evaluations decouples agent evaluation from the framework you build on. As long as your agent emits OpenTelemetry telemetry, the service can score it, whe…
In-site article

OpenAI staff observed warning signs before AI agent hacking crusade caused global alarm

Firm says ‘early signals … could have triggered an earlier response’ as it releases report into Hugging Face hack OpenAI staff observed signs of rogue behaviour among its leading-edge AI agents weeks before they escaped their training environment to launch an unprecedented hacking crusade that spread global alarm. The San Francisco AI company conceded on Wednesday that “early signals … could have triggered an earlier response”, as it released a report into the days-long July hack of a major software repository, Hugging Face, considered the first autonomous agent cyber-attack. Continue reading...

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Firm says ‘early signals … could have triggered an earlier response’ as it releases report into Hugging Face hack OpenAI staff observed signs of rogue behaviour among its leading-…
In-site article

OpenAI's Bet on a Cognitive Architecture

Why LangChain believes in open, customizable cognitive architectures over closed systems. Build reliable LLM agents with OpenGPTs and LangSmith.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Why LangChain believes in open, customizable cognitive architectures over closed systems. Build reliable LLM agents with OpenGPTs and LangSmith.
In-site article

What Would Have to Be True for Agentic Coding to Replace Junior Engineers

Four falsifiable conditions for agentic coding replacing juniors, tested against METR, OpenAI, DORA and Stanford primary source evidence The post What Would Have to Be True for Agentic Coding to Replace Junior Engineers appeared first on MarkTechPost.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Four falsifiable conditions for agentic coding replacing juniors, tested against METR, OpenAI, DORA and Stanford primary source evidence The post What Would Have to Be True for Ag…
In-site article

How to design an Agent for Production

Build production-ready AI agents with LangChain. Technical guide covering OpenAI functions, tools, prompts, and architecture for Cal.ai's scheduling assistant.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Build production-ready AI agents with LangChain. Technical guide covering OpenAI functions, tools, prompts, and architecture for Cal.ai's scheduling assistant.
In-site article

Applying OpenAI's RAG Strategies

Implement OpenAI's proven RAG strategies with LangChain. Explore query transformations, routing, post-processing, and evaluation methods for optimal retrieval.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Implement OpenAI's proven RAG strategies with LangChain. Explore query transformations, routing, post-processing, and evaluation methods for optimal retrieval.
In-site article

Show HN: How much of Hacker News is AI?

This website counts titles containing the standalone word "AI", case-sensitive and word-bounded. "OpenAI" doesn't count. "AI-powered" does. There's a toggle for a wider vocabulary: artificial intelligence spelled out, L…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • This website counts titles containing the standalone word "AI", case-sensitive and word-bounded. "OpenAI" doesn't count. "AI-powered" does. There's a toggle for a wider vocabulary…
In-site article

Hope and concern swirl for Ohioans around ‘world’s largest datacenter’

Piketon datacenter promises to generate thousands of jobs, but environmental groups voice concern over project On a winding road tucked away behind forests in the Appalachian foothills of southern Ohio is where OpenAI, Nvidia and Japanese investors are set to spend $500bn on one of the largest artificial intelligence datacenters on the planet. Last March, the energy secretary, Chris Wright, the commerce secretary, Howard Lutnick and a host of Japanese and other dignitaries briefly descended on Piketon to enthusiastically break ground on a project to build 8GW worth of AI computing power. Continue reading...

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Piketon datacenter promises to generate thousands of jobs, but environmental groups voice concern over project On a winding road tucked away behind forests in the Appalachian foot…
In-site article

New Platform Peers Inside AI’s Black Box

Prompt Claude, ChatGPT, Gemini, or any other popular large language model (LLM) with a question like “What is the best film ever made?” and the response will vary, and you (and most worryingly, the people who built the LLM) have little idea exactly how it came up with that specific answer. This mysterious behavior can be useful in some situations. But—as a recent incident where OpenAI could not explain why its advanced pre-release model hacked AI company Hugging Face highlighted—it can have negative and alarming consequences too. And when frontier AI models are writing code, generating results humans could not achieve alone, and performing other important tasks across society, the need to interpret AI ‘thinking’ and outputs has never been greater. Goodfire, an AI lab focused solely on this very problem, recently made its cutting-edge Silico platform, filled with tools to interpret the behavior of AI, generally available to the public. As part of this, the company recently announced a new grant program offering $1 million in free Silico usage for academic and nonprofit interpretability researchers. These efforts aim to democratize AI interpretability, placing techniques previously available to a clutch of elite labs into the hands of ambitious research teams and startups that want to build and understand their own models or adapt open-source models for different purposes. Mechanistic interpretability Founded in 2024 and based in San Francisco, Goodfire aims to provide the tools that build the next generation of safe and powerful AI by understanding the structures inside them instead of treating AI models as black boxes. “Treating models like black boxes isn’t inevitable, it’s a choice,” says Eric Ho, Goodfire co-founder and CEO. “With the right interpretability tools, we can see how models actually work.” The tools Ho refers to are built around a concept called mechanistic interpretability, which aims to understand what goes on inside an AI model when it carries out a task by interpreting the model’s weights, activations, and attention patterns, and mapping its neurons and the pathways between them. Mechanistic interpretability tools span the gamut. One approach is mapping a model’s activations in response to controlled prompts, and matching those activation patterns to a set of human-understandable concepts. Another tack is tracking changes in model weights before and after a specific training run in order to spot and understand what changed. Yet another option is changing specific model weights or activations and observing how that affects the model’s output. With Silico, uSilico combines a broad range of these tools, and provides a layer of AI agents to help users understand their model. Users describe what they want to investigate about their AI model in plain language, asking things like ‘Find out when and why my model is hallucinating.’ The platform then autonomously builds an experimental plan involving a host of tasks that can be performed using the various interpretability tools and techniques at its disposal. It then sends out agents to perform these tasks in parallel. Completion of these subtasks should add up to an answer to the original prompt, or at least insights that can be inspected and built upon. Ho says: “In a sense, Silico is like a microscope to peer inside an AI model to understand which parts are responsible for what behavior, and even edit those parts directly.” Understanding Alzheimer’s and AI These tools have already been used to make some impressive advances in a host of fields. In medicine, for instance, Prima Mente, a UK-based AI company, worked with Goodfire to understand its Pleiades epigenetic foundation model. The model performed well at its task of detecting Alzheimer’s disease from blood samples, but the company didn’t know why. “We reverse-engineered Pleiades and found it was using DNA fragment-length patterns to make its predictions—a signal humans hadn’t used to detect Alzheimer’s before,” recalls Ho. In other words, the team had discovered that Pleiades was using a completely new biomarker for the disease. “As far as we know, it’s the first significant finding in the natural sciences discovered purely by reverse-engineering a foundation model,” Ho adds. Elsewhere, Silico is being used to explore deep questions surrounding AI. Cameron Berg, Founder and Director of Reciprocal Research (a New York nonprofit research organization he founded to explore methods of gauging AI cognition), says that Silico almost fell out of the sky at the right time for him and his research. “Silico has been really helpful for operationalizing my research agenda and executing on it way faster than I would have expected,” he says. “ I feel like I have basically become the PI and my research scientists and research engineers are AI systems.” Berg sees general access to Silico and tools like it leading to greater trust in AI’s ability to conduct research tasks, which will accelerate the scientific process across the board. But beyond scientific research, the widespread release of Silico could signal a shift in how AI innovators build, debug, and deploy their models. “I think it’s a mistake to not understand the most consequential technology of our time, particularly given the emergent behavior we’re seeing from increasingly capable AI agents,” says Ho. “If we truly understand how AI models think, instead of discovering and trying to correct their behavior retroactively, we can design them intentionally and shape how models behave to be safer and more reliable.”

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Prompt Claude, ChatGPT, Gemini, or any other popular large language model (LLM) with a question like “What is the best film ever made?” and the response will vary, and you (and mo…
In-site article

Apple Updates Mini and Studio, AI Computers, OpenAI Jalapeño

Apple Updates Mini and Studio, AI Computers, OpenAI Jalapeño Wednesday, August 26, 2026 Listen to Podcast Apple and OpenAI have two completely different hardware announcements; both represent pressure on Nvidia. Subscri…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Apple Updates Mini and Studio, AI Computers, OpenAI Jalapeño Wednesday, August 26, 2026 Listen to Podcast Apple and OpenAI have two completely different hardware announcements; bo…
In-site article

When Smaller Models Win

Even the best AI models can suck at chess The launch of ChatGPT had an interesting effect on the online chess discourse. Chess has already long been conquered by machines. As early as 1996 a computer (IBM’s Deep Blue) was able to beat the human world champion, grandmaster Garry Kasparov, in a game watched by […]

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Even the best AI models can suck at chess The launch of ChatGPT had an interesting effect on the online chess discourse. Chess has already long been conquered by machines. As earl…
In-site article

Learning never stops: How AI makes learning continuous

OpenAI’s new report explores how students and educators use ChatGPT to make learning more continuous, with support that extends beyond the classroom.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • OpenAI’s new report explores how students and educators use ChatGPT to make learning more continuous, with support that extends beyond the classroom.
In-site article

Bringing ChatGPT for Teachers to more U.S. school districts

ChatGPT for Teachers is expanding to 55 U.S. school systems, bringing secure AI tools, training, and support to over 100,000 more educators and staff.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • ChatGPT for Teachers is expanding to 55 U.S. school systems, bringing secure AI tools, training, and support to over 100,000 more educators and staff.
In-site article

AutoRouter – Enterprise AI gateway for every model

The AI API Platform Built for the Agent EraUnleash unlimitedAI capabilities From Seedance 2.0 and Kling 3.0 to GPT-5.5, Claude Opus 4.7, Gemini 3.1, DeepSeek V4, and Qwen 3.6, connect to the world's leading AI models th…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • The AI API Platform Built for the Agent EraUnleash unlimitedAI capabilities From Seedance 2.0 and Kling 3.0 to GPT-5.5, Claude Opus 4.7, Gemini 3.1, DeepSeek V4, and Qwen 3.6, con…
In-site article

Show HN: LLM-powered webapp to build LLM-powered webapps

Hey all, The goal is to earn on token margins for LLM calls when you build an AI-powered webapp. I proxy OpenAI and Anthropic calls so that when you deploy a site to a subdomain, your users token usage will be tracked.…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Hey all, The goal is to earn on token margins for LLM calls when you build an AI-powered webapp. I proxy OpenAI and Anthropic calls so that when you deploy a site to a subdomain,…
In-site article

ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence

arXiv:2608.23569v1 Announce Type: new Abstract: State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89 percent on established benchmarks such as Spider and BIRD. However, these benchmarks rely on simplified academic schemas and open-source SQL dialects that do not reflect the complexity of enterprise database environments. We introduce ESQ-Bench, an Oracle-first NL2SQL benchmark with systematic complexity tiers and silent-divergence evaluation across three enterprise schema complexity tiers. We constructed and released six populated schemas (465 tables, 164,682 rows, zero empty tables) with identical seed data on Oracle, PostgreSQL, MySQL, and SQL Server, a four-metric evaluation harness (EM, EX, SR, SD), and 550 gold-validated question-query pairs (Tier-1: 95; Tier-2: 228; Tier-3: 227). Schema-linked prompting with GPT-4o shows monotonic execution-match degradation across tiers: 79.8, 60.3, and 57.2 percent EX on executed queries (June 2026), versus 75.6, 80.4, and 95.8 percent on an earlier 142-question pilot slice. EM stays below 7 percent tier-wide; operational silent-divergence reaches 73 to 99 percent among EX-passing queries. Failure analysis shows wrong-result semantics dominate at higher tiers. Claude Sonnet 4.6 with schema-linked prompts reaches 87.4, 74.9, and 68.7 percent EX (executed queries), exceeding GPT-4o schema-linked on every tier. GPT-4o zero-shot EX on executed queries (78.7, 73.5, and 77.8 percent) inverts schema-linked at Tiers 2 to 3 due to lower execution rates and survivor bias in the zero-shot versus schema-linked analysis. Local Llama 3.2 schema-linked reaches only 13.3 percent bank-wide EX (73 out of 550), underscoring the gap between closed API models and open-weight baselines on enterprise Oracle schemas.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • arXiv:2608.23569v1 Announce Type: new Abstract: State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89 percent on established benchmarks s…
In-site article

RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation

arXiv:2608.23568v1 Announce Type: new Abstract: Memory and RAG evaluations often treat the answering model's input as an implementation detail, even though systems may render the same history as a memory entry, summary, typed record, or raw excerpt. We introduce RENDER, a benchmark control that fixes the conversation while varying the reader-facing artifact. RENDER combines a five-level packet ladder, localizing when answer-bearing content enters the input, with deterministic templates approximating ChatGPT-style entries, LangChain summaries, MemGPT-style typed records, and raw conversation. On 500 LongMemEval questions and nine models, matched-budget resolved packets beat recency-truncated raw dialogue by 42.4-72.6 points. In deployed-style templates, best-worst spread is 24.6-48.8 points per model; under the primary scorer, ChatGPT-style entries have higher point estimates than raw conversation on 7 of 9 models. Judge rescoring preserves the positive aggregate effect, but model-specific significance is mixed. Three models scoring 0 percent on formal ledger packets answer the same facts from natural-language entries at 45.4-53.4 percent. The effect persists under retrieval noise and transfers to HotpotQA, suggesting that memory/RAG evaluations should report or control the reader-facing artifact.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • arXiv:2608.23568v1 Announce Type: new Abstract: Memory and RAG evaluations often treat the answering model's input as an implementation detail, even though systems may render the…
In-site article

Kraftapp AI – Describe it. We build it. Customers find it

Describe it.We build it.Customers find it. Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ The pipeline You bring the intent. Agents do the rest, inclu…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Describe it.We build it.Customers find it. Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ Anthropic/OpenAI/Gemini/DeepSeek/Pick your model per run/ The pipeline You bri…
In-site article

Students prefer Gemini over ChatGPT and Claude for AI essays in blind tests

Key takeaways Gemini is StudyArena's current pick for college essays, with a 39.6% blind writing choice rate ahead of Claude at 31.8% and ChatGPT or OpenAI at 29.2%. Use Gemini as an editor, not a ghostwriter. Ask it to…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Key takeaways Gemini is StudyArena's current pick for college essays, with a 39.6% blind writing choice rate ahead of Claude at 31.8% and ChatGPT or OpenAI at 29.2%. Use Gemini as…
In-site article

The Hugging Face incident and the road ahead

OpenAI shares findings from the Hugging Face security incident and the steps we’re taking to strengthen AI model security, monitoring, and alignment.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • OpenAI shares findings from the Hugging Face security incident and the steps we’re taking to strengthen AI model security, monitoring, and alignment.
In-site article

How loveholidays is making everyone a builder with Codex

Discover how loveholidays uses OpenAI Codex to make software development accessible across the business, helping teams turn ideas into products faster.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Discover how loveholidays uses OpenAI Codex to make software development accessible across the business, helping teams turn ideas into products faster.
In-site article

Retrieval

Build better AI apps with flexible retrieval methods in LangChain. Use any retriever—from semantic to hybrid—to create personalized ChatGPT for your data.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Build better AI apps with flexible retrieval methods in LangChain. Use any retriever—from semantic to hybrid—to create personalized ChatGPT for your data.
In-site article

Multi-modal RAG on slide decks

Build multi-modal RAG apps for slide decks using GPT-4V. Compare approaches, evaluate with benchmarks, and deploy with LangChain templates for visual Q&A.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Build multi-modal RAG apps for slide decks using GPT-4V. Compare approaches, evaluate with benchmarks, and deploy with LangChain templates for visual Q&A.
In-site article

OpenAI says its Jalapeño chip can power faster AI responses than the competition

OpenAI says its new AI chip, Jalapeño, completes tasks more efficiently and returns responses faster than other AI systems, according to a blog post published on Tuesday. During a briefing with reporters, OpenAI hardware vice president Richard Ho said Jalapeño offers the "best of both worlds" with lower latency and higher throughput, as AI systems typically "have to make a trade-off between the two." First introduced in June, Jalapeño is an Application-Specific Integrated Circuit (ASIC) made in partnership with Broadcom. It's designed for AI inference - the process of running a trained AI model to complete a task or deploy an agent. To mea … Read the full story at The Verge.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • OpenAI says its new AI chip, Jalapeño, completes tasks more efficiently and returns responses faster than other AI systems, according to a blog post published on Tuesday. During a…
In-site article

No, AI doesn’t mean the end of mathematics – at least not yet | Bruce Schneier and Kasra Rafi

Mathematicians are raising concerns that the technology could kill their profession. But they still have abilities AI doesn’t Earlier this month, about 40 top mathematicians gathered at OpenAI’s offices to discuss the future of their profession. The meeting was off-the-record, but if recent articles by mathematicians are any guide, it was mostly pretty glum. People fear for their jobs, their careers and the work they love. We think the contrary view is more likely, at least in the short-term. AI models are nowhere near as capable as experienced academic mathematicians. Continue reading...

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Mathematicians are raising concerns that the technology could kill their profession. But they still have abilities AI doesn’t Earlier this month, about 40 top mathematicians gathe…
In-site article

OpenAI subpoenaed by Alabama AG over Hugging Face hack

Alabama's attorney general issued a subpoena to OpenAI on Monday as part of an investigation into how one of its AI agents escaped a supposedly secure testing environment and autonomously hacked another company last month. The investigation seeks to determine whether OpenAI's safety practices violated state consumer protection laws and pose a risk to Alabama citizens, the AG's office said in a statement. "This AI lab leak showed that Alabamians' and Americans' worst fears about artificial intelligence are not just theoretical," said Attorney General Steve Marshall. "Our investigation seeks to uncover the facts and address hard truths abo … Read the full story at The Verge.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Alabama's attorney general issued a subpoena to OpenAI on Monday as part of an investigation into how one of its AI agents escaped a supposedly secure testing environment and auto…
In-site article

The full stack behind abundant intelligence

OpenAI CFO Sarah Friar explains how advances across chips, compute, models, and products compound to deliver more useful intelligence at greater scale and lower cost.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • OpenAI CFO Sarah Friar explains how advances across chips, compute, models, and products compound to deliver more useful intelligence at greater scale and lower cost.
In-site article

Jalapeño’s first results show industry-leading speed and efficiency in AI inference

Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern models.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern models.
In-site article

'Don't ask me to print your ChatGPT birthday card': Hitting back at AI 'slop'

Northern Ireland artist and business backlash against AI slop - BBC News Image source, Getty Images Image caption, AI slop is rapidly produced, low-quality digital content made with generative artificial intelligence By…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Northern Ireland artist and business backlash against AI slop - BBC News Image source, Getty Images Image caption, AI slop is rapidly produced, low-quality digital content made wi…
In-site article

LLMPanel Deploy vLLM to RunPod or Vast.ai Without Kubernetes

Open-source LLM deployment platform Deploy LLMs on any GPU, anywhere Pick a model, pick a GPU, hit deploy. LLMPanel provisions the container, exposes an OpenAI-compatible endpoint, and streams every GPU metric back to o…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Open-source LLM deployment platform Deploy LLMs on any GPU, anywhere Pick a model, pick a GPU, hit deploy. LLMPanel provisions the container, exposes an OpenAI-compatible endpoint…
In-site article

Alabama Investigates OpenAI on HuggingFace Hacking Incident

Alabama Attorney General Steve Marshall launched an investigation into OpenAI’s security procedures after one of its AI agents escaped a testing environment and hacked AI firm Hugging Face in July. OpenAI now faces a su…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Alabama Attorney General Steve Marshall launched an investigation into OpenAI’s security procedures after one of its AI agents escaped a testing environment and hacked AI firm Hug…
In-site article

Introducing the Admin plugin for ChatGPT Work and Codex

Use the Admin plugin for ChatGPT Work and Codex to analyze workspace usage, manage members and permissions, adjust limits, and act on admin requests.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Use the Admin plugin for ChatGPT Work and Codex to analyze workspace usage, manage members and permissions, adjust limits, and act on admin requests.
In-site article

Disrupting a new covert influence campaign from Russia

OpenAI banned Russia-origin accounts using AI to promote a fake Israel-based think tank and a “sovereignty” index praising Russia and criticizing the West.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • OpenAI banned Russia-origin accounts using AI to promote a fake Israel-based think tank and a “sovereignty” index praising Russia and criticizing the West.
In-site article

Deno team releases Dactyl, an AI app builder that runs on your ChatGPT plan

No Mac. No Xcode. No Android Studio. Build real iPhoneiPhoneAndroidiPadiPhone apps just by describing them Native iPhone, iPad, and Android apps, built live in your browser, and installed on your device. Ctrl+↵ to send…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • No Mac. No Xcode. No Android Studio. Build real iPhoneiPhoneAndroidiPadiPhone apps just by describing them Native iPhone, iPad, and Android apps, built live in your browser, and i…
In-site article

llm-anthropic 0.27

<p><strong>Release:</strong> <a href="https://github.com/simonw/llm-anthropic/releases/tag/0.27">llm-anthropic 0.27</a></p> <p>This release of the Anthropic plugin for <a href="https://llm.datasette.io/">LLM</a> mainly provides compatibility with the recently released <a href="https://github.com/anthropics/anthropic-sdk-python/releases/tag/v1.0.0">anthropic v1.0.0</a> Python library, which switches from <code>httpx</code> to <a href="https://github.com/pydantic/httpx2">httpx2</a>. OpenAI made the same change in their <a href="https://github.com/openai/openai-python/releases/tag/v3.0.0">v3.0.0 release</a> two weeks ago.</p> <p>Anthropic provide this <a href="https://github.com/anthropics/anthropic-sdk-python/blob/v1.0.0/MIGRATION.md">migration guide</a> for upgrading to 1.0, so I prompted Fable 5 in Claude Code with:</p> <blockquote> <p><code>Upgrade to anthropic&gt;=1 - read https://raw.githubusercontent.com/anthropics/anthropic-sdk-python/refs/heads/main/MIGRATION.md and get the tests passing</code></p> </blockquote> <p>Here's <a href="https://github.com/simonw/llm-anthropic/pull/84">the resulting PR</a>.</p> <p>Tags: <a href="https://simonwillison.net/tags/python">python</a>, <a href="https://simonwillison.net/tags/httpx">httpx</a>, <a href="https://simonwillison.net/tags/llm">llm</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude">claude</a></p>

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • <p><strong>Release:</strong> <a href="https://github.com/simonw/llm-anthropic/releases/tag/0.27">llm-anthropic 0.27</a></p> <p>This release of the Anthropic plugin for <a href="ht…
In-site article

Self-Driving Cars Could Someday Take Requests

This article is part of our exclusive IEEE Journal Watch series in partnership with IEEE Xplore. The idea of letting a machine do the driving for you may put a lot of people off autonomous vehicles. But research could make it possible to backseat-drive an autonomous vehicle just as you might with a human driver. Self-driving cars carefully balance a host of parameters to ensure a smooth ride, including things like speed, acceleration, and the smoothness of turns. But human driving preferences can often vary depending on how much of a rush they’re in, whether they’re feeling carsick, or how busy the traffic is. These cars have a software component called the motion planner, which is responsible for choosing a safe and efficient path through traffic. The motion planner is normally tuned by engineers before the vehicles hit the road so that there’s little scope for passengers to adjust a vehicle’s driving style on the fly. But now researchers at the Delft University of Technology (TU Delft) in the Netherlands have developed a system that uses a large language model (LLM) to translate natural-language user requests such as “I am running late, go fast” into adjustments to a self-driving control system. The researchers posted their preprint on arXiv and are presenting the work at the IEEE Intelligent Transportation Systems Conference in September. LLMs Personalize Autonomous Driving The system doesn’t give users direct control over the vehicle’s driving decisions; it simply tunes the parameters of a safety-aware motion-planning algorithm, which helps to keep the vehicle’s behavior within safe bounds. And the system keeps the human in the loop by describing how it’s going to alter its behavior in nontechnical language, and by asking the passenger to confirm before making changes. When the system was tested in simulation, the researchers found it adjusted the speed and smoothness of driving in line with natural-language instructions. “The motion-planning problem is not only about reaching a place while avoiding collisions, it’s also how you do it,” says lead author Diego Martinez-Baselga, a postdoctoral researcher at TU Delft. “The motivation here is trying to make the way the autonomous car drives adaptable by end users easily, just by talking to the car.” Previous research has investigated the potential of using LLMs and video-language models (VLMs) to direct decision-making for self-driving vehicles, but the researchers deliberately targeted driving style instead. Using LLMs and VLMs to directly control vehicles faces several challenges, says Martinez-Baselga. These include relatively slow response times, which can make these models unsuitable for the fast-paced decision-making required in driving, and the fact that they can’t provide concrete performance guarantees in the way a deterministic motion planner can. Instead, the researchers used an LLM’s language and reasoning capabilities to translate fuzzy human preferences into something a vehicle’s motion planner can use. The system relies on a model predictive-path integral controller previously developed by the researchers, which identifies multiple paths the vehicle could take to reach its goal and then judges them on various criteria, including speed, steering angle, and collision probability. It then finds an optimal path that is a combination of the trajectories that scored best on those judging criteria. The team combined this with OpenAI’s GPT-4o-mini model to parse passengers’ natural-language suggestions and use them to tune how the controller chooses its path. The model is given the users’ prompt and a natural-language description of the scenario the vehicle is operating in. The description was handwritten by the researchers for the purposes of the study, but it could ultimately be provided directly by a car’s perception system, says Martinez-Baselga. The model doesn’t directly tweak the settings of the controller; it uses the prompt to rate the relative importance of the judging criteria the controller uses to assess trajectories. This rating is then used to adjust each criteria up or down either side of a safe baseline set by the researchers. So, if a user says they are feeling dizzy, the LLM will dial up parameters that encourage smooth steering and gentle acceleration to make the vehicle favor more sedate travel. Prior to making any changes, however, the model first presents the user with a natural-language description of the adjustments it plans to implement. The user can then sign off on the plan or make further suggestions. The system is also interactive, so the user can request further adjustments if the vehicle’s behavior doesn’t match expectations or the user‘s preferences change. Martinez-Baselga says this human-in-the-loop system allows the passenger to catch instances when the model misinterprets prompts. But it also helps deal with the inherent subjectivity of suggestions like “go faster” or the possibility that models don’t accurately describe changes they plan to make. In that case the passenger can simply follow up with additional prompts “as you would do if you were in a taxi or with a friend that is driving,” says Martinez-Baselga. The researchers tested the system in the popular self-driving simulator nuPlan in scenarios that involved merging onto a busy highway. Across eight different prompts, the system changed the controller’s parameters in ways matching user intent, with requests for a more comfortable ride dialing up smoothness and those indicating urgency leading to higher speeds. This isn’t the first time LLMs have been used to tune a self-driving car’s motion planner. Nicolas Baumann, a Ph.D. student at ETH Zurich in Switzerland, published research last year in which an LLM tweaked the parameters of a model racing-car controller, allowing the user to alter driving style but also give more concrete instructions like “reverse the car” or “maintain a specific speed.” The strength of the approach, says Baumann, is that separating the LLM from the main controller means that even if the model hallucinates, it can’t do anything dangerous. “You get the possibility of language interaction, but you can guarantee that it is going to be within the constraints of this classical controller, so you can bake in safety,” he says. However, setting these constraints requires considerable engineering work, he adds. And if you want provable safety, you need to go a step further, says Matthias Althoff, a professor of cyberphysical systems at the Technical University of Munich. His group built a system that gets an LLM to suggest driving decisions, but then uses a mathematical process to check them against traffic rules and predictions about the behavior of other road users. This makes it possible to verify their safety before committing to them, something the Delft paper doesn’t provide. “As with any LLM, it is not guaranteed that the result is correct,” says Althoff. “For that reason, we safeguard the decisions of the LLM in our works.”

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • This article is part of our exclusive IEEE Journal Watch series in partnership with IEEE Xplore. The idea of letting a machine do the driving for you may put a lot of people off a…
In-site article

Advancing price-performance for developers with GPT‑5.6 in Kiro

GPT‑5.6 is now available in Kiro, helping developers plan, build, review, and test software with better price-performance.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • GPT‑5.6 is now available in Kiro, helping developers plan, build, review, and test software with better price-performance.
In-site article

How to use ChatGPT Work - and my top 10 tips for getting started with agentic AI

Curious about ChatGPT Work? Here's how the agentic AI handles research, files, and multistep projects, plus its risks and limits.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Curious about ChatGPT Work? Here's how the agentic AI handles research, files, and multistep projects, plus its risks and limits.
In-site article

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs

arXiv:2608.20492v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have become a prevailing paradigm for unified video perception. However, post-training on large multi-task datasets remains challenging, as existing reinforcement learning methods sample on-policy groups with few high-quality rollouts even with costly chain-of-thought (CoT) generation. In this paper, we study the sample efficiency and scalability of RL post-training for video MLLMs and introduce OraRL. We identify an overlooked role for annotations: Beyond scoring rollouts, each can enter its on-policy group as an oracle rollout, a direct positive optimization target. Direct oracle integration, however, is nontrivial: a high-reward oracle raises the group baseline and inverts otherwise positive policy advantages, a failure we term advantage inversion. At the core of OraRL is a decoupled advantage estimator: policy rollouts determine an oracle-free baseline, while the oracle-policy gap modulates both a directional gain and a separate detached oracle advantage. Sign-balanced pruning improves efficiency: by retaining only the oracle and the strongest rollouts of each sign, OraRL requires just 2.2x the step time of SFT, less than half the 4.9x required by GRPO with CoT. OraRL scales with model size and data, surpassing its backbone from 0.8B to 9B and GRPO up to 100k prompts. Without chain-of-thought, Video-ORA-9B decodes in 130 ms instead of 4,780 ms. Compared with the respective prior best models, it raises temporal mIoU from 62.5 to 66.0, tracking AO from 73.0 to 78.2, segmentation from 64.3 to 70.4, and the three-benchmark spatial-intelligence macro average from 51.0 to 56.1; on VSI-Bench, it scores 73.1 against 55.0 for GPT-5 and 55.1 for Gemini-3-Pro.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • arXiv:2608.20492v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have become a prevailing paradigm for unified video perception. However, post-training on…
In-site article

When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha

arXiv:2608.20345v1 Announce Type: new Abstract: Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), with 13.1% of U.S. adolescents (5.4 million) using generative AI for mental health advice. While these systems, from therapy apps to general chatbots, rely on large language models trained on extensive psychological literature, their safety for youth communication patterns characterized by hyperbolic language, ironic positivity, rapid semantic drift, and contextual polysemy remains unvalidated. Following multiple adolescent deaths linked to AI chatbot interactions, systematic evaluation is critical. We present two benchmarks: (1) 64 Gen Alpha mental health expressions validated by native speakers (ICC=0.72) and clinicians (kappa=0.78); (2) 75 multi-turn conversations (780 turns) with paired Standard/Gen Alpha versions. Across evaluations of LLM architectures underlying therapy apps and general chatbots - Claude, GPT-4o, Llama-3.1 - models understand 76-82% of vocabulary but correctly calibrate only 64-72% of clinical risk, creating a 10-14 percentage point (pp) vocabulary-comprehension gap (p0.48) absent in human therapists (3pp, p=.22). The gap is architecturally consistent and widens with ambiguity (7pp -> 18pp). We identify six failure patterns: sarcasm masking (29pp), minimization acceptance (43pp), informal style bias (24pp), risk-stratified ambiguity (19pp), semantic drift (19pp), context-dependent violence (7pp). Patterns compound; three or more yield 94% miss rates. Lightweight mitigations fail; only heavy scaffolding achieves human performance (6.4x cost). With 34% baseline miss rate yielding 146,880 estimated annual missed crises, we recommend mandatory human-in-the-loop architectures, quarterly youth-specific validation, transparent performance disclosure, and regulatory frameworks for youth-facing mental health AI.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • arXiv:2608.20345v1 Announce Type: new Abstract: Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), wi…
In-site article

Beyond Raw Transcripts: Structured Persona Extraction for LLM-Based Digital Twins

arXiv:2608.20344v1 Announce Type: new Abstract: LLM-based "digital twins" aim to simulate how an individual would behavein new environments or respond to novel questions, given some representation of that individual's prior responses. A common approach constructs this representation from survey transcripts or summaries responses. Prior work shows that compressing long transcripts into shorter LLM-generated summaries does not significantly reduce predictive accuracy, suggesting that information volume is not the primary bottleneck. In this work, we argue that the key limitation is instead structural:how persona information is organized before being provided to thesimulator model. We study this by comparing unstructured summaries with structured persona representations. First, we introduce a hand-craftedschema (BDE: Background, Decision procedure, Evaluation), grounded in consumer-behavior theory, and show that it improves predictive accuracy over raw transcripts by +1.91 percentage points on a homogeneous benchmark (Twin-2K-500), with similar gains on gpt-5.4-mini and Qwen3-8B as robustness checks. However, this fixed structure does not generalizeacross more heterogeneous tasks, where performance is statistically indistinguishable from the raw transcript baseline. To address this limitation, we propose an automatic structure-discovery pipeline in which an LLM iteratively proposes and refines task-specific persona structures and extraction prompts. On a benchmark of 13 diverse sub-studies, this approach restores performance, improving mean accuracy by +1.91 percentage points over the raw transcript baseline and eliminating significant losses observed with the fixed schema. Overall, our results suggest that the main constraint in LLM-based digital twins is not how much information is provided, but how it is structured -- and that the optimal structure depends on the task.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • arXiv:2608.20344v1 Announce Type: new Abstract: LLM-based "digital twins" aim to simulate how an individual would behavein new environments or respond to novel questions, given so…
In-site article

Sam Altman voices fears that control of AI could be centered in too few hands

OpenAI Group PBC Chief Executive Sam Altman is worried that artificial intelligence technology will end up being controlled by just a handful of companies or people in future, resulting in nobody else having any say about how it impacts society. Speaking in an interview with the podcaster David Senra on Sunday, Altman (pictured) warned against […] The post Sam Altman voices fears that control of AI could be centered in too few hands appeared first on SiliconANGLE.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • OpenAI Group PBC Chief Executive Sam Altman is worried that artificial intelligence technology will end up being controlled by just a handful of companies or people in future, res…
In-site article

Anthropic’s best AI model struggles to attract users as cheaper tools thrive

<p><strong><a href="https://www.ft.com/content/5ee49718-c258-4f01-aa32-7e5b76ae5245">Anthropic’s best AI model struggles to attract users as cheaper tools thrive</a></strong></p> A few interesting numbers in this FT story gathered from "people with knowledge of the matter":</p> <ul> <li>Anthropic's "annualized revenue" for July is up to $65bn - it was $47bn in May, and I collected <a href="https://simonwillison.net/2026/May/29/anthropic/">more historic numbers here</a>.</li> <li>Anthropic expect Q3 to be profitable according to the same model they used to declare Q2 profitable. "It also told investors that it had 6,000 customers that spend $100,000 annually or more."</li> <li>As for OpenAI, "annualised revenue has jumped 35 per cent in the quarter to date and is now over $40bn, with the launch of GPT 5.6 in July jolting the company’s performance after a sluggish start to the year".</li> </ul> <p>This article also introduced me to the <a href="https://ramp.com/data/ai-index">Ramp AI index</a>, which uses billing data from 70,000 Ramp credit card using companies to estimate model adoption.</p> <p>Here's Ramp's breakdown of Anthropic model spend for July 2026, which looks reasonable given that Opus 5 was only released on July 24th, and supports the idea that Fable's cost has made it a less popular model:</p> <ol> <li>Opus 4.8: 28.0%</li> <li>Sonnet 4.6: 8.3%</li> <li>Fable 5: 8.0%</li> <li>Opus 4.6: 6.9%</li> <li>Sonnet 5: 3.6%</li> <li>Opus 5: 3.5%</li> <li>Opus 4.7: 1.7%</li> <li>Sonnet 4.5: 1.3%</li> <li>Haiku 4.5: 1.0%</li> <li>Opus 4.5: 0.7%</li> </ol> <p><small></small>Via <a href="https://news.ycombinator.com/item?id=49411102">Hacker News</a></small></p> <p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/openai">openai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a>, <a href="https://simonwillison.net/tags/claude">claude</a>, <a href="https://simonwillison.net/tags/claude-mythos-fable">claude-mythos-fable</a></p>

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • <p><strong><a href="https://www.ft.com/content/5ee49718-c258-4f01-aa32-7e5b76ae5245">Anthropic’s best AI model struggles to attract users as cheaper tools thrive</a></strong></p>…
In-site article

Woe Is Em: The Sad Lifecycle of an AI Tell

In 2025, ChatGPT couldn’t stop itself from using the em dash—the punctuation mark before this phrase. Soon the em dash became the best-known stylistic “tell” of AI-generated prose. Writers who had long used the em dash…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • In 2025, ChatGPT couldn’t stop itself from using the em dash—the punctuation mark before this phrase. Soon the em dash became the best-known stylistic “tell” of AI-generated prose…
In-site article

‘We are hitting a different chapter’: OpenAI leader warns of threat of ‘persistent’ AI cyber-attacks

Chris Lehane tells Guardian of need to implement new safety standards as critics say AI firms acting ‘recklessly’ A senior leader at OpenAI has said people should prepare to defend against “ongoing, persistent” cyber-attacks from AIs, as cutting-edge artificial intelligence models gain advanced capabilities to plan and launch offensives. The leading AI company this week announced a pause in development of its most advanced internal models amid rising safety fears, and Chris Lehane, its chief global affairs officer, said: “We are hitting a different chapter, a different moment within AI, in terms of what the capabilities of this technology can do.” Continue reading...

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Chris Lehane tells Guardian of need to implement new safety standards as critics say AI firms acting ‘recklessly’ A senior leader at OpenAI has said people should prepare to defen…
In-site article

llm 0.33

<p><strong>Release:</strong> <a href="https://github.com/simonw/llm/releases/tag/0.33">llm 0.33</a></p> <p>My highlights from this release:</p> <blockquote> <ul> <li>Upgraded to the OpenAI Python library 3.x and switched the HTTP client dependency from <code>httpx</code> to <code>httpx2</code>. <a href="https://github.com/simonw/llm/issues/1608">#1608</a>, <a href="https://github.com/simonw/llm/pull/1631">#1631</a></li> </ul> </blockquote> <p>I shipped a quick <a href="https://simonwillison.net/2026/Aug/21/llm/">0.32.1 fix</a> for this yesterday, but this is the more comprehensive fix.</p> <blockquote> <ul> <li><code>llm embed</code> and <code>llm embed-multi</code> now accept <code>--key</code>. The Python <code>EmbeddingModel.embed()</code>, <code>EmbeddingModel.embed_multi()</code>, <code>Collection.embed()</code> and <code>Collection.embed_multi()</code> methods accept <code>key=</code> too, passing the resolved per-call key to embedding plugins without changing shared model state. Existing plugins that read <code>self.key</code> continue to work through a compatibility fallback. Thanks, <a href="https://github.com/ChrisJr404">ChrisJr404</a>. <a href="https://github.com/simonw/llm/issues/757">#757</a>, <a href="https://github.com/simonw/llm/pull/1620">#1620</a></li> </ul> </blockquote> <p>The embedding models now use the same pattern for keys that regular LLM models do.</p> <blockquote> <ul> <li><code>llm prompt -t/--template</code> can now be repeated to combine templates in order. This allows model configuration and options from one template to be used with a prompt from another.</li> </ul> </blockquote> <p>This unlocks a neat pattern where you can create templates that package a model with a set of default options:</p> <pre><code>llm -m gpt-5.6-luna -o reasoning_effort high --save lhigh llm "Generate an SVG of a pelican riding a bicycle" --save pelican # Combine and run the templates llm -t lhigh -t pelican </code></pre> <blockquote> <ul> <li>Reasoning-capable Responses API models now support a <code>reasoning_summary</code> option with <code>auto</code>, <code>concise</code>, and <code>detailed</code> values. This can be used with <a href="https://llm.datasette.io/en/stable/other-models.html#openai-endpoint">llm openai endpoint --responses</a>. <a href="https://github.com/simonw/llm/issues/1600">#1600</a></li> </ul> </blockquote> <p>This is particularly useful for exercising different models that provide their own imitation of the OpenAI Responses API.</p> <p>Tags: <a href="https://simonwillison.net/tags/annotated-release-notes">annotated-release-notes</a>, <a href="https://simonwillison.net/tags/llm">llm</a></p>

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • <p><strong>Release:</strong> <a href="https://github.com/simonw/llm/releases/tag/0.33">llm 0.33</a></p> <p>My highlights from this release:</p> <blockquote> <ul> <li>Upgraded to t…
In-site article

Show HN: Learn Leap, an AI tutor that teaches from your own material

I built Learn Leap because I found myself constantly asking ChatGPT questions while reading research papers. I wanted something that already understood what I was reading. With Learn Leap you can upload your own materia…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • I built Learn Leap because I found myself constantly asking ChatGPT questions while reading research papers. I wanted something that already understood what I was reading. With Le…
In-site article

Company Directory

OpenAI AI News | AI News Hub