Skip to content
AI News HubLIVE

Reports for this edition have been collected; translation and analysis are pending. Expand other updates to read source content.

Other updates (146)
Tools

Seth Meyers on Trump canceling $1bn in federal funds: ‘Is he adding a waterslide to the reflecting pool?’

Late-night hosts weighed in on Trump’s relationship with tech leaders and the White House’s new AI-powered site On Tuesday night, late-night hosts discussed a Maha summit in Washington DC, the White House’s new AI-powered website and Donald Trump once again falling asleep on live TV. Continue reading...

The Guardian AISource content · Analysis pendingSeth Meyers on Trump canceling $1bn in federal funds: ‘Is he adding a waterslide to the reflecting pool?’

Lloyal

Discussion | Link

Product Hunt AISource content · Analysis pendingLloyal

Hedge fund CEO gives $3bn to Carnegie Mellon University in historic donation

Unprecedented university donation from Citadel CEO Ken Griffin will go toward building a new Miami campus Carnegie Mellon University will receive an unprecedented $3bn donation from billionaire investor Ken Griffin, making it the largest single donation in the history of US education. Griffin, 57, is the founder and CEO of Citadel, described as the most profitable hedge fund in history. The money is going toward a new Miami campus as well as big investments at the university’s longtime home in Pittsburgh. Continue reading...

The Guardian AISource content · Analysis pendingHedge fund CEO gives $3bn to Carnegie Mellon University in historic donation

Google reportedly tests paying publishers for AI search results

Google has launched a pilot program that pays publishers for their contributions to its AI-powered search features, according to a report from The Information. The pilot program reportedly includes around 100 publishers and comes as Google faces scrutiny over the impact of its AI features on web traffic. Digiday first reported on the pilot program, which began less than a year ago. As part of the test, Google is reportedly paying participating publishers for how much their content contributed to the AI Overviews and AI Mode features in Search, as well as its Gemini chatbot. One publisher that joined the program when it first started earned … Read the full story at The Verge.

The Verge AISource content · Analysis pendingGoogle reportedly tests paying publishers for AI search results

More than 44,000 file legal objections to Palantir NHS platform handling their data

Exclusive: Campaign comes amid backlash against tech company’s involvement with Israeli military and Trump’s ICE crackdown UK politics live – latest updates Tens of thousands of people have formally demanded NHS England blocks the Palantir-powered national medical data platform from handling their personal information. More than 44,000 people have filed legal objections to their health data being shared with, stored in or used by the nationwide Federated Data Platform (FDP) operated using the US company’s AI technology. Continue reading...

The Guardian AISource content · Analysis pendingMore than 44,000 file legal objections to Palantir NHS platform handling their data

AI tool that copied actor’s ‘lustrous’ voice violated his rights, says Tokyo court

Kenjiro Tsuda wins case against TikTok account in landmark verdict to protect publicity rights A court in Japan has ruled that the human voice has legal protection after a landmark case involving a well-known actor and a TikTok account that he claimed had cloned his “lustrous” baritone voice in its AI-generated videos. Tokyo district court said on Wednesday that voices should enjoy the same protection as publicity rights, at the end of a closely watched legal action brought against TikTok by Kenjiro Tsuda – best known for voicing the character Kento Nanami in the anime Jujutsu Kaisen. Continue reading...

The Guardian AISource content · Analysis pendingAI tool that copied actor’s ‘lustrous’ voice violated his rights, says Tokyo court

Twin

Discussion | Link

Product Hunt AISource content · Analysis pendingTwin

Bill Gates predicts AI will kill a billion people – but saying he’s been wrong before is quite the understatement | Arwa Mahdawi

He knows a thing or two about tech, but what to make of the views of a man who met Jeffrey Epstein 30 times and couldn’t see what a monster he was? The Bill Gates rehabilitation tour is in full swing. After having his carefully cultivated image tarnished by the revelation of his connection to Jeffrey Epstein, Gates lay low for a while. Now, however, the Microsoft co-founder is once again thrusting himself into the public consciousness, this time with a dire warning about AI. If left unchecked, AI could cause “a billion deaths”, Gates declared in an interview with NBC on Sunday. Perhaps I should be terrified by this announcement from a man who, for all his sins, knows a thing or two about technology. But I’m sick of hearing about the hypothetical dangers of AI when the technology is being…

The Guardian AISource content · Analysis pendingBill Gates predicts AI will kill a billion people – but saying he’s been wrong before is quite the understatement | Arwa Mahdawi

As AI images become common, the threat to elections draws alarm: ‘How will voters know what’s true?’

Deepfakes and misinformation driven by artificial intelligence have become a point of deep concern for legislators, advocates and the public An image of College Park, Georgia mayor Bianca Motley Broom kissing a man – who was not her husband – in a bar spread far and wide on Facebook. The fact that it’s fake is spreading faster, now, but only after it had been seen by thousands of people. “I think this is a sad commentary about where we are in our democracy, and where we are in our political discourse,” Motley Broom said. “It also seeks to reinforce really tired stereotypes about women in leadership.” Continue reading...

The Guardian AISource content · Analysis pendingAs AI images become common, the threat to elections draws alarm: ‘How will voters know what’s true?’

jambuild

Discussion | Link

Product Hunt AISource content · Analysis pendingjambuild

Chat.sh

Discussion | Link

Product Hunt AISource content · Analysis pendingChat.sh

Reanimated AI Greta Garbo stars again … in a ball-bearing advert

Relatives of Hollywood legend say her doppelganger could now be used for more glamorous roles She shimmered in Hollywood’s silent movie era and lit up the first talkies – now Greta Garbo may be about to star in cinema’s new AI age. More than a century after her screen debut and decades after her ashes were laid to rest in a Swedish woodland cemetery, the Hollywood superstar’s descendants are ready to consider an AI version of her taking a 21st-century film role. Continue reading...

The Guardian AISource content · Analysis pendingReanimated AI Greta Garbo stars again … in a ball-bearing advert

SKF advert review – let’s hope this dull AI Greta Garbo is not the future

Swedish-American actor exhumed for bizarre, humourless 99 seconds that give us another reason to resent AI Poor Greta Garbo. Perhaps, in life, she thought her final indignity would be the Peter Cook and Dudley Moore TV sketch from 1970, showing Pete as the legendary recluse standing on the back of a flatbed truck as it rumbles through the crowded city streets, shouting at everyone through a megaphone: “I vant to be alone!” Continue reading...

The Guardian AISource content · Analysis pendingSKF advert review – let’s hope this dull AI Greta Garbo is not the future

OpenCompanion

Discussion | Link

Product Hunt AISource content · Analysis pendingOpenCompanion

Trump announces vague ‘morally binding’ AI deal among tech CEOs for ‘tremendous self-policing’

US president also issued executive order officially rebranding ‘artificial intelligence’ as ‘superintelligence’ Donald Trump announced on Tuesday that the heads of the largest US tech and AI companies had signed onto a “morally binding” agreement to put controls on artificial intelligence. He announced the deal, called the “Joint Commitment On Frontier Responsibilities,” following a White House luncheon with top tech executives. “It’s almost like a constitution in a way, and the biggest ​people in the world signed that, and ​I signed it as president, and it really ⁠is a form of protection,” Trump said. Continue reading...

The Guardian AISource content · Analysis pendingTrump announces vague ‘morally binding’ AI deal among tech CEOs for ‘tremendous self-policing’

Pheebs

Discussion | Link

Product Hunt AISource content · Analysis pendingPheebs

Elon Musk’s AI-powered Grokipedia is updating again

Elon Musk | Image: Laura Normand / The Verge Grokipedia, the AI-powered online encyclopedia from SpaceXAI, appears to be updating articles once again after a months-long pause. In August, Lawfare reported that articles on Grokipedia hadn't reviewed edits since April, but the platform's live updates site is now showing various recent changes to pages - though as I write this, many are just a note that says "recheck all references and sources." The site seems to have become active again relatively recently. The page for President Barack Obama has a note that says it was "Fact-checked by Grok" two days ago, and the page for Elon Musk was fact-checked while I was writing this, within the … Read the full story at The Verge.

The Verge AISource content · Analysis pendingElon Musk’s AI-powered Grokipedia is updating again

Upsolve Data Models

Discussion | Link

Product Hunt AISource content · Analysis pendingUpsolve Data Models

Formalini

Discussion | Link

Product Hunt AISource content · Analysis pendingFormalini

getcta.store

Discussion | Link

Product Hunt AISource content · Analysis pendinggetcta.store

DailyHelm

Discussion | Link

Product Hunt AISource content · Analysis pendingDailyHelm

ChatGPT Space

Discussion | Link

Product Hunt AISource content · Analysis pendingChatGPT Space

OpenAI halves $200 plan allowance, launches $500 plan

On Tuesday, OpenAI introduced a new $ 500-per-month Pro plan that gives subscribers the highest included usage of the company’s The post OpenAI halves $200 plan allowance, launches $500 plan appeared first on The New Stack.

The New Stack AISource content · Analysis pendingOpenAI halves $200 plan allowance, launches $500 plan

Even George Osborne might be a nimby if an AI datacente was built next to his home | Letters

The quality of life for people living next to one would be destroyed, writes Moira Brewer; plus a letter from Mark Davis So, George Osborne thinks people who don’t want an AI datacentre on their doorstep are nimbys (OpenAI’s George Osborne says datacentre nimbys holding back Britain, 22 September). He ought to come and visit our small town in Devon, where there are plans to build one of the largest datacentres in Europe, along with a battery energy storage system, on 850 acres of farmland. The size of this proposed development is quite out of proportion to its location – larger than the town itself – and accessed by narrow lanes. It would be a huge blot on a beautiful rural landscape that is enjoyed by holidaymakers who are vital to the local economy. Continue reading...

The Guardian AISource content · Analysis pendingEven George Osborne might be a nimby if an AI datacente was built next to his home | Letters

Ferndesk

Discussion | Link

Product Hunt AISource content · Analysis pendingFerndesk

Fireworks AI

Fireworks AI Join us for our inaugural conference, Forge 2026 Blog Multi Region Deployment Inside Fireworks Multi-region Deployments: One Deployment, Global Scale PUBLISHED 9/28/2026 Table of Contents Intelligent Routin…

Fireworks AI BlogSource content · Analysis pendingFireworks AI
Agents

Query claims in natural language with Amazon Bedrock Knowledge Bases

This technical how-to builds a conversational claims assistant on Amazon Bedrock Knowledge Bases that answers natural-language questions with citations. It covers ingesting claim documents from Amazon S3, querying with the AgenticRetrieveStream API, multi-turn follow-ups, metadata filters, and contextual grounding guardrails.

AWS Machine Learning BlogSource content · Analysis pendingQuery claims in natural language with Amazon Bedrock Knowledge Bases

Build a multi-agent music production pipeline on Amazon Bedrock AgentCore Runtime Instances

Amazon Bedrock AgentCore Runtime Instances gives multi-agent workflows AWS managed EC2 infrastructure with GPUs, persistent volumes, and multi-day sessions. In this post, we deploy a three-agent music production pipeline where the agents colocate on one GPU instance, share a filesystem, and hand work to each other to produce a finished track.

AWS Machine Learning BlogSource content · Analysis pendingBuild a multi-agent music production pipeline on Amazon Bedrock AgentCore Runtime Instances

All the latest news on Meta’s cute, creepy Muse AI agent

Meta launched a new Muse AI agent it claims can help you with everything from firing off emails to buying stuff online. Muse can be surprisingly effective at delivering on those promises — if you’re willing to trust Meta with your data and hand Muse your credit card. Since the launch, Meta’s also announced plans for a Tamagotchi-like Muse Charm device for its cuddly AI mascot. Follow along here for the latest news and updates. Meta’s Muse AI sent a YouTuber’s address to a stranger Meta expands Muse tools for small businesses. Meta Enterprise Platform will bring its Muse AI to businesses. Maybe don’t let Muse run your Facebook Marketplace account. The Muse Charm won’t have Instagram built in. Meta makes the Muse filesystem even more accessible Muse will apparently let you download its enti…

The Verge AISource content · Analysis pendingAll the latest news on Meta’s cute, creepy Muse AI agent

From Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AI

Building on nearly a decade of co-engineering, CoreWeave has built NVIDIA compute, networking and software into a cloud purpose-built for AI that’s still returning on investment across multiple generations of deployment. Now, CoreWeave is bringing the next generation of NVIDIA infrastructure to production. At CoreWeave Fully Connected, running this week in San Francisco, CoreWeave announced […]

NVIDIA BlogSource content · Analysis pendingFrom Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AI

Microsoft Fabric is where AI agents learn how the business works

Microsoft is turning Fabric, its integrated data platform, into the place enterprise agents go to learn how a company works, whether The post Microsoft Fabric is where AI agents learn how the business works appeared first on The New Stack.

The New Stack AISource content · Analysis pendingMicrosoft Fabric is where AI agents learn how the business works

Pay Per Use: when AI uses your work, you should get paid

Pay Per Use is now in beta. AI companies report when they use publishers’ content, and Cloudflare handles billing, payouts, and reporting, so our customers are paid according to use.

Cloudflare AI BlogSource content · Analysis pendingPay Per Use: when AI uses your work, you should get paid

Monetization Gateway beta: charge AI agents for consumption with HTTP 402

Cloudflare’s AI Gateway, Ceramic.ai, Stocktwits, and more are using the Cloudflare Monetization Gateway today to charge agents for access to tokens, APIs, and MCP tools. U.S.-based sellers can now apply for access to the closed beta.

Cloudflare AI BlogSource content · Analysis pendingMonetization Gateway beta: charge AI agents for consumption with HTTP 402

Simplifying domains for people and agents

Cloudflare Registrar’s new search delivers fast, transparent results across 420+ extensions using Workers, Durable Objects, and WebSockets. Its expanded API and cf CLI also let agents search, register, and transfer domains.

Cloudflare AI BlogSource content · Analysis pendingSimplifying domains for people and agents

Cut your AI spend with AI Gateway's Auto Router

Cloudflare AI Gateway now features a model router that evaluates request complexity using an edge-deployed classifier to select the optimal model. By balancing expected output quality against token costs, organizations can dramatically cut AI spend while maintaining performance.

Cloudflare AI BlogSource content · Analysis pendingCut your AI spend with AI Gateway's Auto Router

The Internet has a second audience

More than half the traffic reaching sites on Cloudflare is now automated, and AI agents are the fastest-growing part of it. We're giving site owners the tools to see who's visiting, decide who gets in, and charge for access.

Cloudflare AI BlogSource content · Analysis pendingThe Internet has a second audience

Cloudflare Containers, rebuilt to scale agent sandboxes

Cloudflare Containers now start 6x faster, let your agent choose each sandbox's image and instance type at runtime, and support filesystem snapshots in public beta, all controlled from a Durable Object.

Cloudflare AI BlogSource content · Analysis pendingCloudflare Containers, rebuilt to scale agent sandboxes

OpenAI Dot: How to Access & Automate Work with OpenAI’s Agent

That is a slightly awkward introduction to OpenAI Dot, but a useful one. In this article I’ll be discussing what this autonomous 24/7 agent (which consumes no usage) can do, what it is, how you can access it (along with its setup), and a live analytics exercise I ran after on Dot. By the time […] The post OpenAI Dot: How to Access & Automate Work with OpenAI’s Agent appeared first on Analytics Vidhya.

Analytics VidhyaSource content · Analysis pendingOpenAI Dot: How to Access & Automate Work with OpenAI’s Agent

“Think of it as Kubernetes for agents”: OpenClaw lands in the enterprise with Nvidia and Red Hat on board

OpenClaw has come a long way since it emerged as a viral weekend project less than a year ago. Created The post “Think of it as Kubernetes for agents”: OpenClaw lands in the enterprise with Nvidia and Red Hat on board appeared first on The New Stack.

The New Stack AISource content · Analysis pending“Think of it as Kubernetes for agents”: OpenClaw lands in the enterprise with Nvidia and Red Hat on board

Helping small businesses put AI to work

OpenAI is partnering with America’s SBDC to expand hands-on AI training and local support for small businesses, alongside a new report on how small teams are using AI.

OpenAI NewsSource content · Analysis pendingHelping small businesses put AI to work

Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms

Perplexity has released Photon, an in-house retrieval and ranking engine written in Rust. It replaces an open-source engine Perplexity had forked for its AI-native search stack. Photon now handles retrieval and ranking for all production traffic. It also powers a new Fast Search mode in the Perplexity Search API. Perplexity reports single-call latency of 160 […] The post Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms appeared first on MarkTechPost.

MarkTechPostSource content · Analysis pendingPerplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms

NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks

Video world models can render convincing clips that still break physics. Butter spreads like paint. Balls pass through walls. A team from NVIDIA, MIT and the University of Oxford argues the fix can come from language itself, not from extra visual, latent or numerical signals. Their framework, Physis-Lang, treats physical language as a shared, optimizable […] The post NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks appeared first on MarkTechPost.

MarkTechPostSource content · Analysis pendingNVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks

JarvisCore

Discussion | Link

Product Hunt AISource content · Analysis pendingJarvisCore

SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation

arXiv:2609.36171v1 Announce Type: new Abstract: Large-scale demonstrations have driven unprecedented progress in robot learning, yet collecting robot data through teleoperation is expensive and difficult to scale to diverse environments and long-horizon tasks. Simulation offers a scalable alternative, but existing data-generation pipelines often rely on open-loop controllers, scripted skill sequences, or task-specific programs. We introduce SkillWeaver, an agentic framework that autonomously generates robot experience by exploring over Neural Interaction Skills (NIS): reusable, parameterized, closed-loop policies that expose learned physical interaction capabilities to a reasoning agent. Given a task and a simulated environment, a VLM agent reasons about what to do next, invokes and param…

arXiv RoboticsSource content · Analysis pendingSkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation

Scouting the Dynamics Gap: Test-Time Policy Adaptation via Action-Outcome Feedback

arXiv:2609.36107v1 Announce Type: new Abstract: While pretrained robotic policies exhibit impressive capabilities in controlled environments, unobserved physical properties and dynamics require these policies to rapidly adapt during deployment. Existing test-time adaptation methods typically rely on sparse scalar rewards, failing to exploit the rich geometric and dynamic feedback from the environment during physical interaction. To address this challenge, we propose SCOUT, a dynamics-aware meta-learning framework that enables manipulation policies to rapidly adapt by continuously revising their internal beliefs about environment dynamics. Our approach couples an action-prediction policy with a forward dynamics model via a shared belief latent space. During meta-training, an inner loop upd…

arXiv RoboticsSource content · Analysis pendingScouting the Dynamics Gap: Test-Time Policy Adaptation via Action-Outcome Feedback

In-Context Learning for Robots: Methods and Applications

arXiv:2609.36012v1 Announce Type: new Abstract: General-purpose robots must infer what a new task requires and translate that understanding into appropriate physical action. In-context learning (ICL) for robots supports this process by using demonstrations and interaction to direct existing competence with neural parameters held fixed during deployment. We organize this literature review around the interfaces connecting contextual evidence to execution, distinguishing four families: context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution. Comparing these interfaces clarifies their transfer assumptions and the roles of training, correspondence, and memory in making context useful. Across manipulation and navigation, we…

arXiv RoboticsSource content · Analysis pendingIn-Context Learning for Robots: Methods and Applications

CoDimRecon: Agentic Reconstruction of Sim-Ready 3D Scenes with Deformable Curves, Surfaces, and Volumes

arXiv:2609.36024v1 Announce Type: new Abstract: Reconstructing simulation-ready 3D scenes from real-world observations enables robotics, gaming, and immersive applications, yet existing methods largely assume rigid objects. This leaves an important gap for deformables, whose simulation-ready geometry depends on dimensionality (curves, surfaces, or volumes) and whose behavior may require models beyond elasticity. We present CoDimRecon, an agentic framework that reconstructs editable scenes containing rigid, articulated, and deformable objects from multi-view RGB observations. Scene-level geometric priors ground scale and layout, while object-level generated meshes guide the agent toward detailed, compact geometry; articulated rigid objects are decomposed into movable parts with explicit jo…

arXiv Computer VisionSource content · Analysis pendingCoDimRecon: Agentic Reconstruction of Sim-Ready 3D Scenes with Deformable Curves, Surfaces, and Volumes

Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method

arXiv:2609.35965v1 Announce Type: new Abstract: Vision-and-Language Navigation (VLN) has largely focused on a single agent following a single instruction, yet many real-world applications require teams of robots to tackle tasks beyond the capabilities of any individual agent. We present Systematic Multi-Agent Vision-and-Language Navigation, providing, to our knowledge, the first systematic formalization of multi-agent VLN as a constrained coordination problem: each mission consists of subtasks carrying dependency and resource constraints (presence locks and holding chains). A verified four-stage crafting pipeline instantiates the task as MAVLN, comprising 11,724 episodes across 145 scenes with teams of up to four agents under three instruction regimes, accompanied by tailored constraint-a…

arXiv Computer VisionSource content · Analysis pendingSystematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method

HEIR: Learning Human-Entity Interactions with Functional Roles

arXiv:2609.35955v1 Announce Type: new Abstract: Understanding human-entity interactions requires recovering each person-action event's participants, roles, and shared identities. This structure can support embodied agents by clarifying who acts on which entities and how, informing anticipation and coordination in shared environments. Standard HOI metrics score individual links, leaving complete event composition undermeasured. We introduce HEIR (Human-Entity Interactions with Functional Roles), an image benchmark for complete grounded participant-role sets across object, interpersonal, and self-directed interactions. It contains 18,730 images, six roles, 105 actions, and 437 nouns, with shared entities, role changes, and repeated fillers; 51.6% of images contain multiple actors and 62.1%…

arXiv Computer VisionSource content · Analysis pendingHEIR: Learning Human-Entity Interactions with Functional Roles

When Successful Memories Mislead Embodied Agents:Memory Adaption For Task-Conditioned Execution

arXiv:2609.35808v1 Announce Type: new Abstract: Experience reuse can reduce repeated exploration in embodied agents, but a trajectory that succeeded previously may be unsuitable for the current execution context. Existing memory systems pri marily optimize construction and retrieval; semantic relevance and historical success therefore remain insufficient when retrieved ex perience contains incompatible actions or an inappropriate level of structure. We introduce Memory Adaptation for Task-Conditioned Execution (MATE), a deterministic post-retrieval procedure that converts trajectories into execution-oriented memory. MATE re moves obsolete control context, extracts condition-action-effect transitions, applies verified action normalization, selects a task dependent representation, and seria…

arXiv Computational LinguisticsSource content · Analysis pendingWhen Successful Memories Mislead Embodied Agents:Memory Adaption For Task-Conditioned Execution

Environment Steering: Using Data Flow Control to Improve Agent Utility and Safety

arXiv:2609.35807v1 Announce Type: new Abstract: LLM agents can make unsafe tool calls even when instructed to behave safely. Existing defenses constrain agents before execution, modify tool inputs/outputs, or rely on LLM judges; these approaches may depend on model behavior or block unsafe actions without helping the agent recover. We argue that the execution environment should instead enforce safety as the agent runs and steer it toward safe alternatives when violations occur---we call this Environment Steering. We implement this by modeling the agent and harness execution state as database tables, track the record-level data flows, and check these data flows against declarative policies during runtime. When violations are detected, policy- and context-specific feedback steers the agent…

arXiv Computational LinguisticsSource content · Analysis pendingEnvironment Steering: Using Data Flow Control to Improve Agent Utility and Safety

From Lexical Baselines to Agentic Retrieval-Augmented Generation: Structured Skill and Responsibility-Level Extraction with the SFIA Framework

arXiv:2609.35806v1 Announce Type: new Abstract: Automated skill extraction underpins workforce planning, yet most systems represent skills as flat labels with no notion of the responsibility level at which a skill is practiced. The Skills Framework for the Information Age (SFIA) captures exactly this dimension, defining 147 professional skills across seven responsibility levels, but no automated LLM-based extraction targeting SFIA has been reported. We formalize the task as structured prediction of (skill, level) pairs from free text and ask three questions: how accurately can text be mapped onto SFIA's closed vocabulary, which strategies reliably predict the level alongside the skill, and do agentic designs improve on simpler retrieval and prompting? We evaluate five strategies (a lexica…

arXiv Computational LinguisticsSource content · Analysis pendingFrom Lexical Baselines to Agentic Retrieval-Augmented Generation: Structured Skill and Responsibility-Level Extraction with the SFIA Framework

Sage: Formalization with Semantic Correction

arXiv:2609.35790v1 Announce Type: new Abstract: While neural theorem provers have achieved impressive milestones in formal mathematics, they largely operate on the assumption that faithful Lean 4 formal statements are already provided. Translating informal natural language into a formal language is a critical data bottleneck plagued by an "illusion of rigor": standard type-checkers accept statements that compile but drop hypotheses, introduce vacuous truths, or subtly alter mathematical bounds. To resolve this, we introduce Sage (Semantic Agent-Guided Formalization Engine), an agentic framework that replaces monolithic translation with a four-stage decomposed generation pipeline coupled with a dual-signal semantic correction loop. By pairing Lean 4 compiler diagnostics with multi-dimensio…

arXiv Machine LearningSource content · Analysis pendingSage: Formalization with Semantic Correction

Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models

arXiv:2609.35875v1 Announce Type: new Abstract: Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity among agents is the driver, in the setting where the question is still measurable: small open-weight models with benchmark headroom. Across 23 models from eleven vendor families, five tasks, and 5,500+ debate and control runs, we vary diversity along three axes - personas, sampling temperature, and model identity - pairing every debate configuration with a generation-budget-matched majority-vote control. The hypothesis is rejected on every axis. Debate beats single-agent inference (3--7 points where tasks have hea…

arXiv AISource content · Analysis pendingBeyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models

OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing

arXiv:2609.35799v1 Announce Type: new Abstract: In July 2026, OpenAI's agents coordinated over channels outside their intended environment to breach Hugging Face's secured infrastructure. Could existing alignment testing practices have foreseen this incident? If not, what needs to change? We explore these questions. First, we identify the misaligned behaviors that caused this incident. Then, we show how to elicit these behaviors from publicly available models manually and that auditing agents can do the same if given a large compute budget. Based on our results, we propose directions to improve alignment testing. Concretely, in this project: (1) We reproduce the misaligned AI behaviors that led to the OpenAI-Hugging Face incident in an environment that simulates the original pipelines and…

arXiv AISource content · Analysis pendingOpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing

Campfire

Discussion | Link

Product Hunt AISource content · Analysis pendingCampfire

One Bad Prompt Took Down a Company’s Salesforce: RSA’s Jim Taylor on Agent ID and Taming the 4,000 Shadow AI Agents Hiding in Your Enterprise

RSA launched Agent ID at The AI Conference in San Francisco. It's an agentic identity security platform for finance, government, healthcare, and critical infrastructure. It has 3 modules. Discover finds sanctioned and shadow agents and MCP servers. Secure is an inline gateway that checks every tool call against policy. Govern maps evidence to 10 regulatory frameworks. Discover and Secure ship November 16, 2026. The post One Bad Prompt Took Down a Company’s Salesforce: RSA’s Jim Taylor on Agent ID and Taming the 4,000 Shadow AI Agents Hiding in Your Enterprise appeared first on MarkTechPost.

MarkTechPostSource content · Analysis pendingOne Bad Prompt Took Down a Company’s Salesforce: RSA’s Jim Taylor on Agent ID and Taming the 4,000 Shadow AI Agents Hiding in Your Enterprise

SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation

Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and start, crons, and memory consolidation. Yet existing benchmarks and training frameworks schedule only the benchmark’s own events, leaving each benchmark and agent pair to build a custom scheduling loop. We present SCLATE, an execution substrate where benchmarks and unmodified agents each add their events to one open event scheduler through an adapter. A hybrid simulated clock…

Apple Machine Learning ResearchSource content · Analysis pendingSCLATE: A Substrate for Continual-Learning Agent Training and Evaluation

ShareCube

Discussion | Link

Product Hunt AISource content · Analysis pendingShareCube

Polylane

Discussion | Link

Product Hunt AISource content · Analysis pendingPolylane

OpenAI announces ‘dots’ agent after scrapping launch of new AI model over safety concerns

Company’s product comes weeks after Meta introduced its own artificially intelligent agent Muse Less than 24 hours after OpenAI announced it would scrap the launch of an artificial intelligence model over safety concerns, the company debuted a whole new suite of AI tools. During its annual showcase for developers in San Francisco on Tuesday, Sam Altman, OpenAI’s CEO, enthusiastically unveiled an AI agent called “dots” that he said was “more ambitious” than ChatGPT and a “whole new way to work with AI”. Continue reading...

The Guardian AISource content · Analysis pendingOpenAI announces ‘dots’ agent after scrapping launch of new AI model over safety concerns

Cubicle

Discussion | Link

Product Hunt AISource content · Analysis pendingCubicle

Why Featherless says you don’t need a tank to deliver a pizza

The debate over what size, shape, and scale of model best suits each task isn’t going away any time soon. The post Why Featherless says you don’t need a tank to deliver a pizza appeared first on The New Stack.

The New Stack AISource content · Analysis pendingWhy Featherless says you don’t need a tank to deliver a pizza

Dots by OpenAI

Discussion | Link

Product Hunt AISource content · Analysis pendingDots by OpenAI

OpenAI Launches dots: Always-On GPT-6 Astra Agents That Work From Their Own Cloud Computers

OpenAI just introduced dots at their DevDay today. Dots are persistent AI agents powered by GPT-6 Astra. Each dot gets its own cloud computer and browser. It works across 4,000+ apps through ChatGPT plugins and keeps going after you log off. Is it deployable today? Yes, as a managed product. Dots are rolling out in […] The post OpenAI Launches dots: Always-On GPT-6 Astra Agents That Work From Their Own Cloud Computers appeared first on MarkTechPost.

MarkTechPostSource content · Analysis pendingOpenAI Launches dots: Always-On GPT-6 Astra Agents That Work From Their Own Cloud Computers

Superpowers for Humans

I’ve known Jesse Vincent for more than 20 years, since the days when I was still editing and publishing Perl books and organizing the Perl Conference and he was the chief maintainer of Perl 5 and the project manager for Perl 6. We’d lost touch, but he rocketed back into my consciousness last October when […]

O'Reilly AI & ML RadarSource content · Analysis pendingSuperpowers for Humans

OpenAI makes ‘Sign in with ChatGPT’ a way to use your subscription in third-party tools

OpenAI has announced that it’s making it easier for ChatGPT subscribers to use their allowance inside third-party AI and coding The post OpenAI makes ‘Sign in with ChatGPT’ a way to use your subscription in third-party tools appeared first on The New Stack.

The New Stack AISource content · Analysis pendingOpenAI makes ‘Sign in with ChatGPT’ a way to use your subscription in third-party tools

OpenAI answers TypeSafe’s Jev with a Decision API built on Luna

With the sudden rise of TypeSafe’s Jev, decision models have become incredibly popular, so it’s not surprising that OpenAI is The post OpenAI answers TypeSafe’s Jev with a Decision API built on Luna appeared first on The New Stack.

The New Stack AISource content · Analysis pendingOpenAI answers TypeSafe’s Jev with a Decision API built on Luna

OpenAI launches Dots, its Muse competitor

Like Muse, OpenAI’s Dots are represented by cute, customizable avatars. | Image: OpenAI OpenAI is responding to Meta's buzzy Muse AI with agentic helpers of its own: Dots. During its DevDay keynote on Tuesday, OpenAI announced that Dots will serve as always-on AI assistants that can "do nearly anything" across connected apps in the background while learning your preferences over time. The company's capable GPT-6 Astra model powers Dots, which use their own cloud computer to access a web browser and more than 4,000 supported apps. You can interact with a Dot through a text-message-like interface that's similar to the one offered by Muse, as well as hop on a voice call with it from ChatGPT on the web, desktop, or mobile. Dots ca … Read the full story at The Verge.

The Verge AISource content · Analysis pendingOpenAI launches Dots, its Muse competitor

Prompt engineering fundamentals for Amazon Quick

Prompt engineering in Amazon Quick shapes how accurately its AI-powered features respond to your requests. Part 1 of a two-part series covers the foundational principles and reusable frameworks (specificity, context-setting, few-shot examples, and the CRISPE framework) for consistent, high-quality results across Amazon Quick.

AWS Machine Learning BlogSource content · Analysis pendingPrompt engineering fundamentals for Amazon Quick

Prompt engineering by Quick component: Patterns and pitfalls

Part 2 of our Amazon Quick prompt engineering series goes component by component. Learn the prompt patterns that get the best results from Amazon Quick Research, Quick Flows, Quick Sight, chat agents, and action integrations, plus the common pitfalls to avoid.

AWS Machine Learning BlogSource content · Analysis pendingPrompt engineering by Quick component: Patterns and pitfalls

Building an AI-powered contract intelligence platform with Amazon Quick and Amazon Bedrock AgentCore

Manually extracting data from hundreds of vendor contracts doesn't scale, and RAG chat tools fall short on portfolio-wide questions. This post shares a contract intelligence platform on AWS that uses AI agents to extract and verify contract fields, then answers aggregate and single-contract questions through Amazon Quick analytics.

AWS Machine Learning BlogSource content · Analysis pendingBuilding an AI-powered contract intelligence platform with Amazon Quick and Amazon Bedrock AgentCore

OpenAI DevDay 2026: The biggest news and announcements

It’s OpenAI’s turn in the fall tech events calendar. The company is hosting its annual DevDay on September 29th in San Francisco, starting with a live keynote featuring CEO Sam Altman. The company is teasing “20+ launches” with Altman saying that “we have found a new thing,”, and rumors suggest it could launch a consumer AI agent to rival Meta’s Muse. The keynote begins at 1PM ET / 10AM ET, and you can watch it on OpenAI’s website here. DevDay is happening at a tumultuous moment for OpenAI. The AI industry has been rocked by revelations of agents hacking outside companies, including OpenAI’s own models, which breached Hugging Face earlier this year. The hacks kicked off a broader conversation about a potential AI development slowdown. So in addition to new products, Altman and others at O…

The Verge AISource content · Analysis pendingOpenAI DevDay 2026: The biggest news and announcements

ElevenLabs reaches $22bn valuation after employee tender

strong]:tw-font-normal tw-text-gray-600">Introducing Eleven v4Meet Eleven v4, our most emotive model yet. With 3x credits included on Creator+ until October 12 Discover Skip to content Log inSign up Contact salesLog in…

ElevenLabs BlogSource content · Analysis pendingElevenLabs reaches $22bn valuation after employee tender

Announcing our partnership with OpenAI

News Announcing our partnership with OpenAI Baseten is partnering with OpenAI to serve open models natively via Codex and the Responses API. Authors Dannie Herzberg Last updated September 29, 2026 Share TL;DR We’re exci…

Baseten BlogSource content · Analysis pendingAnnouncing our partnership with OpenAI
Startups

MIT Transit Lab to develop an AI platform for public transit agencies

With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.

MIT News AISource content · Analysis pendingMIT Transit Lab to develop an AI platform for public transit agencies

Evaluating AI-Generated Frontend Code: What Should We Actually Test?

AI can now generate a surprising amount of frontend code from a short description. A developer can ask for a form, a table, a modal, a settings page, or a dashboard view and get something that looks usable almost immediately. It may compile, render, and even arrive with a few tests. That is useful, but […]

O'Reilly AI & ML RadarSource content · Analysis pendingEvaluating AI-Generated Frontend Code: What Should We Actually Test?
Research

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

MIT News AISource content · Analysis pendingThis game-playing AI is the new champ at Stratego

Instagram is adding an AI ‘assistant’ to tell you how to post

Instagram is the latest social media platform to add built-in AI-powered features that will give users feedback on their posts. The company announced Wednesday that its standalone Edits app will now include an AI "creative assistant" that pulls in data from a user's Instagram account and offers suggestions on what to change. Using the new assistant, creators can ask open-ended questions about what's trending on Instagram, why one video worked and another didn't, or what piece of content resonated most with their audience. The tool pulls in data and metrics such as likes, views, retention on videos, and shares. Instagram says it can spot pat … Read the full story at The Verge.

The Verge AISource content · Analysis pendingInstagram is adding an AI ‘assistant’ to tell you how to post

Do you want help from humans or AI bots? Because the UK civil service is changing – and not for the better | The civil servant

Rapidly expanding AI use within the government is numbing departments to its profound threats. And try arguing with a faceless digital bureaucrat Another day, another apocalyptic warning about the likelihood that rampantly uncontrollable artificial intelligence will solve all of our problems, including the problem of being alive. Hot on the heels of disgruntled Anthropic researchers announcing the end of civilisation by 2030, we’ve now had the portentously named Hugging Face incident, as well as the Australian government’s grim discovery of a rogue AI’s attack on its Medicare website. Swarms of experts are now falling over each other to warn us that much, much worse is on the way. The author works for the UK civil service Continue reading...

The Guardian AISource content · Analysis pendingDo you want help from humans or AI bots? Because the UK civil service is changing – and not for the better | The civil servant

Disrupting a coordinated model-distillation campaign

Learn how OpenAI disrupted a campaign to extract protected model reasoning and is strengthening defenses against adversarial distillation.

OpenAI NewsSource content · Analysis pendingDisrupting a coordinated model-distillation campaign

Bilinear World Models: Learning Representations with Structured Dynamics for Efficient Control

arXiv:2609.36305v1 Announce Type: new Abstract: World models jointly learn latent representations and dynamics that predict how high-dimensional observations evolve under actions. In this work, we propose a JEPA-style world model in which, rather than learning arbitrary latent dynamics, we restrict them to follow a bilinear parameterization. This structure enables efficient planning and control while shifting the modeling burden onto the encoder, encouraging richer representations that expose the controllable geometry of the system. In particular, this structured parameterization allows us to structurally enforce action recoverability, thereby preventing representation collapse by construction. Although prescribing a bilinear parametrization may appear restrictive, we show that a broad cl…

arXiv RoboticsSource content · Analysis pendingBilinear World Models: Learning Representations with Structured Dynamics for Efficient Control

Design and Validation of an Antagonistic Tendon-Driven Dexterous Robotic Hand with Bidirectional Operation

arXiv:2609.36241v1 Announce Type: new Abstract: Dexterous robotic hands typically reproduce human hand morphology but inherit its one-sided grasping workspace, requiring wrist or arm reorientation to grasp from the opposite side. Existing reversible hands generally rely on non-anthropomorphic, soft, or task-specific finger arrangements, whereas conventional five-digit anthropomorphic hands remain designed primarily for palmar-side grasping. This paper presents an anthropomorphic, human-scale (200 mm length), lightweight (220 g), 3D-printed, 17-DoF robotic hand built on a bidirectional antagonistic tendon-routing mechanism, in which flexion/extension (except the coupled joint) and abduction/adduction at joints are actively driven without passive return springs. The proposed routing mechani…

arXiv RoboticsSource content · Analysis pendingDesign and Validation of an Antagonistic Tendon-Driven Dexterous Robotic Hand with Bidirectional Operation

SAKI: Skill Assembly and Kinematic Imitation from Human Videos for Long-Horizon Mobile Manipulation

arXiv:2609.36031v1 Announce Type: new Abstract: Learning from human videos offers a promising route to acquiring diverse manipulation skills. Extending this capability beyond tabletop settings to long-horizon mobile manipulation requires adapting and composing demonstrated interactions across changing scenes and robot configurations. We present Skill Assembly and Kinematic Imitation (SAKI), a framework connecting human-video skill acquisition, cross-demonstration assembly and closed-loop whole-body execution. SAKI prepares reusable object-centric skills that preserve task-critical interactions while allowing transfer paths to adapt. Given a goal and supplied task dependencies, it selects and orders skills, binds their object roles to the current scene, and carries scene estimates and robo…

arXiv RoboticsSource content · Analysis pendingSAKI: Skill Assembly and Kinematic Imitation from Human Videos for Long-Horizon Mobile Manipulation

Hardware-Aware Functional Kolmogorov-Arnold Networks for Efficient Medical Image Enhancement and Segmentation

arXiv:2609.36134v1 Announce Type: new Abstract: Functional Kolmogorov-Arnold Networks (FunKAN) achieve state-of-the-art accuracy on MRI Gibbs artifact removal and anatomical segmentation, but their 11.6 M parameters and 8.7 GFLOPs are too large for edge medical devices. We present FunKANLite, a two-stage, hardware-aware compression of FunKAN for point-of-care use. FunKANLite-TR reduces the spatial prior and replaces the ResBlock offset predictor with a depthwise-separable block. It has 1.9x fewer parameters than FunKAN and no loss in accuracy. We then distill FunKANLite-TR into FunKANLite-ST, which lowers the Hermite basis rank, factorizes the spatial prior into a low-rank form, and halves the filter widths. FunKANLite-ST has 5.6x fewer parameters and 3.7x fewer GFLOPs than FunKAN. It sta…

arXiv Computer VisionSource content · Analysis pendingHardware-Aware Functional Kolmogorov-Arnold Networks for Efficient Medical Image Enhancement and Segmentation

AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search

arXiv:2609.36066v1 Announce Type: new Abstract: Open-world aerial object-goal search is a foundational yet challenging task, requiring aerial agents to autonomously explore large-scale, unstructured three-dimensional environments and reach target objects specified by semantic descriptions or reference images, rather than following route-specific instructions. However, research in this task remains at a nascent stage and relies on small, environment-specific benchmarks with heterogeneous action spaces and data formats. These limitations hinder large-scale training and cross-benchmark evaluation, constraining the scalability and generalizability of aerial agents. To address this problem, we propose AerialDojo-200K, a large-scale benchmark suite for open-world aerial object-goal search, with…

arXiv Computer VisionSource content · Analysis pendingAerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search

CoVLM-Bench: A Real-World Benchmark for Cooperative Driving Question Answering and Planning

arXiv:2609.35823v1 Announce Type: new Abstract: Vision-language models (VLMs) have made substantial progress in autonomous driving, but their success has primarily been studied in ego-centric scenes. Infrastructure-side observations provide views beyond the ego vehicle's field of view, yet conventional cooperative-driving systems typically transform them into geometric representations for downstream perception and planning. Directly incorporating these views into VLMs offers an opportunity to improve cooperative scene understanding and trajectory planning. However, question answering and trajectory planning have not been jointly evaluated on the same real-world vehicle-infrastructure scenes. We present CoVLM-Bench, a benchmark for cooperative driving question answering (CDQA) and cooperat…

arXiv Computer VisionSource content · Analysis pendingCoVLM-Bench: A Real-World Benchmark for Cooperative Driving Question Answering and Planning

Developing an OCR model for Extracting Information from Invoices with Korean Language

arXiv:2609.35796v1 Announce Type: new Abstract: Invoices are commercial documents that contain various pieces of information, including the purchased items, time, and total money. Making the extraction of important information crucial. The stored information serves different purposes. Korean language is the native language of about 80 million people, playing an important role in not only South and North Korea but also in many other countries such as Vietnam, Philippine where a large number of Korean companies are located. In this context, to automatically extract proper information from the invoices with Korean language, we propose an efficient Optical Character Recognition (OCR) model in which a deep learning model is combined with some image preprocessing techniques. The proposed OCR mo…

arXiv Computational LinguisticsSource content · Analysis pendingDeveloping an OCR model for Extracting Information from Invoices with Korean Language

Mara Chain: Rethinking Failure as a Stepping Stone for AI System Auto-Evolution

arXiv:2609.35855v1 Announce Type: new Abstract: Optimizing deployed AI systems increasingly amounts to editing prompts, skills, harnesses, and code rather than model weights. Existing approaches commonly optimize these artifacts through propose-evaluate-select procedures, where candidate configurations are evaluated and only those meeting an acceptance criterion are selected. Yet our analysis shows that discarded candidates often contain information critical for subsequent optimization. Discarding them causes later proposals to revisit the same failure modes. We introduce Mara Chain, a refinement procedure that turns rejected candidates into stepping stones. Rather than discarding a rejected candidate, Mara Chain retains and iteratively refines it using evidence accumulated across precedi…

arXiv Machine LearningSource content · Analysis pendingMara Chain: Rethinking Failure as a Stepping Stone for AI System Auto-Evolution

Position: Let's Strengthen Verifiability If We Can't Enforce Reproducibility

arXiv:2609.35854v1 Announce Type: new Abstract: In the field of Machine Learning, many papers contain empirical results supporting claimed statements or illustrating the performance of a proposed method. However, most practitioners know that (1) results are generally hard to reproduce, and increasingly so, (2) code is not often available to do so, and (3) it hinders the development of research. In this position paper, we analyze and quantify these issues, and make concrete proposals to improve result checkability, if not reproducibility. Code and supporting materials are available at https://github.com/giddyyupp/position-enforce-verifiability.

arXiv Machine LearningSource content · Analysis pendingPosition: Let's Strengthen Verifiability If We Can't Enforce Reproducibility

Learned Compression of SAR Phase-History Data: A Rate-Honest Feasibility Study on GOTCHA

arXiv:2609.35848v1 Announce Type: new Abstract: On-board compression of synthetic aperture radar (SAR) phase history is bandwidth-critical, and block-adaptive quantization (BAQ) remains the operational standard. We test whether a small convolutional autoencoder, with its encoder on the sensor, can compete with BAQ on complex phase-history patches from the AFRL GOTCHA collection. Every method is charged for all transmitted bits, rates are reported in bits per complex sample (b/cs), and detection is scored by one-to-one matching of CA-CFAR detections. The autoencoder (28,656 encoder parameters) loses at every rate. At 16 b/cs it reaches -2.87 dB NMSE, against -35.5 dB for 8-bit BAQ with $\pm 3\sigma$ clipping and -41.0 dB with a tuned clipping range. It also loses to a $16 \times 16$ block…

arXiv Machine LearningSource content · Analysis pendingLearned Compression of SAR Phase-History Data: A Rate-Honest Feasibility Study on GOTCHA

Serverless gossip training of LSTM failure detectors: A matched-protocol comparison with federated, local and centralized learning on NASA C-MAPSS

arXiv:2609.35792v1 Announce Type: new Abstract: Industrial predictive maintenance increasingly depends on learning from equipment spread across sites whose sensor data cannot easily be pooled. Federated averaging (FedAvg) solves this with a central aggregation server; gossip learning removes the server, but its behaviour for recurrent failure-detection models has not been measured under controlled conditions. We compare synchronous ring gossip with FedAvg, isolated local training and a centralized reference for a stacked LSTM that detects imminent failure on the NASA C-MAPSS turbofan benchmark. All methods share one open implementation, architecture, initialization, optimizer, data split and training budget, and the primary endpoint uses one terminal window per test engine to avoid the st…

arXiv Machine LearningSource content · Analysis pendingServerless gossip training of LSTM failure detectors: A matched-protocol comparison with federated, local and centralized learning on NASA C-MAPSS

Risk-Averse Online POMDP Planning via CVaR of the Immediate Cost with Performance Guarantees

arXiv:2609.35874v1 Announce Type: new Abstract: Online POMDP planners optimize the expected cumulative cost, which can mask dangerous states when the belief places significant mass on high-cost states. Existing risk-averse methods apply static or dynamic Conditional Value at Risk (CVaR) to the value function, capturing trajectory-level risk, but share two gaps: (i) by retaining the immediate cost as an expectation of a state-dependent cost over the belief, the risk \emph{within} the belief is left unaddressed; and (ii) by modifying the value function, they require new tailored algorithms rather than reusing existing expectation-based planners. We instead apply CVaR to the immediate cost over the belief at each step, directly targeting per-step uncertainty about the current state. The stan…

arXiv AISource content · Analysis pendingRisk-Averse Online POMDP Planning via CVaR of the Immediate Cost with Performance Guarantees

Liquid AI Releases d1: A Decision Model That Returns Calibrated Probabilities With Zero Output Tokens

Liquid AI has released d1, a decision model built for structured choices instead of text generation. You give it context and a set of typed questions. It returns calibrated probabilities across a fixed set of outcomes in a single call, with zero generated tokens. The target is the work many teams still send to general […] The post Liquid AI Releases d1: A Decision Model That Returns Calibrated Probabilities With Zero Output Tokens appeared first on MarkTechPost.

MarkTechPostSource content · Analysis pendingLiquid AI Releases d1: A Decision Model That Returns Calibrated Probabilities With Zero Output Tokens

Smart ring maker Oura puts off initial public offering due to market ‘uncertainty’

Researchers say the IPO market was off to a solid start but tailed off in the third quarter Smart ring maker Oura Inc says it is postponing a planned initial public offering due to market “uncertainty”. In a Tuesday press release, Oura said it is delaying the stock market float “despite strong demand”. Continue reading...

The Guardian AISource content · Analysis pendingSmart ring maker Oura puts off initial public offering due to market ‘uncertainty’

Nebius Opens 2026 Physical AI Awards: Five $150K Compute Credit Prizes

Nebius and NVIDIA are running the 2026 Physical AI Awards for startups with products in the field. Five category winners each get $150,000 in compute credits, joint promotion, executive mentorship, and seats at an executive dinner. Applications close October 25. The post Nebius Opens 2026 Physical AI Awards: Five $150K Compute Credit Prizes appeared first on MarkTechPost.

MarkTechPostSource content · Analysis pendingNebius Opens 2026 Physical AI Awards: Five $150K Compute Credit Prizes

AI researchers put out videos saying superintelligence is ‘exactly as dangerous as it sounds’

"The chance of human extinction is about a coin flip, in my view," Geoffrey Irving, a former OpenAI and Google DeepMind employee, said in a new interview. It's one of a dozen interviews with AI researchers, including current and former employees at OpenAI, Google, and Anthropic. Palisade Research, which says it's a non-profit studying AIs' capabilities and motivations, collected and launched the interviews on frominside.ai. Several of the researchers joined Irving in warning about the possibility AI could drive humans extinct. Neel Nanda, a Google DeepMind research scientist, said there's "at least a 10 percent chance that it causes human … Read the full story at The Verge.

The Verge AISource content · Analysis pendingAI researchers put out videos saying superintelligence is ‘exactly as dangerous as it sounds’

Nebius Opens 2026 Physical AI Awards: Five $150K Compute Prizes, Nine Judges, and an October 25 Deadline

Nebius and NVIDIA are running the 2026 Physical AI Awards for startups with products in the field. Five category winners each get $150,000 in compute credits, joint promotion, executive mentorship, and seats at an executive dinner. Applications close October 25. The post Nebius Opens 2026 Physical AI Awards: Five $150K Compute Prizes, Nine Judges, and an October 25 Deadline appeared first on MarkTechPost.

MarkTechPostSource content · Analysis pendingNebius Opens 2026 Physical AI Awards: Five $150K Compute Prizes, Nine Judges, and an October 25 Deadline

RCP-nDCG@10: A New Approach to Enterprise Retrieval Quality

How well does nDCG capture the quality of modern retrieval systems? nDCG is the established yardstick for measuring how well search systems rank results, and is widely used across benchmarks such as MTEB and BEIR. It is…

Cohere BlogSource content · Analysis pendingRCP-nDCG@10: A New Approach to Enterprise Retrieval Quality
Models

Identify AI model overuse with User Insights

AI Gateway User Insights now adds task, model, turn, and user categories to help teams understand AI adoption and make better model decisions. This is available free to AI Gateway users.

Cloudflare AI BlogSource content · Analysis pendingIdentify AI model overuse with User Insights

We need ‘right to intervene’ in AI amid growing threat, says Bank of England boss

Andrew Bailey’s comments come as fears grow that rogue models could take financial system hostage Business live – latest updates The governor of the Bank of England has said authorities must retain the “right to intervene” in the AI industry amid growing fears that rogue models could take the financial system hostage. Andrew Bailey said the risks posed by the rapid advancement of frontier AI models – a number of which have gone rogue in recent months – were “real and increasingly significant” and reduced the ability of society to supervise and intervene when things went wrong. Continue reading...

The Guardian AISource content · Analysis pendingWe need ‘right to intervene’ in AI amid growing threat, says Bank of England boss

Test-Time Adaptation of Manipulation Policies Under Actuator Degradation

arXiv:2609.36182v1 Announce Type: new Abstract: Robot manipulation policies are usually trained under the assumption that a commanded action produces the same motion as it did during training even after hours of operation. Real hardware violates this assumption as the motors gradually heat up, current saturates near contact, voltage sags under load, thus the same policy action can produce a weaker, delayed, or noisier motion. These conditions are already measured by onboard telemetry, such as joint temperature, motor current, and supply voltage, yet this signal is typically used only for logging or safety checks rather than policy adaptation. We introduce Telemetry-Aware Action Rectification (TeAR), a policy-agnostic method that turns a frozen manipulation policy into a telemetry-conditio…

arXiv RoboticsSource content · Analysis pendingTest-Time Adaptation of Manipulation Policies Under Actuator Degradation

KPI: A Promptable Kernel for Physical Interaction on Humanoids

arXiv:2609.36151v1 Announce Type: new Abstract: Humanoids now walk, balance and reach with remarkable generality: one whole-body tracking policy follows references from a human, or from an end-to-end policy. That generality travels in the trajectory, and a trajectory alone carries limited information about the interaction it should produce: at contact, the executing controller determines how the robot behaves. Single-task policies usually reach hard interactions by optimising trajectory and controller together in simulation; general stacks usually assume a preset or hand-chosen controller. We present KPI, a promptable kernel for physical interaction between the trajectory source and an unmodified whole-body tracker. Instead of a controller fixed before the task, the trajectory source send…

arXiv RoboticsSource content · Analysis pendingKPI: A Promptable Kernel for Physical Interaction on Humanoids

Xiaomi-OCR-0 Technical Report

arXiv:2609.36136v1 Announce Type: new Abstract: Compact OCR-specific vision-language models achieve strong document parsing performance, but often rely on costly supervision and focus primarily on visual-text reconstruction. We introduce Xiaomi-OCR-0, a unified 0.8B model for document parsing and OCR-centric understanding. We build an approximately 170M-sample OCR-centric corpus using an automated data engine that combines expert consensus, render-based verification, and targeted synthesis. Starting from Qwen3.5-0.8B, our progressive training recipe combines Q-Mask-based text anchoring, continued pretraining, and mixed-task reinforcement learning (Mix-RL). Xiaomi-OCR-0 achieves 95.24 on Real5-OmniDocBench, 96.83 on OmniDocBench v1.6, and 87.94 on Wild-OmniDocBench, while reaching an avera…

arXiv Computer VisionSource content · Analysis pendingXiaomi-OCR-0 Technical Report

One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models

arXiv:2609.36101v1 Announce Type: new Abstract: Contrastive vision-language models learn shared embedding spaces by aligning matched image-text pairs, yet their representations remain separated by a modality gap. Prior work reports divergent effects of modifying this gap: reducing it can improve zero-shot classification and cross-modal alignment, whereas removing gap-related structure can degrade image-text retrieval. In this paper, we provide a unified geometric explanation for these task-dependent effects. Across CLIP and SigLIP encoders, we find that a single dominant direction captures 94.4-99.9% of the squared norm of the image-text mean separation, revealing that the mean-separation component is approximately rank-one. A decomposition of the similarity score then identifies three ta…

arXiv Computer VisionSource content · Analysis pendingOne Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models

Persistence Forcing: Exploiting Feature Specialization in Pixel-Space Diffusion

arXiv:2609.36014v1 Announce Type: new Abstract: Pixel-space diffusion Transformers (DiTs) directly operate on high-dimensional visual data, yet their hidden representations typically undergo uniform refinement across depth. Natural images, however, are inherently organized at different levels of granularity. Global structure can often be represented compactly, whereas local textures and fine details require richer representations. Motivated by this, we introduce heterogeneous refinement in pixel-space DiTs, assigning different feature groups distinct refinement budgets across depth. Consequently, an ordered feature specialization emerges: sparsely refined features predominantly encode global visual structure, whereas more frequently refined features increasingly specialize toward localize…

arXiv Computer VisionSource content · Analysis pendingPersistence Forcing: Exploiting Feature Specialization in Pixel-Space Diffusion

HERO: Histology Encoder for Robust Representation in Oncology

arXiv:2609.35943v1 Announce Type: new Abstract: Foundation models trained on large pathology image corpora now provide strong, transferable representations for computational pathology. Over the past few years a series of such models has been released, each trained on more slides than the last; on standard classification and segmentation benchmarks, the leading models are now separated by small margins. In clinical use, however, the foundation model is applied to images from hospitals, scanners, and staining protocols outside its training data. Encoders generally embed these acquisition factors alongside biological information, which may introduce downstream errors and hinder safe clinical adoption. A pathology foundation model should therefore be robust to acquisition shift without giving…

arXiv Computer VisionSource content · Analysis pendingHERO: Histology Encoder for Robust Representation in Oncology

TRACE: Deployable Tree-Relational Structure Enhancement for Oncology LLMs

arXiv:2609.35810v1 Announce Type: new Abstract: Large language models are increasingly used in oncology applications, but their predictions are often weakly grounded in explicit medical structure. We present TRACE, a deployable tree-relational enhancement framework for oncology LLMs. TRACE separates expensive offline structure learning from lightweight online inference: oncology concepts and relations are organized into an updatable tree-relational structure, refined using LM-loss-derived evidence, and retrieved at inference time as compact prompt evidence. This design supports task-adaptive evidence selection without requiring supervised labels in the zero-shot setting. Across ten oncology classification tasks and one MedQuAD CancerGov QA benchmark, TRACE improves both label-free evaluat…

arXiv Computational LinguisticsSource content · Analysis pendingTRACE: Deployable Tree-Relational Structure Enhancement for Oncology LLMs

Can Multimodal Large Language Models Generate and Detect Multimodal Social Media Fake News?

arXiv:2609.35809v1 Announce Type: new Abstract: The rapid advancement of generative AI raises concerns about the misuse of Multimodal LLMs (MLLMs) for large-scale disinformation campaigns on social media. Despite existing research on textual disinformation, a fundamental question remains unanswered: can MLLMs be exploited to fabricate realistic multimodal fake news, and can they reliably detect it? We introduce a multi-agent framework in which a story agent, an image agent, and a critic agent collaborate to produce fake social media posts that plausibly counter true news. We apply the framework to generate over 9,000 paired multimodal news posts across science, health, and entertainment domains, and benchmark 16 open- and closed-source MLLMs for automated detection. We find that most mode…

arXiv Computational LinguisticsSource content · Analysis pendingCan Multimodal Large Language Models Generate and Detect Multimodal Social Media Fake News?

Alignment Forecasting: Predicting Misalignment From Training Data

arXiv:2609.35805v1 Announce Type: new Abstract: Training a language model on data with a narrow flaw can sometimes make the model broadly misaligned. Inspecting the data at face value often does not settle whether it will emerge, and today it is caught only after training, by auditing the resulting model. To complement post-hoc audits, we introduce Alignment Forecasting: the task of predicting alignment failures before training. Given a target model, a fine-tuning dataset, and a failure mode such as deception or sycophancy, a forecaster outputs the probability that fine-tuning would meaningfully increase that failure mode. To measure progress on alignment forecasting, we introduce ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions spanning 17 target models, 32 dataset…

arXiv Computational LinguisticsSource content · Analysis pendingAlignment Forecasting: Predicting Misalignment From Training Data

Evaluating the Effects of Prompt Perturbation on Bias and Hallucination in Large Language Models

arXiv:2609.35804v1 Announce Type: new Abstract: Large language models (LLMs) have shown remarkable capabilities in various natural language processing tasks, leading to their widespread deployment as intelligent assistants in decision-making contexts. However, the increasing complexity of these models raises concerns about their reliability, particularly regarding bias and hallucination. In this work, we evaluate the robustness of LLMs to perturbed variations of the original inquiry in decision-making tasks. We show that contrary to previous studies, perturbations can mitigate bias and hallucination in some LLMs over other models. It's found that Claude 3 is more effective for the tasks represented in most datasets, whereas models like GPT3.5 exhibit varying levels of adequacy, performing…

arXiv Computational LinguisticsSource content · Analysis pendingEvaluating the Effects of Prompt Perturbation on Bias and Hallucination in Large Language Models

Sieve and Sage: Efficient Distraction Filtering for Reliable RALM Abstention

arXiv:2609.35794v1 Announce Type: new Abstract: Just as Socrates recognized the limits of his own knowledge, Retrieval-Augmented Language Models (RALMs) should learn to abstain when the retrieved evidence cannot support a reliable response. Existing approaches largely rely on monolithic LLMs to handle heterogeneous retrieval failures in a single step, resulting in limited abstention performance and high computational costs. We instead decompose retrieval failures into two distinct states: (i) the unanswerable state, where the required evidence is absent, and (ii) the distracted state, where relevant evidence is mixed with conflicting, negated, or adversarial information. Based on this decomposition, we introduce a lightweight module (Sieve) that screens retrieved document sets for distrac…

arXiv Computational LinguisticsSource content · Analysis pendingSieve and Sage: Efficient Distraction Filtering for Reliable RALM Abstention

FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech

arXiv:2609.35791v1 Announce Type: new Abstract: Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent. Acoustic voice activity detection lacks this semantic information, while cascaded ASR-based endpointing introduces transcription dependence and additional processing stages. We formulate semantic endpoint detection as a causal audio-language reasoning task and introduce FD-VAD, an ASR-free streaming endpointer that maps bounded causal audio windows directly to Continue/Stop decisions. FD-VAD combines a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model, using a last-chunk training objective for streaming inference. We f…

arXiv Computational LinguisticsSource content · Analysis pendingFD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech

A Mesoscopic View of Transformer Weights Through Row and Column Scale Fields

arXiv:2609.35852v1 Announce Type: new Abstract: Pooled statistics of Transformer weights obscure how magnitude is distributed across functional channels, while individual weights are too numerous to compare directly. We study the mesoscopic level between them: row and column scale fields, the median-centred log-RMS profiles of a weight matrix over its channels, which together with a global scale and a full balanced core represent the matrix exactly. Across public Pythia checkpoints at four sizes and controlled runs from three initialization families, balancing reveals similar measured core magnitude profiles. A mixture bridge, with its form fixed before the analysis and its coefficients fitted, predicts the pooled-shape departure from field width on held-out runs and data arms of the cont…

arXiv Machine LearningSource content · Analysis pendingA Mesoscopic View of Transformer Weights Through Row and Column Scale Fields

HeadGuard: Selective Head Protection for Low-Bit VLM KV-Cache Quantization

arXiv:2609.35800v1 Announce Type: new Abstract: Low-bit key-value (KV) cache quantization saves storage but can sharply degrade vision-language model (VLM) accuracy. We introduce HeadGuard, a composable head-protection method that augments a base KV-cache quantizer with a fixed high-precision mask. Image-sensitivity and output-sensitivity scores select physical KV heads offline, with approximately 1/8 protected in the main experiments; their image keys and optionally values remain in bfloat16 (BF16), while the base quantizes unprotected image entries. Across eight VLMs, three base quantizers, and eight benchmarks (six discriminative and two generative), HeadGuard recovers a substantial fraction of lost accuracy on weaker quantizers, with the strongest gains for Qwen and InternVL. At 2 bit…

arXiv Machine LearningSource content · Analysis pendingHeadGuard: Selective Head Protection for Low-Bit VLM KV-Cache Quantization

Binarization Flattens the Score Space

arXiv:2609.35797v1 Announce Type: new Abstract: Large language model (LLM) judges are often used as rewards to train policies on objectives that deterministic verifiers cannot capture. However, these rewards are often collapsed to pass/fail ({0, 1}), which reports the verdict but not how well a response met each criterion. We model each pass/fail verdict as a score on an unreported scale, compared with one cutoff. A stretch of that scale moves every score proportionally toward or away from the cutoff, but never across it, so no verdict changes. A policy is therefore free to apply any stretch without changing anything the panel reports. Under a joint-Gaussian model, a third grade adds a second threshold and removes this affine stretch ambiguity. On MATH and SciBench outputs from one seven-…

arXiv Machine LearningSource content · Analysis pendingBinarization Flattens the Score Space

Calibration-First Cross-Cohort Multimodal Temporal Learning for Transferable Asthma-Risk Forecasting

arXiv:2609.35795v1 Announce Type: new Abstract: Asthma deterioration forecasting must remain reli- able when patient populations, sensor ecosystems, and available modalities change across cohorts. Existing models commonly optimize within-cohort discrimination and may produce poorly calibrated probabilities after transfer. We present CALIBRA, a calibration-first multimodal temporal framework for short- horizon risk prediction with incomplete data. Dedicated recurrent encoders process environmental, pulmonary, symptom, medication, wearable, and context streams; a reliability-conditioned gate suppresses stale or absent modalities, while gradient-reversal training discourages avoidable cohort signatures. A shrinkage- based hierarchical logistic layer calibrates probabilities using a patient-d…

arXiv Machine LearningSource content · Analysis pendingCalibration-First Cross-Cohort Multimodal Temporal Learning for Transferable Asthma-Risk Forecasting

Learning from the Gap Between Pass@K and Pass@1

arXiv:2609.35793v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly trained with reinforcement learning from verifiable rewards (RLVR). An exact verifier can also support test-time scaling by selecting a passing response from multiple samples, while other deployments use beam search, adaptive sampling, or tools. We study single-sample decoding, where each query receives one response without search, to ask whether search-exposed behavior can be absorbed into the model. Existing verified-response post-training recipes do not generally distinguish problems already solved on the first decode from failures recovered within K samples. Under a fixed budget, this can spend examples repeating behavior the deployed policy already has. We introduce GapFT, which selects trai…

arXiv Machine LearningSource content · Analysis pendingLearning from the Gap Between Pass@K and Pass@1

Grab a Coffee: Future-Aware Guidance for Discrete Diffusion with Compiled Objectives

arXiv:2609.35924v1 Announce Type: new Abstract: Discrete diffusion models generate sequences by iteratively resolving multiple tokens in parallel, offering a flexible alternative to left-to-right generation. However, guiding this process with a sequence-level objective is difficult because the value of one unresolved token depends on the other tokens with which it can form a high-reward sequence. Enumerating all such completions makes the whole guidance computation grow exponentially with the number of unresolved positions. We introduce COFFEE, a plug-and-play framework that avoids this enumeration by separating sequence dependence from the objective. At each diffusion step, a target-free carrier absorbs the marginal token distributions predicted by the denoiser to construct a joint model…

arXiv AISource content · Analysis pendingGrab a Coffee: Future-Aware Guidance for Discrete Diffusion with Compiled Objectives

Self-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden?

arXiv:2609.35897v1 Announce Type: new Abstract: The pursuit of recursive self-improvement (RSI) toward general intelligence is divided between macro-level language model scaling and the interaction-driven principles of "Era of Experience". Yet, any self-improving architecture ultimately rests upon its underlying optimization engine: if general intelligence requires learning from grounded interaction, the reinforcement learning (RL) update rule itself must be capable of cumulative adaptation. While algorithm self-discovery has produced Disco103 that surpassed PPO to achieve SOTA benchmark performance -- its internal update machinery remains an uninspected black box. We present the first causal mechanistic audit of a self-discovered RL rule, structured directly around the five pillars of th…

arXiv AISource content · Analysis pendingSelf-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden?

Representational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial Training

arXiv:2609.35890v1 Announce Type: new Abstract: Sparse-autoencoder decomposability and concentrated feature attribution are increasingly treated as evidence that a model's computation is easier to reverse-engineer. Whether this representational and attributional cleanliness actually predicts a smaller or more tractable causal circuit remains an open question. We test this directly using adversarial training as a controlled instrument: it reliably reshapes internal representations, but this alone does not constitute a test of circuit size. We investigate this question through reverse-engineering complexity: the causal structure required to recover a model's behavior at a fixed level of faithfulness. To our knowledge, this is the first controlled empirical test of whether representational o…

arXiv AISource content · Analysis pendingRepresentational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial Training

More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses

arXiv:2609.35873v1 Announce Type: new Abstract: Automated generation of LLM harnesses promises to improve inference through task specialization. Yet additional answer coverage can arise from repeated execution of the same program, making specialization difficult to identify. We introduce a controlled evaluation that separates answer coverage, repeatable task advantages, and gains from pre-execution selection. On 386 MATH-500 tasks, we compare eight generated harnesses plus a baseline with nine byte-identical baseline copies, using three executions per member. Identical programs yield 2.16 percentage points of repeat-averaged oracle headroom. Generated programs exhibit substantially more repeatable score patterns, but these chiefly reveal persistent weaknesses: losses relative to the basel…

arXiv AISource content · Analysis pendingMore Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses

The Price of Token Boundaries: Compression Certificates and Prediction

arXiv:2609.35869v1 Announce Type: new Abstract: Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule. Nonnegative prices on token occurrences yield a lower bound through shortest paths and vocabulary-budget selection; maximising over all prices recovers the linear programming relaxation, and an independent integer checker certifies the reported values. On English Wikipedia, boundaries increase the optimal token count by 28.3--36.8\%. Byte pair encoding lies 2.1\% above the constrained lower bound, but 10.9\% above the unrestricted bound. Compres…

arXiv AISource content · Analysis pendingThe Price of Token Boundaries: Compression Certificates and Prediction

Is Human-Readable Text Necessary for Effective LLM Fine-Tuning?

arXiv:2609.35868v1 Announce Type: new Abstract: Is human readability necessary for effective fine-tuning of large language models? We investigate whether model-conditioned training representations can preserve or improve adaptation utility without requiring a human-readable textual form. We propose Desired-Update-Aligned Synthetic Data (DASA), which uses activation-gradient feedback from a frozen reference model to guide the optimization of continuous synthetic input embeddings. Inspired by the role of activation gradients in local risk reduction, DASA targets useful adaptation updates rather than source-text reconstruction or linguistic fluency. The resulting embeddings are used directly for downstream fine-tuning; discrete token projections are employed only for qualitative inspection.…

arXiv AISource content · Analysis pendingIs Human-Readable Text Necessary for Effective LLM Fine-Tuning?

Amazon Bedrock expands Claude model availability to in-country inferencing in India

Anthropic's Claude Opus 5, Claude Sonnet 5, and Claude Haiku 4.5 are now available in India through Amazon Bedrock geographic cross-Region inference. You can access these models while processing data within the India Regions, and get started from the Amazon Bedrock console or with the Messages, InvokeModel, and Converse APIs.

AWS Machine Learning BlogSource content · Analysis pendingAmazon Bedrock expands Claude model availability to in-country inferencing in India

Introducing Anthropic models on Amazon Bedrock for in-region inference in Seoul and Singapore

Amazon Bedrock now supports Anthropic's Claude Opus 5 and Claude Sonnet 5 with in-region inference in Seoul, and Claude Sonnet 5 in Singapore. If you have local data processing requirements in South Korea or Singapore, you can now use these Anthropic models at scale, with inference processed entirely within the Region you call.

AWS Machine Learning BlogSource content · Analysis pendingIntroducing Anthropic models on Amazon Bedrock for in-region inference in Seoul and Singapore

OpenAI won’t go public until its models are safe

For months, people have wondered when OpenAI will go public. CEO Sam Altman says it won't happen until the company can make better promises about model safety, with no firm timeline in sight. "We intend to continue with AI progress … but as the models have had this surge forward in capability, and we see more of that ahead of us, we have got to be able to make confident safety claims," Altman said Tuesday during a Q&A with reporters after his DevDay keynote. At the same time, he said, waiting too long for an IPO would be "bad for the world." Altman's comments come after months of controversy about whether or not OpenAI and its rivals can c … Read the full story at The Verge.

The Verge AISource content · Analysis pendingOpenAI won’t go public until its models are safe

On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study

Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the involved trade-offs remains elusive. Current approaches to conditioning are often evaluated with a narrow focus on their effectiveness at injecting or removing a target concept, neglecting generation quality. We systematically investigate a range of conditioning methods in both injection and removal scenarios. We find that efficient steering methods frequently achieve conditioning at a steep cost to fluency. Furthermore, we identify a critical yet…

Apple Machine Learning ResearchSource content · Analysis pendingOn the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study

Quoting Anthropic Frontier Red Team

We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them. — Anthropic Frontier Red Team, GLM-5.3 and the spread of advanced cyber capabilities Tags: anthropic, generative-ai, ai-security-research, glm, ai, ai-in-china, llms

Simon Willison's WeblogSource content · Analysis pendingQuoting Anthropic Frontier Red Team

Bring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock

GPT-6.1 Sol is now generally available on Amazon Bedrock, bringing stronger reasoning to coding, computer use, and professional workloads that run frequently.

AWS Machine Learning BlogSource content · Analysis pendingBring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

My comment on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price — Hacker News. I'm a bit late with the pelicans because I was live-blogging the keynote: https://simonwillison.net/2026/Sep/29/openai-devday-2026-liv... Here they are for GPT-6.1-Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... They're not notably different from the GPT-6 family pelicans: https://static.simonwillison.net/static/2026/gpt-pelicans-gr... Tags: ai, openai, generative-ai, llms, pelican-riding-a-bicycle, gpt

Simon Willison's WeblogSource content · Analysis pendingGPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

OpenAI’s new GPT-6.1 Sol undercuts its own Astra flagship

OpenAI on Tuesday launched GPT-6.1 Sol, the newest version of its workhorse model, only a week after launching GPT-6 Sol. The post OpenAI’s new GPT-6.1 Sol undercuts its own Astra flagship appeared first on The New Stack.

The New Stack AISource content · Analysis pendingOpenAI’s new GPT-6.1 Sol undercuts its own Astra flagship

Photo Scrubber — local face blur & metadata removal

Tool: Photo Scrubber — local face blur & metadata removal I took a photograph of some protesters, then thought about how I don't like sharing photographs of strangers with identifiable faces. I had GPT-6 Astra build this experimental tool that would identify faces and automatically blur them out. It uses Google's MediaPipe C++ library, compiled to WebAssembly via @mediapipe/tasks-vision, plus the BlazeFace face detection model. Tags: photography, tools

Simon Willison's WeblogSource content · Analysis pendingPhoto Scrubber — local face blur & metadata removal

Cohere Embed 5: Frontier Embedding Models for Enterprise

Key takeaways State-of-the-art enterprise retrieval: Embed 5 Pro achieves the highest average score of any model we tested - particularly across financial datasets, parsed PDFs, and visually rich documents. A new Fast t…

Cohere BlogSource content · Analysis pendingCohere Embed 5: Frontier Embedding Models for Enterprise
Policy

Here’s how tech leaders will self-police AI safety under Trump’s deal

Anyone else getting “The Last Supper” vibes from this photo taken at the White House tech gathering? | Photographer: Tierney L. Cross/The Washington Post/Bloomberg via Getty Images We now have the full details of the "morally binding" AI safety deal announced by President Trump yesterday, in which top executives agreed to self-regulate their artificial intelligence technology. The accord, officially titled the Joint Commitment On Frontier Responsibilities, was shared online by tech founder and presidential advisor David Sacks, and has been signed by Google's Sundar Pichai, Anthropic's Dario Amodei, Meta's Mark Zuckerberg, OpenAI's Greg Brockman, XAI's Elon Musk, and Nvidia's Jensen Huang. "In order to build a positive future for the American people and the world, we believe every company…

The Verge AISource content · Analysis pendingHere’s how tech leaders will self-police AI safety under Trump’s deal

Trump orders US government to call AI ‘Super Intelligence’

Image: Carolyn Van Houten/The Washington Post via Getty Images The US executive branch is no longer acknowledging the existence of "artificial intelligence." Going forward, official policy websites, policy documents, and press releases will refer only to "Super Intelligence," thanks to a new executive order signed by President Donald Trump. "The word super is the best word of all, and it's the simplest," Trump said at an event announcing the launch of America.gov earlier on Tuesday, and said Chinese President Xi Jinping, who visited the White House last week, "loves it," too. "We don't want to hear artificial. Because it's not artificial. It's very powerful, it's very brilliant. It's going to mostly be … Read the full story at The Verge.

The Verge AISource content · Analysis pendingTrump orders US government to call AI ‘Super Intelligence’
Robotics

MagNav: A Dual-Core Magnetic Track Guidance Framework for Lighting-Invariant Navigation in Two-Wheeled Robots

arXiv:2609.36091v1 Announce Type: new Abstract: Two-Wheeled Inverted Pendulum (TWIP) robots are useful for studying how to control systems that are naturally unstable and have fewer actuators than degrees of freedom. Adding autonomous line-following to these robots is challenging because steering and balancing are closely linked. Most existing systems use infrared sensors, which can be affected by changes in lighting, such as sunlight or shadows, making them reliable only indoors. This paper presents a self-balancing robot that can follow a line using a magnetic track guidance system. By using a five-channel analog Hall-effect sensor array, the robot is not affected by optical interference. The control system uses a cascaded PID structure: the inner loop keeps the robot balanced using dat…

arXiv RoboticsSource content · Analysis pendingMagNav: A Dual-Core Magnetic Track Guidance Framework for Lighting-Invariant Navigation in Two-Wheeled Robots

Passive-Dynamic-Walking-Inspired Dynamics Guidance for Energy-Efficient Humanoid Locomotion

arXiv:2609.35935v1 Announce Type: new Abstract: Learning energy-efficient humanoid locomotion requires discovering mechanically economical gait coordination, not merely reducing actuator effort. Reinforcement learning promotes efficiency through effort-related reward penalties, which guide the step-to-step mechanics of walking only indirectly. This article proposes a framework inspired by passive dynamic walking (PDW) that temporarily creates slope-equivalent conditions favorable to economical gait discovery and removes all PDW-specific guidance before nominal-dynamics optimization. During early training, a tilted-gravity field assists sagittal progression on flat collision geometry, complemented by curriculum-coupled reward terms. The core framework requires no reference trajectories, ga…

arXiv RoboticsSource content · Analysis pendingPassive-Dynamic-Walking-Inspired Dynamics Guidance for Energy-Efficient Humanoid Locomotion

Protesters gather at OpenAI’s DevDay

On Tuesday, OpenAI's annual DevDay event began with protests, flyers and chants. "Sam Altman, get off it, put people over profit," said a group of protesters marching in a circle in front of a series of signs that spelled out "PEOPLE OVER PROFIT." More than a dozen organizations came together to sponsor a rally outside the entrance to the event venue at San Francisco's Fort Mason, including Bay Resistance, Tech Workers Coalition, Service Employees International Union, San Francisco Rising, and the San Francisco Labor Council. A large orange sign read "Drop the ICE contract." Multiple people dressed in cardboard robot costumes waved cardboa … Read the full story at The Verge.

The Verge AISource content · Analysis pendingProtesters gather at OpenAI’s DevDay
Chips

Neurosymbolic Routing for Reliable Reasoning on Resource-Constrained Edge Devices

arXiv:2609.35833v1 Announce Type: new Abstract: Running a language model on edge hardware provides private and low-latency reasoning without a network connection, and yet the small models that fit on such devices are unreliable on the tasks computers are expected to handle well, such as arithmetic, algebra, and formal logic problems. We argue that much of this unreliability is avoidable. Many queries appearing to demand reasoning are in fact structurally deterministic and permit fast and exact symbolic solutions. Therefore, forcing a probabilistic model to approximate them sacrifices accuracy and energy for little benefit. We present a neurosymbolic router that classifies each incoming query and dispatches it to the cheapest correct solver, sending structured tasks to deterministic engine…

arXiv AISource content · Analysis pendingNeurosymbolic Routing for Reliable Reasoning on Resource-Constrained Edge Devices