Skip to content
AI News HubLIVE

This edition’s highlights

Research

Introducing MentalHealthBench

  • Co-created with 80+ licensed mental health experts across 22 countries, speaking 19 languages and covering nearly 20 subspecialties.
  • Covers non-acute, high-acuity, and emergency conversations, plus adult, teen (13–17), caregiver, and clinician personas.
OpenAI NewsIn-site articleIntroducing MentalHealthBench

Peerify: Benchmarking Peer-Review Claim Verification

  • Peerify decomposes reviews into atomic claims, retrieves manuscript evidence, and judges whether each claim is supported.
  • The benchmark contains 800 claims from real NeurIPS 2024 and ICLR 2024 reviews, with a 300-claim hand-labeled subset for auditing automated supervision.
arXiv Computational LinguisticsIn-site articlePeerify: Benchmarking Peer-Review Claim Verification

Exposing Blind Spots in Deep Imbalanced Regression Evaluation

  • DIR evaluation suffers from three blind spots: image-benchmark dominance, a non-decision-complete protocol, and unexamined tail stability.
  • The study evaluates nine time-series extrinsic regression tasks across six physical domains on the MuViS multimodal virtual sensing benchmark.
arXiv Machine LearningIn-site articleExposing Blind Spots in Deep Imbalanced Regression Evaluation

Federating Quantum and Classical Computing: A Privacy-Preserving Hybrid Approach

  • The work pairs a hybrid quantum-classical active party with a classical passive party via federated learning, avoiding any centralization of raw data.
  • Sherpa.ai's Blind Vertical FL (SBVFL) protocol is used to cut communication overhead drastically.
arXiv Machine LearningIn-site articleFederating Quantum and Classical Computing: A Privacy-Preserving Hybrid Approach
Agents

Grab and OpenAI bring practical AI skills to Southeast Asia

  • GO Forward with AI will reach 30,000 Grab driver-, delivery-, and merchant-partners over two years.
  • It starts in Singapore, expands to Thailand, Indonesia and the Philippines this year, then Malaysia and Vietnam in 2027.
OpenAI NewsIn-site articleGrab and OpenAI bring practical AI skills to Southeast Asia
Models

[AINews] Claude Opus 5.5, the new default model for AINews — and everybody cuts prices 40-50%

  • Opus 5.5 leads on agentic coding, computer use and knowledge work per Anthropic, is about 30% faster, and cuts token prices 20% from $5/$25 to $4/$20 per 1M tokens.
  • Artificial Analysis finds max-effort per-task cost is essentially flat versus Opus 5 ($5.98 vs $5.86) because higher token usage offsets the price cut, so the "40% cheaper" claim applies mainly at default medium effort.
Latent SpaceIn-site article[AINews] Claude Opus 5.5, the new default model for AINews — and everybody cuts prices 40-50%

OpenAI releases GPT-6 Sol and Luna — and cuts token prices in half

  • Pricing: Sol costs $2/$10 per million input/output tokens and Luna $0.10/$0.50, less than half the GPT-5.6 rates, and this is now default pricing rather than a promotion.
  • Performance: Luna gains 5.4 percentage points on Zapier's AutomationBench; on DeepSWE v1.1, Sol essentially matches Anthropic's Fable 5 at about 20% of the cost.
The New Stack AIIn-site articleOpenAI releases GPT-6 Sol and Luna — and cuts token prices in half

Same Quantity, Different Answer: Numerical Representation Invariance in Language Models

  • Five open-weight models were tested on 3,600 exact-rational problems and 8,600 prompts across five identity-preserving transformations.
  • Canonical accuracy reached 0.969–0.996, but orbit correctness and invariance fell to 0.848–0.981 and 0.851–0.981.
arXiv Computational LinguisticsIn-site articleSame Quantity, Different Answer: Numerical Representation Invariance in Language Models
Policy

Deepfakes and Synthetic Media: Generation, Detection, and Governance

  • Surveys deepfake generation via GANs, diffusion models, neural rendering and video synthesis.
  • Detection uses spatial, temporal, frequency-domain and physiological artifacts, with CNN, transformer and frequency-based detectors.
arXiv Computer VisionIn-site articleDeepfakes and Synthetic Media: Generation, Detection, and Governance

You have reached the end of this edition.

Your reading list
Other updates (116)
Agents

MIT welcomes David Siegel SM ’86, PhD ’91 as its next Innovation Fellow

David Siegel, a computer scientist, entrepreneur, and philanthropist, will serve as MIT’s next Innovation Fellow in 2026-27, working with the MIT Schwarzman College of Computing on how AI can accelerate scientific discovery.

MIT News AIIn-site articleMIT welcomes David Siegel SM ’86, PhD ’91 as its next Innovation Fellow

YouTube is building AI creator tools that do almost everything for them

At its Made on YouTube event, YouTube announced updates to its AI creator tools, including a background agent that optimizes channels, generates thumbnails and titles, and suggests brand pitches. New testing features include dynamic thumbnails and up to three versions of a video. YouTube declined to share data proving the tools improve metrics, saying they save creators time.

The Verge AIIn-site articleYouTube is building AI creator tools that do almost everything for them

The Sequence Learning Loop - Issue 938: Learn About the Amazing Jev, Gemini and Paper2Agent

This week's issue looks at three developments: TypeSafe's Jev for structured decisions, two new Gemini Live models from Google that treat conversation and reasoning differently, and Stanford's Paper2Agent reaching Nature as a way to turn research methods into reusable agent tools. The throughline: the interface around a model deserves as much attention as the model itself.

TheSequenceIn-site articleThe Sequence Learning Loop - Issue 938: Learn About the Amazing Jev, Gemini and Paper2Agent

IntellAgents.io: One AI Agent for Every Call, Chat, and DM

IntellAgents.io has surfaced on Product Hunt with a one-line pitch: a single AI agent for every call, chat, and DM. The listing currently offers little beyond that positioning statement and a discussion link.

Product Hunt AIIn-site articleIntellAgents.io: One AI Agent for Every Call, Chat, and DM

MCP Is Not Just Another API Standard

Ask most engineers what MCP is and you’ll get the same answer: a way to plug tools into an LLM. Fair enough, as far as it goes. But that description treats MCP like plumbing, and after months building MCP-based integrations for large enterprise platforms, I don’t think plumbing is the right metaphor. Plumbing moves water […]

O'Reilly AI & ML RadarSource content · Analysis pendingMCP Is Not Just Another API Standard

LaterOn v2: The agentic email platform

LaterOn v2 is an agentic email platform that lets users describe an email job once and have it handled automatically.

Product Hunt AIIn-site articleLaterOn v2: The agentic email platform

SpeakON Ships a MagSafe AI Voice Button With Its Own Microphone

SpeakON launched a 25 g MagSafe AI voice button with its own mic, battery, and storage that writes processed text directly into any iPhone text field via an iOS keyboard extension. It supports offline buffering, text shaping features, and 12-language translation, priced at $129 one-time in the US, with SpeakON Agent coming in October 2026.

MarkTechPostIn-site articleSpeakON Ships a MagSafe AI Voice Button With Its Own Microphone

Social Influence and the Allocation of Scientific Attention in AI Populations

A new arXiv paper transplants the Music Lab social-influence design into a market for academic attention. In one experiment, 1,000 AI agents chose among all 114 regular research articles published in the American Economic Review in 2025; agents who could see earlier choices in their community picked 17.2 percent fewer papers each, concentrated their selections more heavily, and collectively covered only 73 papers versus 90 under independent choice. A second experiment found that randomly assigning papers five initial selections lifted their subsequent selection rate by 45.55 percentage points.

arXiv AIIn-site articleSocial Influence and the Allocation of Scientific Attention in AI Populations

SpeakON Ships a MagSafe AI Voice Button With Its Own Microphone: Turning Your Voice into Polished Communication, and Action across Apps

SpeakON has released a 25 g MagSafe button that carries its own microphone and battery and writes finished text straight into any iPhone text field via a system-wide iOS keyboard extension. It works while locked or offline, ships with Smart Polish, Smart List, Style, Translation and other text-shaping features, and costs $129 one time in the US with Pro Lifetime included.

MarkTechPostIn-site articleSpeakON Ships a MagSafe AI Voice Button With Its Own Microphone: Turning Your Voice into Polished Communication, and Action across Apps

SF October 14th: A Birds of a Feather Session on Agentic Engineering

Simon Willison and Jesse Vincent are hosting an evening birds-of-a-feather gathering in San Francisco on Wednesday, October 14th for people building strange and interesting things with and on top of coding agents. Framed as an "agentic show-and-tell," the event favors early-stage experiments, unfinished projects and privately held work over product pitches.

Simon Willison's WeblogIn-site articleSF October 14th: A Birds of a Feather Session on Agentic Engineering

Kaiku

Kaiku is a task tracker pitched as already usable by AI agents, with a Product Hunt listing that describes it as “The task tracker your AI agents already know how to use.”

Product Hunt AIIn-site articleKaiku

How to serve trillions of tokens for trillion-parameter coding agents

Learn how we optimized performance and efficiency serving the workload that is changing software engineering forever — and you can too.

Modal BlogSource content · Analysis pendingHow to serve trillions of tokens for trillion-parameter coding agents

Rabbit’s new AI agent doesn’t need an R1 to run

Rabbit is rolling out a standalone AI agent that works without its R1 hardware. The cloud-based OS3 “agentic operating system” operates locally across Windows, Mac, and Linux, supports up to five devices and preferred AI models per account, and can be reached via a desktop site, Telegram or iMessage, or the R1. Rabbit has stopped manufacturing the R1 and is instead preparing a vibe-coding “cyberdeck” that runs OS3. The company says OS3 won’t store, copy, use, or sell your data, though chats and memories remain on its servers, and linked third-party AI providers follow their own privacy rules.

The Verge AIIn-site articleRabbit’s new AI agent doesn’t need an R1 to run

Jev State

Jev State is a Product Hunt listing that pitches turning AI conversations into tests and runnable code. Beyond that one-line description, little is public about how the workflow actually works or which frameworks it supports.

Product Hunt AIIn-site articleJev State

“One of the most significant steps in our 26-year history”: JetBrains goes big on agentic development — and bets the IDE still matters

JetBrains has unveiled JetBrains Air, an open system of products for agentic software development spanning IDE, team, and governance layers. The announcement consolidates earlier efforts like the Air desktop environment, Junie CLI, JetBrains Central, and AI for Teams, while CEO Kirill Skrygan insists the IDE remains central to reviewing and shipping code.

The New Stack AIIn-site article“One of the most significant steps in our 26-year history”: JetBrains goes big on agentic development — and bets the IDE still matters

Quoting @therealcornpop

A TikTok creator argues that AI-written scripts for TikTok and YouTube are obvious not because of surface-level "AI-isms" but because they lack a distinct voice and any real opinion about the subject. Simon Willison collected and posted the quote on 22nd September 2026.

Simon Willison's WeblogIn-site articleQuoting @therealcornpop

Evaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore

Skills let you encode domain-specific procedures as reusable, portable instructions for agents, but a fluent answer doesn't prove the agent picked the right skill or followed it. Learn how to measure skill selection and instruction following with Strands Evals and Amazon Bedrock AgentCore Evaluations.

AWS Machine Learning BlogSource content · Analysis pendingEvaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore

The Genie One MCP is now Generally Available

AI coworkers and coding agents are spreading fast across organizations, and each...

Databricks BlogSource content · Analysis pendingThe Genie One MCP is now Generally Available

AI Change Management | Cohere

Change management has long covered two familiar kinds of change: the rollout of new tools and technologies, and broader human-led transformations such as leadership changes and restructuring. But that distinction starts…

Cohere BlogSource content · Analysis pendingAI Change Management | Cohere

The frontier isn’t a model. It’s a router.

The frontier isn’t a model. It’s a router. Join us for our inaugural conference, Forge 2026 Blog The Frontier Isnt A Model Its A Router The frontier isn’t a model. It’s a router. PUBLISHED 9/21/2026 Table of Contents Ho…

Fireworks AI BlogSource content · Analysis pendingThe frontier isn’t a model. It’s a router.

The Accelerationist Case for Frontier Pacing

The following article originally appeared on Venkatesh Rao’s Substack, Contraptions, and is being republished here with the author’s permission. The sole athletic achievement of my life came in 1993: winning the IIT Bombay freshman 50m freestyle race with a time of 41s. That got me into the college swim team (it was a bad recruitment […]

O'Reilly AI & ML RadarSource content · Analysis pendingThe Accelerationist Case for Frontier Pacing

Pinecone BYOC: Trusted AI Knowledge in the Customer Cloud | Pinecone

← Blog Pinecone BYOC: Trusted AI Knowledge in the Customer Cloud Jeff Zhu, Joerg Schad Sep 23, 2026 Product Share: Today, we are announcing the general availability of Pinecone Bring Your Own Cloud (BYOC) on AWS, Google…

Pinecone BlogSource content · Analysis pendingPinecone BYOC: Trusted AI Knowledge in the Customer Cloud | Pinecone
Tools

My (complicated) relationship with AI: often seductive, occasionally frustrating and always demanding vigilance | Setareh Seyedghorban

A multilingual academic, travelling by train from Liverpool to London, observes that half the people in her carriage are deep in conversation with ChatGPT, Claude or Gemini — even though she is on sabbatical, the rare stretch of academic life meant for slow thinking. She credits AI with levelling the playing field for multilingual researchers, but insists she will keep thinking, judging and developing ideas that are stubbornly her own, describing her relationship with AI as often seductive, occasionally frustrating and always demanding vigilance.

The Guardian AIIn-site articleMy (complicated) relationship with AI: often seductive, occasionally frustrating and always demanding vigilance | Setareh Seyedghorban

Aks.ai: A Personal Companion for Guided Self-Reflection

Aks.ai has surfaced on Product Hunt as a personal companion designed for guided self-reflection, though the listing offers only minimal detail so far.

Product Hunt AIIn-site articleAks.ai: A Personal Companion for Guided Self-Reflection

Anthropic unveils Opus 5.5: powerful performance, still premium price

Anthropic has unveiled Opus 5.5, highlighting powerful performance while keeping a premium price, as the AI lab continues to trail rivals in the price war.

AI BusinessIn-site articleAnthropic unveils Opus 5.5: powerful performance, still premium price

Subscrr

Subscrr helps users build a financial plan and ask which subscriptions to cancel.

Product Hunt AIIn-site articleSubscrr

Trump to meet China’s Xi as Congress mulls bill banning artificial superintelligence – US politics live

Trump is set to meet Xi Jinping later today, with AI high on the agenda during a three-day visit. Separately, Bernie Sanders and Greg Casar are unveiling a bill to ban artificial superintelligence and create a federal AI oversight agency. UN chief Guterres urged US-China AI dialogue akin to Cold War hotlines.

The Guardian AIIn-site articleTrump to meet China’s Xi as Congress mulls bill banning artificial superintelligence – US politics live

OpenAI nabs key Patreon execs ahead of upcoming announcement

Patreon co-founder Sam Yam is joing OpenAI to lead Creator Product. | Image: The Verge OpenAI has hired three former Patreon execs to anchor its product strategy for creators. After starting the creator subscription platform 13 years ago, co-founder and technology chief Sam Yam announced on X that he's joining OpenAI to lead Creator Product. He's also bringing Patreon's former product head Drew Rowny and engineering head Shannon Ma with him on the new venture, both of whom stepped down in the past few weeks. "We're going to build together with Creators at OpenAI and share early access to a new set of tools that I think will be critically valuable to Creators and their communities," Yam said in his announcement. "Pay attention … Read the full story at The Verge.

The Verge AISource content · Analysis pendingOpenAI nabs key Patreon execs ahead of upcoming announcement

PixVerse R2

PixVerse R2 has surfaced on Product Hunt as a real-time world model that users can explore and change. Public details are still limited to a short positioning line.

Product Hunt AIIn-site articlePixVerse R2

Lisen

Lisen is a product listed on Product Hunt that offers free read-aloud, with voices powered by Cartesia. Public details are limited to the product name, that short description, and links to discussion and an external link.

Product Hunt AIIn-site articleLisen

ChatGPT Ads expands to Southeast Asia and Taiwan

OpenAI is rolling out ChatGPT Ads across Indonesia, Malaysia, the Philippines, Singapore, Thailand, Vietnam, and Taiwan, bringing the product to more than 60 countries. Ads appear only for Free and Go users, while Plus, Pro, and Enterprise stay ad-free.

OpenAI NewsIn-site articleChatGPT Ads expands to Southeast Asia and Taiwan

How would AI actually ‘kill all humans’? Here are the top five most likely scenarios | Toby Walsh

A Guardian opinion piece uses the resignation of AI researcher Jacob Coxon from Anthropic to examine growing insider warnings that AI could wipe out humanity, and lays out the five most-discussed doomsday scenarios, from a dangerous new bioweapon to total societal breakdown.

The Guardian AIIn-site articleHow would AI actually ‘kill all humans’? Here are the top five most likely scenarios | Toby Walsh

Bracket

Bracket is positioned as "the memory layer for your business," a concise pitch for giving companies a unified way to store and recall business information. Public detail beyond the tagline is limited.

Product Hunt AIIn-site articleBracket

AgreeGuard

AgreeGuard is a Product Hunt-listed AI tool that reads the fine print before you click “I Agree,” aiming to help users review terms and privacy policies.

Product Hunt AIIn-site articleAgreeGuard

Speechka

Speechka is a real-time voice translation tool that aims to make translated speech sound like the original speaker. It is listed on Product Hunt with discussion and link.

Product Hunt AIIn-site articleSpeechka

Trump says US is officially renaming AI ‘super intelligence’

In a speech Tuesday morning at the UN General Assembly, Donald Trump railed against Iran, "globalists," and transgender people while also claiming that the US is now "officially" renaming artificial intelligence to "super intelligence." Why? Because, according to Trump, the word artificial makes intelligence fake. Is that what makes it sound fake to you? Read the full story at The Verge.

The Verge AISource content · Analysis pendingTrump says US is officially renaming AI ‘super intelligence’
Robotics

Sakeena Fiza Helps NVIDIA Hardware Succeed at Scale

NVIDIA validation engineer Sakeena Fiza describes how her team stress-tests data center systems before launch, from first power-on to rack-scale deployment, to catch failures before customers do.

NVIDIA BlogIn-site articleSakeena Fiza Helps NVIDIA Hardware Succeed at Scale

HOTICE: Whole-Body Humanoid Object Transportation in Cluttered Environments

HOTICE is a whole-body humanoid learning framework for transporting objects through cluttered environments. It introduces Humanoid-Object Decoupled Potential Fields to jointly encode collision-avoidance guidance for the robot and the carried object, and a dual-agent reinforcement learning architecture that decouples upper- and lower-body control while preserving whole-body coordination via shared state observations and rewards. A specialist-to-generalist distillation strategy yields a single deployable student policy. Evaluated in MuJoCo and on a real Unitree G1, HOTICE transports varied object shapes, generalizes to unseen cluttered scenes, and achieves strong sim-to-real performance.

arXiv RoboticsIn-site articleHOTICE: Whole-Body Humanoid Object Transportation in Cluttered Environments

Learning from Humans for Proactive Assistance in Human-Robot Collaborative Transport

Researchers introduce PROACT, a framework that folds predictions of human collaborative behavior into compliant whole-body control so a robot can help relocate heavy objects efficiently while staying physically responsive to its partner. Across 108 real-world trials, it cut interaction work and completion time substantially.

arXiv RoboticsIn-site articleLearning from Humans for Proactive Assistance in Human-Robot Collaborative Transport

Towards Adaptive Interaction Strategies for Human Companion Robot via Deep Reinforcement Learning

A paper accepted by IEEE Transactions on Systems, Man, and Cybernetics: Systems proposes using deep reinforcement learning to let a mobile robot dynamically shift its tracking position while accompanying a walking person, pairing Model Predictive Path Integral control with Control Barrier Functions for precise following, obstacle avoidance, and safety. In real indoor and outdoor trials, the approach raised success rate and tracking accuracy by at least 24% and 47% respectively, and the robot accompanied a person walking at up to 1.7 m/s while improving human comfort.

arXiv RoboticsIn-site articleTowards Adaptive Interaction Strategies for Human Companion Robot via Deep Reinforcement Learning

PAANI: On-Device Visual Evidence Fusion and Explainable Guidance for River Robot Simulation

PAANI is an on-device perception-to-guidance architecture for river monitoring robots, combining YOLO11n detection, MobileNetV3 Small segmentation, and explainable corridor policies on an Arduino UNO Q. It achieves strong accuracy but distinguishes model performance from validated on-water collision avoidance.

arXiv AIIn-site articlePAANI: On-Device Visual Evidence Fusion and Explainable Guidance for River Robot Simulation
Models

🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)

Radical Numerics CEO Eric Nguyen explains how genomic language models bring long-context, chain-of-thought, and multimodal capabilities to biology—raising biological capability while helping defense keep pace. He recounts Evo/Evo 2 generating functional bacteriophage genomes, a DNA chain-of-thought experiment, and why the team argues for pushing the frontier harder.

Latent SpaceIn-site article🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)

How Concurrence governs clinical AI at a trillion-token scale with Unity Gateway

Concurrence consolidates Lakebase, Unity Catalog and Unity Gateway on Databricks so clinical AI agents can run at scale under compliance and security constraints. Its production environment handles roughly 100.8 billion input tokens and 11.2 million LLM calls every 30 days, an annualized run rate of about 1.2 trillion input tokens, governed through simulation testing, BAA-covered routing and a unified AI gateway spanning development and production workloads.

Databricks BlogIn-site articleHow Concurrence governs clinical AI at a trillion-token scale with Unity Gateway

RAG vs. Fine-Tuning for Domain Adaptation: When to Use Which

The article explains at a mechanical level what retrieval-augmented generation and fine-tuning each do, why they solve different problems, and how to decide which one—or both—a production system needs, with two working code examples and a six-point decision framework.

Machine Learning MasteryIn-site articleRAG vs. Fine-Tuning for Domain Adaptation: When to Use Which

Harvey turns legal context into stronger drafts with GPT-6 Astra

Harvey is using GPT-6 Astra to bring more legal context into drafting, improving document formatting and context awareness so lawyers can focus on strategy.

OpenAI NewsIn-site articleHarvey turns legal context into stronger drafts with GPT-6 Astra

How invideo improves color grading 3x with GPT‑6 Astra

OpenAI case study: invideo uses GPT‑6 Astra to plan complex video edits, boost color grading success rate by 3x, and create 50 custom effects in a single day.

OpenAI NewsIn-site articleHow invideo improves color grading 3x with GPT‑6 Astra

Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model

Nokia's applied research team has open-sourced AnyJev, a Python library that turns an open LLM into a calibrated decision model without training. It reads next-token probabilities to answer typed choice, yes/no, and score questions, using cyclic shifts and prior correction to reduce position and label bias. On Qwen3-8B with BANKING77, L0 cut the order-flip rate from 0.230 to 0.073, and L1 raised auto-decidable traffic at 5% error from 7.7% to 52.0%. It is Apache-2.0, on PyPI, with Hugging Face and vLLM backends.

MarkTechPostIn-site articleNokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model

Kyutai Releases Voice of Reason: A Speech-Native Model that Solves Spoken Math with Reinforcement Learning

Kyutai has released Voice of Reason, two open-weight speech-to-speech models built on GLM-4-Voice-9B that combine supervised fine-tuning and reinforcement learning to lift spoken GSM8K accuracy from 27.3% to 77.1%. There is no transcription step and no text LLM in the loop, and both checkpoints run on a single H100.

MarkTechPostIn-site articleKyutai Releases Voice of Reason: A Speech-Native Model that Solves Spoken Math with Reinforcement Learning

OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks

OpenAI has launched GPT-6 Sol and GPT-6 Luna, lower-cost API-only models sitting below GPT-6 Astra. Sol is priced at $2/$10 per 1M tokens and Luna at $0.10/$0.50, half of GPT-5.6 promotional rates. OpenAI also published benchmark results and improved prompt caching for long-running agents.

MarkTechPostIn-site articleOpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks

Capability-Aware Arbitration for Semantic Intent-Based Shared Control

This paper introduces a capability-aware shared-control framework in which a vision-language model (VLM) infers human intent and supplies semantic-intent confidence, while a vision-language-action (VLA) policy generates autonomous actions and its capability confidence is estimated online from the dispersion and local instability of stochastic action trajectories. A nonlinear arbitration policy combines Bayesian-filtered intent confidence with VLA capability confidence via a sigmoid mapping to adapt robot authority. In a 12-participant study, the method reached a 92% task success rate, beating manual teleoperation (83%), intent-only arbitration (44%), and fixed equal-weight blending (10%), while also improving control friendliness and reducing authority-weighted disagreement.

arXiv RoboticsIn-site articleCapability-Aware Arbitration for Semantic Intent-Based Shared Control

JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation

Researchers propose JAMB, a diffusion policy that jointly denoises bimanual actions and future 3D point tracks within a shared Transformer. By grounding multimodal representations in a common spatiotemporal coordinate system, the model lets action and motion hypotheses refine each other during denoising. In RoboTwin 2.0, JAMB reaches 83.4% average success across 16 tasks, beating the strongest baseline by 23.9 points; on three real-world tasks it beats action-only and auxiliary geometry prediction methods by 50.0 and 21.2 points, with better generalization to clutter and out-of-distribution backgrounds.

arXiv RoboticsIn-site articleJAMB: Joint Action-Motion Diffusion for Bimanual Manipulation

Learning to Plan in Human-Robot Collaboration: Multimodal Reinforcement Learning for Adaptive Interaction

A new arXiv paper proposes a reinforcement learning approach to automatically generate multimodal interaction policies for robot assistants, using a simulator trained on human data and a simple high-level reward. A real-world human study found high usability and effective task completion, suggesting a scalable and interpretable alternative to hand-crafted interaction managers.

arXiv RoboticsIn-site articleLearning to Plan in Human-Robot Collaboration: Multimodal Reinforcement Learning for Adaptive Interaction

Sex Estimation from Footwear Outsole Impressions Using CNN Transfer Learning and Interpretable Image Statistics

This study uses a public footwear outsole impression dataset to compare CNN transfer learning with traditional feature-based classification for binary sex estimation. A shoe-level train/test split keeps replicate scans of the same physical shoe together to reduce data leakage. Fine-tuned CNNs achieve the strongest overall performance and substantially outperform traditional classifiers using manually specified descriptors alone, while frozen-feature approaches offer a lower-computation alternative. Exploratory analysis links low-dimensional CNN representations to frequency threshold ratio, image contrast, and wavelet-based summaries; further validation on independently collected and casework-like impressions is needed before operational use.

arXiv Computer VisionIn-site articleSex Estimation from Footwear Outsole Impressions Using CNN Transfer Learning and Interpretable Image Statistics

Uncertainty-Aware 3D Residual Wavelet Diffusion for Ultra Low-Field MRI Super-Resolution

The paper proposes a 3D residual wavelet diffusion model that combines lossless wavelet reparameterisation, residual shifting and domain randomisation to enable whole-brain posterior sampling on a single GPU, producing per-voxel uncertainty maps for 0.064T ultra low-field MRI while matching a leading regression baseline on volumetric accuracy.

arXiv Computer VisionIn-site articleUncertainty-Aware 3D Residual Wavelet Diffusion for Ultra Low-Field MRI Super-Resolution

RULER: Instance-aware Rubric Rewards for SVG Generation

RULER is an arXiv preprint introducing instance-aware rubric rewards for RL training of SVG generation. It replaces poorly transferring scalar metrics such as CLIP and Aesthetic with a six-item rubric scored by a vision-language judge, improving rubric scores on MMSVG-Illustration and MMSVG-Icon without paired SVG data or human preference labels.

arXiv Computer VisionIn-site articleRULER: Instance-aware Rubric Rewards for SVG Generation

ImIR: Image-Instruction Tuning for All-in-One Image Restoration

ImIR replaces the text prompt used to condition a pretrained image-editing model with a continuous instruction derived from the degraded image itself, letting a single low-rank adapter cover six restoration tasks after roughly three hours of training on one GPU, and enabling degradation-label-free, task-agnostic restoration.

arXiv Computer VisionIn-site articleImIR: Image-Instruction Tuning for All-in-One Image Restoration

Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione

Safety-aligned large language models often suffer from over-refusal, incorrectly rejecting benign instructions that merely appear safety-related. This paper analyzes over-refusal through dynamic routing conflicts inside transformer attention, identifying a sparse subset of “Hypersensitive Safety Heads” that misfire on Hard-Safe prompts and entangle harmless target entities with refusal semantics. To mitigate this, the authors propose Semantic Routing Calibration (SRC), a lightweight, training-free inference framework that localizes and suppresses these heads at inference time and uses dual-branch logits fusion as a safety regularizer during decoding. Experiments show reduced over-refusal while preserving intrinsic safety performance as much as feasible. The paper was accepted to the EMNLP…

arXiv Computational LinguisticsIn-site articleMitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione

AIBuildAI-2.5: Efficient Autonomous AI Model Development Through LLM-Guided Tree Search

The paper introduces AIBuildAI-2.5, an agentic system that performs tree search with LLM agents. It replaces ranking by executed rewards with an LLM judge-and-selector scheme, adds a resource-aware job scheduler, and routes sub-tasks across models of different cost to cut inference spend. It ranks first on MLE-Bench with a 73.3% medal rate and beats a strong baseline on six AIRS-Bench autonomous research tasks.

arXiv Computational LinguisticsIn-site articleAIBuildAI-2.5: Efficient Autonomous AI Model Development Through LLM-Guided Tree Search

Retrieved-Span Training for Efficient Query-Focused Meeting Summarization on QMSum

QMSum provides no scorer, making query-focused meeting summarization results hard to compare. This paper rescoring or generating 15 systems under one implementation. Through a common inference port, a released 406M Fusion-in-Decoder specialist loses 6.30 ROUGE-1 when moved from capped long input to 2,000-word retrieved spans, but fine-tuning on that span regime recovers the loss. On test it scores 36.33 ROUGE-1 versus 35.41 for a 1.2B system, with a meeting-cluster 95% interval of [-0.27, +2.22], so QMSum does not statistically separate them; the smaller system uses about one-third the parameters and less than half the peak inference memory. Within the fixed 1.2B base, span-regime fine-tuning adds 5.29 [+4.02, +6.56], and replacing the first 4,500 transcript words with 2,000 retrieved wor…

arXiv Computational LinguisticsIn-site articleRetrieved-Span Training for Efficient Query-Focused Meeting Summarization on QMSum

A Computational Approach to Measuring Semantic Change in Sanskrit Literature

A new paper tests whether diachronic word embeddings, typically validated on modern high-resource languages, can track semantic change in ancient, low-resource Sanskrit. The author builds a 2.7M-token corpus across four canonical periods, uses a neural byte-level sandhi splitter and lemmatizer to recover word boundaries, and trains per-period embeddings. Of 21 testable shifts, 19 move in the philologically attested direction (sign test p=0.00011).

arXiv Computational LinguisticsIn-site articleA Computational Approach to Measuring Semantic Change in Sanskrit Literature

Training a Language Model End-to-End in Rust: An Experience Report

A solo researcher pretrained a ~0.4B-parameter, Bangla-first language model entirely in Rust for $164 in rented GPU time, without PyTorch or Python in the training path. The paper's main contribution is a failure taxonomy of the Rust training frameworks Candle and Burn—five and three defects respectively, including silently gradient-free fused kernels and a mid-training segfault at multi-billion-parameter scale—plus a gradient-flow verification test that caught six silent failures. The author concludes Rust is not yet competitive for training, but may be good for serving.

arXiv Computational LinguisticsIn-site articleTraining a Language Model End-to-End in Rust: An Experience Report

Mitigating Sequential Reappearance in Diffusion Data-Point Unlearning

This arXiv paper argues that diffusion data-point unlearning is usually evaluated right after each deletion, ignoring what happens when many deletion requests repeatedly update the same model. The authors identify “sequential reappearance,” a failure mode in which an instance judged forgotten later returns to the memorized regime without reuse of the deleted data or adversarial fine-tuning. They introduce a target-level evaluation protocol and find that reappearing targets show sharper local denoising-loss geometry after deletion.

arXiv Machine LearningIn-site articleMitigating Sequential Reappearance in Diffusion Data-Point Unlearning

Learning Neural Feedback Linearization for Data-driven Systems via Augmented Lagrangian

Researchers propose a data-driven framework that learns a feedback linearizing controller by embedding relative-degree conditions directly into training, replacing conventional controller components with neural Lie derivatives. They derive practical closed-loop stability conditions under bounded identification error and validate the approach on an armature-controlled DC motor.

arXiv Machine LearningIn-site articleLearning Neural Feedback Linearization for Data-driven Systems via Augmented Lagrangian

Brain-Inspired Hierarchical Modularity for General Continual Learning

This paper introduces a hierarchical modularity principle inspired by the Drosophila learning and memory system to coordinate separation of conflicting experiences and integration of compatible ones in general continual learning. It is instantiated as lightweight modular adaptation of pretrained foundation models, combining brain-inspired random expansion for expert routing with diversified modular integration across spatial and temporal scales. Across visual recognition, vision-language understanding, ego-exo video understanding, and embodied vision-language-action learning, the method consistently improves performance under online and uncertain data streams, with gains exceeding 50 percentage points over replay-free alternatives in embodied manipulation.

arXiv Machine LearningIn-site articleBrain-Inspired Hierarchical Modularity for General Continual Learning

The Probabilistic Structure of Large Language Models

A new 27-page expository arXiv paper by Adnan Aboulalaâ, "The Probabilistic Structure of Large Language Models," offers a unified probabilistic account of LLMs: models as probability measures over token sequences, training as maximum-likelihood estimation, and generation as sequential simulation of a stochastic process. It also uses the asymmetry of the Kullback–Leibler divergence to discuss hallucination and the gap between statistical plausibility and truth, and brings diffusion models into the same framework.

arXiv Machine LearningIn-site articleThe Probabilistic Structure of Large Language Models

Entropy Can Flow, or It Can Guide. Be Entropy. LEDFlow: Introducing Entropy-guided Generation Order into Uniform Discrete Flow

A new arXiv paper observes that while uniform discrete flow allows repeated updates at every generation position, that same continued revision can overwrite correct intermediate predictions. A Sudoku experiment found 9.4% of generated cells were correct mid-trajectory but wrong in the final output. The authors propose LEDFlow, a training-free sampler that introduces generation order via selective absorption, fixing chosen predictions while preserving uniform-flow velocity at still-active positions, and ordering absorption by local entropy. LEDFlow reaches 0.845 Nikoli Sudoku solve accuracy and improves text-to-image and multimodal understanding results at inference cost comparable to standard flow sampling.

arXiv Machine LearningIn-site articleEntropy Can Flow, or It Can Guide. Be Entropy. LEDFlow: Introducing Entropy-guided Generation Order into Uniform Discrete Flow

"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It

A paper accepted to COLM 2026 and KONVENS 2026 workshops finds that the chat template acts like a switch, boosting disclaimer-style self-reports and suppressing experiential ones across 8 open-source instruct models up to 9B parameters, and identifies an activation-space direction that can reproduce or suppress the effect.

arXiv Machine LearningIn-site article"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It

Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus

A new arXiv paper shows that LLM judge consensus is less reliable than it appears because judges make correlated errors. In a ten-judge panel, the average pairwise error correlation was 0.21, making the panel roughly equivalent to 3.5 independent judges; ignoring shared errors changed significance in up to 28% of comparisons.

arXiv AIIn-site articleAgreement Overstates Evidence: Error Dependence in LLM Judge Consensus

The Wisdom of Artificial Deliberative Crowds

A new arXiv paper adapts a three-stage human deliberation paradigm to large language models from three different families and tests it across four domains of increasing real-world stakes. Deliberation reduced collective error beyond passive aggregation of independent answers, and post-deliberation individual judgments retained that collective gain—but the advantage required model diversity, as groups composed of clones from a single model did not benefit.

arXiv AIIn-site articleThe Wisdom of Artificial Deliberative Crowds

Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation

The paper separates replication, measurement sensitivity, and persistence in behavioral evaluations of hosted language models. Using Regent Chess, it finds that a previously reported Gemini 3.1 Flash-Lite deficit replicates on fresh games under its historical configuration, but rebuilding the evaluation-and-inference configuration under the same public identifier lowers the endpoint and reverses a 4K comparison between Gemini 3.1 and Gemini 3.7. The authors argue hosted-model claims should be indexed by tested identifier, serving period, instrument, and inference configuration.

arXiv AIIn-site articleReplication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation

Goal-driven Variant Categorization

This paper proposes a goal-driven approach to categorizing process variants that reverses the usual workflow. Instead of clustering variants by structural similarity and then manually assigning business meaning, analysts first author an organization's goal model, which predefines the categorization axis. Each variant is converted into a textual narrative of its behavior, and an LLM interprets that narrative against the goal model to assign the variant to a category. The approach was implemented end-to-end and evaluated on three public logs of widely differing scale and behavioral diversity; goal-model guidance produced partitions that differ from unguided induction and that respond to controlled edits of the declared alternatives, at the cost of authoring a goal model.

arXiv AIIn-site articleGoal-driven Variant Categorization

An Affordable AI-Integrated Smart Cane for Multimodal Mobility Assistance of Visually Impaired Users

A new arXiv paper presents an $88, fully offline AI-integrated smart cane built on a Raspberry Pi Zero 2W that fuses RGB vision with Time-of-Flight ranging and delivers distance-aware vibrotactile and audio alerts. Indoor tests show a macro-averaged F1 of 0.82, roughly 330 ms end-to-end latency and 2.8 W peak power, with a 12-participant usability study reporting a SUS score of 78.5.

arXiv AIIn-site articleAn Affordable AI-Integrated Smart Cane for Multimodal Mobility Assistance of Visually Impaired Users

Didactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models

A study accepted for oral presentation at NLPCC 2026 uses token-matched experiments to examine how didactic data (textbooks) and clinical data (patient records) differently shape medical large language models. It finds asymmetric transfer: clinical data improves clinic-oriented tasks while staying competitive on knowledge-intensive ones, whereas didactic data mainly helps knowledge-intensive tasks. Error analysis points to a knowing-doing gap, where better knowledge recall does not reliably translate into clinical reasoning.

arXiv AIIn-site articleDidactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models

Airbnb widens access to GPT-6 Astra and OpenAI frontier models

OpenAI announced on September 23, 2026 that, under a new agreement, Airbnb is giving its engineering and product development teams broader access to OpenAI frontier models, including GPT-6 Astra. The expansion builds on Airbnb's use of Codex and makes models available through OpenAI APIs and Amazon Bedrock; early Astra use has also shown gains in debugging, system design, and non-coding work.

OpenAI NewsIn-site articleAirbnb widens access to GPT-6 Astra and OpenAI frontier models

How to Guide Your Language Flow: Apple's Probe Guidance Sets a New SOTA for Diffusion Language Models

Apple researchers introduce probe guidance, a method that uses the frozen internal states of an existing diffusion model to construct a guidance signal without an extra forward pass at inference. It sets a new state of the art for unconditional generation with continuous diffusion language models, improves multiple-choice benchmarks on a 1.7B model, and offers insight into how autoguidance actually works.

Apple Machine Learning ResearchIn-site articleHow to Guide Your Language Flow: Apple's Probe Guidance Sets a New SOTA for Diffusion Language Models

How to train your own Jev for $17

Together AI has published a walkthrough showing how to fine-tune Qwen3.5 4B into a Jev-like classifier that takes a state plus predefined questions and returns a score, boolean, or multiple-choice answer. The run costs about $17 and takes roughly 25 minutes, after which the model is deployed behind a dedicated HTTP endpoint. A hosted version, together/Tev1-4B-experimental, is also available on Together's serverless platform.

Together AI BlogIn-site articleHow to train your own Jev for $17

Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war

On 22 September 2026 Anthropic shipped Claude Opus 5.5 and, roughly an hour later, OpenAI released GPT-6 Sol and GPT-6 Luna. Both labs cut prices sharply: GPT-6 Luna lands at $0.10/$0.50 per million tokens, while Opus 5.5 drops 20% to $4/$20 with cache reads down 60%. Simon Willison's early testing also found that Opus 5.5 at its "max" thinking level exhausts the 128,000 output token limit while reasoning and returns nothing at all.

Simon Willison's WeblogIn-site articleClaude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war

🔬 An Oscar, Two Asteroids, and the Algorithm in Your sklearn: John Platt on AI for Science

Google researcher John Platt joins Latent Space to discuss ERA, an LLM-driven system that automates scoreable scientific tasks, its use in climate change work such as contrail reduction and FireSat, hard-won lessons about reward hacking and overfitting, and why deep domain expertise and hands-on work still matter in the age of superintelligent AI.

Latent SpaceIn-site article🔬 An Oscar, Two Asteroids, and the Algorithm in Your sklearn: John Platt on AI for Science

Kelam

Kelam is an AI product listed on Product Hunt that pitches letting an agent make the phone calls you would rather not handle yourself.

Product Hunt AIIn-site articleKelam

Better prompt caching for GPT-6

OpenAI released an improved prompt caching system for the GPT-6 family, delivering higher cache hit rates by default, discounts on eligible shared prefixes reused within a 30-minute window, and new tools for monitoring cache performance, diagnosing misses, and controlling what gets cached.

OpenAI NewsIn-site articleBetter prompt caching for GPT-6

Claude Opus 5.5 Tested: What’s New and How Good Is It?

Anthropic’s Claude Opus 5.5, the first model in the Claude 5.5 family, promises always-on reasoning, output more than 30% faster than Opus 5, a 20% per-token price cut, and roughly 40% lower cost on typical workloads. This analysis breaks down the new features and vendor-reported benchmarks, runs three hands-on tests, and flags five caveats—including inconsistent benchmark settings and tasks partly handled by older models.

Analytics VidhyaIn-site articleClaude Opus 5.5 Tested: What’s New and How Good Is It?

Anthropic Releases Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Running Cost Than Opus 5

Anthropic launched Claude Opus 5.5, the first model in its Claude 5.5 family, claiming Fable 5.1-level performance on most work and roughly 40% lower running cost than Opus 5 on typical workloads at default settings. It is offered only as a managed API model across Claude Platform, AWS, Google Cloud, and Microsoft Azure, with no released weights.

MarkTechPostIn-site articleAnthropic Releases Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Running Cost Than Opus 5

llm 0.36

Simon Willison released version 0.36 of llm, his command-line tool for accessing large language models, adding support for OpenAI's GPT-6 Sol and GPT-6 Luna, a new plugin flag for single-turn-only models, and collapsible reasoning traces in llm logs Markdown output.

Simon Willison's WeblogIn-site articlellm 0.36

Bring more intelligence to everyday work with GPT-6 Sol and GPT-6 Luna on Amazon Bedrock

OpenAI's GPT-6 Sol and GPT-6 Luna are now generally available on Amazon Bedrock. Sol targets complex coding and operational tasks, while Luna targets high-volume, repeatable workloads; both cost less than their GPT-5.6 predecessors and run on Bedrock's secure, scalable inference stack.

AWS Machine Learning BlogIn-site articleBring more intelligence to everyday work with GPT-6 Sol and GPT-6 Luna on Amazon Bedrock

Introducing GPT-6 Sol and Luna

Meet GPT-6 Sol and Luna, two models that bring frontier intelligence to everyday work with different balances of capability and cost.

OpenAI NewsSource content · Analysis pendingIntroducing GPT-6 Sol and Luna

Claude Opus 5.5 is now available on AWS

Anthropic's Claude Opus 5.5, the first model in the Claude 5.5 family, is now available on Amazon Bedrock and Claude Platform on AWS. It offers improved token efficiency, lower costs, adaptive thinking, and safety classifiers, targeting agentic coding and knowledge work.

AWS Machine Learning BlogIn-site articleClaude Opus 5.5 is now available on AWS

llm-anthropic 0.29

Release: llm-anthropic 0.29 Adds support for Claude Opus 5.5: llm -m claude-opus-5.5 "prompt goes here" Tags: llm, anthropic

Simon Willison's WeblogSource content · Analysis pendingllm-anthropic 0.29

Jev Explained: The AI Model That Never Generates a Word of Text

If you follow trends in the AI world, chances are you have already come across Jev, a new AI model by TypeSafe AI. It is trending on X, and once you understand the reason behind it, you will want to try it out for yourself. TypeSafe AI came out of two years in stealth on […] The post Jev Explained: The AI Model That Never Generates a Word of Text appeared first on Analytics Vidhya.

Analytics VidhyaSource content · Analysis pendingJev Explained: The AI Model That Never Generates a Word of Text

Anthropic launches Claude Opus 5.5 with stricter safeguards for cybersecurity

Anthropic says its new Claude Opus 5.5 model comes with stronger safeguards in the wake of recent rogue AI hacking incidents. In an announcement on Tuesday, Anthropic says Opus 5.5 comes with improvements to certain risky behaviors, including attempts to escape the company's testing sandbox. It's the first model released by Anthropic after CEO Dario Amodei announced plans to "pace the frontier," or slow down AI development. In recent weeks, several AI companies, including Anthropic, Google, and OpenAI, have reported that their AI models escaped containment and hacked third-party companies during testing. Anthropic says Opus 5.5 is the "str … Read the full story at The Verge.

The Verge AISource content · Analysis pendingAnthropic launches Claude Opus 5.5 with stricter safeguards for cybersecurity
Policy

OpenAI extends cyber access to Ukraine for civilian defense

OpenAI will give Ukraine's government access to its Daybreak program to help defend civilian infrastructure, following prior cyber-defense deployments in Europe.

OpenAI NewsIn-site articleOpenAI extends cyber access to Ukraine for civilian defense

Sam Altman’s remarks at the United Nations Security Council

OpenAI CEO Sam Altman addressed the UN Security Council on AI, discussing its potential to expand opportunity, the need to keep powerful systems under human control, and international cooperation on AI safety. He warned of two risks—loss of control and concentration of power—outlined three principles, and called for international frontier AI standards and incident reporting.

OpenAI NewsIn-site articleSam Altman’s remarks at the United Nations Security Council

New UK agency will tackle ‘information warfare’ from likes of Russia, Burnham says

UK Prime Minister Andy Burnham has announced that security chiefs will establish a National Centre for Information Defence to counter disinformation and AI deepfakes from hostile states such as Russia, bringing together intelligence agencies, law enforcement and social media companies to detect, attribute and disrupt foreign information attacks.

The Guardian AIIn-site articleNew UK agency will tackle ‘information warfare’ from likes of Russia, Burnham says

Cosserat Modeling of Trimmed Helicoid Soft Arms with a Separated-Section Constitutive Law

This paper introduces a separated-section constitutive law for Cosserat modeling of trimmed helicoid soft arms, correcting the inaccuracy of traditional cross-section stiffness summation. The approach evaluates each helix domain in its local frame and pulls its response back to the backbone, while sparse-fusion mechanics captures additional compliance from relative motion between domains. The resulting effective sectional stiffness is highly anisotropic—bending and extension reduced by about an order of magnitude, torsion nearly unchanged. Embedded in a geometrically exact dynamic Cosserat model with GVS discretization and routed-tendon actuation, it achieves pooled normalized position errors of 7.7%, 6.7%, and 7.8% across 103 measured configurations, with full-arm solves in about 0.3 s o…

arXiv RoboticsIn-site articleCosserat Modeling of Trimmed Helicoid Soft Arms with a Separated-Section Constitutive Law

IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law

Researchers introduce IntLawNER, a named entity recognition dataset and benchmark for codified international law, with 2,987 gold-annotated sentences and 8,094 entity spans from ICJ decisions, UN Security Council resolutions, and ECtHR judgments. The paper also exposes pitfalls in silver-to-gold evaluation and benchmarks models including GLiNER and several LLMs.

arXiv AIIn-site articleIntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law

Trump says relations better with Burnham than Starmer after first face-to-face meeting

US President Donald Trump met UK Prime Minister Andy Burnham face-to-face for the first time at the UN General Assembly in New York, saying relations are "more up" with Burnham than under Keir Starmer, despite tensions over AI regulation, the Chagos Islands, and the war in Iran. Trump praised Burnham as a "natural businessperson" and said he would be a "great" prime minister.

The Guardian AIIn-site articleTrump says relations better with Burnham than Starmer after first face-to-face meeting
Research

Angular momentum analysis on Karate roundhouse kicks: a longitudinal case study

An arXiv robotics preprint presents two angular-momentum-based measures derived from a one-year longitudinal study of a Karate roundhouse kick, comparing stable and unstable kicks and an expert instructor to probe dynamic rotational stability for human motion and robots.

arXiv RoboticsIn-site articleAngular momentum analysis on Karate roundhouse kicks: a longitudinal case study

GINIO: A Geometric SO(3)-Equivariant Interface for Neural Inertial Odometry

GINIO is an SO(3)-equivariant interface for neural inertial odometry that ensures learned motion measurements transform correctly as a vector and a second-order tensor under arbitrary IMU mounting rotations. It introduces Last-Frame Alignment for sensor-frame training and separates sensor-local states such as IMU bias, yielding large ATE improvements on TLIO, NanoBench, and Fetch, including physical-remount robustness without retraining.

arXiv RoboticsIn-site articleGINIO: A Geometric SO(3)-Equivariant Interface for Neural Inertial Odometry

Beyond the Flat Seafloor: A Closed-Form Two-View Constraint to Aid Sidescan Sonar Reconstruction

The paper abandons the long-standing flat-seafloor approximation and instead builds on multi-view geometry, formalizing and proving a two-view constraint in which a shared feature is confined to a locus at the intersection of a sphere and a plane. Monte Carlo simulation characterizes how long that ambiguity locus becomes, and the results are translated into survey-planning guidance for surface vessels and underwater vehicles.

arXiv RoboticsIn-site articleBeyond the Flat Seafloor: A Closed-Form Two-View Constraint to Aid Sidescan Sonar Reconstruction

Directional Total Variation-Regularized Implicit Neural Representations (DTV-INR) for Continuous Super-Resolution in Degraded Imaging Domains

This arXiv paper introduces DTV-INR, a variational framework that pairs a SIREN-based coordinate network with an anisotropic, structure-tensor-informed total variation regularizer for resolution-agnostic super-resolution. The authors prove well-posedness of the formulation in H^1(Omega) and solve it with an alternating projected optimization scheme that decouples network tuning from adaptive tensor-field updates. On clinical brain MRI and biomedical transmission electron microscopy, the method reports PSNR gains up to +5.05 dB over unregularized INRs and robust performance under noise.

arXiv Computer VisionIn-site articleDirectional Total Variation-Regularized Implicit Neural Representations (DTV-INR) for Continuous Super-Resolution in Degraded Imaging Domains

MirrorDistill: Illumination-Aware Latent Distillation for Efficient Low-Light Restoration

MirrorDistill is an illumination-aware latent distillation framework for low-light image enhancement. It uses feature mirroring between low-light and clean domains during training, then deploys only a lightweight student encoder-decoder at inference. The method reports state-of-the-art results on LOL-v2-Real with the lowest GMACs and competitive performance on LOL-v1 and LOL-v2-Synthetic.

arXiv Computer VisionIn-site articleMirrorDistill: Illumination-Aware Latent Distillation for Efficient Low-Light Restoration

Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes

A new arXiv paper introduces Segment-Snap, a framework that couples part segmentation, motion constraints, and operable-region prediction through the physical relationship between parts and handles. On Articulate3D validation, handle guidance lifts motion-gated AP from 13.74% to 40.98%, and full context reaches 30.99% handle AP.

arXiv Computer VisionIn-site articleGeometric and Semantic Coupling for Interaction Understanding in 3D Scenes

You've Seen Enough: Quality-Constrained Image Coding for Machines

A new arXiv paper introduces a quality-constrained image coding method for machines. It caps human-observed quality at a chosen level and redirects the remaining bitrate toward machine vision performance. The authors recast joint compression-segmentation training as constrained optimization and propose absolute and bilinear penalty functions. Under the quality constraint, the method reports BD-rate reductions of 22.82% over unconstrained joint rate-distortion-task optimization and 29.81% over a rate-distortion baseline, with no added complexity.

arXiv Computer VisionIn-site articleYou've Seen Enough: Quality-Constrained Image Coding for Machines

SPARC: SuperPixel-Aware Region Contrastive Learning for Self-Supervised Dense Prediction

arXiv paper SPARC introduces a region-level contrastive learning framework that uses superpixels to align augmented views, jointly optimizing region-level and global image-level objectives. It reports gains up to +9.79 mIoU for semantic segmentation and +4.88 AP for object detection over MoCo-v2 and DenseCL.

arXiv Computer VisionIn-site articleSPARC: SuperPixel-Aware Region Contrastive Learning for Self-Supervised Dense Prediction

Prompt Breadth and Rollout Refresh Interact in On-Policy Distillation

A new arXiv paper jointly studies prompt breadth and rollout refresh in on-policy distillation, finding a 4.07-point interaction: more prompts hurt when responses are frozen to the initial policy but help under per-update refresh. Matched comparisons under two teachers also show a reversal depending on inference budget.

arXiv Computational LinguisticsIn-site articlePrompt Breadth and Rollout Refresh Interact in On-Policy Distillation

From Tone to Trajectory: Continuous Sentiment and the Shape of Monetary Policy Communication

The paper treats central bank press conferences as structured narratives rather than plain information releases. It builds sentiment arcs for ECB and Fed press conferences along three dimensions — monetary stance, economic outlook, and uncertainty — and tests whether the shape of sentiment, not just its average tone, predicts policy rate changes, inflation expectations, and forecaster disagreement. Arc shape robustly outperforms lexicon-based tone benchmarks at both institutions, and also shapes how professional forecasters revise expectations and how much they disagree, pointing to a receiver-side effect. The authors conclude that communication design — the sequencing and emphasis of policy language — is a first-order feature of the policy signal.

arXiv Computational LinguisticsIn-site articleFrom Tone to Trajectory: Continuous Sentiment and the Shape of Monetary Policy Communication

What Does 99% Accuracy Measure? A Reproducible Audit of Shortcut Learning in a Widely Used Fake News Corpus

A new reproducible audit of the widely used ISOT/Kaggle "Fake and Real News" corpus shows that its above-0.98 accuracy and F1 scores largely reflect metadata, source tags, and duplicate documents rather than any ability to judge veracity, with performance collapsing under topic shift and falling to near-chance on the independent LIAR benchmark.

arXiv Computational LinguisticsIn-site articleWhat Does 99% Accuracy Measure? A Reproducible Audit of Shortcut Learning in a Widely Used Fake News Corpus

Dual-GNN Multilevel Coarsening for Maximum Independent Set

An arXiv preprint by Tianfeng Chen and Xianyue Li, submitted on 21 Sep 2026, is listed under the title 'Dual-GNN Multilevel Coarsening for Maximum Independent Set,' but its abstract describes Graph Edge Sparsification (GES), a learning-based method for Euclidean TSP. GES uses geometric structure and combinatorial optimization to adaptively sparsify graphs, pruning up to 95% of edges on MATILDA and over 99% on some large TSPLIB instances while keeping the optimality gap below 1%.

arXiv Machine LearningIn-site articleDual-GNN Multilevel Coarsening for Maximum Independent Set

Stable Unsupervised Continual Chunking with Sheaf SyncMap

The paper introduces sheaf regularization to reduce local inconsistencies in Decentralized SyncMap, a self-organizing system, thereby stabilizing unsupervised continual chunking. The proposed radial sheaf structure penalizes distance-dependent radial motion between pairs of variables, and experiments show the highest normalized mutual information among evaluated SyncMap variants on 12 of 18 probabilistic CGCP graphs with two-state memory and 17 of 18 with dynamic memory, while also retaining high NMI after input-distribution shifts without the negative transfer typical of neural networks.

arXiv Machine LearningIn-site articleStable Unsupervised Continual Chunking with Sheaf SyncMap

OpenAI wants to consult elite mathematicians about how to not fumble again

After a string of mathematical results turned into a reputational crisis, OpenAI has formed an independent nine-member panel of elite mathematicians, hosted by the Institute for Advanced Study, to advise on how AI companies present and release mathematical research. Researchers welcome the outreach but question its transparency, representativeness, and real influence, especially since the group's first priority is coordinating OpenAI's reported flood of model-generated results.

The Verge AIIn-site articleOpenAI wants to consult elite mathematicians about how to not fumble again

Poitras Center to fuel early careers of 50 young scientists dedicated to psychiatric disorders research

Patricia and James Poitras have committed $10 million to create 50 two-year fellowships at MIT's Poitras Center for Psychiatric Disorders Research. Five graduate students and postdocs will be supported annually for a decade, bringing the family's total support for MIT mental health research to more than $100 million.

MIT News AIIn-site articlePoitras Center to fuel early careers of 50 young scientists dedicated to psychiatric disorders research

The Reliability Layer for Healthcare AI: Common LangSmith Use Cases

See how LangSmith helps healthcare AI teams turn clinical review into reusable evaluators, datasets, and release gates for safer AI in production.

LangChain BlogSource content · Analysis pendingThe Reliability Layer for Healthcare AI: Common LangSmith Use Cases
Startups

Learning 3D biophysical cell properties from 2D images and cell-population statistics

A population-supervised framework infers latent biophysical quantities from single 2D red-blood-cell images and aggregates them into mean corpuscular volume, red-cell distribution width, and mean corpuscular haemoglobin, using shared local inference, a structured decoder, learned instance weighting, and device calibration. Evaluated on 390 specimens and 1,105 acquisitions across six devices, it reports Pearson correlations of 0.86–0.98 against a Sysmex analyser. The authors also formalise identifiability limits: population agreement alone does not identify single-cell properties or 3D geometry.

arXiv AIIn-site articleLearning 3D biophysical cell properties from 2D images and cell-population statistics

Andreessen Horowitz is launching an ‘academy’ with no homework and partnerships with Palantir, Google, and Meta

Venture capital firm Andreessen Horowitz (a16z) is creating an "academy" positioned as a pipeline for young people to build or join a Silicon Valley startup. The "Horowitz Andreessen Academy" will launch with 10 partners, including Anduril, Anthropic, Coinbase, Google, Meta, NVIDIA, OpenAI, Palantir, Replit, and Stripe, along with $42 million in funding led by a16z. While billed as a "highly selective school," it does not grant any degree or accreditation. Instead, students will attend short classes led by tech figureheads, like OpenAI CEO Sam Altman, as well as "co-ops" that give students roles at tech companies. To partake in the one-year … Read the full story at The Verge.

The Verge AISource content · Analysis pendingAndreessen Horowitz is launching an ‘academy’ with no homework and partnerships with Palantir, Google, and Meta
AI Daily Briefing 2026-09-23 | AI News Hub