Skip to content
AI News HubLIVE

Reports for this edition have been collected; translation and analysis are pending. Expand other updates to read source content.

Other updates (132)
Agents

The Interfaces Are Arriving

The most consequential AI news of the past year came from a standards body. In December 2025, Anthropic donated the Model Context Protocol to the newly formed Agentic AI Foundation, a directed fund under the Linux Foundation cofounded by Anthropic, Block, and OpenAI, with support from Google, Microsoft, AWS, Cloudflare, and Bloomberg. Six months earlier, […]

O'Reilly AI & ML RadarSource content · Analysis pendingThe Interfaces Are Arriving

Weave Router 2.0

Discussion | Link

Product Hunt AISource content · Analysis pendingWeave Router 2.0

Don't sleep on wrapture

Graham Dumpleton's new monkey patching package wrapture is shaping up to be an indispensable tool for Python developers. I'm not sure why I've seen so little buzz about it! Graham has been posting new tutorials for it almost daily since the initial release on August 31st. Here's everything he's published so far: Introducing wrapture - a new monkey patching library that serves both testing and observability (think New Relic style tracing) at the same time. Unit testing with wrapture - how to use it for the same kinds of thing as unittest.mock. Recording calls with wrapture - recording method calls as timelines and processing and displaying them as trees. Phased behaviour in wrapture - arranging patched methods to change behavior across multiple calls. Beyond callables in wrapture - monkey…

Simon Willison's WeblogSource content · Analysis pendingDon't sleep on wrapture

Open-Source AI & Open Models Reading List

How to get up to speed on open models and their implications.

Interconnects (Nathan Lambert)Source content · Analysis pendingOpen-Source AI & Open Models Reading List

Fine-Tuning Agentic AI: A Practical Guide

In this article, you will learn how to fine-tune an agentic AI system holistically, covering all four critical dials: training data, parameter-efficient fine-tuning, runtime hyperparameters,...

Machine Learning MasterySource content · Analysis pendingFine-Tuning Agentic AI: A Practical Guide

Operating Mode as Runtime State: A Contract for Enterprise

During a service incident, a customer-remediation workflow is moved onto an emergency route because the situation is critical and the team needs a fast resolution. Approvals are shortened, a priority queue is opened, and an on-call agent is cleared to use an alternate procedure until the service recovers. The incident ends, but the route stays […]

O'Reilly AI & ML RadarSource content · Analysis pendingOperating Mode as Runtime State: A Contract for Enterprise

Rapidly scaling online storage to serve over 1 billion ChatGPT users

Learn how OpenAI evolved Habitat from a Python library into a globally distributed storage platform serving 1 billion ChatGPT users and 22M requests per second.

OpenAI NewsSource content · Analysis pendingRapidly scaling online storage to serve over 1 billion ChatGPT users

How to Build the Unified Data Foundation Drug Discovery AI Depends On

This article is sponsored by CDD Vault and was written, edited, and published in alignment with our Emerj sponsored content guidelines. Learn more about our thought leadership and content creation services on our Emerj Media Services page.​ Drug discovery is one of the slowest, costliest processes in enterprise R&D. Developing a single FDA-approved therapy typically […]

Emerj AI ResearchSource content · Analysis pendingHow to Build the Unified Data Foundation Drug Discovery AI Depends On

Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

Sakana AI has released Fugu Max and Fugu Ultra v2, 2 models built on the same learned orchestration architecture. Fugu Max routes tasks to lean open and specialized models, including NVIDIA Nemotron, at $2/$6 per 1M tokens. Fugu Ultra v2 targets peak capability, scoring 48.3 on Chartography and 74.3 on DeepSWE. The post Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration appeared first on MarkTechPost.

MarkTechPostSource content · Analysis pendingSakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

LTLDiff: Finite Linear Temporal Logic-Guided Data Generation and Diffusion Policies for Multi-agent Robotic Manipulation

arXiv:2609.11043v1 Announce Type: new Abstract: Multi-agent robotic manipulation tasks require coordination among agents to satisfy task-level temporal, logical, and safety constraints. Recently, diffusion policies have been used to perform the task. However, they still suffer from desynchronization, incorrect action ordering, and coordination failures in tasks that require simultaneous or sequential multi-agent interaction. Therefore, LTLDiff is proposed as a framework that combines Finite Linear Temporal Logic (LTLf) specification learning for both the generation of demonstrations and learning via diffusion policies. Each task has a specific LTLf formula that is learned from a set of natural language instructions using a large-scale language model. To enable a fixed-dimensional vector e…

arXiv RoboticsSource content · Analysis pendingLTLDiff: Finite Linear Temporal Logic-Guided Data Generation and Diffusion Policies for Multi-agent Robotic Manipulation

Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations

arXiv:2609.09448v1 Announce Type: new Abstract: As agentic systems getting adopted rapidly in safety critical applications, it is vital to measure the confidence associated with the agentic actions. In comparison to the traditional machine learning systems, agentic workflows have complex failure modes with planning, tool invocation and dynamic environment interactions. In this paper, we investigate whether model's internal representations provide stronger signals of eventual task success in multi-turn agentic setups. We introduce two complementary methods: Latent Trajectory Dynamics (LTD), which summarizes changes in residual-stream representations across an an interaction trajectory, and the Action Representation Probe (ARP), which predicts success from representations formed at action d…

arXiv AISource content · Analysis pendingDo Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations

Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration

arXiv:2609.09418v1 Announce Type: new Abstract: World Action Models (WAMs) couple predictive world modeling with action generation, allowing anticipated future states to guide agent behavior. Although WAMs are rapidly advancing embodied AI, general-purpose counterparts remain largely unexplored in games. Existing game-oriented approaches often combine action-conditioned world models with external policies and reward functions to realize WAM-like decision-making, yet they operate mainly in 2D visual observation space and do not instantiate persistent 3D geometry. Extending this paradigm to 3D games introduces a distinct challenge. In autonomous driving and robotics, the physical environment exists independently of the model, providing a persistent 3D world in which selected actions can be…

arXiv AISource content · Analysis pendingValerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration

Decision-Focused Active Learning for Scale-Aware Critical-Materials Recovery

arXiv:2609.09413v1 Announce Type: new Abstract: Choosing a recovery process for scale-up requires connecting laboratory results with product requirements, process costs, and scale effects. We analyze records from Pacific Northwest National Laboratory's Computer Intelligence for Critical Element Recovery and Optimization (CICERO) workflow for autonomous selective precipitation. Active learning uses prior results to choose experiments. In a conditional retrospective benchmark with fitted models and recycled neodymium-iron-boron (NdFeB) magnet records, active learning finds the best recorded result with fewer experiments than nonadaptive space filling. Enrichment is the selected rare-earth-to-iron ratio relative to that in the feed. Adaptive policies reach the recorded enrichment maximum by…

arXiv AISource content · Analysis pendingDecision-Focused Active Learning for Scale-Aware Critical-Materials Recovery

The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents

arXiv:2609.09395v1 Announce Type: new Abstract: Language models act through tools, yet practical agents face libraries containing thousands of interfaces. We introduce the tool menu as the short, ordered subset of available tools shown to an agent before execution. The agent can call only tools in this menu. Multi-step tasks require the final action and the prerequisite tools that create its inputs in a usable order. Current constructors rank tools by request relevance, which can surface the final action while omitting or delaying less obvious producers. We introduce the state path, a pre-execution route from the observable request state to the desired outcome, and propose State-Path Tool Menu to learn it. Our framework treats the menu as an execution prior over these routes. Its encoder…

arXiv AISource content · Analysis pendingThe Menu Is an Execution Prior: State-Path Tool Menus for Online Agents

An Autonomous GeoAI Agent for Arctic Eco-Navigation

arXiv:2609.09374v1 Announce Type: new Abstract: Arctic maritime navigation is becoming increasingly important as changing sea-ice conditions expand seasonal accessibility while simultaneously introducing substantial operational, environmental, and community risks. Arctic route planning is inherently a multi-criteria problem: routes that improve vessel safety or efficiency may increase exposure to sea ice, sensitive ecosystems, or nearby communities. Existing routing methods prioritize travel time, fuel use, and navigational risk, often overlooking ecological and community impacts. We introduce a human-in-the-loop, multi-agent GeoAI system for Arctic eco-navigation that integrates operational, physical, ecological, and community-related criteria within a unified routing framework. Multiple…

arXiv AISource content · Analysis pendingAn Autonomous GeoAI Agent for Arctic Eco-Navigation

Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks

arXiv:2609.09233v1 Announce Type: new Abstract: How can language model agents effectively leverage libraries of reusable knowledge to solve long-horizon tasks? Recent work has increasingly focused on agent skills: reusable capabilities represented as skill packages, i.e., multi-file bundles containing instructions, scripts, and other resources that help agents perform specific tasks. Agent skills are typically executed by loading their skill instructions into an agent's context and relying on the agent to follow them. As task horizons grow, however, this approach becomes increasingly brittle, because reasoning quality degrades as more information accumulates in the context window. We investigate an alternative approach in which skill packages are instead invoked as subagents. Rather than…

arXiv AISource content · Analysis pendingSubagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks

Adaptive Entangled Game Modules in Artificial General Intelligence

arXiv:2609.09226v1 Announce Type: new Abstract: We introduce a probability-wave framework for modeling the collective behavior of interacting adaptive agents, deriving testable eigenmodes through a generalized behavioral intelligence (GBI) nonlocal probability-wave equation. This framework captures a broad range of human intelligence behaviors with analytical mechanisms and offers an indirect method to examine the Liu-Chen-Ao (LCA) hypothesis of nonlocal entangled nerve fibers in the brain through collective trader behaviors. Our empirical analysis of Chinese intraday stock market data demonstrates that adaptive entangled game modes explain 82-94% (89% overall) of observed decision patterns, a sharp contrast to the predictions of neoclassical finance based on independent rational agents.…

arXiv AISource content · Analysis pendingAdaptive Entangled Game Modules in Artificial General Intelligence

OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

arXiv:2609.09203v1 Announce Type: new Abstract: Existing benchmarks for autonomous AI scientists evaluate only final outputs---generated code, hypotheses, or papers---yet discard the reasoning process by which those outputs were obtained. This makes it impossible to audit scientific methodology, diagnose failure modes, or distinguish systematic reasoning from fortunate guessing. We present \textbf{OpenDiscoveryTrace}, a public dataset of 558 complete AI scientific agent trajectories that captures how models reason, not just what they produce. Each trajectory records a structured 9-field-per-step trace---including thoughts, tool calls, observations, errors, revision triggers, and self-reported confidence---as models execute 124 scientific tasks spanning drug discovery, materials science, g…

arXiv AISource content · Analysis pendingOpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

Any Nix package, live in your browser

Any Nix package, live in your browser Farid Zakaria calls this his "magnum opus of Nix work", and I can see why. trynix.dev provides a qemu-wasm powered x86_64 Linux virtual machine running entirely in your browser through WebAssembly. That VM can then be booted with any Nix package from the past 13 years. They are URL addressable, so you can navigate to this page: https://trynix.dev/?pkg=python3%403.6.2 Then click "Load" and get an interactive shell against a virtual machine running Python 3.6.2 from 2017. Farid is building all sorts of neat things on top of this. One recent example: Review a pull request by booting it introduces trynix-preview, described like this: GitHub action that comments a link on a pull request which lets you boot the PR’s build in the browser using https://trynix…

Simon Willison's WeblogSource content · Analysis pendingAny Nix package, live in your browser

Cadenya

Discussion | Link

Product Hunt AISource content · Analysis pendingCadenya

AWS open-sources Pizza Bot: email-style inbox for background AI agents

Amazon Web Services (AWS) has released a new open-source application dubbed Pizza Bot, which gives developers an email-style inbox for The post AWS open-sources Pizza Bot: email-style inbox for background AI agents appeared first on The New Stack.

The New Stack AISource content · Analysis pendingAWS open-sources Pizza Bot: email-style inbox for background AI agents

OzBrain

Discussion | Link

Product Hunt AISource content · Analysis pendingOzBrain

GitHub Copilot app for Beginners: Using the diff, terminal, and browser

Checking agent-generated code usually means hopping between tabs. Learn how to view diffs, run terminal commands, and preview web apps side by side in the GitHub Copilot app. The post GitHub Copilot app for Beginners: Using the diff, terminal, and browser appeared first on The GitHub Blog.

GitHub AI & MLSource content · Analysis pendingGitHub Copilot app for Beginners: Using the diff, terminal, and browser

OpenAI Launches the Agents API in Public Beta, Putting the Codex Harness Behind One API Call

OpenAI has released the Agents API in public beta. It gives developers the same harness and infrastructure that run Codex. OpenAI hosts and maintains the harness. Developers run the agent’s compute in an OpenAI-managed sandbox, their own infrastructure, or a partner sandbox. Is it deployable? Yes. It is live for all developers in public beta. […] The post OpenAI Launches the Agents API in Public Beta, Putting the Codex Harness Behind One API Call appeared first on MarkTechPost.

MarkTechPostSource content · Analysis pendingOpenAI Launches the Agents API in Public Beta, Putting the Codex Harness Behind One API Call

Native is now the future of mobile at Shopify

Native is now the future of mobile at Shopify Shopify are moving from React Native back to separate Swift and Kotlin codebases for their native apps, for the exact reason you would expect: We decided to switch from native to React Native in 2020 for three reasons: Stop building the same features twice Allow developers to work across the stack Spend less time chasing feature parity and more time shipping value [...] Native still means building and maintaining software on two platforms, that cost has not disappeared. What changed is that agents can now do enough of the implementation, translation, testing, and review work that it’s no longer the deciding factor it was in 2020. It's a well-written post, which gives full credit to React Native as a great platform for the six years they were u…

Simon Willison's WeblogSource content · Analysis pendingNative is now the future of mobile at Shopify

“Six tools, one harness”: Salesforce loops together a six-pack of favorites

Salesforce introduced its Salesforce Enterprise AI Harness on Thursday as a formalized amalgamation of the AI harness concepts and infrastructure The post “Six tools, one harness”: Salesforce loops together a six-pack of favorites appeared first on The New Stack.

The New Stack AISource content · Analysis pending“Six tools, one harness”: Salesforce loops together a six-pack of favorites

Amazon Quick is now generally available on desktop

Your teams get an AI assistant that handles real work while your data stays in your environment and your conversations stay private Today, the Amazon Quick desktop application is generally available on macOS and Windows. We’re also adding a new activity feed to the mobile experience on iOS and Android that consolidates email, calendar, CRM, […]

AWS Machine Learning BlogSource content · Analysis pendingAmazon Quick is now generally available on desktop

A Candid Abacus AI Review: The All-in-One AI Platform for Professionals & Enterprises

If you’re paying for ChatGPT, Claude, and another AI tool simultaneously, this review is for you. It covers what an AI platform like Abacus AI actually includes, how the credit system works in practice, and whether it genuinely replaces your current stack or just adds to it.

KDnuggetsSource content · Analysis pendingA Candid Abacus AI Review: The All-in-One AI Platform for Professionals & Enterprises

Build an end-to-end RFI questionnaire workflow using Amazon Quick Automate

Learn how to build an end-to-end RFI questionnaire workflow with Amazon Quick Automate. Read a multi-tab RFI workbook from Amazon S3, use natural-language prompts to extract and structure the questionnaire data, refine the workflow through conversation, and write clean CSV output back to Amazon S3 — cutting development from days to hours.

AWS Machine Learning BlogSource content · Analysis pendingBuild an end-to-end RFI questionnaire workflow using Amazon Quick Automate

When Content Is Free, Trust Is the Product

There is more technical content available today than any human being could read in a thousand lifetimes. Every topic has a dozen YouTube videos, three Substack posts, a GitHub repo, and a Reddit thread, most created in the last six months and, in many cases, technically accurate. And yet most of the professionals I talk […]

O'Reilly AI & ML RadarSource content · Analysis pendingWhen Content Is Free, Trust Is the Product

Physical AI Takes the Wheel: How the World’s Robotaxi Leaders Are Building With NVIDIA Technologies

The global robotaxi market — physical AI’s first commercial breakthrough — is projected to reach $400 billion by 2035, with over 6 million commercial vehicles in operation as driverless fleets are already moving people through some of the world’s busiest and most complex streets. Deploying a driverless vehicle is one challenge. Scaling a fleet is […]

NVIDIA BlogSource content · Analysis pendingPhysical AI Takes the Wheel: How the World’s Robotaxi Leaders Are Building With NVIDIA Technologies

Introducing Projects · Cursor

Blog / product Today we're launching Projects in Cursor. Projects lets you take on larger bodies of work, such as a feature, a migration, or a full app. It maintains context over months of work, delegates tasks to thous…

Cursor BlogSource content · Analysis pendingIntroducing Projects · Cursor

Blaxel is joining Baseten to build the future of agentic infrastructure

News Blaxel is joining Baseten to build the future of agentic infrastructure Baseten has acquired Blaxel Authors Amir Haghighat Tuhin Srivastava Paul Sinai Last updated September 10, 2026 Share Today, Blaxel is joining…

Baseten BlogSource content · Analysis pendingBlaxel is joining Baseten to build the future of agentic infrastructure

Gen-1 Slides: Opus 5-level decks at a fraction of the cost

Gen-1 Slides: Opus 5-level decks at a fraction of the cost Join us for our inaugural conference, Forge 2026 Blog Gen 1 Slides Opus 5 Level Decks At A Fraction Of The Cost Gen-1 Slides: Opus 5-level decks at a fraction o…

Fireworks AI BlogSource content · Analysis pendingGen-1 Slides: Opus 5-level decks at a fraction of the cost
Models

Mistral Bets Enterprise AI Will Be About Control, Not Just Intelligence

The French AI lab is using a $3B fundraise to sell control over AI infrastructure, not just model power -- a shift in direction that could matter to U.S. firms in Europe too.

AI BusinessSource content · Analysis pendingMistral Bets Enterprise AI Will Be About Control, Not Just Intelligence

Soft-deprecating re.match()

Soft-deprecating re.match() Python has a concept of soft deprecation, where APIs are marked as "should no longer be used to write new code" without any promise/threat to remove them in the future. Python 3.15 release manager Hugo van Kemenade describes how in the upcoming 3.15 release soft deprecation has come for the venerable but deeply confusing re.match() function. It's now available with the much clearer alternative re.prefixmatch() name - reflecting how it anchors at the beginning of the string but not the end. Most of the time you probably want re.search() (match this pattern anywhere in the string) or re.fullmatch() (match the entire string) instead. Via Lobste.rs Tags: python, regular-expressions

Simon Willison's WeblogSource content · Analysis pendingSoft-deprecating re.match()

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

Cohere has released North Small Translate, an open-weight Mixture-of-Experts model built for machine translation across 50 languages. It uses 25B of its 218B parameters per token and scores 83.6 on Cohere's WMT26 evaluation. Weights are free for non-commercial use, with commercial access through Cohere Model Vault or RWS Language Weaver. The post Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages appeared first on MarkTechPost.

MarkTechPostSource content · Analysis pendingCohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation

Google Research has released ToolGrad, an ACL 2026 Findings framework that inverts tool-use dataset generation: it builds a verified API chain first, then writes the matching user query. Guided by textual "gradients" from a 4-module propose-execute-select-update loop, ToolGrad reaches a 99.8% pass rate on ToolBench versus 63.8% for DFS search. Gemma-3-12B fine-tuned on only 500 samples scores 83.1 on BFCL, next to Gemini 2.5 Pro at 83.2. Code, dataset, and models are public under Apache-2.0. The post Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation appeared first on MarkTechPost.

MarkTechPostSource content · Analysis pendingGoogle Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation

ObstaDiff: Generalizable Diffusion Policy Learning via Obstacle-aware Representations

arXiv:2609.10918v1 Announce Type: new Abstract: Imitation learning has achieved impressive results in robotic manipulation, yet most existing approaches assume clean backgrounds and lack explicit mechanisms for obstacle-aware motion generation. Extending such policies to cluttered, real-world scenes with unstructured obstacles remains a key generalization challenge. We present ObstaDiff, a decomposed diffusion-policy framework with a lightweight obstacle-aware visual encoder. ObstaDiff extracts a structured target-obstacle-background representation, enabling the downstream alignment policy to generate end-effector trajectories toward a target-centered bottleneck pose while reasoning about surrounding obstacles. We evaluate ObstaDiff on 61 real-robot greenhouse trials per method (366 execu…

arXiv RoboticsSource content · Analysis pendingObstaDiff: Generalizable Diffusion Policy Learning via Obstacle-aware Representations

IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies

arXiv:2609.10915v1 Announce Type: new Abstract: Vision-language-action (VLA) policies leverage pretrained vision-language backbones to achieve strong cross-task generalization. A leading design couples this backbone with a dedicated continuous action head trained via diffusion or flow matching. However, such heads rely on iterative multi-step sampling, for example 10 Euler steps in $\pi_{0.5}$. This creates an inference bottleneck that produces stop-and-go movement in the robot and slower task completion. We introduce IMLE-VLA, which replaces the iterative action head with a single-step conditional generator trained via conditional Implicit Maximum Likelihood Estimation (cIMLE). The cIMLE objective promotes multimodal action coverage, avoiding the mode collapse of naive regression heads w…

arXiv RoboticsSource content · Analysis pendingIMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies

ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

arXiv:2609.10895v1 Announce Type: new Abstract: Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as the decision coreof household robots. Existing evaluations, however, probe intuitive physics passively through question answering over videos, or target deliberate, long-horizon tasks such as navigation and rearrangement; none measure whether a model can turn physical understanding into immediate, safety-critical action. We introduce ReactHuman, the first physics-grounded benchmark for human-like reactive decision-making, in which the evaluated MLLM acts as the brain of a simulated humanoid facing sudden household hazards; i…

arXiv RoboticsSource content · Analysis pendingReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

HuRo: Robotizing Human Videos for Scalable VLA Pretraining

arXiv:2609.10706v1 Announce Type: new Abstract: Human video datasets have emerged as a compelling alternative to expensive real-robot data, offering rich diversity at scale. To bridge the human-to-robot embodiment gap, existing approaches either robotize videos in task-matched settings or address observation and action alignment separately at scale. In this work, we systematically examine whether robotized human videos can provide effective and scalable supervision for pretraining vision-language-action (VLA) policies. To this end, we develop a robotization pipeline that converts heterogeneous human videos into robot-aligned observations and action trajectories while inferring missing intermediate signals across annotation levels. Using this pipeline, we construct the HuRo dataset, compri…

arXiv RoboticsSource content · Analysis pendingHuRo: Robotizing Human Videos for Scalable VLA Pretraining

Overpainting: Localized Context-aware Diffusion Image Editing

arXiv:2609.10811v1 Announce Type: new Abstract: We present "overpainting", an image editing operation which offers both control over the location of the edit and awareness of the previous content in that location. The overpainted area is given by a trimap, where white-annotated pixels must be edited, gray-annotated pixels may be edited, and black-annotated pixels must not be edited. This enables both precise and loose control, depending on user intent. We implement overpainting by adapting a pretrained image editing diffusion model using a combination of joint attention and low-rank adaption across input images with attention-dropout to balance the information flow between noise, source and mask images. We present a novel, automated, training data generation pipeline that (1) generates a…

arXiv Computer VisionSource content · Analysis pendingOverpainting: Localized Context-aware Diffusion Image Editing

TrajFusionNet+: Transformer-Based Prediction of Pedestrian Crossing Intention via Fusion of Trajectory Representations and Scene Graphs

arXiv:2609.10806v1 Announce Type: new Abstract: The pedestrian crossing intention task involves predicting whether pedestrians are likely to cross the road from the point of view of an autonomous vehicle. We introduce TrajFusionNet+, a novel transformer-based model for pedestrian crossing intention prediction. TrajFusionNet+ combines sequential and visual representations of pedestrian trajectory with a graph-based representation of the scene context in order to predict pedestrian crossing intention. The proposed architecture builds upon our previous model, TrajFusionNet, and comprises three branches: a Sequence Attention Module (SAM), which processes a sequential representation of past and predicted pedestrian trajectories; a Visual Attention Module (VAM), which utilizes a visual represen…

arXiv Computer VisionSource content · Analysis pendingTrajFusionNet+: Transformer-Based Prediction of Pedestrian Crossing Intention via Fusion of Trajectory Representations and Scene Graphs

GRADE: Single-Frame Generative Radar Depth Estimation Under Visual Degradation

arXiv:2609.10756v1 Announce Type: new Abstract: Dense 3D depth perception fails under smoke, fog, and darkness because optical sensors cannot penetrate airborne particulates. mmWave radar remains usable and measures range accurately under these conditions, but its small aperture limits angular resolution. We present GRADE, which grounds a pretrained generative prior in single-frame radar geometry to estimate high-fidelity metric depth. GRADE first maps raw 4D radar spectra to coarse metric depth. A latent diffusion backbone then recovers structural detail while conditioning every denoising step on this estimate. A pixel-space adapter uses residual camera cues when available and is trained across clear, smoke-degraded, and occluded inputs so the full output approaches the radar-conditioned…

arXiv Computer VisionSource content · Analysis pendingGRADE: Single-Frame Generative Radar Depth Estimation Under Visual Degradation

Meta-Learning for Data-Efficient Plant Growth Estimation via Vision Transformers and Fuzzy Clustering

arXiv:2609.10749v1 Announce Type: new Abstract: Accurate plant growth estimation is essential for greenhouse monitoring, yet obtaining labeled data remains costly and time-consuming. To address this, we propose a few-shot regression framework that combines Vision Transformer (ViT) feature embeddings, clustering-based task construction, and gradient-based meta-learning, and show that task construction in embedding space is a primary driver of performance. The approach leverages an unlabeled image pool to organize data into structured tasks using fuzzy c-means clustering, enabling efficient learning from a small number of labeled samples. We systematically evaluate meta-learning methods and show that second-order methods (e.g., Model-Agnostic Meta-Learning variants such as MAML++) outperfor…

arXiv Computer VisionSource content · Analysis pendingMeta-Learning for Data-Efficient Plant Growth Estimation via Vision Transformers and Fuzzy Clustering

MHE-Former: Multi-Hypothesis Transformers via Entropy Maximization for 3D Mesh Recovery

arXiv:2609.10743v1 Announce Type: new Abstract: Monocular 3D hand and body mesh recovery often suffers from severe occlusion and ambiguity. Traditional deterministic methods typically regress a single optimal solution, leading to overconfident predictions. In this paper, we introduce an exploration--exploitation paradigm for ambiguous mesh recovery with multi-hypothesis learning and selection. Specifically, during exploration, based on our probabilistic formulation and entropy maximization, we propose a novel multi-hypothesis method referred to as MHE-Former. It is a Transformer-based multi-hypothesis framework, ensuring high training efficiency and label friendliness while generating plausible and diverse hypotheses. During exploitation, we propose Hypothesis Selection, a context-aware p…

arXiv Computer VisionSource content · Analysis pendingMHE-Former: Multi-Hypothesis Transformers via Entropy Maximization for 3D Mesh Recovery

AcFlow: Controlling Text-to-Image Diffusion Transformers via Learned Conditional Activation Flow

arXiv:2609.10723v1 Announce Type: new Abstract: Text-to-image diffusion transformers (DiTs) are powerful generators, yet direct prompting provides limited control interface for style intensity and can fail to suppress unwanted concepts. To enable these controls, we introduce AcFlow, an inference-time controller that transports intermediate layer image-token activations through a learned concept-conditioned velocity field while keeping the base DiT frozen. A textual concept description specifies the desired intervention, while the integration horizon provides a continuous control parameter. The field produces token-varying, activation-dependent updates. With parameters shared across concepts within each task family, the field supports fine-grained descriptions and generalizes to concepts u…

arXiv Computer VisionSource content · Analysis pendingAcFlow: Controlling Text-to-Image Diffusion Transformers via Learned Conditional Activation Flow

LLM-Anchored Paralinguistic Enrichment for Alzheimer's Disease Detection

arXiv:2609.10896v1 Announce Type: new Abstract: Speech-based automatic detection of Alzheimer's disease (AD) provides a non-invasive and scalable approach to early cognitive screening. AD affects both lexical-semantic organization and speech production, including atypical pauses and word elongations. However, existing methods have yet to fully integrate these paralinguistic cues with linguistic content. We propose LLM-Anchored Paralinguistic Enrichment (LAPE), which enriches LLM-derived linguistic representations with paralinguistic cues through three coordinated innovations. The first is prosodic event textualization, which enables the LLM to model pauses and elongations jointly with lexical content by encoding them as explicit markers with bounded duration-aware repetition. The second i…

arXiv Computational LinguisticsSource content · Analysis pendingLLM-Anchored Paralinguistic Enrichment for Alzheimer's Disease Detection

Does Linguistic Structure Enrichment Enhance Coherence Assessment? Not With Current Architectures

arXiv:2609.10893v1 Announce Type: new Abstract: Recent advances in large language models have transformed human-computer interaction. Despite their fluency, these models often produce texts that are grammatically correct but semantically incoherent, containing contradictions or disruptions in logical flow. This work investigates whether enriching text with syntactic and rhetorical information can improve incoherence prediction. Our experiments and analysis show that plain texts achieved higher accuracy because the added information was structurally and syntactically incompatible with the language model's architecture. Additionally, to demonstrate the practical importance of coherence assessment, we performed zero-shot experiments on a Brazilian disinformation dataset, suggesting that text…

arXiv Computational LinguisticsSource content · Analysis pendingDoes Linguistic Structure Enrichment Enhance Coherence Assessment? Not With Current Architectures

Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models

arXiv:2609.10830v1 Announce Type: new Abstract: When a language model finds a sentence unusually cheap to predict, it is tempting to conclude that the sentence was in its training data. Almost every published test of that inference has had to guess which sentences were in the training data, the members, and which were not. This paper removes the guessing. Two model families, OLMo-2 and Pythia, publish their pretraining corpora, and a public index over those corpora returns the exact number of times any sentence appeared in each. Those counts make three questions answerable directly. The answers form a pincer, closing from two sides. At the duplication levels ordinary text actually has, five models from 1B to 13B parameters carry at most a faint trace of their own exposure. We measure that…

arXiv Computational LinguisticsSource content · Analysis pendingDetectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models

Larger Context Window, Fewer Overcorrections: Optimizing Prompts and Batching for Minimal-Edit Grammatical Error Correction

arXiv:2609.10810v1 Announce Type: new Abstract: Minimal-edit Grammatical Error Correction (GEC) is a challenging task for zero- and few-shot prompted Large Language Models (LLMs), which systematically overcorrect and degrade $F_{0.5}$ by rewriting well-formed spans. While fine-tuning provides an effective solution, it imposes substantial infrastructure demands. We introduce a prompt-based approach that closes the gap to fine-tuned models through three advances in GEC prompting methodology. First, we introduce taxonomy-based instructions to enforce minimal-edit constraints with a comprehensive list of grammatical error rules, equipping the LLM with a bounded, metric-aligned scope of correctable edits, which benefits the strongest models while remaining model-dependent overall. Second, we s…

arXiv Computational LinguisticsSource content · Analysis pendingLarger Context Window, Fewer Overcorrections: Optimizing Prompts and Batching for Minimal-Edit Grammatical Error Correction

Analyzing Traditional and Neural Approaches to Multilingual Readability Assessment

arXiv:2609.10792v1 Announce Type: new Abstract: Transformer-based models excel at Automatic Readability Assessment (ARA), yet feature-based models remain in active use because their predictions tie back to linguistic properties. This matters because readability labels are subjective and rater-dependent, so high accuracy on noisy ground truth may reflect surface patterns rather than the linguistic structure that defines difficulty. We test whether transformers internalize the same features as traditional models across Arabic, English, French, Hindi, and Russian using the ReadMe++ dataset. Shapley Additive Explanations (SHAP) identify the features driving traditional classifiers, which we then use as TCAV concept sets to probe multilingual XLM-R and language-specific encoders. Transformers…

arXiv Computational LinguisticsSource content · Analysis pendingAnalyzing Traditional and Neural Approaches to Multilingual Readability Assessment

Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

arXiv:2609.10758v1 Announce Type: new Abstract: Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu language as a representative low-resource language. We generate Urdu-Stories, a corpus of 93 stories generated using three contemporary LLMs (GPT-5.1, Qwen-3-Max, DeepSeek-3.1). We manually annotate the errors present in them under a nine-label linguistic, semantic, and cultural taxonomy. Our notable findings suggest that LLMs often make basic errors of grammar and semantics. The stories lack coherence, have unnatural repetitio…

arXiv Computational LinguisticsSource content · Analysis pendingMultilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking

arXiv:2609.10745v1 Announce Type: new Abstract: Multimodal entity linking grounds entity mentions in text and images to knowledge-base entries. These systems degrade on rare entities, but prior work measures rarity primarily through popularity-based metrics such as pageviews. We broaden this view using knowledge-graph structural metrics that capture how well an entity is documented and connected. These metrics identify many rare entities that popularity metrics miss. Across the resulting rare-entity slices, state-of-the-art accuracy drops by 15.4-39.9%, showing that different rarity definitions expose different failure modes. To address these failures, we introduce a simple, training-free framework in which a reasoning-capable vision-language model iteratively searches and reasons over Wi…

arXiv Computational LinguisticsSource content · Analysis pendingThink Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking

NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

arXiv:2609.10715v1 Announce Type: new Abstract: We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducing an explicit and more challenging concept-level objective while preserving standard token-level autoregressive generation. NCP-ArchPreview builds a latent space by constructing a product-quantized concept vocabulary directly from its hidden states, and subsequently learns to predict future concepts via a dedicated Concept Module. These predicted concepts are then fed back to the token level to guide subsequent generation, with NTP and NCP trained jointly end-to-end.…

arXiv Computational LinguisticsSource content · Analysis pendingNCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement

arXiv:2609.10702v1 Announce Type: new Abstract: Learning from limited text requires models to use context, generalize to new inputs, and retain useful capabilities. Qiushi Engine conducted a long-horizon, end-to-end autonomous research program on BabyLM 2026 Strict-Small, within 10 million corpus words and 100 million cumulative word presentations. Three stages connected frontier advancement, principle discovery, and principle-guided model improvement. Stage I combined compact restatements, budget reinvestment, and residual incremental learning to build a frontier model. Stage II found that exact repetition and aligned restatement produce different patterns of context use, depending on target relations and prediction windows. In controlled tasks, recovering familiar performance did not en…

arXiv Computational LinguisticsSource content · Analysis pendingData-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement

A Bellman Optimality Equation for Plasticity

arXiv:2609.10776v1 Announce Type: new Abstract: In continual reinforcement learning, carefully managing the stability-plasticity tradeoff remains a core challenge. Recent work by Abel et al. (2025) formalized this dilemma by defining plasticity as the generalized directed information from an agent's observations to its actions, and empowerment as the generalized directed information from its actions to its observations. This formulation successfully reframes the traditional stability-plasticity tradeoff as an empowerment-plasticity tradeoff. However, while extensive literature exists on optimizing for empowerment, there is currently no research addressing the optimization of plasticity under this new definition. This paper presents preliminary work toward optimizing plasticity within Mark…

arXiv Machine LearningSource content · Analysis pendingA Bellman Optimality Equation for Plasticity

The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes

arXiv:2609.10739v1 Announce Type: new Abstract: A truth probe fitted where truthful reporting and a task's prescribed action coincide cannot distinguish those targets from its fitting labels alone. We call this failure of semantic identification perfect aliasing. In a controlled binary reporting game, truth and prescribed-action probes fitted on compliant contexts solve the same optimization. On rival contexts their labels are complements, forcing their AUROCs to sum to one; this identity holds across 751 cell-layer pairs to floating-point precision. We separate prescribed output symbols from semantic action using randomized codebooks, then separate truth from prescribed action by fitting on mixed compliant and rival contexts. For a reward-trained Gemma-2-9B policy that answers falsely on…

arXiv Machine LearningSource content · Analysis pendingThe Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes

GEOSTEER: Geodesic Optimization for Activation Steering in Large Language Models

arXiv:2609.10658v1 Announce Type: new Abstract: Activation steering provides a lightweight way to control large language models (LLMs) by modifying their hidden activations at inference time. Among these approaches, norm-preserving steering aims to change model behavior without altering the activation norm, reducing the risk of representation collapse and degradation. However, existing norm-preserving methods are limited by predefined steering trajectories and by their reliance on one-step updates, which may fail to capture the complex structure of activation distributions. We propose GeoSteer, an optimization-based method for norm-preserving activation steering. GeoSteer formulates steering as a Riemannian optimization problem and updates activations through a sequence of small geodesic…

arXiv Machine LearningSource content · Analysis pendingGEOSTEER: Geodesic Optimization for Activation Steering in Large Language Models

Zero-shot rib design: merging training-free generative prior with topology optimization

arXiv:2609.10643v1 Announce Type: new Abstract: Natural load-bearing patterns such as leaf venation, trabecular bone, and spider webs achieve high stiffness per unit mass, yet classical topology optimizers rarely reach such geometries, and few let engineers express structural design intent through natural language. This work treats a frozen text-to-image diffusion model as a training-free source of design knowledge and distills it into the physics loop of density-based topology optimization via score distillation sampling, so that a text prompt becomes an explicit, machine-interpretable representation of engineer intent. The prompt-induced generative gradient and the finite element sensitivity are combined at every iteration, letting physics decide which prompt-induced features survive. I…

arXiv Machine LearningSource content · Analysis pendingZero-shot rib design: merging training-free generative prior with topology optimization

Halo: Improving forecast accuracy through heteroscedastic estimation

arXiv:2609.10589v1 Announce Type: new Abstract: Heteroscedastic forecasting, where a network estimates a scale parameter alongside a location parameter, is normally motivated by uncertainty quantification. This paper shows it also improves the point estimate, in contrast to reported negative results for heteroscedastic estimation outside time series. Halo is a modification that reuses an existing deep forecaster's architecture, giving it a second output for the scale of its implied distribution and training it under the matching negative log likelihood. Adapting three state-of-the-art models --- a transformer, a graph network paired with a variational autoencoder, and a single-layer convolutional network --- under both Gaussian and Laplacian losses demonstrates the phenomenon. On the five…

arXiv Machine LearningSource content · Analysis pendingHalo: Improving forecast accuracy through heteroscedastic estimation

M3-Former: Multimodal Transformer with Mixture-of-Experts for Long-Term Vessel Trajectory Prediction

arXiv:2609.10559v1 Announce Type: new Abstract: To address the challenges of behavioral multimodality, limited semantic utilization, and long-term error accumulation in vessel trajectory prediction, this paper proposes M3-Former, a multimodal trajectory prediction framework enhanced by large language models (LLMs). The proposed framework incorporates vessel static attributes and navigational intent as semantic priors for long-term trajectory modeling. Specifically, a unified multimodal representation space is constructed, in which static semantic information is encoded by a pre-trained LLM and aligned with dynamic trajectory features through self-attention. To jointly capture global route planning and local motion variations, a dual-granularity Mixture-of-Experts (MoE) architecture is int…

arXiv Machine LearningSource content · Analysis pendingM3-Former: Multimodal Transformer with Mixture-of-Experts for Long-Term Vessel Trajectory Prediction

XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?

arXiv:2609.09428v1 Announce Type: new Abstract: Evaluating the quality of explanations produced by explainable AI (XAI) methods remains challenging because existing approaches often rely on subjective human judgment, limiting reproducibility, scalability, and comparability between studies. We examine whether LLMs can serve as a reproducible and scalable mechanism to make comparative assessments of the quality of XAI explanations. We introduce XAI-Arena, an LLM-as-a-judge framework for scalable, reproducible, multidimensional, and stakeholder-sensitive evaluation of XAI explanation quality. XAI-Arena then allows us to compare XAI explanations along various dimensions, namely, perceived simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and ove…

arXiv AISource content · Analysis pendingXAI-Arena: Can LLMs Assess the Quality of XAI Explanations?

Together AI expands fine-tuning service with more models, live metrics, and finer controls

Together Fine-Tuning adds the latest open-weight models, live experiment tracking, Expert LoRA, early stopping, tokenized dataset previews, pre-flight validation, and lower training prices on selected models.

Together AI BlogSource content · Analysis pendingTogether AI expands fine-tuning service with more models, live metrics, and finer controls

SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign

Proteins are fundamental to biological processes, with their function determined by the complex interplay between the amino acid sequence and the three-dimensional structure. Developing generative models capable of understanding this intrinsically multi-modal relationship is crucial for fields like drug discovery and protein engineering. Existing models often rely on a multi-stage training process where autoencoders that tokenize data into latent representations are trained in a first stage. Secondly, a generative model is trained on the latent representation of the autoencoder(s), i.e…

Apple Machine Learning ResearchSource content · Analysis pendingSimpleDesign: A Joint Model for Protein Sequence and Structure Codesign

Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth references. This paradigm suffers from the “one-to-many” nature of video description, where high-quality captions are often penalized for lexical mismatches or valid shifts in visual focus. Furthermore, such assessments are typically one-dimensional, failing to provide a fine-grained analysis of caption quality. To address this, we redefine caption quality via information fidelity: A caption must maximize the coverage…

Apple Machine Learning ResearchSource content · Analysis pendingPutting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation

Sign language processing systems have traditionally operated at the sentence level, ignoring critical discourse phenomena fundamental to sign language comprehension. We introduce DiscoSign, a computational approach for discourse-aware text to sign language gloss translation grounded in linguistic research. We address three key phenomena within our modular Large Language Model (LLM)-based translation framework: (i) spatial coreference resolution, where entities maintain consistent spatial locations throughout discourse; (ii) Question-Answer Clauses (QACs), pseudocleft structures serving…

Apple Machine Learning ResearchSource content · Analysis pendingDiscoSign: Discourse-Aware Text to Sign Language Gloss Translation

Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster

Production LLM applications rarely receive a question nobody has asked before. Support assistants and RAG pipelines field the same intents thousands of times a day, each phrased differently, and most stacks treat every phrasing as a fresh, fully billed request. Redis LangCache is a fully managed semantic caching service that sits between the application and […] The post Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster appeared first on MarkTechPost.

MarkTechPostSource content · Analysis pendingMeet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster

Reduce LLM latency with prefix-aware routing on Amazon SageMaker Inference

Amazon SageMaker Inference now offers prefix-aware routing, a routing strategy that sends requests sharing the same prompt prefix to the same instance so the KV cache stays warm. In benchmarks on Llama 3.1 70B, it reduced P50 time-to-first-token by up to 77% and raised KV cache hit rates from about 25% to over 80%.

AWS Machine Learning BlogSource content · Analysis pendingReduce LLM latency with prefix-aware routing on Amazon SageMaker Inference

Video and image search in Amazon Bedrock Knowledge Base using Marengo 3.0

TwelveLabs Marengo Embed 3.0 is now generally available as an embedding model in Amazon Bedrock Knowledge Bases, bringing fully managed natural language search to video, image, and audio content. This walkthrough shows how to build a knowledge base powered by Marengo 3.0 and run semantic queries against your media.

AWS Machine Learning BlogSource content · Analysis pendingVideo and image search in Amazon Bedrock Knowledge Base using Marengo 3.0

GPT Images 2.5 promises edits that leave the rest of your image alone

When OpenAI launched GPT Images 2.5 this week, the company promised better results for a common editing task: changing one The post GPT Images 2.5 promises edits that leave the rest of your image alone appeared first on The New Stack.

The New Stack AISource content · Analysis pendingGPT Images 2.5 promises edits that leave the rest of your image alone

Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

This week, Mistral announced it raised €3 billion in a Series D funding round, pushing its post-money valuation past €21 The post Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it. appeared first on The New Stack.

The New Stack AISource content · Analysis pendingMistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video

Manufacturing floors, warehouses and production lines rarely stay fixed — tasks change, layouts shift and new products arrive, and most robots can’t keep up without significant reprogramming. Skild AI’s new S1 robot foundation model helps address this, designed to learn previously unseen, long-horizon tasks from a single video demonstration. The model, launched last week, uses […]

NVIDIA BlogSource content · Analysis pendingSkild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video

Model-agnostic PII detection with LLMs

A configurable, model-agnostic detector that turns any large language model on Amazon Bedrock into a PII detector. Because the entities to detect live in a prompt rather than in code, one detector adapts to new entity types without retraining, and it outperforms an off-the-shelf tool across five public corpora and nine LLM-based detectors.

AWS Machine Learning BlogSource content · Analysis pendingModel-agnostic PII detection with LLMs

Introducing North Small Translate: A leading sovereign open-weight machine translation model

Today, we're releasing North Small Translate, a mixture-of-experts machine translation model with strong performance across 50+ languages. Across WMT26 benchmarks,¹ North Small Translate achieves an 83.6 score across al…

Cohere BlogSource content · Analysis pendingIntroducing North Small Translate: A leading sovereign open-weight machine translation model
Tools

Meta says it’s changing AI suggestions after posing invasive personal questions

Meta says it's making changes to the prompts suggested by its AI chatbot after a viral video showed it digging for personal information about a woman's young daughters, as reported earlier by Futurism. In a statement to The Verge, Meta spokesperson Dina El-Kassaby says the company "missed the mark," adding that "the feature never should have prompted the individual with questions like that." Last week, Instagram user Kalie Robins posted a video explaining how Meta AI presented her with an invasive suggestion after cross-posting a clip to Facebook. The AI prompt, "Who is the child passenger?" appeared beneath a video of her and her child sin … Read the full story at The Verge.

The Verge AISource content · Analysis pendingMeta says it’s changing AI suggestions after posing invasive personal questions

Digested week: royal hairlines – Prince George’s hair has such brio as he starts at Eton | Emma Brockes

Plus, your toilet flush can give you norovirus, AI doom-mongering and a drag Ann Droid In a news cycle to make everyone older than gen Z feel ancient, we start the week anticipating the 25th anniversary of 9/11 at the end. Social media floods with poignant interviews and images from lower Manhattan that day, triggering memories for the rest of us of where, who and how young we were. (Ridiculously, I was having my eyebrows done in Hendon, north London, and came out to find a white van pulled over at a crazy angle to the curb, doors flung open, radio cranked up, and people stopped in the street to listen. I remember looking over the city and wondering if I’d seen jets, or explosions.) Continue reading...

The Guardian AISource content · Analysis pendingDigested week: royal hairlines – Prince George’s hair has such brio as he starts at Eton | Emma Brockes

5 Python Techniques for Efficient Resource Orchestration

This article explains 5 Python techniques for efficient resource orchestration and sticks to what's stable today, 3.11 and later for the core techniques, with one 3.14-specific tool called out explicitly as requiring that version

KDnuggetsSource content · Analysis pending5 Python Techniques for Efficient Resource Orchestration

izzit

Discussion | Link

Product Hunt AISource content · Analysis pendingizzit

UK economy defies forecasts with surprise 0.4% growth in July

Welcome boost for chancellor with unexpected rise put down to rapid growth of AI in services sector outweighing Iran war fallout The UK economy grew in July as the rapid growth of AI appeared to outweigh the economic damage from the Iran war, in a welcome boost for the John Healey before next month’s budget. Figures from the Office for National Statistics (ONS) showed a surprise 0.4% increase in gross domestic product (GDP), compared with 0.3% growth in June. City economists had forecast zero growth. Continue reading...

The Guardian AISource content · Analysis pendingUK economy defies forecasts with surprise 0.4% growth in July

claudebill

Discussion | Link

Product Hunt AISource content · Analysis pendingclaudebill

Accordio

Discussion | Link

Product Hunt AISource content · Analysis pendingAccordio

Jackalope

Discussion | Link

Product Hunt AISource content · Analysis pendingJackalope

Devin Voice

Discussion | Link

Product Hunt AISource content · Analysis pendingDevin Voice

Anthropic details bad actors’ efforts to misuse its AI for bioweapons

Report comes two days after former employee quit claiming company’s models could cause human extinction by 2030 Criminals, state-sponsored groups, spyware vendors, scientists and propagandists have attempted to use Anthropic’s powerful artificial intelligence models to design missiles and bombs, create deadly pathogens and surveil dissidents, according to a threat intelligence report the company published on Thursday. “The cases we share here aren’t typical misuse, but rather examples of the most notable and novel threat activity we’ve identified to date,” Anthropic wrote in its 154-page report. “We’re publishing this work because we believe we have a responsibility to disclose malicious misuse of our services.” Continue reading...

The Guardian AISource content · Analysis pendingAnthropic details bad actors’ efforts to misuse its AI for bioweapons

Cognition's SWE-2

Discussion | Link

Product Hunt AISource content · Analysis pendingCognition's SWE-2

Slack can now vibe-code interactive charts and reports inside chats

A new feature coming to Slack will allow you to build interactive reports, polls, dashboards, presentations, microsites, and other tools directly inside a chat. With Slackforce Surfaces, you can describe to Slackbot what you need, and it will use AI to gather information from relevant conversations and connected apps, like Google Drive or Salesforce, to create it. Once Slackbot creates a Surface, you can share it with colleagues and pin it to channels, allowing other people to view it, interact with it, and leave comments. In one example shared by Slack, a user asks Slackbot for help creating an arcade-themed visualization of AI token usage … Read the full story at The Verge.

The Verge AISource content · Analysis pendingSlack can now vibe-code interactive charts and reports inside chats

ABrush

Discussion | Link

Product Hunt AISource content · Analysis pendingABrush

sizeless

Discussion | Link

Product Hunt AISource content · Analysis pendingsizeless

chat-recall

Discussion | Link

Product Hunt AISource content · Analysis pendingchat-recall

Google to Invest $15B in Finland’s AI Infrastructure

The tech giant simultaneously revealed a nuclear power contract with Finnish operator Fortum, its first outside of the U.S.

AI BusinessSource content · Analysis pendingGoogle to Invest $15B in Finland’s AI Infrastructure

Feature Engineering in Scikit-Learn: A KDnuggets Cheat Sheet

Once feature engineering lives inside a Pipeline, each step is fitted on training data only, and the model is scored what it actually earned. And that is the idea behind this new cheat sheet.

KDnuggetsSource content · Analysis pendingFeature Engineering in Scikit-Learn: A KDnuggets Cheat Sheet

3 ways to prep for your next big race with Search

Search can help runners get race-day ready with registration alerts, tailored training plans, and more.

Google AI BlogSource content · Analysis pending3 ways to prep for your next big race with Search

ElevenLabs and Universal Music Group enter strategic agreement

Skip to content Log inSign up Contact salesLog in Sign up First-of-its-kind collaboration spans licensing and product development New platform, to be launched by ElevenLabs, will enable fans to create remixes, mashups a…

ElevenLabs BlogSource content · Analysis pendingElevenLabs and Universal Music Group enter strategic agreement
Research

Lifesaving Lincoln Laboratory device wins 2026 Excellence in Technology Transfer Award

The handheld catheterization device AI-GUIDE, created by Lincoln Laboratory and Massachusetts General Hospital, promises improved health outcomes for injured service members and civilians.

MIT News AISource content · Analysis pendingLifesaving Lincoln Laboratory device wins 2026 Excellence in Technology Transfer Award

Scientists just made quantum computer operations 1,000 times faster

Researchers have found a way to perform certain quantum operations more than 1,000 times faster, cutting thousands of repeated control cycles down to just one. The advance could reduce errors and bring reliable, fault-tolerant quantum computers closer to reality.

ScienceDaily AISource content · Analysis pendingScientists just made quantum computer operations 1,000 times faster

Anthropologic

Discussion | Link

Product Hunt AISource content · Analysis pendingAnthropologic

We must pause risky AI research while we still have the power to do so | Gaby Hinsliff

Warnings of AI’s existential threat to humanity are piling up – it’s time to listen and take them deadly seriously Another day, another horseman of the apocalypse galloping over the horizon. Lately we have heard from so many AI doomers – tech whistleblowers popping up to warn that their work is probably going to kill us – that we’re becoming almost blase about it. Humanity wiped out within a decade? Well, only if another world war or the climate crisis doesn’t get us first. Since it’s never clear whether the tech threat is real, or just a twisted form of hype from an industry that drums up investment by making their products sound more powerful than they really are, most of us settle for trying not to think about it too hard. But something about the AI researcher Jacob Coxon’s very public…

The Guardian AISource content · Analysis pendingWe must pause risky AI research while we still have the power to do so | Gaby Hinsliff

Planning along Differentiable Charts of Constraint Manifolds with General-Purpose IK Solvers

arXiv:2609.10905v1 Announce Type: new Abstract: Planning trajectories for robot manipulators under kinematic equality constraints restricts feasible motions to a measure-zero submanifold of the configuration space, requiring special algorithmic treatment. A promising strategy is parametrizing the set of feasible configurations using analytic inverse kinematics (IK). Bespoke analytic IK functions can be written to be differentiable, a necessary property for gradient-based trajectory optimization. But the vast majority of IK functions are computed by automated meta-solvers like IKFast, and are difficult to modify for differentiability. We present a new approach for computing gradients of analytic IK parameterizations: we leverage the inverse function theorem to recover the desired gradients…

arXiv RoboticsSource content · Analysis pendingPlanning along Differentiable Charts of Constraint Manifolds with General-Purpose IK Solvers

Expressive Robotic Pianist: Mastering Complex Piano Repertoire with Graph-Mimic and Musical Dynamics

arXiv:2609.10844v1 Announce Type: new Abstract: Enabling robots to perform musical instruments with human-level expressivity represents a frontier in bridging the gap between mechanical precision and artistic interpretation. Despite advances in robotic dexterity, replicating the fluid finger transitions and nuanced dynamic control characteristic of human pianists remains a significant challenge. Through a reinforcement learning-based control framework, we demonstrate that a dexterous robotic hand can achieve high-fidelity performance across a diverse piano repertoire. Central to our approach is a graph-based optimization strategy that guides the robot to generate natural pre-press and key-press fingering strategies that closely resemble human movement patterns. To achieve expressive sound…

arXiv RoboticsSource content · Analysis pendingExpressive Robotic Pianist: Mastering Complex Piano Repertoire with Graph-Mimic and Musical Dynamics

Lie-Algebraic Bell Recurrences for Arbitrary-Order Twist Jets and Parallel-Mechanism Closure

arXiv:2609.10748v1 Announce Type: new Abstract: This paper develops an arbitrary-order kinematic construction that links serial propagation, parallel-mechanism closure, and rigid-platform point fields within one dual screw framework. A cylindrical joint is retained as one native physical block, with revolute and prismatic joints obtained as special cases. For each fixed joint axis, ordinary Bell polynomials organize the derivatives of the exponential factor; across a chain, the noncommuting factors remain in their physical order. Initial-frame prefix and terminal-resolved covariant formulas then produce equivalent representations of the serial twist jet. For a parallel mechanism, repeated Leibniz differentiation, with joint-level derivatives organized by Bell polynomials, yields an arbitr…

arXiv RoboticsSource content · Analysis pendingLie-Algebraic Bell Recurrences for Arbitrary-Order Twist Jets and Parallel-Mechanism Closure

How Much Velocity Does Off-Ball Space Value Need? A Broadcast-Viewport Benchmark

arXiv:2609.10801v1 Announce Type: new Abstract: Velocity-aware pitch control is standard, but under a broadcast viewport half the players are off screen and on-screen velocities come from a drifting calibration. We ask at which layer of broadcast off-ball analysis velocity changes the answer. Inheriting our off-screen imputation protocol (three Metrica matches, 44 m viewport, block-bootstrap CIs), we score four velocity regimes -- none, viewport-legal observed, true-for-visible, true-for-all -- against a velocity-aware ground truth at three layers: imputation, the control surface, and team verdicts. Velocity is nearly useless for imputation (-0.2 pp against a 12--14 pp velocity-free surface MAE), first-order for the surface (-1.5 to -1.8 pp, 11--15% of that MAE), and ten times smaller for…

arXiv Computer VisionSource content · Analysis pendingHow Much Velocity Does Off-Ball Space Value Need? A Broadcast-Viewport Benchmark

Two-Parameter Flow Map Learning for Continuous-Time Diffeomorphic Image Registration

arXiv:2609.10789v1 Announce Type: new Abstract: Diffeomorphic image registration is central to medical image analysis, enabling anatomically consistent alignment across subjects. Most learning-based diffeomorphic methods model autonomous ODEs(ordinary differential equations) by parameterizing a stationary velocity field and recovering deformations via scaling-and-squaring. While non-autonomous ODEs with time-dependent velocities increase expressiveness, existing approaches rely on numerical integration to implicitly enforce flow structure that entangles model expressiveness with discretization accuracy. We propose a framework to directly learn the continuous-time solution of a non-autonomous ODE formulated as a two-parameterflow map. By enforcing cocycle consistency, a fundamental structu…

arXiv Computer VisionSource content · Analysis pendingTwo-Parameter Flow Map Learning for Continuous-Time Diffeomorphic Image Registration

Shedding Light: A Benchmark for Evaluating Lighting Understanding in Generative Image Models

arXiv:2609.10787v1 Announce Type: new Abstract: Accurate modelling of illumination is central to realistic image synthesis and scene understanding. Yet, there is little exploration into whether image generative models are good at this task or whether physical plausibility remains a key challenge for them. Clearly, significant progress has been made in realistic image synthesis, but do models truly understand lighting in a physically accurate manner? To answer this question, this work proposes a benchmark to assess the lighting understanding and harmonisation capabilities of generative models. Our key insight is that evaluating lighting understanding for such models only requires testing how well they insert novel objects into real photographs whilst maintaining consistent illumination. To…

arXiv Computer VisionSource content · Analysis pendingShedding Light: A Benchmark for Evaluating Lighting Understanding in Generative Image Models

Rethinking Handwritten Character Recognition

arXiv:2609.10572v1 Announce Type: new Abstract: Non-Latin handwritten character recognition (HCR) remains understudied. Dominant methods consider it as generic image classification, which uses model scale to implicitly learn stroke structure. Structural-prior efficiency---the principle that explicitly encoding script-geometric regularities as architectural inductive biases can be both more accurate and require fewer parameters. We introduce GraphemeNet, a unified multi-script architecture, governed by two orthogonal binary axes. Axis 1 operationalises stroke-level geometric regularity via Persistent Scaffold Injection (PSI): a script-specific asymmetric convolution injects a stroke scaffold as a weighted residual at every encoder stage, continuously anchoring learned features to script ge…

arXiv Computer VisionSource content · Analysis pendingRethinking Handwritten Character Recognition

CMNIE: An Information Extraction Benchmark for Chinese Military News

arXiv:2609.10722v1 Announce Type: new Abstract: Structured extraction from Chinese military news supports intelligence analysis, decision-making, and knowledge base construction. However, existing resources provide limited support for joint informa?tion extraction in this domain, especially when events, event arguments, entities, and relations must be modeled together. We present CMNIE, an information extraction benchmark for Chinese military news. Extend?ing military-domain resources beyond document-level event annotations, CMNIE jointly annotates event triggers, event arguments, named enti?ties, and entity relations under a unified domain schema. The dataset contains 13,000 instances collected from public Chinese military news, with manual annotations for 7 event types, 10 argument role…

arXiv Computational LinguisticsSource content · Analysis pendingCMNIE: An Information Extraction Benchmark for Chinese Military News

Adaptive Margin Ordinal Loss: Penalizing Center-Class Hedging in Ordinal Classification

arXiv:2609.10752v1 Announce Type: new Abstract: Standard cross-entropy loss causes neural networks trained on ordinal classification tasks to hedge predictions toward center classes, a failure mode we term \emph{center-class hedging}. This occurs because predicting the middle class minimizes expected symmetric loss, making it the path of least resistance regardless of the true label. Existing ordinal losses address related problems such as large-error penalization and rank consistency, but none directly suppresses center-class hedging as a function of where the true label lies relative to the ordinal center. We propose the Adaptive Margin Ordinal Loss (AMOL), a multiplicative weight applied to per-class loss terms of the form $m(k,y) = 1 + \alpha \cdot (1 - |k-c|/c) \cdot (|y-c|/c)$, wher…

arXiv Machine LearningSource content · Analysis pendingAdaptive Margin Ordinal Loss: Penalizing Center-Class Hedging in Ordinal Classification

Conformal Calibration Transfer

arXiv:2609.10737v1 Announce Type: new Abstract: Conformal prediction converts point predictions into set-valued predictions with coverage guarantees under exchangeability between calibration and deployment data. We study conformal calibration transfer, where this requirement fails because labeled calibration is available only in a source space, while prediction sets are needed in a target space linked to the source through unlabeled paired observations (e.g., paired modalities or sensor changes). We propose Transported Conformal Calibration (TCC): we transport labeled source calibration into the target space using the paired data, and then correct residual post-transport mismatch using only unlabeled target inputs. We instantiate this correction with two complementary methods: TCC-KS, whi…

arXiv Machine LearningSource content · Analysis pendingConformal Calibration Transfer

Artificial Intelligence Algorithms for the Detection of Pathologies Related to Lung Cancer through Image Analysis using Convolutional Neural Networks and Data Augmentation: a systematic mapping of the literature

arXiv:2609.10652v1 Announce Type: new Abstract: Lung cancer is one of the leading causes of death worldwide, and its early diagnosis is crucial to improving patients prognosis and quality of life. However, the process of interpreting medical images for the detection of lung cancer is complex and requires trained experts. In this context, artificial intelligence (AI) and deep learning (DL) emerge as potential tools to automate and optimize image analysis. The objective of this work is to review the most recent and relevant applications of AI and DL in the field of radiology for the detection of lung cancer. To this end, an exhaustive search was carried out in scientific databases such as PubMed,IEEEXPLORE, Scopus and Web of Science, and 96 articles published from 2015 to the present addres…

arXiv Machine LearningSource content · Analysis pendingArtificial Intelligence Algorithms for the Detection of Pathologies Related to Lung Cancer through Image Analysis using Convolutional Neural Networks and Data Augmentation: a systematic mapping of the literature

Byzantine-Robust Federated Fire Detection with a Rotating Coordinator

arXiv:2609.10647v1 Announce Type: new Abstract: We study the application of federated learning (FL) to indoor fire detection. Such fire-detection systems use edge cameras that record sensitive footage which cannot easily be collected at a central server. Existing federated solutions leave three practical obstacles unaddressed: limited uplink bandwidth, Byzantine (malicious or faulty) clients, and unconditional trust in a single, permanently fixed aggregation server. Our main contributions address all three. In particular, we provide (i) a curated indoor fire-detection dataset assembled from eight public sources; (ii) an edge-deployable detector whose model updates are compressed up to 10 time with only a small loss in balanced accuracy; and (iii) a semi-decentralized Byzantine-robust FL m…

arXiv Machine LearningSource content · Analysis pendingByzantine-Robust Federated Fire Detection with a Rotating Coordinator

Gradland: On Phenomenal Experience, Differentiated Across Many Dimensions

arXiv:2609.09306v1 Announce Type: new Abstract: This paper investigates the hypothesis that the first-order structure of physical interactions, i.e. gradients or Jacobians, characterizes the structure of phenomenal experience. It does so in an idealized world inhabited by neural networks, Gradland, where the physics are known and the functions are (mostly) differentiable. The paper introduces two measures of Jacobian structure: effective rank and cohesion, based on Kirchhoff complexity. Applying the measures to a series of worked examples shows the hypothesis accounts for: (1) the duration of experience, that it can prolong over hundreds of milliseconds; (2) the difference between what is experienced vividly and obscurely; (3) the experience of texture; (4) the blooming buzzing confusion…

arXiv AISource content · Analysis pendingGradland: On Phenomenal Experience, Differentiated Across Many Dimensions

Datasette 1.0a39 and 0.65.4 security releases

Datasette 1.0a39 and 0.65.4 security releases Today we're releasing two new security patch versions of Datasette: 1.0a39 and 0.65.4 - one for the current alpha series and one for the stable 0.65.x family. These are security fixes which you should apply if you are running a Datasette instance on the public web - in particular if that instance mixes both public and private tables. Following issues reported by Sevban Dönmez, Alex Garcia and I ran an extensive audit of Datasette using Claude Fable 5.1, GPT-5.6, and GPT-6 Astra. We then spent almost a week collaborating on and reviewing the fixes. They helped find some very subtle bugs. We'll be incorporating security audits by frontier models into all of our development work going forward. Alex came up with a way of splitting the work which I…

Simon Willison's WeblogSource content · Analysis pendingDatasette 1.0a39 and 0.65.4 security releases

datasette-publish-fly 1.4

Release: datasette-publish-fly 1.4 Sets force_https=true in fly.toml. #31 Fix for Volume could not be found bug. #32 Compatible with app-scoped deploy tokens. #34 Tags: datasette, fly

Simon Willison's WeblogSource content · Analysis pendingdatasette-publish-fly 1.4

github-to-sqlite 2.9.1

Release: github-to-sqlite 2.9.1 Fix for compatibility with sqlite-utils 4.x. #85 Tags: github, sqlite

Simon Willison's WeblogSource content · Analysis pendinggithub-to-sqlite 2.9.1

datasette 0.65.4

Release: datasette 0.65.4 See Datasette 1.0a39 and 0.65.4 security releases on the Datasette blog. Tags: security, datasette

Simon Willison's WeblogSource content · Analysis pendingdatasette 0.65.4

Reduce inference cold starts on Amazon SageMaker HyperPod with model caching

Amazon SageMaker HyperPod now supports model caching for inference, which pre-loads model weights and container images onto cluster nodes so pods read from local NVMe storage instead of downloading over the network. Learn how model caching cuts cold starts from tens of minutes to seconds, how it works, and how to enable it.

AWS Machine Learning BlogSource content · Analysis pendingReduce inference cold starts on Amazon SageMaker HyperPod with model caching

Schools are catching onto Big Tech’s playbook

It's the hot new thing in tech, and it's where all the jobs are. Students who don't learn to use it fall behind. And to help them catch up in time, its creators are graciously providing the resources and curriculum for learning it, often pro bono. That's the narrative AI companies are pitching schools on right now, but according to New York Times education technology reporter Natasha Singer, it's also how the tech industry has built its influence in classrooms for the past 15 years. In a new book, Coding Kids, Singer details how tech companies took a central role in promoting computers and coding skills in education, promising computer sci … Read the full story at The Verge.

The Verge AISource content · Analysis pendingSchools are catching onto Big Tech’s playbook
Robotics
Chips

Testing Between the Test Cases: Proving End-to-End Steering in Conditions You Never Drove

arXiv:2609.10951v1 Announce Type: new Abstract: AI-based automated vehicle testing is challenging because a model that passes every test condition can still fail in the real world. Formal verification offers a way to directly address this gap. On a simulated highway and an arterial road we trained two small end-to-end steering networks each in CARLA, one on clear conditions alone and one on clear, fog, night and low sun. All four models were driven against a 2.19 ft lane-departure budget. Without driving again, we used bound propagation, a formal method that reads the trained weights, to compute how far steering can drift at every disturbance strength between two captured images. One calculation covers more than a campaign could drive: on the arterial it spans 133 poses, where ten intensi…

arXiv RoboticsSource content · Analysis pendingTesting Between the Test Cases: Proving End-to-End Steering in Conditions You Never Drove

NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput and 58.5K Residues per GPU-Hour on 8xH100

NVIDIA has detailed BioNeMo Inference Runtime (BioIR), a Python library that accelerates biomolecular structure-prediction models on NVIDIA GPUs while staying in plain PyTorch. In a matched benchmark on 1,000 human dimer targets across 8xH100 GPUs, BioIR-accelerated Boltz-2 delivered 58.5K successfully folded residues per GPU-hour versus 20.2K for a torch-compiled open-source implementation, a 2.90x gain. The runtime optimizes at 3 layers: custom kernel selection, CUDA Graph capture, and Ray-based replica scaling that places 1 full model copy per GPU. BioIR already powered the AlphaFold Database expansion, generating about 31 million candidate protein complexes across 4,777 proteomes. The post NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput and 58…

MarkTechPostSource content · Analysis pendingNVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput and 58.5K Residues per GPU-Hour on 8xH100

Red Hat AI 3.5 tackles the GPU queue that can stall AI pilots

Red Hat released Red Hat AI 3.5 this week, a move designed to let software engineering teams run AI with The post Red Hat AI 3.5 tackles the GPU queue that can stall AI pilots appeared first on The New Stack.

The New Stack AISource content · Analysis pendingRed Hat AI 3.5 tackles the GPU queue that can stall AI pilots
Startups

When Information is Worth the Risk: Behavioral Valuation for Hazardous Robotic Exploration

arXiv:2609.10726v1 Announce Type: new Abstract: Hazardous robotic exploration requires robots to map spatial risks, such as unsafe terrain, radiation, fire, mines, or structural damage, while operating where collecting information can itself cause failure. A highly informative path may expose the robot to hazards, terminate execution, and prevent future observations. Hazardous exploration therefore requires deciding not only where uncertainty is largest, but when reducing it is worth the risk. This paper introduces a valuation-layer view of this problem. We keep the belief update, sensor model, physical risk model, and finite-horizon informative planner fixed, and change only the scalar objective used to rank feasible paths. Within this framework, we introduce a risk-augmented Behavioral…

arXiv RoboticsSource content · Analysis pendingWhen Information is Worth the Risk: Behavioral Valuation for Hazardous Robotic Exploration

More Anthropic researchers warn of AI’s perils as Musk terms fears a ‘psyop’

Insiders at the firm fear tech’s advancement could cause human extinction while others are calling it a ‘setup’ A day after a former researcher at Anthropic made an apocalyptic declaration about artificial intelligence, more researchers and staff members at the AI startup publicly agreed with him and posted their own dire warnings. In response, Elon Musk and other conservative figures on X, formerly Twitter, called the chorus of concerns a “setup” and a “psyop”. Continue reading...

The Guardian AISource content · Analysis pendingMore Anthropic researchers warn of AI’s perils as Musk terms fears a ‘psyop’

AI Coding Startup Cognition Now Valued at $48B

Interest in automated generative AI coding has exploded along with the vendor’s growth.

AI BusinessSource content · Analysis pendingAI Coding Startup Cognition Now Valued at $48B
Policy

The AI Safety Crunch and How Enterprises Should Deal With It

The competitive landscape, particularly the geopolitical clash between the U.S. and China, complicates the AI safety dilemma.

AI BusinessSource content · Analysis pendingThe AI Safety Crunch and How Enterprises Should Deal With It