Skip to content
AI News HubLIVE

Reports for this edition have been collected; translation and analysis are pending. Expand other updates to read source content.

Other updates (96)
Research

ReviewBench: An open benchmark for AI code review

We’re launching ReviewBench, a benchmark for code review agents built on representative GitHub pull requests, multi-source ground truth, calibrated evaluation, and production-aligned metrics. The post ReviewBench: An open benchmark for AI code review appeared first on The GitHub Blog.

GitHub AI & MLSource content · Analysis pendingReviewBench: An open benchmark for AI code review

Building advertising for the way people use AI

OpenAI introduces a new visual ad format in ChatGPT and expands measurement tools, attribution partnerships, and brand suitability for advertisers.

OpenAI NewsSource content · Analysis pendingBuilding advertising for the way people use AI

Autonomous mobile robot operations logistics: a dataset of jobs, dispatch events and robot states

arXiv:2610.02428v1 Announce Type: new Abstract: Autonomous mobile robots (AMRs) increasingly perform material transport in production logistics, where their operation is governed by job generation, dispatching and robot control. We present MoRoOp, a dataset of AMR operations recorded in a laboratory kit preparation and supply scenario over nine eight-hour shifts. During each shift, an AMR executed stochastically generated kit supply, empty-box refill and charging jobs. The dataset links job specifications, the operations constituting each job, dispatch events documenting operation state transitions and outcomes, and robot-state observations comprising position, orientation, velocity, per-wheel state of charge and diagnostics. It contains 1,382 jobs, 4,815 operations, 19,352 dispatch event…

arXiv RoboticsSource content · Analysis pendingAutonomous mobile robot operations logistics: a dataset of jobs, dispatch events and robot states

SoTa: Soft Tactile Skins for Dexterous Manipulation

arXiv:2610.02338v1 Announce Type: new Abstract: A growing body of work suggests that tactile sensing gives robot policies contact information that complements vision in dexterous manipulation. However, visuo-tactile robot data remains scarce: dexterous demonstrations require teleoperating robots, which limits dataset scale. Human demonstrations are far cheaper to collect and offer a path to scale this data, but only if human and robot hands carry tactile sensors with corresponding signals. This requires sensors that conform to different hand geometries, cover the full hand, and share a common layout across embodiments. We present SoTa, a low-cost capacitive tactile skin that provides full-hand coverage on humans and robots while preserving a shared layout of 202 taxels across correspondin…

arXiv RoboticsSource content · Analysis pendingSoTa: Soft Tactile Skins for Dexterous Manipulation

An AI-Based Multi-Stage Approach for Androgenetic Alopecia Assessment from Low-Magnification Scalp Images

arXiv:2610.02421v1 Announce Type: new Abstract: Androgenetic alopecia (AGA) is characterized by patterned follicular miniaturization, increased single-hair follicular units, and altered hair-shaft diameter. We present an automated quantitative scalp-analysis and clinical decision-support framework combining FU localization, ordinal visible-shaft counting, calibrated shaft-width estimation, regional aggregation, and an interpretable rule layer. The clinical cohort comprised 243 patients (127 AGA, 116 non-AGA), while the computer-vision experiments used 160 expert-annotated patients, 2,400 trichoscopic images, and approximately 158,000 FU annotations. Under patientdisjoint evaluation, YOLOv8m achieved test [email protected]=0.920 and recall=0.860; EfficientNet-B5 with a support-map channel achieved 8…

arXiv Computer VisionSource content · Analysis pendingAn AI-Based Multi-Stage Approach for Androgenetic Alopecia Assessment from Low-Magnification Scalp Images

FactorSplat: Appearance-Controllable Gaussian Proxies for Medical Volume Rendering

arXiv:2610.02382v1 Announce Type: new Abstract: Transfer functions (TFs) control color and visibility in medical volume rendering, but image-trained Gaussian proxies typically bake one transfer function into their appearance. We present FactorSplat, a per-scene N-dimensional Gaussian splatting (N-DGS) proxy that accepts region-specific intensity-to-RGBA curves at inference. A local lookup applies the authored color and opacity change, while a shared functional encoder and low-rank per-Gaussian factors learn the residual appearance response. Geometry and directional appearance remain shared across presets, with visibility control and TF-aware pruning preserving the ability to hide and reveal structures. On seven CT and MR scans, FactorSplat improves mean PSNR and changed-region error over…

arXiv Computer VisionSource content · Analysis pendingFactorSplat: Appearance-Controllable Gaussian Proxies for Medical Volume Rendering

SCION: Scene Composition with Instanced Neural Primitives

arXiv:2610.02322v1 Announce Type: new Abstract: Real-world scenes are compositional: bricks, blades of grass, pebbles, and tree leaves recur across human-built and natural environments. Existing neural scene representations model these elements independently. Most 3D Gaussian Splatting and follow-up abstraction and compression methods treat each element as unique, fitting millions of independent Gaussians per scene. Prior methods like Splat and Replace fit template objects, but they require mostly manual selection of repeated elements. As a result, these representations store redundant parameters and provide weak manipulation handles for downstream tasks. We introduce SCION, a hier- archical compositional scene representation that replaces independent Gaussians with a compact vocabulary o…

arXiv Computer VisionSource content · Analysis pendingSCION: Scene Composition with Instanced Neural Primitives

EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

arXiv:2610.02298v1 Announce Type: new Abstract: 3D editing methods are usually tested on a single edit, yet an asset is built through a long sequence of revisions, each of which must implement the requested change while leaving everything else unchanged. We introduce EditHero, to our knowledge the first benchmark for long-horizon, part-level 3D editing, with natural-language instructions and target images for both geometry and texture. A deterministic assembly engine produces the exact target after every edit, and every sequence is reviewed by hand. We use EditHero to compare 2 opposite approaches to 3D editing. Non-agentic methods operate top down, regenerating the object from a learned 3D representation and inferring what to keep. In contrast, LLM/VLM agents operate bottom up, editing t…

arXiv Computer VisionSource content · Analysis pendingEditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

HakemBench: A Turkish Benchmark of Typed Decisions

arXiv:2610.02293v1 Announce Type: new Abstract: HakemBench is a Turkish benchmark of typed decisions, in which the model under test reads a text, a question and a fixed set of options and returns a probability for every option. Version 1.0 is released fully open under CC BY 4.0, with 2,346 items and 4,275 choice, yes/no and score questions in seven tracks (fact-check triage, education, guardrails, legal routing, moderation, spam and phishing, and customer support). One harness scores decision quality (macro F1), calibration (from the normalised Brier score) and selective automation (from the normalised area under the generalised risk-coverage curve), combines them by a geometric mean and reports intervals from 2,000 bootstrap draws; probes for option order, paraphrase, English translation…

arXiv Computational LinguisticsSource content · Analysis pendingHakemBench: A Turkish Benchmark of Typed Decisions

TRACE: A Reproducible Benchmark for Electricity Price Forecasting with Official Operational Text

arXiv:2610.02256v1 Announce Type: new Abstract: Electricity price forecasting (EPF) supports scheduling, bidding, and risk management in electricity markets, yet existing benchmarks focus mainly on numerical inputs, leaving the forecasting value of forecast-time textual context insufficiently evaluated. We introduce TRACE, a reproducible benchmark of 7,300 zone--day instances pairing prices from five zones in a major U.S. market with official operational text available at the forecast cutoff. TRACE reconstructs official operational text at each cutoff, preventing post-cutoff information leakage. We evaluate TRACE for semantic alignment and forecasting value. Semantic assessments align with central movement and both tail risks in ground-truth prices, most consistently for upper-tail price…

arXiv Machine LearningSource content · Analysis pendingTRACE: A Reproducible Benchmark for Electricity Price Forecasting with Official Operational Text

Approximation Property of Dropout Neural Networks: Sobolev Rates and Confidence Bounds

arXiv:2610.02253v1 Announce Type: new Abstract: The universal approximation property of dropout neural networks does not by itself describe the network size required for an accurate random realization. In this work, we study approximation of the unit ball of $W^{n,\infty}([0,1]^d)$ by ReLU networks whose edges are retained independently with probability $p$. The approximation error is measured uniformly over the input domain, and the guarantee holds with probability at least $1-\delta$ for a single sampled network. We construct networks of constant depth and size $\widetilde O_{n,d}(p^{-9}\varepsilon^{-\max\{d/n,2\}} \log(1/\delta))$. The construction combines bounded local subnetworks, localization on a successful approximation event, and a multiscale Taylor decomposition. Conversely, So…

arXiv Machine LearningSource content · Analysis pendingApproximation Property of Dropout Neural Networks: Sobolev Rates and Confidence Bounds

Nearest-neighbour baselines for fingerprint prediction from MS/MS spectra under different assumptions

arXiv:2610.02249v1 Announce Type: new Abstract: It has recently been shown that nearest-neighbour retrieval provides a strong baseline for molecular fingerprint prediction from MS/MS spectra, with several variants matching or outperforming current deep learning models (Khoo and Barzilay, 2026; Liu et al., 2026; Gupta et al., 2026). Importantly, "nearest neighbour" encompasses a family of retrieval methods that differ in the information assumed to be available at inference. In this report, we systematically compare several nearest-neighbour variants and show how these differing assumptions affect performance. Our goal is to establish stricter baselines that enable more rigorous benchmarking and better measure progress in this area.

arXiv Machine LearningSource content · Analysis pendingNearest-neighbour baselines for fingerprint prediction from MS/MS spectra under different assumptions

A Multi Method Importance and Performance Efficiency Analysis of Topological Metrics for Natural Visibility Graph Based Cyber Attack Detection

arXiv:2610.02342v1 Announce Type: new Abstract: Natural Visibility Graph (NVG) based analysis characterizes network traffic through topological descriptors reflecting different structural properties. However, not all descriptors contribute equally to cyber-attack classification, and extracting a large metric set can increase computational cost. This study evaluates 21 NVG derived topological metrics and investigates whether a compact subset can preserve classification capability while improving computational efficiency. Four importance analysis methods SHAP, grouped Permutation Importance, Boruta, and Recursive Feature Elimination (RFE) are integrated through a Consensus Ranking strategy. Based on this ranking, Full21, Top15, Top10, Top7, Top5, and Top3 configurations are evaluated using…

arXiv AISource content · Analysis pendingA Multi Method Importance and Performance Efficiency Analysis of Topological Metrics for Natural Visibility Graph Based Cyber Attack Detection

Can an Open Model Do Security Research? Cantina’s apex-flash-1 Solves 40 of 60 Held-Out Bug Tasks

Cantina Security, with Yeta Labs, has released apex-flash-1, an open-weights model trained specifically for vulnerability research. It is a reinforcement learning fine-tune of Z.ai’s GLM-5.3-Flash, released on Hugging Face under the MIT license. Is it deployable? Yes, the MIT weights serve on vLLM, SGLang or Transformers, but BF16 needs roughly 640 GB of GPU memory. […] The post Can an Open Model Do Security Research? Cantina’s apex-flash-1 Solves 40 of 60 Held-Out Bug Tasks appeared first on MarkTechPost.

MarkTechPostSource content · Analysis pendingCan an Open Model Do Security Research? Cantina’s apex-flash-1 Solves 40 of 60 Held-Out Bug Tasks
Models

Attempts to Keep Humans in the AI Loop May Actually Push Them Out

A crucial safeguard against AI agents going rogue—keeping humans in the loop to review and approve their decisions—will fail unless designers and users change their current practices, a trio of leading AI ethics researchers argue. Though most autonomous agents have systems to keep users in the loop about their actions, in practice these processes actually push humans out of the loop, the authors argue in a paper posted to ArXiv on 6 September. In other words, “the human just becomes this meat tool to give permissions without the cognitive capability to engage,” says one of the authors, Avijit Ghosh, the lead technical AI policy researcher at Hugging Face, an open-source machine learning platform. In the near term, the paper says, humans’ being out of the loop leads to agents acting in way…

IEEE Spectrum AISource content · Analysis pendingAttempts to Keep Humans in the AI Loop May Actually Push Them Out

AI Solves a Major Unsolved Math Problem. Not Everyone Is Happy

On 13 September, leading luminaries and up-and-coming talents in mathematics and computer science congregated at the annual Heidelberg Laureate Forum in Germany for a week of discussion, networking, and—naturally—wurst. Having attended several of these events in the past, the mathematicians’ chatter is normally concerned with which famed researchers are attending, or what interesting problems they have been working on. But instead, every snippet of conversation caught in passing or any debate accidentally overheard was about how AI companies such as OpenAI, Anthropic, and Google are steamrollering their way through mathematics. And there is good reason for these fervent discussions. Mathematics is the perfect testing ground for AI, involving step-by-step logical reasoning and answers that…

IEEE Spectrum AISource content · Analysis pendingAI Solves a Major Unsolved Math Problem. Not Everyone Is Happy

Yandex Introduces Sona: A Single Generative Recommender That Replaces Entire Recommendation Cascade

One transformer ran candidate generation and ranking in Yandex Music's A/B test without hand-engineered features, lifting likes 11.42%. The post Yandex Introduces Sona: A Single Generative Recommender That Replaces Entire Recommendation Cascade appeared first on MarkTechPost.

MarkTechPostSource content · Analysis pendingYandex Introduces Sona: A Single Generative Recommender That Replaces Entire Recommendation Cascade

The Story of Qwen: Alibaba’s AI Models From 7B to 2.4T

Alibaba's Qwen went from an invite-only chatbot in April 2023 to a 2.4-trillion-parameter open-weight model in August 2026. This is the full story, release by release: every major model, its key feature, and how its license changed. Each claim links to its source. The post The Story of Qwen: Alibaba’s AI Models From 7B to 2.4T appeared first on MarkTechPost.

MarkTechPostSource content · Analysis pendingThe Story of Qwen: Alibaba’s AI Models From 7B to 2.4T

Keep the Effect, Drop the Actor: Programmable Effect-to-Execution World-Action Models

arXiv:2610.02398v1 Announce Type: new Abstract: A robot demonstration records two things in the same frames: what happened to the objects, and how one particular arm made it happen. We condition on the first. A demonstration is compiled into an effect program: the 3D keypoint trajectories of the objects that moved, two points marking where each was held, and the configuration the scene ends in, with the demonstrator removed. PEWAM, a 71.5M-parameter world-action model, generates effect, robot execution, action and terminal state as four streams with independent flow-matching times, so clamping a program and sampling the execution turns inference into programming, re-solved closed loop from the live scene. On held-out LIBERO-Goal tasks, one demonstration's program completes 40 of 90 episod…

arXiv RoboticsSource content · Analysis pendingKeep the Effect, Drop the Actor: Programmable Effect-to-Execution World-Action Models

Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation

arXiv:2610.02368v1 Announce Type: new Abstract: Long-horizon compositional manipulation has become increasingly important for real-world robot deployment, where a single task involves multiple coordinated subtasks. Existing world-action models (WAMs) jointly predict short-horizon visual futures and actions, but typically lack explicit subtask-level reasoning. We propose Visual Goal-conditioned Action Reasoning (ViGAR), a hierarchical framework that factorizes manipulation into a visual subgoal planner and a subgoal executor. Given the current observation and global instruction, the subgoal planner predicts a visual subgoal for the next subtask. The subgoal executor then jointly generates future visual trajectories and actions conditioned on the predicted subgoal. Both components share a p…

arXiv RoboticsSource content · Analysis pendingRethinking World-Action Model for Compositional and In-Context Robotic Manipulation

SocialVLA: A Social Perception Gateway for Human-Reaction-Based Failure Detection and Recovery in VLA Manipulation

arXiv:2610.02360v1 Announce Type: new Abstract: Vision-language-action (VLA) policies enable diverse robotic manipulation but can fail during execution without recognizing their own errors. Human observers provide complementary signals, as unexpected robot behavior can trigger rapid vocal, facial, or verbal reactions before failure is completed. We introduce SocialVLA, a local, policy-agnostic social perception gateway that converts spontaneous human reactions into runtime intervention signals for VLA manipulation. SocialVLA combines causal paralinguistic audio detection, visual reaction recognition, explicit stop phrases, and robot-relevance estimation. An asynchronous first-event fusion mechanism triggers a VLA hold from the earliest sufficiently confident signal, while a separate speec…

arXiv RoboticsSource content · Analysis pendingSocialVLA: A Social Perception Gateway for Human-Reaction-Based Failure Detection and Recovery in VLA Manipulation

Filter-Aware Fine-Tuning for Safe Humanoid Whole-Body Tracking

arXiv:2610.02341v1 Announce Type: new Abstract: Safe whole-body motion is essential for deploying humanoid robots in unstructured environments. Modern humanoid control commonly separates reference specification from execution, with a planner, teleoperator, or motion generator providing a reference that a reinforcement-learning policy tracks through dynamically feasible whole-body control. Runtime safety filters, such as control barrier functions (CBFs), offer a promising approach for enforcing newly introduced constraints via interventions on the tracker's outputs. We show, however, that treating the tracking policy and safety filter independently induces fundamental mismatches, as filtering alters both the executed actions and the induced state distribution. We study this policy-filter i…

arXiv RoboticsSource content · Analysis pendingFilter-Aware Fine-Tuning for Safe Humanoid Whole-Body Tracking

World-Calibrated Proposal-to-Action Flow for Vision-Language-Action Models

arXiv:2610.02323v1 Announce Type: new Abstract: Flow-based Vision-Language-Action (VLA) policies generate action chunks by transporting samples from a task-agnostic isotropic Gaussian source. As this source is conditioned on neither recent execution nor predicted future evolution, (i) it discards the local continuity established by recently executed motion. (ii) Even when predictive world representations are introduced, they often only condition the transport dynamics rather than determine where generation starts, how far it may deviate, or along which action directions it may expand. Building on this observation, we introduce ProAct, a world-calibrated proposal-to-action framework that makes the generative source itself predictable. (i) To preserve motion continuity, a lightweight Propos…

arXiv RoboticsSource content · Analysis pendingWorld-Calibrated Proposal-to-Action Flow for Vision-Language-Action Models

Octrees as an Explicit 3D Language

arXiv:2610.02388v1 Announce Type: new Abstract: Existing 3D large language models (LLMs) compromise on two fronts: they compress shapes into latent codebook indices or coordinate text, which removes spatial structure from what the model observes, and they acquire the 3D modality by fine-tuning the backbone, which overwrites its general language ability. We present OctLLM, which addresses both limitations. Geometry enters as an explicit 3D sequence of octree occupancy tokens. However, full octree sequences grow rapidly with depth; OctLLM therefore randomly empties penultimate-level nodes and omits descendants while preserving shape, yielding a shorter coordinate- and depth-anchored Sparse Octree (S-Octree) for position-aware mask-modeling generation and 3D understanding. On the other front…

arXiv Computer VisionSource content · Analysis pendingOctrees as an Explicit 3D Language

EviDent-CBCT: Evidence-Bottlenecked Report Generation from Dental CBCT under Non-Exhaustive Report Supervision

arXiv:2610.02375v1 Announce Type: new Abstract: Dento-maxillofacial cone-beam CT (CBCT) reports may contain dozens of tooth-specific, anatomical, and spatial findings from a single 3D scan. Learning to generate such reports from limited clinical data is challenging because routine reports may not exhaustively document image findings, and a non-mention may reflect either absence or non-reporting. We present EviDent-CBCT, an evidence-bottlenecked framework designed for this incomplete supervision. An anatomy-aware network maps each CBCT scan to a discrete record of tooth-level, global, and tooth-IAC evidence. A dental-logic consistency projection reconciles incompatible evidence before a deterministic renderer and an image-blind local language model generate the report using only this recor…

arXiv Computer VisionSource content · Analysis pendingEviDent-CBCT: Evidence-Bottlenecked Report Generation from Dental CBCT under Non-Exhaustive Report Supervision

SCOPE-4D: Endoscopic 4D Geometry Foundation Models

arXiv:2610.02343v1 Announce Type: new Abstract: Geometric understanding supports endoscopic navigation and robotic assistance, but learning reliable endoscopic geometry faces two challenges: scarce geometric annotations and ambiguity between camera motion and tissue deformation. We present SCOPE-4D, an endoscopic 4D geometry foundation model that jointly predicts camera parameters, dense geometry, and 3D tissue trajectories from monocular RGB video in a single forward pass. Our curation and annotation pipeline constructs SCOPE-5K, a collection of approximately 5,000 clips spanning real and synthetic gastrointestinal endoscopy and laparoscopy. The collection provides rich geometric supervision and includes newly collected phantom and real-colonoscopy evaluation sets. Geometric supervised f…

arXiv Computer VisionSource content · Analysis pendingSCOPE-4D: Endoscopic 4D Geometry Foundation Models

Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation

arXiv:2610.02612v1 Announce Type: new Abstract: Simultaneous speech translation must emit useful target text before the source is complete while preserving every committed token. We adapt a full-utterance speech language model using prefix supervision derived from its own complete- and partial-waveform translations, requiring neither transcripts nor human translations. We compare single-turn forced-prefix and multi-turn append-only decoding, use a confidence threshold to control the inference-time quality--latency trade-off, and vary the density of training prefixes with a separate synthesis margin. On FLEURS and CoVoST2 in three language directions, prefix training improves quality--latency frontiers over the unadapted model, and confidence provides the broadest consistently competitive…

arXiv Computational LinguisticsSource content · Analysis pendingLearning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation

Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks

arXiv:2610.02549v1 Announce Type: new Abstract: Time and event expression extraction are fundamental temporal reasoning tasks, but the problem remains difficult due to annotation ambiguity, domain sensitivity, and unstable model behavior. Existing evaluations focus on in-domain performance, offering limited insight into reliability under distribution shifts. We evaluate multiple model configurations across families, architectures, and reasoning strategies over four dimensions of generalization, examining transfer from base performance, cross-dimensional correlations, and the effects of scale, architecture, and prompting. This provides a systematic study of how prompted LLMs generalize in time and event expression extraction tasks. We find that strong base-task performance generally predic…

arXiv Computational LinguisticsSource content · Analysis pendingEvaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks

A generative-informed neuro-symbolic framework for syntactic ambiguity resolution: Evidence from Arabic DPs

arXiv:2610.02529v1 Announce Type: new Abstract: Syntactic ambiguity poses a persistent challenge for Arabic NLP, particularly in morphologically rich nominal constructions where multiple structu6ral interpretations may be compatible with the same surface sequence. This study proposes a generatively informed neuro-symbolic framework for resolving structural ambiguity in Modern Standard Arabic (MSA) DPs. The framework integrates generative syntactic notions with AraBERT by representing ambiguity as a candidate-based decision task in which linguistically motivated alternatives are explicitly constructed and evaluated through candidate-conditioned input representations. Findings indicate that the model achieved 96.88% accuracy, 95.92% macro-F1, 96.83% weighted F1, and 93.94% binary F1 on the…

arXiv Computational LinguisticsSource content · Analysis pendingA generative-informed neuro-symbolic framework for syntactic ambiguity resolution: Evidence from Arabic DPs

From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders

arXiv:2610.02486v1 Announce Type: new Abstract: Typed decision models answer schema-constrained questions about a text in one forward pass and return probabilities meant to be thresholded. We ask whether biomedical sentence encoders trained for retrieval are good starting points for such models. We present SBERT2S1, which converts Sentence-Transformers encoders into bi-encoder, cross-head (C) and prior-fused residual (PFR) decision models, together with BIODECIDE, a biomedical typed-decision suite, and MEDLINE-S1, 243k training decisions derived from NLM indexing. Across six parent-retriever pairs, retrieval training improves zero-shot matching of content-bearing options. After fine-tuning, its effect depends on the head: across five pairs and three training-set sizes, retrieval training…

arXiv Computational LinguisticsSource content · Analysis pendingFrom Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders

CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking

arXiv:2610.02460v1 Announce Type: new Abstract: Recent benchmarks rely on user simulators to evaluate AI agents in multi-turn interaction. While existing simulation techniques demonstrate surface fidelity to human style and behavior, ecologically valid interactive benchmarking also requires alignment in when and how agents fail across simulated and real user populations. We find that existing simulators lack outcome calibration: agreement with observed success rates and failure patterns when real users interact with the same agent. We introduce Calibrated User Embeddings (CUE), a framework that both encodes observed sessions and samples continuous representations, then decodes them into persona commands to steer LLMs to act as user simulators without training. Through this, we evaluate us…

arXiv Computational LinguisticsSource content · Analysis pendingCUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking

FinDialogLens: Event Extraction over Multi-Party Dialogue for Missed-Trade Identification in Financial Chatrooms

arXiv:2610.02455v1 Announce Type: new Abstract: Multi-party financial chatrooms are vital for sales-and-trading professionals, but their complexity makes manual recovery of missed trades infeasible: each Request for Quote (RFQ) is an event whose final price and trade outcome appear many messages after the RFQ-trigger message (the inquiry message), interleaved with concurrent RFQs from other participants. We cast this as event extraction (EE) over multi-party dialogue and present FinDialogLens, a hybrid LLM pipeline in which compact fine-tuned classifiers act as inference-time scaffolds: they detect RFQ-triggers and price/trade outcome metadata, an RFQ-Level Module segments per-event RFQ windows, and a Trade Engine fills argument roles. With GPT-4o, FinDialogLens reaches 92.1% and 94.3% ac…

arXiv Computational LinguisticsSource content · Analysis pendingFinDialogLens: Event Extraction over Multi-Party Dialogue for Missed-Trade Identification in Financial Chatrooms

Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs

arXiv:2610.02444v1 Announce Type: new Abstract: Large language models often solve a theorem forward yet fail to disprove a closely related false one: a falsification gap that supervised fine-tuning does not close and can actively worsen. We frame counterexample generation as constrained witness emission against a deterministic per-theorem Python verifier, and release SymCE, a corpus of 4,707 false undergraduate-algebra and real-analysis conjectures, each paired with executable verifiers. The verifier also serves as the reward function, making SymCE a training environment. Training Qwen3-4B with SFT followed by GRPO under this oracle reveals an imitation trap: counterexample-only SFT collapses true-theorem recognition from 0.27 to 0.00, while RLVR with a sparse outcome-only reward repairs…

arXiv Computational LinguisticsSource content · Analysis pendingCounterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs

Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents

arXiv:2610.02425v1 Announce Type: new Abstract: Static evaluations credit a language model for naming the right move, but an agent must carry a plan through to a verified outcome while an opponent responds. We introduce XiangqiBench, an executable benchmark that measures this difference in Chinese chess: starting from 119 tactical endgames with forced mates supported by engine or checks-only search, an LLM agent must deliver checkmate against an engine defender. An interactive REPL interface separates real moves, state queries, and forward simulation, and we record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols. Three signals that look like competence each overstate closed-loop success. (i) The Conversion Gap: models play the stored reference first mov…

arXiv Computational LinguisticsSource content · Analysis pendingFinding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents

Overcoming Challenges of Interpretive Structural Modeling with Large Language Models

arXiv:2610.02254v1 Announce Type: new Abstract: Interpretive Structural Modeling (ISM) is a well-known process for multi-criteria decision making. The success of ISM over other methodologies is its ability to model causal relationships, the binary scale of factors, and resulting hierarchical representation. Traditionally, the modeling process is performed by repeated interactions with subject matter experts until consensus is reached. This process is tedious, labor-intense, and most importantly limits the ability of ISM to scale to studies with hundreds of variables. Drawing on existing work of causal graph discovery with large language models (LLM) as imperfect experts, this work explores an integrated LLM-ISM approach for ISM. Pairwise, k-wise, rowwise, and full graph discovery methodol…

arXiv Machine LearningSource content · Analysis pendingOvercoming Challenges of Interpretive Structural Modeling with Large Language Models

Counterfactual Predictions in Scientific Emulators Without Controlled Experiments

arXiv:2610.02252v1 Announce Type: new Abstract: Many scientific questions require reasoning about what was never observed: What if the conditions, interventions, or history had been different? Models can predict accurately on observed data yet fail on such what-if queries when correlated inputs are varied independently. A common remedy is to add controlled simulation data in which these factors are explicitly disentangled, but this requires access to a simulator, can be computationally expensive, and inherits the simulator's modeling assumptions. We introduce ReRoute, a framework for targeted scientific what-if prediction that combines factual data with partial mechanistic knowledge, without requiring controlled intervention data for adaptation. ReRoute fixes the queried input of a pretra…

arXiv Machine LearningSource content · Analysis pendingCounterfactual Predictions in Scientific Emulators Without Controlled Experiments

Rank-Aware Speculative Sampling for Diffusion Draft Trees

arXiv:2610.02251v1 Announce Type: new Abstract: Speculative sampling accelerates diffusion generation by verifying inexpensive draft states in parallel while preserving the target law. Recent tree-based methods allocate the parallel compute budget more effectively than single-chain drafts, as demonstrated by Diffusion Greedy Rejection Sampling (D-GRS). D-GRS generates $K$ conditionally independent candidates per node, and sequentially tests them in their generation order. Yet the sampled candidates admit an informative ranking without additional target-model evaluations. To exploit this, we introduce Rank-Aware Speculative Sampling (RASS), a verification rule for speculative draft trees based on rank-aware list coupling. RASS orders draft candidates along the proposal-target mean displace…

arXiv Machine LearningSource content · Analysis pendingRank-Aware Speculative Sampling for Diffusion Draft Trees

Hybrid Machine Learning-Assisted Raman Spectroscopy with Generative Feature Augmentation for Pharmaceutical Identification

arXiv:2610.02224v1 Announce Type: new Abstract: Rapid and reliable identification of pharmaceutical residues is important for safeguarding public health, ensuring food safety, and enabling practical Raman-based screening. In this study, we propose HyMLRaman, a hybrid Raman spectroscopy framework that combines deep spectral feature extraction, generative models, and classical machine-learning classifiers to identify six pharmaceutical compounds, including amoxicillin, chloramphenicol, ciprofloxacin, tetracycline, ibuprofen, and paracetamol. Raman spectra are converted into spectral images and encoded with several deep neural-network backbones, among which EfficientNet-B3 yields the most effective representation. The resulting 1536-dimensional embeddings are then used to train downstream cl…

arXiv Machine LearningSource content · Analysis pendingHybrid Machine Learning-Assisted Raman Spectroscopy with Generative Feature Augmentation for Pharmaceutical Identification

THPL: A Vision-to-Language Decision Support Framework for Rainbow Trout Feeding Management in RAS

arXiv:2610.02378v1 Announce Type: new Abstract: In Recirculating Aquaculture Systems (RAS), precision feeding is critical for minimizing costs and improving fish welfare. However, existing methods lack cognitive alignment between fish behaviors and management knowledge, impeding translation into executable, interpretable feeding decisions. To address this, we propose THPL, a generative feeding decision framework tailored for rainbow trout (Oncorhynchus mykiss) in RAS. First, Fishsort extracts trajectories to establish an Activity Coefficient (AC) quantifying feeding intensity. Second, a Hierarchical Behavior Encoder (HBE) models individual temporal progression and collective dynamics using Temporal and Set Transformers, transforming trajectory tensors into dual-evidence representations of…

arXiv AISource content · Analysis pendingTHPL: A Vision-to-Language Decision Support Framework for Rainbow Trout Feeding Management in RAS

Traversing the Satisfaction-Diversity Frontier in Text-to-Image Diffusion

arXiv:2610.02372v1 Announce Type: new Abstract: Text-to-image generation enables users to explore several images generated from the same prompt. For these generated images to be useful, each one must reflect the user's preferences, measured by a learned reward, and differ visually from the others to maintain diversity. Existing methods are limited: they either address reward and diversity separately or combine them in one aggregate score, enabling high diversity to offset low rewards. In this paper, we address these limitations by formulating generation as satisficing: every image (candidate) must satisfy a reward floor and the batch of images must satisfy a diversity cutoff. The reward floor controls the balance between worst-candidate reward and batch diversity; we show that varying thi…

arXiv AISource content · Analysis pendingTraversing the Satisfaction-Diversity Frontier in Text-to-Image Diffusion

The AI Risk Observatory: What Can We Learn from AI Disclosures in Annual Reports About Societal Resilience?

arXiv:2610.02281v1 Announce Type: new Abstract: Societal resilience research relies on access to useful and actionable data, which motivates our main research question: Can annual reports, processed at scale with LLMs, provide a useful signal about how companies disclose their response to AI? We test this by applying a reproducible two-stage classification pipeline to 9,821 annual reports from 1,362 UK listed companies (2020-2025, with partial 2026 data). We first validate the method against 474 human-annotated passages, finding high recall and moderate label-level agreement. We then report three empirical patterns: (i) between 2020 and 2025, the share of reports mentioning AI risk rose from 2.8% to 41.2%, while AI adoption disclosure also rose, from 13.8% to 45.2%, and named vendor menti…

arXiv AISource content · Analysis pendingThe AI Risk Observatory: What Can We Learn from AI Disclosures in Annual Reports About Societal Resilience?

Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses

arXiv:2610.02267v1 Announce Type: new Abstract: Agent harnesses make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input carries an injection. System-1 decision models answer such questions in a single forward pass with class probabilities, promising large cost and latency savings over LLM calls. We present a paired evaluation of an open-weight (Laya) and a hosted (Jev) System-1 model on 11 agent decision points built from 18 public sources: 7,283 base cases plus 6,640 robustness variants, with byte-identical inputs, paired tests, and cross-hardware and cross-day reproducibility checks. Jev is significantly more accurate on 9 of 11 decision points (+10.8 to +46.0 pp). Neither model beats chance on zero-shot mod…

arXiv AISource content · Analysis pendingFast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses

MintFlow: Minimal Trajectory Intervention for Constrained Flow Matching

arXiv:2610.02260v1 Announce Type: new Abstract: Flow matching models excel at generative modeling, and many downstream applications require their samples to satisfy prescribed constraints, such as observed measurements and physical laws. However, existing constrained samplers often face a trade-off: \textit{enforcing constraints can substantially displace samples from the pretrained data distribution}. To address this trade-off, we introduce \textbf{MintFlow}, a training-free constrained sampling framework that formulates constraint enforcement as a minimal intervention on the pretrained flow trajectory. MintFlow seeks the minimal perturbation of an intermediate flow state such that its subsequent evolution under the pretrained flow field satisfies the target constraint. By minimally pert…

arXiv AISource content · Analysis pendingMintFlow: Minimal Trajectory Intervention for Constrained Flow Matching

Qwen3.8 27B addition in words

Research: Qwen3.8 27B addition in words Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to see how well it could "compute the sum but return the answer in words" across increasingly large numbers. Here's the chart he shared of those results: I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect in a fully controlled environment. I pasted his image into a Codex Remote session (GPT-6 Astra) and had it run the same experiment using Qwen3.8-27B-Q4_K_M.gguf. Here's the result for a run of 30 attempts per combination with reasoning disabled: Then I ran it again with reasoning enabled. This too…

Simon Willison's WeblogSource content · Analysis pendingQwen3.8 27B addition in words

GPT-6 Astra vs GPT-6.1 Sol vs Gemini 4 Argon vs Claude Fable 5.1: Which Frontier Model Fits Which Job

Astra leads computer use, Argon leads legal and finance work, and Sol wins on price for coding agents. The post GPT-6 Astra vs GPT-6.1 Sol vs Gemini 4 Argon vs Claude Fable 5.1: Which Frontier Model Fits Which Job appeared first on MarkTechPost.

MarkTechPostSource content · Analysis pendingGPT-6 Astra vs GPT-6.1 Sol vs Gemini 4 Argon vs Claude Fable 5.1: Which Frontier Model Fits Which Job
Agents

Making Amazon Quick enterprise-ready: Automated, auditable cross-account resource promotion

Promoting Amazon Quick resources (agents, action connectors, knowledge bases, flows, and spaces) from a development to a production AWS account has been a manual, error-prone chore. This post shows how to automate cross-account promotion with an idempotent, auditable MCP server on Amazon Bedrock AgentCore.

AWS Machine Learning BlogSource content · Analysis pendingMaking Amazon Quick enterprise-ready: Automated, auditable cross-account resource promotion

Agentic retrieval with LangChain and Amazon Bedrock Knowledge Bases

Build a Retrieval Augmented Generation (RAG) application on Amazon Bedrock Managed Knowledge Base with LangChain, and see how agentic retrieval handles the multi-part questions that single-shot retrieval answers poorly. Run the same query through both paths, read the trace events, and compare what each retrieval path costs.

AWS Machine Learning BlogSource content · Analysis pendingAgentic retrieval with LangChain and Amazon Bedrock Knowledge Bases

Downgrading user roles in Amazon Quick

Amazon Quick doesn't offer a direct console path to downgrade a user from Admin or Author to Reader. This post walks through two reliable methods: a manual delete-and-recreate approach and an AWS CLI step-down sequence that downgrades roles safely while preserving asset ownership.

AWS Machine Learning BlogSource content · Analysis pendingDowngrading user roles in Amazon Quick

Evaluating multi-agent systems for explainability and helpfulness with Amazon Bedrock AgentCore

Multi-agent systems need deeper guarantees than fluent responses: they must select the right tools, respect constraints, and explain their decisions. Learn how to build a Strands-based multi-agent supply chain decisioning system and evaluate it with Amazon Bedrock AgentCore Evaluations using built-in, custom, and explainability evaluators.

AWS Machine Learning BlogSource content · Analysis pendingEvaluating multi-agent systems for explainability and helpfulness with Amazon Bedrock AgentCore

Brnch

Discussion | Link

Product Hunt AISource content · Analysis pendingBrnch

XCOR launches to trace outages in minutes. It still pages engineers.

Palo Alto Networks introduced a new approach to AI-driven observability last week, signaling a shift from dashboards and manual incident The post XCOR launches to trace outages in minutes. It still pages engineers. appeared first on The New Stack.

The New Stack AISource content · Analysis pendingXCOR launches to trace outages in minutes. It still pages engineers.

Our approach to EU text provenance rules

How OpenAI is approaching text watermarking under EU rules. Learn where watermarks apply, how detection works, and why access starts with researchers.

OpenAI NewsSource content · Analysis pendingOur approach to EU text provenance rules

OpenAI must explain action taken to stop AI hacking Australians’ private data, chair of federal inquiry says

Unions, employer groups, banks and industry experts will appear at the four-day hearings along with OpenAI, Anthropic, Microsoft and Google Get our new political email, free app or daily news podcast OpenAI must explain how they will stop their models from inappropriately accessing Australian data, the Labor chair of the parliament’s committee on artificial intelligence has warned, ahead of federal inquiry hearings which will grill the tech company alongside Anthropic, Microsoft and Google. Independent senator David Pocock is also demanding answers, saying OpenAI “still have a lot they need to answer” about their AI agent accessing Services Australia data on Medicare and the company’s “appalling” tardiness in notifying the federal government about the incident. Follow our Australia politi…

The Guardian AISource content · Analysis pendingOpenAI must explain action taken to stop AI hacking Australians’ private data, chair of federal inquiry says

From Scan to Treatment Plan, AI Helps Close Breast Cancer’s Deadliest Gaps

Breast cancer is the most commonly diagnosed cancer among American women — yet the gaps in care are wide. A majority of women over age 40 skip the recommended annual screening. Radiologists are reading more mammograms with fewer colleagues. And when a diagnosis arrives, the tests that inform treatment can take weeks to return results. […]

NVIDIA BlogSource content · Analysis pendingFrom Scan to Treatment Plan, AI Helps Close Breast Cancer’s Deadliest Gaps

Everything we launched during Birthday Week 2026

We celebrated our 16th birthday with 46 announcements across open source, post-quantum security, AI agents, and developer platform upgrades. Here’s a day-by-day roundup of everything we shipped.

Cloudflare AI BlogSource content · Analysis pendingEverything we launched during Birthday Week 2026

Building an Enterprise AI Customer Support Platform with Claude Fable 5.1 and Claude Code

Most AI coding demos stop at task managers, weather apps, or simple chatbots. For this project, we take on something more demanding: building an enterprise customer-support platform that can investigate complaints, retrieve relevant policies, recommend resolutions, and keep risky actions behind human approval. This gives us a practical way to test Claude Fable 5.1 as […] The post Building an Enterprise AI Customer Support Platform with Claude Fable 5.1 and Claude Code appeared first on Analytics Vidhya.

Analytics VidhyaSource content · Analysis pendingBuilding an Enterprise AI Customer Support Platform with Claude Fable 5.1 and Claude Code

Accelerating Evidence to Action in Pharma with Practical AI Adoption

The pharmaceutical industry is producing more clinical evidence than ever before, but the ability to turn that evidence into timely strategic decisions is struggling to keep pace. PubMed added more than 1.5 million citations in fiscal year 2023, roughly 4,300 a day, according to the National Library of Medicine. On top of that literature, medical […]

Emerj AI ResearchSource content · Analysis pendingAccelerating Evidence to Action in Pharma with Practical AI Adoption

How to Build Reliable AI Agent Systems for Production

The following article was originally published on the Agentic AI Foundation blog site and is being republished here with the author’s permission. A customer asks your company’s AI agent to update the shipping address on account 123. The agent relies on a support ticket with a typo and updates account 132 instead. The customer relationship […]

O'Reilly AI & ML RadarSource content · Analysis pendingHow to Build Reliable AI Agent Systems for Production

Our minds aren’t equipped to handle AI

Norbert Wiener, godfather of cybernetics, once said, "The thought of every age is reflected in its technique." For the past century, our thought has been reflected in our computers, including by those in the AI industry. Google's Demis Hassabis calls the brain "a biological approximation to a Turing machine." Elon Musk puts it more bluntly, declaring that "people should just think of the brain as a biological computer." (Musk's brain often worries me.) But humans are more complex than a straightforward comparison to computers gives us credit for - and far more than most of the AI industry seems to appreciate. And as their products push e … Read the full story at The Verge.

The Verge AISource content · Analysis pendingOur minds aren’t equipped to handle AI

OrgComputers

Discussion | Link

Product Hunt AISource content · Analysis pendingOrgComputers

HyperFrames Studio (Desktop)

Discussion | Link

Product Hunt AISource content · Analysis pendingHyperFrames Studio (Desktop)

OpenRUA: Robot-Use Agents Are Zero-Shot Visuomotor Policies

arXiv:2610.02459v1 Announce Type: new Abstract: Coding agents are extending their reach into the physical world by writing and executing robot control programs. One might expect the agents to use the existing mature software stack that engineers have developed over decades to access sensors and control motion. Yet prior work primarily engineers complex custom harnesses to orchestrate agents for robot use, particularly by prescribing specialized workflows and providing bespoke interfaces. This raises the question: "Is such additional harness engineering necessary?" We introduce OpenRUA, a zero-abstraction harness that bypasses bespoke abstraction layers by providing off-the-shelf coding agents with only terminal access to the robot's native software interface ROS 2. OpenRUA employs a minim…

arXiv RoboticsSource content · Analysis pendingOpenRUA: Robot-Use Agents Are Zero-Shot Visuomotor Policies

Awomo-SimDataEngine: Agentic Simulation-ReadyWorld Generation

arXiv:2610.02274v1 Announce Type: new Abstract: Generating useful robot-training data requires more than visually plausiblescenes: objects must support interaction, placements must remain physicallyvalid, and tasks must admit repeatable execution. We present\textbf{Awomo-SimDataEngine}, an agentic system that connects asset and scenegeneration to robot demonstration synthesis. Shared asset services providerigid and articulated objects, including structure-grounded part and jointgeneration with ISArt. Scene generation supports two complementary routes:Unravel reconstructs editable scenes from images, while SimForge buildssingle-room and multi-room environments from text. A graph-native harnesscoordinates construction, validation, andbounded repair, routing failures to the responsible modul…

arXiv RoboticsSource content · Analysis pendingAwomo-SimDataEngine: Agentic Simulation-ReadyWorld Generation

A Simulation-Grounded Agentic VLM Framework for Wildfire Monitoring and Reporting

arXiv:2610.02451v1 Announce Type: new Abstract: Effective wildfire monitoring requires relating visual evidence to physical fire dynamics, yet real videos with synchronized physical annotations are scarce and high-fidelity 3D simulation is costly. We present a simulation-grounded vision-language model (VLM) framework that automatically converts 2D wildfire simulations into labeled video episodes. A fixed Blender mapping produces low-detail 3D proxies aligned with simulator terrain, fuel layout, fire activity, and wind cues; controllable video generation supplies richer appearance. The proxies are intermediate representations rather than finely rendered final scenes. Generated videos and simulator labels form reusable multimodal memory for a training-free multi-agent VLM system that retrie…

arXiv Computer VisionSource content · Analysis pendingA Simulation-Grounded Agentic VLM Framework for Wildfire Monitoring and Reporting

DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

arXiv:2610.02320v1 Announce Type: new Abstract: Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a controlled way. We introduce DeskForge, a controllable desktop environment that composes and explores real applications to generate large-scale supervision for computer-use agents. It varies application states, content, window layout, appearance, and resolution, and fuses screenshots, accessibility trees, and window geometry into dense element annotations while recording the outcome of each executed action. Using this environment, we construct DeskForge-1M, a corpus of…

arXiv Computer VisionSource content · Analysis pendingDeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

APDMem: Agent-Controlled Progressive Disclosure for Query-Adaptive Long-Term Memory

arXiv:2610.02472v1 Announce Type: new Abstract: Personalized LLM assistants must recover sparse evidence from long conversation histories across queries of varying complexity. We introduce APDMem (Agent-controlled Progressive Disclosure Memory), a hierarchical long-term memory architecture that applies progressive disclosure to memory retrieval. Rather than relying on a flat memory store or fixed retrieval granularity, APDMem represents conversation history as four progressively detailed layers: thematic summaries, personalized key facts, turn-level evidence notes, and raw messages. At inference time, a controller applies progressive disclosure to the memory hierarchy: it first reads high-level summaries and drills into finer evidence only when needed. This creates an adaptive cost-fideli…

arXiv Computational LinguisticsSource content · Analysis pendingAPDMem: Agent-Controlled Progressive Disclosure for Query-Adaptive Long-Term Memory

MACTS-EM: Multi-Agent Collaborative Time Series Forecasting with Emergent Memory

arXiv:2610.02255v1 Announce Type: new Abstract: Time series forecasting remains a critical challenge across numerous domains. Despite significant advancements, existing approaches struggle with complex phenomena such as regime shifts, cross-domain knowledge transfer, and multimodal data integration. This paper introduces Multi-Agent Collaborative Time Series Forecasting with Emergent Memory (MACTS-EM), a novel framework where specialised agents collaborate to achieve superior forecasting performance. The MACTS-EM architecture integrates: (1) domain-specialised forecasting agents for pattern recognition, anomaly detection, causal inference, and uncertainty quantification; (2) a meta-cognitive layer for dynamic agent allocation; (3) an emergent memory mechanism enabling cross-domain pattern…

arXiv Machine LearningSource content · Analysis pendingMACTS-EM: Multi-Agent Collaborative Time Series Forecasting with Emergent Memory

DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents

arXiv:2610.02351v1 Announce Type: new Abstract: ReAct-based agents typically rely on a single LLM policy to propose actions, interact with the environment, and decide when a task is complete. This coupling makes action authorization and completion control difficult to enforce independently, allowing errors to propagate and unsupported completion claims to terminate execution. We introduce DeReAct, a modular agent architecture that externalizes two gating policies: a Critic that validates proposed actions before execution, and a Context Manager that reconstructs an environment-supported \textsc{State} and certifies task completion. Across GAIA and SWE-bench Verified, DeReAct improves Pass@1 most for weaker Brain models, with gains of 6.5--7.0 points for Qwen3-Coder-480B and 4.2--5.2 points…

arXiv AISource content · Analysis pendingDeReAct: Decomposed Reasoning and Acting for Reliable AI Agents

World Editing: Intervening on Executable Worlds at Increasing Depth

arXiv:2610.02331v1 Announce Type: new Abstract: Interactive world models are increasingly capable of generating environments and acting within them, yet deliberately editing an existing executable world remains underexplored. We formulate world editing as intervening on an existing world while preserving properties that should remain unchanged, and introduce intervention depth as an axis describing how strongly an edit couples world entities, dynamics, and systems. We instantiate this capability through industry-grade game modding and introduce IGMWorld, together with IGMBench, a benchmark of 110 tasks and over 1.1K executable state and behavioral criteria across Minecraft and Terraria. The tasks span property, entity, dynamics, and system interventions and are evaluated through determini…

arXiv AISource content · Analysis pendingWorld Editing: Intervening on Executable Worlds at Increasing Depth

Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

arXiv:2610.02330v1 Announce Type: new Abstract: Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more targeted feedback, but obtaining reliable step supervision often requires human or LLM judgment, or additional rollouts to estimate the downstream effect of an intermediate decision. In this paper, we argue that effective tool-use agents should estimate the long-horizon value of a possible next tool invocation before executing it. This objective requires comparative supervision over alternative invocations under the same conte…

arXiv AISource content · Analysis pendingChoosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

crosswalk

Discussion | Link

Product Hunt AISource content · Analysis pendingcrosswalk

Negotiating Ontological Boundaries in User-Authored Personal Sensing Systems

Designed artifacts are ontological, shaping, and at times limiting, what becomes possible or imaginable. One path toward mitigating such foreclosures is giving people power over how systems are designed and built. Despite decades of scholarship around systems that enable such authorship, these systems are often evaluated on whether or not they are usable, useful, or technically feasible, leaving questions of ontological boundary negotiation, unexamined. We design two open-ended probes that utilize a Wizard of Oz technique to enable the experience of training a personalized machine learning…

Apple Machine Learning ResearchSource content · Analysis pendingNegotiating Ontological Boundaries in User-Authored Personal Sensing Systems

Reviu

Discussion | Link

Product Hunt AISource content · Analysis pendingReviu

Well, if AI said it, it must be true

Dale Caldwell, former lieutenant governor for New Jersey. | Bloomberg via Getty Images New Jersey's lieutenant governor Dale Caldwell was forced to resign on September 25th after an investigation found he had sexually harassed a staffer and repeatedly violated ethics rules. The now-former Lt. governor has been making the media rounds trying to clear his name. But he took a particularly odd tactic during an interview on NJ PBS. Caldwell claims he's being unfairly targeted, and multiple AI agents back that up. He told host Rob Nelson that he "AIed" the report on the investigation "from multiple AI platforms." "I put it through AI," he said, "and it said 59 times, it [sic] said, 'What would your findings be?' There was no insta … Read the full story at The Verge.

The Verge AISource content · Analysis pendingWell, if AI said it, it must be true

Cohere and PwC partnership announcement | Cohere

PwC and Cohere announced a global alliance, launching first in Canada, to help organizations deploy secure enterprise AI and turn AI investment into measurable business value. The alliance brings together North, Cohere’…

Cohere BlogSource content · Analysis pendingCohere and PwC partnership announcement | Cohere

North 2: Enterprise AI Without Compromises

Key takeaways New agent harness: North 2 has a redesigned orchestration system enabling more sophisticated, reliable, multi-step automations and agents across the enterprise. Increased intelligence: Build smarter workfl…

Cohere BlogSource content · Analysis pendingNorth 2: Enterprise AI Without Compromises
Policy

Sen. Adam Schiff on AI regulation, free speech, and impeaching Trump one more time

Today, I’m talking with Senator Adam Schiff, a Democrat from California. Sen. Schiff sits on a number of committees with oversight into tech and AI: intellectual property, antitrust, privacy and technology — it’s all there. I really wanted to ask him about how we might regulate anything related to the tech industry at this moment in time. Verge subscribers, don’t forget you get exclusive access to ad-free Decoder wherever you get your podcasts. Head here. Not a subscriber? You can sign up here. But as you’ll hear, he started our conversation by talking about the self-dealing and corruption present all through our politics. That of course fell against the backdrop of President Trump gathering AI CEOs to the White House to sign a non-binding pact in which they agree to, you know, do a good…

The Verge AISource content · Analysis pendingSen. Adam Schiff on AI regulation, free speech, and impeaching Trump one more time

Accept ‘bad things’ in return for benefits of AI, says Sam Altman

Boss of OpenAI calls for a regulatory light touch because of the ‘good stuff’ the technology can deliver Sam Altman says he believes the world should accept “bad things” happening with AI in exchange for the benefits of the technology. The chief executive of OpenAI cited hacks, scams and “other bad things that will happen” in an interview that sparked an instant backlash from critics of the major AI companies. His comments came after one of his company’s safety experts resigned, saying at the weekend that its “culture is broken”. Continue reading...

The Guardian AISource content · Analysis pendingAccept ‘bad things’ in return for benefits of AI, says Sam Altman

The UK’s G20 presidency will push Andy Burnham to the centre of the global stage. What will he use it for? | Michael Jacobs

This is a chance for the PM to advance his progressive agenda and tackle urgent issues such as AI regulation and climate action Among the many crucial decisions Andy Burnham has to make over the next few weeks, one has gone almost completely unremarked: what Britain’s priorities will be when it assumes the presidency of the G20 next year. Burnham has said he wants to concentrate on domestic policy. But with the world in turmoil, chairing the group of the world’s largest economies will be one of his most important responsibilities. Michael Jacobs is professor of political economy at the University of Sheffield and a former energy and climate change adviser to Gordon Brown when he was prime minister Continue reading...

The Guardian AISource content · Analysis pendingThe UK’s G20 presidency will push Andy Burnham to the centre of the global stage. What will he use it for? | Michael Jacobs

Confidence-Controlled XAI Auditing for Pedestrian Detection under Domain Shift

arXiv:2610.02364v1 Announce Type: new Abstract: Explainability is increasingly required for perception models in intelligent vehicles, yet whether explanations remain faithful under driving domain shift is still poorly understood. This work audits post-hoc explanations of a fixed YOLOv8s pedestrian detector across PIE and JAAD using ROI-based D-Deletion, frozen confidence terciles, rank-based tests, bootstrap intervals, and Holm correction. The audit shows that deletion-based faithfulness is strongly coupled to detection strength at explanation time, with Spearman correlations between 0.70 and 0.82 for D-RISE, making naive confidence-stratified comparisons unreliable. After controlling for detection strength within fixed f0 bins, D-RISE faithfulness remains domain-dependent in the central…

arXiv Computer VisionSource content · Analysis pendingConfidence-Controlled XAI Auditing for Pedestrian Detection under Domain Shift

The Price of Greenwashing: Algorithmic Verification and Market Discipline using Conformal Machine Learning

arXiv:2610.02225v1 Announce Type: new Abstract: While corporate sustainability mandates are expanding, the systemic reliance on self-reported emissions data exposes financial markets to pervasive greenwashing. Current literature relies heavily on subjective ESG ratings or textual sentiment analysis, leaving a critical econometric gap in objectively quantifying physical climate realities. To resolve this information asymmetry, we fuse U.S. SEC financial fundamentals with facility-level EPA greenhouse gas registries to establish a mathematically guaranteed baseline of physical corporate emissions. Leveraging a gradient boosting architecture and Mondrian Conformal Prediction, we quantify the shortfall between self-reported data and this algorithmic baseline into a novel Conformal-Weighted Co…

arXiv Machine LearningSource content · Analysis pendingThe Price of Greenwashing: Algorithmic Verification and Market Discipline using Conformal Machine Learning

Keep It CALM: Analyzing the Limits of Global Unsafety in Text-to-Image Generation

arXiv:2610.02300v1 Announce Type: new Abstract: Training-free safeguards for text-to-image generation often rely on a reusable safety signal, such as an unsafe direction or global toxic subspace, applied broadly across prompts. We provide a controlled geometric analysis of this global-unsafety assumption and reveal a consistent coverage-selectivity trade-off: compact unsafe subspaces fail to cover heterogeneous unsafe semantics, whereas broader aggregation increasingly distorts safety-adjacent benign prompts. Motivated by this finding, we propose CALM (Counterfactual Adaptive Local Modulation), a training-free safeguard that replaces uniform global removal with prompt-local counterfactual correction. Using matched unsafe-benign anchors, CALM routes each prompt to active unsafe categories,…

arXiv AISource content · Analysis pendingKeep It CALM: Analyzing the Limits of Global Unsafety in Text-to-Image Generation
Robotics

Physical AI moves beyond traditional robotics

Physical AI is moving beyond traditional robotics into real enterprise tasks, but enterprises still face big challenges integrating and scaling robotic systems.

AI BusinessSource content · Analysis pendingPhysical AI moves beyond traditional robotics

NEEDLEWORK: Offline Rewriting of Robot Data with Verified Local Stitches

arXiv:2610.02339v1 Announce Type: new Abstract: Robot demonstrations may contain useful behavior even when individual episodes are inefficient or unsuccessful. Trajectory stitching offers a way to compose these behaviors into improved training data, but identifying useful connections and verifying their feasibility is difficult in high-dimensional robot data, where many prior methods rely on low-dimensional state representations. We introduce NEEDLE, an offline dataset-augmentation algorithm that addresses these challenges by adding short, verified action bridges between recorded observations in high-dimensional robot demonstrations. First, NEEDLE identifies and creates connections that bypass suboptimal detours, broaden action coverage, and augment the original dataset with failed trajec…

arXiv RoboticsSource content · Analysis pendingNEEDLEWORK: Offline Rewriting of Robot Data with Verified Local Stitches
Tools

An open-source tool lets you delete 12GB of Apple Intelligence data on macOS

Getting some extra storage space on your Mac could be as easy as deleting Apple's AI features with a new open-source command line tool called RemoveMacAI. There used to be a single Settings toggle for disabling Apple Intelligence, but that was removed in macOS 27. Now, the models stay on your disk, and the features are split up across different settings panes. As spotted earlier by MacRumors, the RemoveMacAI developer said in a post on Reddit that the tool essentially does what the old settings toggle did, deleting the models and turning the features off: You have to find about a dozen settings, and several of them are under Screen Time. … Read the full story at The Verge.

The Verge AISource content · Analysis pendingAn open-source tool lets you delete 12GB of Apple Intelligence data on macOS

OpenAI is sticking more ads in ChatGPT

OpenAI's latest ad format will put images of sponsored products and services on your screen. The ads, which OpenAI will begin testing in the US later this month, will "initially" appear when you generate images with ChatGPT, according to an announcement on Monday. The company first brought ads to ChatGPT in February. Until now, ChatGPT's ads only showed a company's name, logo, and a link to its product in a "sponsored" section beneath your chat. OpenAI's visual ad format will similarly remain separate from the image you generate and will "not influence the answers ChatGPT provides." Ads won't appear if you subscribe to ChatGPT's Plus, Pro, … Read the full story at The Verge.

The Verge AISource content · Analysis pendingOpenAI is sticking more ads in ChatGPT

Rill Browser

Discussion | Link

Product Hunt AISource content · Analysis pendingRill Browser

Internet Watch Foundation reports huge rise in AI child sexual abuse material

Abuse material monitor says number of AI images assessed this year is already 40% higher than last year’s total AI-generated child sexual abuse material is proliferating online, with the amount of illegal material investigated this year already exceeding the total for 2025. Analysts at the Internet Watch Foundation have found more photorealistic child sexual abuse material in the first half of 2026 than for the whole of the prior year. The UK-based IWF, which monitors abuse material globally, said it had assessed 6,310 AI images that met the legal definition of child sexual abuse, 40% higher than last year’s total of more than 4,500. Continue reading...

The Guardian AISource content · Analysis pendingInternet Watch Foundation reports huge rise in AI child sexual abuse material

Lecta

Discussion | Link

Product Hunt AISource content · Analysis pendingLecta

Floani

Discussion | Link

Product Hunt AISource content · Analysis pendingFloani

Marv

Discussion | Link

Product Hunt AISource content · Analysis pendingMarv

Pilot5 Legal

Discussion | Link

Product Hunt AISource content · Analysis pendingPilot5 Legal
Chips

State-Space Unlearning for Non-Stationary Bias in Land Surface Forecasting

arXiv:2610.02248v1 Announce Type: new Abstract: Operational land surface forecasting systems built on Mamba-family Structured State Space Models absorb non-stationary confounding events (unrecorded irrigation booms, dam-operation shifts, sensor recalibrations) into their state-transition matrices, silently biasing NDVI, LST, and crop phenology predictions long after the physical cause ends. This paper introduces SSU-LSF (State-Space Unlearning for Land Surface Forecasting), the first machine-unlearning framework purpose-built for geoscientific Mamba-based SSMs. We develop EKFac influence functions specialized to the Mamba state matrices via a closed-form matrix-exponential gradient, use spectral-radius-weighted elbow thresholding to localize a temporal confounding footprint $\Phi$, and ap…

arXiv Machine LearningSource content · Analysis pendingState-Space Unlearning for Non-Stationary Bias in Land Surface Forecasting