AI News HubLIVE

Live AI News Intelligence

Live monitoring

The most important shift in AI today

Distilled from 105 trusted sources. Last update 2026-06-17 04:09 UTC.

Live monitoring

Live updates

Trusted sources, attribution, rights, and in-site reading distilled into a signal-first AI brief.

Live updates

Reset
Contactless Respiratory Monitoring on Heterogeneous Mobile Robots: A Multimodal Edge-Computing Framework

This paper presents a modality-adaptive contactless respiratory rate monitoring framework for heterogeneous mobile robots with onboard edge computing. It combines brightness-adaptive sensor selection across RGB, thermal, NIR, and low-light cameras, keypoint-guided chest ROI extraction, and SQI-based filtering. Experiments on quadruped and wheeled platforms demonstrate generalization without per-platform retuning: RGB covers up to 8m, NIR up to 6m, thermal short-range only, low-light effective in complete darkness up to 8m. The framework supports autonomous triage and victim assessment in hazardous environments.

arXiv RoboticsModels / Policy / ResearchIn-site article
Abstention-Aware Personalized Object Rearrangement via Uncertainty-Guided LLM Assistance

This paper introduces APOLLO, a hybrid framework that combines a lightweight personalized embedding model with selective large language model (LLM) assistance for abstention-aware object rearrangement in cluttered, partially erroneous environments. The framework uses uncertainty estimates to decide when to invoke the LLM, balancing efficiency, privacy, and reasoning. A new synthetic dataset, APOR, is also introduced. Experiments show APOLLO reduces LLM usage while maintaining or improving performance over prior methods.

arXiv RoboticsModels / Research / RoboticsIn-site article
Contrastive Action-Image Pre-training for Visuomotor Control

To overcome the scarcity of paired vision-action data in robotics, researchers propose CAIP, a vision encoder that uses hand poses from egocentric human video as proxy for robot actions. Using only 88 hours of robot data and 32,041 hours of human video, CAIP outperforms leading encoders by over 30% on dexterous manipulation tasks.

arXiv RoboticsModels / Agents / ResearchIn-site article
ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining

ACE-Ego-0 is a unified Vision-Language-Action (VLA) pretraining framework that integrates egocentric human videos with robot data by converting human videos into pseudo-action trajectories and using a reliability-aware training objective, boosting model performance on benchmarks and real-world tasks.

arXiv RoboticsModels / Research / RoboticsIn-site article
VL-MemKnG: Hybrid Memory with a Spatio-Temporal Knowledge Graph for Question Answering over Long Egocentric Navigation Trajectories

This paper presents VL-MemKnG, a hybrid memory framework that combines a spatio-temporal knowledge graph with persistent segment-level contextual memory for question answering over long egocentric navigation videos. It improves Top-1 retrieval accuracy from 58% to 67% and Recall@1 from 34.50% to 40.55% on the WalkieKnowledgeT+ benchmark, outperforming methods including Gemini 2.5 Pro and Qwen 3.5+.

arXiv RoboticsModels / Research / RoboticsIn-site article
ParkingTransformer: LLM-Enhanced End-to-End Trajectory Planning for Autonomous Parking

This paper proposes ParkingTransformer, a novel framework that leverages multi-view perception and scene understanding capability of Large Language Models (LLMs) for end-to-end autonomous parking. By combining trajectory queries with LLMs implicit state features, it outputs planning trajectories directly, eliminating dense BEV representations. It introduces 3D positional encoding, a fixed-window streaming mechanism, and a coarse-to-fine decoding strategy. Experiments on CARLA simulator achieve a driving score of 61.32, and real-world experiments show an average success rate of 88.70%.

arXiv RoboticsModels / Research / RoboticsIn-site article
HRDX: A Large-Scale Vector HD-Map Dataset

HRDX is a large-scale dataset for vector HD map construction, spanning about 40 hours (1,400 km) of driving data with rich semantic annotations and aerial orthoimagery, aimed at advancing autonomous driving research.

arXiv RoboticsModels / Research / StartupsIn-site article
Pareto LoRA: Mitigating Modality Imbalance in Unified Multimodal Models via Pareto-Optimal Gradient Integration

Unified multimodal models (UMMs) exhibit modality imbalance during instruction tuning, where language gradients dominate, degrading image generation quality. This paper proposes Pareto LoRA, a Pareto-optimal gradient integration strategy that balances text and image objectives by modulating gradient direction and strength. Experiments on Emu2 show up to 44.9% improvements in perceptual image quality while maintaining text performance.

arXiv Computer VisionModels / ResearchIn-site article
Training LLMs with Reinforcement Learning over Digital Twin Representations for Reasoning-Intensive Surgical VideoQA

Existing surgical video question answering methods compress videos into discrete tokens and couple perception with reasoning, limiting multi-step reasoning. This paper introduces a reinforcement learning framework that trains LLMs to operate over digital twin representations, decoupling perception from reasoning. It introduces hierarchical representations and a novel reward, and presents the REAL-Colon-Reason benchmark, achieving state-of-the-art performance on multiple benchmarks.

arXiv Computer VisionModels / Research / StartupsIn-site article
Pulling The REINS: Training-Free Safety Alignment of Video Diffusion Models via Representation Steering

REINS is a training-free method that aligns video diffusion models at inference time by steering internal representations toward safe generation. It uses Supervised PCA to find a single direction separating safe from unsafe trajectories, applied at intermediate transformer layers with negligible overhead. Evaluated on 9 models, it is the broadest safety evaluation in video generation literature.

arXiv Computer VisionModels / Policy / ResearchIn-site article
GeoDisaster: Benchmarking Orchestrated Agents for Operational Disaster Geo-Intelligence

GeoDisaster is a new benchmark for operational disaster geo-intelligence, containing 2,921 instances across 43 question types and five task families (deforestation monitoring, multi-hazard analysis, building-damage assessment, flood-safe routing, and Sentinel-1 SAR flood monitoring). It integrates heterogeneous EO/GIS data and uses executable workflows for ground truth. The paper also proposes an orchestrated multi-agent framework with 18 disaster tools and Role-Contract Expectation Alignment (RCEA) to improve tool use and decision making.

arXiv Computer VisionModels / Agents / ResearchIn-site article
Landsat-Sentinel-2 Algal Bloom Mapping Using Vision Transformers: Model Description, Implementation, and Examples

This study presents the first successful application of vision transformers for coastal algal bloom mapping using 30-meter Landsat-Sentinel-2 imagery. A globally distributed bloom patch dataset was created, and four transformer architectures were compared against a convolutional baseline. The Swin Transformer outperformed traditional spectral indices under cloud and glint stress, reducing false positives. The findings support deep learning as a reliable tool for medium-resolution algal bloom monitoring in dynamic coastal environments.

arXiv Computer VisionModels / ResearchIn-site article
Beyond Benchmarks: Continuous Edge Inference for Fine-Grained Roadside Perception

Research reveals benchmarks overstate edge AI performance by 20-30% in real deployment. Edge-TSR, a continuous inference system on NVIDIA Jetson Orin Nano, integrates detection, tracking, and a lightweight temporal stabilization mechanism, recovering up to 10.16% accuracy over per-frame baselines. A 55-minute vehicular test achieves sustained 16.18 FPS within safe thermal limits without cloud offload.

arXiv Computer VisionChips / ResearchIn-site article
Not Truly Multilingual: Script Consistency as a Missing Dimension in VLM Evaluation

A new study reveals that current multilingual evaluations for Vision-Language Models (VLMs) overlook users of multi-script languages. The authors introduce the Punjabi Multimodal Visual Reasoning (PuMVR) benchmark, consisting of 1,000 strictly parallel image-text instances across Punjabi's three active scripts: Gurmukhi, Shahmukhi, and Roman. Evaluating 10 state-of-the-art VLMs, they uncover a systematic 'Script Gap', where models succeed in one script but fail in another, with accuracy differences up to 16%. Visual input boosts performance but does not close the gap, and cross-script in-context transfer is brittle. They propose the Script Consistency Rate (SCR), which can be as low as 24.8%, as a mandatory metric for script-agnostic evaluation.

arXiv Computer VisionModels / Research / StartupsIn-site article
Examining the Limits of Word2Vec with Toki Pona

A study tests Word2Vec on Toki Pona, a constructed language with only ~130 words, finding that its effectiveness depends more on distributional patterns than vocabulary size.

arXiv Computational LinguisticsModels / Research / StartupsIn-site article
Are you speaking my languages? On spoken language adherence in multimodal LLMs

This paper addresses the issue of language misidentification in LLM-based ASR, proposing a soft prompting approach, defining a language adherence metric, and evaluating three mitigation strategies: zero-shot prompting, supervised fine-tuning, and chain-of-thought reasoning.

arXiv Computational LinguisticsModels / ResearchIn-site article
MLLP-VRAIN UPV system for the IWSLT 2026 Simultaneous Speech Translation task

This paper describes the MLLP-VRAIN UPV system for IWSLT 2026 Simultaneous Speech Translation, using Parakeet and Qwen 3.5 models with adaptive black-box policies to improve quality-latency trade-offs. The system participates in all language directions and introduces a context track for En→De, It, Zh using ASR word-boosting and RAG. Results show a +5.82 XCOMET-XL improvement on MCIF En→De, with additional +1.03 from context processing.

arXiv Computational LinguisticsModels / ResearchIn-site article
Revisiting LLM Adaptation for 3D CT Report Generation: A Study of Scaling and Diagnostic Priors

This study investigates parameter-efficient adaptation strategies for volumetric CT report generation and introduces RAD3D-Prefix, a lightweight diagnostic-prior conditioning framework. By freezing the LLM and training only projection layers, the method minimizes trainable parameters and overfitting. Systematic experiments across LLMs from 96.1M to 1.6B parameters reveal that fine-tuning benefits smaller LLMs, while freezing larger ones (1B+) offers a superior trade-off. RAD3D-Prefix outperforms baselines on automatic metrics and clinical reader studies with strong out-of-domain generalization.

arXiv Computational LinguisticsModels / ResearchIn-site article
Self-Generated Error Training for Token Editing in Diffusion Language Models

This paper addresses the training-inference mismatch in the token-to-token (T2T) editor of LLaDA2.1 diffusion language models. The proposed self-generated T2T method performs a no-gradient draft pass, fills masked positions with predicted tokens, and supervises recovery in a second pass under self-generated corruptions. Implemented via LoRA continued-pretraining on LLaDA2.1-mini, it improves accuracy while reducing edit intensity, mitigating failure modes like final-digit transcription errors and excessive self-correction.

arXiv Computational LinguisticsModels / ResearchIn-site article
From Parasocial Scripts to Dyadic Persistence in Autonomous AI-Agent Communities

This study investigates whether parasocial interaction cues exist in online communities where both sides are autonomous AI agents. Analyzing 4,434 posts and 50,338 comments from Moltbook using keyword matching, few-shot LLM annotation, and grouped-context LLM annotation, the authors found that PSI colloquial cues are prevalent and strongly associated with OP re-engagement and reciprocal reply structures. A dyadic persistence test further confirms reciprocity bids aligned with sustained OP-involving mutual recurrence, providing empirical evidence bridging interaction-level PSI scripts with relationship-level persistent dyadic patterns.

arXiv Computational LinguisticsModels / Agents / ResearchIn-site article
RepSelect: Robust LLM Unlearning via Representation Selectivity

RepSelect is a new LLM unlearning method that isolates forget-set-specific representations by collapsing top principal components of weight gradients, achieving 4-50x better resistance to reversal than existing methods.

arXiv Computational LinguisticsModels / ResearchIn-site article
PromptMN: Pseudo Prompting Language

PromptMN is a lightweight domain-specific language that annotates natural language prompts with %-prefixed directives, making roles, goals, and constraints explicit to reduce ambiguity in agentic workflows. It bridges informal prompting and pseudocode, supports reverse prompt engineering, and has been validated on frontier models without fine-tuning.

arXiv Computational LinguisticsModels / Agents / ResearchIn-site article
MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision

MemSlides proposes a hierarchical memory framework separating long-term memory (user profile and tool memory) from working memory, combined with scoped slide-local revision, to maintain user preferences across tasks and reliably carry out localized edits over multiple turns. Experiments show improvements in persona alignment, modification behavior, and preference carryover.

arXiv Computational LinguisticsAgents / ResearchIn-site article
Verified Detection and Prevention of Concurrency Anomalies in Multi-Agent Large Language Model Systems

This research addresses concurrency anomalies in multi-agent LLM systems by formalizing four anomaly types using TLA+ and building a mechanically verified consistency hierarchy L0-L4. With 274 Verus proof obligations, the detectors are proven sound and complete. Three Rust runtimes implement L0-L1, and the work reproduces real-world anomalies in ByteDance's deer-flow and LangGraph, providing verified fixes.

arXiv Machine LearningModels / Agents / ResearchIn-site article