Chatr is an AI tool that handles customer support, collects feedback, books appointments, and sells products.
Live AI News Intelligence
Live monitoring
Live updates
Trusted sources, attribution, rights, and in-site reading distilled into a signal-first AI brief.
Live updates
This paper presents a modality-adaptive contactless respiratory rate monitoring framework for heterogeneous mobile robots with onboard edge computing. It combines brightness-adaptive sensor selection across RGB, thermal, NIR, and low-light cameras, keypoint-guided chest ROI extraction, and SQI-based filtering. Experiments on quadruped and wheeled platforms demonstrate generalization without per-platform retuning: RGB covers up to 8m, NIR up to 6m, thermal short-range only, low-light effective in complete darkness up to 8m. The framework supports autonomous triage and victim assessment in hazardous environments.
A new study proposes a transformer-based warm-start method for sequential convex programming (SCP) in the terminal approach of a space manipulator to a tumbling target, reducing iterations by 28% and runtime by 23% while preserving optimal control cost.
This paper introduces APOLLO, a hybrid framework that combines a lightweight personalized embedding model with selective large language model (LLM) assistance for abstention-aware object rearrangement in cluttered, partially erroneous environments. The framework uses uncertainty estimates to decide when to invoke the LLM, balancing efficiency, privacy, and reasoning. A new synthetic dataset, APOR, is also introduced. Experiments show APOLLO reduces LLM usage while maintaining or improving performance over prior methods.
To overcome the scarcity of paired vision-action data in robotics, researchers propose CAIP, a vision encoder that uses hand poses from egocentric human video as proxy for robot actions. Using only 88 hours of robot data and 32,041 hours of human video, CAIP outperforms leading encoders by over 30% on dexterous manipulation tasks.
ACE-Ego-0 is a unified Vision-Language-Action (VLA) pretraining framework that integrates egocentric human videos with robot data by converting human videos into pseudo-action trajectories and using a reliability-aware training objective, boosting model performance on benchmarks and real-world tasks.
This paper presents VL-MemKnG, a hybrid memory framework that combines a spatio-temporal knowledge graph with persistent segment-level contextual memory for question answering over long egocentric navigation videos. It improves Top-1 retrieval accuracy from 58% to 67% and Recall@1 from 34.50% to 40.55% on the WalkieKnowledgeT+ benchmark, outperforming methods including Gemini 2.5 Pro and Qwen 3.5+.
This paper proposes ParkingTransformer, a novel framework that leverages multi-view perception and scene understanding capability of Large Language Models (LLMs) for end-to-end autonomous parking. By combining trajectory queries with LLMs implicit state features, it outputs planning trajectories directly, eliminating dense BEV representations. It introduces 3D positional encoding, a fixed-window streaming mechanism, and a coarse-to-fine decoding strategy. Experiments on CARLA simulator achieve a driving score of 61.32, and real-world experiments show an average success rate of 88.70%.
HRDX is a large-scale dataset for vector HD map construction, spanning about 40 hours (1,400 km) of driving data with rich semantic annotations and aerial orthoimagery, aimed at advancing autonomous driving research.
A preliminary approach using large language models (LLMs) to automatically generate robot semantic abstractions by transforming URDF models into populated ontologies, with majority voting and validation to ensure reliability.
SierpinskiCam enhances video retaking from a single monocular video by augmenting geometry-based guidance with Sierpinski dome texture cues, enabling robust performance even under large viewpoint changes.
We propose OR3, a method that converts OR video clips into action-driven digital twins (ActDT) and uses an LLM to generate hypothetical ActDTs from queries for imagination-based retrieval, achieving 57.6% R@1 and 77.3% R@5 on a benchmark of 276 implicit queries.
Unified multimodal models (UMMs) exhibit modality imbalance during instruction tuning, where language gradients dominate, degrading image generation quality. This paper proposes Pareto LoRA, a Pareto-optimal gradient integration strategy that balances text and image objectives by modulating gradient direction and strength. Experiments on Emu2 show up to 44.9% improvements in perceptual image quality while maintaining text performance.
Existing surgical video question answering methods compress videos into discrete tokens and couple perception with reasoning, limiting multi-step reasoning. This paper introduces a reinforcement learning framework that trains LLMs to operate over digital twin representations, decoupling perception from reasoning. It introduces hierarchical representations and a novel reward, and presents the REAL-Colon-Reason benchmark, achieving state-of-the-art performance on multiple benchmarks.
REINS is a training-free method that aligns video diffusion models at inference time by steering internal representations toward safe generation. It uses Supervised PCA to find a single direction separating safe from unsafe trajectories, applied at intermediate transformer layers with negligible overhead. Evaluated on 9 models, it is the broadest safety evaluation in video generation literature.
GeoDisaster is a new benchmark for operational disaster geo-intelligence, containing 2,921 instances across 43 question types and five task families (deforestation monitoring, multi-hazard analysis, building-damage assessment, flood-safe routing, and Sentinel-1 SAR flood monitoring). It integrates heterogeneous EO/GIS data and uses executable workflows for ground truth. The paper also proposes an orchestrated multi-agent framework with 18 disaster tools and Role-Contract Expectation Alignment (RCEA) to improve tool use and decision making.
This study presents the first successful application of vision transformers for coastal algal bloom mapping using 30-meter Landsat-Sentinel-2 imagery. A globally distributed bloom patch dataset was created, and four transformer architectures were compared against a convolutional baseline. The Swin Transformer outperformed traditional spectral indices under cloud and glint stress, reducing false positives. The findings support deep learning as a reliable tool for medium-resolution algal bloom monitoring in dynamic coastal environments.
Research reveals benchmarks overstate edge AI performance by 20-30% in real deployment. Edge-TSR, a continuous inference system on NVIDIA Jetson Orin Nano, integrates detection, tracking, and a lightweight temporal stabilization mechanism, recovering up to 10.16% accuracy over per-frame baselines. A 55-minute vehicular test achieves sustained 16.18 FPS within safe thermal limits without cloud offload.
A new framework combining multi-scale CNN, spectral attention, BiSpectral Mamba, and quantum-inspired learning achieves 84.83% accuracy in hyperspectral crop classification, addressing challenges like high dimensionality and class imbalance.
A new study reveals that current multilingual evaluations for Vision-Language Models (VLMs) overlook users of multi-script languages. The authors introduce the Punjabi Multimodal Visual Reasoning (PuMVR) benchmark, consisting of 1,000 strictly parallel image-text instances across Punjabi's three active scripts: Gurmukhi, Shahmukhi, and Roman. Evaluating 10 state-of-the-art VLMs, they uncover a systematic 'Script Gap', where models succeed in one script but fail in another, with accuracy differences up to 16%. Visual input boosts performance but does not close the gap, and cross-script in-context transfer is brittle. They propose the Script Consistency Rate (SCR), which can be as low as 24.8%, as a mandatory metric for script-agnostic evaluation.
A study tests Word2Vec on Toki Pona, a constructed language with only ~130 words, finding that its effectiveness depends more on distributional patterns than vocabulary size.
This paper addresses the issue of language misidentification in LLM-based ASR, proposing a soft prompting approach, defining a language adherence metric, and evaluating three mitigation strategies: zero-shot prompting, supervised fine-tuning, and chain-of-thought reasoning.
This paper describes the MLLP-VRAIN UPV system for IWSLT 2026 Simultaneous Speech Translation, using Parakeet and Qwen 3.5 models with adaptive black-box policies to improve quality-latency trade-offs. The system participates in all language directions and introduces a context track for En→De, It, Zh using ASR word-boosting and RAG. Results show a +5.82 XCOMET-XL improvement on MCIF En→De, with additional +1.03 from context processing.
This study investigates parameter-efficient adaptation strategies for volumetric CT report generation and introduces RAD3D-Prefix, a lightweight diagnostic-prior conditioning framework. By freezing the LLM and training only projection layers, the method minimizes trainable parameters and overfitting. Systematic experiments across LLMs from 96.1M to 1.6B parameters reveal that fine-tuning benefits smaller LLMs, while freezing larger ones (1B+) offers a superior trade-off. RAD3D-Prefix outperforms baselines on automatic metrics and clinical reader studies with strong out-of-domain generalization.
This paper addresses the training-inference mismatch in the token-to-token (T2T) editor of LLaDA2.1 diffusion language models. The proposed self-generated T2T method performs a no-gradient draft pass, fills masked positions with predicted tokens, and supervises recovery in a second pass under self-generated corruptions. Implemented via LoRA continued-pretraining on LLaDA2.1-mini, it improves accuracy while reducing edit intensity, mitigating failure modes like final-digit transcription errors and excessive self-correction.
This study investigates whether parasocial interaction cues exist in online communities where both sides are autonomous AI agents. Analyzing 4,434 posts and 50,338 comments from Moltbook using keyword matching, few-shot LLM annotation, and grouped-context LLM annotation, the authors found that PSI colloquial cues are prevalent and strongly associated with OP re-engagement and reciprocal reply structures. A dyadic persistence test further confirms reciprocity bids aligned with sustained OP-involving mutual recurrence, providing empirical evidence bridging interaction-level PSI scripts with relationship-level persistent dyadic patterns.
RepSelect is a new LLM unlearning method that isolates forget-set-specific representations by collapsing top principal components of weight gradients, achieving 4-50x better resistance to reversal than existing methods.
PromptMN is a lightweight domain-specific language that annotates natural language prompts with %-prefixed directives, making roles, goals, and constraints explicit to reduce ambiguity in agentic workflows. It bridges informal prompting and pseudocode, supports reverse prompt engineering, and has been validated on frontier models without fine-tuning.
MemSlides proposes a hierarchical memory framework separating long-term memory (user profile and tool memory) from working memory, combined with scoped slide-local revision, to maintain user preferences across tasks and reliably carry out localized edits over multiple turns. Experiments show improvements in persona alignment, modification behavior, and preference carryover.
This research addresses concurrency anomalies in multi-agent LLM systems by formalizing four anomaly types using TLA+ and building a mechanically verified consistency hierarchy L0-L4. With 274 Verus proof obligations, the detectors are proven sound and complete. Three Rust runtimes implement L0-L1, and the work reproduces real-world anomalies in ByteDance's deer-flow and LangGraph, providing verified fixes.