跳到主要內容
AI News HubLIVE

AI 新聞即時情報

即時監測

今日 AI 世界最重要的變化

來源目錄收錄 105 個來源 · 本頁最新發布時間 2026-10-05 12:00 UTC+8。

部分報道的翻譯與分析尚未完成,已標示來源內容。精選優先顯示處理完成的報道。

即時監測

即時更新

即時追蹤可信來源,保留出處、權限和站內閱讀模式,把噪音壓成可讀情報。

即時更新

重設

2026-10-05

待翻譯:NEEDLEWORK: Offline Rewriting of Robot Data with Verified Local Stitches

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02339v1 Announce Type: new Abstract: Robot demonstrations may contain useful behavior even when individual episodes are inefficient or unsuccessful. Trajectory stitching offers a way to compose these behaviors into improved training data, but identifying useful connections and verifying their feasibility is difficult in high-dimensional robot data, where many prior methods rely on low-dimensional state representations. We introduce NEEDLE, an offline dataset-augmentation algorithm that addresses these challenges by adding short, verified action bridges between recorded observations in high-dimensional robot demonstrations. First, NEEDLE identifies and creates connections that bypass suboptimal detours, broaden action coverage, and augment the origina…

arXiv Robotics機械人 / 研究來源內容 · 翻譯待補全待翻譯:NEEDLEWORK: Offline Rewriting of Robot Data with Verified Local Stitches
待翻譯:SoTa: Soft Tactile Skins for Dexterous Manipulation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02338v1 Announce Type: new Abstract: A growing body of work suggests that tactile sensing gives robot policies contact information that complements vision in dexterous manipulation. However, visuo-tactile robot data remains scarce: dexterous demonstrations require teleoperating robots, which limits dataset scale. Human demonstrations are far cheaper to collect and offer a path to scale this data, but only if human and robot hands carry tactile sensors with corresponding signals. This requires sensors that conform to different hand geometries, cover the full hand, and share a common layout across embodiments. We present SoTa, a low-cost capacitive tactile skin that provides full-hand coverage on humans and robots while preserving a shared layout of 20…

arXiv Robotics研究 / 創業融資 / 機械人來源內容 · 翻譯待補全待翻譯:SoTa: Soft Tactile Skins for Dexterous Manipulation
待翻譯:World-Calibrated Proposal-to-Action Flow for Vision-Language-Action Models

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02323v1 Announce Type: new Abstract: Flow-based Vision-Language-Action (VLA) policies generate action chunks by transporting samples from a task-agnostic isotropic Gaussian source. As this source is conditioned on neither recent execution nor predicted future evolution, (i) it discards the local continuity established by recently executed motion. (ii) Even when predictive world representations are introduced, they often only condition the transport dynamics rather than determine where generation starts, how far it may deviate, or along which action directions it may expand. Building on this observation, we introduce ProAct, a world-calibrated proposal-to-action framework that makes the generative source itself predictable. (i) To preserve motion cont…

arXiv Robotics模型 / 研究 / 機械人來源內容 · 翻譯待補全待翻譯:World-Calibrated Proposal-to-Action Flow for Vision-Language-Action Models
待翻譯:Awomo-SimDataEngine: Agentic Simulation-ReadyWorld Generation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02274v1 Announce Type: new Abstract: Generating useful robot-training data requires more than visually plausiblescenes: objects must support interaction, placements must remain physicallyvalid, and tasks must admit repeatable execution. We present\textbf{Awomo-SimDataEngine}, an agentic system that connects asset and scenegeneration to robot demonstration synthesis. Shared asset services providerigid and articulated objects, including structure-grounded part and jointgeneration with ISArt. Scene generation supports two complementary routes:Unravel reconstructs editable scenes from images, while SimForge buildssingle-room and multi-room environments from text. A graph-native harnesscoordinates construction, validation, andbounded repair, routing failu…

arXiv RoboticsAgent / 研究 / 創業融資來源內容 · 翻譯待補全待翻譯:Awomo-SimDataEngine: Agentic Simulation-ReadyWorld Generation
待翻譯:A Simulation-Grounded Agentic VLM Framework for Wildfire Monitoring and Reporting

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02451v1 Announce Type: new Abstract: Effective wildfire monitoring requires relating visual evidence to physical fire dynamics, yet real videos with synchronized physical annotations are scarce and high-fidelity 3D simulation is costly. We present a simulation-grounded vision-language model (VLM) framework that automatically converts 2D wildfire simulations into labeled video episodes. A fixed Blender mapping produces low-detail 3D proxies aligned with simulator terrain, fuel layout, fire activity, and wind cues; controllable video generation supplies richer appearance. The proxies are intermediate representations rather than finely rendered final scenes. Generated videos and simulator labels form reusable multimodal memory for a training-free multi-…

arXiv Computer VisionAgent / 模型 / 研究來源內容 · 翻譯待補全待翻譯:A Simulation-Grounded Agentic VLM Framework for Wildfire Monitoring and Reporting
待翻譯:An AI-Based Multi-Stage Approach for Androgenetic Alopecia Assessment from Low-Magnification Scalp Images

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02421v1 Announce Type: new Abstract: Androgenetic alopecia (AGA) is characterized by patterned follicular miniaturization, increased single-hair follicular units, and altered hair-shaft diameter. We present an automated quantitative scalp-analysis and clinical decision-support framework combining FU localization, ordinal visible-shaft counting, calibrated shaft-width estimation, regional aggregation, and an interpretable rule layer. The clinical cohort comprised 243 patients (127 AGA, 116 non-AGA), while the computer-vision experiments used 160 expert-annotated patients, 2,400 trichoscopic images, and approximately 158,000 FU annotations. Under patientdisjoint evaluation, YOLOv8m achieved test [email protected]=0.920 and recall=0.860; EfficientNet-B5 with a su…

arXiv Computer Vision研究 / 創業融資來源內容 · 翻譯待補全待翻譯:An AI-Based Multi-Stage Approach for Androgenetic Alopecia Assessment from Low-Magnification Scalp Images
待翻譯:Octrees as an Explicit 3D Language

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02388v1 Announce Type: new Abstract: Existing 3D large language models (LLMs) compromise on two fronts: they compress shapes into latent codebook indices or coordinate text, which removes spatial structure from what the model observes, and they acquire the 3D modality by fine-tuning the backbone, which overwrites its general language ability. We present OctLLM, which addresses both limitations. Geometry enters as an explicit 3D sequence of octree occupancy tokens. However, full octree sequences grow rapidly with depth; OctLLM therefore randomly empties penultimate-level nodes and omits descendants while preserving shape, yielding a shorter coordinate- and depth-anchored Sparse Octree (S-Octree) for position-aware mask-modeling generation and 3D under…

arXiv Computer Vision模型 / 研究來源內容 · 翻譯待補全待翻譯:Octrees as an Explicit 3D Language
待翻譯:FactorSplat: Appearance-Controllable Gaussian Proxies for Medical Volume Rendering

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02382v1 Announce Type: new Abstract: Transfer functions (TFs) control color and visibility in medical volume rendering, but image-trained Gaussian proxies typically bake one transfer function into their appearance. We present FactorSplat, a per-scene N-dimensional Gaussian splatting (N-DGS) proxy that accepts region-specific intensity-to-RGBA curves at inference. A local lookup applies the authored color and opacity change, while a shared functional encoder and low-rank per-Gaussian factors learn the residual appearance response. Geometry and directional appearance remain shared across presets, with visibility control and TF-aware pruning preserving the ability to hide and reveal structures. On seven CT and MR scans, FactorSplat improves mean PSNR an…

arXiv Computer Vision研究來源內容 · 翻譯待補全待翻譯:FactorSplat: Appearance-Controllable Gaussian Proxies for Medical Volume Rendering
待翻譯:EviDent-CBCT: Evidence-Bottlenecked Report Generation from Dental CBCT under Non-Exhaustive Report Supervision

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02375v1 Announce Type: new Abstract: Dento-maxillofacial cone-beam CT (CBCT) reports may contain dozens of tooth-specific, anatomical, and spatial findings from a single 3D scan. Learning to generate such reports from limited clinical data is challenging because routine reports may not exhaustively document image findings, and a non-mention may reflect either absence or non-reporting. We present EviDent-CBCT, an evidence-bottlenecked framework designed for this incomplete supervision. An anatomy-aware network maps each CBCT scan to a discrete record of tooth-level, global, and tooth-IAC evidence. A dental-logic consistency projection reconciles incompatible evidence before a deterministic renderer and an image-blind local language model generate the…

arXiv Computer Vision模型 / 研究 / 創業融資來源內容 · 翻譯待補全待翻譯:EviDent-CBCT: Evidence-Bottlenecked Report Generation from Dental CBCT under Non-Exhaustive Report Supervision
待翻譯:Confidence-Controlled XAI Auditing for Pedestrian Detection under Domain Shift

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02364v1 Announce Type: new Abstract: Explainability is increasingly required for perception models in intelligent vehicles, yet whether explanations remain faithful under driving domain shift is still poorly understood. This work audits post-hoc explanations of a fixed YOLOv8s pedestrian detector across PIE and JAAD using ROI-based D-Deletion, frozen confidence terciles, rank-based tests, bootstrap intervals, and Holm correction. The audit shows that deletion-based faithfulness is strongly coupled to detection strength at explanation time, with Spearman correlations between 0.70 and 0.82 for D-RISE, making naive confidence-stratified comparisons unreliable. After controlling for detection strength within fixed f0 bins, D-RISE faithfulness remains dom…

arXiv Computer Vision政策 / 研究 / 創業融資來源內容 · 翻譯待補全待翻譯:Confidence-Controlled XAI Auditing for Pedestrian Detection under Domain Shift
待翻譯:SCOPE-4D: Endoscopic 4D Geometry Foundation Models

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02343v1 Announce Type: new Abstract: Geometric understanding supports endoscopic navigation and robotic assistance, but learning reliable endoscopic geometry faces two challenges: scarce geometric annotations and ambiguity between camera motion and tissue deformation. We present SCOPE-4D, an endoscopic 4D geometry foundation model that jointly predicts camera parameters, dense geometry, and 3D tissue trajectories from monocular RGB video in a single forward pass. Our curation and annotation pipeline constructs SCOPE-5K, a collection of approximately 5,000 clips spanning real and synthetic gastrointestinal endoscopy and laparoscopy. The collection provides rich geometric supervision and includes newly collected phantom and real-colonoscopy evaluation…

arXiv Computer Vision模型 / 研究 / 創業融資來源內容 · 翻譯待補全待翻譯:SCOPE-4D: Endoscopic 4D Geometry Foundation Models
待翻譯:SCION: Scene Composition with Instanced Neural Primitives

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02322v1 Announce Type: new Abstract: Real-world scenes are compositional: bricks, blades of grass, pebbles, and tree leaves recur across human-built and natural environments. Existing neural scene representations model these elements independently. Most 3D Gaussian Splatting and follow-up abstraction and compression methods treat each element as unique, fitting millions of independent Gaussians per scene. Prior methods like Splat and Replace fit template objects, but they require mostly manual selection of repeated elements. As a result, these representations store redundant parameters and provide weak manipulation handles for downstream tasks. We introduce SCION, a hier- archical compositional scene representation that replaces independent Gaussians…

arXiv Computer Vision研究來源內容 · 翻譯待補全待翻譯:SCION: Scene Composition with Instanced Neural Primitives
待翻譯:DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02320v1 Announce Type: new Abstract: Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a controlled way. We introduce DeskForge, a controllable desktop environment that composes and explores real applications to generate large-scale supervision for computer-use agents. It varies application states, content, window layout, appearance, and resolution, and fuses screenshots, accessibility trees, and window geometry into dense element annotations while recording the outcome of each executed action. Using this environment, we construc…

arXiv Computer VisionAgent / 模型 / 研究來源內容 · 翻譯待補全待翻譯:DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents
待翻譯:EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02298v1 Announce Type: new Abstract: 3D editing methods are usually tested on a single edit, yet an asset is built through a long sequence of revisions, each of which must implement the requested change while leaving everything else unchanged. We introduce EditHero, to our knowledge the first benchmark for long-horizon, part-level 3D editing, with natural-language instructions and target images for both geometry and texture. A deterministic assembly engine produces the exact target after every edit, and every sequence is reviewed by hand. We use EditHero to compare 2 opposite approaches to 3D editing. Non-agentic methods operate top down, regenerating the object from a learned 3D representation and inferring what to keep. In contrast, LLM/VLM agents…

arXiv Computer Vision研究 / 模型 / Agent來源內容 · 翻譯待補全待翻譯:EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling
待翻譯:Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02612v1 Announce Type: new Abstract: Simultaneous speech translation must emit useful target text before the source is complete while preserving every committed token. We adapt a full-utterance speech language model using prefix supervision derived from its own complete- and partial-waveform translations, requiring neither transcripts nor human translations. We compare single-turn forced-prefix and multi-turn append-only decoding, use a confidence threshold to control the inference-time quality--latency trade-off, and vary the density of training prefixes with a separate synthesis margin. On FLEURS and CoVoST2 in three language directions, prefix training improves quality--latency frontiers over the unadapted model, and confidence provides the broade…

arXiv Computational Linguistics模型 / 研究來源內容 · 翻譯待補全待翻譯:Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation
待翻譯:Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02549v1 Announce Type: new Abstract: Time and event expression extraction are fundamental temporal reasoning tasks, but the problem remains difficult due to annotation ambiguity, domain sensitivity, and unstable model behavior. Existing evaluations focus on in-domain performance, offering limited insight into reliability under distribution shifts. We evaluate multiple model configurations across families, architectures, and reasoning strategies over four dimensions of generalization, examining transfer from base performance, cross-dimensional correlations, and the effects of scale, architecture, and prompting. This provides a systematic study of how prompted LLMs generalize in time and event expression extraction tasks. We find that strong base-task…

arXiv Computational Linguistics模型 / 研究 / 創業融資來源內容 · 翻譯待補全待翻譯:Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks
待翻譯:A generative-informed neuro-symbolic framework for syntactic ambiguity resolution: Evidence from Arabic DPs

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02529v1 Announce Type: new Abstract: Syntactic ambiguity poses a persistent challenge for Arabic NLP, particularly in morphologically rich nominal constructions where multiple structu6ral interpretations may be compatible with the same surface sequence. This study proposes a generatively informed neuro-symbolic framework for resolving structural ambiguity in Modern Standard Arabic (MSA) DPs. The framework integrates generative syntactic notions with AraBERT by representing ambiguity as a candidate-based decision task in which linguistically motivated alternatives are explicitly constructed and evaluated through candidate-conditioned input representations. Findings indicate that the model achieved 96.88% accuracy, 95.92% macro-F1, 96.83% weighted F1,…

arXiv Computational Linguistics模型 / 研究 / 創業融資來源內容 · 翻譯待補全待翻譯:A generative-informed neuro-symbolic framework for syntactic ambiguity resolution: Evidence from Arabic DPs
待翻譯:From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02486v1 Announce Type: new Abstract: Typed decision models answer schema-constrained questions about a text in one forward pass and return probabilities meant to be thresholded. We ask whether biomedical sentence encoders trained for retrieval are good starting points for such models. We present SBERT2S1, which converts Sentence-Transformers encoders into bi-encoder, cross-head (C) and prior-fused residual (PFR) decision models, together with BIODECIDE, a biomedical typed-decision suite, and MEDLINE-S1, 243k training decisions derived from NLM indexing. Across six parent-retriever pairs, retrieval training improves zero-shot matching of content-bearing options. After fine-tuning, its effect depends on the head: across five pairs and three training-se…

arXiv Computational Linguistics模型 / 研究來源內容 · 翻譯待補全待翻譯:From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders
待翻譯:APDMem: Agent-Controlled Progressive Disclosure for Query-Adaptive Long-Term Memory

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02472v1 Announce Type: new Abstract: Personalized LLM assistants must recover sparse evidence from long conversation histories across queries of varying complexity. We introduce APDMem (Agent-controlled Progressive Disclosure Memory), a hierarchical long-term memory architecture that applies progressive disclosure to memory retrieval. Rather than relying on a flat memory store or fixed retrieval granularity, APDMem represents conversation history as four progressively detailed layers: thematic summaries, personalized key facts, turn-level evidence notes, and raw messages. At inference time, a controller applies progressive disclosure to the memory hierarchy: it first reads high-level summaries and drills into finer evidence only when needed. This cre…

arXiv Computational LinguisticsAgent / 模型 / 研究來源內容 · 翻譯待補全待翻譯:APDMem: Agent-Controlled Progressive Disclosure for Query-Adaptive Long-Term Memory
待翻譯:CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02460v1 Announce Type: new Abstract: Recent benchmarks rely on user simulators to evaluate AI agents in multi-turn interaction. While existing simulation techniques demonstrate surface fidelity to human style and behavior, ecologically valid interactive benchmarking also requires alignment in when and how agents fail across simulated and real user populations. We find that existing simulators lack outcome calibration: agreement with observed success rates and failure patterns when real users interact with the same agent. We introduce Calibrated User Embeddings (CUE), a framework that both encodes observed sessions and samples continuous representations, then decodes them into persona commands to steer LLMs to act as user simulators without training.…

arXiv Computational Linguistics模型 / 研究 / Agent來源內容 · 翻譯待補全待翻譯:CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking
待翻譯:FinDialogLens: Event Extraction over Multi-Party Dialogue for Missed-Trade Identification in Financial Chatrooms

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02455v1 Announce Type: new Abstract: Multi-party financial chatrooms are vital for sales-and-trading professionals, but their complexity makes manual recovery of missed trades infeasible: each Request for Quote (RFQ) is an event whose final price and trade outcome appear many messages after the RFQ-trigger message (the inquiry message), interleaved with concurrent RFQs from other participants. We cast this as event extraction (EE) over multi-party dialogue and present FinDialogLens, a hybrid LLM pipeline in which compact fine-tuned classifiers act as inference-time scaffolds: they detect RFQ-triggers and price/trade outcome metadata, an RFQ-Level Module segments per-event RFQ windows, and a Trade Engine fills argument roles. With GPT-4o, FinDialogLen…

arXiv Computational Linguistics模型 / 研究來源內容 · 翻譯待補全待翻譯:FinDialogLens: Event Extraction over Multi-Party Dialogue for Missed-Trade Identification in Financial Chatrooms
待翻譯:Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02444v1 Announce Type: new Abstract: Large language models often solve a theorem forward yet fail to disprove a closely related false one: a falsification gap that supervised fine-tuning does not close and can actively worsen. We frame counterexample generation as constrained witness emission against a deterministic per-theorem Python verifier, and release SymCE, a corpus of 4,707 false undergraduate-algebra and real-analysis conjectures, each paired with executable verifiers. The verifier also serves as the reward function, making SymCE a training environment. Training Qwen3-4B with SFT followed by GRPO under this oracle reveals an imitation trap: counterexample-only SFT collapses true-theorem recognition from 0.27 to 0.00, while RLVR with a sparse…

arXiv Computational Linguistics模型 / 研究來源內容 · 翻譯待補全待翻譯:Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs
待翻譯:Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02425v1 Announce Type: new Abstract: Static evaluations credit a language model for naming the right move, but an agent must carry a plan through to a verified outcome while an opponent responds. We introduce XiangqiBench, an executable benchmark that measures this difference in Chinese chess: starting from 119 tactical endgames with forced mates supported by engine or checks-only search, an LLM agent must deliver checkmate against an engine defender. An interactive REPL interface separates real moves, state queries, and forward simulation, and we record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols. Three signals that look like competence each overstate closed-loop success. (i) The Conversion Gap: models play th…

arXiv Computational Linguistics模型 / Agent / 研究來源內容 · 翻譯待補全待翻譯:Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents
待翻譯:HakemBench: A Turkish Benchmark of Typed Decisions

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02293v1 Announce Type: new Abstract: HakemBench is a Turkish benchmark of typed decisions, in which the model under test reads a text, a question and a fixed set of options and returns a probability for every option. Version 1.0 is released fully open under CC BY 4.0, with 2,346 items and 4,275 choice, yes/no and score questions in seven tracks (fact-check triage, education, guardrails, legal routing, moderation, spam and phishing, and customer support). One harness scores decision quality (macro F1), calibration (from the normalised Brier score) and selective automation (from the normalised area under the generalised risk-coverage curve), combines them by a geometric mean and reports intervals from 2,000 bootstrap draws; probes for option order, par…

arXiv Computational Linguistics研究 / 模型 / Agent來源內容 · 翻譯待補全待翻譯:HakemBench: A Turkish Benchmark of Typed Decisions
待翻譯:TRACE: A Reproducible Benchmark for Electricity Price Forecasting with Official Operational Text

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02256v1 Announce Type: new Abstract: Electricity price forecasting (EPF) supports scheduling, bidding, and risk management in electricity markets, yet existing benchmarks focus mainly on numerical inputs, leaving the forecasting value of forecast-time textual context insufficiently evaluated. We introduce TRACE, a reproducible benchmark of 7,300 zone--day instances pairing prices from five zones in a major U.S. market with official operational text available at the forecast cutoff. TRACE reconstructs official operational text at each cutoff, preventing post-cutoff information leakage. We evaluate TRACE for semantic alignment and forecasting value. Semantic assessments align with central movement and both tail risks in ground-truth prices, most consis…

arXiv Machine Learning研究 / 模型來源內容 · 翻譯待補全待翻譯:TRACE: A Reproducible Benchmark for Electricity Price Forecasting with Official Operational Text
待翻譯:MACTS-EM: Multi-Agent Collaborative Time Series Forecasting with Emergent Memory

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02255v1 Announce Type: new Abstract: Time series forecasting remains a critical challenge across numerous domains. Despite significant advancements, existing approaches struggle with complex phenomena such as regime shifts, cross-domain knowledge transfer, and multimodal data integration. This paper introduces Multi-Agent Collaborative Time Series Forecasting with Emergent Memory (MACTS-EM), a novel framework where specialised agents collaborate to achieve superior forecasting performance. The MACTS-EM architecture integrates: (1) domain-specialised forecasting agents for pattern recognition, anomaly detection, causal inference, and uncertainty quantification; (2) a meta-cognitive layer for dynamic agent allocation; (3) an emergent memory mechanism e…

arXiv Machine LearningAgent / 模型 / 研究來源內容 · 翻譯待補全待翻譯:MACTS-EM: Multi-Agent Collaborative Time Series Forecasting with Emergent Memory
待翻譯:Overcoming Challenges of Interpretive Structural Modeling with Large Language Models

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02254v1 Announce Type: new Abstract: Interpretive Structural Modeling (ISM) is a well-known process for multi-criteria decision making. The success of ISM over other methodologies is its ability to model causal relationships, the binary scale of factors, and resulting hierarchical representation. Traditionally, the modeling process is performed by repeated interactions with subject matter experts until consensus is reached. This process is tedious, labor-intense, and most importantly limits the ability of ISM to scale to studies with hundreds of variables. Drawing on existing work of causal graph discovery with large language models (LLM) as imperfect experts, this work explores an integrated LLM-ISM approach for ISM. Pairwise, k-wise, rowwise, and f…

arXiv Machine Learning模型 / 研究來源內容 · 翻譯待補全待翻譯:Overcoming Challenges of Interpretive Structural Modeling with Large Language Models
待翻譯:Approximation Property of Dropout Neural Networks: Sobolev Rates and Confidence Bounds

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02253v1 Announce Type: new Abstract: The universal approximation property of dropout neural networks does not by itself describe the network size required for an accurate random realization. In this work, we study approximation of the unit ball of $W^{n,\infty}([0,1]^d)$ by ReLU networks whose edges are retained independently with probability $p$. The approximation error is measured uniformly over the input domain, and the guarantee holds with probability at least $1-\delta$ for a single sampled network. We construct networks of constant depth and size $\widetilde O_{n,d}(p^{-9}\varepsilon^{-\max\{d/n,2\}} \log(1/\delta))$. The construction combines bounded local subnetworks, localization on a successful approximation event, and a multiscale Taylor d…

arXiv Machine Learning研究來源內容 · 翻譯待補全待翻譯:Approximation Property of Dropout Neural Networks: Sobolev Rates and Confidence Bounds
待翻譯:Counterfactual Predictions in Scientific Emulators Without Controlled Experiments

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02252v1 Announce Type: new Abstract: Many scientific questions require reasoning about what was never observed: What if the conditions, interventions, or history had been different? Models can predict accurately on observed data yet fail on such what-if queries when correlated inputs are varied independently. A common remedy is to add controlled simulation data in which these factors are explicitly disentangled, but this requires access to a simulator, can be computationally expensive, and inherits the simulator's modeling assumptions. We introduce ReRoute, a framework for targeted scientific what-if prediction that combines factual data with partial mechanistic knowledge, without requiring controlled intervention data for adaptation. ReRoute fixes t…

arXiv Machine Learning模型 / 研究來源內容 · 翻譯待補全待翻譯:Counterfactual Predictions in Scientific Emulators Without Controlled Experiments
待翻譯:Rank-Aware Speculative Sampling for Diffusion Draft Trees

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.02251v1 Announce Type: new Abstract: Speculative sampling accelerates diffusion generation by verifying inexpensive draft states in parallel while preserving the target law. Recent tree-based methods allocate the parallel compute budget more effectively than single-chain drafts, as demonstrated by Diffusion Greedy Rejection Sampling (D-GRS). D-GRS generates $K$ conditionally independent candidates per node, and sequentially tests them in their generation order. Yet the sampled candidates admit an informative ranking without additional target-model evaluations. To exploit this, we introduce Rank-Aware Speculative Sampling (RASS), a verification rule for speculative draft trees based on rank-aware list coupling. RASS orders draft candidates along the p…

arXiv Machine Learning模型 / 政策 / 研究來源內容 · 翻譯待補全待翻譯:Rank-Aware Speculative Sampling for Diffusion Draft Trees