跳到主要內容
AI News HubLIVE

AI 新聞即時情報

即時監測

今日 AI 世界最重要的變化

來源目錄收錄 105 個來源 · 本頁最新發布時間 2026-10-02 19:52 UTC+8。

部分報道的翻譯與分析尚未完成,已標示來源內容。精選優先顯示處理完成的報道。

即時監測

即時更新

即時追蹤可信來源,保留出處、權限和站內閱讀模式,把噪音壓成可讀情報。

即時更新

重設

2026-10-02

待翻譯:Amazon writes scary blog warning communities not to block data centers

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Amazon is calling for people to support AI data center projects, or risk irreparable harm to the US economy and national security. In a more than 3,000 word blog posted today, Amazon web services CEO Matt Garman pushed back on public concerns around the impact that data centers may have on jobs, power demands, and the environment, saying "our nation can't afford to lose" the race for AI dominance. "With any change, there will be important questions raised, but there will also be misinformation and outright lies, and in the age of social media and 24/7 news, myths take hold faster than ever before," said Garman. "In fact, this build out is s … Read the full story at The Verge.

The Verge AI創業融資來源內容 · 翻譯待補全待翻譯:Amazon writes scary blog warning communities not to block data centers
待翻譯:Coding Agents Love Decision Records

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The following article originally appeared on Duncan Davidson’s blog and is being republished here with the author’s permission. Decision records give coding agents durable project context—as long as they don’t turn every decision into a courtroom transcript. Architectural Decision Records (ADRs) help human teams establish rules and carry context forward in software projects. They capture […]

O'Reilly AI & ML RadarAgent來源內容 · 翻譯待補全待翻譯:Coding Agents Love Decision Records
待翻譯:AI weapons systems are already here. Algorithms must not decide who lives and dies | Kenneth Roth

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Israel’s conduct in Gaza shows the danger of using AI in warfare. Humans must remain at the center of wartime decisions As tech leaders warn about the threats of artificial intelligence, most conjure up images of AI agents running amok, such as hacking systems used to run our critical infrastructure, banks or even militaries. Yet one serious danger is mentioned less frequently – the threat posed by AI – empowered lethal weapons. Fully autonomous versions of these weapons are colloquially known as killer robots. Israel’s conduct in Gaza shows that, in an important respect, such AI weapon systems are already here. An algorithm has been deciding who lives and who dies. Continue reading...

The Guardian AIAgent / 機械人來源內容 · 翻譯待補全待翻譯:AI weapons systems are already here. Algorithms must not decide who lives and dies | Kenneth Roth
待翻譯:Takweem AI

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI工具來源內容 · 翻譯待補全待翻譯:Takweem AI
待翻譯:AI music maker Suno now generates spoken words

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Suno is branching out from the world of AI music, launching a new feature that generates spoken voices based on scripts or prompted descriptions. Speech is now available in public beta across Suno's web and mobile platforms, and allows you to simultaneously generate voiceovers and background music to accompany them. "Music will always be at the heart of Suno and what we build. At the same time, our vision has always extended to other forms of human expression," Suno chief product officer, Jack Brody, said in the announcement. "Today, we're expanding what's possible in Suno with Speech: the first audio model that generates voice and music to … Read the full story at The Verge.

The Verge AI模型來源內容 · 翻譯待補全待翻譯:AI music maker Suno now generates spoken words
待翻譯:HireKey

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI工具來源內容 · 翻譯待補全待翻譯:HireKey
待翻譯:ParakeetAI 2.0

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI工具來源內容 · 翻譯待補全待翻譯:ParakeetAI 2.0
待翻譯:[AINews] Pi 1.0, Pi Durable, and AIE NYC

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:the minimalist harness goes stable... and TypeScript!

Latent SpaceAgent / 研究來源內容 · 翻譯待補全待翻譯:[AINews] Pi 1.0, Pi Durable, and AIE NYC
待翻譯:AWS Strands Labs Releases Strands Decider 2B: An Open Source Decision Model That Picks Options in About 115 ms

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AWS's Strands Agents team released Strands Decider 2B, an Apache-2.0 decision model built on Qwen3.5-2B-Base. It returns choices, yes/no probabilities and scores with calibrated confidence in one forward pass, never text. It runs at a 115 ms median on an RTX 3090 and scores 0.723 on the JevBench public set, which makes it a fast local option for routing, tool selection and guardrails in AI agents. The post AWS Strands Labs Releases Strands Decider 2B: An Open Source Decision Model That Picks Options in About 115 ms appeared first on MarkTechPost.

MarkTechPost模型 / Agent / 芯片來源內容 · 翻譯待補全待翻譯:AWS Strands Labs Releases Strands Decider 2B: An Open Source Decision Model That Picks Options in About 115 ms
待翻譯:Aster by AsterWise

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AIAgent來源內容 · 翻譯待補全待翻譯:Aster by AsterWise
待翻譯:Codync

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI工具來源內容 · 翻譯待補全待翻譯:Codync
待翻譯:DeepJEPA: Scaling World Models from Within

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00368v1 Announce Type: new Abstract: World-model planners typically scale outward by rolling farther, sampling more trajectories, or optimizing longer, while assigning the same computation to every imagined transition. We show that making every transition uniformly deeper wastes computation and can degrade planning because useful refinement is concentrated at a small set of decision-critical events. We introduce DeepJEPA, a weight-tied joint-embedding predictive world model that treats transition depth as an inner test-time scaling axis and learns when another recurrent update is worth computing for each candidate and rollout step. Across five visual-control settings, DeepJEPA improves or matches the strongest fixed-depth planner while averaging only…

arXiv Robotics模型 / 研究 / 機械人來源內容 · 翻譯待補全待翻譯:DeepJEPA: Scaling World Models from Within
待翻譯:DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00360v1 Announce Type: new Abstract: Reinforcement learning (RL) for dexterous manipulation must discover finger-object contacts and then control the object precisely; the action noise that serves the first goal can interfere with the second. In trajectory-guided settings such as ViViDex, where RL refine hand-object trajectories from human video, our baseline PPO runs end near their initial action noise after 5M steps, motivating explicit control of exploration scale. DexPolicy makes that scale an explicit function of training steps, annealing from broad to narrow exploration while holding loss, architecture, reward, and optimizer settings fixed. We study three policy-optimization settings: PPO, critic-free GRPO continuation, and a flow-parameterized…

arXiv Robotics研究 / 機械人來源內容 · 翻譯待補全待翻譯:DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation
待翻譯:IndoorBEV: A Lightweight Real-Time LiDAR BEV Perception System for Indoor Mobile Robots

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00355v1 Announce Type: new Abstract: Efficient indoor LiDAR perception is challenging because mobile robots must understand cluttered three-dimensional environments under strict latency and memory constraints. Existing point-based and voxel-based methods often incur substantial computational overhead, whereas conventional bird's-eye-view (BEV) representations improve efficiency at the cost of discarding vertical geometric information. We present IndoorBEV, a lightweight LiDAR perception framework that mitigates this tradeoff through a height-aware BEV representation and geometry-conditioned feature fusion. IndoorBEV summarizes the vertical point distribution in each BEV cell using statistical height features and multi-frequency height encoding, allow…

arXiv Robotics機械人 / 芯片 / 研究來源內容 · 翻譯待補全待翻譯:IndoorBEV: A Lightweight Real-Time LiDAR BEV Perception System for Indoor Mobile Robots
待翻譯:Retrospective Open-Vocabulary Memory for Long-Term Object Search

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00330v1 Announce Type: new Abstract: Long-term object search requires learning where objects usually appear from repeated but uneven observations of a changing environment. We formulate retrospective open-vocabulary memory as probabilistic inference from censored observations, where the key idea is to reason with evidence per opportunity: a detection or non-detection should influence belief only in proportion to the robot's opportunity to observe the corresponding location. We introduce ECROM, which uses this principle to estimate long-term prevalence for concepts specified only at query time and converts the resulting belief directly into an active-search prior. To evaluate this problem, we introduce a controlled long-term benchmark in ten HM3D home…

arXiv Robotics研究 / 機械人來源內容 · 翻譯待補全待翻譯:Retrospective Open-Vocabulary Memory for Long-Term Object Search
待翻譯:DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00317v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. Sequence-level reinforcement learning can address this limitation, but typically requires policy rollouts and closed-loop interaction, which are costly for real-robot manipulation. We introduce DriftOPD, a teacher-free, rollout-free framework for sequence-level on-policy distillation of continuous VLA action experts. We show that the sequence-level reverse Kullback-Leibler (KL) divergence decomposes into a chunk-leve…

arXiv Robotics模型 / 研究 / 機械人來源內容 · 翻譯待補全待翻譯:DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies
待翻譯:Real-Time Whole-Body Safe Motion Generation for Multi-Segment Tendon-Driven Continuum Robots

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00220v1 Announce Type: new Abstract: Real-time motion generation for tendon-driven continuum robots requires accurate modeling of nonuniform bending and whole-body collision avoidance. This paper presents a unified actuation-space framework for planar multi-segment tendon-driven continuum robots. An energy-based variable-curvature model captures spatially varying tendon spacing and bending stiffness and provides analytical Jacobians for differential inverse kinematics and safety monitoring. A multipoint CBF-QP enforces backbone clearance under obstacle motion and actuation-velocity bounds, while its decision dimension depends only on the number of independently actuated segments. The model closely agrees with GVS references, with a maximum curvature…

arXiv Robotics機械人 / 政策 / 研究來源內容 · 翻譯待補全待翻譯:Real-Time Whole-Body Safe Motion Generation for Multi-Segment Tendon-Driven Continuum Robots
待翻譯:HumanoidTTT: Test-Time Capability Reuse for Efficient Humanoid Control

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00198v1 Announce Type: new Abstract: Recent advances in motion generation and whole-body tracking have enabled humanoid robots to execute increasingly diverse motions, yet the same motion capabilities may be requested repeatedly during continual deployment. Reliable reuse is challenging because intervening motions can change the robot's entry state, making previously successful motions unsafe to replay blindly. Meanwhile, validated capabilities accumulate during deployment, while bounded storage requires deciding which ones are worth retaining. To address these challenges, we present HumanoidTTT, a framework for test-time capability reuse in continual humanoid control. Specifically, we introduce Selective Full-Motion Reuse, which authorizes direct re…

arXiv Robotics機械人 / 研究來源內容 · 翻譯待補全待翻譯:HumanoidTTT: Test-Time Capability Reuse for Efficient Humanoid Control
待翻譯:Probabilistic Plan Legibility with Off-the-shelf Planners

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00065v1 Announce Type: new Abstract: Legible planning is the creation of plans that best disambiguate their goals from a set of other candidates from an observer's perspective. In this paper we propose a method for legible planning for arbitrary PDDL domains, by extending previous research on legibility to classical planning without requiring to construct ad-hoc planners. We also discuss how the observer perspective may be estimated through a second order theory of mind that connects the planner's and the observer's task spaces. Our solution can for example be deployed in human-robot teaming scenarios, where an autonomous robot in a team can implicitly communicate its goal by producing legible plans. We present benchmark results on several PDDL plann…

arXiv Robotics研究 / 機械人來源內容 · 翻譯待補全待翻譯:Probabilistic Plan Legibility with Off-the-shelf Planners
待翻譯:Multi-Reference Path Tracking Control for an Agricultural Tractor with Nonlinear Model Predictive Control

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00057v1 Announce Type: new Abstract: Guiding a tractor along a predefined reference path is a key component of precision agriculture. This study develops a path tracking controller based on Nonlinear Model Predictive Control, which incorporates multiple segments of a piecewise-linear reference path directly into the objective function. In addition, methods for selecting viable reference segments from the full path are presented. The control system is evaluated during a field test with a tractor controlled via the Tractor Implement Management steering interface. The NMPC solver converged on average after 3.45 ms and tracked the curved reference path with a mean absolute cross-track error of 6.1 cm.

arXiv Robotics研究 / 機械人來源內容 · 翻譯待補全待翻譯:Multi-Reference Path Tracking Control for an Agricultural Tractor with Nonlinear Model Predictive Control
待翻譯:Bounded-Fidelity Sim-as-Demo-Stage: Mocap Handoff for Governance Benchmarks

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00008v1 Announce Type: new Abstract: Sim-to-real research pursues physics fidelity as a primary objective: simulators are judged by how closely they reproduce real-world contact dynamics. For governance benchmarking of LLM-driven robots, where the simulator demonstrates that an admission/policy/contract/audit pipeline behaves correctly, contact fidelity at object handoffs (grasp, carry, place) becomes a liability: contact-force integration noise injects audit-chain divergence that is structurally unrelated to the governance property under test. We propose bounded-fidelity sim-as-demo-stage, a design pattern that suppresses contact physics within explicitly bracketed handoff envelopes while preserving full dynamics elsewhere. The construction uses MuJ…

arXiv Robotics政策 / 研究 / 模型來源內容 · 翻譯待補全待翻譯:Bounded-Fidelity Sim-as-Demo-Stage: Mocap Handoff for Governance Benchmarks
待翻譯:A Framework for Egocentric and Exocentric Procedural Understanding via Temporal Segmentation and Semantic Abstraction

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00069v1 Announce Type: new Abstract: Long-horizon ego/exo data contains rich procedural evidence, but are redundant, noisy, and costly to process or retain. We propose a compact framework that converts continuous multimodal workplace video into a structured Procedural State Memory, implemented as a Work Environment Model (WEM). Inspired by event segmentation theory, we detect boundaries using changes in visual context, location, motion, narration, gaze/object interaction, and optional exocentric workspace evidence, rather than fixed windows or visual novelty alone. Each segment is abstracted into an evidence-linked event card containing actor, interval, location, action, objects/tools, pre/post state, confidence, and provenance. These event cards inc…

arXiv Computer Vision模型 / 研究 / 創業融資來源內容 · 翻譯待補全待翻譯:A Framework for Egocentric and Exocentric Procedural Understanding via Temporal Segmentation and Semantic Abstraction
待翻譯:Robust Online Aero-Engine Blade Defect Detection via Dual-Alignment Test-Time Adaptation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00067v1 Announce Type: new Abstract: Reliable visual inspection is essential for quality assurance in aero-engine blade manufacturing, where defect appearance may vary across production lines, imaging conditions, blade poses, and surface backgrounds. Such domain shifts cause a mismatch between training and deployment data and degrade the reliability of deep defect detectors in online inspection. This problem is particularly challenging because aero-engine blade images usually contain sparse defects, making pseudolabel-based adaptation vulnerable to noisy or missing predictions. To address this issue, we propose Aero-engine Blade Defect Detector (ABDD), an online adaptive detection framework based on test-time adaptation. ABDD introduces a Dual-Alignm…

arXiv Computer Vision研究來源內容 · 翻譯待補全待翻譯:Robust Online Aero-Engine Blade Defect Detection via Dual-Alignment Test-Time Adaptation
待翻譯:Reachability Is Not Generalization: Understanding Verb--Noun Decomposition in Assembly Action Recognition

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00064v1 Announce Type: new Abstract: Assembly actions are compositional: they combine a manipulation with a part or tool. In deployment, systems routinely encounter novel combinations of familiar components, yet an atomic action classifier assigns every unseen combination exactly zero probability by construction. The prevailing solution is verb--noun decomposition, which predicts components separately and recombines them to reach unseen actions. While widely adopted, how decomposition generalizes under compositional shift remains poorly understood. We present a systematic analysis of verb--noun decomposition across three assembly datasets (MECCANO, HAViD, and IMPACT). Although decomposition escapes the atomic ceiling, its generalization extends only…

arXiv Computer Vision研究來源內容 · 翻譯待補全待翻譯:Reachability Is Not Generalization: Understanding Verb--Noun Decomposition in Assembly Action Recognition
待翻譯:DSSR-3D: Decoupled Reasoning for View-Dependent Referring in 3D Gaussians

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00040v1 Announce Type: new Abstract: Recent advances in 3D Gaussian Splatting have enabled open-vocabulary and referring segmentation by distilling semantic knowledge from 2D foundation models into 3D representations. However, existing referring fields embed language features in a globally view-invariant space, making them fundamentally unable to resolve observer-centric spatial relations (e.g., "to the left of") that depend on camera pose. We propose DSSR-3D, an inference-time framework for view-dependent referring segmentation on continuous 3D Gaussian fields, formalized as two interfaces - pose-invariant semantic localization and pose-conditioned spatial reasoning - such that any pair of functions satisfying these constraints yields a valid instan…

arXiv Computer Vision模型 / 研究來源內容 · 翻譯待補全待翻譯:DSSR-3D: Decoupled Reasoning for View-Dependent Referring in 3D Gaussians
待翻譯:Seeing the City or Recognizing the Place? What Street-View Imagery Adds Beyond Existing Urban Data in VLM Urban Sensing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00031v1 Announce Type: new Abstract: Street-view imagery is increasingly used to infer urban attributes, but predictive accuracy alone does not reveal how much a photograph contributes beyond data already available for the same place. We compare image-based predictions with existing urban data across seven attributes from five public resources and three VLMs. The same urban units are evaluated using images, task context, nearby observations, and public records, while image replacements and conflicting records test source reliance. Existing urban data matched or exceeded image-only models for road damage, curb ramps, and house price, while neighbouring official statistics nearly matched the best image result for population. Images were more informativ…

arXiv Computer Vision研究來源內容 · 翻譯待補全待翻譯:Seeing the City or Recognizing the Place? What Street-View Imagery Adds Beyond Existing Urban Data in VLM Urban Sensing
待翻譯:Domain generalization and synthetic data in object detection: the enabler, the probe, and the gap

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00030v1 Announce Type: new Abstract: Object detection models often experience performance degradation when deployed under distribution shifts, caused by for example changes in weather type, operational environment, or object appearance. Domain Generalization (DG) aims to develop models that remain robust under such shifts and generalize well to unseen domains. DG research specifically focused on object detection models is scarce, although these models face additional challenges around localization and multi-scale representations. Synthetic data is a promising tool to support in DG, by enabling large-scale generation of diverse new samples. In this paper, we present an object detection-centric review of DG and examine the role of synthetic data from t…

arXiv Computer Vision研究來源內容 · 翻譯待補全待翻譯:Domain generalization and synthetic data in object detection: the enabler, the probe, and the gap
待翻譯:Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00024v1 Announce Type: new Abstract: Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE -- the benchmark common to all three -- the mid layers encode the ground-truth answer in 68-91% of errors, yet this signal is not causally active for the final prediction: residual-stream patching yields 0% non-trivial flip at the layer level on all three architectures, and on two of three at the per-head level (Qwen: 0/12,600 patched forwards). The lone exception, InternVL3 layer-20 head-2, is a non-vocab, self-attending head whose effect is localized to that specific head (p < 1e-4). Despite the null, the errors separate operationally into th…

arXiv Computer Vision模型 / 研究來源內容 · 翻譯待補全待翻譯:Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null
待翻譯:Spatial Lifting for Dense Prediction

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00017v1 Announce Type: new Abstract: We present Spatial Lifting (SL), a novel methodology for dense prediction tasks. SL operates by lifting standard inputs, such as 2D images, into a higher-dimensional space and subsequently processing them using networks designed for that higher dimension, such as a 3D U-Net. Counterintuitively, this dimensionality lifting allows us to achieve good performance on benchmark tasks compared to conventional approaches, while reducing inference costs and \textbf{drastically lowering the number of model parameters}. The SL framework produces intrinsically structured outputs along the lifted dimension. This emergent structure facilitates dense supervision during training and enables single-forward-pass self-consistency-ba…

arXiv Computer Vision研究來源內容 · 翻譯待補全待翻譯:Spatial Lifting for Dense Prediction