AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Amazon is calling for people to support AI data center projects, or risk irreparable harm to the US economy and national security. In a more than 3,000 word blog posted today, Amazon web services CEO Matt Garman pushed back on public concerns around the impact that data centers may have on jobs, power demands, and the environment, saying "our nation can't afford to lose" the race for AI dominance. "With any change, there will be important questions raised, but there will also be misinformation and outright lies, and in the age of social media and 24/7 news, myths take hold faster than ever before," said Garman. "In fact, this build out is s … Read the full story at The Verge.
AI 新聞即時情報
部分報導的翻譯與分析尚未完成,已標示來源內容。精選優先顯示處理完成的報導。
即時監測
即時更新
即時追蹤可信來源,保留出處、權限和站內閱讀模式,把噪音壓成可讀情報。
即時更新
2026-10-02
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The following article originally appeared on Duncan Davidson’s blog and is being republished here with the author’s permission. Decision records give coding agents durable project context—as long as they don’t turn every decision into a courtroom transcript. Architectural Decision Records (ADRs) help human teams establish rules and carry context forward in software projects. They capture […]
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The opportunity: Personalization as a revenue engineEvery second a shopper spends...
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Israel’s conduct in Gaza shows the danger of using AI in warfare. Humans must remain at the center of wartime decisions As tech leaders warn about the threats of artificial intelligence, most conjure up images of AI agents running amok, such as hacking systems used to run our critical infrastructure, banks or even militaries. Yet one serious danger is mentioned less frequently – the threat posed by AI – empowered lethal weapons. Fully autonomous versions of these weapons are colloquially known as killer robots. Israel’s conduct in Gaza shows that, in an important respect, such AI weapon systems are already here. An algorithm has been deciding who lives and who dies. Continue reading...
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Suno is branching out from the world of AI music, launching a new feature that generates spoken voices based on scripts or prompted descriptions. Speech is now available in public beta across Suno's web and mobile platforms, and allows you to simultaneously generate voiceovers and background music to accompany them. "Music will always be at the heart of Suno and what we build. At the same time, our vision has always extended to other forms of human expression," Suno chief product officer, Jack Brody, said in the announcement. "Today, we're expanding what's possible in Suno with Speech: the first audio model that generates voice and music to … Read the full story at The Verge.
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:the minimalist harness goes stable... and TypeScript!
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AWS's Strands Agents team released Strands Decider 2B, an Apache-2.0 decision model built on Qwen3.5-2B-Base. It returns choices, yes/no probabilities and scores with calibrated confidence in one forward pass, never text. It runs at a 115 ms median on an RTX 3090 and scores 0.723 on the JevBench public set, which makes it a fast local option for routing, tool selection and guardrails in AI agents. The post AWS Strands Labs Releases Strands Decider 2B: An Open Source Decision Model That Picks Options in About 115 ms appeared first on MarkTechPost.
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00368v1 Announce Type: new Abstract: World-model planners typically scale outward by rolling farther, sampling more trajectories, or optimizing longer, while assigning the same computation to every imagined transition. We show that making every transition uniformly deeper wastes computation and can degrade planning because useful refinement is concentrated at a small set of decision-critical events. We introduce DeepJEPA, a weight-tied joint-embedding predictive world model that treats transition depth as an inner test-time scaling axis and learns when another recurrent update is worth computing for each candidate and rollout step. Across five visual-control settings, DeepJEPA improves or matches the strongest fixed-depth planner while averaging only…
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00360v1 Announce Type: new Abstract: Reinforcement learning (RL) for dexterous manipulation must discover finger-object contacts and then control the object precisely; the action noise that serves the first goal can interfere with the second. In trajectory-guided settings such as ViViDex, where RL refine hand-object trajectories from human video, our baseline PPO runs end near their initial action noise after 5M steps, motivating explicit control of exploration scale. DexPolicy makes that scale an explicit function of training steps, annealing from broad to narrow exploration while holding loss, architecture, reward, and optimizer settings fixed. We study three policy-optimization settings: PPO, critic-free GRPO continuation, and a flow-parameterized…
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00355v1 Announce Type: new Abstract: Efficient indoor LiDAR perception is challenging because mobile robots must understand cluttered three-dimensional environments under strict latency and memory constraints. Existing point-based and voxel-based methods often incur substantial computational overhead, whereas conventional bird's-eye-view (BEV) representations improve efficiency at the cost of discarding vertical geometric information. We present IndoorBEV, a lightweight LiDAR perception framework that mitigates this tradeoff through a height-aware BEV representation and geometry-conditioned feature fusion. IndoorBEV summarizes the vertical point distribution in each BEV cell using statistical height features and multi-frequency height encoding, allow…
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00330v1 Announce Type: new Abstract: Long-term object search requires learning where objects usually appear from repeated but uneven observations of a changing environment. We formulate retrospective open-vocabulary memory as probabilistic inference from censored observations, where the key idea is to reason with evidence per opportunity: a detection or non-detection should influence belief only in proportion to the robot's opportunity to observe the corresponding location. We introduce ECROM, which uses this principle to estimate long-term prevalence for concepts specified only at query time and converts the resulting belief directly into an active-search prior. To evaluate this problem, we introduce a controlled long-term benchmark in ten HM3D home…
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00317v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. Sequence-level reinforcement learning can address this limitation, but typically requires policy rollouts and closed-loop interaction, which are costly for real-robot manipulation. We introduce DriftOPD, a teacher-free, rollout-free framework for sequence-level on-policy distillation of continuous VLA action experts. We show that the sequence-level reverse Kullback-Leibler (KL) divergence decomposes into a chunk-leve…
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00220v1 Announce Type: new Abstract: Real-time motion generation for tendon-driven continuum robots requires accurate modeling of nonuniform bending and whole-body collision avoidance. This paper presents a unified actuation-space framework for planar multi-segment tendon-driven continuum robots. An energy-based variable-curvature model captures spatially varying tendon spacing and bending stiffness and provides analytical Jacobians for differential inverse kinematics and safety monitoring. A multipoint CBF-QP enforces backbone clearance under obstacle motion and actuation-velocity bounds, while its decision dimension depends only on the number of independently actuated segments. The model closely agrees with GVS references, with a maximum curvature…
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00198v1 Announce Type: new Abstract: Recent advances in motion generation and whole-body tracking have enabled humanoid robots to execute increasingly diverse motions, yet the same motion capabilities may be requested repeatedly during continual deployment. Reliable reuse is challenging because intervening motions can change the robot's entry state, making previously successful motions unsafe to replay blindly. Meanwhile, validated capabilities accumulate during deployment, while bounded storage requires deciding which ones are worth retaining. To address these challenges, we present HumanoidTTT, a framework for test-time capability reuse in continual humanoid control. Specifically, we introduce Selective Full-Motion Reuse, which authorizes direct re…
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00065v1 Announce Type: new Abstract: Legible planning is the creation of plans that best disambiguate their goals from a set of other candidates from an observer's perspective. In this paper we propose a method for legible planning for arbitrary PDDL domains, by extending previous research on legibility to classical planning without requiring to construct ad-hoc planners. We also discuss how the observer perspective may be estimated through a second order theory of mind that connects the planner's and the observer's task spaces. Our solution can for example be deployed in human-robot teaming scenarios, where an autonomous robot in a team can implicitly communicate its goal by producing legible plans. We present benchmark results on several PDDL plann…
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00057v1 Announce Type: new Abstract: Guiding a tractor along a predefined reference path is a key component of precision agriculture. This study develops a path tracking controller based on Nonlinear Model Predictive Control, which incorporates multiple segments of a piecewise-linear reference path directly into the objective function. In addition, methods for selecting viable reference segments from the full path are presented. The control system is evaluated during a field test with a tractor controlled via the Tractor Implement Management steering interface. The NMPC solver converged on average after 3.45 ms and tracked the curved reference path with a mean absolute cross-track error of 6.1 cm.
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00008v1 Announce Type: new Abstract: Sim-to-real research pursues physics fidelity as a primary objective: simulators are judged by how closely they reproduce real-world contact dynamics. For governance benchmarking of LLM-driven robots, where the simulator demonstrates that an admission/policy/contract/audit pipeline behaves correctly, contact fidelity at object handoffs (grasp, carry, place) becomes a liability: contact-force integration noise injects audit-chain divergence that is structurally unrelated to the governance property under test. We propose bounded-fidelity sim-as-demo-stage, a design pattern that suppresses contact physics within explicitly bracketed handoff envelopes while preserving full dynamics elsewhere. The construction uses MuJ…
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00069v1 Announce Type: new Abstract: Long-horizon ego/exo data contains rich procedural evidence, but are redundant, noisy, and costly to process or retain. We propose a compact framework that converts continuous multimodal workplace video into a structured Procedural State Memory, implemented as a Work Environment Model (WEM). Inspired by event segmentation theory, we detect boundaries using changes in visual context, location, motion, narration, gaze/object interaction, and optional exocentric workspace evidence, rather than fixed windows or visual novelty alone. Each segment is abstracted into an evidence-linked event card containing actor, interval, location, action, objects/tools, pre/post state, confidence, and provenance. These event cards inc…
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00067v1 Announce Type: new Abstract: Reliable visual inspection is essential for quality assurance in aero-engine blade manufacturing, where defect appearance may vary across production lines, imaging conditions, blade poses, and surface backgrounds. Such domain shifts cause a mismatch between training and deployment data and degrade the reliability of deep defect detectors in online inspection. This problem is particularly challenging because aero-engine blade images usually contain sparse defects, making pseudolabel-based adaptation vulnerable to noisy or missing predictions. To address this issue, we propose Aero-engine Blade Defect Detector (ABDD), an online adaptive detection framework based on test-time adaptation. ABDD introduces a Dual-Alignm…
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00064v1 Announce Type: new Abstract: Assembly actions are compositional: they combine a manipulation with a part or tool. In deployment, systems routinely encounter novel combinations of familiar components, yet an atomic action classifier assigns every unseen combination exactly zero probability by construction. The prevailing solution is verb--noun decomposition, which predicts components separately and recombines them to reach unseen actions. While widely adopted, how decomposition generalizes under compositional shift remains poorly understood. We present a systematic analysis of verb--noun decomposition across three assembly datasets (MECCANO, HAViD, and IMPACT). Although decomposition escapes the atomic ceiling, its generalization extends only…
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00040v1 Announce Type: new Abstract: Recent advances in 3D Gaussian Splatting have enabled open-vocabulary and referring segmentation by distilling semantic knowledge from 2D foundation models into 3D representations. However, existing referring fields embed language features in a globally view-invariant space, making them fundamentally unable to resolve observer-centric spatial relations (e.g., "to the left of") that depend on camera pose. We propose DSSR-3D, an inference-time framework for view-dependent referring segmentation on continuous 3D Gaussian fields, formalized as two interfaces - pose-invariant semantic localization and pose-conditioned spatial reasoning - such that any pair of functions satisfying these constraints yields a valid instan…
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00031v1 Announce Type: new Abstract: Street-view imagery is increasingly used to infer urban attributes, but predictive accuracy alone does not reveal how much a photograph contributes beyond data already available for the same place. We compare image-based predictions with existing urban data across seven attributes from five public resources and three VLMs. The same urban units are evaluated using images, task context, nearby observations, and public records, while image replacements and conflicting records test source reliance. Existing urban data matched or exceeded image-only models for road damage, curb ramps, and house price, while neighbouring official statistics nearly matched the best image result for population. Images were more informativ…
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00030v1 Announce Type: new Abstract: Object detection models often experience performance degradation when deployed under distribution shifts, caused by for example changes in weather type, operational environment, or object appearance. Domain Generalization (DG) aims to develop models that remain robust under such shifts and generalize well to unseen domains. DG research specifically focused on object detection models is scarce, although these models face additional challenges around localization and multi-scale representations. Synthetic data is a promising tool to support in DG, by enabling large-scale generation of diverse new samples. In this paper, we present an object detection-centric review of DG and examine the role of synthetic data from t…
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00024v1 Announce Type: new Abstract: Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE -- the benchmark common to all three -- the mid layers encode the ground-truth answer in 68-91% of errors, yet this signal is not causally active for the final prediction: residual-stream patching yields 0% non-trivial flip at the layer level on all three architectures, and on two of three at the per-head level (Qwen: 0/12,600 patched forwards). The lone exception, InternVL3 layer-20 head-2, is a non-vocab, self-attending head whose effect is localized to that specific head (p < 1e-4). Despite the null, the errors separate operationally into th…
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2610.00017v1 Announce Type: new Abstract: We present Spatial Lifting (SL), a novel methodology for dense prediction tasks. SL operates by lifting standard inputs, such as 2D images, into a higher-dimensional space and subsequently processing them using networks designed for that higher dimension, such as a 3D U-Net. Counterintuitively, this dimensionality lifting allows us to achieve good performance on benchmark tasks compared to conventional approaches, while reducing inference costs and \textbf{drastically lowering the number of model parameters}. The SL framework produces intrinsically structured outputs along the lifted dimension. This emergent structure facilitates dense supervision during training and enables single-forward-pass self-consistency-ba…