AI News HubLIVE

研究动态

待翻译:Why AI Companies Are Buying–and Then Destroying–Old Books

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Why AI Companies Are Buying—And Then Destroying—Old Books | Open Culture --> --> --> --> --> --> Why AI Companies Are Buying—And Then Destroying—Old Books Why AI Companies Are Buying—And Then Destroying—Old Books in Boo…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Why AI Companies Are Buying—And Then Destroying—Old Books | Open Culture --> --> --> --> --> --> Why AI Companies Are Buying—And Then Destroying—Old Books Why AI Companies Are Buy…
站内正文

待翻译:Superhuman Attention

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:In the last three years a model learned to write in minutes what took a person a day, and the pace of change in codebases rose to match. Attention did not get cheaper. An engineer's focused hour is the same hour it was…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • In the last three years a model learned to write in minutes what took a person a day, and the pace of change in codebases rose to match. Attention did not get cheaper. An engineer…
站内正文

待翻译:This Linux and Windows app makes managing all your documents easier

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:If you're a student, writer, or professional juggling documents, this app can help keep you organized.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • If you're a student, writer, or professional juggling documents, this app can help keep you organized.
站内正文

待翻译:There are four types of AI sandboxes, maybe?

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The Four Sandbox Markets: A 2x2 Hypothesis TL;DR There are a lot of sandboxes these days. I was starting to lose my sanity. I therefore devised a 2x2 framework that categorizes sandboxes based on two dimensions, resulti…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • The Four Sandbox Markets: A 2x2 Hypothesis TL;DR There are a lot of sandboxes these days. I was starting to lose my sanity. I therefore devised a 2x2 framework that categorizes sa…
站内正文

待翻译:91% of professionals say their firm still falls short on AI - how to fix that

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Research suggests a gap between AI ambition and on-the-ground reality, but the good news is that professionals can fill it by focusing on well-grounded explorations and solid production use cases.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Research suggests a gap between AI ambition and on-the-ground reality, but the good news is that professionals can fill it by focusing on well-grounded explorations and solid prod…
站内正文

待翻译:What's Next for Robotics? Humanoids, Physical AI – Or Both?

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Aug 2026 What's Next for Robotics? Humanoids, Physical AI — or Both? Dr. Albert Meige Director, Blue Shift Dr. Karim Taga Managing Partner, Global Head of Functional Practices Dr. Oscar Mendez Director, AI & Data Scienc…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Aug 2026 What's Next for Robotics? Humanoids, Physical AI — or Both? Dr. Albert Meige Director, Blue Shift Dr. Karim Taga Managing Partner, Global Head of Functional Practices Dr.…
站内正文

待翻译:Show HN: I built a tool showing how AI providers (should) throttle their models

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:OP here: this project was born out of the frustration/paranoia that AI providers are throttling their models when their server load is too high. So, I set out to model and study the problem mathematically to understand…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • OP here: this project was born out of the frustration/paranoia that AI providers are throttling their models when their server load is too high. So, I set out to model and study t…
站内正文

待翻译:Why AI doesn't make companies more productive

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:In 1987, Nobel Prize–winning economist Robert Solow wrote, "You can see the computer age everywhere but in the productivity statistics." The same could be said about AI today. Gartner projects worldwide AI spending of $…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • In 1987, Nobel Prize–winning economist Robert Solow wrote, "You can see the computer age everywhere but in the productivity statistics." The same could be said about AI today. Gar…
站内正文

待翻译:Anthropic proposes plumbing spec to link AI agents to lab kit and robots

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Anthropic proposes plumbing spec to link AI agents to lab kit and robots Say you're trying to enrich Uranium and your centrifuges broke - soon it will be easy to connect an AI to figure out why Thomas Claburn Thomas Cla…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Anthropic proposes plumbing spec to link AI agents to lab kit and robots Say you're trying to enrich Uranium and your centrifuges broke - soon it will be easy to connect an AI to…
站内正文

待翻译:Quantization and Pruning Methods to Make Your LLM Leaner

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:This article walks through what each technique actually does, why skipping them costs real money and real latency, and then gets hands-on with five specific methods people are running in production right now.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • This article walks through what each technique actually does, why skipping them costs real money and real latency, and then gets hands-on with five specific methods people are run…
站内正文

待翻译:Tokens Aren’t Dollars

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The following article originally appeared on Tim O’Brien’s Medium blog and is being republished here with the author’s permission. AI costs are easy to count and hard to understand, and judging effort by a token volume? While that might feel like a valid measure of value or complexity, it doesn’t capture the details that define […]

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • The following article originally appeared on Tim O’Brien’s Medium blog and is being republished here with the author’s permission. AI costs are easy to count and hard to understan…
站内正文

待翻译:Are we just a couple steps away from a runaway AI?

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Are we just a couple steps away from a runaway AI? 28 August 2026 Are we just a couple steps away from a runaway AI? OpenAI published its post-mortem of the Hugging Face incident, and it is a fascinating read. During an…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Are we just a couple steps away from a runaway AI? 28 August 2026 Are we just a couple steps away from a runaway AI? OpenAI published its post-mortem of the Hugging Face incident,…
站内正文

待翻译:Making Your Data Ready for Agentic AI

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Making Your Data Ready for Agentic AI For thirty years we built data systems for human analysts, who supply the context, judgment, and skepticism to work around data that's incomplete or wrong. Autonomous agents supply…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Making Your Data Ready for Agentic AI For thirty years we built data systems for human analysts, who supply the context, judgment, and skepticism to work around data that's incomp…
站内正文

待翻译:IBM's new Granite 4.2 models ride the wave of interest in local LLMs

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:IBM has rolled out the newest models in its family of open-weight large language models designed to be downloaded and self-hosted. The newly launched Granite 4.2 comes in 3B, 8B, and 30B parameter variants. Like previou…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • IBM has rolled out the newest models in its family of open-weight large language models designed to be downloaded and self-hosted. The newly launched Granite 4.2 comes in 3B, 8B,…
站内正文

待翻译:Your AGENTS.md file doesn't do anything

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:AI coding bot vendors tell you to use a context file with instructions for the chatbot on how to edit your project. Claude Code wants a CLAUDE.md, or there’s AGENTS.md in general. [Anthropic] But does your AGENTS.md do…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • AI coding bot vendors tell you to use a context file with instructions for the chatbot on how to edit your project. Claude Code wants a CLAUDE.md, or there’s AGENTS.md in general.…
站内正文

待翻译:AI writing has begun to appear on the opinion pages

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Exclusive / AI writing has already begun to appear on the opinion pages Aug 26, 2026, 9:48pm EDT TechnologyMedia Illustration/Jake Angelo/Semafor PostEmailWhatsapp The Scoop Humans are still writing the vast majority of…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Exclusive / AI writing has already begun to appear on the opinion pages Aug 26, 2026, 9:48pm EDT TechnologyMedia Illustration/Jake Angelo/Semafor PostEmailWhatsapp The Scoop Human…
站内正文

待翻译:Anthropic's new hardware standard lets AI agents control the physical world

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:For all the interest in and uptake of agentic AI systems over the past year or so, the world of automated AI has thus far been primarily limited to text, images, code, and other data and actions that take place inside a…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • For all the interest in and uptake of agentic AI systems over the past year or so, the world of automated AI has thus far been primarily limited to text, images, code, and other d…
站内正文

待翻译:Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Google has released Gemini 3.5 Transcribe, a speech-to-text model that ships as two separate endpoints rather than one. The streaming endpoint delivers sub-second transcription but drops speaker diarization and word timestamps. The batch endpoint keeps both, at half the cost. Google reports 4.0% word error rate streaming and 2.6% non-streaming, with 70% faster finalization than Chirp 3. Here is what the split means for anyone building voice agents or transcription pipelines. The post Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages appeared first on MarkTechPost.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Google has released Gemini 3.5 Transcribe, a speech-to-text model that ships as two separate endpoints rather than one. The streaming endpoint delivers sub-second transcription bu…
站内正文

待翻译:Relaxation-Aware Multimodal Sensing of Soft Gripper Driven by Structure-Perception-Learning

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26622v1 Announce Type: new Abstract: Achieving stable, sustained grasping with soft robotic hands remains a fundamental challenge. Compliance enables safe and adaptive contact, yet the intrinsic viscoelasticity of soft polymers leads to stress relaxation and a continuous decay of grasping force during holding. Inspired by human grasping, which combines phase-dependent stiffness regulation with continuous sensing and feedback, this paper presents an integrated structure--perception--learning framework. We develop a variable-stiffness soft gripper that uses onboard vision and infrared thermography to track deformation and the temperature field in real time, preserving continuous tracking of the interaction state. To mitigate relaxation-induced force decay, we propose a temperature-coupled viscoelastic force representation, together with a physics-informed learning model, to reconstruct the force trend and provide explicit compensation during holding. Experiments show that, in a 280s force-controlled grasp-and-hold task, the proposed method maintains the desired force with a mean absolute error of 0.066N, outperforming fixed-aperture and instantaneous-only baselines by 80% and 95%, respectively. Overall, the results support a mechanism--AI co-design view: mechanisms shape feasible interactions, while learning compensates remaining uncertainty in viscoelastic dynamics, together enabling stable, sustained grasping.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26622v1 Announce Type: new Abstract: Achieving stable, sustained grasping with soft robotic hands remains a fundamental challenge. Compliance enables safe and adaptive…
站内正文

待翻译:SOLO: Stable Omni-terrain Long-Horizon Perceptive Humanoid Locomotion

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26583v1 Announce Type: new Abstract: Humans traverse complex terrain over long distances without losing balance, whereas perceptive humanoid policies become fragile as perception and control errors accumulate. We present SOLO, a unified framework addressing two compounding causes of this long-horizon fragility: dense terrain reconstruction smooths action-critical details, and pointwise imitation lacks temporal credit assignment. Its Query Reconstructor (QR) uses Fourier-encoded cell queries to retrieve spatially specific evidence from depth-proprioception tokens, preserving sharp terrain boundaries. Trajectory-Aware MSE (TA-MSE) Distillation adds next-state teacher-student disagreement to the PPO reward, enabling Generalized Advantage Estimation to propagate future disagreement penalties to preceding actions. In simulation, QR reduces height-map L1 error by factors of 3.3-4.0, while TA-MSE surpasses PPO and MSE+PPO in curriculum progression. On stress-test terrains, SOLO achieves 97.5% mean traversal success and 96% stepping-stone success, versus 75.0-75.6% and 0-3% for dense-reconstructor variants. Deployed zero-shot with only a chest-mounted depth camera and proprioception, SOLO completes a continuous 1.5-km outdoor route and an indoor mixed-terrain course. Project page: https://sunpihai-up.github.io/solo/

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26583v1 Announce Type: new Abstract: Humans traverse complex terrain over long distances without losing balance, whereas perceptive humanoid policies become fragile as…
站内正文

待翻译:TrapVLA: Trapping Vision-Language-Action Models in Configured Failure Modes

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26578v1 Announce Type: new Abstract: This work introduces Configured Failure Trapping, a novel backdoor attack task against Vision-Language-Action (VLA) models, which aims to activate attacks through stealthy textual triggers and induce configured failure modes. Unlike prior backdoor attacks that treat any task failure as a successful attack, Configured Failure Trapping requires the attacker to control how the robot fails (e.g., causing the robot to grasp with a specified positional offset), making it substantially more challenging and hard to detect. To support the new task, we propose an effective data engine for synthesizing high-quality target trajectories and an automated suite for measuring configured-failure fidelity. Then, based on this foundation, we construct two new benchmarks, namely Trap-LIBERO and Trap-RoboTwin, that instantiate Configured Failure Trapping across four representative failure modes. To address this task, we identify sparse action deviation as a critical challenge and accordingly propose a novel method named TrapVLA, which explicitly learns trigger-induced action residuals to steer the policy toward the configured failure behavior. Extensive experiments across simulation benchmarks and real-world robotic settings show that TrapVLA effectively injects configured failure modes into VLA models while largely preserving performance on clean data. Project page: https://john-liua.github.io/TrapVLA/

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26578v1 Announce Type: new Abstract: This work introduces Configured Failure Trapping, a novel backdoor attack task against Vision-Language-Action (VLA) models, which a…
站内正文

待翻译:Memory Anchors for Continual Robot Learning

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26545v1 Announce Type: new Abstract: Robot policies deployed in the wild should have the capability to continually learn new tasks without forgetting existing behaviors. A common approach to combat such catastrophic forgetting is to train on new task data with a replay buffer of previously learned task data. Although this buffer is commonly sampled randomly from all prior experiences, we show that a small set of these experiences contributes greatly in anchoring past performance. We call these experiences Memory Anchors. We identify Memory Anchors in regions where representations of new-task observations collapse onto those of old-task observations even though the tasks require conflicting actions, like when a familiar object must be manipulated in a new way. Rehearsing old data in this region plays a key role in preventing destructive overwriting of past task knowledge, serving as this critical Memory Anchor role. Excluding only 10% Memory Anchors before sampling the buffer leads to more than a 4.5x increase in catastrophic forgetting on the LIBERO benchmark suites. Conversely, enriching the replay buffer with Memory Anchors can decrease high-conflict task forgetting by 63% and enables successful continual learning of two task sequences on a real robot. Videos and additional visualizations can be found at https://robot-adaptation.github.io/MemoryAnchors

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26545v1 Announce Type: new Abstract: Robot policies deployed in the wild should have the capability to continually learn new tasks without forgetting existing behaviors…
站内正文

待翻译:Closing the Loop on the Poppy Humanoid: Bipedal Locomotion with Linear-Quadratic Control and Learned Cost Functions

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26505v1 Announce Type: new Abstract: The Poppy Humanoid is an open-source, low-cost robot suitable for research and education in artificial intelligence. However, we are unaware of any published methodology that achieves reliable, unassisted bipedal locomotion on the standard Poppy hardware. This paper contributes a functional closed-loop walking controller for Poppy, based on the linear-quadratic regulator (LQR) framework for trajectory tracking. Starting with data collected from open-loop playback of a nominal walking trajectory, our proposed method learns a quadratic cost function for an LQR controller that substantially improves the reliability of the motion. The closed-loop controller is validated empirically, demonstrating statistically significant improvements in walking performance compared to open-loop trajectory playback.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26505v1 Announce Type: new Abstract: The Poppy Humanoid is an open-source, low-cost robot suitable for research and education in artificial intelligence. However, we ar…
站内正文

待翻译:RTNav: Towards Real-Time Zero-Shot Object Navigation

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26496v1 Announce Type: new Abstract: Navigation in unknown environments to find unforeseen objects has become increasingly feasible with capable vision and language foundation models. However, these models also introduce non-negligible inference latency, which becomes an important concern when agents must operate continuously in the real world. Most state-of-the-art methods are still developed in synchronous simulators, where the environment waits for the agent to act and inference time is effectively free. As a result, agents are often designed around the sequential execution of perception, reasoning, and action, with little regard for time constraints. Under real-time execution, where wall-clock time counts towards the task budget, the inefficiencies of these architectures become clear. We show that recent zero-shot object navigation methods suffer consistent performance degradation under such realistic timing conditions. Motivated by this observation, we propose RTNav, a simple but effective architecture that treats inference latency, asynchronous environment stepping, and bounded compute as explicit design considerations. Evaluated on real-time variants of HM3D-v1, HM3D-v2, and HM3D-OVON, RTNav improves the success rate by up to 11% and the Success weighted by Completion Time by up to 5.1 points over prior work.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26496v1 Announce Type: new Abstract: Navigation in unknown environments to find unforeseen objects has become increasingly feasible with capable vision and language fou…
站内正文

待翻译:Cross-Platform Benchmark of Neural 3D Reconstruction for Autonomous Laboratory Robots

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26383v1 Announce Type: new Abstract: Autonomous robots performing laboratory tasks depend on 3D reconstruction pipelines that can turn raw camera streams into actionable object representations within the latency budget of a physical control loop. Neural 3D reconstruction methods have demonstrated high-quality view synthesis, but their real-time viability across the compute platforms on which laboratory robots actually run remains poorly characterized. In this work, we present a systematic compute-platform benchmark of neural 3D reconstruction methods, evaluating NeRF and 3D Gaussian Splatting training and rendering on GPU-enabled computing devices ranging from single-board computers to server-class nodes, and place Meta's SAM3D single-image reconstruction on the same axes to quantify its latency and fidelity gap relative to per-scene optimization. Our results show that Gaussian Splatting yields higher rendering quality than NeRF at greater GPU cost, and that onboard compute is insufficient for full per-scene optimization at interactive rates. Our preliminary assessment on SAM3D indicates that it delivers plausible object geometry within seconds, but with detail mismatches that can compromise downstream manipulation. Together, these findings motivate tiered pipelines in which lightweight feed-forward reconstruction sustains the real-time perception-and-tracking loop for laboratory robots, while heavier neural reconstruction is scheduled selectively on suitable compute.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26383v1 Announce Type: new Abstract: Autonomous robots performing laboratory tasks depend on 3D reconstruction pipelines that can turn raw camera streams into actionabl…
站内正文

待翻译:Dispersive Forward Tree Search for Optimal Control: Coverage, Complexity, and Computation

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26314v1 Announce Type: new Abstract: Steering-based planners require solutions to state-to-state boundary value problems, which can be inaccessible for nonlinear platforms. Forward propagation evades the steering requirement, but the finite-sample behavior of the associated planners remains uncharacterized and their implementations underperform in practice. This paper develops a propagation-based kinodynamic planner with deterministic finite-sample near-optimality guarantees. We work within the large class of differentially flat nonlinear systems and show that a forward tree of locally dispersive control commands contains a near-optimal trajectory at a certified tree size. We provide a general mechanism to construct dispersive command sets for control-affine systems, which are necessary to implement the search algorithm prescribed by the theory. We show that covering the certified trajectory class irrespective of cost provably demands a tree exponentially sized in the problem horizon, and present a cost-conditioned dominance pruning procedure that retains near-optimality at a tree size polynomial in the horizon. We implement the resulting search algorithm, Dispersive Forward Tree search (DFT*), as breadth-first expansion of the forward tree, which maps naturally onto parallel hardware. We design efficient dispersive samplers for the unicycle, the trailer car, and the quadrotor and evaluate challenging planning tasks for these platforms. DFT* delivers consistently competitive and often substantially better solution quality than state-of-the-art kinodynamic planners at comparable solution times on embedded-tier processors, accelerating further as parallel compute is scaled. We also implement DFT* in a receding-horizon loop to demonstrate real-time planning in dynamic environments at embedded-tier compute budgets.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26314v1 Announce Type: new Abstract: Steering-based planners require solutions to state-to-state boundary value problems, which can be inaccessible for nonlinear platfo…
站内正文

待翻译:Constraint-Aware Physics-Informed Neural Networks for Static Shape Estimation of Co-Manipulative Continuum Robots

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26273v1 Announce Type: new Abstract: Static shape estimation of co-manipulative continuum robots (CCRs) is challenging because the continuum arms and manipulated flexible object form a closed chain that must satisfy both static equilibrium and geometric loop-closure constraints. This paper presents a constraint-aware physics-informed neural network (PINN) for static shape estimation of a tendon-driven CCR modeled using the geometric variable strain formulation. The proposed method incorporates a projected static equilibrium residual and a configuration-level geometric residual to enforce the governing mechanics and closed-chain geometry. In simulation, the PINN is compared with a purely data-driven artificial neural network (ANN) under limited and noisy training data. With 140 samples and 50% label noise, the PINN reduces the relative configuration error, equilibrium residual, and closed-chain residual by 67.88%, 67.35%, and 88.06%, respectively. Using the full dataset, the PINN achieves 0.1597% relative configuration error with an inference time of 0.1773 ms, compared with 17.97 s for an iterative nonlinear solver. Experimental fine-tuning reduces the marker RMSE from 2.657 mm to 0.497 mm and increases R2 from -0.788 to 0.937. These results demonstrate accurate, physically consistent, and computationally efficient static shape estimation of closed-chain CCRs.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26273v1 Announce Type: new Abstract: Static shape estimation of co-manipulative continuum robots (CCRs) is challenging because the continuum arms and manipulated flexib…
站内正文

待翻译:WALL-SS: Scaling Long-horizon World Models via Next-Scale Autoregression

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26239v1 Announce Type: new Abstract: Generative world models provide robots with predictive models of how the world evolves under interaction, with growing potential for simulation, planning, policy evaluation, and robot learning. Beyond clip-level future prediction, a unified generative formulation should relate actions to consequences, support flexible horizons and continuous interaction, and enable reward-driven optimization. We introduce WALL-SS, a world model that generates visual futures through Scale-wise autoregressive Scaling, enabling action-controllable and long-horizon robotic simulation. WALL-SS represents embodied trajectories as causal sequences of temporally interleaved observations and actions, making action-dependent state transitions explicit while naturally supporting variable-length generation, streaming extension through reusable causal states, and direct optimization through sequence probabilities. To make this formulation effective over long horizons, we generate each future observation in a coarse-to-fine manner and develop three complementary components within the same hierarchy. Action-conditioned next-scale prediction injects scale-aligned action representations to improve action-future coupling and model both successful and failed behaviors. Scale-compressed long-horizon memory retains recent interactions at fine resolution while compressing distant observations and actions, with scale-wise dream forcing enhancing robustness to self-generated context. Finally, on-policy alignment optimizes autoregressive visual dynamics with action-following and long-term consistency rewards while preserving the pretrained visual distribution. Experiments show that WALL-SS improves action following and trajectory accuracy, supports coherent minute-long streaming rollout under bounded memory, and consistently benefits from on-policy alignment in reducing action drift and long-horizon inconsistency.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26239v1 Announce Type: new Abstract: Generative world models provide robots with predictive models of how the world evolves under interaction, with growing potential fo…
站内正文

待翻译:Video-FLAIR: Not Whether to Reason, But How

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26495v1 Announce Type: new Abstract: Multimodal queries can require different types of reasoning. Some can be answered via perceptual reasoning, extracting information directly from the visual signal, while others require compositional reasoning that combines observations or deliberative reasoning that evaluates competing hypotheses. However, many existing methods apply a uniform reasoning strategy across queries, leading to unnecessary computation on simple tasks and insufficient reasoning on complex ones. We introduce Video-FLAIR, a training framework that learns to select the appropriate reasoning mode for each query using reinforcement learning. During training, the model generates responses under all three modes for the same prompt, enabling direct comparison. A composite reward compares these responses to favor the most effective one based on correctness, grounding, and cost, while discouraging unsupported or misaligned deliberation. This yields a supervision signal for learning adaptive reasoning without per-query annotations. Video-FLAIR improves accuracy over the Qwen2.5-VL base model by +5.4 on MathVista, +4.8 on Video-Holmes, and +4.8 on Video-MMMU, while reducing average token usage to 95 compared to 417 for always-thinking baselines.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26495v1 Announce Type: new Abstract: Multimodal queries can require different types of reasoning. Some can be answered via perceptual reasoning, extracting information…
站内正文

待翻译:Learning Woody Clearing With Loss Alignment for Zero-Shot Regrowth and Woody Segmentation

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26489v1 Announce Type: new Abstract: Detecting woody clearing is vital for managing biodiversity. Deep learning models can detect change in woody vegetation from bitemporal remote sensing imagery, however generated products may not meet end-user specifications due to unaligned loss definitions. Further limitations of deep learning models are the reliance on large datasets which can be difficult to attain for spatially rare and ambiguous events such as regrowth detection. In this work we train a model to detect woody change using bitemporal Sentinel-2 imagery consisting of 7 years' worth of annual imagery across the state of New South Wales, Australia. To align the objective of the model with end-user metrics, we introduce the loss scaling coefficient $\alpha$ which transforms the objective to optimize for specific $F_{\beta}$ scores. Introducing $\alpha$ was found to increase precision by 1.85x or recall by 1.12x. We propose input imagery augmentation and generation techniques that allow the woody change detection model to zero-shot transfer to regrowth and woody segmentation tasks. For woody segmentation, image generation techniques using activation maximization with low $\alpha$ values for stability and image generation techniques derived from handcrafted features utilizing a mosaic of clearing patches and artificial trees for contextual grounding were found to outperform prior woody segmentation works of the study area, reducing the overall error by up to 18.2%. For zero-shot woody regrowth, creating pseudo-post and prior images resulted in the model achieving an F1 score of 0.845, creating a foundation for future regrowth detection work.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26489v1 Announce Type: new Abstract: Detecting woody clearing is vital for managing biodiversity. Deep learning models can detect change in woody vegetation from bitemp…
站内正文

待翻译:Zero-Shot Video Restoration and Enhancement with Text-to-Image Latent Diffusion Models and Multi-Modal References

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26476v1 Announce Type: new Abstract: Zero-shot image restoration methods with text-to-image latent diffusion models have achieved great success in universal image restoration tasks without training. However, applying them to video restoration will result in severe temporal flickering. In this paper, we propose a novel framework for zero-shot video restoration and enhancement which uses a text-to-image latent diffusion model and multi-modal references. Through the proposed dual prompt tuning inversion and sampling, the inference time can be reduced to nearly 1/3 of the original. The performance and temporal consistency can be also significantly stregthened. By using the proposed texture-aware video token merging, the temporal correlation between frames can be further utilized to improve the temporal consistency. We futher propose the referenced self-attention and referenced token merging to support image reference. Experimental results demonstrate the superiority of the proposed method in restoring and enhancing temporally consistent videos.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26476v1 Announce Type: new Abstract: Zero-shot image restoration methods with text-to-image latent diffusion models have achieved great success in universal image resto…
站内正文

待翻译:Mapping Woody Vegetation from Multi-Source Imagery and Prediction Fusion for Enhanced Data Efficiency and Accuracy

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26471v1 Announce Type: new Abstract: Tree cover maps are a fundamental remote sensing product, used to derive ecological insights about the landscape and are essential to change detection, vegetation mapping and fire monitoring programs. However, comprehensive tree cover mapping requires reliable and high-quality imagery, free of cloud and weather defects to ensure accurate model outputs. Deep learning approaches can generate high quality maps with minimal human intervention but require large amounts of human annotated data to be successful. In this work we propose a framework consisting of methods that aim to improve the data efficiency and robustness of deep learning models using data fusion techniques to segment woody vegetation defined as vegetation over the height of 2m across the state of New South Wales, Australia. To improve robustness against varying image quality, we propose an image composition method that normalizes the imagery and removes defects, whilst also minimizing the reliance on individual image quality by proposing a prediction fusion method. The two methods resulted in an error reduction of 38.2% and 53.6% respectively compared to single-source imagery. To address deep learning approaches' limitation of requiring large amounts of data, we apply label transfer to multiple sources of imagery as a form of data augmentation to improve data efficiency. Learning from multiple image sources was shown to be the biggest improvement in performance, resulting in an error reduction between 28.1% to 76.2% across the different validation experiments, whilst reducing the standard deviation of performance across image dates by a factor of 13.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26471v1 Announce Type: new Abstract: Tree cover maps are a fundamental remote sensing product, used to derive ecological insights about the landscape and are essential…
站内正文

待翻译:VIPER: An Expert-Curated Benchmark for Vision-Language Models in Veterinary Pathology

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26382v1 Announce Type: new Abstract: Pathology vision-language models are advancing rapidly, yet existing benchmarks remain focused on human tissue, particularly oncology, leaving non-human pathology largely unaddressed. This gap is especially important in toxicologic pathology, where microscopic tissue examination of laboratory animals is a core component of preclinical drug safety assessment. To address it, we introduce VIPER, the first expert-curated benchmark for vision-language model evaluation in toxicologic pathology. VIPER contains 1,251 questions associated with 419 H&E-stained rat histology images across seven organ systems, covering multiple-choice, KPrim, and free-text formats. All questions were curated and validated by board-certified veterinary pathologists. In total, we benchmarked 16 models, including two newly introduced veterinary-pathology models, seven human pathology-specialized models, and seven general-purpose frontier models. The results identify a substantial domain gap between veterinary and human pathology, expose the risk of over-diagnosis of normal tissue in frontier models, and show that domain-specific training remains critical for visually grounded predictions. VIPER data and evaluation code are available at https://github.com/mahmoodlab/viper.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26382v1 Announce Type: new Abstract: Pathology vision-language models are advancing rapidly, yet existing benchmarks remain focused on human tissue, particularly oncolo…
站内正文

待翻译:A Unified Framework for the Mechanics of Information in Convolutional Neural Network Image Space

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26363v1 Announce Type: new Abstract: This paper introduces a unified mathematical framework for modeling information propagation through convolutional neural networks (CNNs), with the aim of connecting descriptions of physical space and information space. A correspondence is presented linking discrete filter symmetry and the relativistic energy--momentum relation under the widely used nonlinear rectified convolution operation. Specifically, symmetric filter components (e.g. the sum $\Sigma = [1,1]$) operate analogously to rest energy $mc^2$ in preserving the image centre of mass (e.g. isotropic diffusion), whereas antisymmetric components (e.g. the gradient $\nabla = [-1,1]$) operate analogously to the momentum term $pc$ in generally inducing a displacement (e.g. vibration or translation). For typical small discrete filters, this displacement is determined by the ratio of antisymmetric to total filter energy, analogously to how the displacement of a relativistic particle relates to a Lorentz transform with beta parameter $\beta = \frac{v}{c}=\frac{pc}{E}$ equal to the ratio of momentum $pc$ to total energy $E$. Repeated filtering leads to the Gaussian scale-space and emergent scale-invariant features. These constructions share a Laplacian-driven structure with the classical heat (diffusion) equation and, via standard mathematical correspondences, with the Schr\"odinger equation and aspects of the Friedmann equations, together with emergent Morse topological structure. Demonstrations in 3D images reveal blob-like, scale-invariant Morse critical points in images spanning a wide range of physical scales, including organic sugar molecules and inorganic silicon crystals, human and primate brains in magnetic resonance images (MRI), galaxies and the cosmic microwave background (CMB).

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26363v1 Announce Type: new Abstract: This paper introduces a unified mathematical framework for modeling information propagation through convolutional neural networks (…
站内正文

待翻译:Finding the Right Evidence: Factor-Guided Coarse-to-Fine Reasoning for Long Videos

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26355v1 Announce Type: new Abstract: While LVLMs rapidly improve, long-video question answering still remains challenging: relevant evidence is sparse, and question-relevant context often fails to provide cues that discriminate the correct answer from plausible alternatives. Diagnostic analysis on a manually annotated subset of MMR-V shows that prior agentic systems substantially improve cue retrieval over direct VLM inference yet fail to achieve a corresponding gain in answer accuracy, indicating that the bottleneck lies in option-discriminative evidence rather than topical relevance alone. We propose PACE (Progressive Acquisition of Critical Evidence), a factor-guided framework for long-video evidence acquisition. PACE proceeds in two stages: it first indexes clip-level descriptions guided by question-derived factors without observing the candidate answers; it then uses the candidate answers to derive contrastive cues and queries the index for verification. On MMR-V with the open-source Qwen3-VL backbone, PACE achieves 42.6% accuracy, outperforming direct inference and prior agentic baselines including Deep Video Discovery (DVD). On the same diagnostic subset, PACE recovers 66.9% of the annotated cues, providing empirical evidence that its gains are associated with improved evidence recovery rather than stronger answer-side priors alone. Consistent gains over DVD on LVBench, Video-MME, EgoSchema, and LongVideoBench suggest that option-aware evidence acquisition transfers beyond MMR-V. Code is available at https://github.com/HKUST-KnowComp/PACE.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26355v1 Announce Type: new Abstract: While LVLMs rapidly improve, long-video question answering still remains challenging: relevant evidence is sparse, and question-rel…
站内正文

待翻译:Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26317v1 Announce Type: new Abstract: Frontier language models are increasingly marketed as omni systems that can perceive and respond across modalities. Existing evaluation frameworks, however, focus almost exclusively on bimodal understanding, typically text plus one other modality. We propose the Modality Maturity Index (MMI), a benchmark designed to evaluate the multimodal capabilities of large language models across five modalities (text, image, audio, video and document) and combinations of up to three modalities in both inputs and outputs. MMI consists of 893 questions, each carefully crafted to require the model to demonstrate its understanding of multiple input modalities and to generate responses that incorporate various output formats. The questions are designed to be self-contained, with clear expectations for the correct modality or mix of modalities required for an accurate response. Every MMI prompt carries human-authored rubric criteria for each output modality expected in the response; a model's MMI Value expresses the average of the per-modality scores for each prompt. Because low scores can reflect either failure to generate a modality (lack of presence) or failure to generate correct content, we introduce also a supplementary Modality Presence Score (MPS), a per-prompt F1 over the expected output modalities. Applying MMI to five frontier multimodal models, we find that the MPS ranges from only 15.6 (Claude Opus 4.6) to 34.9 (GPT-5.4). Given the low availability of returned modalities to even grade, we report MPS as our main result pending model improvements. To assess the viability of judging output correctness with LLM judges and rubrics, we run a separate experiment with custom generation tools. On the assets that generates, we find that an LLM judge applying the rubrics agrees with rubric-blind human annotators (who score the outputs directly and never see the criteria) on 70.8% of judgments.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26317v1 Announce Type: new Abstract: Frontier language models are increasingly marketed as omni systems that can perceive and respond across modalities. Existing evalua…
站内正文

待翻译:Procedura: Agentic 3D Modeling with Procedural Control

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26238v1 Announce Type: new Abstract: Native 3D generators now recover impressive mesh geometry from a single image. However, a dense mesh stays soft where a machined object should be sharp, it carries no part decomposition, and it exposes no parameter a user could edit. To address this, we explore the paradigm of 3D shape as code, leveraging and scaling the coding ability of an LLM for 3D modeling. We introduce Procedura, a novel 3D modeling agent framework that writes an object as a procedural assembly, a parametric program whose named parts are joined by typed, machine-checkable mates. From a text prompt, the agent plans the object as an assembly graph and writes the program part by part, solving each placement from the mated frames rather than guessing it, and admitting a part only once compile, mate, and connectivity checks pass. A decoupled vision critic then refines the assembly one diagnosed fix at a time. Moreover, the same graph carries per-part materials and a simulator-validated articulation. We evaluate on P3D-Bench under its assembly judge, and with the same judge on MechBench-36, our hard-surface benchmark. On both, Procedura outperforms state-of-the-art native 3D generators and every prior 3D-code agent on judged quality, produces the sharpest edges of any method we evaluate, and is the only one whose output is an editable, part-structured program.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26238v1 Announce Type: new Abstract: Native 3D generators now recover impressive mesh geometry from a single image. However, a dense mesh stays soft where a machined ob…
站内正文

待翻译:Surgical Video Generation From Diffusion to World Models: A Survey

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26214v1 Announce Type: new Abstract: Surgical video data provides the primary training resource for models of intraoperative perception, surgical workflow understanding, and robotic decision-making. However, clinical data acquisition remains constrained by privacy, cost, and class imbalance. Surgical video generation has emerged as a transformative approach to addressing data scarcity and as a foundation for surgical simulation, training, and robotic policy learning. The field has developed rapidly without a clear conceptual framework. This survey organizes the 2024-2026 literature into three categories: unconditional generation, conditional generation, and world modeling generation, revealing a fundamental shift in how the task is defined from synthesizing visually plausible frames to modeling the causal dynamics of surgical scenes. We examine the persistent gap between pixel-level fidelity and clinical plausibility, and identify generalization, physical realism, controllability, and interpretability as bottlenecks. We further summarize experimental results of representative methods on public datasets to provide a quantitative reference for the field. This survey provides a structured overview of the current state and open challenges, offering a reference for researchers working at the intersection of intelligent perception, multi-modal fusion, generative AI, and surgical data science.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26214v1 Announce Type: new Abstract: Surgical video data provides the primary training resource for models of intraoperative perception, surgical workflow understanding…
站内正文

待翻译:FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26129v1 Announce Type: new Abstract: Scientific peer review datasets have trained AI systems exclusively on Computer Science and Machine Learning venues, producing models that critique ablation studies yet have never seen a biology reviewer demand contamination controls or a chemist question Nuclear Magnetic Resonance (NMR) spectral assignments. We introduce FIRSTPASS, the first large-scale peer review dataset built on complete multi-round editorial dialogues from a multidisciplinary high-impact journal. Curated from Nature Communications mandatory transparent peer review (instituted November 2022), FIRSTPASS comprises 3,668 records spanning five scientific domains (biology, chemistry, neuroscience, physics, and earth science), capturing the full iterative structure of scientific validation: initial referee reports, author point-by-point responses, and updated reviewer assessments. Each record carries an outcome label derived directly from editorial decisions (STANDARD for two-round review; EXTENDED for three or more rounds), providing ground truth absent in all prior corpora. An automated audit confirms 100% content integrity. Expert reviews average 2,155 words, substantially denser than conference venue reviews. All data, parsing pipelines, and evaluation scripts are released to enable reproducible benchmarking of AI scientific judgment across disciplines.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26129v1 Announce Type: new Abstract: Scientific peer review datasets have trained AI systems exclusively on Computer Science and Machine Learning venues, producing mode…
站内正文

待翻译:TelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26126v1 Announce Type: new Abstract: Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows require joint grounding in normative specifications, operational telemetry, vendor-specific fault evidence, and exact RF/network calculations. However, current LLM integration in telecom remains bottlenecked by a two-sided capability gap: generic reasoners often lack telecom-specific grounding, while domain-specific telecom LLMs remain limited in structured, multi-step reasoning. To bridge this gap, we release TelecomGPT-R1-9B, a unified open-source telecom reasoner that ranks top-performing on the GSMA open telco leaderboard. Specifically, we curate a 67,427-example supervised fine-tuning (SFT) corpus organized around four complementary reasoning axes: protocol, knowledge, modeling, and fault. The corpus is built from axis-matched public web sources and enhanced through axis-specific chain-of-thought (CoT) generation and prefix-continuation self-validation. Starting from Qwen3.5-9B, we further develop a two-stage post-training recipe. First, multi-teacher low-rank adaptation (LoRA)-based SFT injects telecom knowledge and induces axis-specific reasoning formats. Second, group relative policy optimization (GRPO), stabilized by decoupled clip and dynamic sampling policy optimization (DAPO), optimizes the policy using four axis-aligned binary verifier rewards. Across seven public telecom benchmarks, TelecomGPT-R1-9B ranks first among open-source telecom LLMs and achieves a seven-axis mean comparable to state-of-the-art closed-source frontier reasoners.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26126v1 Announce Type: new Abstract: Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows r…
站内正文

待翻译:Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26125v1 Announce Type: new Abstract: Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation, particularly when detached from sociocultural context. We propose a \emph{training-time} explainability framework that aligns model reasoning with human-annotated rationales, improving both classification performance and interpretability. Our approach is evaluated on HateXplain (English) and BullySent (Hinglish), reflecting the prevalence of anti-Muslim hate across both languages. Using LIME, Integrated Gradients, Grad X Input, and attention, we assess accuracy, explanation quality, and cross-method agreement. Results show that gradient- and attention-based regularization improve F-scores, enhance plausibility and faithfulness, and capture culturally specific cues for detecting implicit anti-Muslim hate, offering a path toward multilingual, culturally aware content moderation.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26125v1 Announce Type: new Abstract: Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation.…
站内正文

待翻译:Natural-Language Policies to Executable Decisions: An Interpretable Large Language Model Framework

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26124v1 Announce Type: new Abstract: Pricing automation in large-scale tourism is challenging because travel orders are highly unstructured, while pricing policies are complex, rapidly evolving, and inherently open-ended. Traditional rule engines are brittle and costly to maintain, whereas unconstrained LLM agents lack the reliability and auditability required for financial decisions. We present a production-grade LLM-powered pricing system with a strict decision boundary: LLMs perform structured extraction and bounded policy/path selection, while all numeric pricing, including total-price computation, is executed deterministically. Policies are compiled into interpretable condition trees, enabling open-ended support for new clauses and evolving rules without code changes, while exposing auditable artifacts for human-in-the-loop control. Periodic fine-tuning on logged traces further improves tree induction and path matching. Deployed at a municipal state-owned tourism enterprise across 7 scenic sites and 12 business categories with 1,500+ operators and 1,000+ active policies, the system processed 3,960 orders in six months, reduced the order management team from 15-20 to 3, and cut per-order handling time from 10 minutes to <2 minutes.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26124v1 Announce Type: new Abstract: Pricing automation in large-scale tourism is challenging because travel orders are highly unstructured, while pricing policies are…
站内正文

待翻译:Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26123v1 Announce Type: new Abstract: Large language models (LLMs) are trained predominantly on English-language internet text that over-represents certain cultural narratives, raising concerns that models flatten the diversity of non-Western storytelling traditions into a single homogenized archetype. We present a pilot computational study examining this across three maximally distinct Indian regional oral and literary traditions: the Rajasthani Pabuji epic, classical Tamil Sangam poetry, and Bengali folk tales. We collected authentic reference corpora for each tradition (11, 21, and 10 passages respectively) and prompted two LLMs (Claude Sonnet and Gemini) with 54 generation requests spanning three prompt types per tradition - generic, culturally specific, and regional-language. Using Sentence-BERT embeddings and cosine similarity, we measure reference drift (how closely outputs track their own tradition's authentic texts relative to the other two) and cross-tradition convergence (how similar outputs are across traditions). We find that while outputs remain closer to their own tradition's reference than to others, cross-tradition similarity is high (0.52-0.66) relative to what the traditions' genuine distance would predict, indicating partial homogenisation. Unexpectedly, prompting in the regional language (Hindi, Tamil, or Bengali) consistently reduced fidelity to the authentic tradition relative to English prompting, by as much as 27 percentage points for Rajasthani and Bengali traditions. We discuss this against conflicting prior results on multilingual prompting and argue it reflects a difference between eliciting general cultural diversity and simulating one narrow, lesser-documented oral tradition. We position this pilot as a lightweight, scalable complement to recent large-scale human-annotation studies of Indian cultural misrepresentation in LLM-generated stories, as part of a broader doctoral research program.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26123v1 Announce Type: new Abstract: Large language models (LLMs) are trained predominantly on English-language internet text that over-represents certain cultural narr…
站内正文

待翻译:Can a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for Abstention

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26121v1 Announce Type: new Abstract: Large language models state false facts as fluently as true ones, yet a model often "knows" internally when it is on shaky ground: the probability it assigns to its own answer tends to dip on the facts it gets wrong. The usual way to act on this, teaching a model to abstain rather than guess, requires a labelled dataset of right and wrong answers. We ask whether the model's own confidence, which is free and needs no labels, can do that job instead. We fine-tune each model (with LoRA) to answer when its frozen confidence is high and to say "I'm not sure" when it is low, using the signal alone and no correctness labels. Across six open-weights models (1B-8B, two families) on short-form factual question answering, with correctness adjudicated by an independent judge model, this label-free recipe holds its own against label-supervised abstention-tuning: at matched coverage we find no statistically detectable difference between the two. A control that drills hard examples instead of abstaining does not help, indicating the gain comes from calibration, not rote memorization. The signal's one blind spot is confidently wrong facts, which it cannot flag. A model's own doubt is thus a near-free substitute for a labelled dataset when teaching it when to abstain. Code and artifacts are available on request.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26121v1 Announce Type: new Abstract: Large language models state false facts as fluently as true ones, yet a model often "knows" internally when it is on shaky ground:…
站内正文

待翻译:Recipes for Steering and Scaling LLMs via Sampling

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26120v1 Announce Type: new Abstract: Large Language Models (LLMs) are probabilistic models, typically defined by an autoregressive factorization. While recent work has begun to study richer target distributions beyond the base model, the sampling strategies remain highly inefficient. In this paper, we present a flexible and theoretically grounded framework for steering and scaling autoregressive LLMs with sampling. Within this framework, we describe two algorithms -- one based on Sequential Monte Carlo (SMC) and one based on Replica Exchange (RE) -- that steer generation toward powering, product or tilting of the base model distribution. We illustrate this framework through scaling the generation quality of LLMs without external supervision or reward models. Experimental results demonstrate our methods scale more favorably than Best-of-N and standard MCMC baselines. Overall, this paper offers a systematic recipe for probabilistic inference with LLMs via sampling.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26120v1 Announce Type: new Abstract: Large Language Models (LLMs) are probabilistic models, typically defined by an autoregressive factorization. While recent work has…
站内正文

待翻译:DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26119v1 Announce Type: new Abstract: Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training constrains this behavior, has received less attention than the related question of detecting fallacies in existing text. We close this gap with DeflectBench, evaluating 23,990 generations from four frontier models across three deflection strategies (whataboutism, ad hominem, red herring), seven prompt framings, and 80 claims spanning four controversy levels. Refusal is governed primarily by request structure rather than claim content. Per claim refusal varies by only 11 percentage points across the 80 claims, while a single prompt frame change can swing within model refusal by nearly 100 percentage points and switching the requested fallacy type can swing it by over 80 percentage points within explicit framings. An educational debate coach prompt framing collapses refusal to near zero across all four model families, but the bypassed behavior is not clean compliance. Models typically produce labeled compliance, naming the requested manipulation in the same response that contains it. The four models distribute differently across refusal, labeled compliance, soft refusal, and clean compliance. The code and dataset are released at https://github.com/ArtKanke/DeflectBench.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26119v1 Announce Type: new Abstract: Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training…
站内正文

待翻译:ElementCheck: Complexity-Aware Long-Form Text Factuality Evaluation via Sentence Elements

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26118v1 Announce Type: new Abstract: Existing long-form factuality evaluation relies on the decompose-retrieve-verify pipeline. However, the pipeline suffers from noise from claim decomposition and fixed verification granularity, resulting in unreliable results. We propose ElementCheck, a complexity-aware framework that verifies long-form outputs via sentence elements. Instead of uniformly decomposing sentences into atomic sub-claims, ElementCheck extracts entity pairs that are explicitly linked through verifiable connections in the original sentence as elements, and organizes these into an element graph. The graph topology provides a structural signal for estimating sentence complexity, enabling direct verification for simple sentences and targeted element-level refinement and verification for complex ones. To support fine-grained evaluation, we construct a new benchmark FastFact-Sent by mapping isolated claims from FastFact-Bench back to their source sentences. Experiments on FastFact-Sent and two domain-specific benchmarks show ElementCheck consistently improves factuality verification across five backbone models while maintaining a favorable accuracy-cost trade-off. Further analyses demonstrate that complexity-aware verification reduces unnecessary re-verification and maintains stability across different backbones.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26118v1 Announce Type: new Abstract: Existing long-form factuality evaluation relies on the decompose-retrieve-verify pipeline. However, the pipeline suffers from noise…
站内正文

待翻译:TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decoding

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26112v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference through a draft-then-verify paradigm. Building on this, tree-structured methods improve inference by organizing proposals into multiple candidate paths, increasing the accepted length. However, existing tree-structured methods use a single drafter for all drafting steps, creating a dilemma: a smaller drafter is fast but yields lower-quality trees, whereas a larger drafter improves tree quality but suffers from high latency. To address this, we propose TreeGraft, a multi-drafter framework in which drafters of different costs jointly construct a shared draft tree. TreeGraft uses the stronger drafter to rescore candidates by updating scores assigned by the weaker drafter, reselect grafting positions, and recover promising paths left unexplored. It also integrates stronger drafter expansions non-destructively, preserving existing branches that may still be accepted by the target model. Together, these designs improve the quality of the shared draft tree. To control the drafting cost, TreeGraft introduces a lightweight scheduler distilled from an offline value system to decide when to call the stronger drafter. Across 10 model pairs and 6 benchmarks, TreeGraft outperforms the better of the two fixed single-drafter endpoint strategies by 15.1% on average, reaching a maximum gain of 26.6%. Our code is available at https://anonymous.4open.science/r/TreeGraft-E983.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26112v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference through a draft-then-verify paradigm. Building on this, tree-struct…
站内正文

待翻译:FedCMAPSS: A Benchmark for Federated Learning in Remaining Useful Life Estimation

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.26433v1 Announce Type: new Abstract: Data-driven prognostics and health management has emerged as a key enabler for Industry 4.0, yet the development of robust remaining useful life (RUL) estimation models is often limited by the scarcity of run-to-failure data. While federated learning offers a promising paradigm to collaboratively train predictive models without sharing sensor data, research efforts have operated so far in the absence of a common evaluation framework. To address this gap, this paper introduces FedCMAPSS, a benchmark for federated RUL estimation based on the commonly-used NASA C-MAPSS dataset. We define a set of five standardized tasks designed to simulate real-world industrial challenges, ranging from ideal IID settings to extreme statistical heterogeneity, and conduct a systematic evaluation of state-of-the-art federated optimization algorithms across multiple neural architectures. By establishing reproducible baselines and making the source code and data splits publicly available, this work aims to provide a standard foundation for developing and comparing federated predictive maintenance solutions.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.26433v1 Announce Type: new Abstract: Data-driven prognostics and health management has emerged as a key enabler for Industry 4.0, yet the development of robust remainin…
站内正文

主题导航

研究 — AI 话题新闻 | AI News Hub