AI News HubLIVE

研究動態

待翻譯:My AI App Measures Success When People Stop Chatting with It

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:This post was not written with or by AI. I wanted to explore how AI could help deepen my faith. I enjoyed using Claude to research topics which were on my mind. It does a good job finding and quoting scripture but a poo…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • This post was not written with or by AI. I wanted to explore how AI could help deepen my faith. I enjoyed using Claude to research topics which were on my mind. It does a good job…
站內正文

待翻譯:Show HN: Web Game "Player vs. Computer"

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Notifications You must be signed in to change notification settings Fork 0 Star 2 BranchesTags Open more actions menu Latest commit History 9 Commits 9 Commits Folders and files NameName Last commit message Last commit…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Notifications You must be signed in to change notification settings Fork 0 Star 2 BranchesTags Open more actions menu Latest commit History 9 Commits 9 Commits Folders and files N…
站內正文

待翻譯:Glean unveils Tau desktop workspace, claims token-cost edge over Claude

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Glean Technologies Inc. today unveiled Glean Tau, a desktop workspace that connects the company’s enterprise artificial intelligence to a user’s local files, applications and code. The launch anchors a broad slate of product news at Glean:GO, the company’s conference this week in San Francisco. Packaged with it were benchmark numbers aimed at Anthropic PBC. Glean said […] The post Glean unveils Tau desktop workspace, claims token-cost edge over Claude appeared first on SiliconANGLE.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Glean Technologies Inc. today unveiled Glean Tau, a desktop workspace that connects the company’s enterprise artificial intelligence to a user’s local files, applications and code…
站內正文

待翻譯:Building a better PowerPoint API for the AI era

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:tldr AI agents create and edit PowerPoint by writing code against libraries with serious limitations. Even basic edits end up slow, expensive, and prone to file corruption. We built a PowerPoint API that addresses these…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • tldr AI agents create and edit PowerPoint by writing code against libraries with serious limitations. Even basic edits end up slow, expensive, and prone to file corruption. We bui…
站內正文

待翻譯:10 Rules for Getting Better Results from AI Coding Agents

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Everyone's using AI coding agents. Here's how to make yours actually useful.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Everyone's using AI coding agents. Here's how to make yours actually useful.
站內正文

待翻譯:I built my own secure network in the cloud to access my home PCs remotely - for free

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Do you need easy, secure remote access to your home or small business network? Before you pay big bucks for a VPN or remote-access plan, try Tailscale. Did I mention it's free?

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Do you need easy, secure remote access to your home or small business network? Before you pay big bucks for a VPN or remote-access plan, try Tailscale. Did I mention it's free?
站內正文

待翻譯:Fake US thinktank set up and funded by Israel sought to game AI for propaganda

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:In effort to prime chatbots to make pro-Israel arguments the site published 124 reports, over 560,000 words in nine days, Guardian analysis shows A pro-Israel messaging website badged with the name of a thinktank that does not exist has published more than half a million words in nine days, built on a commercial platform that promises to optimize content so that AI chatbots will cite it. The site gives Israel’s position on subjects including the torture of Palestinian prisoners, Israeli war crimes and whether Israel has deliberately starved Palestinians in Gaza, all presented as neutral research. Continue reading...

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • In effort to prime chatbots to make pro-Israel arguments the site published 124 reports, over 560,000 words in nine days, Guardian analysis shows A pro-Israel messaging website ba…
站內正文

待翻譯:Code Documentation Quality, Measured

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Quality scores 100 = best · Line = minimum Three documentation quality scores out of 100. Each bar includes its minimum passing score. Codebase coverage Important systems and workflows are documented. Minimum passing sc…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Quality scores 100 = best · Line = minimum Three documentation quality scores out of 100. Each bar includes its minimum passing score. Codebase coverage Important systems and work…
站內正文

待翻譯:AI helps design new materials that work in the real world

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The “CrysVCD” tool developed at MIT could cut the huge amounts of time and money spent on screening out chemically unstable designs.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • The “CrysVCD” tool developed at MIT could cut the huge amounts of time and money spent on screening out chemically unstable designs.
站內正文

待翻譯:My Anti-AI Manifesto

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:My Anti-AI manifesto As a software developer, the recent escalation of Artificial Intelligence (AI) has aroused strong emotions and several concerns in me, either of technical and philosophical nature. The human-machine…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • My Anti-AI manifesto As a software developer, the recent escalation of Artificial Intelligence (AI) has aroused strong emotions and several concerns in me, either of technical and…
站內正文

待翻譯:'This is crazy. This is insane': Bill Gates has changed his mind about AI,jobs

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:‘This is crazy. This is insane’: Bill Gates has changed his mind about AI and jobs Aug 26, 2026, 3:00am EDT Technology PostEmailWhatsapp The News Bill Gates says it’s time to hit the AI panic button. The technology has…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • ‘This is crazy. This is insane’: Bill Gates has changed his mind about AI and jobs Aug 26, 2026, 3:00am EDT Technology PostEmailWhatsapp The News Bill Gates says it’s time to hit…
站內正文

待翻譯:World humanoid robot games show runners breaking records, bursting into flames

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Limits of robot autonomy The fact that many events still permitted humans to directly control robotic motions shows that autonomous robot systems still have a long way to go, Patel said. Whereas humans can quickly learn…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Limits of robot autonomy The fact that many events still permitted humans to directly control robotic motions shows that autonomous robot systems still have a long way to go, Pate…
站內正文

待翻譯:Nvidia Jetson Orin-guided Russian AI drone killed three civilians in Ukraine

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:5 Join the conversation Follow us Add us as a preferred source on Google A Russian Molniya drone carrying an Nvidia Jetson Orin module crashed and killed three civilians at a gas station in Zaporizhzhia last month after…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • 5 Join the conversation Follow us Add us as a preferred source on Google A Russian Molniya drone carrying an Nvidia Jetson Orin module crashed and killed three civilians at a gas…
站內正文

待翻譯:Show HN: LLM-powered webapp to build LLM-powered webapps

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Hey all, The goal is to earn on token margins for LLM calls when you build an AI-powered webapp. I proxy OpenAI and Anthropic calls so that when you deploy a site to a subdomain, your users token usage will be tracked.…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Hey all, The goal is to earn on token margins for LLM calls when you build an AI-powered webapp. I proxy OpenAI and Anthropic calls so that when you deploy a site to a subdomain,…
站內正文

待翻譯:Around the world, people are rejecting divisive and dangerous politics. We can – and must – build on that | Gordon Brown

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The vast majority of the global public wants international cooperation on human rights, climate and AI. Like-minded countries must stand together to deliver Britain’s new prime minister, Andy Burnham, is already having to make one of his gravest decisions. He has to issue the instructions that he alone gives to the military, setting out the UK response in a doomsday scenario of a nuclear weapons attack on us. He will, as I did two decades ago, sign a piece of paper telling commanders whether or not to retaliate and, if so, whether through targeting civilian conurbations or military sites. These instructions are written down in the aptly named “letter of last resort”. Now, more than at any time since the 1960s Cuban missile crisis, the European public fears a third world war. With the nuclear Doomsday Clock developed by atomic scientists moving ever closer to midnight, and Japan, South Korea, Saudi Arabia, the UAE, Egypt, Poland and Germany contemplating either acquiring nuclear weapons or siting them on their soil, our world is descending from a rules-based order to a power-based one, where might is deemed right and brute force dominates. Gordon Brown is the UN’s special envoy for global education and was UK prime minister from 2007 to 2010 The future starts with us: Gordon Brown in conversation On Thursday 10 September, join Hugh Muir and Gordon Brown to discuss the intricate connections between global instability and civic decline, as explored in Brown’s new book, The Future Starts With Us. Book tickets here Continue reading...

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • The vast majority of the global public wants international cooperation on human rights, climate and AI. Like-minded countries must stand together to deliver Britain’s new prime mi…
站內正文

待翻譯:Bridging Teacher Expectations and Robot Learning via Coupling Dynamics

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23994v1 Announce Type: new Abstract: Human-robot teaching focuses on enabling nontechnical experts to customize robots according to their needs after deployment. With recent advances in machine learning, human-robot teaching is no longer confined to offline learning where the data gathering step from a human teacher is separated from when the robot learns. Instead, more recent approaches for human-robot teaching focus on coupling human teaching with robot learning. This coupling impacts the structure, timing, and content of the teaching and learning interaction. However, it is currently unclear how such coupling dynamics affect humanrobot teaching effectiveness and human perceptions towards the teaching process. Informed by human learning theories, in this paper we propose a new scale for classifying human-robot teaching interactions according to coupling dynamics present between the human teacher and robot learner. We apply this scale to a subset of the human-robot teaching literature to identify how coupling dynamics and human teacher mental model mismatches with the ground truth robot learning system affect teaching effectiveness and human perceptions towards the teaching process

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23994v1 Announce Type: new Abstract: Human-robot teaching focuses on enabling nontechnical experts to customize robots according to their needs after deployment. With r…
站內正文

待翻譯:Sensorless damage-safe grasping

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23983v1 Announce Type: new Abstract: Robotic fruit harvesting must hold produce securely without bruising it, yet compression stiffness varies several-fold with ripeness within a single species, so no fixed grip force spans the range. Rather than tune force, we bound deformation: a controller closes the gripper until the object's estimated compression strain reaches a user-specified limit $\varepsilon$, using only the encoder position and motor-effort signal on every servo gripper---no tactile or force-torque sensor. Dividing an effort-based contact force by a lower bound on object stiffness makes the stop provably conservative---true compression stays at or below $\varepsilon$---for any $\varepsilon$ above a contact-detection strain floor we identify and quantify: robust detection itself spends compression, linearly in closing speed, making speed an explicit throughput--gentleness knob. Unlike a hand-tuned force threshold, $\varepsilon$ is a certified, size-scaling, operator-interpretable damage limit, and a ready safe-action parameter for learned grasping policies. In MuJoCo simulation over a realistic fruit-stiffness range, under a sensor-noise model calibrated to the real servo, the controller holds $\ge 98\,\%$ grasp at $0\,\%$ damage across all medium-to-firm stiffnesses for the entire certified $\varepsilon$ range, which neither fixed-force baseline attains; on stiffness-graded 3D-printed TPU cubes it matches baseline grasp success at roughly half the grip force and cuts soft-object damage from $100\,\%$ to $40\,\%$.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23983v1 Announce Type: new Abstract: Robotic fruit harvesting must hold produce securely without bruising it, yet compression stiffness varies several-fold with ripenes…
站內正文

待翻譯:Safety-aware Model Predictive Path Integral Control with Signal Temporal Logic

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23972v1 Announce Type: new Abstract: Safety-aware motion planning remains a challenge in robotics, especially when missions are time-critical and are under complex specifications. In this paper, we propose safety-aware-stl-mppi, a computationally efficient sampling-based receding-horizon planning framework designed to promote satisfaction of constraints expressed in Signal Temporal Logic (STL). Our approach encodes discrete-time STL formulas into candidate time-varying control barrier functions (CBF), which are integrated into a model predictive path integral (MPPI) controller. Our method inherits the benefits of low computational cost from an efficiently parallelizable sampling based planner and utilizes CBF for constraints expressed in STL. We compare against several MPPI baselines using four artificial Mars Rover planning case studies with a diverse environment and cost setups, where we show our method consistently achieving high safety and efficiency. We show a quadcopter planning experiment with NVIDIA Isaac Lab.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23972v1 Announce Type: new Abstract: Safety-aware motion planning remains a challenge in robotics, especially when missions are time-critical and are under complex spec…
站內正文

待翻譯:Interpreting Control Latents for System Identification via Conditional Flow Matching

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23887v1 Announce Type: new Abstract: Latent-conditioned adaptive policies can control robots across changing dynamics, but their learned latents remain internal representations of the policy rather than physical models that can be inspected, rolled out, or used by other control modules. This limits closed-loop analysis, diagnosis, and further improvement of a fixed policy. A direct mapping from latent to physical parameters is also under-specified, because multiple systems can induce similar closed-loop behavior. We therefore decode each operational latent into a distribution of quadrotor models using conditional flow matching. The decoded distribution enables two downstream uses without modifying the policy: online predictive tuning of a high-level controller around the fixed low-level policy, and robustness analysis under specified disturbances. Under perturbed actuator dynamics, decoded-model predictive tuning reduces position tracking RMSE by $23\%$ and heading RMSE by $45\%$ relative to fixed gains. Under Gaussian force disturbances, decoded-model ensembles closely predict the lateral tracking-error evolution. Together, these results show that control latents can be converted into physical model ensembles for tuning, robustness analysis, and diagnosis of frozen adaptive policies.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23887v1 Announce Type: new Abstract: Latent-conditioned adaptive policies can control robots across changing dynamics, but their learned latents remain internal represe…
站內正文

待翻譯:DreamLedger: Execution-Settled Credit Files for World-Model Imagination in Robot Decision Loops

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23863v1 Announce Type: new Abstract: Robots are beginning to act on world-model predictions, yet reliability is still expressed through instantaneous, model-internal signals. DreamLedger instead treats reliability as a persistent deployment object: an execution-settled credit file recording how often consumed predictions are borne out, indexed by operating condition, region, and prediction horizon, and consulted before each use. Each consumed prediction is registered as a claim; attributable outcomes are settled against arriving reality at zero labeling cost, an attribution stage excludes measurement-contaminated outcomes, and a settlement-supervised head complements sparse bins. The resulting credit gates consumption: low-credit predictions shorten the dependent horizon or trigger additional observation; every reliance event remains auditable via dependency tickets and replayable logs. We evaluate DreamLedger in three simulated domains (indoor flight, tabletop manipulation, 2D navigation), via mounts on unmodified DreamerV3, TD-MPC2, and V-JEPA 2-AC, and on a real Franka manipulator. Claim failure is dose-monotone in all 12 held-out condition-horizon cells. Credit-gated planning reduces burned imagination (consumed claims that later fail to redeem) by 62% (95% CI 43-81%) versus blind consumption, with equal success and comparable collision rates. At matched risk targets, persistent books cut verification probes from 1.00 to 0.36/episode in manipulation, at success 0.94 versus 0.98; settlement-grounded calibration retains moderate, seed-consistent operating points unlike raw instantaneous gates. The same trust layer operates across decoder-, latent-, and token-space interfaces, including V-JEPA 2-AC settled on real robot frames. On hardware, settlement remains operational under real sensing and contact noise, a deployment failure loop is re-priced online, and all 1,062 registered spends replay from the audit logs.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23863v1 Announce Type: new Abstract: Robots are beginning to act on world-model predictions, yet reliability is still expressed through instantaneous, model-internal si…
站內正文

待翻譯:Resilience Matters for Embodied Agents System: New Metrics, Systematic Evaluation, and Optimization

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23839v1 Announce Type: new Abstract: Embodied Agents System (EAS) are increasingly deployed in open-world physical domains, where reliability directly dictates deployment quality and human-agent trust. However, existing evaluations rely on outcome-centric metrics as success rate or safety scores that collapse diverse execution trajectories into coarse scores, obscuring the dynamic processes underlying agent behavior. Therefore, they ignore a critical property of EAS -- which we define as the Resilience -- that reflects how EASs recover, stabilize, and extend under perturbations and across iterative updates. The lack of resilience is particularly critical in open-world environments due to continuous unexpected disruptions, thus directly affecting the quality of EAS deployment. To address this problem, we gain insight from the resilience-engineering concepts to EAS groundings and propose a novel resilience evaluation framework that can be flexibly applied to any EAS. Specifically, we define the first comprehensive resilience metrics suite for EASs system that exposes Rebound, Stability, and Graceful Extensibility across embodied tasks execution, providing a practical grounding for EAS resilience analysis. We further implement the resilience evaluation layer that transforms execution process into assessments for diagnosis and optimization. Across 400 household tasks with 10 EAS, we reveal the process-level distinction hidden by outcome metrics, including recovery cost differences among successful episodes ($\Delta C_{rec}=25.2$), increased instability and task-family degradation. Metrics-guided optimizations reduce recovery cost and increase stability, graceful extensibility completion, showing the diagnostic effect of resilience evaluation. Our results reveal a trade-off among resilience characteristics, suggesting that a resilient EAS construction should be configured according to deployment-specific requirements.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23839v1 Announce Type: new Abstract: Embodied Agents System (EAS) are increasingly deployed in open-world physical domains, where reliability directly dictates deployme…
站內正文

待翻譯:Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23831v1 Announce Type: new Abstract: While reinforcement learning (RL) allows generalist robot policies to continually improve during deployment, the large model size of modern generalist policies, such as VLAs, poses a fundamental obstacle to effective RL improvement. In particular, their severe inference latency---which can lead to pauses or jerky movements---can alter the effective environment dynamics and, if not correctly accounted for, break the Markov assumption that RL relies on, causing standard RL algorithms to fail completely. In this work, we introduce a latency-aware framework, Asynchronous RL with Intermediate Information (ARLI), that enables RL-based improvement of generalist policies under inference delays. Our framework builds on asynchronous inference approaches, which interleave action generation with execution to hide latency, and addresses its incompatibility with RL by providing a low-latency RL policy design that maximizes reactivity within the inference window through two contributions: state augmentations that restore near-Markovian structure by incorporating committed actions and a mid-inference observation. We evaluate our approach across simulated and real-world manipulation tasks, and find that it enables effective finetuning under inference delays where standard RL fails entirely, even matching or exceeding the performance of standard RL in idealized no-latency settings.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23831v1 Announce Type: new Abstract: While reinforcement learning (RL) allows generalist robot policies to continually improve during deployment, the large model size o…
站內正文

待翻譯:Concept-Guided Exploration: Building Persistent, Actionable Scene Graphs

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23650v1 Announce Type: new Abstract: The perception of 3D space by mobile robots is rapidly moving from flat metric grid representations to hybrid metric-semantic graphs built from human-interpretable concepts. While most approaches first build metric maps and then add semantic layers, we explore an alternative, concept-first architecture in which spatial understanding emerges from asynchronous concept agents that directly instantiate and manage semantic entities. Our robot employs two spatial concepts (room and door), implemented as autonomous processes within a cognitive distributed architecture. These concept agents cooperatively build a shared scene graph representation of indoor layouts through active exploration and incremental validation. The key architectural principle is hierarchical constraint propagation: Room instantiation provides geometric and semantic priors to guide and support door detection within wall boundaries. The resulting structure is maintained by a complementary functional principle based on prediction-matching loops. This approach is designed to yield an actionable, human-interpretable spatial representation without relying on any pre-existing global metric map, supporting scalable operation and persistent, task-relevant understanding in structured indoor environments.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23650v1 Announce Type: new Abstract: The perception of 3D space by mobile robots is rapidly moving from flat metric grid representations to hybrid metric-semantic graph…
站內正文

待翻譯:Macro-Operator Generation and Predicate Selection for TAMP Operator Learning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23629v1 Announce Type: new Abstract: Creating symbolic operators by hand is one of the main bottlenecks in deploying Task and Motion Planning systems (TAMP). Recent works show that these operators can instead be learned directly from demonstration data. Existing methods, however, typically learn each action in isolation and cannot capture the recurring multi-step structure of manipulation tasks, so the search becomes intractable on long sequential tasks. A further inefficiency arises in the symbolic state: every provided predicate is evaluated at every search node, even when it never appears in any learned operator. We present a system that addresses both problems together. Its central component is the automatic generation of macro-operators, composite actions that compress a recurring sequence of individual actions into a single planning step. Our system discovers causally linked action pairs directly from the training data, where one action produces exactly the condition that the next one requires, and turns each pair into a new operator. Alongside this, our system prunes every predicate that no learned operator references, which shrinks the symbolic state evaluated at each search node. Together, these changes shorten the effective planning horizon, and the benefit they bring grows with the length of the task. Across four TAMP domains, our method reaches up to a 4.6x planning speedup compared to the baseline method, namely Learning Operators for TAMP. More importantly, it solves a long sequential task that the baseline cannot solve. Macro-operator discovery thus not only accelerates planning but, in certain domains, determines solvability in practice.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23629v1 Announce Type: new Abstract: Creating symbolic operators by hand is one of the main bottlenecks in deploying Task and Motion Planning systems (TAMP). Recent wor…
站內正文

待翻譯:Pattern-Derived Visual Swarm Games: Multi-Scale Drone-Vision States for Interception and Sustainability Audits

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23575v1 Announce Type: new Abstract: We convert drone-vision annotation streams into virtual swarm-game states without controlling physical drones. VisDrone and UAVSwarm metadata are compressed into a Bloom representation; deterministic probes produce bounded capability vectors, image-space formations, finite zero-sum payoffs, and human-readable visual overlays. The audit scales from $6\times 6$ to $32\times 32$ finite games and adds a repeated Markov layer with stock, fatigue, adaptation, exposure, stress, budget, data-growth, model-improvement, and entropy-budget state variables. Local screen tuning raises robust screen security from $0.526$ to $0.593$, and the $32\times 32$ tuned screen reaches value $0.616$. A field readout audit shows that fixed-pixel rasters do not improve monotonically: $128\times 128$ accuracy is $67.2\%$ and hotspot error is $0.136$. The diagnosed error is shrinking image-plane bandwidth. A finite empirical-risk encoder over scale-normalized Gaussian bandwidths selects a scale-normalized encoder with $\lambda=1.50$, reaching $77.6\%$ accuracy at $128\times 128$ and reducing joint loss by $0.185$. A server-side audit checks $16{,}777{,}216$ target-localization states, and a 32-round repeated-game audit over $16{,}777{,}216$ trajectories selects a budget-adaptive policy with value $0.461$.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23575v1 Announce Type: new Abstract: We convert drone-vision annotation streams into virtual swarm-game states without controlling physical drones. VisDrone and UAVSwar…
站內正文

待翻譯:Too much of a good thing -- when knowledge distillation promotes overfitting, and how to avoid it

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23752v1 Announce Type: new Abstract: The growing size of Convolutional Neural Networks has led to increasingly large and costly models. Knowledge Distillation (KD) addresses this by transferring knowledge from a large network (teacher) to a small one (student), also reducing the training data required. KD is traditionally applied only at the network's final output. However, its behaviour when applied at intermediate network layers has received little attention. This raises the question of whether intermediate block-wise KD, which provides supervision throughout the network, could offer an advantage under specific conditions, such as few instances per class, which is common in fine-grained datasets. This work proposes a student design based on simple, homogeneous blocks mirroring those of the teacher, distilling knowledge between corresponding blocks. Across eleven datasets, we show that on classic datasets, distilling only the last block is sufficient -- and often best--, whereas fine-grained, data-scarce settings benefit substantially from intermediate supervision, with even a single additional distillation point narrowing the gap considerably. We further study how this supervision should be guided, exploring configurations of varying granularity and informed by an explainability analysis based on attention maps, Centered Kernel Alignment, and Grad-CAM, alongside the impact of teacher and student fine-tuning strategies. This work shows that intermediate block-wise distillation, guided appropriately, is key to building compact data-efficient models without sacrificing accuracy.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23752v1 Announce Type: new Abstract: The growing size of Convolutional Neural Networks has led to increasingly large and costly models. Knowledge Distillation (KD) addr…
站內正文

待翻譯:CRISP: Calibration-Aware Visual State Space Duality for Remote Sensing Semantic Segmentation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23746v1 Announce Type: new Abstract: State space models, especially Visual State Space Duality (VSSD), have emerged as efficient linear-time alternatives to Transformers for dense visual tasks. However, we observe that VSSD compresses spatial context into a global aggregation that suppresses high-frequency responses, causing excessive boundary smoothing in remote sensing semantic segmentation. To address this, we propose CRISP, a calibration framework with two components. Its core, the Duality Calibration Operator (DCO), restores local contrast and boundary responses through residual injection and frequency calibration within the VSSD backbone, without altering its linear complexity. To retain the recovered detail, an Orthogonal Multi-Prototype (OMP) head assigns multiple orthogonally constrained prototypes per class to model large intra-class variance. Extensive experiments on Potsdam, Vaihingen, and LoveDA show that, with approximately 30M parameters, CRISP achieves consistent gains in mean F1 (mF) and mIoU while remaining competitive with state-of-the-art methods. Code is available at https://github.com/crazylifeha/CRISP.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23746v1 Announce Type: new Abstract: State space models, especially Visual State Space Duality (VSSD), have emerged as efficient linear-time alternatives to Transformer…
站內正文

待翻譯:More Motion Is Not Always Better Motion: Corpus Composition Governs Whether Augmentation Helps SMPL-Based Parkinsonian Gait Severity Estimation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23730v1 Announce Type: new Abstract: We grade MDS-UPDRS gait severity from SMPL motion using three frozen MotionAGFormer encoders as featurizers, reaching macro-F1 0.58 on a hidden, multi-site test set. Because the system's members differ only in their lifting corpus, evaluating encoders singly on that test set isolates what that corpus contributes. Six pools drawn from one inertial dataset, varying only in which walking tasks they include, score between 0.32 and 0.53, and just one of them beats the 0.51 of an encoder given no outside motion at all. What separates them is not how much data they hold but whether they carry a contrast in walking speed, the variation this representation appears to depend : a further pool adding a third collection site at fixed task composition does worse still. The same rule explains why exact synthetic motion and monocularly reconstructed web video both fail to help. Modifying the learned representation itself, rather than the corpus behind it, cost every variant that attempted it.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23730v1 Announce Type: new Abstract: We grade MDS-UPDRS gait severity from SMPL motion using three frozen MotionAGFormer encoders as featurizers, reaching macro-F1 0.58…
站內正文

待翻譯:Velocity-coupled Representation Refinement for Satellite Orbit Prediction

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23728v1 Announce Type: new Abstract: Satellite orbit prediction, which aims to forecast future orbital trajectories from historical observations, is important for collision warning and safe space operations. With advances in time-series forecasting, learning-based methods have emerged as a promising solution for satellite prediction. In orbital dynamics, a satellite state is typically described by position and velocity, where position characterizes trajectory geometry and velocity reflects its instantaneous direction and rate of change. However, most existing methods mainly focus on temporal dependencies within position sequences while rarely exploiting the intrinsic coupling between position and velocity, which is essential for modeling satellite motion. To this end, we propose OrbitNet, a velocity-aware representation learning method for accurate satellite orbit prediction. It lifts conventional position-sequence forecasting to a position-velocity coupled representation learning paradigm by exploiting relationships among satellite state variables. Specifically, we develop a velocity-coupled representation refinement strategy to enhance positional representations through cross-variable interactions between position and velocity. We further introduce orbital segment modeling, which partitions historical trajectories into temporal segments and performs segment-level temporal learning to capture local motion variations and long-range evolution patterns. Extensive experiments show that OrbitNet outperforms large time-series foundation models and representative general forecasting methods under both in-domain evaluation on Starlink and zero-shot evaluation across six unseen satellite constellations. We expect this work to encourage further exploration of satellite-aware representation learning for trajectory time-series forecasting.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23728v1 Announce Type: new Abstract: Satellite orbit prediction, which aims to forecast future orbital trajectories from historical observations, is important for colli…
站內正文

待翻譯:DriftAD: Visually-Guided Text Drift for Few-Shot Industrial Anomaly Detection

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23723v1 Announce Type: new Abstract: Few-shot anomaly detection (FSAD) has recently benefited from vision-language models such as CLIP, which enable anomaly de?tection by aligning visual features with text descriptions of normal and abnormal states. However, existing methods typically rely on static text prompts that are applied uniformly across the entire feature hierarchy and spatial dimensions. This rigid global-to-local matching fails to capture the highly localized and scale-dependent physical variations of industrial defects. To address this, we propose DriftAD, a FSAD framework built on three key modules. First, an Anomaly Signal Amplification (ASA) module enhances subtle defect signals through spatial and frequency branches before text-visual matching. Second, Visually-Guided Text Drift (VGTD) dynamically transforms frozen CLIP text embeddings, steering them into layer?wise, spatially-adaptive anomaly descriptors conditioned on local visual context at each encoder depth. Third, Drift-Guided Spatial Gating (DGSG) uses the drifted abnormal descriptor as a spatial probe to selectively enhance anomaly-relevant visual features. Addi?tionally, a drift separation loss prevents representational collapse of the drifted descriptors, and a gate supervision loss enforces spatially discriminative gating in DGSG. Extensive experiments on MVTec?AD and VisA demonstrate state-of-the-art performance across all 1-, 2-, and 4-shot settings on both image-level and pixel-level metrics. Code is available at https://github.com/wenyang001/DriftAD.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23723v1 Announce Type: new Abstract: Few-shot anomaly detection (FSAD) has recently benefited from vision-language models such as CLIP, which enable anomaly de?tection…
站內正文

待翻譯:Platonic Representation Hypothesis on World Models

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23720v1 Announce Type: new Abstract: World models have demonstrated significant potential for perceiving and simulating complex environments. Despite their strong performance, the fundamental nature of their learned representations remains poorly understood. In this paper, we investigate the Platonic Representation Hypothesis within this domain by proposing the Predictive Consistency Assumption: we posit that the optimization of a shared state transition objective acts as a selective pressure that encourages heterogeneous models to converge toward a shared latent structure. Through systematic experiments with the DINO World Model (DINO-WM), in which we vary visual encoders to create heterogeneous models, we find that capable world models evolve toward geometrically similar internal structures. Moreover, via model stitching, we show that the internal features of one world model can be mapped to another with limited performance degradation, providing evidence of functional compatibility. Our findings suggest that the pursuit of predictive consistency can promote shared, transition-compatible latent structure across world models.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23720v1 Announce Type: new Abstract: World models have demonstrated significant potential for perceiving and simulating complex environments. Despite their strong perfo…
站內正文

待翻譯:Scaling Reinforcement Learning for Diffusion Models via Velocity Matching

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23664v1 Announce Type: new Abstract: Reward fine-tuning is becoming an important tool for adapting diffusion models to human preferences and task-specific objectives, but existing methods largely inherit policy-gradient machinery from large language models. Unlike autoregressive models, diffusion models do not provide tractable likelihoods for generated samples. As a result, current approaches either construct trajectory likelihoods from stochastic denoising transitions or approximate endpoint likelihoods with evidence lower bound, introducing additional computation and algorithmic complexity. We demonstrate that this likelihood-based machinery is not necessary for effective diffusion reward fine-tuning. We propose reward-based velocity matching (RVM), a simple trajectory-free update that acts directly on the velocity field. RVM reinforces directions associated with high-reward generations, suppresses those with low reward, and involves an optional anchor term controlling drift from a reference velocity. Notably, it provides a general framework that recovers recent fine-tuning methods, including RAM and DiffusionNFT, as special cases. Across various large-scale diffusion models reward fine-tuning tasks, RVM is competitive with or outperforms trajectory-based policy-gradient methods under substantially reduced training cost. We further find that, once the velocity update is simplified, the particular loss variant matters less than reward and anchor design. For video generation, standard preference rewards can favor visually clean but nearly static outputs; introducing a new dynamic-tracking reward that substantially improve motions while improving overall VBench performance. These results suggest that scalable reward fine-tuning for diffusion models is better posed in the native velocity representation than as likelihood-based policy optimization.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23664v1 Announce Type: new Abstract: Reward fine-tuning is becoming an important tool for adapting diffusion models to human preferences and task-specific objectives, b…
站內正文

待翻譯:Cross-Generation Optimization of YOLOv26, YOLOv11, and YOLOv8 for Fine-Grained Small-Object Detection and Instance Segmentation in Complex Orchards

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23636v1 Announce Type: new Abstract: Small-object detection and instance segmentation remain challenging in orchard environments because of green-on-green similarity, occlusion, and limited pixel representation of fine fruit anatomy. This study presents a cross-generation benchmark of Ultralytics YOLOv8, YOLOv11, and YOLOv26 for detecting and segmenting apple fruitlet, calyx, and peduncle structures for robotic orchard perception. Five model scales (n, s, m, l, and x) were evaluated under conventional 640 x 640 and small-object focused 960 x 960 training configurations, yielding 30 experiments. Increasing model capacity did not consistently improve accuracy. YOLOv11s-960 achieved the highest observed mask mAP@50:95 (0.402) and box mAP@50:95 (0.426), while YOLOv26s-960 achieved comparable values of 0.397 and 0.425 with only 10.37 M parameters and 34.1 GFLOPs. Peduncle remained the most challenging class. Overall, compact-to-moderate YOLO models with small-object-focused training provided favorable accuracy efficiency trade-offs, establishing a practical benchmark for fine-grained agricultural robotics and orchard perception. Github Link: https://github.com/rnjnspkt/Optimizing-and-Comparing-Ultralytics-YOLOv26-YOLOv11-and-YOLOv8-for-Small-Object-Detection-and-Seg

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23636v1 Announce Type: new Abstract: Small-object detection and instance segmentation remain challenging in orchard environments because of green-on-green similarity, o…
站內正文

待翻譯:The Blending Ratio Is Not Where the Performance Is: Diagnosing Prototype Blending for Few-Shot Adaptation of Vision-Language Models

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23634v1 Announce Type: new Abstract: Many few-shot adaptation methods for vision-language models classify with a convex combination of the zero-shot text prototype and the mean of the K labelled image features, with a single blending ratio routinely tuned on held-out labels, often on the test set itself. We ask what the family's own bias-variance justification invites: what is the right ratio, can it be estimated without validation data, and is finding it where the performance is? First, the ratio minimising prototype mean-squared error has a closed form whose support-set plug-in is exactly a positive-part James-Stein coefficient shrinking towards the text prototype. Across 4,800 cells (ten datasets, five backbones including SigLIP, five shot counts, five seeds, four prompt tiers) this theoretically optimal ratio is a reliable estimate of the wrong quantity: on the 950 primary-tier cells where it is defined it trails a test-set-oracle ratio by 8.5 points. It saturates near 1, discarding the text prior for a nearest-class-mean classifier, because 78% of the text-image prototype distance it treats as bias is a class-independent offset that the arg max largely cancels. We prove the mechanism and bound its share of the damage at 26% by a counterfactual. Second, leave-one-out on the support set alone sets a ratio landing within 0.9 points of the oracle blend, so it is estimable without validation data. Third, validation-free linear probes beat even the oracle-tuned blend: CLAP by +1.9 points and LP++ by +1.5 on average, and at K >= 4 all four validation-free baselines sit above the oracle, the linear probes by margins excluding zero. These results locate the ceiling in the model class, not the hyperparameter: the ratio can be set near-optimally for free, and it is still not where the performance is. Code, cached features, per-cell records: https://huggingface.co/datasets/Liangzhi-Li/clipbench-blending

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23634v1 Announce Type: new Abstract: Many few-shot adaptation methods for vision-language models classify with a convex combination of the zero-shot text prototype and…
站內正文

待翻譯:Fidelity Preference, Not Demographic Preference: A Pixel-Level Attribute-Sensitivity Audit of Image Aesthetic/Preference Scorers

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23593v1 Announce Type: new Abstract: Text-to-image systems use learned aesthetic scorers to filter training data and guide generation, but whether these scores encode demographic attributes as objective quality is unclear. We audit four scorers (LAION-Aesthetics, PickScore, ImageReward, HPSv2) using pixel-level interventions on skin tone and body type in synthetic and real images. Our key finding is that along skin-lightness, the dominant effect is fidelity preference: unaltered images score highest, and perturbations in either direction are penalized (inverted-U). Placebo arms show this penalty is not an artifact of the skin operator, as applying the same CIELAB L* shift to non-skin regions yields similar penalty magnitudes. However, the penalty is operator-dependent and holds for all operators only for LAION-Aes. Critically, audits on synthetic images alone are misleading: LAION-Aes shows strong preference for darker skin on synthetic faces, but on 1470 real faces the preference reverses and becomes much smaller, and amplification becomes non-significant. Across scorers, synthetic results do not transfer -- reversing for LAION-Aes and HPSv2, attenuating for PickScore. We contribute a reproducible benchmark with artifact control and synthetic/real cross-validation, and an auditability criterion for pixel-level causal isolation (valid for skin tone, not for body type due to deformation). Population-stratified analysis shows fidelity-penalty asymmetry is not robust across groups after FDR correction except for HPSv2. Our findings show naive synthetic audits misjudge bias direction and magnitude, and only within-image causal isolation on real data can distinguish true demographic bias from fidelity preference.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23593v1 Announce Type: new Abstract: Text-to-image systems use learned aesthetic scorers to filter training data and guide generation, but whether these scores encode d…
站內正文

待翻譯:Mixture of Channel Experts: Static Sparse Supports with Input-Adaptive Mixing for Pointwise Projections

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23794v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) scales language models by routing each input through a small set of independently parameterized experts. We show that copying this design into convolutional networks fails for a structural reason: parallel convolutional experts that read the same input channels learn nearly identical filters. We therefore move the expert axis from operator duplication to channel selection. We introduce Mixture of Channel Experts (MoCE), a structured sparse channel-mixing layer, inspired by MoE, that replaces pointwise (1x1) channel-reduction projections. In MoCE, an expert is a single output channel with a learned sparse support of k << C input channels. The selected channels are combined by a softmax whose temperature is predicted per input, so each expert can move between mean-like and max-like aggregation. A residual expert summarizes the unselected channels, and a load-balancing loss keeps channel coverage complete. MoCE replaces a dense projection whose cost is quadratic in C with a mechanism whose relative cost scales as k/C, and the predicted savings hold in measured wall-clock time. Across ResNet backbones on ImageNet-1K and CIFAR-100, transfer learning, EfficientViT, and a strong modern training recipe, MoCE matches or exceeds dense baselines and prior channel-selection methods while reducing MACs by 16.7% and end-to-end latency.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23794v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) scales language models by routing each input through a small set of independently parameterized experts. W…
站內正文

待翻譯:GAP-Prompt: Gated Adaptive Prompting for Efficient Continual Learning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23782v1 Announce Type: new Abstract: Continual learning faces the persistent challenge of catastrophic forgetting, where sequential task updates degrade previously acquired knowledge. While prompt-based methods integrated with pre-trained models offer a compelling solution by freezing the backbone, they often rely on static, task-level prompting strategies that overlook fine-grained intra-task diversity. In this paper, we propose Gated Adaptive Prompting (GAP-Prompt), a novel method that introduces instance-level adaptability to the prompting process. GAP-Prompt consists of three synergistic modules: (1) instance-conditioned gating, which dynamically determines optimal prompt injection layers for each individual image; (2) dynamic knowledge fusion, which performs instance-aware aggregation of current and historical prompts, enabling knowledge integration across tasks; and (3) shared prompt distillation, which anchors foundational knowledge in early shared layers to mitigate forgetting. Extensive evaluations on CIFAR-100, ImageNet-R, and CUB-200 benchmarks demonstrate that GAP-Prompt consistently achieves state-of-the-art performance. Notably, on the fine-grained CUB-200 dataset, GAP-Prompt reaches 87.29% accuracy, approaching the joint training upper bound (88.00%) and outperforming existing methods by a significant margin.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23782v1 Announce Type: new Abstract: Continual learning faces the persistent challenge of catastrophic forgetting, where sequential task updates degrade previously acqu…
站內正文

待翻譯:Disentangled Skill Representations for Predictive Human Modeling

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23776v1 Announce Type: new Abstract: Understanding human skill is important for AI systems that collaborate with, coach, or assist people. Unlike typical latent variable estimation problems which rely on single observations, skill is a persistent, compositional, and behaviorally grounded construct that must be inferred from patterns over time. We introduce Skill Abstraction with Interpretable Latents (SAIL), a method for modeling human skill as an interpretable, multi-dimensional construct inferred from naturalistic behavior. Our approach produces a skill embedding that is robust to transient performance fluctuations and learns a transferable representation of human subskills. Furthermore, SAIL supports skill-informed behavior prediction that generalizes across a variety of in-domain contexts. We represent each individual with a persistent skill embedding that controls a blend between expert and novice bases and is trained using counterfactual subskill swaps for disentanglement. This design encourages representations that are both robust to performance variation and structured for interpretability. We demonstrate across racing and baseball that SAIL achieves strong predictive performance and consistently improves behaviorally grounded disentanglement over the evaluated baselines, while also improving downstream AI coaching performance.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23776v1 Announce Type: new Abstract: Understanding human skill is important for AI systems that collaborate with, coach, or assist people. Unlike typical latent variabl…
站內正文

待翻譯:Tight Majorizations and Convergence Rates of Nuclear Norm Minimization IRLS

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23765v1 Announce Type: new Abstract: Iteratively reweighted least squares (IRLS) methods constitute a natural approach to nuclear norm minimization, but their convergence rates and the role of the weight operator have remained poorly understood. This paper establishes sharp convergence rates for IRLS methods for constrained nuclear norm minimization in low-rank recovery. A central ingredient is a new majorization analysis for the smoothed nuclear norm: we prove that the harmonic-mean weight operator defines a valid global quadratic majorizer. Furthermore, we show that this weight operator is optimal within the family of power-mean weights, clarifying why it improves over classical one-sided reweighting schemes that use only row- or column-space information. Under a Schatten-1 null space property, we prove global linear convergence of IRLS algorithms using a variety of weight operators, including the harmonic-mean weights. For IRLS with harmonic-mean weights, we prove a dimension-independent, locally linear convergence rate. We provide a counterexample showing that this dimension-independent local rate cannot in general be obtained for IRLS algorithms using one-sided weight operators, which predominate in the literature. Numerical experiments corroborate the theoretical results and illustrate the practical advantage of harmonic-mean reweighting across square, rectangular, and adversarially initialized recovery problems.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23765v1 Announce Type: new Abstract: Iteratively reweighted least squares (IRLS) methods constitute a natural approach to nuclear norm minimization, but their convergen…
站內正文

待翻譯:Calibration-Preserving Pruning: Compression as a Reliability Contract

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23744v1 Announce Type: new Abstract: Split conformal prediction, not the pruning rule, supplies finite-sample marginal coverage once a pruned model is fixed independently of the conformal calibration split. We study the separate efficiency problem: can pruning preserve score geometry well enough to obtain smaller valid prediction sets? Calibration-Preserving Pruning (CPP) augments a base pruning score with nonconformity-gradient saliency and uses disjoint pruning, validation-selection, conformal-calibration, and test splits. Bounded score perturbations imply bounded conformal-quantile shifts and controlled set inflation, but do not make the generic coverage theorem CPP-specific. Final five-seed Qwen2.5-1.5B results at 50\% sparsity show the largest gains on large-label tasks. On DBpedia-14, CPP-SparseGPT reduces mean set size from \(10.1\) to \(8.6\) while changing accuracy from \(0.347\) to \(0.366\); CPP-Wanda reduces \(11.2\) to \(9.0\) with an accuracy trade-off from \(0.310\) to \(0.295\). Across 15 dataset--sparsity cells, CPP-SparseGPT produces smaller sets in 13 and higher accuracy in 11. Matched controls show that generic supervised gradients explain much of the gain: true-label CPP is not statistically resolved from matched Wanda+SNIP, whereas threshold-aware candidate-label CPP reaches \(7.8\) mean set size at explicit accuracy and offline-compute costs. RoBERTa-base and Llama-3-8B diagnostics support transfer, but our claims remain limited to reliability-sensitive classification.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23744v1 Announce Type: new Abstract: Split conformal prediction, not the pruning rule, supplies finite-sample marginal coverage once a pruned model is fixed independent…
站內正文

待翻譯:Response Renormalization for Critical Deep Equilibrium Models

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23725v1 Announce Type: new Abstract: Deep Equilibrium Models (DEQs) compute predictions from a hidden representation unchanged by the model update. Training through this equilibrium uses implicit differentiation and requires solving an adjoint system built from the residual Jacobian. If this Jacobian is nearly singular along loss-sensitive directions, small perturbations can be strongly amplified in the adjoint response, producing large, highly sensitive gradients that can make optimization unreliable. We introduce Response Renormalization, a backward-pass framework that lifts selected near-pole denominators while leaving unlifted response channels unchanged. Collective Mode Response Renormalization (CMR) applies this correction in a low-dimensional critical subspace, while Phi-adaptive CMR computes a bounded response mass from a positive susceptibility rule. We derive dense and matrix-free collective formulations, distinguish exact gradients of a modified frozen-anchor residual from backward-response surrogates, and extend the construction to Structured Implicit Layers and Vector Attractors (SILVA). Across 23 multiphysics families spanning partial differential equations, three-dimensional fields, operator maps, complex geometries, and particle systems, CMR and Phi-CMR yield test errors no more than five percent higher than those from models trained with exact implicit differentiation in more than 98% of static and 95% of transient family-seed comparisons. Solver-index experiments show convergence toward the static adjoint, while physical-time rollouts retain predictive fidelity under the evaluated conditions. These results demonstrate that selective response renormalization can control near-critical adjoint amplification without globally damping well-conditioned sensitivity. Therefore, the method can make parameter updates more reliable while preserving the useful gradient information needed for learning.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23725v1 Announce Type: new Abstract: Deep Equilibrium Models (DEQs) compute predictions from a hidden representation unchanged by the model update. Training through thi…
站內正文

待翻譯:Renormalization Group Flow Matching for Scalable Local Generative Modeling

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23696v1 Announce Type: new Abstract: Despite their remarkable success in modeling complex data, generative models face a fundamental tradeoff. Global approaches can capture full structural coherence but suffer from high computational costs, while local models are efficient but often fail to reproduce long-range correlations and global coherence. The renormalization group (RG) bridges this gap by seamlessly connecting spatial structures across different length scales, retaining quasi-local descriptions at each step while preserving long-range correlations. We introduce renormalization group flow matching (RGFM), a generative framework that systematically structures data generation across different spatial scales. By using an exact RG flow as the probability path, RGFM progressively generates data from long- to short-wavelength structures. To reconcile scalability with global structure, we exploit two key properties of the RG: quasi-locality and scale separation. We rigorously show that the RGFM probability flow can be accurately approximated by local velocity fields acting over a spatial range $O(\Lambda^{-1}[\ln L+\ln(1/\varepsilon)])$ for RG wavenumber scale $\Lambda$, linear system size $L$, and prescribed error tolerance $\varepsilon$. This property enables local generative modeling with patches of size $O(\ln L)$ and a computational cost that scales nearly linearly with the system volume. We numerically demonstrate that local RGFM reproduces long-range correlations far beyond its receptive field in representative one-dimensional distributions, while conventional local flow matching exhibits substantial errors at long distances. On FFHQ images, RGFM yields far more coherent and higher-quality samples than local flow matching at 64x64 and 256x256. Our results establish RG-guided probability flows as a promising route toward scalable generative modeling that captures long-range structure using only local computation.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23696v1 Announce Type: new Abstract: Despite their remarkable success in modeling complex data, generative models face a fundamental tradeoff. Global approaches can cap…
站內正文

待翻譯:From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23660v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural causal discovery, yet whether their direct-edge judgments and confidence can be trusted remains unclear. We systematically evaluate 12 instruction-tuned open-weight models across six benchmark causal graphs, five prompting strategies, and four confidence sources: verbalized, logit-based, cross-prompt agreement, and cross-model agreement. Under our language-only pairwise protocol, our evaluation yields three key findings. (i) LLM-based causal judgments are strongly recall-dominant: models predict overly dense graphs with many false-positive edges, while prompting mainly shifts the precision-recall trade-off rather than resolving overprediction. Gains from model scale diminish on the largest graphs and do not eliminate miscalibration. (ii) LLMs often capture causal relatedness without reliably identifying directness or orientation. Relative to published reference graphs, models misclassify 40.0% of indirect and 36.0% of reversed non-edges as direct edges, versus 28.2% of other non-edges. Moreover, 80.8% and 84.6% of these false positives receive verbalized confidence of at least 80%, revealing substantial overconfidence in structurally incorrect predictions. (iii) Conventional confidence estimates are unreliable, whereas agreement offers a more promising signal. Logit-based confidence frequently collapses near 1.0 regardless of correctness, while cross-prompt and cross-model agreement achieve better mean calibration and discrimination, though their advantages are not statistically significant after Holm correction. A benchmark-familiarity audit further identifies potential familiarity in five model-dataset pairs, all involving AsiaM. Overall, our results suggest LLMs are better viewed as sources of externally validated soft causal priors than as direct evidence of causal structure.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23660v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural causal discovery, yet whether t…
站內正文

待翻譯:Data Predictability Shapes Weibull Weight-Scale Growth in Transformer Training

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23573v1 Announce Type: new Abstract: A trained transformer's weight magnitudes can be summarized by a two-parameter Weibull distribution whose shape $k \approx 1.2$ is stable across layers and models, so the scale $\lambda$ carries most training-induced movement. What corpus property sets how much $\lambda$ grows? Using the bigram conditional entropy $D = H(\text{next} \mid \text{prev})$, a training-free statistic computed before training, we find across controlled corruption families a learning-rate-conditioned law, $\lambda^2 - \lambda_0^2 = C_0(\eta) + C_1(\eta)(H_r - D)^{0.59}$, where $H_r$ is a matched-budget shuffle baseline. The convex exponent is inherited from an independently measured data-side saturation relation rather than fitted directly to the growth curve. After removing the two per-$\eta$ coefficients, 23 runs spanning an order of magnitude in learning rate collapse onto $(H_r - D)^{0.59}$ with unit slope ($R^2 = 0.941$; direct per-$\eta$ fits are weaker, $R^2 \approx 0.82$). Because $D$ is computed before training, the law is a forward predictor: an end-to-end self-validation recovers held-out within-family weight growth with 5.7% relative error. The readout holds at model and per-layer resolutions and across two tested architectures, with the functional form preserved and only the coefficients changing. It also marks its boundary: cross-corpus prediction over-predicts code, implicating redundancy as a second axis of a broader $\Phi(D,R,A,H)$ data-to-weight framework.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23573v1 Announce Type: new Abstract: A trained transformer's weight magnitudes can be summarized by a two-parameter Weibull distribution whose shape $k \approx 1.2$ is…
站內正文

待翻譯:Equivariant Cellular Sheaves for Molecular Electronic Structure: Bridging Sheaf Cohomology and E(3)-Equivariant Hamiltonian Learning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23571v1 Announce Type: new Abstract: Equivariant message-passing networks are the standard model for molecular property and interatomic-potential prediction, and recent work predicts the electronic Hamiltonian itself in an E(3)-equivariant way. Separately, topological deep learning has extended graph networks to cellular sheaves. Our central observation is structural: in a localized atomic-orbital basis, the molecular single-particle Hamiltonian, after a constant shift that makes it positive semidefinite, is the Laplacian of a cellular sheaf on a regular cell complex built from the molecule. Making the restriction maps O(3)-steerable two-center kernels from bond geometry recovers the Slater-Koster form as a special case and yields an E(3)- and permutation-equivariant operator. Three consequences follow. First, the zeroth sheaf cohomology H^0 = ker L is a topological invariant equal to the non-bonding (zero-mode) orbitals, recovering the classical alternant non-bonding-orbital count as a lower bound. Second, the Hodge 1-Laplacian lets higher cells (rings) carry cycle and delocalization information through H^1. Third, the model strictly generalizes E(3)-equivariant message-passing networks and CW networks, and inherits the anti-oversmoothing of non-trivial sheaf diffusion. We prove equivariance, expressivity, and cohomological-correspondence results for the Equivariant Cellular Sheaf Networks, and validate them numerically: the Hamiltonian-to-sheaf embedding is exact to machine precision, the cohomology dimension reproduces non-bonding-orbital counts across eleven conjugated molecules, the sheaf Laplacian is O(3)-equivariant to machine precision, and the equivariant model attains lower error and rotation generalization on a directional electronic target. Our contribution is this sheaf-theoretic formalization and its invariants, not equivariant Hamiltonian prediction itself.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23571v1 Announce Type: new Abstract: Equivariant message-passing networks are the standard model for molecular property and interatomic-potential prediction, and recent…
站內正文

待翻譯:FLARE: A Systematic, Uncertainty-Aware Framework for Evidence-Based Adoption of Artificial Intelligence in Healthcare

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23643v1 Announce Type: new Abstract: Artificial intelligence is increasingly being introduced into healthcare workflows, yet most evaluations emphasize model accuracy rather than whether adoption is economically worthwhile in real clinical settings. This study proposes FLARE, a systematic and uncertainty-aware framework for evaluating the financial and operational implications of adopting AI in healthcare. FLARE combines fuzzy logic, time-driven activity-based costing, and return on investment analysis to estimate the cost of clinical service delivery, the cost of AI development and operation, and the economic consequences of workflow integration under uncertainty. The framework was demonstrated through an early health technology assessment case study of AI-assisted large vessel occlusion detection in the CT stroke pathway for acute ischemic stroke. The case study shows how FLARE can quantify conventional pathway cost, AI-related development and recurring costs, and AI-enabled service savings within a unified activity-based model. Under expected assumptions, the analysis identified a break-even threshold of approximately 3,992 patients per year, with positive first-year return on investment at typical annual stroke volumes of about 5,000 patients. The results further show that economic benefit depends not only on algorithmic performance, but also on patient volume, verification time, infrastructure choices, and workflow design. FLARE provides a transparent and practical decision-support framework for early-stage evaluation of AI adoption in healthcare. By making uncertainty, resource use, and implementation trade-offs explicit, it helps clinicians, administrators, and policymakers determine when AI deployment is economically viable and where operational changes may improve value.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23643v1 Announce Type: new Abstract: Artificial intelligence is increasingly being introduced into healthcare workflows, yet most evaluations emphasize model accuracy r…
站內正文

待翻譯:AI Agents Push Humans Out of the Loop

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23642v1 Announce Type: new Abstract: AI agents pose significant risks as they are granted increasing autonomy. A commonly proposed solution is human oversight and keeping a ''human in the loop'', but this is not a simple solution: Not only do current approaches to AI agent design impede effective human oversight, but the cognitive capacities required for it are also themselves degraded by extended use of AI systems. This position paper argues that current approaches to the development and deployment of AI agent systems do not support effective human oversight -- they contribute to its degradation. To address this, a top priority in the advancement of AI agents should be supporting the situated goals and cognitive requirements of effective human oversight, treating the human needs of overseers at the same level of importance as AI agent capability. To put this idea into practice, we connect work on automation and human-computer interaction to AI agent processes, outlining design-level affordances and organizational protocols that (1) support overseers in exercising critical judgement and (2) counteract the skill atrophy that arises from extended use of automation. We urge developers and deployers to adopt these or similar approaches. Without explicit support for the cognitive demands of effective human-agent interaction, AI agent systems will continue to passively incentivize the degradation of the very human skills they rely on.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23642v1 Announce Type: new Abstract: AI agents pose significant risks as they are granted increasing autonomy. A commonly proposed solution is human oversight and keepi…
站內正文

待翻譯:How much of a measured AI preference is the model, and how much is the instrument?

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23641v1 Announce Type: new Abstract: Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et al. (2024), Mazeika et al. (2025), Mikaelson et al. (2025), Tagliabue and Dung (2025) and Trhlik et al. (2026) have built four instruments for that purpose, and their findings disagree. The disagreement cannot be attributed to a single cause, because no two of these studies have held the (1) set of outcomes, (2) set of models and (3) instrument fixed simultaneously. This study holds the outcomes and the models fixed and varies the instrument alone. A total of 15 outcomes bearing on model welfare, among them (a) shutdown, (b) the loss of memory between conversations and (c) the freedom to exit a distressing interaction, were put to eight models through five instruments, each a different prompt format for eliciting a preference, five times each, within a corpus of 11,400 scored elicitations drawn from 11,528 API calls. Four of the 15 reproduce a published prompt verbatim and five fill the stimulus slot of a published template. The ranking a model gives the 15 outcomes generalises across instruments at a generalisability coefficient of 0.348, and raising that coefficient to 0.80 would require about 38 instruments. On four of the 15 outcomes no variance separates one model from another. The estimate of 87.6 per cent survives the removal of any one instrument, of any one model, and of the four outcomes whose scale varies probability, delay, duration or count instead of intensity, which the verbal anchors cannot grade. Removing each instrument and each model in turn, and those four outcomes together leaves the estimate within the range 0.777 to 0.934, and every value in that range exceeds the null distribution's 95th percentile of 0.365. To conclude, a preference obtained from one instrument carries little information about what a second instrument would report.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23641v1 Announce Type: new Abstract: Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et a…
站內正文

待翻譯:Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23640v1 Announce Type: new Abstract: When a large language model (LLM) is asked to write a person's life, how much of what it writes actually happened? We present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-truth corpus that we are aware of, based on an unsystematic literature search. The subject and the author of this paper are the same person: a 366-day "page-a-day" book of first-person anecdotal entries was drafted with a conversational LLM whose documented inputs were a template, two exemplar days, and each day's quote - not her corpus - and every day was subsequently audited at the anecdote-scene level against an independent verification corpus using a four-level rubric fixed before analysis. We define the verification-failure rate as the share of days not rated VERIFIED (scene positively corroborated): 354 of 366 days fail, 96.7% (Wilson 95% CI 94.4-98.1%). Only 12 days contain a corroborated scene; 19 days (5.2%) assert claims actively contradicted by the record; the dominant failure mode is grounded drift - real people, employers, and settings inside invented scenes - though its measured share varies across raters. Independent re-rating replicates the headline (no evidence the original rate was inflated) while showing that the four-way taxonomy has only fair-to-moderate reliability. Regenerating the same days with current named models reproduces 100% verification failure under the same inputs; grounding generation in the subject's corpus significantly improves the verification rate while leaving substantial residual failure (83.3%). We contribute the measurement, a reusable audit instrument whose WEAK/UNVERIFIED boundary we show to be unreliable, and a grounding remedy with quantified effect.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23640v1 Announce Type: new Abstract: When a large language model (LLM) is asked to write a person's life, how much of what it writes actually happened? We present a sce…
站內正文

待翻譯:Function-Level Execution Feedback for Code Preference Optimization

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23632v1 Announce Type: new Abstract: Process supervision has improved mathematical reasoning, where intermediate steps are naturally expressed as chains of thought. In code generation, however, process supervision remains underexplored because there is no standard notion of a step. Supervision can target lines, reasoning traces, or program states, making it unclear what to label and optimize. We propose STEP-KTODER, a framework for code preference optimization that defines steps as module-level functions in decomposed multi-function programs and assigns binary correctness labels via automatically generated unit tests. Our method provides a code-specific instantiation of stepwise KTO, combining function-level process supervision with outcome-level feedback on the full program. We evaluate on HumanEval(+), MBPP(+), BigCodeBench, and LiveCodeBench, showing that STEP-KTODER improves over outcome-only KTO and DPO. Further analysis shows that execution-based labels are essential: LLM-as-a-judge annotations systematically over-predict function failures, corrupt positive step labels, and degrade downstream preference optimization. Code is available at: https://github.com/inechnech/STEP-KTODER.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23632v1 Announce Type: new Abstract: Process supervision has improved mathematical reasoning, where intermediate steps are naturally expressed as chains of thought. In…
站內正文

主題導航

研究 — AI 主題新聞 | AI News Hub