跳到主要內容
AI News HubLIVE

本期報導已收集,譯文與分析尚待補全。可展開其餘更新查看來源內容。

其餘更新(132 條)
Agent

待翻譯:The Interfaces Are Arriving

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The most consequential AI news of the past year came from a standards body. In December 2025, Anthropic donated the Model Context Protocol to the newly formed Agentic AI Foundation, a directed fund under the Linux Foundation cofounded by Anthropic, Block, and OpenAI, with support from Google, Microsoft, AWS, Cloudflare, and Bloomberg. Six months earlier, […]

O'Reilly AI & ML Radar來源內容 · 翻譯待補全待翻譯:The Interfaces Are Arriving

待翻譯:Weave Router 2.0

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:Weave Router 2.0

待翻譯:Don't sleep on wrapture

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯: Graham Dumpleton's new monkey patching package wrapture is shaping up to be an indispensable tool for Python developers. I'm not sure why I've seen so little buzz about it! Graham has been posting new tutorials for it almost daily since the initial release on August 31st. Here's everything he's published so far: Introducing wrapture - a new monkey patching library that serves both testing and observability (think New Relic style tracing) at the same time. Unit testing with wrapture - how to use it for the same kinds of thing as unittest.mock. Recording calls with wrapture - recording method calls as timelines and processing and displaying them as trees. Phased behaviour in wrapture - arranging patched methods to change behavior across multiple calls. Beyond ca…

Simon Willison's Weblog來源內容 · 翻譯待補全待翻譯:Don't sleep on wrapture

待翻譯:Open-Source AI & Open Models Reading List

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:How to get up to speed on open models and their implications.

Interconnects (Nathan Lambert)來源內容 · 翻譯待補全待翻譯:Open-Source AI & Open Models Reading List

待翻譯:Fine-Tuning Agentic AI: A Practical Guide

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:In this article, you will learn how to fine-tune an agentic AI system holistically, covering all four critical dials: training data, parameter-efficient fine-tuning, runtime hyperparameters,...

Machine Learning Mastery來源內容 · 翻譯待補全待翻譯:Fine-Tuning Agentic AI: A Practical Guide

待翻譯:Operating Mode as Runtime State: A Contract for Enterprise

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:During a service incident, a customer-remediation workflow is moved onto an emergency route because the situation is critical and the team needs a fast resolution. Approvals are shortened, a priority queue is opened, and an on-call agent is cleared to use an alternate procedure until the service recovers. The incident ends, but the route stays […]

O'Reilly AI & ML Radar來源內容 · 翻譯待補全待翻譯:Operating Mode as Runtime State: A Contract for Enterprise

待翻譯:Rapidly scaling online storage to serve over 1 billion ChatGPT users

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Learn how OpenAI evolved Habitat from a Python library into a globally distributed storage platform serving 1 billion ChatGPT users and 22M requests per second.

OpenAI News來源內容 · 翻譯待補全待翻譯:Rapidly scaling online storage to serve over 1 billion ChatGPT users

待翻譯:How to Build the Unified Data Foundation Drug Discovery AI Depends On

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:This article is sponsored by CDD Vault and was written, edited, and published in alignment with our Emerj sponsored content guidelines. Learn more about our thought leadership and content creation services on our Emerj Media Services page.​ Drug discovery is one of the slowest, costliest processes in enterprise R&D. Developing a single FDA-approved therapy typically […]

Emerj AI Research來源內容 · 翻譯待補全待翻譯:How to Build the Unified Data Foundation Drug Discovery AI Depends On

待翻譯:Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Sakana AI has released Fugu Max and Fugu Ultra v2, 2 models built on the same learned orchestration architecture. Fugu Max routes tasks to lean open and specialized models, including NVIDIA Nemotron, at $2/$6 per 1M tokens. Fugu Ultra v2 targets peak capability, scoring 48.3 on Chartography and 74.3 on DeepSWE. The post Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration appeared first on MarkTechPost.

MarkTechPost來源內容 · 翻譯待補全待翻譯:Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

待翻譯:LTLDiff: Finite Linear Temporal Logic-Guided Data Generation and Diffusion Policies for Multi-agent Robotic Manipulation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.11043v1 Announce Type: new Abstract: Multi-agent robotic manipulation tasks require coordination among agents to satisfy task-level temporal, logical, and safety constraints. Recently, diffusion policies have been used to perform the task. However, they still suffer from desynchronization, incorrect action ordering, and coordination failures in tasks that require simultaneous or sequential multi-agent interaction. Therefore, LTLDiff is proposed as a framework that combines Finite Linear Temporal Logic (LTLf) specification learning for both the generation of demonstrations and learning via diffusion policies. Each task has a specific LTLf formula that is learned from a set of natural language instructions using a large-scale language model. To enable…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:LTLDiff: Finite Linear Temporal Logic-Guided Data Generation and Diffusion Policies for Multi-agent Robotic Manipulation

待翻譯:Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09448v1 Announce Type: new Abstract: As agentic systems getting adopted rapidly in safety critical applications, it is vital to measure the confidence associated with the agentic actions. In comparison to the traditional machine learning systems, agentic workflows have complex failure modes with planning, tool invocation and dynamic environment interactions. In this paper, we investigate whether model's internal representations provide stronger signals of eventual task success in multi-turn agentic setups. We introduce two complementary methods: Latent Trajectory Dynamics (LTD), which summarizes changes in residual-stream representations across an an interaction trajectory, and the Action Representation Probe (ARP), which predicts success from repres…

arXiv AI來源內容 · 翻譯待補全待翻譯:Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations

待翻譯:Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09418v1 Announce Type: new Abstract: World Action Models (WAMs) couple predictive world modeling with action generation, allowing anticipated future states to guide agent behavior. Although WAMs are rapidly advancing embodied AI, general-purpose counterparts remain largely unexplored in games. Existing game-oriented approaches often combine action-conditioned world models with external policies and reward functions to realize WAM-like decision-making, yet they operate mainly in 2D visual observation space and do not instantiate persistent 3D geometry. Extending this paradigm to 3D games introduces a distinct challenge. In autonomous driving and robotics, the physical environment exists independently of the model, providing a persistent 3D world in wh…

arXiv AI來源內容 · 翻譯待補全待翻譯:Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration

待翻譯:Decision-Focused Active Learning for Scale-Aware Critical-Materials Recovery

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09413v1 Announce Type: new Abstract: Choosing a recovery process for scale-up requires connecting laboratory results with product requirements, process costs, and scale effects. We analyze records from Pacific Northwest National Laboratory's Computer Intelligence for Critical Element Recovery and Optimization (CICERO) workflow for autonomous selective precipitation. Active learning uses prior results to choose experiments. In a conditional retrospective benchmark with fitted models and recycled neodymium-iron-boron (NdFeB) magnet records, active learning finds the best recorded result with fewer experiments than nonadaptive space filling. Enrichment is the selected rare-earth-to-iron ratio relative to that in the feed. Adaptive policies reach the rec…

arXiv AI來源內容 · 翻譯待補全待翻譯:Decision-Focused Active Learning for Scale-Aware Critical-Materials Recovery

待翻譯:The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09395v1 Announce Type: new Abstract: Language models act through tools, yet practical agents face libraries containing thousands of interfaces. We introduce the tool menu as the short, ordered subset of available tools shown to an agent before execution. The agent can call only tools in this menu. Multi-step tasks require the final action and the prerequisite tools that create its inputs in a usable order. Current constructors rank tools by request relevance, which can surface the final action while omitting or delaying less obvious producers. We introduce the state path, a pre-execution route from the observable request state to the desired outcome, and propose State-Path Tool Menu to learn it. Our framework treats the menu as an execution prior ove…

arXiv AI來源內容 · 翻譯待補全待翻譯:The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents

待翻譯:An Autonomous GeoAI Agent for Arctic Eco-Navigation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09374v1 Announce Type: new Abstract: Arctic maritime navigation is becoming increasingly important as changing sea-ice conditions expand seasonal accessibility while simultaneously introducing substantial operational, environmental, and community risks. Arctic route planning is inherently a multi-criteria problem: routes that improve vessel safety or efficiency may increase exposure to sea ice, sensitive ecosystems, or nearby communities. Existing routing methods prioritize travel time, fuel use, and navigational risk, often overlooking ecological and community impacts. We introduce a human-in-the-loop, multi-agent GeoAI system for Arctic eco-navigation that integrates operational, physical, ecological, and community-related criteria within a unified…

arXiv AI來源內容 · 翻譯待補全待翻譯:An Autonomous GeoAI Agent for Arctic Eco-Navigation

待翻譯:Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09233v1 Announce Type: new Abstract: How can language model agents effectively leverage libraries of reusable knowledge to solve long-horizon tasks? Recent work has increasingly focused on agent skills: reusable capabilities represented as skill packages, i.e., multi-file bundles containing instructions, scripts, and other resources that help agents perform specific tasks. Agent skills are typically executed by loading their skill instructions into an agent's context and relying on the agent to follow them. As task horizons grow, however, this approach becomes increasingly brittle, because reasoning quality degrades as more information accumulates in the context window. We investigate an alternative approach in which skill packages are instead invoke…

arXiv AI來源內容 · 翻譯待補全待翻譯:Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks

待翻譯:Adaptive Entangled Game Modules in Artificial General Intelligence

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09226v1 Announce Type: new Abstract: We introduce a probability-wave framework for modeling the collective behavior of interacting adaptive agents, deriving testable eigenmodes through a generalized behavioral intelligence (GBI) nonlocal probability-wave equation. This framework captures a broad range of human intelligence behaviors with analytical mechanisms and offers an indirect method to examine the Liu-Chen-Ao (LCA) hypothesis of nonlocal entangled nerve fibers in the brain through collective trader behaviors. Our empirical analysis of Chinese intraday stock market data demonstrates that adaptive entangled game modes explain 82-94% (89% overall) of observed decision patterns, a sharp contrast to the predictions of neoclassical finance based on i…

arXiv AI來源內容 · 翻譯待補全待翻譯:Adaptive Entangled Game Modules in Artificial General Intelligence

待翻譯:OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09203v1 Announce Type: new Abstract: Existing benchmarks for autonomous AI scientists evaluate only final outputs---generated code, hypotheses, or papers---yet discard the reasoning process by which those outputs were obtained. This makes it impossible to audit scientific methodology, diagnose failure modes, or distinguish systematic reasoning from fortunate guessing. We present \textbf{OpenDiscoveryTrace}, a public dataset of 558 complete AI scientific agent trajectories that captures how models reason, not just what they produce. Each trajectory records a structured 9-field-per-step trace---including thoughts, tool calls, observations, errors, revision triggers, and self-reported confidence---as models execute 124 scientific tasks spanning drug dis…

arXiv AI來源內容 · 翻譯待補全待翻譯:OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

待翻譯:Oracle says AI will save it from the SaaSpocalypse, not bring it on

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:It’s a better interface, can speed installations, and drives IaaS sales too

The Register AI + ML來源內容 · 翻譯待補全待翻譯:Oracle says AI will save it from the SaaSpocalypse, not bring it on

待翻譯:Any Nix package, live in your browser

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯: Any Nix package, live in your browser Farid Zakaria calls this his "magnum opus of Nix work", and I can see why. trynix.dev provides a qemu-wasm powered x86_64 Linux virtual machine running entirely in your browser through WebAssembly. That VM can then be booted with any Nix package from the past 13 years. They are URL addressable, so you can navigate to this page: https://trynix.dev/?pkg=python3%403.6.2 Then click "Load" and get an interactive shell against a virtual machine running Python 3.6.2 from 2017. Farid is building all sorts of neat things on top of this. One recent example: Review a pull request by booting it introduces trynix-preview, described like this: GitHub action that comments a link on a pull request which lets you boot the PR’s build in the…

Simon Willison's Weblog來源內容 · 翻譯待補全待翻譯:Any Nix package, live in your browser

待翻譯:Cadenya

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:Cadenya

待翻譯:AWS open-sources Pizza Bot: email-style inbox for background AI agents

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Amazon Web Services (AWS) has released a new open-source application dubbed Pizza Bot, which gives developers an email-style inbox for The post AWS open-sources Pizza Bot: email-style inbox for background AI agents appeared first on The New Stack.

The New Stack AI來源內容 · 翻譯待補全待翻譯:AWS open-sources Pizza Bot: email-style inbox for background AI agents

待翻譯:OzBrain

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:OzBrain

待翻譯:GitHub Copilot app for Beginners: Using the diff, terminal, and browser

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Checking agent-generated code usually means hopping between tabs. Learn how to view diffs, run terminal commands, and preview web apps side by side in the GitHub Copilot app. The post GitHub Copilot app for Beginners: Using the diff, terminal, and browser appeared first on The GitHub Blog.

GitHub AI & ML來源內容 · 翻譯待補全待翻譯:GitHub Copilot app for Beginners: Using the diff, terminal, and browser

待翻譯:OpenAI Launches the Agents API in Public Beta, Putting the Codex Harness Behind One API Call

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:OpenAI has released the Agents API in public beta. It gives developers the same harness and infrastructure that run Codex. OpenAI hosts and maintains the harness. Developers run the agent’s compute in an OpenAI-managed sandbox, their own infrastructure, or a partner sandbox. Is it deployable? Yes. It is live for all developers in public beta. […] The post OpenAI Launches the Agents API in Public Beta, Putting the Codex Harness Behind One API Call appeared first on MarkTechPost.

MarkTechPost來源內容 · 翻譯待補全待翻譯:OpenAI Launches the Agents API in Public Beta, Putting the Codex Harness Behind One API Call

待翻譯:Native is now the future of mobile at Shopify

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯: Native is now the future of mobile at Shopify Shopify are moving from React Native back to separate Swift and Kotlin codebases for their native apps, for the exact reason you would expect: We decided to switch from native to React Native in 2020 for three reasons: Stop building the same features twice Allow developers to work across the stack Spend less time chasing feature parity and more time shipping value [...] Native still means building and maintaining software on two platforms, that cost has not disappeared. What changed is that agents can now do enough of the implementation, translation, testing, and review work that it’s no longer the deciding factor it was in 2020. It's a well-written post, which gives full credit to React Native as a great platform…

Simon Willison's Weblog來源內容 · 翻譯待補全待翻譯:Native is now the future of mobile at Shopify

待翻譯:“Six tools, one harness”: Salesforce loops together a six-pack of favorites

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Salesforce introduced its Salesforce Enterprise AI Harness on Thursday as a formalized amalgamation of the AI harness concepts and infrastructure The post “Six tools, one harness”: Salesforce loops together a six-pack of favorites appeared first on The New Stack.

The New Stack AI來源內容 · 翻譯待補全待翻譯:“Six tools, one harness”: Salesforce loops together a six-pack of favorites

待翻譯:Amazon Quick is now generally available on desktop

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Your teams get an AI assistant that handles real work while your data stays in your environment and your conversations stay private Today, the Amazon Quick desktop application is generally available on macOS and Windows. We’re also adding a new activity feed to the mobile experience on iOS and Android that consolidates email, calendar, CRM, […]

AWS Machine Learning Blog來源內容 · 翻譯待補全待翻譯:Amazon Quick is now generally available on desktop

待翻譯:A Candid Abacus AI Review: The All-in-One AI Platform for Professionals & Enterprises

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:If you’re paying for ChatGPT, Claude, and another AI tool simultaneously, this review is for you. It covers what an AI platform like Abacus AI actually includes, how the credit system works in practice, and whether it genuinely replaces your current stack or just adds to it.

KDnuggets來源內容 · 翻譯待補全待翻譯:A Candid Abacus AI Review: The All-in-One AI Platform for Professionals & Enterprises

待翻譯:Build an end-to-end RFI questionnaire workflow using Amazon Quick Automate

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Learn how to build an end-to-end RFI questionnaire workflow with Amazon Quick Automate. Read a multi-tab RFI workbook from Amazon S3, use natural-language prompts to extract and structure the questionnaire data, refine the workflow through conversation, and write clean CSV output back to Amazon S3 — cutting development from days to hours.

AWS Machine Learning Blog來源內容 · 翻譯待補全待翻譯:Build an end-to-end RFI questionnaire workflow using Amazon Quick Automate

待翻譯:When Content Is Free, Trust Is the Product

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:There is more technical content available today than any human being could read in a thousand lifetimes. Every topic has a dozen YouTube videos, three Substack posts, a GitHub repo, and a Reddit thread, most created in the last six months and, in many cases, technically accurate. And yet most of the professionals I talk […]

O'Reilly AI & ML Radar來源內容 · 翻譯待補全待翻譯:When Content Is Free, Trust Is the Product

待翻譯:Physical AI Takes the Wheel: How the World’s Robotaxi Leaders Are Building With NVIDIA Technologies

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The global robotaxi market — physical AI’s first commercial breakthrough — is projected to reach $400 billion by 2035, with over 6 million commercial vehicles in operation as driverless fleets are already moving people through some of the world’s busiest and most complex streets. Deploying a driverless vehicle is one challenge. Scaling a fleet is […]

NVIDIA Blog來源內容 · 翻譯待補全待翻譯:Physical AI Takes the Wheel: How the World’s Robotaxi Leaders Are Building With NVIDIA Technologies

待翻譯:How a researcher uses Codex and ChatGPT to search for new antimicrobial molecules

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:César de la Fuente’s lab uses Codex and ChatGPT to search living and extinct genomes for antimicrobial candidates to fight drug-resistant infections.

OpenAI News來源內容 · 翻譯待補全待翻譯:How a researcher uses Codex and ChatGPT to search for new antimicrobial molecules

待翻譯:Introducing Projects · Cursor

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Blog / product Today we're launching Projects in Cursor. Projects lets you take on larger bodies of work, such as a feature, a migration, or a full app. It maintains context over months of work, delegates tasks to thous…

Cursor Blog來源內容 · 翻譯待補全待翻譯:Introducing Projects · Cursor

待翻譯:Blaxel is joining Baseten to build the future of agentic infrastructure

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:News Blaxel is joining Baseten to build the future of agentic infrastructure Baseten has acquired Blaxel Authors Amir Haghighat Tuhin Srivastava Paul Sinai Last updated September 10, 2026 Share Today, Blaxel is joining…

Baseten Blog來源內容 · 翻譯待補全待翻譯:Blaxel is joining Baseten to build the future of agentic infrastructure

待翻譯:Gen-1 Slides: Opus 5-level decks at a fraction of the cost

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Gen-1 Slides: Opus 5-level decks at a fraction of the cost Join us for our inaugural conference, Forge 2026 Blog Gen 1 Slides Opus 5 Level Decks At A Fraction Of The Cost Gen-1 Slides: Opus 5-level decks at a fraction o…

Fireworks AI Blog來源內容 · 翻譯待補全待翻譯:Gen-1 Slides: Opus 5-level decks at a fraction of the cost
模型

待翻譯:Mistral Bets Enterprise AI Will Be About Control, Not Just Intelligence

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The French AI lab is using a $3B fundraise to sell control over AI infrastructure, not just model power -- a shift in direction that could matter to U.S. firms in Europe too.

AI Business來源內容 · 翻譯待補全待翻譯:Mistral Bets Enterprise AI Will Be About Control, Not Just Intelligence

待翻譯:Soft-deprecating re.match()

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯: Soft-deprecating re.match() Python has a concept of soft deprecation, where APIs are marked as "should no longer be used to write new code" without any promise/threat to remove them in the future. Python 3.15 release manager Hugo van Kemenade describes how in the upcoming 3.15 release soft deprecation has come for the venerable but deeply confusing re.match() function. It's now available with the much clearer alternative re.prefixmatch() name - reflecting how it anchors at the beginning of the string but not the end. Most of the time you probably want re.search() (match this pattern anywhere in the string) or re.fullmatch() (match the entire string) instead. Via Lobste.rs Tags: python, regular-expressions

Simon Willison's Weblog來源內容 · 翻譯待補全待翻譯:Soft-deprecating re.match()

待翻譯:Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Cohere has released North Small Translate, an open-weight Mixture-of-Experts model built for machine translation across 50 languages. It uses 25B of its 218B parameters per token and scores 83.6 on Cohere's WMT26 evaluation. Weights are free for non-commercial use, with commercial access through Cohere Model Vault or RWS Language Weaver. The post Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages appeared first on MarkTechPost.

MarkTechPost來源內容 · 翻譯待補全待翻譯:Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

待翻譯:Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Google Research has released ToolGrad, an ACL 2026 Findings framework that inverts tool-use dataset generation: it builds a verified API chain first, then writes the matching user query. Guided by textual "gradients" from a 4-module propose-execute-select-update loop, ToolGrad reaches a 99.8% pass rate on ToolBench versus 63.8% for DFS search. Gemma-3-12B fine-tuned on only 500 samples scores 83.1 on BFCL, next to Gemini 2.5 Pro at 83.2. Code, dataset, and models are public under Apache-2.0. The post Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation appeared first on MarkTechPost.

MarkTechPost來源內容 · 翻譯待補全待翻譯:Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation

待翻譯:ObstaDiff: Generalizable Diffusion Policy Learning via Obstacle-aware Representations

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10918v1 Announce Type: new Abstract: Imitation learning has achieved impressive results in robotic manipulation, yet most existing approaches assume clean backgrounds and lack explicit mechanisms for obstacle-aware motion generation. Extending such policies to cluttered, real-world scenes with unstructured obstacles remains a key generalization challenge. We present ObstaDiff, a decomposed diffusion-policy framework with a lightweight obstacle-aware visual encoder. ObstaDiff extracts a structured target-obstacle-background representation, enabling the downstream alignment policy to generate end-effector trajectories toward a target-centered bottleneck pose while reasoning about surrounding obstacles. We evaluate ObstaDiff on 61 real-robot greenhouse…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:ObstaDiff: Generalizable Diffusion Policy Learning via Obstacle-aware Representations

待翻譯:IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10915v1 Announce Type: new Abstract: Vision-language-action (VLA) policies leverage pretrained vision-language backbones to achieve strong cross-task generalization. A leading design couples this backbone with a dedicated continuous action head trained via diffusion or flow matching. However, such heads rely on iterative multi-step sampling, for example 10 Euler steps in $\pi_{0.5}$. This creates an inference bottleneck that produces stop-and-go movement in the robot and slower task completion. We introduce IMLE-VLA, which replaces the iterative action head with a single-step conditional generator trained via conditional Implicit Maximum Likelihood Estimation (cIMLE). The cIMLE objective promotes multimodal action coverage, avoiding the mode collapse…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies

待翻譯:ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10895v1 Announce Type: new Abstract: Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as the decision coreof household robots. Existing evaluations, however, probe intuitive physics passively through question answering over videos, or target deliberate, long-horizon tasks such as navigation and rearrangement; none measure whether a model can turn physical understanding into immediate, safety-critical action. We introduce ReactHuman, the first physics-grounded benchmark for human-like reactive decision-making, in which the evaluated MLLM acts as the brain of a simulated humanoid facing…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

待翻譯:HuRo: Robotizing Human Videos for Scalable VLA Pretraining

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10706v1 Announce Type: new Abstract: Human video datasets have emerged as a compelling alternative to expensive real-robot data, offering rich diversity at scale. To bridge the human-to-robot embodiment gap, existing approaches either robotize videos in task-matched settings or address observation and action alignment separately at scale. In this work, we systematically examine whether robotized human videos can provide effective and scalable supervision for pretraining vision-language-action (VLA) policies. To this end, we develop a robotization pipeline that converts heterogeneous human videos into robot-aligned observations and action trajectories while inferring missing intermediate signals across annotation levels. Using this pipeline, we constr…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:HuRo: Robotizing Human Videos for Scalable VLA Pretraining

待翻譯:Overpainting: Localized Context-aware Diffusion Image Editing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10811v1 Announce Type: new Abstract: We present "overpainting", an image editing operation which offers both control over the location of the edit and awareness of the previous content in that location. The overpainted area is given by a trimap, where white-annotated pixels must be edited, gray-annotated pixels may be edited, and black-annotated pixels must not be edited. This enables both precise and loose control, depending on user intent. We implement overpainting by adapting a pretrained image editing diffusion model using a combination of joint attention and low-rank adaption across input images with attention-dropout to balance the information flow between noise, source and mask images. We present a novel, automated, training data generation pi…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:Overpainting: Localized Context-aware Diffusion Image Editing

待翻譯:TrajFusionNet+: Transformer-Based Prediction of Pedestrian Crossing Intention via Fusion of Trajectory Representations and Scene Graphs

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10806v1 Announce Type: new Abstract: The pedestrian crossing intention task involves predicting whether pedestrians are likely to cross the road from the point of view of an autonomous vehicle. We introduce TrajFusionNet+, a novel transformer-based model for pedestrian crossing intention prediction. TrajFusionNet+ combines sequential and visual representations of pedestrian trajectory with a graph-based representation of the scene context in order to predict pedestrian crossing intention. The proposed architecture builds upon our previous model, TrajFusionNet, and comprises three branches: a Sequence Attention Module (SAM), which processes a sequential representation of past and predicted pedestrian trajectories; a Visual Attention Module (VAM), whic…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:TrajFusionNet+: Transformer-Based Prediction of Pedestrian Crossing Intention via Fusion of Trajectory Representations and Scene Graphs

待翻譯:GRADE: Single-Frame Generative Radar Depth Estimation Under Visual Degradation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10756v1 Announce Type: new Abstract: Dense 3D depth perception fails under smoke, fog, and darkness because optical sensors cannot penetrate airborne particulates. mmWave radar remains usable and measures range accurately under these conditions, but its small aperture limits angular resolution. We present GRADE, which grounds a pretrained generative prior in single-frame radar geometry to estimate high-fidelity metric depth. GRADE first maps raw 4D radar spectra to coarse metric depth. A latent diffusion backbone then recovers structural detail while conditioning every denoising step on this estimate. A pixel-space adapter uses residual camera cues when available and is trained across clear, smoke-degraded, and occluded inputs so the full output appr…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:GRADE: Single-Frame Generative Radar Depth Estimation Under Visual Degradation

待翻譯:Meta-Learning for Data-Efficient Plant Growth Estimation via Vision Transformers and Fuzzy Clustering

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10749v1 Announce Type: new Abstract: Accurate plant growth estimation is essential for greenhouse monitoring, yet obtaining labeled data remains costly and time-consuming. To address this, we propose a few-shot regression framework that combines Vision Transformer (ViT) feature embeddings, clustering-based task construction, and gradient-based meta-learning, and show that task construction in embedding space is a primary driver of performance. The approach leverages an unlabeled image pool to organize data into structured tasks using fuzzy c-means clustering, enabling efficient learning from a small number of labeled samples. We systematically evaluate meta-learning methods and show that second-order methods (e.g., Model-Agnostic Meta-Learning varian…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:Meta-Learning for Data-Efficient Plant Growth Estimation via Vision Transformers and Fuzzy Clustering

待翻譯:MHE-Former: Multi-Hypothesis Transformers via Entropy Maximization for 3D Mesh Recovery

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10743v1 Announce Type: new Abstract: Monocular 3D hand and body mesh recovery often suffers from severe occlusion and ambiguity. Traditional deterministic methods typically regress a single optimal solution, leading to overconfident predictions. In this paper, we introduce an exploration--exploitation paradigm for ambiguous mesh recovery with multi-hypothesis learning and selection. Specifically, during exploration, based on our probabilistic formulation and entropy maximization, we propose a novel multi-hypothesis method referred to as MHE-Former. It is a Transformer-based multi-hypothesis framework, ensuring high training efficiency and label friendliness while generating plausible and diverse hypotheses. During exploitation, we propose Hypothesis…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:MHE-Former: Multi-Hypothesis Transformers via Entropy Maximization for 3D Mesh Recovery

待翻譯:AcFlow: Controlling Text-to-Image Diffusion Transformers via Learned Conditional Activation Flow

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10723v1 Announce Type: new Abstract: Text-to-image diffusion transformers (DiTs) are powerful generators, yet direct prompting provides limited control interface for style intensity and can fail to suppress unwanted concepts. To enable these controls, we introduce AcFlow, an inference-time controller that transports intermediate layer image-token activations through a learned concept-conditioned velocity field while keeping the base DiT frozen. A textual concept description specifies the desired intervention, while the integration horizon provides a continuous control parameter. The field produces token-varying, activation-dependent updates. With parameters shared across concepts within each task family, the field supports fine-grained descriptions a…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:AcFlow: Controlling Text-to-Image Diffusion Transformers via Learned Conditional Activation Flow

待翻譯:LLM-Anchored Paralinguistic Enrichment for Alzheimer's Disease Detection

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10896v1 Announce Type: new Abstract: Speech-based automatic detection of Alzheimer's disease (AD) provides a non-invasive and scalable approach to early cognitive screening. AD affects both lexical-semantic organization and speech production, including atypical pauses and word elongations. However, existing methods have yet to fully integrate these paralinguistic cues with linguistic content. We propose LLM-Anchored Paralinguistic Enrichment (LAPE), which enriches LLM-derived linguistic representations with paralinguistic cues through three coordinated innovations. The first is prosodic event textualization, which enables the LLM to model pauses and elongations jointly with lexical content by encoding them as explicit markers with bounded duration-aw…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:LLM-Anchored Paralinguistic Enrichment for Alzheimer's Disease Detection

待翻譯:Does Linguistic Structure Enrichment Enhance Coherence Assessment? Not With Current Architectures

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10893v1 Announce Type: new Abstract: Recent advances in large language models have transformed human-computer interaction. Despite their fluency, these models often produce texts that are grammatically correct but semantically incoherent, containing contradictions or disruptions in logical flow. This work investigates whether enriching text with syntactic and rhetorical information can improve incoherence prediction. Our experiments and analysis show that plain texts achieved higher accuracy because the added information was structurally and syntactically incompatible with the language model's architecture. Additionally, to demonstrate the practical importance of coherence assessment, we performed zero-shot experiments on a Brazilian disinformation d…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:Does Linguistic Structure Enrichment Enhance Coherence Assessment? Not With Current Architectures

待翻譯:Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10830v1 Announce Type: new Abstract: When a language model finds a sentence unusually cheap to predict, it is tempting to conclude that the sentence was in its training data. Almost every published test of that inference has had to guess which sentences were in the training data, the members, and which were not. This paper removes the guessing. Two model families, OLMo-2 and Pythia, publish their pretraining corpora, and a public index over those corpora returns the exact number of times any sentence appeared in each. Those counts make three questions answerable directly. The answers form a pincer, closing from two sides. At the duplication levels ordinary text actually has, five models from 1B to 13B parameters carry at most a faint trace of their o…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models

待翻譯:Larger Context Window, Fewer Overcorrections: Optimizing Prompts and Batching for Minimal-Edit Grammatical Error Correction

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10810v1 Announce Type: new Abstract: Minimal-edit Grammatical Error Correction (GEC) is a challenging task for zero- and few-shot prompted Large Language Models (LLMs), which systematically overcorrect and degrade $F_{0.5}$ by rewriting well-formed spans. While fine-tuning provides an effective solution, it imposes substantial infrastructure demands. We introduce a prompt-based approach that closes the gap to fine-tuned models through three advances in GEC prompting methodology. First, we introduce taxonomy-based instructions to enforce minimal-edit constraints with a comprehensive list of grammatical error rules, equipping the LLM with a bounded, metric-aligned scope of correctable edits, which benefits the strongest models while remaining model-dep…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:Larger Context Window, Fewer Overcorrections: Optimizing Prompts and Batching for Minimal-Edit Grammatical Error Correction

待翻譯:Analyzing Traditional and Neural Approaches to Multilingual Readability Assessment

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10792v1 Announce Type: new Abstract: Transformer-based models excel at Automatic Readability Assessment (ARA), yet feature-based models remain in active use because their predictions tie back to linguistic properties. This matters because readability labels are subjective and rater-dependent, so high accuracy on noisy ground truth may reflect surface patterns rather than the linguistic structure that defines difficulty. We test whether transformers internalize the same features as traditional models across Arabic, English, French, Hindi, and Russian using the ReadMe++ dataset. Shapley Additive Explanations (SHAP) identify the features driving traditional classifiers, which we then use as TCAV concept sets to probe multilingual XLM-R and language-spec…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:Analyzing Traditional and Neural Approaches to Multilingual Readability Assessment

待翻譯:Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10758v1 Announce Type: new Abstract: Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu language as a representative low-resource language. We generate Urdu-Stories, a corpus of 93 stories generated using three contemporary LLMs (GPT-5.1, Qwen-3-Max, DeepSeek-3.1). We manually annotate the errors present in them under a nine-label linguistic, semantic, and cultural taxonomy. Our notable findings suggest that LLMs often make basic errors of grammar and semantics. The stories lack coheren…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

待翻譯:Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10745v1 Announce Type: new Abstract: Multimodal entity linking grounds entity mentions in text and images to knowledge-base entries. These systems degrade on rare entities, but prior work measures rarity primarily through popularity-based metrics such as pageviews. We broaden this view using knowledge-graph structural metrics that capture how well an entity is documented and connected. These metrics identify many rare entities that popularity metrics miss. Across the resulting rare-entity slices, state-of-the-art accuracy drops by 15.4-39.9%, showing that different rarity definitions expose different failure modes. To address these failures, we introduce a simple, training-free framework in which a reasoning-capable vision-language model iteratively…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking

待翻譯:NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10715v1 Announce Type: new Abstract: We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducing an explicit and more challenging concept-level objective while preserving standard token-level autoregressive generation. NCP-ArchPreview builds a latent space by constructing a product-quantized concept vocabulary directly from its hidden states, and subsequently learns to predict future concepts via a dedicated Concept Module. These predicted concepts are then fed back to the token level to guide subsequent generation, with NTP and NCP…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

待翻譯:Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10702v1 Announce Type: new Abstract: Learning from limited text requires models to use context, generalize to new inputs, and retain useful capabilities. Qiushi Engine conducted a long-horizon, end-to-end autonomous research program on BabyLM 2026 Strict-Small, within 10 million corpus words and 100 million cumulative word presentations. Three stages connected frontier advancement, principle discovery, and principle-guided model improvement. Stage I combined compact restatements, budget reinvestment, and residual incremental learning to build a frontier model. Stage II found that exact repetition and aligned restatement produce different patterns of context use, depending on target relations and prediction windows. In controlled tasks, recovering fam…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement

待翻譯:A Bellman Optimality Equation for Plasticity

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10776v1 Announce Type: new Abstract: In continual reinforcement learning, carefully managing the stability-plasticity tradeoff remains a core challenge. Recent work by Abel et al. (2025) formalized this dilemma by defining plasticity as the generalized directed information from an agent's observations to its actions, and empowerment as the generalized directed information from its actions to its observations. This formulation successfully reframes the traditional stability-plasticity tradeoff as an empowerment-plasticity tradeoff. However, while extensive literature exists on optimizing for empowerment, there is currently no research addressing the optimization of plasticity under this new definition. This paper presents preliminary work toward optim…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:A Bellman Optimality Equation for Plasticity

待翻譯:The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10739v1 Announce Type: new Abstract: A truth probe fitted where truthful reporting and a task's prescribed action coincide cannot distinguish those targets from its fitting labels alone. We call this failure of semantic identification perfect aliasing. In a controlled binary reporting game, truth and prescribed-action probes fitted on compliant contexts solve the same optimization. On rival contexts their labels are complements, forcing their AUROCs to sum to one; this identity holds across 751 cell-layer pairs to floating-point precision. We separate prescribed output symbols from semantic action using randomized codebooks, then separate truth from prescribed action by fitting on mixed compliant and rival contexts. For a reward-trained Gemma-2-9B po…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes

待翻譯:GEOSTEER: Geodesic Optimization for Activation Steering in Large Language Models

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10658v1 Announce Type: new Abstract: Activation steering provides a lightweight way to control large language models (LLMs) by modifying their hidden activations at inference time. Among these approaches, norm-preserving steering aims to change model behavior without altering the activation norm, reducing the risk of representation collapse and degradation. However, existing norm-preserving methods are limited by predefined steering trajectories and by their reliance on one-step updates, which may fail to capture the complex structure of activation distributions. We propose GeoSteer, an optimization-based method for norm-preserving activation steering. GeoSteer formulates steering as a Riemannian optimization problem and updates activations through a…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:GEOSTEER: Geodesic Optimization for Activation Steering in Large Language Models

待翻譯:Zero-shot rib design: merging training-free generative prior with topology optimization

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10643v1 Announce Type: new Abstract: Natural load-bearing patterns such as leaf venation, trabecular bone, and spider webs achieve high stiffness per unit mass, yet classical topology optimizers rarely reach such geometries, and few let engineers express structural design intent through natural language. This work treats a frozen text-to-image diffusion model as a training-free source of design knowledge and distills it into the physics loop of density-based topology optimization via score distillation sampling, so that a text prompt becomes an explicit, machine-interpretable representation of engineer intent. The prompt-induced generative gradient and the finite element sensitivity are combined at every iteration, letting physics decide which prompt…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:Zero-shot rib design: merging training-free generative prior with topology optimization

待翻譯:Halo: Improving forecast accuracy through heteroscedastic estimation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10589v1 Announce Type: new Abstract: Heteroscedastic forecasting, where a network estimates a scale parameter alongside a location parameter, is normally motivated by uncertainty quantification. This paper shows it also improves the point estimate, in contrast to reported negative results for heteroscedastic estimation outside time series. Halo is a modification that reuses an existing deep forecaster's architecture, giving it a second output for the scale of its implied distribution and training it under the matching negative log likelihood. Adapting three state-of-the-art models --- a transformer, a graph network paired with a variational autoencoder, and a single-layer convolutional network --- under both Gaussian and Laplacian losses demonstrates…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:Halo: Improving forecast accuracy through heteroscedastic estimation

待翻譯:M3-Former: Multimodal Transformer with Mixture-of-Experts for Long-Term Vessel Trajectory Prediction

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10559v1 Announce Type: new Abstract: To address the challenges of behavioral multimodality, limited semantic utilization, and long-term error accumulation in vessel trajectory prediction, this paper proposes M3-Former, a multimodal trajectory prediction framework enhanced by large language models (LLMs). The proposed framework incorporates vessel static attributes and navigational intent as semantic priors for long-term trajectory modeling. Specifically, a unified multimodal representation space is constructed, in which static semantic information is encoded by a pre-trained LLM and aligned with dynamic trajectory features through self-attention. To jointly capture global route planning and local motion variations, a dual-granularity Mixture-of-Exper…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:M3-Former: Multimodal Transformer with Mixture-of-Experts for Long-Term Vessel Trajectory Prediction

待翻譯:XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09428v1 Announce Type: new Abstract: Evaluating the quality of explanations produced by explainable AI (XAI) methods remains challenging because existing approaches often rely on subjective human judgment, limiting reproducibility, scalability, and comparability between studies. We examine whether LLMs can serve as a reproducible and scalable mechanism to make comparative assessments of the quality of XAI explanations. We introduce XAI-Arena, an LLM-as-a-judge framework for scalable, reproducible, multidimensional, and stakeholder-sensitive evaluation of XAI explanation quality. XAI-Arena then allows us to compare XAI explanations along various dimensions, namely, perceived simplicity, clarity, task adequacy, trust calibration, actionability, transpa…

arXiv AI來源內容 · 翻譯待補全待翻譯:XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?

待翻譯:Together AI expands fine-tuning service with more models, live metrics, and finer controls

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Together Fine-Tuning adds the latest open-weight models, live experiment tracking, Expert LoRA, early stopping, tokenized dataset previews, pre-flight validation, and lower training prices on selected models.

Together AI Blog來源內容 · 翻譯待補全待翻譯:Together AI expands fine-tuning service with more models, live metrics, and finer controls

待翻譯:SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Proteins are fundamental to biological processes, with their function determined by the complex interplay between the amino acid sequence and the three-dimensional structure. Developing generative models capable of understanding this intrinsically multi-modal relationship is crucial for fields like drug discovery and protein engineering. Existing models often rely on a multi-stage training process where autoencoders that tokenize data into latent representations are trained in a first stage. Secondly, a generative model is trained on the latent representation of the autoencoder(s), i.e…

Apple Machine Learning Research來源內容 · 翻譯待補全待翻譯:SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign

待翻譯:Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth references. This paradigm suffers from the “one-to-many” nature of video description, where high-quality captions are often penalized for lexical mismatches or valid shifts in visual focus. Furthermore, such assessments are typically one-dimensional, failing to provide a fine-grained analysis of caption quality. To address this, we redefine caption quality via information fidelity: A caption must maximize the coverage…

Apple Machine Learning Research來源內容 · 翻譯待補全待翻譯:Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

待翻譯:DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Sign language processing systems have traditionally operated at the sentence level, ignoring critical discourse phenomena fundamental to sign language comprehension. We introduce DiscoSign, a computational approach for discourse-aware text to sign language gloss translation grounded in linguistic research. We address three key phenomena within our modular Large Language Model (LLM)-based translation framework: (i) spatial coreference resolution, where entities maintain consistent spatial locations throughout discourse; (ii) Question-Answer Clauses (QACs), pseudocleft structures serving…

Apple Machine Learning Research來源內容 · 翻譯待補全待翻譯:DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation

待翻譯:Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Production LLM applications rarely receive a question nobody has asked before. Support assistants and RAG pipelines field the same intents thousands of times a day, each phrased differently, and most stacks treat every phrasing as a fresh, fully billed request. Redis LangCache is a fully managed semantic caching service that sits between the application and […] The post Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster appeared first on MarkTechPost.

MarkTechPost來源內容 · 翻譯待補全待翻譯:Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster

待翻譯:Reduce LLM latency with prefix-aware routing on Amazon SageMaker Inference

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Amazon SageMaker Inference now offers prefix-aware routing, a routing strategy that sends requests sharing the same prompt prefix to the same instance so the KV cache stays warm. In benchmarks on Llama 3.1 70B, it reduced P50 time-to-first-token by up to 77% and raised KV cache hit rates from about 25% to over 80%.

AWS Machine Learning Blog來源內容 · 翻譯待補全待翻譯:Reduce LLM latency with prefix-aware routing on Amazon SageMaker Inference

待翻譯:Video and image search in Amazon Bedrock Knowledge Base using Marengo 3.0

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:TwelveLabs Marengo Embed 3.0 is now generally available as an embedding model in Amazon Bedrock Knowledge Bases, bringing fully managed natural language search to video, image, and audio content. This walkthrough shows how to build a knowledge base powered by Marengo 3.0 and run semantic queries against your media.

AWS Machine Learning Blog來源內容 · 翻譯待補全待翻譯:Video and image search in Amazon Bedrock Knowledge Base using Marengo 3.0

待翻譯:GPT Images 2.5 promises edits that leave the rest of your image alone

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:When OpenAI launched GPT Images 2.5 this week, the company promised better results for a common editing task: changing one The post GPT Images 2.5 promises edits that leave the rest of your image alone appeared first on The New Stack.

The New Stack AI來源內容 · 翻譯待補全待翻譯:GPT Images 2.5 promises edits that leave the rest of your image alone

待翻譯:Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:This week, Mistral announced it raised €3 billion in a Series D funding round, pushing its post-money valuation past €21 The post Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it. appeared first on The New Stack.

The New Stack AI來源內容 · 翻譯待補全待翻譯:Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

待翻譯:Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Manufacturing floors, warehouses and production lines rarely stay fixed — tasks change, layouts shift and new products arrive, and most robots can’t keep up without significant reprogramming. Skild AI’s new S1 robot foundation model helps address this, designed to learn previously unseen, long-horizon tasks from a single video demonstration. The model, launched last week, uses […]

NVIDIA Blog來源內容 · 翻譯待補全待翻譯:Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video

待翻譯:Model-agnostic PII detection with LLMs

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:A configurable, model-agnostic detector that turns any large language model on Amazon Bedrock into a PII detector. Because the entities to detect live in a prompt rather than in code, one detector adapts to new entity types without retraining, and it outperforms an off-the-shelf tool across five public corpora and nine LLM-based detectors.

AWS Machine Learning Blog來源內容 · 翻譯待補全待翻譯:Model-agnostic PII detection with LLMs

待翻譯:Introducing North Small Translate: A leading sovereign open-weight machine translation model

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Today, we're releasing North Small Translate, a mixture-of-experts machine translation model with strong performance across 50+ languages. Across WMT26 benchmarks,¹ North Small Translate achieves an 83.6 score across al…

Cohere Blog來源內容 · 翻譯待補全待翻譯:Introducing North Small Translate: A leading sovereign open-weight machine translation model
工具

待翻譯:Meta says it’s changing AI suggestions after posing invasive personal questions

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Meta says it's making changes to the prompts suggested by its AI chatbot after a viral video showed it digging for personal information about a woman's young daughters, as reported earlier by Futurism. In a statement to The Verge, Meta spokesperson Dina El-Kassaby says the company "missed the mark," adding that "the feature never should have prompted the individual with questions like that." Last week, Instagram user Kalie Robins posted a video explaining how Meta AI presented her with an invasive suggestion after cross-posting a clip to Facebook. The AI prompt, "Who is the child passenger?" appeared beneath a video of her and her child sin … Read the full story at The Verge.

The Verge AI來源內容 · 翻譯待補全待翻譯:Meta says it’s changing AI suggestions after posing invasive personal questions

待翻譯:From Spaghetti Code to Clean Python: A Beginner’s Guide

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Learn how to refactor messy Python code into clean, maintainable functions.

KDnuggets來源內容 · 翻譯待補全待翻譯:From Spaghetti Code to Clean Python: A Beginner’s Guide

待翻譯:Digested week: royal hairlines – Prince George’s hair has such brio as he starts at Eton | Emma Brockes

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Plus, your toilet flush can give you norovirus, AI doom-mongering and a drag Ann Droid In a news cycle to make everyone older than gen Z feel ancient, we start the week anticipating the 25th anniversary of 9/11 at the end. Social media floods with poignant interviews and images from lower Manhattan that day, triggering memories for the rest of us of where, who and how young we were. (Ridiculously, I was having my eyebrows done in Hendon, north London, and came out to find a white van pulled over at a crazy angle to the curb, doors flung open, radio cranked up, and people stopped in the street to listen. I remember looking over the city and wondering if I’d seen jets, or explosions.) Continue reading...

The Guardian AI來源內容 · 翻譯待補全待翻譯:Digested week: royal hairlines – Prince George’s hair has such brio as he starts at Eton | Emma Brockes

待翻譯:5 Python Techniques for Efficient Resource Orchestration

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:This article explains 5 Python techniques for efficient resource orchestration and sticks to what's stable today, 3.11 and later for the core techniques, with one 3.14-specific tool called out explicitly as requiring that version

KDnuggets來源內容 · 翻譯待補全待翻譯:5 Python Techniques for Efficient Resource Orchestration

待翻譯:izzit

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:izzit

待翻譯:UK economy defies forecasts with surprise 0.4% growth in July

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Welcome boost for chancellor with unexpected rise put down to rapid growth of AI in services sector outweighing Iran war fallout The UK economy grew in July as the rapid growth of AI appeared to outweigh the economic damage from the Iran war, in a welcome boost for the John Healey before next month’s budget. Figures from the Office for National Statistics (ONS) showed a surprise 0.4% increase in gross domestic product (GDP), compared with 0.3% growth in June. City economists had forecast zero growth. Continue reading...

The Guardian AI來源內容 · 翻譯待補全待翻譯:UK economy defies forecasts with surprise 0.4% growth in July

待翻譯:claudebill

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:claudebill

待翻譯:Accordio

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:Accordio

待翻譯:Jackalope

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:Jackalope

待翻譯:Devin Voice

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:Devin Voice

待翻譯:Anthropic details bad actors’ efforts to misuse its AI for bioweapons

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Report comes two days after former employee quit claiming company’s models could cause human extinction by 2030 Criminals, state-sponsored groups, spyware vendors, scientists and propagandists have attempted to use Anthropic’s powerful artificial intelligence models to design missiles and bombs, create deadly pathogens and surveil dissidents, according to a threat intelligence report the company published on Thursday. “The cases we share here aren’t typical misuse, but rather examples of the most notable and novel threat activity we’ve identified to date,” Anthropic wrote in its 154-page report. “We’re publishing this work because we believe we have a responsibility to disclose malicious misuse of our services.” Continue reading...

The Guardian AI來源內容 · 翻譯待補全待翻譯:Anthropic details bad actors’ efforts to misuse its AI for bioweapons

待翻譯:Cognition's SWE-2

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:Cognition's SWE-2

待翻譯:Slack can now vibe-code interactive charts and reports inside chats

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:A new feature coming to Slack will allow you to build interactive reports, polls, dashboards, presentations, microsites, and other tools directly inside a chat. With Slackforce Surfaces, you can describe to Slackbot what you need, and it will use AI to gather information from relevant conversations and connected apps, like Google Drive or Salesforce, to create it. Once Slackbot creates a Surface, you can share it with colleagues and pin it to channels, allowing other people to view it, interact with it, and leave comments. In one example shared by Slack, a user asks Slackbot for help creating an arcade-themed visualization of AI token usage … Read the full story at The Verge.

The Verge AI來源內容 · 翻譯待補全待翻譯:Slack can now vibe-code interactive charts and reports inside chats

待翻譯:ABrush

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:ABrush

待翻譯:sizeless

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:sizeless

待翻譯:chat-recall

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:chat-recall

待翻譯:Google to Invest $15B in Finland’s AI Infrastructure

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The tech giant simultaneously revealed a nuclear power contract with Finnish operator Fortum, its first outside of the U.S.

AI Business來源內容 · 翻譯待補全待翻譯:Google to Invest $15B in Finland’s AI Infrastructure

待翻譯:Feature Engineering in Scikit-Learn: A KDnuggets Cheat Sheet

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Once feature engineering lives inside a Pipeline, each step is fitted on training data only, and the model is scored what it actually earned. And that is the idea behind this new cheat sheet.

KDnuggets來源內容 · 翻譯待補全待翻譯:Feature Engineering in Scikit-Learn: A KDnuggets Cheat Sheet

待翻譯:3 ways to prep for your next big race with Search

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Search can help runners get race-day ready with registration alerts, tailored training plans, and more.

Google AI Blog來源內容 · 翻譯待補全待翻譯:3 ways to prep for your next big race with Search

待翻譯:ElevenLabs and Universal Music Group enter strategic agreement

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Skip to content Log inSign up Contact salesLog in Sign up First-of-its-kind collaboration spans licensing and product development New platform, to be launched by ElevenLabs, will enable fans to create remixes, mashups a…

ElevenLabs Blog來源內容 · 翻譯待補全待翻譯:ElevenLabs and Universal Music Group enter strategic agreement
研究

待翻譯:Lifesaving Lincoln Laboratory device wins 2026 Excellence in Technology Transfer Award

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The handheld catheterization device AI-GUIDE, created by Lincoln Laboratory and Massachusetts General Hospital, promises improved health outcomes for injured service members and civilians.

MIT News AI來源內容 · 翻譯待補全待翻譯:Lifesaving Lincoln Laboratory device wins 2026 Excellence in Technology Transfer Award

待翻譯:Scientists just made quantum computer operations 1,000 times faster

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Researchers have found a way to perform certain quantum operations more than 1,000 times faster, cutting thousands of repeated control cycles down to just one. The advance could reduce errors and bring reliable, fault-tolerant quantum computers closer to reality.

ScienceDaily AI來源內容 · 翻譯待補全待翻譯:Scientists just made quantum computer operations 1,000 times faster

待翻譯:Anthropologic

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:Anthropologic

待翻譯:We must pause risky AI research while we still have the power to do so | Gaby Hinsliff

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Warnings of AI’s existential threat to humanity are piling up – it’s time to listen and take them deadly seriously Another day, another horseman of the apocalypse galloping over the horizon. Lately we have heard from so many AI doomers – tech whistleblowers popping up to warn that their work is probably going to kill us – that we’re becoming almost blase about it. Humanity wiped out within a decade? Well, only if another world war or the climate crisis doesn’t get us first. Since it’s never clear whether the tech threat is real, or just a twisted form of hype from an industry that drums up investment by making their products sound more powerful than they really are, most of us settle for trying not to think about it too hard. But something about the AI research…

The Guardian AI來源內容 · 翻譯待補全待翻譯:We must pause risky AI research while we still have the power to do so | Gaby Hinsliff

待翻譯:Planning along Differentiable Charts of Constraint Manifolds with General-Purpose IK Solvers

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10905v1 Announce Type: new Abstract: Planning trajectories for robot manipulators under kinematic equality constraints restricts feasible motions to a measure-zero submanifold of the configuration space, requiring special algorithmic treatment. A promising strategy is parametrizing the set of feasible configurations using analytic inverse kinematics (IK). Bespoke analytic IK functions can be written to be differentiable, a necessary property for gradient-based trajectory optimization. But the vast majority of IK functions are computed by automated meta-solvers like IKFast, and are difficult to modify for differentiability. We present a new approach for computing gradients of analytic IK parameterizations: we leverage the inverse function theorem to r…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:Planning along Differentiable Charts of Constraint Manifolds with General-Purpose IK Solvers

待翻譯:Expressive Robotic Pianist: Mastering Complex Piano Repertoire with Graph-Mimic and Musical Dynamics

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10844v1 Announce Type: new Abstract: Enabling robots to perform musical instruments with human-level expressivity represents a frontier in bridging the gap between mechanical precision and artistic interpretation. Despite advances in robotic dexterity, replicating the fluid finger transitions and nuanced dynamic control characteristic of human pianists remains a significant challenge. Through a reinforcement learning-based control framework, we demonstrate that a dexterous robotic hand can achieve high-fidelity performance across a diverse piano repertoire. Central to our approach is a graph-based optimization strategy that guides the robot to generate natural pre-press and key-press fingering strategies that closely resemble human movement patterns.…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:Expressive Robotic Pianist: Mastering Complex Piano Repertoire with Graph-Mimic and Musical Dynamics

待翻譯:Lie-Algebraic Bell Recurrences for Arbitrary-Order Twist Jets and Parallel-Mechanism Closure

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10748v1 Announce Type: new Abstract: This paper develops an arbitrary-order kinematic construction that links serial propagation, parallel-mechanism closure, and rigid-platform point fields within one dual screw framework. A cylindrical joint is retained as one native physical block, with revolute and prismatic joints obtained as special cases. For each fixed joint axis, ordinary Bell polynomials organize the derivatives of the exponential factor; across a chain, the noncommuting factors remain in their physical order. Initial-frame prefix and terminal-resolved covariant formulas then produce equivalent representations of the serial twist jet. For a parallel mechanism, repeated Leibniz differentiation, with joint-level derivatives organized by Bell p…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:Lie-Algebraic Bell Recurrences for Arbitrary-Order Twist Jets and Parallel-Mechanism Closure

待翻譯:How Much Velocity Does Off-Ball Space Value Need? A Broadcast-Viewport Benchmark

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10801v1 Announce Type: new Abstract: Velocity-aware pitch control is standard, but under a broadcast viewport half the players are off screen and on-screen velocities come from a drifting calibration. We ask at which layer of broadcast off-ball analysis velocity changes the answer. Inheriting our off-screen imputation protocol (three Metrica matches, 44 m viewport, block-bootstrap CIs), we score four velocity regimes -- none, viewport-legal observed, true-for-visible, true-for-all -- against a velocity-aware ground truth at three layers: imputation, the control surface, and team verdicts. Velocity is nearly useless for imputation (-0.2 pp against a 12--14 pp velocity-free surface MAE), first-order for the surface (-1.5 to -1.8 pp, 11--15% of that MAE…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:How Much Velocity Does Off-Ball Space Value Need? A Broadcast-Viewport Benchmark

待翻譯:Two-Parameter Flow Map Learning for Continuous-Time Diffeomorphic Image Registration

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10789v1 Announce Type: new Abstract: Diffeomorphic image registration is central to medical image analysis, enabling anatomically consistent alignment across subjects. Most learning-based diffeomorphic methods model autonomous ODEs(ordinary differential equations) by parameterizing a stationary velocity field and recovering deformations via scaling-and-squaring. While non-autonomous ODEs with time-dependent velocities increase expressiveness, existing approaches rely on numerical integration to implicitly enforce flow structure that entangles model expressiveness with discretization accuracy. We propose a framework to directly learn the continuous-time solution of a non-autonomous ODE formulated as a two-parameterflow map. By enforcing cocycle consis…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:Two-Parameter Flow Map Learning for Continuous-Time Diffeomorphic Image Registration

待翻譯:Shedding Light: A Benchmark for Evaluating Lighting Understanding in Generative Image Models

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10787v1 Announce Type: new Abstract: Accurate modelling of illumination is central to realistic image synthesis and scene understanding. Yet, there is little exploration into whether image generative models are good at this task or whether physical plausibility remains a key challenge for them. Clearly, significant progress has been made in realistic image synthesis, but do models truly understand lighting in a physically accurate manner? To answer this question, this work proposes a benchmark to assess the lighting understanding and harmonisation capabilities of generative models. Our key insight is that evaluating lighting understanding for such models only requires testing how well they insert novel objects into real photographs whilst maintaining…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:Shedding Light: A Benchmark for Evaluating Lighting Understanding in Generative Image Models

待翻譯:Rethinking Handwritten Character Recognition

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10572v1 Announce Type: new Abstract: Non-Latin handwritten character recognition (HCR) remains understudied. Dominant methods consider it as generic image classification, which uses model scale to implicitly learn stroke structure. Structural-prior efficiency---the principle that explicitly encoding script-geometric regularities as architectural inductive biases can be both more accurate and require fewer parameters. We introduce GraphemeNet, a unified multi-script architecture, governed by two orthogonal binary axes. Axis 1 operationalises stroke-level geometric regularity via Persistent Scaffold Injection (PSI): a script-specific asymmetric convolution injects a stroke scaffold as a weighted residual at every encoder stage, continuously anchoring l…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:Rethinking Handwritten Character Recognition

待翻譯:CMNIE: An Information Extraction Benchmark for Chinese Military News

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10722v1 Announce Type: new Abstract: Structured extraction from Chinese military news supports intelligence analysis, decision-making, and knowledge base construction. However, existing resources provide limited support for joint informa?tion extraction in this domain, especially when events, event arguments, entities, and relations must be modeled together. We present CMNIE, an information extraction benchmark for Chinese military news. Extend?ing military-domain resources beyond document-level event annotations, CMNIE jointly annotates event triggers, event arguments, named enti?ties, and entity relations under a unified domain schema. The dataset contains 13,000 instances collected from public Chinese military news, with manual annotations for 7 e…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:CMNIE: An Information Extraction Benchmark for Chinese Military News

待翻譯:Adaptive Margin Ordinal Loss: Penalizing Center-Class Hedging in Ordinal Classification

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10752v1 Announce Type: new Abstract: Standard cross-entropy loss causes neural networks trained on ordinal classification tasks to hedge predictions toward center classes, a failure mode we term \emph{center-class hedging}. This occurs because predicting the middle class minimizes expected symmetric loss, making it the path of least resistance regardless of the true label. Existing ordinal losses address related problems such as large-error penalization and rank consistency, but none directly suppresses center-class hedging as a function of where the true label lies relative to the ordinal center. We propose the Adaptive Margin Ordinal Loss (AMOL), a multiplicative weight applied to per-class loss terms of the form $m(k,y) = 1 + \alpha \cdot (1 - |k-…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:Adaptive Margin Ordinal Loss: Penalizing Center-Class Hedging in Ordinal Classification

待翻譯:Conformal Calibration Transfer

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10737v1 Announce Type: new Abstract: Conformal prediction converts point predictions into set-valued predictions with coverage guarantees under exchangeability between calibration and deployment data. We study conformal calibration transfer, where this requirement fails because labeled calibration is available only in a source space, while prediction sets are needed in a target space linked to the source through unlabeled paired observations (e.g., paired modalities or sensor changes). We propose Transported Conformal Calibration (TCC): we transport labeled source calibration into the target space using the paired data, and then correct residual post-transport mismatch using only unlabeled target inputs. We instantiate this correction with two comple…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:Conformal Calibration Transfer

待翻譯:Artificial Intelligence Algorithms for the Detection of Pathologies Related to Lung Cancer through Image Analysis using Convolutional Neural Networks and Data Augmentation: a systematic mapping of the literature

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10652v1 Announce Type: new Abstract: Lung cancer is one of the leading causes of death worldwide, and its early diagnosis is crucial to improving patients prognosis and quality of life. However, the process of interpreting medical images for the detection of lung cancer is complex and requires trained experts. In this context, artificial intelligence (AI) and deep learning (DL) emerge as potential tools to automate and optimize image analysis. The objective of this work is to review the most recent and relevant applications of AI and DL in the field of radiology for the detection of lung cancer. To this end, an exhaustive search was carried out in scientific databases such as PubMed,IEEEXPLORE, Scopus and Web of Science, and 96 articles published fro…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:Artificial Intelligence Algorithms for the Detection of Pathologies Related to Lung Cancer through Image Analysis using Convolutional Neural Networks and Data Augmentation: a systematic mapping of the literature

待翻譯:Byzantine-Robust Federated Fire Detection with a Rotating Coordinator

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10647v1 Announce Type: new Abstract: We study the application of federated learning (FL) to indoor fire detection. Such fire-detection systems use edge cameras that record sensitive footage which cannot easily be collected at a central server. Existing federated solutions leave three practical obstacles unaddressed: limited uplink bandwidth, Byzantine (malicious or faulty) clients, and unconditional trust in a single, permanently fixed aggregation server. Our main contributions address all three. In particular, we provide (i) a curated indoor fire-detection dataset assembled from eight public sources; (ii) an edge-deployable detector whose model updates are compressed up to 10 time with only a small loss in balanced accuracy; and (iii) a semi-decentr…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:Byzantine-Robust Federated Fire Detection with a Rotating Coordinator

待翻譯:Gradland: On Phenomenal Experience, Differentiated Across Many Dimensions

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09306v1 Announce Type: new Abstract: This paper investigates the hypothesis that the first-order structure of physical interactions, i.e. gradients or Jacobians, characterizes the structure of phenomenal experience. It does so in an idealized world inhabited by neural networks, Gradland, where the physics are known and the functions are (mostly) differentiable. The paper introduces two measures of Jacobian structure: effective rank and cohesion, based on Kirchhoff complexity. Applying the measures to a series of worked examples shows the hypothesis accounts for: (1) the duration of experience, that it can prolong over hundreds of milliseconds; (2) the difference between what is experienced vividly and obscurely; (3) the experience of texture; (4) the…

arXiv AI來源內容 · 翻譯待補全待翻譯:Gradland: On Phenomenal Experience, Differentiated Across Many Dimensions

待翻譯:Datasette 1.0a39 and 0.65.4 security releases

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯: Datasette 1.0a39 and 0.65.4 security releases Today we're releasing two new security patch versions of Datasette: 1.0a39 and 0.65.4 - one for the current alpha series and one for the stable 0.65.x family. These are security fixes which you should apply if you are running a Datasette instance on the public web - in particular if that instance mixes both public and private tables. Following issues reported by Sevban Dönmez, Alex Garcia and I ran an extensive audit of Datasette using Claude Fable 5.1, GPT-5.6, and GPT-6 Astra. We then spent almost a week collaborating on and reviewing the fixes. They helped find some very subtle bugs. We'll be incorporating security audits by frontier models into all of our development work going forward. Alex came up with a way…

Simon Willison's Weblog來源內容 · 翻譯待補全待翻譯:Datasette 1.0a39 and 0.65.4 security releases

待翻譯:datasette-publish-fly 1.4

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯: Release: datasette-publish-fly 1.4 Sets force_https=true in fly.toml. #31 Fix for Volume could not be found bug. #32 Compatible with app-scoped deploy tokens. #34 Tags: datasette, fly

Simon Willison's Weblog來源內容 · 翻譯待補全待翻譯:datasette-publish-fly 1.4

待翻譯:github-to-sqlite 2.9.1

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯: Release: github-to-sqlite 2.9.1 Fix for compatibility with sqlite-utils 4.x. #85 Tags: github, sqlite

Simon Willison's Weblog來源內容 · 翻譯待補全待翻譯:github-to-sqlite 2.9.1

待翻譯:datasette 0.65.4

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯: Release: datasette 0.65.4 See Datasette 1.0a39 and 0.65.4 security releases on the Datasette blog. Tags: security, datasette

Simon Willison's Weblog來源內容 · 翻譯待補全待翻譯:datasette 0.65.4

待翻譯:Reduce inference cold starts on Amazon SageMaker HyperPod with model caching

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Amazon SageMaker HyperPod now supports model caching for inference, which pre-loads model weights and container images onto cluster nodes so pods read from local NVMe storage instead of downloading over the network. Learn how model caching cuts cold starts from tens of minutes to seconds, how it works, and how to enable it.

AWS Machine Learning Blog來源內容 · 翻譯待補全待翻譯:Reduce inference cold starts on Amazon SageMaker HyperPod with model caching

待翻譯:Schools are catching onto Big Tech’s playbook

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:It's the hot new thing in tech, and it's where all the jobs are. Students who don't learn to use it fall behind. And to help them catch up in time, its creators are graciously providing the resources and curriculum for learning it, often pro bono. That's the narrative AI companies are pitching schools on right now, but according to New York Times education technology reporter Natasha Singer, it's also how the tech industry has built its influence in classrooms for the past 15 years. In a new book, Coding Kids, Singer details how tech companies took a central role in promoting computers and coding skills in education, promising computer sci … Read the full story at The Verge.

The Verge AI來源內容 · 翻譯待補全待翻譯:Schools are catching onto Big Tech’s playbook
機器人

待翻譯:The Sequence Opinion - Issue 931: Robotics Is Waiting for Its ChatGPT Moment

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Generalist models are expanding what robots can do. The real breakthrough will be how easily we can teach them something new.

TheSequence來源內容 · 翻譯待補全待翻譯:The Sequence Opinion - Issue 931: Robotics Is Waiting for Its ChatGPT Moment
晶片

待翻譯:Testing Between the Test Cases: Proving End-to-End Steering in Conditions You Never Drove

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10951v1 Announce Type: new Abstract: AI-based automated vehicle testing is challenging because a model that passes every test condition can still fail in the real world. Formal verification offers a way to directly address this gap. On a simulated highway and an arterial road we trained two small end-to-end steering networks each in CARLA, one on clear conditions alone and one on clear, fog, night and low sun. All four models were driven against a 2.19 ft lane-departure budget. Without driving again, we used bound propagation, a formal method that reads the trained weights, to compute how far steering can drift at every disturbance strength between two captured images. One calculation covers more than a campaign could drive: on the arterial it spans…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:Testing Between the Test Cases: Proving End-to-End Steering in Conditions You Never Drove

待翻譯:NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput and 58.5K Residues per GPU-Hour on 8xH100

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:NVIDIA has detailed BioNeMo Inference Runtime (BioIR), a Python library that accelerates biomolecular structure-prediction models on NVIDIA GPUs while staying in plain PyTorch. In a matched benchmark on 1,000 human dimer targets across 8xH100 GPUs, BioIR-accelerated Boltz-2 delivered 58.5K successfully folded residues per GPU-hour versus 20.2K for a torch-compiled open-source implementation, a 2.90x gain. The runtime optimizes at 3 layers: custom kernel selection, CUDA Graph capture, and Ray-based replica scaling that places 1 full model copy per GPU. BioIR already powered the AlphaFold Database expansion, generating about 31 million candidate protein complexes across 4,777 proteomes. The post NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz…

MarkTechPost來源內容 · 翻譯待補全待翻譯:NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput and 58.5K Residues per GPU-Hour on 8xH100

待翻譯:Red Hat AI 3.5 tackles the GPU queue that can stall AI pilots

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Red Hat released Red Hat AI 3.5 this week, a move designed to let software engineering teams run AI with The post Red Hat AI 3.5 tackles the GPU queue that can stall AI pilots appeared first on The New Stack.

The New Stack AI來源內容 · 翻譯待補全待翻譯:Red Hat AI 3.5 tackles the GPU queue that can stall AI pilots
創業融資

待翻譯:When Information is Worth the Risk: Behavioral Valuation for Hazardous Robotic Exploration

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.10726v1 Announce Type: new Abstract: Hazardous robotic exploration requires robots to map spatial risks, such as unsafe terrain, radiation, fire, mines, or structural damage, while operating where collecting information can itself cause failure. A highly informative path may expose the robot to hazards, terminate execution, and prevent future observations. Hazardous exploration therefore requires deciding not only where uncertainty is largest, but when reducing it is worth the risk. This paper introduces a valuation-layer view of this problem. We keep the belief update, sensor model, physical risk model, and finite-horizon informative planner fixed, and change only the scalar objective used to rank feasible paths. Within this framework, we introduce…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:When Information is Worth the Risk: Behavioral Valuation for Hazardous Robotic Exploration

待翻譯:More Anthropic researchers warn of AI’s perils as Musk terms fears a ‘psyop’

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Insiders at the firm fear tech’s advancement could cause human extinction while others are calling it a ‘setup’ A day after a former researcher at Anthropic made an apocalyptic declaration about artificial intelligence, more researchers and staff members at the AI startup publicly agreed with him and posted their own dire warnings. In response, Elon Musk and other conservative figures on X, formerly Twitter, called the chorus of concerns a “setup” and a “psyop”. Continue reading...

The Guardian AI來源內容 · 翻譯待補全待翻譯:More Anthropic researchers warn of AI’s perils as Musk terms fears a ‘psyop’

待翻譯:AI Coding Startup Cognition Now Valued at $48B

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Interest in automated generative AI coding has exploded along with the vendor’s growth.

AI Business來源內容 · 翻譯待補全待翻譯:AI Coding Startup Cognition Now Valued at $48B
政策

待翻譯:The AI Safety Crunch and How Enterprises Should Deal With It

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The competitive landscape, particularly the geopolitical clash between the U.S. and China, complicates the AI safety dilemma.

AI Business來源內容 · 翻譯待補全待翻譯:The AI Safety Crunch and How Enterprises Should Deal With It