跳到主要內容
AI News HubLIVE

本期報道已收集,譯文與分析尚待補全。可展開其餘更新查看來源內容。

其餘更新(128 條)
Agent

待翻譯:Agent Evaluation Metric for multi-turn conversations

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Multi-turn agents fail in ways single-turn evaluation misses: one early mistake corrupts every later turn. This post introduces the Agent Evaluation Metric (AEM), a decomposable, turn-level way to measure agent quality, applied to its first dimension, correctness, to pinpoint the turn that caused a failure and separate it from the turns that inherited it.

AWS Machine Learning Blog來源內容 · 翻譯待補全待翻譯:Agent Evaluation Metric for multi-turn conversations

待翻譯:How AvioBook builds turnaround insights from operational data with Amazon Bedrock AgentCore

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AvioBook, a Thales Group Company, prototyped Connected Analytics on Amazon Bedrock AgentCore to turn AvioBook Connect's operational data into plain-language, evidence-based answers for airline managers and dispatchers, helping them find and act on the causes of flight turnaround delays.

AWS Machine Learning Blog來源內容 · 翻譯待補全待翻譯:How AvioBook builds turnaround insights from operational data with Amazon Bedrock AgentCore

待翻譯:Now everyone can put data to work

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Meet the Data agent in ChatGPT Work. Connect company data, uncover insights, and build interactive dashboards with AI using natural language.

OpenAI News來源內容 · 翻譯待補全待翻譯:Now everyone can put data to work

待翻譯:Sure, Meta’s AI Muse works, but it sure creeps me out

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:I turned my Muse assistant into a purple cat. | Screenshot: The Verge Meta has launched its new Muse assistant, marking the company's first real foray into AI-powered productivity tools. The company says its AI agent can "take the busywork off your plate" by helping you with online shopping, emails, trip-planning, and more. I decided to try out the new tool and see how well it performed - especially from a company that previously prioritized entertainment over productivity. Though the AI assistant generally worked as I expected it to, the unnerving amount of information it autonomously gleaned about me largely overshadowed my experience. One of the first tasks I assigned Muse, which carries out actions using … Read the full story at The Verge.

The Verge AI來源內容 · 翻譯待補全待翻譯:Sure, Meta’s AI Muse works, but it sure creeps me out

待翻譯:7 Steps to Become a Forward Deployed Engineer in 2026

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:FDEs are becoming some of the most in-demand engineers in AI. Here’s the 7-step roadmap to becoming one in 2026.

KDnuggets來源內容 · 翻譯待補全待翻譯:7 Steps to Become a Forward Deployed Engineer in 2026

待翻譯:Why the current tech backlash feels different

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:This interview has been lightly edited for length and clarity. Nick Statt: Hello and welcome to Decoder, Nilay’s show about big ideas and other problems. This is Nick Statt, senior producer. And I’m joined by our brand-new supervising producer, Greg Ott. Greg Ott: Good day, everyone. And Hi, Nilay. Nilay is here too. He is the person who hosts the show, and his name is also in the show. So it makes sense. Nilay Patel: It’s true. We don’t consistently say my name in the show enough. We should do it all the time. I should do it. Welcome to Decoder with Nilay Patel. GO: Like Nick was saying, this is Nilay’s show and we are doing a mailbag episode. This is where we go through all the feedback. Because you listen to the end of every episode, we know you do. We do me…

The Verge AI來源內容 · 翻譯待補全待翻譯:Why the current tech backlash feels different

待翻譯:Improving Lakebase Postgres compute cache

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The disaggregated storage model of Lakebase Postgres provides a feature rich, flexible...

Databricks Blog來源內容 · 翻譯待補全待翻譯:Improving Lakebase Postgres compute cache

待翻譯:How Credit Genie keeps codebase docs fresh with OpenWiki

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:See how Credit Genie uses OpenWiki to automate repo documentation, reduce tribal knowledge, and give engineers and coding agents searchable codebase context.

LangChain Blog來源內容 · 翻譯待補全待翻譯:How Credit Genie keeps codebase docs fresh with OpenWiki

待翻譯:5 Useful Python Scripts to Automate CSV Processing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Automate common CSV tasks with these 5 Python scripts for cleaning, validating, transforming, and processing CSV files using the standard library.

KDnuggets來源內容 · 翻譯待補全待翻譯:5 Useful Python Scripts to Automate CSV Processing

待翻譯:How to Combine Traditional Machine Learning with Agentic Reasoning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:In this article, you will learn where traditional machine learning reaches its limits, what agentic reasoning adds, and how combining the two produces AI systems...

Machine Learning Mastery來源內容 · 翻譯待補全待翻譯:How to Combine Traditional Machine Learning with Agentic Reasoning

待翻譯:AI floods security teams with flaws — business context sets priorities

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:A security researcher testing a 300-person B2B company with a global footprint discovered an internet-exposed database with weak authentication during The post AI floods security teams with flaws — business context sets priorities appeared first on The New Stack.

The New Stack AI來源內容 · 翻譯待補全待翻譯:AI floods security teams with flaws — business context sets priorities

待翻譯:Mathematicians want proof OpenAI didn’t use their work

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Another researcher is challenging OpenAI about the data driving its increasingly impressive array of mathematical discoveries. Just days after a bitter row erupted over whether the company's models benefited from unpublished work, a second mathematician has come forward accusing the AI giant of unethical and "dishonest" behavior and a lack of transparency about the origins of its training data. In a series of posts on Mastodon, mathematician Andreas Thom raised concerns that interactions he and his colleagues had had with the ChatGPT chatbot before OpenAI's triumphant announcement may have contributed to its success in the field. One of the … Read the full story at The Verge.

The Verge AI來源內容 · 翻譯待補全待翻譯:Mathematicians want proof OpenAI didn’t use their work

待翻譯:Building a Reliable Foundation for Agentic AI in SMBs

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:This article is sponsored by Salesforce and was written, edited, and published in alignment with our Emerj sponsored content guidelines. Learn more about our thought leadership and content creation services on our Emerj Media Services page. Customer-facing organizations now face a widening capacity gap driven by escalating multi‑channel demand and the constraints of human-only workflows […]

Emerj AI Research來源內容 · 翻譯待補全待翻譯:Building a Reliable Foundation for Agentic AI in SMBs

待翻譯:Bodhan AI Releases Four Indic Models for OCR, Translation and Speech

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:A Hindi lesson can mix English terms (loan words), scanned tables and handwritten equations. Making that content searchable, translating it and reading it aloud requires several kinds of AI. Bodhan AI and AI4Bharat’s four new models target those jobs across Indian languages. Released in September 2026, the models cover document parsing, translation, speech recognition and speech generation, with support for mixed languages and scripts. In this […] The post Bodhan AI Releases Four Indic Models for OCR, Translation and Speech appeared first on Analytics Vidhya.

Analytics Vidhya來源內容 · 翻譯待補全待翻譯:Bodhan AI Releases Four Indic Models for OCR, Translation and Speech

待翻譯:Agentic AI-enabled Semantic Commissioning of a Cognitive Digital Twin for Reconfigurable Manufacturing

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09503v1 Announce Type: new Abstract: Rapid bespoke commissioning of the Cognitive Digital Twin (CDT) is a major challenge in reconfigurable manufacturing. Traditional digital twin (DT) construction methods primarily focus on geometric reconstruction, often neglecting the deep semantic integration and functional interoperability necessary for autonomous reasoning. This paper proposes an agent-based, AI-driven workflow to automate end-to-end CDT debugging. The system utilises LangGraph as a multi-agent orchestration engine to achieve dual-path synthesis: the semantic path extracts technical specifications from unstructured documents using Retrieval Augmented Generation (RAG), while the functional path autonomously discovers and binds to real-time indus…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:Agentic AI-enabled Semantic Commissioning of a Cognitive Digital Twin for Reconfigurable Manufacturing

待翻譯:AgenticGen: Reward-Guided Agentic Video Generation for Advertising

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09187v1 Announce Type: new Abstract: Advertising video generation is not only a video synthesis task, but also a product-conditioned reasoning problem whose success is measured by online business metrics. Recent video foundation models can generate realistic clips from multimodal conditions, yet they do not optimize how a product should be transformed into an effective advertisement or how future generation should be improved from online business feedback. To close this loop, we propose AgenticGen, a reward-guided agentic framework that decomposes advertising video generation into two trainable reasoning stages, strategy selection and draft generation, thereby exposing optimization targets that online business feedback can supervise. AgenticGen learn…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:AgenticGen: Reward-Guided Agentic Video Generation for Advertising

待翻譯:Endogenous Exploration in Reinforcement Learning with Intrinsic Curiosity

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05650v1 Announce Type: new Abstract: We propose a reinforcement learning framework in which exploration is driven by intrinsic curiosity, designed for scenarios where environments are non-stationary and rewards are sparse, delayed, uninformative, or absent. In our model, action selection is guided by a combination of external rewards and an epistemic motivation mechanism that biases the agent toward structured exploratory directions. The central hypothesis is that effective exploration emerges at intermediate levels of incoherence, while performance degrades under both overly rigid and overly disordered dynamics. To test this idea, we implement the framework on top of a Liquid State Machine (LSM) substrate and evaluate it on two standard benchmarks:…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:Endogenous Exploration in Reinforcement Learning with Intrinsic Curiosity

待翻譯:AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05435v1 Announce Type: new Abstract: Modern language agents are expected to operate over long horizons: they ask follow-up questions, reuse worked examples, handle tool feedback, and adapt to delayed consequences. Most evaluations still reset the agent after a prompt or score only the final state of one trajectory. AhaBench asks a more operational question: when a fixed model receives useful experience, does its later behavior improve under a related evaluation condition where the obvious support has been removed, changed, or delayed? The suite contains three components. Aha-Puzzle tests no-hint exploration after solved hidden-state puzzles; Aha-Euler turns Project-Euler-style mathematical ideas into generated taught/held-out tasks with exact validat…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning

待翻譯:Compiling VGDL into Causal Models

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05459v1 Announce Type: new Abstract: Reinforcement learning and large language models often struggle to accurately capture the causal mechanics of game environments. Standard reinforcement learning agents tend to rely on spurious correlations, while large language models are prone to hallucinating game rules. Although causal reinforcement learning improves interpretability, there is currently no formal methodology to map complex game mechanics directly into causal models. To address this, we propose a deterministic framework that compiles games specified in the Video Game Description Language into Dynamic Structural Causal Models. Rather than inferring causal structures from gameplay traces or noisy large language models' outputs, our methodology dir…

arXiv AI來源內容 · 翻譯待補全待翻譯:Compiling VGDL into Causal Models

待翻譯:AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05446v1 Announce Type: new Abstract: We introduce AutoFyn, an agent harness inspired by the Expert Iteration algorithm, adapting a frozen model across many rounds by updating persistent state from verified reward signals rather than model weights. Each round begins from a fresh model session, and durable information is reintroduced only through explicit interfaces such as persistent memory files, reports, and repository state. Within a round, an orchestrator explores, plans and builds many alternative approaches with specialized agents, while a task-grounded verifier verifies the work and supplies an objective reward for measuring progress. This reward is distilled back into the persistent state, which updates the effective policy for the next round.…

arXiv AI來源內容 · 翻譯待補全待翻譯:AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents

待翻譯:When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05441v1 Announce Type: new Abstract: Long-term memory for LLM agents is evaluated today by conversational recall benchmarks (LoCoMo, LongMemEval), which measure question answering over dialogue history, not whether remembered facts change what a tool-using agent does. We present MERIT (Memory Evaluation for Realistic Instrumented Tasks), a benchmark and harness that measures the marginal utility of memory for task-executing agents under explicit cost accounting. MERIT provides episodic tool-use tasks in three domains whose dependence on earlier-episode facts is verified by an automated leak check; a difficulty ladder ending in updated-fact recall; controlled memory corruption; and full token and dollar metering of every memory operation. Across 23,44…

arXiv AI來源內容 · 翻譯待補全待翻譯:When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents

待翻譯:[AINews] not much happened today

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:a quiet day

Latent Space來源內容 · 翻譯待補全待翻譯:[AINews] not much happened today

待翻譯:LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:LandingAI has shipped Agentic Document Extraction Gen2, a rebuild of its document stack on the DPT-3 model family. Chunks are retired in favor of a document, page and block tree. DPT-3 Pro grounds to the line, DPT-3 Verity grounds to the word with a confidence score, and Parse billing now counts output characters instead of flat pages. Gen1 code will not run against Gen2 endpoints. The post LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity appeared first on MarkTechPost.

MarkTechPost來源內容 · 翻譯待補全待翻譯:LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

待翻譯:Introducing the Agents API

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Build and launch cloud agents with the Agents API, a managed service powered by the Codex harness for orchestration, long-running sessions, and tool use.

OpenAI News來源內容 · 翻譯待補全待翻譯:Introducing the Agents API

待翻譯:.blend URL Viewer

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯: Tool: .blend URL Viewer I'm continuing to have a lot of fun with GPT-6 Astra and Blender (see my TIL). As a big fan of the Imperial Fabergé Easter eggs, I've always thought it would be fun to make some new ones that celebrate popular culture. Yesterday I decided to try out the new ChatGPT Images 2.5 by running this prompt: Generate a photo of a faberge egg that's themed after the TV show Pluribus - research first It gave me this - honestly not bad for a first attempt! Then, just to see what would happen, I pasted that image into Codex running GPT-6 Astra (high) and prompted: Use your blender local skill to create a blender model of this faverge egg (Here's the skill file, which I created like this.) It churned away for 17m51s and built me several .blend files.…

Simon Willison's Weblog來源內容 · 翻譯待補全待翻譯:.blend URL Viewer

待翻譯:Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Google has open-sourced Mantis, a stack-agnostic toolkit of security review skills for AI coding agents. It runs the full vulnerability lifecycle: sweep the code, filter false positives, reproduce the bug in a sandbox, patch it, re-attack the patch, then score the risk. Apache 2.0, and documented as demonstration-only. The post Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities appeared first on MarkTechPost.

MarkTechPost來源內容 · 翻譯待補全待翻譯:Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities

待翻譯:OpenAI’s sly mathematical breakthrough sends a chill through academia

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Open AI CEO Sam Altman speaks during the G20 Innovation Ministerial. | (Photo by Matt RAMEY / AFP via Getty Images) OpenAI's announcement Tuesday that it has solved one of mathematics' legendary Millennium Prize problems should have been a moment of triumph. The result is both an undeniable achievement and a striking demonstration of just how rapidly AI is transforming mathematics. But before it was even formally announced, the breakthrough had been complicated by the unusual circumstances that prompted OpenAI to pursue the problem: After hearing other researchers were making progress, it seems to have thrown its considerable resources into a last-minute effort to beat them to the punch. The ensuing controversy has surfaced allegations of scooping, spying … Rea…

The Verge AI來源內容 · 翻譯待補全待翻譯:OpenAI’s sly mathematical breakthrough sends a chill through academia

待翻譯:Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests.

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AI models now power all manner of agents, from coding assistants that write and debug software to customer service systems The post Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests. appeared first on The New Stack.

The New Stack AI來源內容 · 翻譯待補全待翻譯:Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests.

待翻譯:hob

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:hob

待翻譯:ICYMI: What landed for AI builders in August 2026

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:A recap of August 2026 launches for AI builders across Amazon Bedrock, Amazon Bedrock AgentCore, and Strands: million-token context for OpenAI models, cross-Region inference, agents that run for up to 14 days on dedicated compute, expanded AWS GovCloud availability, and Strands Robots for physical deployment.

AWS Machine Learning Blog來源內容 · 翻譯待補全待翻譯:ICYMI: What landed for AI builders in August 2026

待翻譯:How Heurist Finance built an AI-native investment workbench on Amazon Bedrock AgentCore

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Learn how Heurist built Heurist Finance, a conversational AI investment workbench, on Amazon Bedrock AgentCore. This customer story shows how AgentCore payments, Identity, Memory, Code Interpreter, and Observability let a small team buy premium market data per query, isolate analysis in a sandbox, and keep every action auditable.

AWS Machine Learning Blog來源內容 · 翻譯待補全待翻譯:How Heurist Finance built an AI-native investment workbench on Amazon Bedrock AgentCore

待翻譯:Five AI Questions We're Hearing from Financial Services Leaders

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Last year at Sibos Frankfurt, the question was whether AI works. This year: can your...

Databricks Blog來源內容 · 翻譯待補全待翻譯:Five AI Questions We're Hearing from Financial Services Leaders

待翻譯:Connections: managed credentials and per-caller identity for Managed Deep Agents

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Learn how Connections in Managed Deep Agents securely manage credentials, support per-user OAuth, and let agents act with each caller’s identity.

LangChain Blog來源內容 · 翻譯待補全待翻譯:Connections: managed credentials and per-caller identity for Managed Deep Agents

待翻譯:NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:At the IBC conference, running Sept. 11-14 in Amsterdam, the creative, technology and business communities are coming together to turn ideas into action and discuss innovations across the media and entertainment industries. More than 44,000 attendees from 170+ countries are gathering to explore 1,300+ exhibitions in 14+ halls and outdoor spaces, with over 600 speakers […]

NVIDIA Blog來源內容 · 翻譯待補全待翻譯:NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC

待翻譯:Recreating a 70-year love story frame by frame

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discover how filmmakers and Google DeepMind used AI to recreate a couple's unrecorded past in the short film "Love, Rendered."

Google AI Blog來源內容 · 翻譯待補全待翻譯:Recreating a 70-year love story frame by frame

待翻譯:Training API now generally available | Fireworks

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Training API now generally available | Fireworks Join us for our inaugural conference, Forge 2026 Blog Train Past The Frontier Training API Now Generally Available Train past the frontier: Training API now generally ava…

Fireworks AI Blog來源內容 · 翻譯待補全待翻譯:Training API now generally available | Fireworks

待翻譯:Own the Outer Loop

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The following article originally appeared on Addy Osmani’s blog and is being republished here with the author’s permission. In the past year, the conversation around agentic engineering has moved to harnesses and loops, fleets and software factories. My 2 cents is engineers need to own the outer loop—the accountability for these systems. This only gets […]

O'Reilly AI & ML Radar來源內容 · 翻譯待補全待翻譯:Own the Outer Loop
工具

待翻譯:Universal Music is launching an AI music platform with ElevenLabs

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Universal Music Group is launching a new AI-powered platform that will allow users to draw from its catalog of licensed music to create song remixes, mashups, and new takes on tracks, according to an announcement on Thursday. The record label is developing the platform through a multi-year licensing agreement with ElevenLabs, a company that specializes in AI voice and music generation. Artists can choose whether to participate in UMG and ElevenLabs' upcoming platform, which marks yet another AI deal for the record label. UMG is currently developing an AI music platform with Udio and has struck AI licensing deals with Spotify, Nvidia, and Kl … Read the full story at The Verge.

The Verge AI來源內容 · 翻譯待補全待翻譯:Universal Music is launching an AI music platform with ElevenLabs

待翻譯:ChatHop

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:ChatHop

待翻譯:d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AI inference chipmaker d-Matrix today announced it will use NVIDIA NVLink Fusion to connect its next-generation Raptor XPUs to NVIDIA’s AI infrastructure platform — joining a growing roster of ecosystem partners. By connecting Raptor to NVIDIA NVLink scale-up and Spectrum-X scale-out networking, the NVIDIA MGX rack architecture and the broader NVIDIA AI platform, NVLink Fusion […]

NVIDIA Blog來源內容 · 翻譯待補全待翻譯:d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment

待翻譯:Boots on the Ground: ‘WARDOGS’ Goes All Out on GeForce NOW at Early-Access Launch

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Gear up: The latest PC games and major updates are ready to play on GeForce NOW this week. WARDOGS drops onto the cloud at early-access launch, alongside the Valheim 1.0 Deep North update and Bus Simulator 27 — part of nine new titles joining the cloud. The newest PC releases can demand serious hardware, storage […]

NVIDIA Blog來源內容 · 翻譯待補全待翻譯:Boots on the Ground: ‘WARDOGS’ Goes All Out on GeForce NOW at Early-Access Launch

待翻譯:OpenAI not on track to reduce risk of ‘catastrophic’ loss of control, says board member

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:US government adviser Paul Christiano warns of risks to AI industry as he joins OpenAI’s non-profit foundation OpenAI is not on track to reduce risks of “catastrophic” loss of control to an acceptable level, a member of its non-profit board has warned, amid spreading public and political concern that super-advanced AIs could one day wipe out humanity. Paul Christiano, a US government technology adviser, said “there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term.” Continue reading...

The Guardian AI來源內容 · 翻譯待補全待翻譯:OpenAI not on track to reduce risk of ‘catastrophic’ loss of control, says board member

待翻譯:Driverless cars are taking us on a road to nowhere | Adrian Chiles

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Some inventions, like the dishwasher, improve our lives. But driverless cars will lobotomise us I’ve found a hill to die on. Driverless cars. Why? Really, why? If driverless cars are the answer, then what is the question? Who asked for them? Oh yes, they’re fine for the novelty value, for the lols, for the clicks. We probably know someone who’s been to California or China and sent home a video of their driverless journey. Fine fun for feeble minds, if you ask me. And feeble-minded is how we’re all going to end up if this madness takes hold. It’s been a slippery slope since the invention of automatic transmission, rendering our left legs and arms redundant and sparing us the trouble of feeling our way up and down through the gears. Then there was the coming of s…

The Guardian AI來源內容 · 翻譯待補全待翻譯:Driverless cars are taking us on a road to nowhere | Adrian Chiles

待翻譯:Neopress

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:Neopress

待翻譯:Visiby

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:Visiby

待翻譯:Suno v6

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:Suno v6

待翻譯:Modeinspect

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:Modeinspect

待翻譯:Anomalo

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:Anomalo

待翻譯:Suno releases its first AI music model made with record industry help

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Suno's new v6 AI music model is its first made with support from the record industry. Suno's Jack Brody told The Verge that v6 was "trained from the ground up, with a new set of data that does not include the same data that our previous models were trained on." The data includes content licensed from partners Warner Music Group, BMG, and Believe, as well as "user data." It's unclear whether that means v6 training data is completely free of dubiously obtained content. One of the big changes in v6 is that there are actually three different models: v6, v6-wild, and v6-mini. Mini is the model available for free to all. It's focused on fast, res … Read the full story at The Verge.

The Verge AI來源內容 · 翻譯待補全待翻譯:Suno releases its first AI music model made with record industry help

待翻譯:Hyrax AI

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:Hyrax AI

待翻譯:Desert Ant Labs

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:Desert Ant Labs

待翻譯:Apple’s new iPhone camera mode promises to prove your photo isn’t AI

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Apple is launching a new way to prove that the picture you took isn't manipulated by AI. A new feature, called "Reference Image," will arrive with the iPhone 18 Pro lineup later this month and is supposed to use the device's new camera sensor to "sign every pixel it sees." The iPhone 18 Pro and Pro Max will only authenticate photos when placed into Reference mode. Once the camera captures signed sensor data, Apple says its Private Cloud Compute will develop it "into an unalterable reference image" that can be viewed in the Photos app. You'll be able to compare the reference image with other versions of the photo to see if any edits were mad … Read the full story at The Verge.

The Verge AI來源內容 · 翻譯待補全待翻譯:Apple’s new iPhone camera mode promises to prove your photo isn’t AI

待翻譯:Apple unveils folding iPhone Duo as new CEO takes center stage

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The passport-size iPhone Duo serves as the first big test for John Ternus, who took over for Tim Cook last week Apple has unveiled the iPhone Duo, the company’s first foldable phone and the most significant change to the iPhone since the smartphone’s introduction in 2007. The nearly $2,000 device takes inspiration from iPads’ wide screens and combines them with the portability and camera features of iPhones. When opened, it is the company’s largest phone display and thinnest phone. The Duo opens and closes like a passport to double its screen size. Continue reading...

The Guardian AI來源內容 · 翻譯待補全待翻譯:Apple unveils folding iPhone Duo as new CEO takes center stage

待翻譯:Curie

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Discussion | Link

Product Hunt AI來源內容 · 翻譯待補全待翻譯:Curie

待翻譯:Get ready for the game with new football features in Search

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Track live game feeds, explore detailed stats, and get custom fantasy recommendations directly in Search this season.

Google AI Blog來源內容 · 翻譯待補全待翻譯:Get ready for the game with new football features in Search
政策

待翻譯:Unifying governance across engines and catalogs in the Open Lakehouse

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:In our previous posts, we showed how open table formats, open APIs and unified governance...

Databricks Blog來源內容 · 翻譯待補全待翻譯:Unifying governance across engines and catalogs in the Open Lakehouse

待翻譯:Paul Christiano joins OpenAI Foundation Board

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Paul Christiano joins the OpenAI Foundation Board and its Safety and Security Committee, bringing experience in AI alignment, safety, and standards.

OpenAI News來源內容 · 翻譯待補全待翻譯:Paul Christiano joins OpenAI Foundation Board
研究

待翻譯:47,000 job listings reveal the engineering roles that AI is creating

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Every major transformation in tech has led to roles merging, then new ones emerging. Friction between developers and operations drove The post 47,000 job listings reveal the engineering roles that AI is creating appeared first on The New Stack.

The New Stack AI來源內容 · 翻譯待補全待翻譯:47,000 job listings reveal the engineering roles that AI is creating

待翻譯:The Sequence Chat - Issue 930: Arena’s Anastasios Angelopoulos on Chatbot Arena, Evaluation, and What Models Actually Measure

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:From Berkeley’s Chatbot Arena to Agent Arena: preference rankings, cost-per-task frontiers, and the hard problems in measuring real-world AI utility.

TheSequence來源內容 · 翻譯待補全待翻譯:The Sequence Chat - Issue 930: Arena’s Anastasios Angelopoulos on Chatbot Arena, Evaluation, and What Models Actually Measure

待翻譯:We have started losing control of AI. It’s time to shut it down | Garrison Lovely

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:What sounds like the overwrought penultimate episode in a sci-fi series about AI doom is now our reality On Tuesday, a former OpenAI researcher quit his job at Anthropic, warning that “neither company is acting responsibly” and that “the people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt.” As someone who’s reported on AI risk for years, this wasn’t news to me. But the outpouring of alarm suggests a much wider public is properly confronting this ludicrous situation for the first time. Continue reading...

The Guardian AI來源內容 · 翻譯待補全待翻譯:We have started losing control of AI. It’s time to shut it down | Garrison Lovely

待翻譯:“AI factories are among the most complex systems ever built”: Nvidia and Palantir turn Nvidia’s supply chain into a proving ground for sovereign AI

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Nvidia and Palantir announced on Thursday that they’re working together to bring “sovereign AI to critical supply chains,” kicking off The post “AI factories are among the most complex systems ever built”: Nvidia and Palantir turn Nvidia’s supply chain into a proving ground for sovereign AI appeared first on The New Stack.

The New Stack AI來源內容 · 翻譯待補全待翻譯:“AI factories are among the most complex systems ever built”: Nvidia and Palantir turn Nvidia’s supply chain into a proving ground for sovereign AI

待翻譯:Expanding AI access and cyber defense for federal, state, local, and tribal governments

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:OpenAI and GSA will offer eligible federal, state, local, and tribal governments $0 license fees, 50% off usage, and expanded cyber defense support.

OpenAI News來源內容 · 翻譯待補全待翻譯:Expanding AI access and cyber defense for federal, state, local, and tribal governments

待翻譯:Compact Visuotactile World Models for Lifting: Prediction, Reward Alignment, and Force Constraints

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09597v1 Announce Type: new Abstract: Accurate contact prediction is useful for robotic manipulation only if it supports effective decisions. We investigate this connection using a compact, randomly initialized visuotactile world model, trajectory-level uncertainty calibration, and behavior-initialized actor-critic learning in imagination. On 160 MuJoCo Lift episodes, adding touch reduces endpoint-force prediction error from 1.058 to 0.228 N and interval-peak error from 2.724 to 0.523 N across three training seeds. However, tactile persistence achieves lower errors of 0.095 and 0.498 N, respectively. Two exploratory control rounds comprise 680 executions on 40 independent test initial conditions. A matched reward revision on fresh test environments in…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:Compact Visuotactile World Models for Lifting: Prediction, Reward Alignment, and Force Constraints

待翻譯:AccelMPC: High-Rate, Low-Power FPGA-Accelerated Model Predictive Control for Tiny Drones

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09380v1 Announce Type: new Abstract: Unlocking the potential of tiny aerial robots requires order of magnitude improvements in the performance of embedded edge control. In particular, although recent cached model predictive control (MPC) solvers can handle the fast system dynamics and complex constraints required for agile drone flight, their computational demands remain prohibitive for resource-constrained robots, forcing prior implementations to operate at reduced control rates. AccelMPC overcomes this challenge through an end-to-end co-design approach that jointly optimizes the solver algorithm, numerical representation, hardware mapping, and physical integration. AccelMPC pairs a co-designed FPGA-accelerated alternating direction method of multip…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:AccelMPC: High-Rate, Low-Power FPGA-Accelerated Model Predictive Control for Tiny Drones

待翻譯:Learning to Fly: Stable Vision-Guided UAV Servoing with Compact Target-Centric Cues and Reinforcement Learning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09234v1 Announce Type: new Abstract: Vision-guided reinforcement learning for Unmanned Aerial Vehicles (UAVs) remains challenging due to unstable policy optimisation, aggressive exploration, and the cost of high-dimensional visual perception. In this work, we investigate long-horizon UAV visual servoing using compact target-centric cues combined with low-dimensional sensor measurements. Rather than learning directly from RGB images, lightweight target segmentation provides image-space offsets and relative depth, which are combined with quadrotor velocity and projected-gravity measurements into a compact 12D policy observation. We compare Direct PPO with three matched-budget curriculum strategies: a Visual curriculum that progressively expands target…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:Learning to Fly: Stable Vision-Guided UAV Servoing with Compact Target-Centric Cues and Reinforcement Learning

待翻譯:The Living Library: Transforming Archival Collections into Conversational Knowledge Systems -- Lessons from the Theodore Roosevelt Presidential Library

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09368v1 Announce Type: new Abstract: We present the Living Library, an end-to-end framework for transforming fragmented digital archives into governed, conversational, in-person exhibit experiences. Developed and deployed at the Theodore Roosevelt Presidential Library, the framework comprises four layers: digitization and corpus creation, AI-powered processing, retrieval and reasoning, and an optional embodied conversational interface. The first three layers aggregate a 300,000-record collection, apply OCR and structured metadata enrichment for expert curatorial review, and publish records to a hybrid dense/semantic index. Expert review is conducted through the Archivist App, a curator-facing interface that supports correction of AI-generated transcr…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:The Living Library: Transforming Archival Collections into Conversational Knowledge Systems -- Lessons from the Theodore Roosevelt Presidential Library

待翻譯:DensePol: Dense-Angle Polarization Dataset for Learning-Based Polarimetric Vision

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09359v1 Announce Type: new Abstract: Polarimetric vision is gaining increasing attention because it provides physical cues about scene shape, material, and reflection that are difficult to recover from RGB alone. Recent work has therefore explored predicting polarization directly from conventional RGB images; however, the fidelity of these methods strongly depends on the polarization supervision used for training. Most existing datasets rely on Division-of-Focal-Plane (DoFP) cameras with four spatially interleaved analyzer orientations, which provide limited angular redundancy and introduce interpolation and instantaneous-field-of-view errors. We introduce DensePol, a high-redundancy RGB--polarization dataset based on Division-of-Time (DoT) acquisiti…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:DensePol: Dense-Angle Polarization Dataset for Learning-Based Polarimetric Vision

待翻譯:The Mutations of Machine Speech

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09496v1 Announce Type: new Abstract: Algorithmic outputs now populate the digital environments through which contemporary life is organized. The role of law in facilitating and constituting (rather than merely responding to) these processes is gaining increasing traction across scholarly accounts. This inquiry traces the evolution of algorithmic outputs attending to their legal underpinnings and social implications, surfacing the mutations of machine speech. The first mutation redefined speech as data to be queried: search engines transformed the web from a space of information retrieval into an economic regime of algorithmic visibility. The second mutation reframed speech as engagement: social media platforms fused moderation with amplification, tur…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:The Mutations of Machine Speech

待翻譯:Benchmarking Hybrid Deep Research Across Database Querying and Web Search

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09410v1 Announce Type: new Abstract: While autonomous agents have made significant strides in "deep research" by iteratively navigating the open web to synthesize information, real-world problem-solving is rarely confined to a single environment. Complex analytical tasks inherently require agents to weave together evidence from both ambiguous unstructured text (e.g., the open web) and highly precise structured data (e.g., relational databases). However, existing benchmarks evaluate these modalities in isolation, failing to capture the critical "handoff" - the ability to preserve constraints when moving evidence between systems. We introduce HybridDeepResearch, to our knowledge the first deep-research benchmark that requires both web search and SQL to…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:Benchmarking Hybrid Deep Research Across Database Querying and Web Search

待翻譯:StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09264v1 Announce Type: new Abstract: Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition math, such as the IMO and Putnam, that poorly represent field-specific applications. We introduce StochBench, a Lean 4 benchmark of 450 graduate stochastic-processes problems at varying abstraction levels, each paired with its natural-language source. Addressing a field underrepresented in Mathlib, it covers finite and countable Markov chains, renewal processes, random walks, martingales, stopping times, queues, Brownian motion, stochastic calculus, weak convergence, and Poisson and continuous-time Markov processes. Our Opus 4.8-based agent achieves a 34.9% proof rate (157/450) under a 15-minute pe…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean

待翻譯:HB-PVI: A Hierarchical Bayesian Personalization and Value-of-Information Framework for Complex Activity Recognition

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05582v1 Announce Type: new Abstract: Personalization can improve activity-recognition performance, but participant-specific gains are heterogeneous, and every additional calibration label has an acquisition cost. This study presents HB-PVI, a hierarchical Bayesian personalization and value-of-information framework jointly modeling participant heterogeneity, the benefit and harm of four personalization mechanisms, and the economic value of an additional label, for the 47-participant MUSIC-CAR complex-activity cohort. A leakage-safe, leave-one-participant-out evaluation combines a sequential-Monte-Carlo participant-effect updater with a Student-$t$ hierarchical gain model and a one-step expected-value-of-sample-information (EVSI) stopping rule. Adapter…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:HB-PVI: A Hierarchical Bayesian Personalization and Value-of-Information Framework for Complex Activity Recognition

待翻譯:Multi-granularity Adaptive Hypergraph Representation Learning via Granular-ball

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05574v1 Announce Type: new Abstract: Hypergraph representation learning aims to capture high-order information in graphs by constructing hyperedges that simultaneously connect multiple nodes. These hyperedges adapt to the graph's topological features, facilitating the extraction of high-order relationships at multiple granularities. Most prior work relies on predefined definitions to generate hyperedges, overlooking the diversity in graph topological structures and the multi-granularity characteristics of hyperedges. As a result, this limits their ability to effectively and adaptively discover high-order relationships and efficiently process complex structural information. To address this limitation, we propose a novel framework called \underline{M}u…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:Multi-granularity Adaptive Hypergraph Representation Learning via Granular-ball

待翻譯:When Do Options Help? Policy Necrosis and Redundant Coverage in Option-Critic

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05508v1 Announce Type: new Abstract: Option-critic learns options: sub-policies together with a learned rule for when each one hands control back. Its headline result is that performance improves as options are added. We explain that result, with theory and experiment. First, the termination rule option-critic learns by maximising return contributes nothing. When the termination test and the policy that picks options read the same values, the test fires at every step, so the learned rule is identical to always terminating. When that policy explores and the test does not, as in option-critic itself, the rule can block the exploration; there are instances where it suffers $\Omega(T)$ regret while always terminating holds to $O(\log T)$. Forcing termina…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:When Do Options Help? Policy Necrosis and Redundant Coverage in Option-Critic

待翻譯:PGP-Clinical-TimeKAN: Prior-Guided Joint Probabilistic Forecasting of Clinical Trajectories

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05488v1 Announce Type: new Abstract: Clinical deterioration unfolds through coupled, partially observed trajectories, not a single diagnostic label. We introduce PGP-Clinical-TimeKAN, a trajectory-first framework for joint probabilistic forecasting of multivariate physiology. It combines missingness-aware temporal encoders, a soft organ-system prior, patient-specific relations, nonlinear Kolmogorov-Arnold messages, and a low-rank multivariate Student-t head. We evaluate 24-hour histories and six-hour forecasts on a frozen MIMIC-IV-derived cohort of 6,882 patients and 54,694 windows. Across five seeds and 13 models, PGP-Clinical-TimeKAN obtains the second-lowest normalized MAE (0.37727 +/- 0.00029) and the lowest RMSE (0.52656 +/- 0.00034). It reduces…

arXiv AI來源內容 · 翻譯待補全待翻譯:PGP-Clinical-TimeKAN: Prior-Guided Joint Probabilistic Forecasting of Clinical Trajectories

待翻譯:RAPID: Reliability-Aware Pair Importance Distillation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05481v1 Announce Type: new Abstract: Inter example relational distillation transfers a teacher's representation geometry by matching relations among examples within a mini batch. Computing all pairs has quadratic complexity in the batch size, whereas uniform subsampling may use a limited relation budget inefficiently. We introduce Reliability Aware Pair Importance Distillation, or RAPID, which separates a reliability gated relational target from a full support adaptive pair proposal. Reliability determines which teacher relations are emphasized, while calibrated teacher entropy and detached student-teacher residuals determine which relations are evaluated. Exact inverse proposal correction makes the loss and gradient estimators conditionally unbiased…

arXiv AI來源內容 · 翻譯待補全待翻譯:RAPID: Reliability-Aware Pair Importance Distillation

待翻譯:CriticGen: Generation-Aware Evaluation as Actionable Feedback

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05439v1 Announce Type: new Abstract: Current evaluation methods for large language models are coarse-grained and decoupled from generation, producing generic explanations that fail to provide actionable feedback for model improvement. We propose CriticGen, a fine-grained, generation-aware evaluation framework that turns evaluation into actionable control for answer improvement. CriticGen first generates sample-specific evaluation dimensions and scoring criteria under high-level categories such as subjective, objective, and self-derived constraints. These criteria then serve as a dynamic rubric for jointly producing a score, a reason, an executable refinement suggestion, and a refined answer. This rubric-conditioned refinement process enables models t…

arXiv AI來源內容 · 翻譯待補全待翻譯:CriticGen: Generation-Aware Evaluation as Actionable Feedback

待翻譯:Quoting Calif Research

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯: Today, we're releasing a demo of WeWorm, the first zero-click worm to spread through WeChat calls across iOS and Android. [...] The victim does not need to answer the call, or interact with their phone at all. Even if they do answer, they hear nothing, and the exploit still succeeds. [...] Working with AI, our team found the bug and wrote the first remote code execution (RCE) exploit in about two days. Building the worm took one more week. A worm at this scale used to be the kind of thing that took a larger team months. AI can already do most of the work here. Our team provided the judgment about what to target and how to test it safely. — Calif Research, WeWorm Tags: ai-security-research, ai, llms, security, generative-ai

Simon Willison's Weblog來源內容 · 翻譯待補全待翻譯:Quoting Calif Research

待翻譯:Lawmakers blast AI companies after researcher warns of human extinction by 2030

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Former Anthropic employee Jacob Coxon said AI will become ‘superhuman systems’ that can cause human extinction by the end of the decade Just a day after three Anthropic researchers warned that artificial intelligence could kill off humanity within the decade, lawmakers have begun lashing out about the risks of the burgeoning technology. Ted Cruz, a republican senator from Texas, said in an interview on ABC’s The View, that AI poses a “catastrophic risk” and that he “read that whole tweet thread that that developer put out. It was highly concerning. Continue reading...

The Guardian AI來源內容 · 翻譯待補全待翻譯:Lawmakers blast AI companies after researcher warns of human extinction by 2030

待翻譯:To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We ported ThunderKittens to NVIDIA's Vera Rubin NVL72 and rebuilt our NVFP4 GEMM around the new hardware, taking it from 42% of roofline to over 22 PFLOPS — competitive with cuBLAS and CuTe DSL. Here is what changed in the ISA and how we used it.

Together AI Blog來源內容 · 翻譯待補全待翻譯:To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

待翻譯:MIT Schwarzman College of Computing launches pilot to help educators teach AI across disciplines

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:A weeklong summer workshop brought higher education faculty to campus to explore how AI and machine learning materials can be adapted for their classrooms.

MIT News AI來源內容 · 翻譯待補全待翻譯:MIT Schwarzman College of Computing launches pilot to help educators teach AI across disciplines

待翻譯:How Baseten makes pyannote’s diarization models 9.6x faster

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Model performance How Baseten makes pyannote’s diarization models 9.6x faster Using quality-aware quantization, index-based clustering, and scheduling optimizations Authors Matte Lim Ansel Erol Last updated September 9,…

Baseten Blog來源內容 · 翻譯待補全待翻譯:How Baseten makes pyannote’s diarization models 9.6x faster
模型

待翻譯:DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Long-horizon agents have turned LLM serving into an input-heavy workload. Repeated prefills and million-token contexts leave KV caches that strain HBM, SSD capacity, and bandwidth. DeepSeek AI built its newest release around that exact bottleneck. DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters, 196B additional Engram parameters, and a 1M-token context window. It […] The post DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse appeared first on MarkTechPost.

MarkTechPost來源內容 · 翻譯待補全待翻譯:DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse

待翻譯:Introducing ChatGPT for Financial Services

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Introducing ChatGPT for Financial Services, combining built-in financial data and GPT-6 Astra for research, modeling, and client-ready materials.

OpenAI News來源內容 · 翻譯待補全待翻譯:Introducing ChatGPT for Financial Services

待翻譯:Weatherwatch: AI model beats standard methods at predicting cyclones

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Artificial intelligence system delivers a three-day forecast as accurate as previous two-day predictions Google’s WeatherNext AI model is better than existing systems at forecasting cyclones, according to a paper published in Nature. The AI-based technology gives an extra day of warning, providing a three-day forecast for hurricanes or typhoons that is as accurate as previous two-day predictions. Continue reading...

The Guardian AI來源內容 · 翻譯待補全待翻譯:Weatherwatch: AI model beats standard methods at predicting cyclones

待翻譯:No Free Checker: A Survey of Verifiers for Robot Policies

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09250v1 Announce Type: new Abstract: A verifier for robot policies reads a candidate behavior and returns a score for how well it did, used both to evaluate vision-language-action policies and to train them. Verifiers range from success detectors and reward models to runtime monitors, safety filters, and temporal-logic specifications. We survey roughly 150 verifiers and compare them along two properties. Availability is how much a verdict costs, how early in a rollout the verdict arrives, and how often a verdict can be asked for. Availability rises as verdicts get cheaper, earlier, and denser. Credibility is how much a high score tells us about the task. Credibility falls as the judgment becomes gameable and self-serving. We group the verifiers by wh…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:No Free Checker: A Survey of Verifiers for Robot Policies

待翻譯:Geometry Conditioning in an Embodied SLM: Training Controls and Robustness Diagnostics in a 0.8B Hybrid Model

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09213v1 Announce Type: new Abstract: We study how physical-state inputs affect a 0.8B hybrid language model adapted for manipulation with 6.2M trainable parameters. Six conditions are trained on three LIBERO-Spatial tasks and evaluated over three seeds and 540 held-out rollouts. Conditioning recurrent decay gates on geometric increments yields 28.9% success, compared with 36.7% when those increments are shuffled during training and 24.4% without explicit object/goal geometry. Both geometry policies receive correct inputs at evaluation. A token adapter using the same increments scores 27.8%; differences vary across seeds and remain inconclusive. Token-clock conditioning scores 11.1%, including one seed that fails to converge. In separate robustness te…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:Geometry Conditioning in an Embodied SLM: Training Controls and Robustness Diagnostics in a 0.8B Hybrid Model

待翻譯:OmniPoint: Universal Monocular Metric Pointcloud from Any Camera

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09394v1 Announce Type: new Abstract: Recovering metric 3D geometry from monocular images is a fundamental computer vision task, yet current methods remain heavily fragmented by fixed camera model assumptions and inflexible input schemes. We present OmniPoint, a unified framework designed to generalize metric reconstruction across diverse imaging sensors, including pinhole, fisheye, and equirectangular projections, while accommodating varying geometric priors. To overcome projection rigidity, OmniPoint abandons conventional planar depth regression. It instead adopts a decoupled ray and distance representation alongside a decoupled training objective, explicitly separating the camera projection model from the scene structure. To address the severe scar…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:OmniPoint: Universal Monocular Metric Pointcloud from Any Camera

待翻譯:Video-MOPD: Multi-Teacher On-Policy Distillation for Video Understanding

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09300v1 Announce Type: new Abstract: Video understanding demands a convergence of complementary capabilities across perception, temporal understanding, and complex reasoning, which are difficult to jointly optimize within a single model. We introduce Video-MOPD-8B, an open-weight model dedicated to video understanding tasks. To fundamentally enhance its capabilities, we conduct targeted reinforcement learning (RL) optimization across three core domains: video temporal grounding (VTG), general video comprehension, and video STEM reasoning. We then unify their complementary capabilities via Multi-Teacher On-Policy Distillation (MOPD), which consolidates expert knowledge by supervising student-generated trajectories with routed teacher feedback. We furt…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:Video-MOPD: Multi-Teacher On-Policy Distillation for Video Understanding

待翻譯:MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09206v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) often struggle with hallucinations, thus hindering their reliable practical applications. Existing attention-based mitigation methods mainly rely on indirect signals (e.g., attention weights) that fail to accurately reflect the actual information shift underlying hallucination generation. In this paper, we propose HEAL, Head-lEvel information disentAnglement and caLibration for identifying and mitigating hallucinations. HEAL first employs causal noise intervention on multi-head outputs to filter out causally redundant heads. Subsequently, it disentangles information distribution within the remaining heads via the counterfactual Difference-in-Differences, categorizing heads…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads

待翻譯:Lensless Gaze Is Not Private by Default: Auditing Identity Leakage Across Disclosure Surfaces

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09188v1 Announce Type: new Abstract: Lensless near-eye sensing is often described as privacy-friendly because its coded measurements are visually unintelligible. Yet visual unintelligibility reflects human interpretation, not what a learned adversary can recover. We therefore treat identity privacy as a systems property of disclosure surfaces: representations crossing sensing, storage, computation, and output boundaries. We audit a simulated lensless gaze pipeline under a 36-subject known-gallery closed-set identification protocol with a fixed, known PSF; privacy from an unknown or varying optical key is outside our scope. Reported accuracies are empirical attack success rates under matched linear and MLP probes and do not upper-bound stronger advers…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:Lensless Gaze Is Not Private by Default: Auditing Identity Leakage Across Disclosure Surfaces

待翻譯:M2LG-DG: A Multi-modal Local-Global Domain Generalization Framework for Cross-site Major Depressive Disorder Classification

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09186v1 Announce Type: new Abstract: Classification models based on resting-state functional magnetic resonance imaging (rs-fMRI) often show lower performance at imaging sites not included during model development, which can limit their use in clinical settings. Domain generalization (DG) addresses this issue by learning representations from source sites that remain effective for unseen target sites. However, existing DG approaches for psychiatric disorder classification commonly rely on a single imaging modality and may not fully account for site-specific acquisition effects on the learned representation space. Subjects scanned at the same site share scanner hardware, acquisition settings, and preprocessing characteristics, which can cause represent…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:M2LG-DG: A Multi-modal Local-Global Domain Generalization Framework for Cross-site Major Depressive Disorder Classification

待翻譯:Integrating Unimodal and Vision-Language Representations in Latent Space for Multi-Label Chest X-Ray Classification

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09185v1 Announce Type: new Abstract: Multi-label chest X-ray classification has attracted considerable attention in recent years, with the effective use of visual representations and clinical semantic knowledge playing an important role. This study proposes a framework that combines unimodal representations from RAD-DINO with vision--language representations from BioViL-T for the classification of 14 labels in the MIMIC-CXR-JPG dataset. The RAD-DINO and BioViL-T embeddings and their combined representation are refined separately in latent space before being normalized and fused across the three branches. In addition to improving classification performance, the study aims to clarify the role of each embedding source and the degree to which they comple…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:Integrating Unimodal and Vision-Language Representations in Latent Space for Multi-Label Chest X-Ray Classification

待翻譯:Evidence-Order Calibration for Selective Visual Reasoning under Progressive Loss of Question-Critical Evidence

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09184v1 Announce Type: new Abstract: Vision-language model (VLM) confidence may change in aggregate when visual evidence is degraded while remaining structurally inconsistent within individual examples. We study answer-level reliability along five-step, question-conditioned evidence-loss trajectories. Using a frozen Qwen2.5-VL-3B-Instruct model, we construct 176 accepted GQA-derived trajectories (880 masking conditions) by progressively masking scene-graph-localized question-critical regions. Native sequence confidence has an evidence monotonicity violation rate (EMVR) of 0.436, and 92.0% of trajectories contain at least one adjacent violation. A matched non-critical-region control shows that full critical masking reduces accuracy by 28.2 percentage…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:Evidence-Order Calibration for Selective Visual Reasoning under Progressive Loss of Question-Critical Evidence

待翻譯:TEFM: Token-Efficient Faithful Modeling for Structured Data

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09552v1 Announce Type: new Abstract: In this paper, we solve two fundamental obstacles in applying LLMs to critical domains: token efficiency and faithfulness. To address both constraints jointly, we present TEFM (Token-Efficient Faithful Modeling), a framework designed for structured data analysis in critical domains. TEFM achieves token efficiency by compressing lengthy structured observations into compact Behavioral Code tokens, dramatically reducing token consumption with minimal information loss. Moreover, TEFM enables faithful rationalization through a dual-fidelity objective that jointly optimizes code-level reconstruction and prediction-level fidelity, identifying minimal sufficient feature subsets grounded in input data. Comprehensive experi…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:TEFM: Token-Efficient Faithful Modeling for Structured Data

待翻譯:Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09425v1 Announce Type: new Abstract: Educational data filters have become a practical way to improve language-model pre-training, but most filters treat educational value as a single scalar property. This may be too broad for some applications, especially if the data set already features a high density of educational material. Useful learning material needs to be accurate, engaging, well structured, and appropriate for the intended audience and application (e.g. learner- vs teacher-facing). Following QuRating (Wettig et al. 2024), we introduce Edu-QuRating: a pipeline for multi-dimensional educational data scoring and curation. Edu-QuRating defines education-specific rubrics, uses an LLM judge to label sampled document pairs and distills those pairwi…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements

待翻譯:Do LLMs Make More Mistakes If They Do Not Believe the Input Data?

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09363v1 Announce Type: new Abstract: Large language models (LLMs) are prone to hallucinating or misinterpreting facts, which impairs their usability in retrieval-augmented generation or data-to-text systems. We analyse how faithfulness of LLMs to provided context depends on how plausible they perceive the context to be (context-memory conflict). To better identify error patterns, we make use of the increased difficulty of non-English and low-resource language text generation and input data based on local knowledge, only partially captured in models' parametric knowledge. We let the models generate text in English, Czech, Slovak and Upper Sorbian from factual (FA), counterfactual (CFA) and fictional (FI) RDF triples containing local Czech and Slovak d…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:Do LLMs Make More Mistakes If They Do Not Believe the Input Data?

待翻譯:Auditable Emergency Triage for Maternal and Newborn Care in India

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09356v1 Announce Type: new Abstract: At Noora Health, our nurses answer more than 50,000 medical queries per month on our WhatsApp-based service that provides caregivers with on-demand support. Their most time-critical task is emergency triage: deciding which queries need immediate in-person attention. To support them, we built a system that uses a large language model (LLM) to classify whether a message is an emergency and provide a rationale for interpretability. But the system was opaque: analyzing mistakes meant reading reasoning chains for each message, which is infeasible at our scale. Prompt changes meant re-running a full evaluation to prevent regressions, which was both costly and operationally challenging. Clinicians follow a decision tree…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:Auditable Emergency Triage for Maternal and Newborn Care in India

待翻譯:SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09349v1 Announce Type: new Abstract: Modern LLMs demonstrate impressive multilingual performance, yet standard benchmarks primarily reward selecting correct answers rather than evaluating genuine factual understanding. We introduce Systematic Wikidata-based Object-Relation Distortion (SWORD), a benchmark that evaluates whether models consistently reject factual errors across languages. SWORD generates syntactically well-formed but factually incorrect statements in eight widely spoken languages through controlled perturbations of Wikidata triples, ranging from random entity substitutions to semantically plausible property-based selections. Our distortion-based evaluation surfaces two critical insights that remain entirely obscured by conventional benc…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection

待翻譯:Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09338v1 Announce Type: new Abstract: Speculative decoding is critical for accelerating LLM inference. However, the speedup is fragile: drafters are typically trained against a narrow distribution for a single target model, and their acceptance rate collapses under workload shifts. This is a striking inversion of modern LLM development, where target models are valued precisely for the broad generalization they acquire through large-scale pretraining. We argue that the natural remedy, pretraining, has been hard to apply to drafters because existing recipes are target-specific: the drafter consumes the target's hidden states and is distilled on the target's logits, so pretraining must be repeated for each target. We introduce Osprey, which instead boots…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding

待翻譯:X-CoSD: Communication-Efficient Cross-Vocabulary Collaborative Speculative Decoding

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09166v1 Announce Type: new Abstract: This paper investigates collaborative speculative decoding (CoSD), a distributed large language model (LLM) inference framework in which an on-device small language model (SLM) drafts candidate tokens and a server LLM verifies them. Existing CoSD methods assume a shared vocabulary between the SLM and the LLM and incur substantial communication load because residual resampling requires token distribution exchange between the user device and the edge server. To address these limitations, we propose cross-vocabulary CoSD (X-CoSD), a lossless and communication-efficient CoSD framework for heterogeneous SLM-LLM vocabularies. X-CoSD is built on hybrid resampling (HR), which splits residual resampling across the common-v…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:X-CoSD: Communication-Efficient Cross-Vocabulary Collaborative Speculative Decoding

待翻譯:GraphNOSE: A Graph Transformer in Olfaction

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05694v1 Announce Type: new Abstract: Predicting olfactory qualities from molecular structure is an open problem in chemoinformatics. Although linear models can link molecular features to odor descriptors, they often fail when extrapolating to novel chemical scaffolds, extreme molecular weights, or complex odor mixtures. To address this, we introduce GraphNOSE, an open-source graph transformer framework that predicts multi-label odor descriptors from simplified molecular-input line-entry system (SMILES) strings for single molecules and binary mixtures. By integrating positional and structural encodings within a transformer-based graph architecture, GraphNOSE achieves strong performance with six times fewer parameters than standard graph neural network…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:GraphNOSE: A Graph Transformer in Olfaction

待翻譯:Connecting Score Matching, Maximum Likelihood, and Expectation-Maximization in Mixed Linear Regression

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05688v1 Announce Type: new Abstract: We study variance-preserving diffusion of the response in mixed linear regression (MLR) with unknown mixing weights. Our analysis separates the statistical guarantees of score matching from the loss geometry and optimization signal at a fixed diffusion noise level. The KL divergence links the denoising score matching objective integrated over the diffusion path with the likelihood and a terminal discrepancy. Under mild regularity conditions and terminal schedule, the resulting estimator converges up to the ground truth parameters of MLR, and its scaled error converges to the Gaussian limit of the maximum-likelihood estimator. At a fixed scale of the diffusion noise level, we derive a decomposition linking the scor…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:Connecting Score Matching, Maximum Likelihood, and Expectation-Maximization in Mixed Linear Regression

待翻譯:PAC-Private Autoregressive Generation: Calibrating Noise to Ensemble Disagreement

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05676v1 Announce Type: new Abstract: Language models adapted on private text are often served through APIs, so privacy leakage occurs through generated outputs rather than exposed weights. Private prediction protects these releases. Methods such as PMixED incur privacy cost at each release and increasingly rely on the public model over long horizons. PAC privacy instead calibrates noise to output variability across possible secrets, adding less noise when predictions are stable. To our knowledge, PAC-private prediction has not previously been extended from classification to autoregressive generation. We construct $m=128$ overlapping worlds from the private corpus, with each record appearing in exactly $m/2$ worlds, and train one adapter per world ove…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:PAC-Private Autoregressive Generation: Calibrating Noise to Ensemble Disagreement

待翻譯:Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05658v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being explored for automating SystemVerilog Assertion (SVA) generation, yet most evaluations report correctness on a single syntactic representation of an input. Such point accuracy does not reveal whether a model's correct output is stable when the same RTL behavior is written differently. This paper presents a controlled metamorphic evaluation of LLM-based SVA generation under semantics-preserving RTL transformations. Starting from the VERT dataset, we construct a quality-filtered conditional-control pool and a stratified 40-program evaluation set containing 295 assignment behaviors. We evaluate two open code models, Qwen2.5-Coder-7B and DeepSeek-Coder-V2-Lite, with…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations

待翻譯:Capsule Lens: Locating and Tracking Concept Geometry in Model Representations

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05575v1 Announce Type: new Abstract: Understanding how concepts are encoded in the internal representations of machine learning models is a central problem in mechanistic interpretability, essential both for the science of deep learning and for the trustworthy deployment of increasingly capable models. Existing approaches to interpret model representations mainly map representations onto more interpretable spaces and do not directly characterize how concepts occupy representation space; various hypotheses have been proposed, but often lack of rigorous validation and largely focus on static representations. In this work, we introduce Capsule Lens, a framework that matches the region a concept occupies with a simple, trackable geometric form, a capsule…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:Capsule Lens: Locating and Tracking Concept Geometry in Model Representations

待翻譯:SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05505v1 Announce Type: new Abstract: Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations of large language models (LLMs) typically examine review stages in isolation. We introduce SciLitBench, a multi-stage benchmark spanning title and abstract screening, full-text screening, and schema-guided data extraction, with 42,981 retrieved records, 1,012 full texts, and annotations for 888 included papers. Across 22 open-weight LLMs from six model families, explicit inclusion and exclusion criteria improve title and abstract screening $F_2$ by 28.8\%, while researcher-authored rationales improve full-text screening by 15\%. Data extraction reveals a different reliability regime: performance declines from…

arXiv AI來源內容 · 翻譯待補全待翻譯:SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews

待翻譯:ARC-Bench: Closed-Loop Replanning Masks Broken Action Ranking in Frozen JEPA World Models

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05461v1 Announce Type: new Abstract: Reward-free latent world models plan by scoring candidate actions with distances in a frozen latent space: an action is preferred if its predicted future embedding lands closer to the goal embedding. This silently assumes that latent closeness is action-rankable, i.e., that ordering candidates by latent distance agrees with ordering them by true cost. We audit this assumption directly. We introduce ARC-Bench, a no-leak, fixed-candidate protocol that measures whether frozen JEPA-style objectives rank candidate actions correctly, and apply it to official released JEPA-WM checkpoints across navigation and manipulation-style control. The assumption fails, severely and structurally: on the official manipulation audits…

arXiv AI來源內容 · 翻譯待補全待翻譯:ARC-Bench: Closed-Loop Replanning Masks Broken Action Ranking in Frozen JEPA World Models

待翻譯:Damage-Aware Bandit Pruning for Vision and Language Transformers

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05448v1 Announce Type: new Abstract: Structured post-training pruning of transformers requires selecting complete functional units whose suppression causes limited degradation. We formulate structured-unit selection for language and vision transformers as a damage-aware multi-armed bandit problem under a fixed candidate-evaluation budget. Attention heads and MLP channel groups are temporarily masked on calibration batches. Paired damage is the masked loss minus the base loss on the same batch, reducing batch-to-batch variation. A smooth bounded reward drives either a UCB-style policy or fractional-Beta Thompson Sampling, and the final mask is constructed sequentially by adding one unit at each step. The selected units are functionally zeroed in the o…

arXiv AI來源內容 · 翻譯待補全待翻譯:Damage-Aware Bandit Pruning for Vision and Language Transformers

待翻譯:Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.05437v1 Announce Type: new Abstract: Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., `do not steal'). However, social intelligence depends not only on norm recognition, but also on anticipating who will enforce it and how (e.g., public shame or even imprisonment). These second-order expectations, known as metanorms, govern how people respond when social rules are broken. We introduce a novel framework for evaluating metanorm reasoning in Large Language Models (LLMs) along two dimensions: emotional appraisal and behavioral response, and propose new classification tasks, namely, predicting self-regulation in violators, and other-regulation in observer…

arXiv AI來源內容 · 翻譯待補全待翻譯:Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models

待翻譯:Build more natural voice experiences with GPT‑Live‑1 in the API

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:GPT‑Live‑1 brings natural, full-duplex voice conversations to the API, with stronger instruction following, custom voices, and telephony support.

OpenAI News來源內容 · 翻譯待補全待翻譯:Build more natural voice experiences with GPT‑Live‑1 in the API

待翻譯:Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Learn how to deploy Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter open-weight model, on Amazon SageMaker HyperPod with vLLM. This walkthrough covers cluster provisioning, NVFP4 quantization, and an OpenAI-compatible endpoint with built-in reasoning, tool calling, and native MTP speculative decoding.

AWS Machine Learning Blog來源內容 · 翻譯待補全待翻譯:Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

待翻譯:“It could kill us all”: what Anthropic’s own researchers really think about superintelligence

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:On Tuesday evening, Anthropic pretraining researcher Jacob Coxon announced on X that he’d resigned. Within hours, two of his colleagues The post “It could kill us all”: what Anthropic’s own researchers really think about superintelligence appeared first on The New Stack.

The New Stack AI來源內容 · 翻譯待補全待翻譯:“It could kill us all”: what Anthropic’s own researchers really think about superintelligence

待翻譯:Microsoft has new AI privacy rules for schools

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Microsoft agreed to a set of safety and privacy principles for AI in schools a week after two major school systems announced a ban on student-facing AI. In a new agreement with the American Federation of Teachers (AFT), the second-largest teachers union in the US, and its New York City affiliate the United Federation of Teachers (UFT), Microsoft committed to ten principles that can be contractually enforced by school districts that adopt them. The terms include pledging not to train AI models on student or educator data, limiting the amount of data Microsoft collects in the first place and disclosing to families how its tools work in plain … Read the full story at The Verge.

The Verge AI來源內容 · 翻譯待補全待翻譯:Microsoft has new AI privacy rules for schools

待翻譯:Cloudera and Mistral Partner for Sovereign Enterprise AI

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We’ve spoken with many of the world’s largest enterprises across regulated industries like financial services, manufacturing and telecommunications. One of their common strategic partners is Cloudera, providing them wit…

Mistral AI News來源內容 · 翻譯待補全待翻譯:Cloudera and Mistral Partner for Sovereign Enterprise AI

待翻譯:Cloudera and Mistral Partner for Sovereign Enterprise AI

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:We’ve spoken with many of the world’s largest enterprises across regulated industries like financial services, manufacturing and telecommunications. One of their common strategic partners is Cloudera, providing them wit…

Mistral AI News來源內容 · 翻譯待補全待翻譯:Cloudera and Mistral Partner for Sovereign Enterprise AI

待翻譯:Modernizing complex legacy code with AI agents | Mistral

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Thinking Summary Mistral helped a European energy operator migrate 40,000 lines of Fortran 77 to C++. Learn how it was done, and the lessons to carry forward. Legacy scientific codebases accumulate over decades, and whe…

Mistral AI News來源內容 · 翻譯待補全待翻譯:Modernizing complex legacy code with AI agents | Mistral

待翻譯:Modernizing complex legacy code with AI agents | Mistral

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Thinking Summary Mistral helped a European energy operator migrate 40,000 lines of Fortran 77 to C++. Learn how it was done, and the lessons to carry forward. Legacy scientific codebases accumulate over decades, and whe…

Mistral AI News來源內容 · 翻譯待補全待翻譯:Modernizing complex legacy code with AI agents | Mistral
機械人

待翻譯:Actuator Dynamics Curricula for Narrow-Viability Tasks in Legged Robot Learning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09492v1 Announce Type: new Abstract: Reinforcement learning has produced capable controllers across a broad range of legged-robot tasks, but a subset of these tasks fail to converge under standard training: those for which most exploration trajectories terminate before producing useful gradient signal. To address such tasks we introduce the \emph{Actuator Dynamics Curriculum}, a procedure that initializes joint stiffness at a high value and anneals it toward the system-identified value as completed episode lengths grow. Using a cart-pole system as a representative example, we show that higher closed-loop joint natural frequency under critical damping enlarges the viability kernel of the underlying Markov Decision Process, increasing the fraction of i…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:Actuator Dynamics Curricula for Narrow-Viability Tasks in Legged Robot Learning

待翻譯:A Decade of Bayesian Optimization for Controller Tuning and Robot Learning: Tutorial, Review, and Future Prospects

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09403v1 Announce Type: new Abstract: In the past decade, Bayesian optimization (BO) has emerged as a powerful and adaptable framework for automatic controller tuning and robot learning. This article offers a comprehensive overview of the state-of-the-art in BO, designed to support both researchers and practitioners in understanding recent advancements, practical applications, and future research directions. We begin by adopting a practitioner's perspective, illustrating how to effectively set up BO through a representative controller tuning example. We position BO within the broader context of learning paradigms, ranging from deep reinforcement learning to data-driven control, and highlight scenarios where BO is most advantageous. Next, we discuss th…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:A Decade of Bayesian Optimization for Controller Tuning and Robot Learning: Tutorial, Review, and Future Prospects

待翻譯:Design and Attitude Control of an Underwater Quadruped Robot

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09217v1 Announce Type: new Abstract: Legged robots are versatile on land, but their use in underwater environments remains limited. Extending quadruped locomotion to water enables amphibious mobility with applications in inspection, environmental monitoring and disaster response. This paper presents the design, modeling, and experimental validation of a reproducible underwater quadruped robot. The robot is built around custom waterproof motor housings machined from polyoxymethylene plastic, which use off-the-shelf O-rings and dynamic shaft seals. A simplified model is derived to describe the dynamics of this underwater legged system, capturing how drag forces on spherical end effectors transmit torque to the floating base. Building on this model, a c…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:Design and Attitude Control of an Underwater Quadruped Robot

待翻譯:Identifying Habit, Physics, and Nuisance in Robot World Models

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09210v1 Announce Type: new Abstract: Teleoperated demonstrations are often multimodal even when the underlying dynamics are nearly deterministic given the executed action. We argue that this multimodality typically mixes three factors--operator habit in action selection, shared physics, and observation nuisance--and that entangled next-observation predictors absorb all three. We formalize the split with a structural causal model a=g(h,z,u), z'=f(z,a), o=r(z,c), and test it with complementary interventions: replacing or shuffling actions at fixed state sharply increases next-state error, whereas appearance and camera changes should not; habit-aware reverse scoring improves ranking of feasible pasts without rewriting the dynamics. The associated adapta…

arXiv Robotics來源內容 · 翻譯待補全待翻譯:Identifying Habit, Physics, and Nuisance in Robot World Models

待翻譯:Neura Robotics, Seco Partner to Scale Physical AI in Europe

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The companies will develop compute modules using Qualcomm Dragonwing processors, as Neura looks to scale physical AI in Europe.

AI Business來源內容 · 翻譯待補全待翻譯:Neura Robotics, Seco Partner to Scale Physical AI in Europe
芯片

待翻譯:Introducing preemptible compute: the same compute, half the price

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Together GPU Clusters now supports preemptible compute: the same GPU capacity at a flat 50% of the on-demand rate, with a five-minute drain window.

Together AI Blog來源內容 · 翻譯待補全待翻譯:Introducing preemptible compute: the same compute, half the price

待翻譯:Read the Apple document explaining how new listening features still protect your privacy

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:At Wednesday's iPhone Duo launch event, Apple announced a handful of new Siri AI Audio Intelligence features, including Siri Recap, Live Rewind, Sound Recognition, and Music Recognition. Alongside its announcement, Apple released a document laying out how it plans to balance AI "ambient listening" and users' privacy. It says the raw audio from the new features "is handled within dedicated hardware, is not saved as a file, and is not accessible to the operating system, apps, or Apple." Apple explains that Audio Intelligence relies on the Secure Exclave in the S11 chip that powers the new Apple Watch Series 12 and Apple Watch Ultra 4: The … Read the full story at The Verge.

The Verge AI來源內容 · 翻譯待補全待翻譯:Read the Apple document explaining how new listening features still protect your privacy

待翻譯:Qualcomm Forges AI Chip Deal with Amazon

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The chipmaker is competing with Nvidia, the world’s dominant producer of AI processors.

AI Business來源內容 · 翻譯待補全待翻譯:Qualcomm Forges AI Chip Deal with Amazon