本文にスキップ
AI News HubLIVE

モデルの最新ニュース

翻訳待ち:The Sequence Radar - Issue 948: Last Week in AI: Billions in Funding, Hundreds of Papers, One Big Leap?

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Reflection and Mistral expand the model race, TypeSafe and Arena raise the stakes, and OpenAI puts mathematical discovery under scrutiny.

TheSequence原典の内容 · 翻訳・分析待ち翻訳待ち:The Sequence Radar - Issue 948: Last Week in AI: Billions in Funding, Hundreds of Papers, One Big Leap?

翻訳待ち:Fixated on AI, the world has forgotten what is staring it in the face: nuclear annihilation | Simon Tisdall

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:This is no scoffing matter. The accelerating proliferation of advanced nuclear weapons is bringing ‘end times’ closer by the day Humankind has fearfully and sometimes joyfully anticipated the end of the world ever since it began. The Bible’s Book of Revelation speaks of “end times”, a period of tribulation, war, famine, plague, floods and environmental catastrophe that may seem a less absurd prospect today than perhaps it once did. Islam, too, posits a fateful day of judgment or day of reckoning. Christian evangelicals believe that, in the final rapture, a chosen few – the righteous or the “elect” – will be saved while sinners burn. It’s a minority view. Matthew’s gospel predicts most people will defy, ignore or ridicule warnings of imminent Armaged…

The Guardian AI原典の内容 · 翻訳・分析待ち翻訳待ち:Fixated on AI, the world has forgotten what is staring it in the face: nuclear annihilation | Simon Tisdall

翻訳待ち:Satya Nadella says we should assume all AI models are ‘compromised’

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:In a lengthy post on X, Microsoft's CEO laid out his views on the dangers posed by highly advanced AI models and how to confront those risks. Nadella says we can no longer accept a world where AI is treated as a "set of nested black boxes" whose advice and actions we simply accept or reject. He calls for building a more transparent system where models can be contained, observed, and leaves behind "tamper-proof human readable evidence." Many of his recommendations align with what we've heard from others in the industry: timely incident disclosure, independent audits, verifiable data, and containment. It's on that last point that he appears t … Read the full story at The Verge.

The Verge AI原典の内容 · 翻訳・分析待ち翻訳待ち:Satya Nadella says we should assume all AI models are ‘compromised’

翻訳待ち:Sakana AI’s LLM Peer Review System Catches 73% of Core-Claim Errors

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Sakana AI’s TMLR paper introduces Multi-Layered Review, a 3-agent Claude-based reviewer, and a 1,164-error Contradiction Benchmark. MLR caught 73.43% of core-claim errors, versus 14.81% for the best prior system. The post Sakana AI’s LLM Peer Review System Catches 73% of Core-Claim Errors appeared first on MarkTechPost.

MarkTechPost原典の内容 · 翻訳・分析待ち翻訳待ち:Sakana AI’s LLM Peer Review System Catches 73% of Core-Claim Errors

翻訳待ち:Building AI for Reliable Execution: Lessons From Industrial Robotics

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Inside Standard Bots’ AI stack, pretrained models learn factory tasks from demonstrations and improve through corrections from real deployments.

Latent Space原典の内容 · 翻訳・分析待ち翻訳待ち:Building AI for Reliable Execution: Lessons From Industrial Robotics

翻訳待ち:Microsoft AI Releases Microsoft-Decision-1: A Qwen3.5-9B Decision-Scoring Model

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Microsoft has released Microsoft-Decision-1, a decision model for routing, classification, verification and agent control. Microsoft-Decision-1 is a decision-scoring model that returns a calibrated probability for each fixed answer option instead of generated text. It is post-trained from Alibaba’s Qwen3.5-9B and available now in Microsoft Foundry and OpenRouter. . TL;DR What is a decision model? A […] The post Microsoft AI Releases Microsoft-Decision-1: A Qwen3.5-9B Decision-Scoring Model appeared first on MarkTechPost.

MarkTechPost原典の内容 · 翻訳・分析待ち翻訳待ち:Microsoft AI Releases Microsoft-Decision-1: A Qwen3.5-9B Decision-Scoring Model

翻訳待ち:Quoting The New York Times

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要: Anthropic detailed the activity of its A.I. agents in a blog post on Friday, without naming the targeted websites. But two sources with knowledge of the incidents said Anthropic’s A.I. agents had submitted 20 visa applications through a form available on the State Department’s website. All the applications were incomplete and were not processed, they said. — The New York Times, Anthropic Agents Tried to Fill Out Visa Forms on State Dept. Website Tags: accidental-cyberattacks, anthropic, generative-ai, ai, llms

Simon Willison's Weblog原典の内容 · 翻訳・分析待ち翻訳待ち:Quoting The New York Times

翻訳待ち:Anthropic’s AI gave Philadelphia police a fake tip about an unsolved homicide

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:An Anthropic AI model provided false information about an unsolved homicide to a Philadelphia Police Department (PPD) tipline, according to a report from 6abc. In a statement released on Friday, the PPD said the AI model sent the tip through PhillyUnsolvedMurders.com on July 18th, but the investigators never reviewed it because it was marked as spam. Anthropic learned its AI model sent the false tip on September 28th and notified the PPD on October 7th. The company said that during testing, its AI model was interacting with "randomly selected websites" and submitted false information through the PPD's tipline, according to the PPD's stateme … Read the full story at The Verge.

The Verge AI原典の内容 · 翻訳・分析待ち翻訳待ち:Anthropic’s AI gave Philadelphia police a fake tip about an unsolved homicide

翻訳待ち:Alibaba Qwen Releases Qwen-Image-2.1-Turbo, an 8-Step 7B Image Model

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Alibaba’s Qwen team has released Qwen-Image-2.1-Turbo, an accelerated checkpoint of its open-weight Qwen-Image-2.1 model. It generates and edits images in 8 denoising steps instead of the base model’s 40-step default. For developers, that means 5x fewer denoising steps on the same 7B architecture, plus a hosted API option. TL;DR What is Qwen-Image-2.1-Turbo? Qwen-Image-2.1-Turbo is an […] The post Alibaba Qwen Releases Qwen-Image-2.1-Turbo, an 8-Step 7B Image Model appeared first on MarkTechPost.

MarkTechPost原典の内容 · 翻訳・分析待ち翻訳待ち:Alibaba Qwen Releases Qwen-Image-2.1-Turbo, an 8-Step 7B Image Model

翻訳待ち:OpenAI Decisions API Hits Public Beta With 10x Faster Typed Answers

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:OpenAI’s Decisions API is now in public beta on GPT-6 Luna. It returns typed probabilities, choices and scores about 10x faster than the Responses API, billing $0.10 per 1M input tokens with no output charges. The post OpenAI Decisions API Hits Public Beta With 10x Faster Typed Answers appeared first on MarkTechPost.

MarkTechPost原典の内容 · 翻訳・分析待ち翻訳待ち:OpenAI Decisions API Hits Public Beta With 10x Faster Typed Answers

翻訳待ち:Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We are expanding the Clef decision model family with Clef-omni, natively processing audio, video, images, and text in a single pipeline. We’ve also lowered Clef-flash pricing and boosted Clef inference speeds by up to 2.0x.

Cloudflare AI Blog原典の内容 · 翻訳・分析待ち翻訳待ち:Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash

翻訳待ち:Master AI Chip Principles With New IEEE Design Program

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Today’s engineers face an unprecedented acceleration in AI hardware complexity, as explained in the recent research article “Revisiting Edge AI: Opportunities and Challenges.” The article examines the rapid growth of edge AI and the challenges it creates, including resource constraints, model architecture limitations, and network demands across edge-AI deployments. The acceleration is driven by a fundamental shift in how modern AI models are built and scaled. As the models have become much larger and more complex, they are computationally more demanding because they contain more parameters and require more calculations. To meet the demands of scaling deep neural networks, the industry is increasingly developing AI chips that are designed for specifi…

IEEE Spectrum AI原典の内容 · 翻訳・分析待ち翻訳待ち:Master AI Chip Principles With New IEEE Design Program

翻訳待ち:Quoting Matthew Green

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要: Everyone is very concerned about being respectable, so I’m going to be the goofball who raises worst-case possibilities. I think there is a 1% chance we live in Minicrypt, and a 15% chance we functionally lose confidence in our existing public-key encryption algorithms. [...] The problem here is that the speed of AI producing surprises, and the speed of human beings replacing standards (even with the very best AI assistance) are just orders of magnitude different. You only recover from a surprise like this if you do the preparation in advance. — Matthew Green, on Twitter. I looked it up and Minicrypt is Russell Impagliazzo’s hypothetical world in which public-key encryption is impossible. Tags: ai-security-research, matthew-green, cryptograph…

Simon Willison's Weblog原典の内容 · 翻訳・分析待ち翻訳待ち:Quoting Matthew Green

翻訳待ち:EmbeddingGemma 2: Text, Code, Images, Video and Audio in One Vector Space

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:EmbeddingGemma 2 launched on October 6, 2026 under Apache 2.0. It is a sub-1B model built on Gemma 4 that maps text, code, images, video and audio into one 768-dimensional space. This article covers the architecture, the benchmarks, and runnable scripts to provide measured results. Specifications Specification EmbeddingGemma 2 Base model Gemma 4 License Apache […] The post EmbeddingGemma 2: Text, Code, Images, Video and Audio in One Vector Space appeared first on Analytics Vidhya.

Analytics Vidhya原典の内容 · 翻訳・分析待ち翻訳待ち:EmbeddingGemma 2: Text, Code, Images, Video and Audio in One Vector Space

翻訳待ち:A new feature for my, blog built using my voice

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要: I shipped a new feature for my blog today: the Newsletters page, which offers an index of all of the newsletters I've sent out, both my free weekly Substack and my monthly sponsors-only updates. I built the feature almost entirely using my voice, chatting away to my laptop while I cooked dinner. Codex voice mode I used the ChatGPT desktop app for this, in the Codex tab, using the voice conversation mode, running against a local development environment. Here's what that looks like: I started the session against my local simonwillisonblog checkout by typing: Start dev server and open in browser This gave me a preview of the site that it would be working on, and meant that I could later ask it to show me the new pages so I could visually track its pro…

Simon Willison's Weblog原典の内容 · 翻訳・分析待ち翻訳待ち:A new feature for my, blog built using my voice

翻訳待ち:Meet the Underdog Saluki 27B: A 2-bit Qwen3.8-27B That Beats the Original at Tool Calling

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Underdog Saluki 27B is a 7.89 GB, 2-bit GGUF of Qwen3.8-27B under Apache 2.0. It beats the 54 GB original on tool calling but gives up ground on competition math and reasoning. The post Meet the Underdog Saluki 27B: A 2-bit Qwen3.8-27B That Beats the Original at Tool Calling appeared first on MarkTechPost.

MarkTechPost原典の内容 · 翻訳・分析待ち翻訳待ち:Meet the Underdog Saluki 27B: A 2-bit Qwen3.8-27B That Beats the Original at Tool Calling

翻訳待ち:Asana cuts model costs 76x in browser tests with GPT-6.1 Sol

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Using GPT-6 Astra in Codex, Asana made its browser agent 76x cheaper and 5x faster in tests to offer customers more capable models.

OpenAI News原典の内容 · 翻訳・分析待ち翻訳待ち:Asana cuts model costs 76x in browser tests with GPT-6.1 Sol

翻訳待ち:“Don’t use ‘open weight’ and ‘open source’ interchangeably”: Percona CEO on why AI terminology matters

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The AI industry has enthusiastically embraced the language of open source, even as some of its most prominent “open” models The post “Don’t use ‘open weight’ and ‘open source’ interchangeably”: Percona CEO on why AI terminology matters appeared first on The New Stack.

The New Stack AI原典の内容 · 翻訳・分析待ち翻訳待ち:“Don’t use ‘open weight’ and ‘open source’ interchangeably”: Percona CEO on why AI terminology matters

翻訳待ち:Last Week in AI #346 - 719 math manuscripts, 2 Western open models, 1 more safety resignation

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:OpenAI publishes hundreds of math proofs from unreleased frontier model, Mistral and Reflection AI launch open-weight models to rival China, and more!

Last Week in AI原典の内容 · 翻訳・分析待ち翻訳待ち:Last Week in AI #346 - 719 math manuscripts, 2 Western open models, 1 more safety resignation

翻訳待ち:Cross-Embodiment Robot Foundation World Models with Latent Actions

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10846v1 Announce Type: new Abstract: The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments. We introduce the Latent Action-Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse embodiments. This unified action space improves the world model's performance when adapted to previously unseen robot embodiments. We compare LAC-WM with an Explicit Action-Conditioned World Model (EAC-WM), which conditions on explicit motion labels. Our results show that explicit action conditioning leads to disjoint action representations across embodiments, limiting downstream performance when adapting t…

arXiv Robotics原典の内容 · 翻訳・分析待ち翻訳待ち:Cross-Embodiment Robot Foundation World Models with Latent Actions

翻訳待ち:Skill-SLM: Agent Skill-driven Small Language Models for Reliable Robot Operation

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10812v1 Announce Type: new Abstract: Small language models (SLMs) have been increasingly adopted for onboard robot operation because they enable intelligent decision-making. However, existing approaches are mainly distillation-oriented and rely on enumerating representative task-solution pairs. This makes dataset construction difficult and limits generalization to diverse robot tasks whose possible forms grow rapidly. This paper proposes Skill-SLM, a framework that reformulates SLM-driven robot operation as a task-decomposition and skill-composition problem. Given a natural language task instruction, Skill-SLM decomposes the task into subtasks, selects appropriate skills from the skill library, and orchestrates the selected skills into ex…

arXiv Robotics原典の内容 · 翻訳・分析待ち翻訳待ち:Skill-SLM: Agent Skill-driven Small Language Models for Reliable Robot Operation

翻訳待ち:Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10810v1 Announce Type: new Abstract: Long-horizon robotic manipulation is often built by chaining independently trained skills. Although each skill can be reliable in isolation, performance degrades sharply when skills are chained: each downstream skill must start from the state its predecessor leaves behind rather than from its training distribution. We study this failure mode, Observation-Space Shift (OSS), and ask what causes these skill-seam failures. Using privileged simulator resets, we find that the dominant shift comes from displaced scene state (e.g., an open drawer or secondary objects left behind by earlier skills), not from the robot's joint configuration or the object the downstream skill manipulates. To test this diagnosis,…

arXiv Robotics原典の内容 · 翻訳・分析待ち翻訳待ち:Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams

翻訳待ち:NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10787v1 Announce Type: new Abstract: Language models trained with long-horizon agentic reinforcement learning can generalize knowledge through reasoning, express precise actions, and pursue goals over many steps, raising the ceiling on what an embodied agent can understand and decide. Physical interaction, however, remains the domain of action policies, which provide dense, low-latency control. We present NavGPT-3, a harness that connects the two models, with an OS-like runtime built above it: reasoning, acting, and monitoring run as threads with their own context, tools, and permissions, while the runtime schedules them and decides which thread controls the robot's motion, so that the robot can react to sudden real-world events through i…

arXiv Robotics原典の内容 · 翻訳・分析待ち翻訳待ち:NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime

翻訳待ち:Masked Generative Motion Planning with Geometry-Guided Token Search

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10646v1 Announce Type: new Abstract: Generative motion planners typically use learned trajectory priors for initial generation, while leaving test-time repair to local continuous refinement. We introduce Masked Generative Motion Planning (MGMP), which extends the learned prior from efficient parallel generation to structural repair. A masked generative transformer generates discrete trajectory candidates in parallel, and Geometry-Guided Token Search (GGTS) uses scene geometry to target where to edit and which prior-supported alternatives to evaluate. This turns refinement into an efficient search over discrete motion alternatives, enabling route-level restructuring beyond local trajectory deformation. MGMP achieves 96% success on Ring Maz…

arXiv Robotics原典の内容 · 翻訳・分析待ち翻訳待ち:Masked Generative Motion Planning with Geometry-Guided Token Search

翻訳待ち:Teaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task Selection

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10601v1 Announce Type: new Abstract: Reinforcement Learning (RL) has enabled legged robots to perform a range of skills in single-task settings. However, applications such as farm robotics or space exploration require diverse skills such as locomotion, digging, or close-range surveying. Training an end-to-end policy to address this problem remains difficult due to challenges such as sample inefficiency and gradient conflict between tasks in multi-task learning. We propose a three-stage method that trains a single policy to perform distinct tasks such as walking, digging, and hopping, and compose them into novel behaviors such as crawling. First, multiple teacher policies are trained using RL on narrowly defined tasks. Then, two additional…

arXiv Robotics原典の内容 · 翻訳・分析待ち翻訳待ち:Teaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task Selection

翻訳待ち:Does Dynamic-Point Filtering Help When Texture Is Scarce? A Controlled Study of ORB-SLAM2 Front-Ends in Synthetic Indoor Scenes

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10564v1 Announce Type: new Abstract: Dynamic-point filters are routinely added to feature-based visual SLAM, and several recent systems argue that removing dynamic features can leave too few static features in low-texture regions. So far, these systems have been evaluated only on texture-rich benchmark sequences. We present a controlled study that isolates this interaction. We render synthetic indoor sequences in which surface texture (four levels, quantified by FAST-corner density and image-gradient entropy) and scene dynamics (three levels) are varied factorially along identical camera trajectories, with stereo, RGB-D, ground-truth poses and dynamic masks. On this grid we compare ORB-SLAM2 without filtering, with an optical-flow and epi…

arXiv Robotics原典の内容 · 翻訳・分析待ち翻訳待ち:Does Dynamic-Point Filtering Help When Texture Is Scarce? A Controlled Study of ORB-SLAM2 Front-Ends in Synthetic Indoor Scenes

翻訳待ち:SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10889v1 Announce Type: new Abstract: Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR). Reinforcement learning with verifiable rewards typically trains both through a single chain-of-thought with a final-answer reward. This gives every CoT token the same sequence-level advantage, failing to distinguish capability specific errors. We propose SPLIT-RL, a staged post-training approach that trains VR and LR in disjoint phases. Because a group's rollouts differ along one capability at a time, the group-relative advantage isolates it, and each phase is optimized using phase-specific reward. We further in…

arXiv Computer Vision原典の内容 · 翻訳・分析待ち翻訳待ち:SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages

翻訳待ち:Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10859v1 Announce Type: new Abstract: Text-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference. However, their ability to generate harmful or undesired content poses significant safety risks. Data-driven unlearning methods suppress targeted generations by fine-tuning model weights using specialized unlearning objectives. Crucially, these objectives implicitly rely on multi-step denoising dynamics, an assumption that breaks down for few-step distilled (FSD) models, resulting in ineffective forgetting. Furthermore, performing unlearning on the non-distilled base model and subsequently re-distilling it to obtain an unlearned FSD model incurs substantial computational and time…

arXiv Computer Vision原典の内容 · 翻訳・分析待ち翻訳待ち:Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models

翻訳待ち:VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10782v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs), but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many become trivial, others remain unsolvable; and the learning signal collapses. We argue that VLM post-training should evolve the visual environment alongside the actor, not just the actor itself. We propose VICO, a co-evolutionary framework in which an actor and an Environment-as-Rewriter (EnvRewriter) are trained jointly: the EnvRewriter edits verifiable image-side structures, such as scene graphs, chart tables, or protected region masks, and r…

arXiv Computer Vision原典の内容 · 翻訳・分析待ち翻訳待ち:VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning

翻訳待ち:SLVR: Structured Latent Visual Reasoning via Human-like Reasoning Flows

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10563v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) often answer visual reasoning questions by relying on linguistic priors rather than task-relevant visual evidence. Textual chain-of-thought reasoning can partially mitigate this issue by encouraging models to decompose visual questions into intermediate evidence-seeking steps, but generating these steps autoregressively increases inference cost. Latent reasoning avoids explicit rationale generation, but existing approaches provide limited control over what intermediate states encode, making it difficult to impose separate supervision for planning, grounding, and evidence selection. We propose Structured Latent Visual Reasoning (SLVR), a training framework that b…

arXiv Computer Vision原典の内容 · 翻訳・分析待ち翻訳待ち:SLVR: Structured Latent Visual Reasoning via Human-like Reasoning Flows

翻訳待ち:Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10871v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong performance across many domains, but their efficiency is limited by the quadratic cost of attention with respect to prompt length. Sparse attention reduces this cost by retaining only a small fraction of query-key interactions to approximate the full attention matrix. However, existing methods are trapped in a mathematically wrong view: they simply keep large scalar entries or high-mass regions of the attention matrix. This treats the attention matrix as a bag of values, ignoring that it is used as a structured matrix whose entries jointly determine the attention output through multiplication with value vectors. We argue that this is the core conceptual issue…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values

翻訳待ち:Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10845v1 Announce Type: new Abstract: A large language model can only use the text that fits in its context window, and it recomputes its internal key-value (KV) state for a prompt every time the prompt is sent. We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it. We ran it on 50,000,000 tokens of real public text, served through vLLM on one NVIDIA H100, with Gemma 4 12B and Gemma 4 31B. Every block we probed was loaded back from the encrypted store with no recompute (100 of 100, at depths from 0 to 50M tokens) on both models. Loading a block was 2.8x to 4.3x faster than recomputing it and…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute

翻訳待ち:Grammar Concept Annotation at Scale: Deployed Fine-Tuned Small Language Models Outperform Prompted Frontier Models

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10827v1 Announce Type: new Abstract: Corrective feedback is among the best-evidenced drivers of second-language acquisition, yet corrections delivered during lessons rarely accumulate into an actionable view of grammar mastery. Prompted frontier models can provide such a view from learner--tutor lesson transcripts, but they are costly at scale. We close this gap by fine-tuning Qwen3.5 small language models (SLMs) on filtered and rebalanced teacher-generated supervision, then deploying an efficient 0.8B model in an end-to-end grammar mastery tracker for all English learners on our platform. Internalizing the annotation contract into adapter weights enables pairing the 0.8B model with a compact matched prompt rather than verbose instruction…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:Grammar Concept Annotation at Scale: Deployed Fine-Tuned Small Language Models Outperform Prompted Frontier Models

翻訳待ち:Lossy Compressive Text Autoencoders

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10738v1 Announce Type: new Abstract: Our work explores learning a compressed latent representation of text, at the intersection of data compression and representation learning. We propose an autoencoder architecture that performs residual downscaling and upscaling of hidden representations along the time axis, with a residual low-dimension discrete bottleneck. We analyze our approach for different quantization methods, training objectives, and datasets. For different levels of compression, we evaluate the similarity between the original and reconstructed text both at the surface-level (BLEU) and at the semantic-level (LLM-based judge). Additionally, we evaluate our models on downstream question-answering and semantic text similarity bench…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:Lossy Compressive Text Autoencoders

翻訳待ち:Large Language Model-Assisted Preparation of Transportation Management Plans: A Case Study with WisDOT WisTMP System

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10650v1 Announce Type: new Abstract: Work zones are critical yet hazardous components of transportation infrastructure, requiring carefully designed Transportation Management Plans (TMPs) to ensure safety and mobility. However, TMP preparation remains labor-intensive and heavily dependent on practitioner expertise. This paper proposes a Large Language Model (LLM)-assisted framework to automate TMP content generation, leveraging the WisDOT WisTMP system as the application context. The framework fine-tunes multiple open-source LLMs across different model scales and deploys them locally to ensure data security. To support model training, we construct a domain-specific dataset from historical WisTMP documents by converting PDF files into stru…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:Large Language Model-Assisted Preparation of Transportation Management Plans: A Case Study with WisDOT WisTMP System

翻訳待ち:Wieszcz-XIX: A 3.1-Billion-Word Corpus of Pre-1918 Polish and Temporally Bounded Language Models Trained From Scratch

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10592v1 Announce Type: new Abstract: Historical Polish is well documented as a language but annotated in machine-readable form only to about a million words for the period this paper covers; the rest sits behind optical character recognition of variable quality. We present Wieszcz-XIX, a corpus of 6.75 billion tokens (about 3.1 billion words) in 294,369 documents, most of them periodical issues, of Polish published from 1800 to 1918, assembled from Wolne Lektury and the Internet Archive by a pipeline that filters, deduplicates, audits for post-1918 leakage and splits at the document level. It is over three orders of magnitude larger than the annotated corpus of the same period, and we quantify its defects: recognition corruption against a…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:Wieszcz-XIX: A 3.1-Billion-Word Corpus of Pre-1918 Polish and Temporally Bounded Language Models Trained From Scratch

翻訳待ち:Diffu-LoRA: A Novel Low-Rank Adaptation for Personalized Diffusion Models

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10550v1 Announce Type: new Abstract: Personalizing text-to-image diffusion models from a few reference images requires preserving subject identity while following prompts that describe new contexts. Full-model fine-tuning is parameter-intensive, whereas low-rank adaptation (LoRA) reduces the number of trainable parameters but leaves open how adaptation capacity should be distributed across layers. We introduce Diffu-LoRA, a parameter-efficient method that learns this allocation through gated low-rank adaptation. Diffu-LoRA inserts trainable low-rank components into the linear layers of Transformer blocks and assigns a learnable gate to each component. Bilevel optimization updates the adaptation weights and gate parameters on separate data…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:Diffu-LoRA: A Novel Low-Rank Adaptation for Personalized Diffusion Models

翻訳待ち:Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language Models

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10623v1 Announce Type: new Abstract: Looped Language Models (LoopLMs) offer a parameter efficient approach to scaling reasoning by reusing shared parameters across recurrent computation steps. Despite their promise, effective post-training of LoopLMs remains challenging. Existing approaches either provide reward based supervision that is sparse or costly to extend across loops, or rely on external teachers or privileged information, leading to limited teacher availability or teacher-student context mismatch. To address these limitations, we introduce LoopOPD, a cross-loop on-policy distillation framework that uses additional recurrent computation within a LoopLM as its own source of supervision. LoopOPD uses a frozen terminal loop policy…

arXiv Machine Learning原典の内容 · 翻訳・分析待ち翻訳待ち:Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language Models

翻訳待ち:When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10616v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) language models produce routing information during inference that may be logged or exposed for monitoring, debugging, load analysis, and safety auditing. Unlike ordinary model outputs, this telemetry reveals a view of the model's internal computation, raising a privacy question: can it reveal whether an example was used to fine-tune the deployed model? We introduce a router-augmented membership inference attack that combines conventional output-side signals with aggregated routing features and applies a membership classifier learned from independently fine-tuned shadow models to the target model. Across three MoE architectures and three data domains, router telemetry consistent…

arXiv Machine Learning原典の内容 · 翻訳・分析待ち翻訳待ち:When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry

翻訳待ち:Temporal transformer CAN encoder with federated lightweight heads for anomaly detection

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10613v1 Announce Type: new Abstract: Modern vehicles rely on large numbers of Electronic Control Units (ECUs) that constantly exchange information over the Controller Area Network (CAN) bus. Due to the rapidity, structure, and repetition of this communication, even slight variations in timing, payload values, or message patterns can point to unusual activity. Whether due to errors, malfunctions, or deliberate interference, these anomalies are frequently subtle and challenging to identify with conventional methods that handle messages separately or rely on manually created rules. Motivated by this gap, we present a privacy-preserving framework for anomaly detection in in-vehicle networks, based on a Temporal Transformer CAN Encoder with Fe…

arXiv Machine Learning原典の内容 · 翻訳・分析待ち翻訳待ち:Temporal transformer CAN encoder with federated lightweight heads for anomaly detection

翻訳待ち:Coverage, Not Difficulty, Sets How Much Synthetic Data an Activation Probe Needs

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10594v1 Announce Type: new Abstract: Activation probes that monitor deployed language models are trained on synthetic conversations, and how many a probe needs is open. We trace learning curves over 10-590 synthetic samples for three monitoring concepts, high-stakes situations, replies harmful to a person, and replies that do not follow the user's instruction, on fourteen held-out evaluation distributions and four probe models, varying the generator LLM and the prompt's detail. The need is set by what is monitored: probes for high-stakes and harmful are within a few hundredths of their plateau from 80 samples on Gemma-3-27B-IT, instruction probes need several times as many, and the ordering holds on three smaller probe models and on real…

arXiv Machine Learning原典の内容 · 翻訳・分析待ち翻訳待ち:Coverage, Not Difficulty, Sets How Much Synthetic Data an Activation Probe Needs

翻訳待ち:SPERA: Spherical Prior EEG Foundation Model with Geometry- and Frequency-Aware Latent Prediction

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10571v1 Announce Type: new Abstract: Electroencephalography (EEG) provides a non-invasive measure of ongoing neural activity, but building general-purpose EEG models remains challenging due to the heterogeneity of subjects, devices, and electrode montages. Existing EEG foundation models predominantly rely on reconstruction-based objectives defined on the observed signal, which contains both neural and non-neural components. We introduce SPERA (Spherical Prior EEG Representation Architecture), an EEG foundation model that adopts the joint-embedding predictive architecture (JEPA) to predict in latent space. SPERA introduces a Legendre-polynomial spatial prior, incorporated into attention to encode varying scalp electrode geometries. Two fur…

arXiv Machine Learning原典の内容 · 翻訳・分析待ち翻訳待ち:SPERA: Spherical Prior EEG Foundation Model with Geometry- and Frequency-Aware Latent Prediction

翻訳待ち:Freeze the Decoder, Heal the Encoder: Parameter-Efficient Adaptation for SVD-Based KV-Cache Compression

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10552v1 Announce Type: new Abstract: Comparing parameter-efficient fine-tuning recipes under a single, shared learning rate is a common but flawed practice: when the arms being compared have very different trainable-parameter counts, a shared rate can simultaneously depress the larger arms' means and inflate their variance, manufacturing a large, seemingly multi-seed-significant advantage for the smallest arm that is not a real effect. We document this confound in a concrete setting: post-hoc SVD-based KV-cache compression, where an already-pretrained model is converted to a low-rank (multi-head-latent-attention-style) cache by factorizing its key/value weights into a down-projection ("encoder") and an up-projection ("decoder"), after whi…

arXiv Machine Learning原典の内容 · 翻訳・分析待ち翻訳待ち:Freeze the Decoder, Heal the Encoder: Parameter-Efficient Adaptation for SVD-Based KV-Cache Compression

翻訳待ち:On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10833v1 Announce Type: new Abstract: We study whether small LLM agents can operate effectively under explicit wall-clock time budgets by both respecting the allocated runtime and using available time productively. We evaluate Qwen3.6-27B on five competitions from MLE-Bench Lite and Qwen3-4B on Zork I (Jericho), two agentic benchmarks where additional computational time can meaningfully improve performance. In the simplest setting, where the budget is stated only in the prompt, agents fail to translate the stated budget into controlled use of time. These failures arise from gaps in time awareness, since the harness provides no timing feedback, but also because they cannot reliably anticipate the duration of actions, and do not have a learn…

arXiv AI原典の内容 · 翻訳・分析待ち翻訳待ち:On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents

翻訳待ち:Plan-and-Patch: Diffusion Language Models for Agentic Planning

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10786v1 Announce Type: new Abstract: Planning is increasingly important for long-horizon agents, where successful execution requires coordinating subgoals, tool use, and intermediate outcomes over many steps. Yet assumptions made during planning may be invalidated by the environment, tools may return unexpected results, or actions may fail. Effective agents must therefore not only generate plans, but also revise them. Such revisions often affect only part of a plan, leaving the preceding and subsequent structure intact. Rather than regenerate the entire plan and risk unnecessary changes, repair can regenerate the affected region conditioned on the preserved prefix and suffix. We introduce Plan-and-Patch, a plan-and-act framework in which…

arXiv AI原典の内容 · 翻訳・分析待ち翻訳待ち:Plan-and-Patch: Diffusion Language Models for Agentic Planning

翻訳待ち:Speaking the Navigator's Language: Trajectory-Grounded Instruction Translation for Frozen Aerial VLN Agents

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10635v1 Announce Type: new Abstract: Aerial vision-and-language navigation (VLN) agents are typically trained on detail-rich, trajectory-aligned commands, whereas users issue short, intent-driven instructions; on a frozen OpenFly navigator, this \emph{instruction gap} drops success rate (SR) from $31.03\%$ to $11.33\%$. To scale translator training, we prompt a language model with human-written style examples to convert original commands into paired, intent-centered Weak commands, which yield $15.27\%$ SR. We introduce the \textbf{Trajectory-Grounded Instruction Translator (TGIT)}, a front-end that keeps the navigator frozen and translates Weak inputs into agent-executable commands by learning from its trajectory outcomes. The resulting W…

arXiv AI原典の内容 · 翻訳・分析待ち翻訳待ち:Speaking the Navigator's Language: Trajectory-Grounded Instruction Translation for Frozen Aerial VLN Agents

翻訳待ち:The Harness as the Only Mutable Surface: Compliance-Bounded Self-Evolution of LLM Agents in Credit Pipelines, with a Measured Admission Gate

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10629v1 Announce Type: new Abstract: Self-improving LLM agents can adapt a credit pipeline to a changed rule, but an agent that rewrites itself destroys the artefact a supervisor reviews: a named change, a recorded test, an approval. We argue that self-evolution is reviewable only if it is confined to the runtime harness (instruction text, tool-call logic and primitive composition) while model weights stay fixed, so that every adaptation is a diff with a cause and a test attached. We give a dual-loop engine built on that bound, with one admission gate that writes a hash-chained record before deployment, and we measure the gate in simulation, with a simulated agent and a seeded-search proposer rather than language models. Across three fami…

arXiv AI原典の内容 · 翻訳・分析待ち翻訳待ち:The Harness as the Only Mutable Surface: Compliance-Bounded Self-Evolution of LLM Agents in Credit Pipelines, with a Measured Admission Gate

翻訳待ち:Synthesis Through Simulation: Generating Coherent Enterprise Data via Scalable Agent-System Interaction

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2610.10549v1 Announce Type: new Abstract: Tool-calling agents have become central to enterprise AI, yet training and evaluating them at scale remains severely constrained due to business and legal restrictions on enterprise systems, data, and database schemas. Tabular data synthesis offers a natural alternative, but its effectiveness is fundamentally limited by structural validity and schema availability, while procedure-based approaches yield the opposite weakness, typically lacking distributional fidelity without per-domain authoring. We introduce **Synthesis Through Simulation** (STS), a **schema--free** data synthesis paradigm in which an LLM agent generates data by executing operations against policy-enforcing APIs within simulated enterp…

arXiv AI原典の内容 · 翻訳・分析待ち翻訳待ち:Synthesis Through Simulation: Generating Coherent Enterprise Data via Scalable Agent-System Interaction

翻訳待ち:ttok 1.0

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要: Release: ttok 1.0 I released ttok 0.4, ran uv tool upgrade ttok, piped a file into the new version... and realized that it was defaulting to the GPT-4 tokenizer when it should very clearly default to GPT-5/GPT-6 instead! I figured switching the default was a reasonable excuse to finally ship a 1.0. OpenAI haven't actually confirmed that GPT-6 uses the same tokenizer as the GPT-5 family yet - there's an angry issue about it - but I found this commit by William Liu which reports on an experiment he ran confirming that the tokenizers are likely the same: All seven GPT models (5.5, 5.6 Sol/Terra/Luna, 6 Astra/Sol/Luna) report 44,794 tokens and match each other on every one of the 31 fixtures. GPT-6 introduces no input-count change on this corpus. Tags:…

Simon Willison's Weblog原典の内容 · 翻訳・分析待ち翻訳待ち:ttok 1.0

翻訳待ち:Western open-weight campaign could give enterprises more control over AI

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:New Western open-weight models could give enterprises an alternative to Chinese and proprietary AI, but greater control also brings new responsibilities.

AI Business原典の内容 · 翻訳・分析待ち翻訳待ち:Western open-weight campaign could give enterprises more control over AI

トピック

モデル AI ニュース | AI News Hub