本文にスキップ
AI News HubLIVE

ソース分布

  • arXiv Computational Linguistics9
  • MarkTechPost7
  • AI Business5
  • arXiv AI5
  • Hacker News AI4
  • Simon Willison's Weblog4
  • NVIDIA Blog3
  • arXiv Machine Learning2

トピック分布

  • モデル43
  • 研究27
  • Agent24
  • チップ6
  • スタートアップ5
  • ツール4
  • 政策4
  • ロボット1

タイムライン

  • 2026-08-253
  • 2026-09-013
  • 2026-10-073
  • 2026-07-132
  • 2026-07-142
  • 2026-08-132
  • 2026-08-202
  • 2026-09-092

最新動向

翻訳待ち:Last Week in AI #346 - 719 math manuscripts, 2 Western open models, 1 more safety resignation

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:OpenAI publishes hundreds of math proofs from unreleased frontier model, Mistral and Reflection AI launch open-weight models to rival China, and more!

Last Week in AI原典の内容 · 翻訳・分析待ち翻訳待ち:Last Week in AI #346 - 719 math manuscripts, 2 Western open models, 1 more safety resignation

Mistral Large 4「Le chonk」登場

Mistral が Mistral Large 4 のプレビュー版を公開した。総パラメータ 1 兆、アクティブパラメータ 490 億で、自社の NVIDIA Grace Blackwell GPU 3,800 基クラスタで学習されている。API 経由のプレビューはすでに利用可能で、オープンウェイト版は今月末に公開予定。Artificial Analysis のスコアは 38 で、Mistral Large 3 の 9 から大幅に改善したものの、最先端からは依然として約 6 か月遅れている。

Simon Willison's Weblogサイト内本文Mistral Large 4「Le chonk」登場

Mistral Large 4:網タイツ姿のアルマジロを火星で描かせてみた

Simon Willison が Hacker News の Mistral Large 4 に関する議論に寄せたコメント。ベンチマーク飽和という皮肉をきっかけに、4つのフロンティアモデルへ同一のナンセンスなSVG生成プロンプトを実際に投げている。

Simon Willison's Weblogサイト内本文Mistral Large 4:網タイツ姿のアルマジロを火星で描かせてみた

Mistral AI、Mistral Large 4(Le Chonk)を公開:1.05兆パラメータのマルチモーダルMoEモデル

Mistral AIはMistral Large 4(内部コードネームLe Chonk)をパブリックプレビューとして公開した。細粒度のMoEで、総パラメータ1.05兆、トークンあたり49Bを活性化、16億パラメータの視覚エンコーダと100万トークンのコンテキストを備える。欧州の自社データセンターでNVIDIA Grace Blackwell GPU 3,800基を用いてゼロから学習された。APIは稼働中で、入力/出力100万トークンあたり1.36/4.18ドル、キャッシュ入力は0.14ドル。重みとライセンスは2026年10月末に公開予定で、セルフホストはまだできない。最も際立つのはサイバーセキュリティで、Cybench 93%、CyberGym-E2E 82%。Mistralは複数のクローズドなフロンティアモデルがタスクを拒否してほぼゼロ点だと主張している。

MarkTechPostサイト内本文Mistral AI、Mistral Large 4(Le Chonk)を公開:1.05兆パラメータのマルチモーダルMoEモデル

翻訳待ち:Building an Enterprise AI Customer Support Platform with Claude Fable 5.1 and Claude Code

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Most AI coding demos stop at task managers, weather apps, or simple chatbots. For this project, we take on something more demanding: building an enterprise customer-support platform that can investigate complaints, retrieve relevant policies, recommend resolutions, and keep risky actions behind human approval. This gives us a practical way to test Claude Fable 5.1 as […] The post Building an Enterprise AI Customer Support Platform with Claude Fable 5.1 and Claude Code appeared first on Analytics Vidhya.

Analytics Vidhya原典の内容 · 翻訳・分析待ち翻訳待ち:Building an Enterprise AI Customer Support Platform with Claude Fable 5.1 and Claude Code

翻訳待ち:How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users. Accelerated by inference optimizations through OpenAI’s models that tap into the capabilities of the NVIDIA Blackwell architecture, Ultrafast offers up to 8x faster token generation than the Astra Standard mode. For developers, […]

NVIDIA Blog原典の内容 · 翻訳・分析待ち翻訳待ち:How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast

翻訳待ち:Forget ‘superintelligence’: error-prone AI nearly sparked world war three this month | Timnit Gebru and Emily M Bender

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The risks of AI aren’t what we think they are, as a recent security incident between China and the United States reveals Amid a barrage of news stories warning about superintelligent machines rendering humanity extinct, a CNN story describing the opposite scenario – one in which the US military’s reliance on brittle chatbots almost brought the US into war with China – went mostly unnoticed by the public. The biggest international AI news of the past three weeks was Anthropic engineer Jacob Coxon’s resignation. According to him, OpenAI and Anthropic are “racing straight towards self-improving superintelligence and gambling with our lives”. Coxon’s description of a “terminator” scenario, a machine becoming much smarter than humanity and deciding to wi…

The Guardian AI原典の内容 · 翻訳・分析待ち翻訳待ち:Forget ‘superintelligence’: error-prone AI nearly sparked world war three this month | Timnit Gebru and Emily M Bender

同じ量、異なる答え:言語モデルにおける数値表現不変性

新しい論文が、数値的に等価な書き換えに対してオープンウェイト言語モデルが同じ正準解を返せるかを検証。3,600 の厳密有理数問題と 8,600 のプロンプト、5 つの変換ファミリーで 5 モデルを評価した結果、正準精度は 0.969–0.996 だが、オービット正解率は 0.848–0.981、オービット不変性は 0.851–0.981 に低下した。厳格パーサーの崩壊の多くは、乗算形の科学表記法が評価器の数値文法外にあることが原因で、推論失敗に見える評価インターフェースの問題だと指摘。Mistral Small 4 は単位変換入力で 0.699 となり、ラベルとちょうど 10 の冪だけ異なる 265 件の誤りを出した。9,000 回呼び出しの別実験では、表現コンセンサスは言い換えコンセンサスを上回らず、誤警報が大幅に増えた。

arXiv Computational Linguisticsサイト内本文同じ量、異なる答え:言語モデルにおける数値表現不変性

翻訳待ち:Why Read a Research Paper When You Can Turn It Into an AI Agent?

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Have you ever read a paper in Science or Nature and thought, “Man, that research was so cool. I wish I could try that method on my own data”—only to spend a week wrestling with someone else’s undocumented repo, broken dependencies, and half-finished readme.txt? Well, now you can, more or less. Say hello to Paper2Agent, a new open-source framework that transforms academic reports into interactive AI agents you can talk to. Give it a paper along with the accompanying codebase, data or other supplementary material, and the system automatically extracts the core workflows, then spins up a tested, runnable toolkit that you can use on your own datasets. The concept may sound a little like Google’s NotebookLM (now called Gemini Notebook), which lets you up…

IEEE Spectrum AI原典の内容 · 翻訳・分析待ち翻訳待ち:Why Read a Research Paper When You Can Turn It Into an AI Agent?

翻訳待ち:From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.17538v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for structured information extraction from documents, yet their behavior under realistic OCR noise remains poorly understood. We present a systematic benchmark of open-source instruction-tuned LLMs for key-value pair (KVP) extraction under both clean-text and noisy OCR conditions. We evaluate representative decoder-only models (Gemma, Mistral, Qwen2.5, LLaMA 3, and DeepSeek) on the FUNSD, CORD, and SROIE benchmarks using both Gold-text annotations and OCR outputs from PaddleOCR, EasyOCR, and Tesseract. A unified evaluation protocol isolates the effects of input quality, model design, and prompting under consistent conditions. The results show that mode…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings

翻訳待ち:Mistral Bets Enterprise AI Will Be About Control, Not Just Intelligence

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The French AI lab is using a $3B fundraise to sell control over AI infrastructure, not just model power -- a shift in direction that could matter to U.S. firms in Europe too.

AI Business原典の内容 · 翻訳・分析待ち翻訳待ち:Mistral Bets Enterprise AI Will Be About Control, Not Just Intelligence

翻訳待ち:OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.09203v1 Announce Type: new Abstract: Existing benchmarks for autonomous AI scientists evaluate only final outputs---generated code, hypotheses, or papers---yet discard the reasoning process by which those outputs were obtained. This makes it impossible to audit scientific methodology, diagnose failure modes, or distinguish systematic reasoning from fortunate guessing. We present \textbf{OpenDiscoveryTrace}, a public dataset of 558 complete AI scientific agent trajectories that captures how models reason, not just what they produce. Each trajectory records a structured 9-field-per-step trace---including thoughts, tool calls, observations, errors, revision triggers, and self-reported confidence---as models execute 124 scientific tasks spann…

arXiv AI原典の内容 · 翻訳・分析待ち翻訳待ち:OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

翻訳待ち:Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:This week, Mistral announced it raised €3 billion in a Series D funding round, pushing its post-money valuation past €21 The post Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it. appeared first on The New Stack.

The New Stack AI原典の内容 · 翻訳・分析待ち翻訳待ち:Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

翻訳待ち:[AINews] OpenAI reports Navier-Stokes singularity find in 88 hours using Astra-next, roughly 10,000 agents and 130B tokens (>$40M), a contender for second ever Millennium Prize awarded

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Overshadowing Cognition's $48B Series E, Mistral's $24B Series D, Meta's Muse agent, and GPT Image 2.5. The most jam packed, feel the AGI day in the history of AI.

Latent Space原典の内容 · 翻訳・分析待ち翻訳待ち:[AINews] OpenAI reports Navier-Stokes singularity find in 88 hours using Astra-next, roughly 10,000 agents and 130B tokens (>$40M), a contender for second ever Millennium Prize awarded

Mistralの新規資金調達、Sovereign AIへの架け橋に

MistralはオープンウェイトのAIスタートアップとして始まったが、欧州市場に対応するため、現在はSovereign AI(主権AI)へと戦略をシフトしている。今回の資金調達は、その移行を支える架け橋になると見られている。

AI Businessサイト内本文Mistralの新規資金調達、Sovereign AIへの架け橋に

翻訳待ち:Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.28623v1 Announce Type: new Abstract: Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating explicit chain-of-thought reasoning before answering. In language models it has been observed that this performance often comes with sycophancy, the tendency of a model to agree with the user over the evidence. However, for LMRMs no reliable method to measure sycophancy yet exists. We bridge this gap by introducing a benchmark and dataset for evaluating LMRM sycophancy when confronted with a wrong answer from a user. Our benchmark pairs four visually grounded datasets spanning mathematical, clinical, temporal, and demographic reasoning with five pressure conditions in single-turn and multi-turn sett…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure

翻訳待ち:SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.28911v1 Announce Type: new Abstract: The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length. We show that uniform KV quantization on a fractional-bit grid does not degrade gracefully: under a prespecified multi-seed statistical protocol, Llama-3.1-8B-Instruct with an affine quantizer is statistically indistinguishable from FP16 KV down to 2.322 code bits/value and collapses at 2.0 bits - a quality cliff in (2.0, 2.322] that reappears in generation-time quantization and multi-turn dialogue and transfers to Mistral-7B. The cliff reframes importance-aware mixed precision: above it, eight model-internal importance indicators are statistically interc…

arXiv Machine Learning原典の内容 · 翻訳・分析待ち翻訳待ち:SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

翻訳待ち:How law firm Gilbert + Tobin governs and scales AI with OpenAI

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:See how Gilbert + Tobin combines CEO-led commitment, rigorous governance, and human accountability to scale ChatGPT Enterprise and Codex across the firm.

OpenAI News原典の内容 · 翻訳・分析待ち翻訳待ち:How law firm Gilbert + Tobin governs and scales AI with OpenAI

翻訳待ち:Understanding ChatGPT Work

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要: OpenAI announced ChatGPT Work on July 9th, and have been furiously iterating on it ever since. It is an extraordinarily confusing and very powerful product. Here's what I've figured out about it so far. ChatGPT Work is actually two products The more interesting version of ChatGPT Work is the one that runs in the cloud. This can be accessed via chatgpt.com or through the ChatGPT mobile apps. Let's call it Work Cloud. If you install the ChatGPT desktop app - the app that used to be called Codex - you gain access to a thing called ChatGPT Work that can access files and run programs directly on your computer. Let's call that one Work Local. This one feels more like regular Codex re-skinned to be less intimidating to non-software-developers. For the res…

Simon Willison's Weblog原典の内容 · 翻訳・分析待ち翻訳待ち:Understanding ChatGPT Work

翻訳待ち:Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Cohere has released Parse (parse-v5.0), a 2.3B-parameter vision language model that converts PDFs, slides and images into Markdown with HTML tables, bounding boxes and image descriptions. It runs at $1.50 per 1,000 pages through the API, or on dedicated Model Vault instances from $2,500 a month. Cohere reports a ParseBench score of 79.2, ahead of Mistral OCR 4, Azure Document Intelligence and Databricks AI Parse — but that figure averages three of the benchmark's five dimensions and drops charts and visual grounding entirely. The post Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown appeared first on MarkTechPost.

MarkTechPost原典の内容 · 翻訳・分析待ち翻訳待ち:Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown

翻訳待ち:Mistral and Saudi Vendor to Advance Sovereign AI in Middle East

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The French AI lab extends its push for regional control of AI from Europe to the Middle East.

AI Business原典の内容 · 翻訳・分析待ち翻訳待ち:Mistral and Saudi Vendor to Advance Sovereign AI in Middle East

翻訳待ち:Comparing Local Tool Calling: Gemma 4 vs. Llama 3 vs. Mistral

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:In this article, you will learn how Gemma 4, Llama 3, and Mistral implement tool calling locally, and what trade-offs each model family presents for...

Machine Learning Mastery原典の内容 · 翻訳・分析待ち翻訳待ち:Comparing Local Tool Calling: Gemma 4 vs. Llama 3 vs. Mistral

翻訳待ち:Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. Why? Consider what happens when an AI agent researches a company for an investment decision. The agent queries financial databases, searches news and filings, invokes a sub-agent to run peer comparisons and model valuations, then synthesizes everything into a […]

NVIDIA Blog原典の内容 · 翻訳・分析待ち翻訳待ち:Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

翻訳待ち:Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.18090v1 Announce Type: new Abstract: Inside a modern language model sits a single internal direction that tracks how positive or negative a sentence feels. We show how to find this valence axis (V-axis) from just 9 emotion category names plus 50 short narrative paragraphs per emotion -- about 1,500 fewer labels than the usual supervised approach -- and that the same direction appears in vision, audio, and human-brain encoders never jointly trained. The recipe: embed nine emotion-anchored story sets in a frozen encoder, take the top principal direction of the nine averaged embeddings. Projecting new inputs onto it captures 93% of supervised performance on SST-2 (Llama-3-8B-Instruct, AUC 0.772 vs. 0.828), correlates with human valence ratin…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities

翻訳待ち:Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.18089v1 Announce Type: new Abstract: Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hausa. This suggests that the refusal mechanism is present in the residual stream but fails to activate for low-resource inputs. Recovering it normally requires labelled target-language data and retraining, neither of which is available at scale for most African languages. We introduce Latent Space Refusal Anchoring (LSR-Anchoring), a training-free method that extracts the refusal direction from English prompts and clamps it onto the residual stream at inference time. The primary variant, Mean-Activation Steering (MAS), operates across the four architectures we tested: Llama-3…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining

翻訳待ち:Mistral AI Strategy

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Mistral AI wants to turn European AI sovereignty from a talking point into a product — one with a service-level agreement attached. The French artificial intelligence company announced Tuesday a three-part expansion of…

Hacker News AI原典の内容 · 翻訳・分析待ち翻訳待ち:Mistral AI Strategy

翻訳待ち:Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.11210v1 Announce Type: new Abstract: Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers almost always fall back on uniform priors. The main reason is that building informative priors from scientific literature is slow and needs both domain and statistical expertise. We present \textbf{Distribird}, an agentic web application that automates this process. Given a parameter name, physical description, and domain context, Distribird deploys a multi-agent pipeline that searches the literature, extracts and weights reported values by domain relevance, and fits a probability distribution via AIC model selection. When no literature is available, t…

arXiv AI原典の内容 · 翻訳・分析待ち翻訳待ち:Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration

翻訳待ち:Mistral Aims to Build 1GB of Compute Capacity by 2030

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The Paris-based vendor continues to build European AI infrastructure.

AI Business原典の内容 · 翻訳・分析待ち翻訳待ち:Mistral Aims to Build 1GB of Compute Capacity by 2030

翻訳待ち:ChatGPT and Gemini both just passed 1 billion users

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:That’s a lot of people chatting with their AI friends all day. | Image: Google For the 14th time, a Google product has hit 1 billion users. Google CEO Sundar Pichai posted on X that a billion people are using Gemini every month, and that Gemini is Google's fastest-growing product ever. A billion users is a huge milestone, but Google isn't the first AI app to hit it. OpenAI's ChatGPT hit the mark a few weeks ago, though the company buried the announcement that "more than 1 billion people are putting ChatGPT to work" in an otherwise anodyne blog post about how people use AI. External data suggested ChatGPT crossed 1 billion users as early as this June, but OpenAI hadn't announced anything until that post on August 6th. … Read the full story at The Ver…

The Verge AI原典の内容 · 翻訳・分析待ち翻訳待ち:ChatGPT and Gemini both just passed 1 billion users

翻訳待ち:Mistral AI Releases Shieldstral 1.0 3B: An Open-Weights Policy-Adaptive Multimodal Safety Classifier Matching Models 7× Its Size

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Mistral AI has released Shieldstral 1.0 3B, an open-weights, policy-adaptive multimodal safety classifier that frames content moderation as a single yes/no question instead of a fixed harm taxonomy. Operators supply the policy as a plain-language query at inference time and get back a calibrated safety score from one forward pass — no retraining required to re-target the model. Built on Ministral-3-3B-Base-2512 with a Pixtral vision encoder and trained on roughly 54.1M samples, it reports 84.9% average F1 on text safety (matching GPT-OSS-Safeguard-20B), 83.8% on multimodal safety, and 91.3% on Mistral's adaptability benchmark — while fitting in 16GB of VRAM under an Apache 2.0 license. The post Mistral AI Releases Shieldstral 1.0 3B: An Open-Weights…

MarkTechPost原典の内容 · 翻訳・分析待ち翻訳待ち:Mistral AI Releases Shieldstral 1.0 3B: An Open-Weights Policy-Adaptive Multimodal Safety Classifier Matching Models 7× Its Size

翻訳待ち:PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing this choice across concepts, models, and tasks. Reported gains are confounded by simultaneous changes in dataset, layer, construction method, and pooling rule, making principled decisions impossible. We introduce PoolBench, a benchmark that isolates pooling as the experimental variable under a fixed evaluation protocol. PoolBench covers 17 concepts, 19 pooling strategies, and 3 open-weight decoder-only models (Llama-3.1-8B, Gemma-2-9B, Mistral-7B), evaluated on a single audi…

arXiv Computational Linguistics原典の内容 · 翻訳・分析待ち翻訳待ち:PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

翻訳待ち:Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.04130v1 Announce Type: new Abstract: Vision-language models for autonomous driving primarily rely on cameras and LiDAR, leaving 4D radar largely unexplored as a standalone perceptual modality despite its robustness to adverse visibility and direct measurement of radial velocity. We introduce Radar4D-VLM, a radar-only temporal vision-language model that reasons from ten consecutive 4D-radar point-cloud sweeps without camera or LiDAR input. Radar4D-VLM extracts geometrically grounded object proposals and organizes radar evidence into a compact hierarchy of object, scene, and kinematic tokens. A parameter-efficient projector maps these tokens into frozen language backbones, while auditable prediction heads jointly model object count, spatial…

arXiv Computer Vision原典の内容 · 翻訳・分析待ち翻訳待ち:Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models

翻訳待ち:New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要: I released LLM 0.32 this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider tools, redesigned content-addressable SQLite logs, new models, and new features enabled by the OpenAI Responses API. I also released new versions of the llm-anthropic, llm-gemini, and llm-openrouter plugins, each with substantial updates of their own. Headline features for LLM CLI users Running LLM against reasoning models now displays their reasoning traces to standard error, so you can see what they are "thinking" without that information being included in the standard output that you might pipe to another tool. Add -R/--hide-reasoning to turn this of…

Simon Willison's Weblog原典の内容 · 翻訳・分析待ち翻訳待ち:New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

デプロイ前にオープンソースLLMを比較する方法

AIモデルハブは、Meta、Alibaba、Google、Mistralなどの主要な開発者によるオープンソースの大規模言語モデルを比較するための集中プラットフォームです。100以上のアクティブなモデルの詳細仕様(コンテキストウィンドウ、アーキテクチャ、パラメータ数、ライセンス、ベンチマークなど)を提供します。

Hacker News AIサイト内本文デプロイ前にオープンソースLLMを比較する方法

Intel TDX上でのNVIDIA H100における機密GPU推論のベンチマーク

新しい研究では、Intel TDX環境下のNVIDIA H100 GPUで機密コンピューティングを有効にした場合の大規模言語モデル推論のパフォーマンスコストを評価。Mistral-7BとQwen3-30B-A3Bモデルを使用し、機密モードでは最初のトークンまでの時間が21.8%〜27.8%増加し、グローバルトークンスループットが17.7%〜21.1%低下した。大規模モデルはより早く飽和に達し、キャパシティ計画の調整が必要であることが示された。

arXiv AIサイト内本文Intel TDX上でのNVIDIA H100における機密GPU推論のベンチマーク

NVIDIA Vera Rubin、パフォーマンス・パーワットを向上、パートナー向けに最低トークンコストを実現

NVIDIA Vera Rubin NVL72の生産が本格化し、CoreWeave、Google Cloud、Microsoft Azure、Oracle Cloud Infrastructureの各パートナーと連携しています。このプラットフォームは極限の共同設計により最高のパフォーマンス・パーワットと最低のトークンコストを実現し、DeepSeek-R1ベンチマークではGrace Blackwell NVL72と比較してメガワットあたりのスループットが10倍向上しました。また、MicrosoftとMistralの提携により欧州のオープンモデル時代を支援します。

NVIDIA Blogサイト内本文NVIDIA Vera Rubin、パフォーマンス・パーワットを向上、パートナー向けに最低トークンコストを実現

2026年に単一24GB GPUで実行可能な最高のローカルLLM:Qwen、Gemma、Mistral、DeepSeek比較

24GB GPUは本格的なローカル推論の実用的な最低ラインです。本ガイドでは、Q4_K_Mで1枚のカードに収まる6つのオープンウェイトモデル(Qwen3.6、Gemma 4、Mistral Small、gpt-oss-20b、DeepSeek-R1-Distill)を比較し、VRAM消費、ライセンス、各モデルの得意分野を解説します。

MarkTechPostサイト内本文2026年に単一24GB GPUで実行可能な最高のローカルLLM:Qwen、Gemma、Mistral、DeepSeek比較

Mistral Vibe for Code vs Claude Code vs Cursor vs Codex:4つのエージェントが1つのスキャフォールドからPRタスクで評価

本記事では、4つの主要なAIコーディングエージェント(Mistral Vibe for Code、Claude Code、Cursor、OpenAI Codex)を、スキャフォールドからプルリクエストまでの実際のワークフローで比較評価しています。Mistral Vibeが22/25でトップ、低コスト、オープンウェイト、セルフホスティングが強み。Claude CodeとCodexは21/25で同点、Cursorは16/25。各ツールの5次元(機能スキャフォールド、テスト生成、PR/非同期ワークフロー、サーフェスカバレッジ、コスト/開放性)における長所と短所を詳述。

MarkTechPostサイト内本文Mistral Vibe for Code vs Claude Code vs Cursor vs Codex:4つのエージェントが1つのスキャフォールドからPRタスクで評価

Mistral AI、単一RGBカメラでロボットが複雑な環境をナビゲートできる8Bモデル「Robostral Navigate」をリリース

Mistral AIは、8Bパラメータの具身ナビゲーションモデルRobostral Navigateを発表しました。このモデルは、LiDARや深度センサーを必要とせず、単一のRGBカメラと自然言語の指示のみでロボットを動かします。R2R-CEの未見環境検証において、ポインティング手法、プレフィックスキャッシュトレーニング、CISPOオンライン強化学習により76.6%の成功率を達成しました。

MarkTechPostサイト内本文Mistral AI、単一RGBカメラでロボットが複雑な環境をナビゲートできる8Bモデル「Robostral Navigate」をリリース

大規模文学コーパスの自動主題索引付け:ヴォルテール全集への機械学習アプローチ

本研究は、機械学習を用いた大規模文学コーパスの自動主題索引付けを探求し、ヴォルテール作品をテストケースとして、さまざまなモデルを比較。最良のMistralシリーズ4ビット量子化モデルはF1スコア0.67を達成し、自動索引の可能性を示した。

arXiv Computational Linguisticsサイト内本文大規模文学コーパスの自動主題索引付け:ヴォルテール全集への機械学習アプローチ

Director: オンライン予測型エキスパート配置による分散MoEサービングの高速化

本論文では、予測駆動のオンラインエキスパート配置によりエンドツーエンドレイテンシを最小化する新しい分散MoEサービングシステムDirectorを提案する。軽量カスケード予測器または低ビット量子化レプリカを用いてエキスパート活性化パターンを予測し、ほぼゼロダウンタイムのマイグレーションモジュールと、多項式時間で(1+ε)近似比を達成する緩和ベースの最適化器を備える。実験では、Mistral、DeepSeek、Qwenなどの人気MoEモデルにおいて、既存手法と比較して11〜55%のレイテンシ削減を実証した。

arXiv Machine Learningサイト内本文Director: オンライン予測型エキスパート配置による分散MoEサービングの高速化

2026年中期AIモデルティアリスト

著者がコーディングと監査の経験に基づき、2026年中期の主要AIモデルを非公式にランク付け。Anthropic Fable、OpenAI Sol、Mistral、Gemini、DeepSeekを対象とし、米国の輸出規制や欧州の視点も含む。

Hacker News AIサイト内本文2026年中期AIモデルティアリスト

Show HN: Google Chat用AIアシスタント - レイアウトを保持してファイル翻訳

AnyFile Translatorは、Google Chat内でファイル、ウェブリンク、テキストを翻訳できるAIアシスタントです。元のレイアウトや書式を保持し、100以上の言語に対応。AIライティング機能も備え、コンテンツの作成と翻訳が可能です。データは暗号化され、処理後に削除されます。

Hacker News AIサイト内本文Show HN: Google Chat用AIアシスタント - レイアウトを保持してファイル翻訳

Amazon Bedrock AgentCore と Mistral AI Studio を使用した本番環境対応の e コマース MCP サーバーの構築と接続

この記事では、Amazon Bedrock AgentCore と Mistral AI Studio を使用して、本番環境対応の e コマース MCP サーバーを構築し接続する方法を詳しく説明します。MCP ツールの実装、2 層 JWT 認証、AWS CDK によるデプロイ、Mistral AI の Vibe との統合、DynamoDB と Cognito を使用したデータと ID 管理のベストプラクティスをカバーしています。

AWS Machine Learning Blogサイト内本文Amazon Bedrock AgentCore と Mistral AI Studio を使用した本番環境対応の e コマース MCP サーバーの構築と接続

タスク品質とシステムパフォーマンスに基づく長コンテキストサービングのKVキャッシュ最適化のベンチマーク

本論文は、KIVI、TurboQuant、SnapKV、CaMなどのKVキャッシュ最適化手法を、Llama-3.1-8B-InstructおよびMistral-7B-Instruct-v0.3モデル上で、マルチドキュメントQA、シングルドキュメントQA、少数ショット学習、要約タスクにおいてワークロードを考慮したベンチマークで評価した。結果は、圧縮率だけではエンドツーエンドのパフォーマンスを予測するには不十分であることを示している。KIVI4はモデル間で最も安定した品質を提供し、SnapKVは長コンテキストスループットで最も強力であり、CaMは特定のQAワークロードで大きな改善を示すが、ワークロードに対する感度が高い。この研究は、KVキャッシュ機構のワークロードを考慮した選択を動機付けている。

arXiv Computational Linguisticsサイト内本文タスク品質とシステムパフォーマンスに基づく長コンテキストサービングのKVキャッシュ最適化のベンチマーク

Mistral AI、Leanstral 1.5 を公開:Apache-2.0ライセンスのLean 4コードエージェントモデル、PutnamBench 672問中587問を解決

Mistral AI は、Lean 4 向けの無料の Apache-2.0 コードエージェントモデル Leanstral 1.5 をリリースしました。119B の mixture-of-experts アーキテクチャで、トークンあたり 6.5B パラメータを活性化し、コンテキスト長は 256k。miniF2F で 100% を達成し、PutnamBench で 587/672 問を解決、FATE-H および FATE-X で新たな SOTA を記録しました。また、実際のバグ発見にも成功し、57 のオープンソースリポジトリから 5 つの未報告バグを特定しました。

MarkTechPostサイト内本文Mistral AI、Leanstral 1.5 を公開:Apache-2.0ライセンスのLean 4コードエージェントモデル、PutnamBench 672問中587問を解決

効率的な小型言語モデルのためのWiolaアーキテクチャ

Wiolaは、GPT、LLaMA、Mistral、Falconなどの既存モデルファミリーとは無関係に、第一原理から構築された完全にオリジナルの小型言語モデル(SLM)アーキテクチャです。螺旋回転位置符号化(SRPE)、ゲート付き層間注意(GCLA)、適応型トークン統合(ATM)、二重ストリームフィードフォワード(DSFF)、WiolaRMSNormの5つの新しいコンポーネントを導入しています。4つのサイズ(120M、360M、700M、1.5Bパラメータ)でリリースされ、HuggingFace Transformersと完全互換です。

arXiv AIサイト内本文効率的な小型言語モデルのためのWiolaアーキテクチャ

基盤なきペルソナ:体制依存性とLLM個別化問題

本論文は、Beckmann & Butlin (2026) によるLLM個別化問題の存在論的枠組みに疑問を呈し、それが未議論の体制間共参照仮定を継承していると論じる。Qwen3-4B-InstructおよびMistral-7B-Instruct-v0.2でのペルソナトポロジー実験を通じて、4つの経験的楔を提示し、この仮定を覆す。そして、体制指標個別化を提案する。すなわち、表象内容の同一性単位は(媒体、体制)対であり、媒体単独ではない。

arXiv Computational Linguisticsサイト内本文基盤なきペルソナ:体制依存性とLLM個別化問題

RoPoLL: ロバストなLLM審査員団

本論文は、Huber汚染モデルの下でLLM Juryを形式化し、単一の審査員が偏ったLLM典型的な方法(モード崩壊、sycoファンシー、安全拒否)で失敗すると、任意の正の汚染に対してPoLLが非有界なバイアスを被ることを示す。審査員のコンセンサスを古典的なロバスト平均推定として捉え、RoPoLLを提案し、幾何中央値を集約関数として使用することで、最適な有限サンプル破綻点1/2を達成する。13の審査員(4B-675B)、3つの報酬モデルベンチマーク、4つの汚染体制(最大50%)での実験により、RoPoLLはすべての偏った汚染タイプでPoLLを凌駕し、38Bの3審査員委員会が30%のバイモーダルランダム汚染下でMistral-Large-3(675B)を1.31倍上回る。

arXiv AIサイト内本文RoPoLL: ロバストなLLM審査員団

企業ナビゲーション