跳到主要內容
AI News HubLIVE

來源分布

  • arXiv Computational Linguistics9
  • MarkTechPost7
  • AI Business5
  • arXiv AI5
  • Hacker News AI4
  • Simon Willison's Weblog4
  • NVIDIA Blog3
  • arXiv Machine Learning2

主題分布

  • 模型43
  • 研究27
  • Agent24
  • 晶片6
  • 創業融資5
  • 工具4
  • 政策3
  • 機器人1

日期線

  • 2026-09-013
  • 2026-09-113
  • 2026-10-073
  • 2026-07-112
  • 2026-07-132
  • 2026-07-142
  • 2026-08-132
  • 2026-08-202

最新動態

待翻譯:Last Week in AI #346 - 719 math manuscripts, 2 Western open models, 1 more safety resignation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:OpenAI publishes hundreds of math proofs from unreleased frontier model, Mistral and Reflection AI launch open-weight models to rival China, and more!

Last Week in AI來源內容 · 翻譯待補全待翻譯:Last Week in AI #346 - 719 math manuscripts, 2 Western open models, 1 more safety resignation

Mistral Large 4 釋出:代號「Le chonk」

Mistral 放出 Mistral Large 4 預覽版:總引數 1 萬億、啟用引數 490 億,在自建的 3,800 塊 NVIDIA Grace Blackwell GPU 叢集上訓練。API 預覽版已上線,開放權重承諾本月底釋出;Artificial Analysis 得分 38,較上一代 Large 3 的 9 分大幅躍升,但整體仍落後前沿約 6 個月。

Simon Willison's Weblog站內正文Mistral Large 4 釋出:代號「Le chonk」

Mistral Large 4:讓四個前沿模型畫“穿漁網襪在火星亂穿馬路的犰狳”

Simon Willison 在 Hacker News 上評論 Mistral Large 4 釋出時,借一句“基準測試已經飽和”的吐槽,用四個前沿模型分別生成同一張荒誕提示詞的 SVG,並在部落格中給出了結果連結。

Simon Willison's Weblog站內正文Mistral Large 4:讓四個前沿模型畫“穿漁網襪在火星亂穿馬路的犰狳”

Mistral AI 釋出 Mistral Large 4(Le Chonk):1.05 萬億引數多模態 MoE 模型

Mistral AI 以公開預覽形式釋出 Mistral Large 4(內部代號 Le Chonk):細粒度 MoE,總引數 1.05 萬億、每 token 啟用 490 億,配 16 億引數視覺編碼器與 100 萬 token 上下文,在歐洲自有資料中心用 3800 塊 NVIDIA Grace Blackwell GPU 從零訓練。API 已上線,輸入/輸出每百萬 token 1.36/4.18 美元,快取輸入 0.14 美元;權重與許可證預計 2026 年 10 月底公佈,暫不能自託管。最突出成績在網路安全:Cybench 93%、CyberGym-E2E 82%,Mistral 稱多家閉源前沿模型因拒絕任務而接近零分。

MarkTechPost站內正文Mistral AI 釋出 Mistral Large 4(Le Chonk):1.05 萬億引數多模態 MoE 模型

待翻譯:Building an Enterprise AI Customer Support Platform with Claude Fable 5.1 and Claude Code

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Most AI coding demos stop at task managers, weather apps, or simple chatbots. For this project, we take on something more demanding: building an enterprise customer-support platform that can investigate complaints, retrieve relevant policies, recommend resolutions, and keep risky actions behind human approval. This gives us a practical way to test Claude Fable 5.1 as […] The post Building an Enterprise AI Customer Support Platform with Claude Fable 5.1 and Claude Code appeared first on Analytics Vidhya.

Analytics Vidhya來源內容 · 翻譯待補全待翻譯:Building an Enterprise AI Customer Support Platform with Claude Fable 5.1 and Claude Code

待翻譯:How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users. Accelerated by inference optimizations through OpenAI’s models that tap into the capabilities of the NVIDIA Blackwell architecture, Ultrafast offers up to 8x faster token generation than the Astra Standard mode. For developers, […]

NVIDIA Blog來源內容 · 翻譯待補全待翻譯:How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast

待翻譯:Forget ‘superintelligence’: error-prone AI nearly sparked world war three this month | Timnit Gebru and Emily M Bender

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The risks of AI aren’t what we think they are, as a recent security incident between China and the United States reveals Amid a barrage of news stories warning about superintelligent machines rendering humanity extinct, a CNN story describing the opposite scenario – one in which the US military’s reliance on brittle chatbots almost brought the US into war with China – went mostly unnoticed by the public. The biggest international AI news of the past three weeks was Anthropic engineer Jacob Coxon’s resignation. According to him, OpenAI and Anthropic are “racing straight towards self-improving superintelligence and gambling with our lives”. Coxon’s description of a “terminator” scenario, a machine becoming much smarter than humanity and deciding to wipe us out, c…

The Guardian AI來源內容 · 翻譯待補全待翻譯:Forget ‘superintelligence’: error-prone AI nearly sparked world war three this month | Timnit Gebru and Emily M Bender

相同數量,不同答案:語言模型中的數值表示不變性

一篇新論文檢驗了開放權重語言模型在數值等價改寫下的答案一致性:在 3,600 道精確有理數問題和 8,600 條提示上,五個模型的正準準確率高達 0.969–0.996,但同一等值軌道的正確率與不變性降至約 0.85–0.98。研究指出,部分所謂推理失敗其實來自評測器的數字語法未覆蓋乘法形式的科學計數法;Mistral Small 4 則在單位換算上出現相差整十次冪的系統性錯誤。另有 9,000 次呼叫實驗顯示,表示共識並未優於複述共識,反而產生更多誤報。

arXiv Computational Linguistics站內正文相同數量,不同答案:語言模型中的數值表示不變性

待翻譯:Why Read a Research Paper When You Can Turn It Into an AI Agent?

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Have you ever read a paper in Science or Nature and thought, “Man, that research was so cool. I wish I could try that method on my own data”—only to spend a week wrestling with someone else’s undocumented repo, broken dependencies, and half-finished readme.txt? Well, now you can, more or less. Say hello to Paper2Agent, a new open-source framework that transforms academic reports into interactive AI agents you can talk to. Give it a paper along with the accompanying codebase, data or other supplementary material, and the system automatically extracts the core workflows, then spins up a tested, runnable toolkit that you can use on your own datasets. The concept may sound a little like Google’s NotebookLM (now called Gemini Notebook), which lets you upload documen…

IEEE Spectrum AI來源內容 · 翻譯待補全待翻譯:Why Read a Research Paper When You Can Turn It Into an AI Agent?

待翻譯:From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.17538v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for structured information extraction from documents, yet their behavior under realistic OCR noise remains poorly understood. We present a systematic benchmark of open-source instruction-tuned LLMs for key-value pair (KVP) extraction under both clean-text and noisy OCR conditions. We evaluate representative decoder-only models (Gemma, Mistral, Qwen2.5, LLaMA 3, and DeepSeek) on the FUNSD, CORD, and SROIE benchmarks using both Gold-text annotations and OCR outputs from PaddleOCR, EasyOCR, and Tesseract. A unified evaluation protocol isolates the effects of input quality, model design, and prompting under consistent conditions. The results show that modern LLMs act…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings

待翻譯:Mistral Bets Enterprise AI Will Be About Control, Not Just Intelligence

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The French AI lab is using a $3B fundraise to sell control over AI infrastructure, not just model power -- a shift in direction that could matter to U.S. firms in Europe too.

AI Business來源內容 · 翻譯待補全待翻譯:Mistral Bets Enterprise AI Will Be About Control, Not Just Intelligence

待翻譯:OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2609.09203v1 Announce Type: new Abstract: Existing benchmarks for autonomous AI scientists evaluate only final outputs---generated code, hypotheses, or papers---yet discard the reasoning process by which those outputs were obtained. This makes it impossible to audit scientific methodology, diagnose failure modes, or distinguish systematic reasoning from fortunate guessing. We present \textbf{OpenDiscoveryTrace}, a public dataset of 558 complete AI scientific agent trajectories that captures how models reason, not just what they produce. Each trajectory records a structured 9-field-per-step trace---including thoughts, tool calls, observations, errors, revision triggers, and self-reported confidence---as models execute 124 scientific tasks spanning drug dis…

arXiv AI來源內容 · 翻譯待補全待翻譯:OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

待翻譯:Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:This week, Mistral announced it raised €3 billion in a Series D funding round, pushing its post-money valuation past €21 The post Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it. appeared first on The New Stack.

The New Stack AI來源內容 · 翻譯待補全待翻譯:Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

待翻譯:[AINews] OpenAI reports Navier-Stokes singularity find in 88 hours using Astra-next, roughly 10,000 agents and 130B tokens (>$40M), a contender for second ever Millennium Prize awarded

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Overshadowing Cognition's $48B Series E, Mistral's $24B Series D, Meta's Muse agent, and GPT Image 2.5. The most jam packed, feel the AGI day in the history of AI.

Latent Space來源內容 · 翻譯待補全待翻譯:[AINews] OpenAI reports Navier-Stokes singularity find in 88 hours using Astra-next, roughly 10,000 agents and 130B tokens (>$40M), a contender for second ever Millennium Prize awarded

Mistral新融資:通往主權AI的橋樑

Mistral最初以開放權重模型起家,如今面對歐洲市場環境,正將重點轉向主權AI。新一輪融資被視為連線其過往開放生態與歐洲AI自主目標的關鍵一步。

AI Business站內正文Mistral新融資:通往主權AI的橋樑

待翻譯:Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.28623v1 Announce Type: new Abstract: Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating explicit chain-of-thought reasoning before answering. In language models it has been observed that this performance often comes with sycophancy, the tendency of a model to agree with the user over the evidence. However, for LMRMs no reliable method to measure sycophancy yet exists. We bridge this gap by introducing a benchmark and dataset for evaluating LMRM sycophancy when confronted with a wrong answer from a user. Our benchmark pairs four visually grounded datasets spanning mathematical, clinical, temporal, and demographic reasoning with five pressure conditions in single-turn and multi-turn settings. We eva…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure

待翻譯:SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.28911v1 Announce Type: new Abstract: The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length. We show that uniform KV quantization on a fractional-bit grid does not degrade gracefully: under a prespecified multi-seed statistical protocol, Llama-3.1-8B-Instruct with an affine quantizer is statistically indistinguishable from FP16 KV down to 2.322 code bits/value and collapses at 2.0 bits - a quality cliff in (2.0, 2.322] that reappears in generation-time quantization and multi-turn dialogue and transfers to Mistral-7B. The cliff reframes importance-aware mixed precision: above it, eight model-internal importance indicators are statistically interchangeable, s…

arXiv Machine Learning來源內容 · 翻譯待補全待翻譯:SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

待翻譯:How law firm Gilbert + Tobin governs and scales AI with OpenAI

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:See how Gilbert + Tobin combines CEO-led commitment, rigorous governance, and human accountability to scale ChatGPT Enterprise and Codex across the firm.

OpenAI News來源內容 · 翻譯待補全待翻譯:How law firm Gilbert + Tobin governs and scales AI with OpenAI

待翻譯:Understanding ChatGPT Work

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯: OpenAI announced ChatGPT Work on July 9th, and have been furiously iterating on it ever since. It is an extraordinarily confusing and very powerful product. Here's what I've figured out about it so far. ChatGPT Work is actually two products The more interesting version of ChatGPT Work is the one that runs in the cloud. This can be accessed via chatgpt.com or through the ChatGPT mobile apps. Let's call it Work Cloud. If you install the ChatGPT desktop app - the app that used to be called Codex - you gain access to a thing called ChatGPT Work that can access files and run programs directly on your computer. Let's call that one Work Local. This one feels more like regular Codex re-skinned to be less intimidating to non-software-developers. For the rest of this ar…

Simon Willison's Weblog來源內容 · 翻譯待補全待翻譯:Understanding ChatGPT Work

待翻譯:Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Cohere has released Parse (parse-v5.0), a 2.3B-parameter vision language model that converts PDFs, slides and images into Markdown with HTML tables, bounding boxes and image descriptions. It runs at $1.50 per 1,000 pages through the API, or on dedicated Model Vault instances from $2,500 a month. Cohere reports a ParseBench score of 79.2, ahead of Mistral OCR 4, Azure Document Intelligence and Databricks AI Parse — but that figure averages three of the benchmark's five dimensions and drops charts and visual grounding entirely. The post Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown appeared first on MarkTechPost.

MarkTechPost來源內容 · 翻譯待補全待翻譯:Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown

待翻譯:Mistral and Saudi Vendor to Advance Sovereign AI in Middle East

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The French AI lab extends its push for regional control of AI from Europe to the Middle East.

AI Business來源內容 · 翻譯待補全待翻譯:Mistral and Saudi Vendor to Advance Sovereign AI in Middle East

待翻譯:Comparing Local Tool Calling: Gemma 4 vs. Llama 3 vs. Mistral

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:In this article, you will learn how Gemma 4, Llama 3, and Mistral implement tool calling locally, and what trade-offs each model family presents for...

Machine Learning Mastery來源內容 · 翻譯待補全待翻譯:Comparing Local Tool Calling: Gemma 4 vs. Llama 3 vs. Mistral

待翻譯:Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. Why? Consider what happens when an AI agent researches a company for an investment decision. The agent queries financial databases, searches news and filings, invokes a sub-agent to run peer comparisons and model valuations, then synthesizes everything into a […]

NVIDIA Blog來源內容 · 翻譯待補全待翻譯:Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

待翻譯:Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.18090v1 Announce Type: new Abstract: Inside a modern language model sits a single internal direction that tracks how positive or negative a sentence feels. We show how to find this valence axis (V-axis) from just 9 emotion category names plus 50 short narrative paragraphs per emotion -- about 1,500 fewer labels than the usual supervised approach -- and that the same direction appears in vision, audio, and human-brain encoders never jointly trained. The recipe: embed nine emotion-anchored story sets in a frozen encoder, take the top principal direction of the nine averaged embeddings. Projecting new inputs onto it captures 93% of supervised performance on SST-2 (Llama-3-8B-Instruct, AUC 0.772 vs. 0.828), correlates with human valence ratings on 11,811…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities

待翻譯:Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.18089v1 Announce Type: new Abstract: Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hausa. This suggests that the refusal mechanism is present in the residual stream but fails to activate for low-resource inputs. Recovering it normally requires labelled target-language data and retraining, neither of which is available at scale for most African languages. We introduce Latent Space Refusal Anchoring (LSR-Anchoring), a training-free method that extracts the refusal direction from English prompts and clamps it onto the residual stream at inference time. The primary variant, Mean-Activation Steering (MAS), operates across the four architectures we tested: Llama-3-8B, Llama-3…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining

待翻譯:Mistral AI Strategy

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Mistral AI wants to turn European AI sovereignty from a talking point into a product — one with a service-level agreement attached. The French artificial intelligence company announced Tuesday a three-part expansion of…

Hacker News AI來源內容 · 翻譯待補全待翻譯:Mistral AI Strategy

待翻譯:Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.11210v1 Announce Type: new Abstract: Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers almost always fall back on uniform priors. The main reason is that building informative priors from scientific literature is slow and needs both domain and statistical expertise. We present \textbf{Distribird}, an agentic web application that automates this process. Given a parameter name, physical description, and domain context, Distribird deploys a multi-agent pipeline that searches the literature, extracts and weights reported values by domain relevance, and fits a probability distribution via AIC model selection. When no literature is available, the system fa…

arXiv AI來源內容 · 翻譯待補全待翻譯:Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration

待翻譯:Mistral Aims to Build 1GB of Compute Capacity by 2030

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The Paris-based vendor continues to build European AI infrastructure.

AI Business來源內容 · 翻譯待補全待翻譯:Mistral Aims to Build 1GB of Compute Capacity by 2030

待翻譯:ChatGPT and Gemini both just passed 1 billion users

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:That’s a lot of people chatting with their AI friends all day. | Image: Google For the 14th time, a Google product has hit 1 billion users. Google CEO Sundar Pichai posted on X that a billion people are using Gemini every month, and that Gemini is Google's fastest-growing product ever. A billion users is a huge milestone, but Google isn't the first AI app to hit it. OpenAI's ChatGPT hit the mark a few weeks ago, though the company buried the announcement that "more than 1 billion people are putting ChatGPT to work" in an otherwise anodyne blog post about how people use AI. External data suggested ChatGPT crossed 1 billion users as early as this June, but OpenAI hadn't announced anything until that post on August 6th. … Read the full story at The Verge.

The Verge AI來源內容 · 翻譯待補全待翻譯:ChatGPT and Gemini both just passed 1 billion users

待翻譯:Mistral AI Releases Shieldstral 1.0 3B: An Open-Weights Policy-Adaptive Multimodal Safety Classifier Matching Models 7× Its Size

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Mistral AI has released Shieldstral 1.0 3B, an open-weights, policy-adaptive multimodal safety classifier that frames content moderation as a single yes/no question instead of a fixed harm taxonomy. Operators supply the policy as a plain-language query at inference time and get back a calibrated safety score from one forward pass — no retraining required to re-target the model. Built on Ministral-3-3B-Base-2512 with a Pixtral vision encoder and trained on roughly 54.1M samples, it reports 84.9% average F1 on text safety (matching GPT-OSS-Safeguard-20B), 83.8% on multimodal safety, and 91.3% on Mistral's adaptability benchmark — while fitting in 16GB of VRAM under an Apache 2.0 license. The post Mistral AI Releases Shieldstral 1.0 3B: An Open-Weights Policy-Adap…

MarkTechPost來源內容 · 翻譯待補全待翻譯:Mistral AI Releases Shieldstral 1.0 3B: An Open-Weights Policy-Adaptive Multimodal Safety Classifier Matching Models 7× Its Size

待翻譯:PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing this choice across concepts, models, and tasks. Reported gains are confounded by simultaneous changes in dataset, layer, construction method, and pooling rule, making principled decisions impossible. We introduce PoolBench, a benchmark that isolates pooling as the experimental variable under a fixed evaluation protocol. PoolBench covers 17 concepts, 19 pooling strategies, and 3 open-weight decoder-only models (Llama-3.1-8B, Gemma-2-9B, Mistral-7B), evaluated on a single audited corpus o…

arXiv Computational Linguistics來源內容 · 翻譯待補全待翻譯:PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

待翻譯:Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.04130v1 Announce Type: new Abstract: Vision-language models for autonomous driving primarily rely on cameras and LiDAR, leaving 4D radar largely unexplored as a standalone perceptual modality despite its robustness to adverse visibility and direct measurement of radial velocity. We introduce Radar4D-VLM, a radar-only temporal vision-language model that reasons from ten consecutive 4D-radar point-cloud sweeps without camera or LiDAR input. Radar4D-VLM extracts geometrically grounded object proposals and organizes radar evidence into a compact hierarchy of object, scene, and kinematic tokens. A parameter-efficient projector maps these tokens into frozen language backbones, while auditable prediction heads jointly model object count, spatial distributio…

arXiv Computer Vision來源內容 · 翻譯待補全待翻譯:Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models

待翻譯:New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯: I released LLM 0.32 this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider tools, redesigned content-addressable SQLite logs, new models, and new features enabled by the OpenAI Responses API. I also released new versions of the llm-anthropic, llm-gemini, and llm-openrouter plugins, each with substantial updates of their own. Headline features for LLM CLI users Running LLM against reasoning models now displays their reasoning traces to standard error, so you can see what they are "thinking" without that information being included in the standard output that you might pipe to another tool. Add -R/--hide-reasoning to turn this off. LLM inclu…

Simon Willison's Weblog來源內容 · 翻譯待補全待翻譯:New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

如何比較部署前的開源大語言模型?

AI模型中心是一個集中平臺,幫助使用者對比來自Meta、阿里巴巴、谷歌、Mistral等領先開發者的開源大語言模型。它提供了超過100個活躍模型的詳細規格,包括上下文視窗、架構、引數規模、許可證和基準測試結果。

Hacker News AI站內正文如何比較部署前的開源大語言模型?

基於Intel TDX的NVIDIA H100機密GPU推理效能基準測試

一項新研究評估了在NVIDIA H100 GPU上啟用機密計算對大型語言模型推理效能的影響。測試使用Mistral-7B和Qwen3-30B-A3B模型,發現機密模式使首令牌延遲平均增加21.8%-27.8%,全域性令牌吞吐量下降17.7%-21.1%,且較大模型更早達到飽和。結果表明機密GPU推理在負載下仍可保持可用吞吐量,但容量規劃需考慮效能損失和早期飽和現象。

arXiv AI站內正文基於Intel TDX的NVIDIA H100機密GPU推理效能基準測試

NVIDIA Vera Rubin:每瓦效能領先,為全球合作伙伴提供最低令牌成本

NVIDIA Vera Rubin NVL72 正加速生產,與 CoreWeave、Google Cloud、Microsoft Azure 和 Oracle Cloud Infrastructure 等合作伙伴共同部署。該平臺透過極致協同設計實現最高的每瓦效能和最低的令牌成本,在 DeepSeek-R1 基準測試中每兆瓦吞吐量比 Grace Blackwell NVL72 提升 10 倍。Vera Rubin 還支援歐洲開放模型時代,與微軟和 Mistral 合作擴充套件 AI 基礎設施。

NVIDIA Blog站內正文NVIDIA Vera Rubin:每瓦效能領先,為全球合作伙伴提供最低令牌成本

Mistral Vibe for Code vs Claude Code vs Cursor vs Codex:四大AI程式設計代理在腳手架到PR任務中的對比評分

本文對比了四種主流的AI程式設計代理:Mistral Vibe for Code、Claude Code、Cursor和OpenAI Codex,針對從功能腳手架到拉取請求的完整工作流進行評分。Mistral Vibe以22/25的總分領先,憑藉成本、開放性和控制力獲勝;Claude Code和Codex並列21/25;Cursor得16/25。文章詳細分析了每個工具在腳手架、測試迴圈、PR及非同步工作流、覆蓋範圍、成本與開放性五個維度的表現。

MarkTechPost站內正文Mistral Vibe for Code vs Claude Code vs Cursor vs Codex:四大AI程式設計代理在腳手架到PR任務中的對比評分

Mistral AI 釋出機器人導航視覺模型

Mistral AI 推出了一款新型視覺模型,機器人僅需一個RGB攝像頭和自然語言指令即可在陌生環境中導航。

AI Business站內正文Mistral AI 釋出機器人導航視覺模型

Mistral AI 釋出 Robostral Navigate:8B 模型僅憑單 RGB 攝像頭讓機器人導航複雜環境

Mistral AI 推出了 Robostral Navigate,一個 8B 引數的具身導航模型。該模型僅使用單個 RGB 攝像頭,無需 LiDAR 或深度感測器,即可根據自然語言指令驅動機器人。在 R2R-CE 驗證未見過的場景中,它達到了 76.6% 的成功率,這得益於其指向方法、字首快取訓練和 CISPO 線上強化學習。

MarkTechPost站內正文Mistral AI 釋出 Robostral Navigate:8B 模型僅憑單 RGB 攝像頭讓機器人導航複雜環境

大型文學語料庫的自動主題索引:伏爾泰全集的機器學習方法

本研究探索利用機器學習自動對大型文學語料庫進行主題索引,以伏爾泰作品為案例,比較了多種模型,其中Mistral系列4位量化模型F1得分達0.67,證明了自動索引的潛力。

arXiv Computational Linguistics站內正文大型文學語料庫的自動主題索引:伏爾泰全集的機器學習方法

Director:透過線上主動專家放置加速分散式MoE服務

本文介紹了Director,一種新的分散式MoE推理系統,透過預測驅動的線上專家放置最佳化,顯著降低端到端延遲。系統採用輕量級級聯預測器或低位元量化副本預測專家啟用模式,結合近乎零停機的線上遷移模組,以及基於鬆弛最佳化的專家放置演算法,在多項式時間內達到(1+ε)近似比。實驗表明,在Mistral、DeepSeek和Qwen等流行MoE模型上,相比現有工作延遲降低11%~55%。

arXiv Machine Learning站內正文Director:透過線上主動專家放置加速分散式MoE服務

2026年中AI模型分級

作者從個人編碼和審計經驗出發,對2026年中的主流AI模型進行非正式分級,涵蓋Anthropic Fable、OpenAI Sol、Mistral、Gemini和DeepSeek等模型,並融入美國出口管制和歐洲視角的評論。

Hacker News AI站內正文2026年中AI模型分級

Show HN: 用於Google Chat的AI助手,翻譯任意檔案並保留佈局

AnyFile Translator 是一款AI翻譯助手,可在Google Chat中直接翻譯檔案、網頁連結和文本,保留原始佈局和格式,支援超過100種語言。它還具備AI寫作功能,可生成並翻譯內容。適合國際團隊和全球客戶使用。

Hacker News AI站內正文Show HN: 用於Google Chat的AI助手,翻譯任意檔案並保留佈局

使用 Amazon Bedrock AgentCore 和 Mistral AI Studio 構建並連線生產級電子商務 MCP 伺服器

本文詳細介紹瞭如何使用 Amazon Bedrock AgentCore 和 Mistral AI Studio 構建並連線一個生產就緒的電子商務 MCP(模型上下文協議)伺服器。內容涵蓋 MCP 工具實現、雙層 JWT 認證、AWS CDK 部署、與 Mistral AI Vibe 整合,以及使用 DynamoDB 和 Cognito 管理資料與身份的最佳實踐。

AWS Machine Learning Blog站內正文使用 Amazon Bedrock AgentCore 和 Mistral AI Studio 構建並連線生產級電子商務 MCP 伺服器

基於任務質量和系統效能的長上下文服務KV快取最佳化基準測試

該論文對KIVI、TurboQuant、SnapKV和CaM等KV快取最佳化技術進行了工作量感知的基準測試,評估了它們在Llama-3.1-8B-Instruct和Mistral-7B-Instruct-v0.3模型上的多文件問答、單文件問答、少樣本學習和摘要任務中的表現。結果表明,壓縮率本身並不能很好地預測端到端效能。KIVI4提供最穩定的質量,SnapKV在長上下文吞吐量方面表現最佳,而CaM在特定問答任務上取得顯著提升,但對工作負載敏感。該研究強調了根據工作負載選擇KV快取機制的必要性。

arXiv Computational Linguistics站內正文基於任務質量和系統效能的長上下文服務KV快取最佳化基準測試

Mistral AI 釋出 Leanstral 1.5:Apache-2.0 許可的 Lean 4 程式碼代理模型,解決 PutnamBench 672 道問題中的 587 道

Mistral AI 釋出了 Leanstral 1.5,這是一個基於 Apache-2.0 許可的 Lean 4 程式碼代理模型。該模型採用 119B 混合專家架構,每令牌啟用 6.5B 引數,上下文長度 256k。它在 miniF2F 上達到 100% 準確率,解決了 PutnamBench 中 587/672 的問題,並在 FATE-H 和 FATE-X 基準測試上實現了新 SOTA。此外,它還能發現真實軟體缺陷,已在 57 個開源倉庫中識別出 5 個未報告的錯誤。

MarkTechPost站內正文Mistral AI 釋出 Leanstral 1.5:Apache-2.0 許可的 Lean 4 程式碼代理模型,解決 PutnamBench 672 道問題中的 587 道

高效小型語言模型的Wiola架構

Wiola是一種全新的小型語言模型架構,從基本原理設計,與GPT、LLaMA、Mistral或Falcon等現有模型無結構關聯。它引入了五種獨立創新的元件:螺旋旋轉位置編碼(SRPE)、門控跨層注意力(GCLA)、自適應令牌合併(ATM)、雙流前饋(DSFF)和WiolaRMSNorm歸一化。模型提供四種規模(120M、360M、700M和1.5B引數),完全相容HuggingFace Transformers生態系統。

arXiv AI站內正文高效小型語言模型的Wiola架構

無基底的個性:體制依賴與LLM個體化問題

本文對Beckmann & Butlin (2026)關於LLM個體化的本體論框架提出質疑,認為其繼承了未論證的跨體制共指假設。透過Qwen3-4B-Instruct和Mistral-7B-Instruct-v0.2上的個性拓撲實驗,作者展示了四個經驗性楔子,共同削弱該假設,並提出體制索引個體化:表徵內容的身份單位是(載體,體制)對,而非僅載體。

arXiv Computational Linguistics站內正文無基底的個性:體制依賴與LLM個體化問題

RoPoLL:魯棒的大語言模型評委團

本文形式化了基於Huber汙染模型的LLM陪審團,並證明即使只有一個評委以LLM典型方式(模式崩潰、諂媚、安全拒絕)產生偏差,任何正汙染都會導致PoLL產生無界偏差。透過將陪審團共識視為經典魯棒均值估計,作者提出RoPoLL,用幾何中位數替換聚合函式,實現了最優有限樣本崩潰點1/2。實驗表明,在13個開源評委(4B-675B)、三個獎勵模型基準和四種腐敗機制(高達50%)下,RoPoLL在每一種有偏腐敗型別上都優於PoLL:在匹配計算量的跨維度攻擊上提升約19%,在重尾拜占庭對手上提升數個數量級。一個38B引數的3評委RoPoLL委員會在30%雙模隨機腐敗下,在HelpSteer-2上以18倍引數優勢超越Mistral-Large-3(675B)1.31倍。

arXiv AI站內正文RoPoLL:魯棒的大語言模型評委團

公司導航