跳到主要内容
AI News HubLIVE

来源分布

  • arXiv Computational Linguistics9
  • MarkTechPost7
  • AI Business5
  • arXiv AI5
  • Hacker News AI4
  • Simon Willison's Weblog4
  • NVIDIA Blog3
  • arXiv Machine Learning2

主题分布

  • 模型43
  • 研究27
  • Agent24
  • 芯片6
  • 创业融资5
  • 工具4
  • 政策3
  • 机器人1

日期线

  • 2026-09-013
  • 2026-09-113
  • 2026-10-073
  • 2026-07-112
  • 2026-07-132
  • 2026-07-142
  • 2026-08-132
  • 2026-08-202

最新动态

待翻译:Last Week in AI #346 - 719 math manuscripts, 2 Western open models, 1 more safety resignation

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:OpenAI publishes hundreds of math proofs from unreleased frontier model, Mistral and Reflection AI launch open-weight models to rival China, and more!

Last Week in AI来源内容 · 翻译待补全待翻译:Last Week in AI #346 - 719 math manuscripts, 2 Western open models, 1 more safety resignation

Mistral Large 4 发布:代号「Le chonk」

Mistral 放出 Mistral Large 4 预览版:总参数 1 万亿、激活参数 490 亿,在自建的 3,800 块 NVIDIA Grace Blackwell GPU 集群上训练。API 预览版已上线,开放权重承诺本月底发布;Artificial Analysis 得分 38,较上一代 Large 3 的 9 分大幅跃升,但整体仍落后前沿约 6 个月。

Simon Willison's Weblog站内正文Mistral Large 4 发布:代号「Le chonk」

Mistral Large 4:让四个前沿模型画“穿渔网袜在火星乱穿马路的犰狳”

Simon Willison 在 Hacker News 上评论 Mistral Large 4 发布时,借一句“基准测试已经饱和”的吐槽,用四个前沿模型分别生成同一张荒诞提示词的 SVG,并在博客中给出了结果链接。

Simon Willison's Weblog站内正文Mistral Large 4:让四个前沿模型画“穿渔网袜在火星乱穿马路的犰狳”

Mistral AI 发布 Mistral Large 4(Le Chonk):1.05 万亿参数多模态 MoE 模型

Mistral AI 以公开预览形式发布 Mistral Large 4(内部代号 Le Chonk):细粒度 MoE,总参数 1.05 万亿、每 token 激活 490 亿,配 16 亿参数视觉编码器与 100 万 token 上下文,在欧洲自有数据中心用 3800 块 NVIDIA Grace Blackwell GPU 从零训练。API 已上线,输入/输出每百万 token 1.36/4.18 美元,缓存输入 0.14 美元;权重与许可证预计 2026 年 10 月底公布,暂不能自托管。最突出成绩在网络安全:Cybench 93%、CyberGym-E2E 82%,Mistral 称多家闭源前沿模型因拒绝任务而接近零分。

MarkTechPost站内正文Mistral AI 发布 Mistral Large 4(Le Chonk):1.05 万亿参数多模态 MoE 模型

待翻译:Building an Enterprise AI Customer Support Platform with Claude Fable 5.1 and Claude Code

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Most AI coding demos stop at task managers, weather apps, or simple chatbots. For this project, we take on something more demanding: building an enterprise customer-support platform that can investigate complaints, retrieve relevant policies, recommend resolutions, and keep risky actions behind human approval. This gives us a practical way to test Claude Fable 5.1 as […] The post Building an Enterprise AI Customer Support Platform with Claude Fable 5.1 and Claude Code appeared first on Analytics Vidhya.

Analytics Vidhya来源内容 · 翻译待补全待翻译:Building an Enterprise AI Customer Support Platform with Claude Fable 5.1 and Claude Code

待翻译:How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users. Accelerated by inference optimizations through OpenAI’s models that tap into the capabilities of the NVIDIA Blackwell architecture, Ultrafast offers up to 8x faster token generation than the Astra Standard mode. For developers, […]

NVIDIA Blog来源内容 · 翻译待补全待翻译:How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast

待翻译:Forget ‘superintelligence’: error-prone AI nearly sparked world war three this month | Timnit Gebru and Emily M Bender

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The risks of AI aren’t what we think they are, as a recent security incident between China and the United States reveals Amid a barrage of news stories warning about superintelligent machines rendering humanity extinct, a CNN story describing the opposite scenario – one in which the US military’s reliance on brittle chatbots almost brought the US into war with China – went mostly unnoticed by the public. The biggest international AI news of the past three weeks was Anthropic engineer Jacob Coxon’s resignation. According to him, OpenAI and Anthropic are “racing straight towards self-improving superintelligence and gambling with our lives”. Coxon’s description of a “terminator” scenario, a machine becoming much smarter than humanity and deciding to wipe us out, c…

The Guardian AI来源内容 · 翻译待补全待翻译:Forget ‘superintelligence’: error-prone AI nearly sparked world war three this month | Timnit Gebru and Emily M Bender

相同数量,不同答案:语言模型中的数值表示不变性

一篇新论文检验了开放权重语言模型在数值等价改写下的答案一致性:在 3,600 道精确有理数问题和 8,600 条提示上,五个模型的正准准确率高达 0.969–0.996,但同一等值轨道的正确率与不变性降至约 0.85–0.98。研究指出,部分所谓推理失败其实来自评测器的数字语法未覆盖乘法形式的科学计数法;Mistral Small 4 则在单位换算上出现相差整十次幂的系统性错误。另有 9,000 次调用实验显示,表示共识并未优于复述共识,反而产生更多误报。

arXiv Computational Linguistics站内正文相同数量,不同答案:语言模型中的数值表示不变性

待翻译:Why Read a Research Paper When You Can Turn It Into an AI Agent?

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Have you ever read a paper in Science or Nature and thought, “Man, that research was so cool. I wish I could try that method on my own data”—only to spend a week wrestling with someone else’s undocumented repo, broken dependencies, and half-finished readme.txt? Well, now you can, more or less. Say hello to Paper2Agent, a new open-source framework that transforms academic reports into interactive AI agents you can talk to. Give it a paper along with the accompanying codebase, data or other supplementary material, and the system automatically extracts the core workflows, then spins up a tested, runnable toolkit that you can use on your own datasets. The concept may sound a little like Google’s NotebookLM (now called Gemini Notebook), which lets you upload documen…

IEEE Spectrum AI来源内容 · 翻译待补全待翻译:Why Read a Research Paper When You Can Turn It Into an AI Agent?

待翻译:From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2609.17538v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for structured information extraction from documents, yet their behavior under realistic OCR noise remains poorly understood. We present a systematic benchmark of open-source instruction-tuned LLMs for key-value pair (KVP) extraction under both clean-text and noisy OCR conditions. We evaluate representative decoder-only models (Gemma, Mistral, Qwen2.5, LLaMA 3, and DeepSeek) on the FUNSD, CORD, and SROIE benchmarks using both Gold-text annotations and OCR outputs from PaddleOCR, EasyOCR, and Tesseract. A unified evaluation protocol isolates the effects of input quality, model design, and prompting under consistent conditions. The results show that modern LLMs act…

arXiv Computational Linguistics来源内容 · 翻译待补全待翻译:From Pixels to Pairs: A Comprehensive Benchmark of LLM-Based Key-Value Extraction in Noisy Document Settings

待翻译:Mistral Bets Enterprise AI Will Be About Control, Not Just Intelligence

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The French AI lab is using a $3B fundraise to sell control over AI infrastructure, not just model power -- a shift in direction that could matter to U.S. firms in Europe too.

AI Business来源内容 · 翻译待补全待翻译:Mistral Bets Enterprise AI Will Be About Control, Not Just Intelligence

待翻译:OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2609.09203v1 Announce Type: new Abstract: Existing benchmarks for autonomous AI scientists evaluate only final outputs---generated code, hypotheses, or papers---yet discard the reasoning process by which those outputs were obtained. This makes it impossible to audit scientific methodology, diagnose failure modes, or distinguish systematic reasoning from fortunate guessing. We present \textbf{OpenDiscoveryTrace}, a public dataset of 558 complete AI scientific agent trajectories that captures how models reason, not just what they produce. Each trajectory records a structured 9-field-per-step trace---including thoughts, tool calls, observations, errors, revision triggers, and self-reported confidence---as models execute 124 scientific tasks spanning drug dis…

arXiv AI来源内容 · 翻译待补全待翻译:OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

待翻译:Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:This week, Mistral announced it raised €3 billion in a Series D funding round, pushing its post-money valuation past €21 The post Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it. appeared first on The New Stack.

The New Stack AI来源内容 · 翻译待补全待翻译:Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

待翻译:[AINews] OpenAI reports Navier-Stokes singularity find in 88 hours using Astra-next, roughly 10,000 agents and 130B tokens (>$40M), a contender for second ever Millennium Prize awarded

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Overshadowing Cognition's $48B Series E, Mistral's $24B Series D, Meta's Muse agent, and GPT Image 2.5. The most jam packed, feel the AGI day in the history of AI.

Latent Space来源内容 · 翻译待补全待翻译:[AINews] OpenAI reports Navier-Stokes singularity find in 88 hours using Astra-next, roughly 10,000 agents and 130B tokens (>$40M), a contender for second ever Millennium Prize awarded

Mistral新融资:通往主权AI的桥梁

Mistral最初以开放权重模型起家,如今面对欧洲市场环境,正将重点转向主权AI。新一轮融资被视为连接其过往开放生态与欧洲AI自主目标的关键一步。

AI Business站内正文Mistral新融资:通往主权AI的桥梁

待翻译:Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.28623v1 Announce Type: new Abstract: Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating explicit chain-of-thought reasoning before answering. In language models it has been observed that this performance often comes with sycophancy, the tendency of a model to agree with the user over the evidence. However, for LMRMs no reliable method to measure sycophancy yet exists. We bridge this gap by introducing a benchmark and dataset for evaluating LMRM sycophancy when confronted with a wrong answer from a user. Our benchmark pairs four visually grounded datasets spanning mathematical, clinical, temporal, and demographic reasoning with five pressure conditions in single-turn and multi-turn settings. We eva…

arXiv Computational Linguistics来源内容 · 翻译待补全待翻译:Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure

待翻译:SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.28911v1 Announce Type: new Abstract: The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length. We show that uniform KV quantization on a fractional-bit grid does not degrade gracefully: under a prespecified multi-seed statistical protocol, Llama-3.1-8B-Instruct with an affine quantizer is statistically indistinguishable from FP16 KV down to 2.322 code bits/value and collapses at 2.0 bits - a quality cliff in (2.0, 2.322] that reappears in generation-time quantization and multi-turn dialogue and transfers to Mistral-7B. The cliff reframes importance-aware mixed precision: above it, eight model-internal importance indicators are statistically interchangeable, s…

arXiv Machine Learning来源内容 · 翻译待补全待翻译:SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

待翻译:How law firm Gilbert + Tobin governs and scales AI with OpenAI

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:See how Gilbert + Tobin combines CEO-led commitment, rigorous governance, and human accountability to scale ChatGPT Enterprise and Codex across the firm.

OpenAI News来源内容 · 翻译待补全待翻译:How law firm Gilbert + Tobin governs and scales AI with OpenAI

待翻译:Understanding ChatGPT Work

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译: OpenAI announced ChatGPT Work on July 9th, and have been furiously iterating on it ever since. It is an extraordinarily confusing and very powerful product. Here's what I've figured out about it so far. ChatGPT Work is actually two products The more interesting version of ChatGPT Work is the one that runs in the cloud. This can be accessed via chatgpt.com or through the ChatGPT mobile apps. Let's call it Work Cloud. If you install the ChatGPT desktop app - the app that used to be called Codex - you gain access to a thing called ChatGPT Work that can access files and run programs directly on your computer. Let's call that one Work Local. This one feels more like regular Codex re-skinned to be less intimidating to non-software-developers. For the rest of this ar…

Simon Willison's Weblog来源内容 · 翻译待补全待翻译:Understanding ChatGPT Work

待翻译:Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Cohere has released Parse (parse-v5.0), a 2.3B-parameter vision language model that converts PDFs, slides and images into Markdown with HTML tables, bounding boxes and image descriptions. It runs at $1.50 per 1,000 pages through the API, or on dedicated Model Vault instances from $2,500 a month. Cohere reports a ParseBench score of 79.2, ahead of Mistral OCR 4, Azure Document Intelligence and Databricks AI Parse — but that figure averages three of the benchmark's five dimensions and drops charts and visual grounding entirely. The post Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown appeared first on MarkTechPost.

MarkTechPost来源内容 · 翻译待补全待翻译:Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown

待翻译:Mistral and Saudi Vendor to Advance Sovereign AI in Middle East

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The French AI lab extends its push for regional control of AI from Europe to the Middle East.

AI Business来源内容 · 翻译待补全待翻译:Mistral and Saudi Vendor to Advance Sovereign AI in Middle East

待翻译:Comparing Local Tool Calling: Gemma 4 vs. Llama 3 vs. Mistral

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:In this article, you will learn how Gemma 4, Llama 3, and Mistral implement tool calling locally, and what trade-offs each model family presents for...

Machine Learning Mastery来源内容 · 翻译待补全待翻译:Comparing Local Tool Calling: Gemma 4 vs. Llama 3 vs. Mistral

待翻译:Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. Why? Consider what happens when an AI agent researches a company for an investment decision. The agent queries financial databases, searches news and filings, invokes a sub-agent to run peer comparisons and model valuations, then synthesizes everything into a […]

NVIDIA Blog来源内容 · 翻译待补全待翻译:Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

待翻译:Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.18090v1 Announce Type: new Abstract: Inside a modern language model sits a single internal direction that tracks how positive or negative a sentence feels. We show how to find this valence axis (V-axis) from just 9 emotion category names plus 50 short narrative paragraphs per emotion -- about 1,500 fewer labels than the usual supervised approach -- and that the same direction appears in vision, audio, and human-brain encoders never jointly trained. The recipe: embed nine emotion-anchored story sets in a frozen encoder, take the top principal direction of the nine averaged embeddings. Projecting new inputs onto it captures 93% of supervised performance on SST-2 (Llama-3-8B-Instruct, AUC 0.772 vs. 0.828), correlates with human valence ratings on 11,811…

arXiv Computational Linguistics来源内容 · 翻译待补全待翻译:Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities

待翻译:Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.18089v1 Announce Type: new Abstract: Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hausa. This suggests that the refusal mechanism is present in the residual stream but fails to activate for low-resource inputs. Recovering it normally requires labelled target-language data and retraining, neither of which is available at scale for most African languages. We introduce Latent Space Refusal Anchoring (LSR-Anchoring), a training-free method that extracts the refusal direction from English prompts and clamps it onto the residual stream at inference time. The primary variant, Mean-Activation Steering (MAS), operates across the four architectures we tested: Llama-3-8B, Llama-3…

arXiv Computational Linguistics来源内容 · 翻译待补全待翻译:Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining

待翻译:Mistral AI Strategy

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Mistral AI wants to turn European AI sovereignty from a talking point into a product — one with a service-level agreement attached. The French artificial intelligence company announced Tuesday a three-part expansion of…

Hacker News AI来源内容 · 翻译待补全待翻译:Mistral AI Strategy

待翻译:Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.11210v1 Announce Type: new Abstract: Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers almost always fall back on uniform priors. The main reason is that building informative priors from scientific literature is slow and needs both domain and statistical expertise. We present \textbf{Distribird}, an agentic web application that automates this process. Given a parameter name, physical description, and domain context, Distribird deploys a multi-agent pipeline that searches the literature, extracts and weights reported values by domain relevance, and fits a probability distribution via AIC model selection. When no literature is available, the system fa…

arXiv AI来源内容 · 翻译待补全待翻译:Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration

待翻译:Mistral Aims to Build 1GB of Compute Capacity by 2030

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The Paris-based vendor continues to build European AI infrastructure.

AI Business来源内容 · 翻译待补全待翻译:Mistral Aims to Build 1GB of Compute Capacity by 2030

待翻译:ChatGPT and Gemini both just passed 1 billion users

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:That’s a lot of people chatting with their AI friends all day. | Image: Google For the 14th time, a Google product has hit 1 billion users. Google CEO Sundar Pichai posted on X that a billion people are using Gemini every month, and that Gemini is Google's fastest-growing product ever. A billion users is a huge milestone, but Google isn't the first AI app to hit it. OpenAI's ChatGPT hit the mark a few weeks ago, though the company buried the announcement that "more than 1 billion people are putting ChatGPT to work" in an otherwise anodyne blog post about how people use AI. External data suggested ChatGPT crossed 1 billion users as early as this June, but OpenAI hadn't announced anything until that post on August 6th. … Read the full story at The Verge.

The Verge AI来源内容 · 翻译待补全待翻译:ChatGPT and Gemini both just passed 1 billion users

待翻译:Mistral AI Releases Shieldstral 1.0 3B: An Open-Weights Policy-Adaptive Multimodal Safety Classifier Matching Models 7× Its Size

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Mistral AI has released Shieldstral 1.0 3B, an open-weights, policy-adaptive multimodal safety classifier that frames content moderation as a single yes/no question instead of a fixed harm taxonomy. Operators supply the policy as a plain-language query at inference time and get back a calibrated safety score from one forward pass — no retraining required to re-target the model. Built on Ministral-3-3B-Base-2512 with a Pixtral vision encoder and trained on roughly 54.1M samples, it reports 84.9% average F1 on text safety (matching GPT-OSS-Safeguard-20B), 83.8% on multimodal safety, and 91.3% on Mistral's adaptability benchmark — while fitting in 16GB of VRAM under an Apache 2.0 license. The post Mistral AI Releases Shieldstral 1.0 3B: An Open-Weights Policy-Adap…

MarkTechPost来源内容 · 翻译待补全待翻译:Mistral AI Releases Shieldstral 1.0 3B: An Open-Weights Policy-Adaptive Multimodal Safety Classifier Matching Models 7× Its Size

待翻译:PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing this choice across concepts, models, and tasks. Reported gains are confounded by simultaneous changes in dataset, layer, construction method, and pooling rule, making principled decisions impossible. We introduce PoolBench, a benchmark that isolates pooling as the experimental variable under a fixed evaluation protocol. PoolBench covers 17 concepts, 19 pooling strategies, and 3 open-weight decoder-only models (Llama-3.1-8B, Gemma-2-9B, Mistral-7B), evaluated on a single audited corpus o…

arXiv Computational Linguistics来源内容 · 翻译待补全待翻译:PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

待翻译:Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.04130v1 Announce Type: new Abstract: Vision-language models for autonomous driving primarily rely on cameras and LiDAR, leaving 4D radar largely unexplored as a standalone perceptual modality despite its robustness to adverse visibility and direct measurement of radial velocity. We introduce Radar4D-VLM, a radar-only temporal vision-language model that reasons from ten consecutive 4D-radar point-cloud sweeps without camera or LiDAR input. Radar4D-VLM extracts geometrically grounded object proposals and organizes radar evidence into a compact hierarchy of object, scene, and kinematic tokens. A parameter-efficient projector maps these tokens into frozen language backbones, while auditable prediction heads jointly model object count, spatial distributio…

arXiv Computer Vision来源内容 · 翻译待补全待翻译:Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models

待翻译:New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译: I released LLM 0.32 this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider tools, redesigned content-addressable SQLite logs, new models, and new features enabled by the OpenAI Responses API. I also released new versions of the llm-anthropic, llm-gemini, and llm-openrouter plugins, each with substantial updates of their own. Headline features for LLM CLI users Running LLM against reasoning models now displays their reasoning traces to standard error, so you can see what they are "thinking" without that information being included in the standard output that you might pipe to another tool. Add -R/--hide-reasoning to turn this off. LLM inclu…

Simon Willison's Weblog来源内容 · 翻译待补全待翻译:New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

如何比较部署前的开源大语言模型?

AI模型中心是一个集中平台,帮助用户对比来自Meta、阿里巴巴、谷歌、Mistral等领先开发者的开源大语言模型。它提供了超过100个活跃模型的详细规格,包括上下文窗口、架构、参数规模、许可证和基准测试结果。

Hacker News AI站内正文如何比较部署前的开源大语言模型?

基于Intel TDX的NVIDIA H100机密GPU推理性能基准测试

一项新研究评估了在NVIDIA H100 GPU上启用机密计算对大型语言模型推理性能的影响。测试使用Mistral-7B和Qwen3-30B-A3B模型,发现机密模式使首令牌延迟平均增加21.8%-27.8%,全局令牌吞吐量下降17.7%-21.1%,且较大模型更早达到饱和。结果表明机密GPU推理在负载下仍可保持可用吞吐量,但容量规划需考虑性能损失和早期饱和现象。

arXiv AI站内正文基于Intel TDX的NVIDIA H100机密GPU推理性能基准测试

NVIDIA Vera Rubin:每瓦性能领先,为全球合作伙伴提供最低令牌成本

NVIDIA Vera Rubin NVL72 正加速生产,与 CoreWeave、Google Cloud、Microsoft Azure 和 Oracle Cloud Infrastructure 等合作伙伴共同部署。该平台通过极致协同设计实现最高的每瓦性能和最低的令牌成本,在 DeepSeek-R1 基准测试中每兆瓦吞吐量比 Grace Blackwell NVL72 提升 10 倍。Vera Rubin 还支持欧洲开放模型时代,与微软和 Mistral 合作扩展 AI 基础设施。

NVIDIA Blog站内正文NVIDIA Vera Rubin:每瓦性能领先,为全球合作伙伴提供最低令牌成本

Mistral Vibe for Code vs Claude Code vs Cursor vs Codex:四大AI编程代理在脚手架到PR任务中的对比评分

本文对比了四种主流的AI编程代理:Mistral Vibe for Code、Claude Code、Cursor和OpenAI Codex,针对从功能脚手架到拉取请求的完整工作流进行评分。Mistral Vibe以22/25的总分领先,凭借成本、开放性和控制力获胜;Claude Code和Codex并列21/25;Cursor得16/25。文章详细分析了每个工具在脚手架、测试循环、PR及异步工作流、覆盖范围、成本与开放性五个维度的表现。

MarkTechPost站内正文Mistral Vibe for Code vs Claude Code vs Cursor vs Codex:四大AI编程代理在脚手架到PR任务中的对比评分

Mistral AI 发布机器人导航视觉模型

Mistral AI 推出了一款新型视觉模型,机器人仅需一个RGB摄像头和自然语言指令即可在陌生环境中导航。

AI Business站内正文Mistral AI 发布机器人导航视觉模型

Mistral AI 发布 Robostral Navigate:8B 模型仅凭单 RGB 摄像头让机器人导航复杂环境

Mistral AI 推出了 Robostral Navigate,一个 8B 参数的具身导航模型。该模型仅使用单个 RGB 摄像头,无需 LiDAR 或深度传感器,即可根据自然语言指令驱动机器人。在 R2R-CE 验证未见过的场景中,它达到了 76.6% 的成功率,这得益于其指向方法、前缀缓存训练和 CISPO 在线强化学习。

MarkTechPost站内正文Mistral AI 发布 Robostral Navigate:8B 模型仅凭单 RGB 摄像头让机器人导航复杂环境

大型文学语料库的自动主题索引:伏尔泰全集的机器学习方法

本研究探索利用机器学习自动对大型文学语料库进行主题索引,以伏尔泰作品为案例,比较了多种模型,其中Mistral系列4位量化模型F1得分达0.67,证明了自动索引的潜力。

arXiv Computational Linguistics站内正文大型文学语料库的自动主题索引:伏尔泰全集的机器学习方法

Director:通过在线主动专家放置加速分布式MoE服务

本文介绍了Director,一种新的分布式MoE推理系统,通过预测驱动的在线专家放置优化,显著降低端到端延迟。系统采用轻量级级联预测器或低比特量化副本预测专家激活模式,结合近乎零停机的在线迁移模块,以及基于松弛优化的专家放置算法,在多项式时间内达到(1+ε)近似比。实验表明,在Mistral、DeepSeek和Qwen等流行MoE模型上,相比现有工作延迟降低11%~55%。

arXiv Machine Learning站内正文Director:通过在线主动专家放置加速分布式MoE服务

2026年中AI模型分级

作者从个人编码和审计经验出发,对2026年中的主流AI模型进行非正式分级,涵盖Anthropic Fable、OpenAI Sol、Mistral、Gemini和DeepSeek等模型,并融入美国出口管制和欧洲视角的评论。

Hacker News AI站内正文2026年中AI模型分级

Show HN: 用于Google Chat的AI助手,翻译任意文件并保留布局

AnyFile Translator 是一款AI翻译助手,可在Google Chat中直接翻译文件、网页链接和文本,保留原始布局和格式,支持超过100种语言。它还具备AI写作功能,可生成并翻译内容。适合国际团队和全球客户使用。

Hacker News AI站内正文Show HN: 用于Google Chat的AI助手,翻译任意文件并保留布局

使用 Amazon Bedrock AgentCore 和 Mistral AI Studio 构建并连接生产级电子商务 MCP 服务器

本文详细介绍了如何使用 Amazon Bedrock AgentCore 和 Mistral AI Studio 构建并连接一个生产就绪的电子商务 MCP(模型上下文协议)服务器。内容涵盖 MCP 工具实现、双层 JWT 认证、AWS CDK 部署、与 Mistral AI Vibe 集成,以及使用 DynamoDB 和 Cognito 管理数据与身份的最佳实践。

AWS Machine Learning Blog站内正文使用 Amazon Bedrock AgentCore 和 Mistral AI Studio 构建并连接生产级电子商务 MCP 服务器

基于任务质量和系统性能的长上下文服务KV缓存优化基准测试

该论文对KIVI、TurboQuant、SnapKV和CaM等KV缓存优化技术进行了工作量感知的基准测试,评估了它们在Llama-3.1-8B-Instruct和Mistral-7B-Instruct-v0.3模型上的多文档问答、单文档问答、少样本学习和摘要任务中的表现。结果表明,压缩率本身并不能很好地预测端到端性能。KIVI4提供最稳定的质量,SnapKV在长上下文吞吐量方面表现最佳,而CaM在特定问答任务上取得显著提升,但对工作负载敏感。该研究强调了根据工作负载选择KV缓存机制的必要性。

arXiv Computational Linguistics站内正文基于任务质量和系统性能的长上下文服务KV缓存优化基准测试

Mistral AI 发布 Leanstral 1.5:Apache-2.0 许可的 Lean 4 代码代理模型,解决 PutnamBench 672 道问题中的 587 道

Mistral AI 发布了 Leanstral 1.5,这是一个基于 Apache-2.0 许可的 Lean 4 代码代理模型。该模型采用 119B 混合专家架构,每令牌激活 6.5B 参数,上下文长度 256k。它在 miniF2F 上达到 100% 准确率,解决了 PutnamBench 中 587/672 的问题,并在 FATE-H 和 FATE-X 基准测试上实现了新 SOTA。此外,它还能发现真实软件缺陷,已在 57 个开源仓库中识别出 5 个未报告的错误。

MarkTechPost站内正文Mistral AI 发布 Leanstral 1.5:Apache-2.0 许可的 Lean 4 代码代理模型,解决 PutnamBench 672 道问题中的 587 道

高效小型语言模型的Wiola架构

Wiola是一种全新的小型语言模型架构,从基本原理设计,与GPT、LLaMA、Mistral或Falcon等现有模型无结构关联。它引入了五种独立创新的组件:螺旋旋转位置编码(SRPE)、门控跨层注意力(GCLA)、自适应令牌合并(ATM)、双流前馈(DSFF)和WiolaRMSNorm归一化。模型提供四种规模(120M、360M、700M和1.5B参数),完全兼容HuggingFace Transformers生态系统。

arXiv AI站内正文高效小型语言模型的Wiola架构

无基底的个性:体制依赖与LLM个体化问题

本文对Beckmann & Butlin (2026)关于LLM个体化的本体论框架提出质疑,认为其继承了未论证的跨体制共指假设。通过Qwen3-4B-Instruct和Mistral-7B-Instruct-v0.2上的个性拓扑实验,作者展示了四个经验性楔子,共同削弱该假设,并提出体制索引个体化:表征内容的身份单位是(载体,体制)对,而非仅载体。

arXiv Computational Linguistics站内正文无基底的个性:体制依赖与LLM个体化问题

RoPoLL:鲁棒的大语言模型评委团

本文形式化了基于Huber污染模型的LLM陪审团,并证明即使只有一个评委以LLM典型方式(模式崩溃、谄媚、安全拒绝)产生偏差,任何正污染都会导致PoLL产生无界偏差。通过将陪审团共识视为经典鲁棒均值估计,作者提出RoPoLL,用几何中位数替换聚合函数,实现了最优有限样本崩溃点1/2。实验表明,在13个开源评委(4B-675B)、三个奖励模型基准和四种腐败机制(高达50%)下,RoPoLL在每一种有偏腐败类型上都优于PoLL:在匹配计算量的跨维度攻击上提升约19%,在重尾拜占庭对手上提升数个数量级。一个38B参数的3评委RoPoLL委员会在30%双模随机腐败下,在HelpSteer-2上以18倍参数优势超越Mistral-Large-3(675B)1.31倍。

arXiv AI站内正文RoPoLL:鲁棒的大语言模型评委团

公司导航