AI News HubLIVE

ソース分布

  • Hacker News AI14
  • AI Business6
  • arXiv Computational Linguistics6
  • MarkTechPost6
  • arXiv AI4
  • NVIDIA Blog2
  • The Decoder2
  • AI Weekly1

トピック分布

  • モデル43
  • Agent26
  • 研究22
  • チップ10
  • 政策9
  • スタートアップ6
  • ロボット4
  • ツール3

タイムライン

  • 2026-06-173
  • 2026-05-272
  • 2026-05-282
  • 2026-06-242
  • 2026-06-302
  • 2026-07-132
  • 2026-07-142
  • 2026-08-132

最新動向

翻訳待ち:Mistral and Saudi Vendor to Advance Sovereign AI in Middle East

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The French AI lab extends its push for regional control of AI from Europe to the Middle East.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • The French AI lab extends its push for regional control of AI from Europe to the Middle East.
サイト内本文

翻訳待ち:Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. Why? Consider what happens when an AI agent researches a company for an investment decision. The agent queries financial databases, searches news and filings, invokes a sub-agent to run peer comparisons and model valuations, then synthesizes everything into a […]

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. Why? Consider what happens when an AI agent researches a company for an inve…
サイト内本文

翻訳待ち:Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.18090v1 Announce Type: new Abstract: Inside a modern language model sits a single internal direction that tracks how positive or negative a sentence feels. We show how to find this valence axis (V-axis) from just 9 emotion category names plus 50 short narrative paragraphs per emotion -- about 1,500 fewer labels than the usual supervised approach -- and that the same direction appears in vision, audio, and human-brain encoders never jointly trained. The recipe: embed nine emotion-anchored story sets in a frozen encoder, take the top principal direction of the nine averaged embeddings. Projecting new inputs onto it captures 93% of supervised performance on SST-2 (Llama-3-8B-Instruct, AUC 0.772 vs. 0.828), correlates with human valence ratings on 11,811 EmoSet images at r=0.636, reaches AUC 0.906 on ESC-50 audio (p12). A 2-parameter classifier trained on text labels transfers to images (AUC 0.961), audio (0.764), and brain recordings (0.828) without target-modality labels; a generic 16-D subspace stays at chance (0.525). The recipe is bounded to continuous attributes -- seven tests on categorical concepts return near-chance -- and steering is family-specific (Llama/Mistral yes, Qwen/Gemma no).

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.18090v1 Announce Type: new Abstract: Inside a modern language model sits a single internal direction that tracks how positive or negative a sentence feels. We show how…
サイト内本文

翻訳待ち:Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.18089v1 Announce Type: new Abstract: Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hausa. This suggests that the refusal mechanism is present in the residual stream but fails to activate for low-resource inputs. Recovering it normally requires labelled target-language data and retraining, neither of which is available at scale for most African languages. We introduce Latent Space Refusal Anchoring (LSR-Anchoring), a training-free method that extracts the refusal direction from English prompts and clamps it onto the residual stream at inference time. The primary variant, Mean-Activation Steering (MAS), operates across the four architectures we tested: Llama-3-8B, Llama-3.1-70B, Mistral-7B-Instruct, and Qwen2.5-7B. On Mistral and Qwen it recovers safety with benign degradation below 0.08. On Llama-3-8B it overcorrects, with Degraded Performance on Legitimate prompts (DPL) reaching 1.00. We address this with SAE-Derived Steering (SDS), which replaces the dense mean-difference direction with a single Sparse Autoencoder (SAE) feature and reduces Kullback-Leibler (KL) divergence by 3.5-7x without benign collapse. Four languages transfer positively, but Arabic fails on every architecture and at every steering magnitude, indicating a geometric mismatch rather than a baseline effect. Massive Multitask Language Understanding (MMLU) accuracy drops remain below 0.35 percentage points at every effective steering magnitude.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.18089v1 Announce Type: new Abstract: Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hau…
サイト内本文

翻訳待ち:Mistral AI Strategy

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Mistral AI wants to turn European AI sovereignty from a talking point into a product — one with a service-level agreement attached. The French artificial intelligence company announced Tuesday a three-part expansion of…

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Mistral AI wants to turn European AI sovereignty from a talking point into a product — one with a service-level agreement attached. The French artificial intelligence company anno…
サイト内本文

翻訳待ち:Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.11210v1 Announce Type: new Abstract: Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers almost always fall back on uniform priors. The main reason is that building informative priors from scientific literature is slow and needs both domain and statistical expertise. We present \textbf{Distribird}, an agentic web application that automates this process. Given a parameter name, physical description, and domain context, Distribird deploys a multi-agent pipeline that searches the literature, extracts and weights reported values by domain relevance, and fits a probability distribution via AIC model selection. When no literature is available, the system falls back to sensible uninformative alternatives, and clearly reports both the evidence behind and the confidence level of every prior it produces. It is designed for the problems where the models have physically interpretable parameters, where domain knowledge exists in the published literature. We evaluate the tool on 24~parameters across 10 scientific domains comparing three open-weight models (Qwen3.6 27B, Gemma 4 31B, Mistral Small 4 119B) with a single-prompt LLM baseline. On prior quality the full pipeline \emph{matches} this baseline. Every prior is traced to the specific papers and values from which it was constructed; a built-in validity layer declines to produce priors for out-of-scope requests, whereas the single-prompt baseline returns confident but unfounded priors for them in 11 of 30~model--parameter cases; and every language-model call runs locally, so no parameter description or unpublished modelling detail is transmitted to a third-party LLM provider (only generated search terms reach the public literature databases). For scientific use, we argue these properties matter more than a marginal improvement in point-estimate accuracy.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.11210v1 Announce Type: new Abstract: Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodologi…
サイト内本文

翻訳待ち:Mistral Aims to Build 1GB of Compute Capacity by 2030

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The Paris-based vendor continues to build European AI infrastructure.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • The Paris-based vendor continues to build European AI infrastructure.
サイト内本文

翻訳待ち:ChatGPT and Gemini both just passed 1 billion users

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:That’s a lot of people chatting with their AI friends all day. | Image: Google For the 14th time, a Google product has hit 1 billion users. Google CEO Sundar Pichai posted on X that a billion people are using Gemini every month, and that Gemini is Google's fastest-growing product ever. A billion users is a huge milestone, but Google isn't the first AI app to hit it. OpenAI's ChatGPT hit the mark a few weeks ago, though the company buried the announcement that "more than 1 billion people are putting ChatGPT to work" in an otherwise anodyne blog post about how people use AI. External data suggested ChatGPT crossed 1 billion users as early as this June, but OpenAI hadn't announced anything until that post on August 6th. … Read the full story at The Verge.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • That’s a lot of people chatting with their AI friends all day. | Image: Google For the 14th time, a Google product has hit 1 billion users. Google CEO Sundar Pichai posted on X th…
サイト内本文

翻訳待ち:Mistral AI Releases Shieldstral 1.0 3B: An Open-Weights Policy-Adaptive Multimodal Safety Classifier Matching Models 7× Its Size

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Mistral AI has released Shieldstral 1.0 3B, an open-weights, policy-adaptive multimodal safety classifier that frames content moderation as a single yes/no question instead of a fixed harm taxonomy. Operators supply the policy as a plain-language query at inference time and get back a calibrated safety score from one forward pass — no retraining required to re-target the model. Built on Ministral-3-3B-Base-2512 with a Pixtral vision encoder and trained on roughly 54.1M samples, it reports 84.9% average F1 on text safety (matching GPT-OSS-Safeguard-20B), 83.8% on multimodal safety, and 91.3% on Mistral's adaptability benchmark — while fitting in 16GB of VRAM under an Apache 2.0 license. The post Mistral AI Releases Shieldstral 1.0 3B: An Open-Weights Policy-Adaptive Multimodal Safety Classifier Matching Models 7× Its Size appeared first on MarkTechPost.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Mistral AI has released Shieldstral 1.0 3B, an open-weights, policy-adaptive multimodal safety classifier that frames content moderation as a single yes/no question instead of a f…
サイト内本文

翻訳待ち:PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing this choice across concepts, models, and tasks. Reported gains are confounded by simultaneous changes in dataset, layer, construction method, and pooling rule, making principled decisions impossible. We introduce PoolBench, a benchmark that isolates pooling as the experimental variable under a fixed evaluation protocol. PoolBench covers 17 concepts, 19 pooling strategies, and 3 open-weight decoder-only models (Llama-3.1-8B, Gemma-2-9B, Mistral-7B), evaluated on a single audited corpus of 37,693 real-text passages. The primary axis is linear separability (D1/AUROC); steered concept prevalence (D2/SCP) and output-level disentanglement (D3) serve as diagnostic axes. The primary finding is decisive: W4_hierarchical reaches a cross-model mean AUROC of 0.7799, while the widely adopted P1_last_token baseline reaches only 0.7640 and is statistically significantly worse (Friedman+Nemenyi, p = 2.0e-36; 77 significant pairs among 18 effective strategies). Rankings are stable across layers (rho = 0.961--0.990). A key negative result: strong detection does not imply strong steering -- D2 and D3 are substantially weaker than D1 for most concepts, indicating a fundamental representational limit rather than a pooling failure. On mid-difficulty concepts, W4_hierarchical outperforms P1_last_token by 0.042--0.113 AUROC; construction method choice (DiffMean vs. REPE) has a larger effect (delta AUROC 0.15) than pooling (delta AUROC 0.016), establishing the correct practical hierarchy. We release the corpus, pre-extracted activations, scorer models, steering vectors, and evaluation code as a reusable protocol for pooling research.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collaps…
サイト内本文

翻訳待ち:Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.04130v1 Announce Type: new Abstract: Vision-language models for autonomous driving primarily rely on cameras and LiDAR, leaving 4D radar largely unexplored as a standalone perceptual modality despite its robustness to adverse visibility and direct measurement of radial velocity. We introduce Radar4D-VLM, a radar-only temporal vision-language model that reasons from ten consecutive 4D-radar point-cloud sweeps without camera or LiDAR input. Radar4D-VLM extracts geometrically grounded object proposals and organizes radar evidence into a compact hierarchy of object, scene, and kinematic tokens. A parameter-efficient projector maps these tokens into frozen language backbones, while auditable prediction heads jointly model object count, spatial distribution, motion state, collision risk, semantic category, and radial velocity. Radar4D-VLM combines proposal-grounded temporal object tokenization, global scene context, and explicit kinematic tokens within a unified frozen-backbone interface. On sequence-isolated K-Radar development validation, its Top-64 proposal recall reaches 98.13% at 4 m, exceeding fixed-lattice and uniform-random controls by 6.40 and 22.83 percentage points, respectively. We further evaluate 24 matched runs spanning eight frozen Qwen, Phi, Mistral, Llama, and Gemma backbones under an identical adaptation budget. The radar-token interface remains compatible across all five language-model families, while matched aligned, permuted, and no-language controls show sensor dependence but no stable direct-head gain from aligned language supervision. These results establish a reproducible foundation for radar-only multimodal scene and motion reasoning while separating interface compatibility from the benefit of language supervision.

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • arXiv:2608.04130v1 Announce Type: new Abstract: Vision-language models for autonomous driving primarily rely on cameras and LiDAR, leaving 4D radar largely unexplored as a standal…
サイト内本文

翻訳待ち:New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:<p>I released <a href="https://llm.datasette.io/en/stable/changelog.html#v0-32">LLM 0.32</a> this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider tools, redesigned content-addressable SQLite logs, new models, and new features enabled by the OpenAI Responses API. I also released new versions of the <code>llm-anthropic</code>, <code>llm-gemini</code>, and <code>llm-openrouter</code> plugins, each with substantial updates of their own.</p> <h4 id="headline-features-for-llm-cli-users">Headline features for LLM CLI users</h4> <p>Running LLM against reasoning models now <strong>displays their reasoning traces</strong> to standard error, so you can see what they are "thinking" without that information being included in the standard output that you might pipe to another tool. Add <code>-R/--hide-reasoning</code> to turn this off.</p> <p><img src="https://static.simonwillison.net/static/2026/best-pelicans.gif" alt="Running llm &quot;think about the best thing about pelicans&quot; in the macOS terminal window - grey text outputs saying Exploring pelican qualities, then after a paragraph of that a white paragraph of text comes out saying: The best thing about pelicans is their wonderfully oversized, practical design: that enormous bill and pouch look comical, but they make pelicans remarkably skilled fishers. Even better, many species cooperate—working together to herd fish before scooping them up. They’re a great mix of goofy, graceful, and surprisingly clever." style="max-width: 100%;" /></p> <p>LLM includes support out-of-the-box for the <strong>GPT-5.6 model family</strong>, and the new default model used with <code>llm "prompt"</code> is now the inexpensive but capable <strong>GPT-5.6 Luna</strong>.</p> <p>LLM calls can now use <strong>server-side tools</strong> from various providers. OpenAI provide <a href="https://llm.datasette.io/en/stable/openai-models.html#code-interpreter">a code execution environment</a> as a server-side tool; LLM can now run prompts that benefit from that like so:</p> <div class="highlight highlight-source-shell"><pre>llm --tool CodeInterpreter <span class="pl-s"><span class="pl-pds">'</span>Show current python and SQLite versions<span class="pl-pds">'</span></span></pre></div> <p>OpenAI also gets a <a href="https://llm.datasette.io/en/stable/openai-models.html#web-search">WebSearch</a> tool.</p> <p>The <a href="https://github.com/simonw/llm-anthropic">llm-anthropic</a> plugin adds <a href="https://github.com/simonw/llm-anthropic/blob/0.26/README.md#web-search">WebSearch</a>, <a href="https://github.com/simonw/llm-anthropic/blob/0.26/README.md#web-fetch">WebFetch</a>, <a href="https://github.com/simonw/llm-anthropic/blob/0.26/README.md#code-execution">CodeExecution</a>, and <a href="https://github.com/simonw/llm-anthropic/blob/0.26/README.md#mcp-connector">AnthropicMCP</a>, which looks like this:</p> <div class="highlight highlight-source-shell"><pre>llm -m claude-sonnet-5 -T <span class="pl-s"><span class="pl-pds">'</span>AnthropicMCP("https://datasette.simonwillison.net/-/mcp")<span class="pl-pds">'</span></span> \ <span class="pl-s"><span class="pl-pds">'</span>how many rows in the blog_blogmark table?<span class="pl-pds">'</span></span></pre></div> <p>That causes Anthropic to execute MCP calls against my new <a href="https://simonwillison.net/2026/Jul/31/stateless-mcp/#datasette-mcp">datasette-mcp</a> plugin as part of a single request/response interaction with their API.</p> <p>The new <strong>llm openai endpoint</strong> command provides a tool for <a href="https://llm.datasette.io/en/stable/other-models.html#run-against-an-endpoint-without-configuring-it">executing prompts against <em>any</em> OpenAI compatible endpoint</a> as a one-liner. These aren't logged, which makes this a handy tool for running one-off prompts against anything that speaks the lingua franca of the LLM API world.</p> <p>Here's how I use that to run prompts against Gemma 4 12B running in my localhost <a href="https://lmstudio.ai">LM Studio</a> API, via <code>uvx</code> (no LLM installation required) and mixing in the <a href="https://github.com/simonw/llm-tools-quickjs">llm-tools-quickjs</a> tool plugin for good measure:</p> <div class="highlight highlight-source-shell"><pre>uvx --with llm-tools-quickjs \ llm openai endpoint http://localhost:1234/v1 -m google/gemma-4-12b \ -T QuickJS <span class="pl-s"><span class="pl-pds">'</span>Use QuickJS to multiply 3434 * 2434<span class="pl-pds">'</span></span> --td</pre></div> <p><img src="https://static.simonwillison.net/static/2026/openai-endpoint-gemma.webp" alt="Output reads Tool call: QuickJS_execute_javascript({'javascript': '3434 * 2434'}) 8358356 The result of 3434 * 2434 is 8,358,356." style="max-width: 100%;" /></p> <h4 id="new-features-in-the-python-api">New features in the Python API</h4> <p>LLM's Python API previously required you to create a conversation and then send messages to it one at a time. This was an abstraction over the true nature of LLMs, where each request carries a complete history of the messages that came before it. That abstraction started to get in the way for some more advanced cases, so the new release introduces a <code>model.prompt(messages=[])</code> parameter that can be used like this:</p> <pre><span class="pl-k">import</span> <span class="pl-s1">llm</span> <span class="pl-k">from</span> <span class="pl-s1">llm</span> <span class="pl-k">import</span> <span class="pl-s1">user</span>, <span class="pl-s1">assistant</span>, <span class="pl-s1">system</span> <span class="pl-s1">model</span> <span class="pl-c1">=</span> <span class="pl-s1">llm</span>.<span class="pl-c1">get_model</span>(<span class="pl-s">"gpt-5.6-luna"</span>) <span class="pl-s1">response</span> <span class="pl-c1">=</span> <span class="pl-s1">model</span>.<span class="pl-c1">prompt</span>(<span class="pl-s1">messages</span><span class="pl-c1">=</span>[ <span class="pl-en">system</span>(<span class="pl-s">"You are a helpful pirate."</span>), <span class="pl-en">user</span>(<span class="pl-s">"What is the capital of France?"</span>), <span class="pl-en">assistant</span>(<span class="pl-s">"Paris, matey."</span>), <span class="pl-en">user</span>(<span class="pl-s">"And Germany?"</span>), ]) <span class="pl-en">print</span>(<span class="pl-s1">response</span>.<span class="pl-c1">text</span>())</pre> <p>LLM previously returned an iterable sequence of strings from each prompt. This worked great when models returned a string response, but failed to predict the weird shape that models would evolve towards. Today many models return a mix of reasoning text, output strings, tool calls, and even image attachments. With LLM 0.32 you can <a href="https://llm.datasette.io/en/stable/python-api.html#structured-messages-and-streaming-events">do this instead</a>:</p> <pre><span class="pl-k">for</span> <span class="pl-s1">event</span> <span class="pl-c1">in</span> <span class="pl-s1">model</span>.<span class="pl-c1">prompt</span>(<span class="pl-s">"Explain cats"</span>).<span class="pl-c1">stream_events</span>(): <span class="pl-k">if</span> <span class="pl-s1">event</span>.<span class="pl-c1">type</span> <span class="pl-c1">==</span> <span class="pl-s">"reasoning"</span>: <span class="pl-en">print</span>(<span class="pl-s">f"[thinking] <span class="pl-s1"><span class="pl-kos">{</span><span class="pl-s1">event</span>.<span class="pl-c1">chunk</span><span class="pl-kos">}</span></span>"</span>, <span class="pl-s1">end</span><span class="pl-c1">=</span><span class="pl-s">""</span>, <span class="pl-s1">flush</span><span class="pl-c1">=</span><span class="pl-c1">True</span>) <span class="pl-k">elif</span> <span class="pl-s1">event</span>.<span class="pl-c1">type</span> <span class="pl-c1">==</span> <span class="pl-s">"text"</span>: <span class="pl-en">print</span>(<span class="pl-s1">event</span>.<span class="pl-c1">chunk</span>, <span class="pl-s1">end</span><span class="pl-c1">=</span><span class="pl-s">""</span>, <span class="pl-s1">flush</span><span class="pl-c1">=</span><span class="pl-c1">True</span>) <span class="pl-k">else</span>: <span class="pl-en">print</span>(<span class="pl-s">f"Other event: <span class="pl-s1"><span class="pl-kos">{</span><span class="pl-s1">event</span><span class="pl-kos">}</span></span>"</span>)</pre> <p>Combine these features and we can <em>finally</em> provide a robust implementation of the semi-standard OpenAI chat completions API, which I've now released as the <a href="https://github.com/simonw/llm-chat-completions-server">llm-chat-completions-server</a> plugin:</p> <div class="highlight highlight-source-shell"><pre>llm install llm-chat-completions-server llm chat-completions-server --port 9000 <span class="pl-c"><span class="pl-c">#</span> Server is now running on http://127.0.0.1:9000/v1</span></pre></div> <p>Now you can run prompts against LLM via that server, using the new <code>llm openai endpoint</code> command!</p> <div class="highlight highlight-source-shell"><pre>llm openai endpoint http://127.0.0.1:9000/v1 <span class="pl-s"><span class="pl-pds">'</span>hello<span class="pl-pds">'</span></span> -m gpt-5.4-mini</pre></div> <p>The bigger challenge with that kind of API concerns logging. If we're going to support the pattern where the message sequence is appended to on every request, ideally we can avoid logging all of that duplicate JSON for every turn.</p> <p>The solution is the new <a href="https://llm.datasette.io/en/stable/logging.html#the-message-store">content-addressable message store</a>, modeled after Git. You can see the new schema for that <a href="https://llm.datasette.io/en/stable/logging.html#sql-schema">in the documentation</a>, but the <code>llm logs</code> and <code>llm logs --json</code> commands have both been upgraded to convert that format back into something that's easy to consume.</p> <h4 id="and-the-rest">And the rest</h4> <p>There is a whole lot more in this release. The <a href="https://llm.datasette.io/en/stable/changelog.html#v0-32">0.32 release notes</a> are pretty comprehensive, and the notes for <a href="https://llm.datasette.io/en/stable/changelog.html#rc2-2026-07-30">0.32rc2</a>, <a href="https://llm.datasette.io/en/stable/changelog.html#rc1-2026-07-30">0.32rc</a>, <a href="https://llm.datasette.io/en/stable/changelog.html#a3-2026-06-09">0.32a3</a>, <a href="https://llm.datasette.io/en/stable/changelog.html#a2-2026-05-12">0.32a2</a>, and <a href="https://llm.datasette.io/en/stable/changelog.html#a0-2026-04-28">0.32a0</a> should fill in any gaps.</p> <p>Existing LLM plugins should all continue to work, but plugins that provide extra models will need to be upgraded to 0.32 in order to participate fully in the new streaming events system. There's a guide to implementing plugins with <a href="https://llm.datasette.io/en/stable/plugins/advanced-model-plugins.html#structured-messages-and-streaming-events">Structured messages and streaming events</a> in the documentation.</p> <p>I've updated some of my own plugins:</p> <ul> <li> <a href="https://github.com/simonw/llm-anthropic/releases/tag/0.26">llm-anthropic 0.26</a> adds support for the Claude 5 family of models, plus <code>WebSearch</code>, <code>WebFetch</code>, <code>CodeExecution</code>, and <code>AnthropicMCP</code> server-side tools.</li> <li> <a href="https://github.com/simonw/llm-gemini">llm-gemini</a> and <a href="https://github.com/simonw/llm-openrouter">llm-openrouter</a> and <a href="https://github.com/simonw/llm-mistral">llm-mistral</a> are nearly there, releases coming soon.</li> </ul> <h4 id="i-guess-llm-is-an-agent-framework-now">I guess LLM is an agent framework now</h4> <p>Quite a few of the lower-level tools changes in this release were driven by the needs of <a href="https://agent.datasette.io/">Datasette Agent</a>. When I started work on LLM, the term "agent" had such a vague definition that I refused to use it. In <a href="https://simonwillison.net/2025/Sep/18/agents/">September 2025</a> I came around to the idea that "<strong>An LLM agent runs tools in a loop to achieve a goal</strong>" is well established enough now that I could stop avoiding the term entirely.</p> <p>Tool chains can now <a href="https://llm.datasette.io/en/stable/python-api.html#python-api-tools-pause">pause for human approval</a> and <a href="https://llm.datasette.io/en/stable/python-api.html#python-api-tools-resume">resume from a stored message history</a> - both needed by Datasette Agent.</p> <p>Looking at LLM today it's beginning to look very agent-shaped to me. There's something neat about having a CLI utility that can mix and match different tools from different sources with different models all as a one-liner, and that includes a Python library powerful enough to build systems like <a href="https://agent.datasette.io/">Datasette Agent</a> and <a href="https://github.com/simonw/llm-coding-agent">llm-coding-agent</a>.</p> <p>Maybe the next version of LLM will bake the concept of an "agent" into the core library. I'm still trying to figure out what that would look like.</p> <p>Tags: <a href="https://simonwillison.net/tags/projects">projects</a>, <a href="https://simonwillison.net/tags/releases">releases</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/openai">openai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/llm">llm</a>, <a href="https://simonwillison.net/tags/anthropic">anthropic</a></p>

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • <p>I released <a href="https://llm.datasette.io/en/stable/changelog.html#v0-32">LLM 0.32</a> this morning, the most significant new version of LLM since the initial launch of the…
サイト内本文

デプロイ前にオープンソースLLMを比較する方法

AIモデルハブは、Meta、Alibaba、Google、Mistralなどの主要な開発者によるオープンソースの大規模言語モデルを比較するための集中プラットフォームです。100以上のアクティブなモデルの詳細仕様(コンテキストウィンドウ、アーキテクチャ、パラメータ数、ライセンス、ベンチマークなど)を提供します。

  • AIモデルハブには100のアクティブなオープンソースLLMが集約されています。
  • モデルはMeta、Alibaba、Google、Mistral、Microsoft、DeepSeekなどの開発者から提供されています。
サイト内本文

Intel TDX上でのNVIDIA H100における機密GPU推論のベンチマーク

新しい研究では、Intel TDX環境下のNVIDIA H100 GPUで機密コンピューティングを有効にした場合の大規模言語モデル推論のパフォーマンスコストを評価。Mistral-7BとQwen3-30B-A3Bモデルを使用し、機密モードでは最初のトークンまでの時間が21.8%〜27.8%増加し、グローバルトークンスループットが17.7%〜21.1%低下した。大規模モデルはより早く飽和に達し、キャパシティ計画の調整が必要であることが示された。

  • 機密コンピューティングはAI推論の実用的な要件になりつつあるが、パフォーマンスコストが生じる。
  • Intel TDX機密インスタンス内のH100 GPUで2つのLLMをテスト。
サイト内本文

NVIDIA Vera Rubin、パフォーマンス・パーワットを向上、パートナー向けに最低トークンコストを実現

NVIDIA Vera Rubin NVL72の生産が本格化し、CoreWeave、Google Cloud、Microsoft Azure、Oracle Cloud Infrastructureの各パートナーと連携しています。このプラットフォームは極限の共同設計により最高のパフォーマンス・パーワットと最低のトークンコストを実現し、DeepSeek-R1ベンチマークではGrace Blackwell NVL72と比較してメガワットあたりのスループットが10倍向上しました。また、MicrosoftとMistralの提携により欧州のオープンモデル時代を支援します。

  • Vera Rubin NVL72の生産が加速、30カ国350以上の工場サイトをカバー
  • メガワットあたりのトークン数が前世代比10倍、100万トークンあたりのコストは1/10
サイト内本文

2026年に単一24GB GPUで実行可能な最高のローカルLLM:Qwen、Gemma、Mistral、DeepSeek比較

24GB GPUは本格的なローカル推論の実用的な最低ラインです。本ガイドでは、Q4_K_Mで1枚のカードに収まる6つのオープンウェイトモデル(Qwen3.6、Gemma 4、Mistral Small、gpt-oss-20b、DeepSeek-R1-Distill)を比較し、VRAM消費、ライセンス、各モデルの得意分野を解説します。

  • 24GBが実用的な最低ライン:無理に70Bを詰め込むのではなく、適切なサイズの20B~35Bモデルを実行すべき。
  • Qwen3.6-27Bが最もバランスの取れたデフォルト、DeepSeek-R1-Distill-Qwen-32Bは約18~20GBで最もタイトなフィット。
サイト内本文

Mistral Vibe for Code vs Claude Code vs Cursor vs Codex:4つのエージェントが1つのスキャフォールドからPRタスクで評価

本記事では、4つの主要なAIコーディングエージェント(Mistral Vibe for Code、Claude Code、Cursor、OpenAI Codex)を、スキャフォールドからプルリクエストまでの実際のワークフローで比較評価しています。Mistral Vibeが22/25でトップ、低コスト、オープンウェイト、セルフホスティングが強み。Claude CodeとCodexは21/25で同点、Cursorは16/25。各ツールの5次元(機能スキャフォールド、テスト生成、PR/非同期ワークフロー、サーフェスカバレッジ、コスト/開放性)における長所と短所を詳述。

  • Mistral Vibe for Codeが22/25で首位、低価格、オープンソースCLI、セルフホスティング可能。
  • Claude CodeとOpenAI Codexは21/25で同点;Claudeは生のコーディング品質、Codexはクロスサーフェス非同期が優位。
サイト内本文

Mistral AI、単一RGBカメラでロボットが複雑な環境をナビゲートできる8Bモデル「Robostral Navigate」をリリース

Mistral AIは、8Bパラメータの具身ナビゲーションモデルRobostral Navigateを発表しました。このモデルは、LiDARや深度センサーを必要とせず、単一のRGBカメラと自然言語の指示のみでロボットを動かします。R2R-CEの未見環境検証において、ポインティング手法、プレフィックスキャッシュトレーニング、CISPOオンライン強化学習により76.6%の成功率を達成しました。

  • Robostral NavigateはMistral AI初の具身ナビゲーション向け8Bモデル。
  • 単一RGBカメラのみでR2R-CE未見環境において76.6%の成功率を達成。
サイト内本文

大規模文学コーパスの自動主題索引付け:ヴォルテール全集への機械学習アプローチ

本研究は、機械学習を用いた大規模文学コーパスの自動主題索引付けを探求し、ヴォルテール作品をテストケースとして、さまざまなモデルを比較。最良のMistralシリーズ4ビット量子化モデルはF1スコア0.67を達成し、自動索引の可能性を示した。

  • 主題索引は大規模文学・歴史版にとって重要だが、手動では労力がかかる。本研究はヴォルテールの『諸国民の風俗と精神』と『百科全書問題』をテストケースとしてMLを適用。
  • タスクはマルチラベル分類として枠組み化。エンコーダベースのモデルからファインチューニングされたLLM(3~1200億パラメータ)まで比較。
サイト内本文

Director: オンライン予測型エキスパート配置による分散MoEサービングの高速化

本論文では、予測駆動のオンラインエキスパート配置によりエンドツーエンドレイテンシを最小化する新しい分散MoEサービングシステムDirectorを提案する。軽量カスケード予測器または低ビット量子化レプリカを用いてエキスパート活性化パターンを予測し、ほぼゼロダウンタイムのマイグレーションモジュールと、多項式時間で(1+ε)近似比を達成する緩和ベースの最適化器を備える。実験では、Mistral、DeepSeek、Qwenなどの人気MoEモデルにおいて、既存手法と比較して11〜55%のレイテンシ削減を実証した。

  • 予測駆動のオンラインエキスパート配置
  • ほぼゼロダウンタイムのエキスパートマイグレーション
サイト内本文

2026年中期AIモデルティアリスト

著者がコーディングと監査の経験に基づき、2026年中期の主要AIモデルを非公式にランク付け。Anthropic Fable、OpenAI Sol、Mistral、Gemini、DeepSeekを対象とし、米国の輸出規制や欧州の視点も含む。

  • Fable(Anthropic)はB評価:流暢だが信頼性に欠け、バグを隠す傾向がある。
  • Sol(OpenAI)はS評価:低レベルコードとテストで信頼できる。
サイト内本文

Show HN: Google Chat用AIアシスタント - レイアウトを保持してファイル翻訳

AnyFile Translatorは、Google Chat内でファイル、ウェブリンク、テキストを翻訳できるAIアシスタントです。元のレイアウトや書式を保持し、100以上の言語に対応。AIライティング機能も備え、コンテンツの作成と翻訳が可能です。データは暗号化され、処理後に削除されます。

  • PDF、Word、PPTなどのファイルをレイアウト保持で翻訳
  • 100以上の言語に対応、チャット内で直接利用可能
サイト内本文

Amazon Bedrock AgentCore と Mistral AI Studio を使用した本番環境対応の e コマース MCP サーバーの構築と接続

この記事では、Amazon Bedrock AgentCore と Mistral AI Studio を使用して、本番環境対応の e コマース MCP サーバーを構築し接続する方法を詳しく説明します。MCP ツールの実装、2 層 JWT 認証、AWS CDK によるデプロイ、Mistral AI の Vibe との統合、DynamoDB と Cognito を使用したデータと ID 管理のベストプラクティスをカバーしています。

  • AgentCore Runtime を利用して MCP サーバーをホストし、コンテナやロードバランサーの管理を不要に。
  • 2 層認証を実装:インフラ層での JWT 検証とアプリケーション層での ID 解決。
サイト内本文

タスク品質とシステムパフォーマンスに基づく長コンテキストサービングのKVキャッシュ最適化のベンチマーク

本論文は、KIVI、TurboQuant、SnapKV、CaMなどのKVキャッシュ最適化手法を、Llama-3.1-8B-InstructおよびMistral-7B-Instruct-v0.3モデル上で、マルチドキュメントQA、シングルドキュメントQA、少数ショット学習、要約タスクにおいてワークロードを考慮したベンチマークで評価した。結果は、圧縮率だけではエンドツーエンドのパフォーマンスを予測するには不十分であることを示している。KIVI4はモデル間で最も安定した品質を提供し、SnapKVは長コンテキストスループットで最も強力であり、CaMは特定のQAワークロードで大きな改善を示すが、ワークロードに対する感度が高い。この研究は、KVキャッシュ機構のワークロードを考慮した選択を動機付けている。

  • KIVI4はモデル間で最も安定した品質を提供する。
  • SnapKVは長コンテキストスループットで最高のパフォーマンスを示す。
サイト内本文

Mistral AI、Leanstral 1.5 を公開:Apache-2.0ライセンスのLean 4コードエージェントモデル、PutnamBench 672問中587問を解決

Mistral AI は、Lean 4 向けの無料の Apache-2.0 コードエージェントモデル Leanstral 1.5 をリリースしました。119B の mixture-of-experts アーキテクチャで、トークンあたり 6.5B パラメータを活性化し、コンテキスト長は 256k。miniF2F で 100% を達成し、PutnamBench で 587/672 問を解決、FATE-H および FATE-X で新たな SOTA を記録しました。また、実際のバグ発見にも成功し、57 のオープンソースリポジトリから 5 つの未報告バグを特定しました。

  • Leanstral 1.5 は、Mistral AI による無料・Apache-2.0 ライセンスの Lean 4 証明エンジニアリングモデル。
  • 119B の mixture-of-experts、トークンあたり 6.5B 活性パラメータ、256k コンテキスト。
サイト内本文

効率的な小型言語モデルのためのWiolaアーキテクチャ

Wiolaは、GPT、LLaMA、Mistral、Falconなどの既存モデルファミリーとは無関係に、第一原理から構築された完全にオリジナルの小型言語モデル(SLM)アーキテクチャです。螺旋回転位置符号化(SRPE)、ゲート付き層間注意(GCLA)、適応型トークン統合(ATM)、二重ストリームフィードフォワード(DSFF)、WiolaRMSNormの5つの新しいコンポーネントを導入しています。4つのサイズ(120M、360M、700M、1.5Bパラメータ)でリリースされ、HuggingFace Transformersと完全互換です。

  • Wiolaは既存モデルと構造的系統を共有しない完全オリジナルのSLMアーキテクチャ。
  • 5つの新規コンポーネント:SRPE、GCLA、ATM、DSFF、WiolaRMSNorm。
サイト内本文

基盤なきペルソナ:体制依存性とLLM個別化問題

本論文は、Beckmann & Butlin (2026) によるLLM個別化問題の存在論的枠組みに疑問を呈し、それが未議論の体制間共参照仮定を継承していると論じる。Qwen3-4B-InstructおよびMistral-7B-Instruct-v0.2でのペルソナトポロジー実験を通じて、4つの経験的楔を提示し、この仮定を覆す。そして、体制指標個別化を提案する。すなわち、表象内容の同一性単位は(媒体、体制)対であり、媒体単独ではない。

  • Beckmann & Butlinの枠組みは、異なる体制間で同じ方向が同じ内容を指すという未証明の仮定に依存している。
  • 実験により、プロンプト抽出ベクトルと微調整盆地の非共線性、架空ペルソナが実在アンカー方向に強くモデルを変位させることなどが示された。
サイト内本文

RoPoLL: ロバストなLLM審査員団

本論文は、Huber汚染モデルの下でLLM Juryを形式化し、単一の審査員が偏ったLLM典型的な方法(モード崩壊、sycoファンシー、安全拒否)で失敗すると、任意の正の汚染に対してPoLLが非有界なバイアスを被ることを示す。審査員のコンセンサスを古典的なロバスト平均推定として捉え、RoPoLLを提案し、幾何中央値を集約関数として使用することで、最適な有限サンプル破綻点1/2を達成する。13の審査員(4B-675B)、3つの報酬モデルベンチマーク、4つの汚染体制(最大50%)での実験により、RoPoLLはすべての偏った汚染タイプでPoLLを凌駕し、38Bの3審査員委員会が30%のバイモーダルランダム汚染下でMistral-Large-3(675B)を1.31倍上回る。

  • PoLL(LLM評価者パネル)は、単一の審査員が典型的なLLMバイアスを示すと、任意の正の汚染に対して非有界なバイアスを生じる。
  • RoPoLLは集約関数を幾何中央値に置き換え、破綻点1/2の最適なロバスト性を達成する。
サイト内本文

Bored People Chat:匿名グローバルチャットルーム、昔のインターネットの純粋さを再現

Bored People Chat は、サインアップ不要、広告なし、ボットなしの匿名グローバルチャットルームです。安全性に重点を置き、AIによるモデレーションを導入。孤独や退屈を感じる人々が交流できる空間を提供します。

  • 匿名、サインアップ不要、広告なし、ボットなしの公共ルーム
  • AIによる自動モデレーションで個人情報をフィルタリング
サイト内本文

エージェントツール使用のベンチマーク

LangChain は、LLM のツール使用能力を評価するための4つの新しいテスト環境を公開しました。関数呼び出し、計画、推論などのスキルをカバーします。テストでは、GPT-4、Claude 2.1、GPT-3.5、およびオープンソースモデル(Mistral 7b など)を比較。主な発見:GPT-4 は関係データタスクで最高得点だが、長い軌道では失敗;Claude 2.1 は3つのタスクで GPT-4 と同等;オープンソースモデルは複数関数の組み合わせが苦手;計画は依然として困難。

  • LangChain が LLM のツール使用を評価する4つのベンチマークを発表(タイプライター(単一・26ツール)、関係データ、マルチバース算数)。
  • GPT-4 は関係データタスクで最高得点だが、単純な長期的タスクでも失敗。
サイト内本文

Mistral AI、OCR 4で非構造化データの課題に取り組む

フランスのスタートアップ企業のモデルには、ユーザーが非構造化データをよりよく理解するためのバウンディングボックスなどの機能が含まれています。

  • Mistral AIがOCR 4モデルをリリース
  • 非構造化データ分析のためのバウンディングボックス機能を搭載
サイト内本文

Mistral OCR 4:RAG、エージェント、エンタープライズ検索パイプラインに引用可能な構造化出力を提供

Mistral AI は2026年6月23日、OCR 4をリリースしました。これは、クリーンなテキスト抽出から構造化ドキュメント出力に移行したものです。各ブロックは、バウンディングボックス、型分類、ページ単位および単語単位の信頼度スコアを返します。このモデルは170の言語をサポートし、単一のセルフホストコンテナで実行され、1つのAPIエンドポイントを通じて引用可能な入力をRAG、エージェント、エンタープライズ検索パイプラインに供給します。

  • OCR 4はテキストだけでなく、バウンディングボックス、型付きブロックラベル、単語単位の信頼度スコアを返します。
  • 10のグループにわたる170の言語をサポートし、希少言語や低リソース言語で向上しています。
サイト内本文

Mistral OCR 4 発表:文書理解の新境地

Mistral OCR 4 は、バウンディングボックス、ブロック分類、信頼度スコアを導入。人間による評価テストで全競合を上回り(勝率72%)、170言語をサポート、単一コンテナでのセルフホストが可能。

  • 独立した評価者がOCR 4を72%の勝率で選好、OlmOCRBenchで最高スコア85.20。
  • テキストに加えて、バウンディングボックス、ブロックタイプ、単語ごとの信頼度スコアを出力。
サイト内本文

Mistral AI、より大規模なモデルファミリーを生産へ

Mistral AIは今年夏に新モデルをリリース予定。それは「太いが疎な」新しいモデルファミリーの始まりとなる。7月から研究、政府などの主要パートナー向けに早期アクセスプログラムを開始。

  • Mistral AIが夏に新モデルを発表
  • 新しいモデルファミリーは太いが疎
サイト内本文

Mistral社のLe Chat、プロンプトに対して国家支援の偽情報を半数繰り返す

ニュースガードの監査によると、Mistral AIのチャットボットLe Chatはイラン戦争に関する虚偽の主張に対して、英語で50%、フランス語で56.6%の確率で偽情報を繰り返した。フランス軍はカスタマイズ版のLe Chat Enterpriseを使用しており、一般向け版とは異なる。

  • ニュースガード監査:Le Chatは英語で50%、フランス語で56.6%の確率で偽情報を繰り返した。
  • 虚偽の主張はロシア、中国、イランの国家系情報源からのもの。
サイト内本文

Vibe が本稼働開始

Mistral は、長期にわたる多段階の作業とコーディングを処理する統一型 AI エージェント「Vibe」をリリース。エンタープライズツールと統合し、ワークモードとコードモードを提供するほか、VS Code 拡張機能と CLI のアップデートも発表。

  • Vibe は Mistral の新しい AI エージェントで、作業とコーディングを統合し、Le Chat からアップグレード。
  • ワークモードでは、エンタープライズ検索、データ分析、ドキュメント合成、スケジューリングなどの複雑なタスクを処理。
サイト内本文

Cohere、エンタープライズ向けに Sovereign AI を販売した後、初のコーディングモデルで開発者をターゲットに

Cohere は、Apache 2.0 ライセンスの初のオープンウェイトコーディングモデル North Mini Code をリリース。AI インフラを所有・管理したい開発者を対象とし、30B MoE モデルは H100 GPU 1枚で動作し、エージェンティックコーディングタスクで Mistral、Qwen、Gemma と競合する。

  • Cohere が North Mini Code を発表。300億パラメータの MoE モデルで、アクティブパラメータは30億、Apache 2.0 で提供。
  • モデルは NVIDIA H100 GPU 1枚で動作し、開発者がセルフホスティングを実用的に利用可能。
サイト内本文

主要AIプラットフォームでのトークン使用量とサブスクリプション追跡

Tokens 4 Breakfast は、Claude、OpenAI、Cursor、Copilot、Gemini、DeepSeek、Mistral などの主要AIプラットフォームにおけるトークン使用量、サブスクリプション費用、レート制限をリアルタイムで追跡するmacOSメニューバーアプリです。開発者が予期しない過剰支出を回避するためのアラート、予算設定、プロバイダー間の可視性を提供します。フリー版は1プロバイダー対応、Pro版は一括$7.99で全機能を解放します。

  • メニューバーにAI使用コスト、レート制限、サブスクリプション費用をリアルタイム表示。
  • Claude、OpenAI、Cursor など主要8プロバイダーに対応。
サイト内本文

Mistral AI、欧州AI推進のため30億ユーロ調達を模索

フランスのAIスタートアップMistral AIは、約30億ユーロの新たな資金調達ラウンドを交渉中で、評価額は約200億ユーロとなっています。

  • Mistral AIが30億ユーロの資金調達を交渉中
  • 評価額は約200億ユーロ
サイト内本文

Scikit-LLMとオープンソースLLMの使用

この記事では、OllamaとScikit-LLM Pythonライブラリを使用して、Llama 3、Mistral、Gemmaなどのローカルでホストされたオープンソース大規模言語モデルを無料でテキスト分類に利用する方法を学びます。

  • Ollamaをインストールし、オープンソースLLMをローカルで実行。
  • Scikit-LLMをローカルのOllamaエンドポイントに設定。
サイト内本文

Mistral Vibe:長期・多段階の作業とコーディングのためのAIエージェント

Mistral Vibeは、長時間実行される多段階の作業やプログラミングタスクに特化したAIエージェントです。Product Huntでの議論を紹介します。

  • Mistral Vibeは長期・多段階のワークフローとコーディングに重点を置いています。
  • Product Huntで公開され、コミュニティで話題に。
サイト内本文

Mistral、欧州はAIインフラ構築に2年の猶予と警告

Mistral AIのサミットで、CEOのArthur Mensch氏は欧州がAIインフラを構築するのに残された時間は2年であり、さもなくば米国の「属国」になると警告した。イベントには多くの参加者が集まり、データ主権やオープンソースモデルへの関心の高まりが示されたが、投資・規模では米国に大きく後れを取っている。

  • Mistral CEO、欧州は2年以内にAIインフラ構築を、さもなくば米国の属国になると警告。
  • サミットに大勢の参加者、欧州の自立したAIエコシステムへの渇望を示す。
サイト内本文

Mistral AI Now Summit パリ現地レポート

パリで開催されたMistral AI Now Summitでの個人的な洞察。Mistralはモデル企業から、自社の計算リソース、モデル、プラットフォーム、コンサルティングを含むフルスタックAIプロバイダーへと進化しています。サミットでは新モデルの発表よりも、ASML、BNPパリバ、アマゾンとのパートナーシップが強調されました。専門的な小型モデル(Document AI、Voxtral、Robostral)が特定のタスクで大型汎用モデルを上回っています。主権とオンプレミス展開は、欧州企業にとって重要な差別化要因です。古代パピルス文書の解読にAIを活用した講演は、人文科学におけるAIの可能性を示しました。

  • Mistralはモデル企業から、自社の計算資源、モデル、プラットフォーム、コンサルティングを備えたフルスタックAIプロバイダーへと変貌している。
  • サミットでは新モデル発表よりも、ASML、BNPパリバ、アマゾンとのパートナーシップに焦点が当てられた。
サイト内本文

Mistral AI、Digital Realtyと提携し欧州AIインフラを拡大

フランスのスタートアップ企業Mistral AIは、Digital Realtyのパリ南キャンパスで10メガワットのコンピューティング能力を確保しました。

  • Mistral AIがDigital Realtyのパリ南キャンパスで10MWの計算能力を確保
  • この提携は欧州のAIインフラ拡大を目指す
サイト内本文

Mistral、LeChatをVibeにブランド変更、チャットボットの未来は本格的なワークエージェントに

Mistral AIは、チャットボット「Le Chat」を「Vibe」に名称変更し、チャット、コーディングエージェント、新しいワークモードを1つのブランドに統合する。ワークモードはGoogle Workspace、Outlook、Slack、GitHubに接続し、メールやレポート、プルリクエストなどのタスクを自律的に処理する。Pro料金は17.99ユーロから14.99ユーロに値下げされたが、具体的な利用制限は明らかにされていない。これにより、OpenAI、Google、Anthropicのエージェント型サービスとの直接的な競争を仕掛ける。

  • Mistral AIがチャットボット「Le Chat」を「Vibe」にブランド変更、チャット、コーディングエージェント、ワークモードを統合。
  • ワークモードはGoogle Workspace、Outlook、Slack、GitHubと連携し、タスクを自律処理。
サイト内本文

Mistral、独自チップの設計を検討とCEOが表明

Mistral AIのCEOアーサー・メンシュ氏は、インフラコスト削減のためカスタムチップの開発を検討していると認め、OpenAIやAnthropicに対抗する。また、フランスに推論専用のデータセンターを新設し、エンタープライズ向けエージェントプラットフォーム「Vibe」を発表した。

  • Mistral AIは独自カスタムチップの設計を検討し、展開コスト削減を目指す。
  • フランスに推論専用の新しいデータセンターを発表。
サイト内本文

AIウィークリー第496号:Anthropicの国防総省モデルが今や誰でも使える

今週のAIニュース:Anthropicがこれまで政府契約業者限定だったMythosモデルを公開、国防総省級AIが誰でも利用可能に。DeepMindのDemis HassabisはAGI実現時期を2029年に前倒し。Starletteフレームワークに重大な認証バイパス脆弱性、数百万のAIエージェントに影響。CrowdStrikeらがGlasswormボットネットを共同撃滅。BNPパリバがMistralと主権AIセキュリティ提携、中国はAlibabaとDeepSeekのトップAIエンジニアの海外渡航を制限。UberはAIトークン予算を4ヶ月で使い切り、ClickUpは2200人を解雇して3000の内部AIエージェントを導入。一方、MITテクノロジーレビューはAI露出職種の失業率が低いと報告、Altmanはホワイトカラー消滅予測を撤回。

  • AnthropicがMythosモデルを公開、NSAや国防総省の能力が標準APIで利用可能に。
  • DeepMindのハサビスCEOがAGI実現を2029年と明言、AlphaProof Nexusの成果を根拠に。
サイト内本文

Mistral AI、Harveyとの提携で法務分野に進出

生成AIベンダーのMistral AIは、Anthropicの法務AI取引を彷彿とさせる動きで、法務業界に進出しています。

  • Mistral AIがHarveyと提携し、法務分野に参入。
  • この動きはAnthropicの法務AI連携を彷彿とさせる。
サイト内本文

企業ナビゲーション

Mistral AI ニュース | AI News Hub