本文にスキップ
AI News HubLIVE

オープンソースモデルの最新ニュース

翻訳待ち:Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Cohere has released North Small Translate, an open-weight Mixture-of-Experts model built for machine translation across 50 languages. It uses 25B of its 218B parameters per token and scores 83.6 on Cohere's WMT26 evaluation. Weights are free for non-commercial use, with commercial access through Cohere Model Vault or RWS Language Weaver. The post Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages appeared first on MarkTechPost.

MarkTechPostサイト内本文翻訳待ち:Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

翻訳待ち:Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Google Research has released ToolGrad, an ACL 2026 Findings framework that inverts tool-use dataset generation: it builds a verified API chain first, then writes the matching user query. Guided by textual "gradients" from a 4-module propose-execute-select-update loop, ToolGrad reaches a 99.8% pass rate on ToolBench versus 63.8% for DFS search. Gemma-3-12B fine-tuned on only 500 samples scores 83.1 on BFCL, next to Gemini 2.5 Pro at 83.2. Code, dataset, and models are public under Apache-2.0. The post Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation appeared first on MarkTechPost.

MarkTechPostサイト内本文翻訳待ち:Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation

翻訳待ち:Reduce LLM latency with prefix-aware routing on Amazon SageMaker Inference

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Amazon SageMaker Inference now offers prefix-aware routing, a routing strategy that sends requests sharing the same prompt prefix to the same instance so the KV cache stays warm. In benchmarks on Llama 3.1 70B, it reduced P50 time-to-first-token by up to 77% and raised KV cache hit rates from about 25% to over 80%.

AWS Machine Learning Blogサイト内本文翻訳待ち:Reduce LLM latency with prefix-aware routing on Amazon SageMaker Inference

翻訳待ち:Reduce inference cold starts on Amazon SageMaker HyperPod with model caching

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Amazon SageMaker HyperPod now supports model caching for inference, which pre-loads model weights and container images onto cluster nodes so pods read from local NVMe storage instead of downloading over the network. Learn how model caching cuts cold starts from tens of minutes to seconds, how it works, and how to enable it.

AWS Machine Learning Blogサイト内本文翻訳待ち:Reduce inference cold starts on Amazon SageMaker HyperPod with model caching

翻訳待ち:Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:This week, Mistral announced it raised €3 billion in a Series D funding round, pushing its post-money valuation past €21 The post Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it. appeared first on The New Stack.

The New Stack AIサイト内本文翻訳待ち:Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

翻訳待ち:Video-MOPD: Multi-Teacher On-Policy Distillation for Video Understanding

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.09300v1 Announce Type: new Abstract: Video understanding demands a convergence of complementary capabilities across perception, temporal understanding, and complex reasoning, which are difficult to jointly optimize within a single model. We introduce Video-MOPD-8B, an open-weight model dedicated to video understanding tasks. To fundamentally enhance its capabilities, we conduct targeted reinforcement learning (RL) optimization across three core domains: video temporal grounding (VTG), general video comprehension, and video STEM reasoning. We then unify their complementary capabilities via Multi-Teacher On-Policy Distillation (MOPD), which consolidates expert knowledge by supervising student-generated trajectories with routed teacher feedb…

arXiv Computer Visionサイト内本文翻訳待ち:Video-MOPD: Multi-Teacher On-Policy Distillation for Video Understanding

翻訳待ち:MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.09206v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) often struggle with hallucinations, thus hindering their reliable practical applications. Existing attention-based mitigation methods mainly rely on indirect signals (e.g., attention weights) that fail to accurately reflect the actual information shift underlying hallucination generation. In this paper, we propose HEAL, Head-lEvel information disentAnglement and caLibration for identifying and mitigating hallucinations. HEAL first employs causal noise intervention on multi-head outputs to filter out causally redundant heads. Subsequently, it disentangles information distribution within the remaining heads via the counterfactual Difference-in-Differences, categor…

arXiv Computer Visionサイト内本文翻訳待ち:MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads

翻訳待ち:TEFM: Token-Efficient Faithful Modeling for Structured Data

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.09552v1 Announce Type: new Abstract: In this paper, we solve two fundamental obstacles in applying LLMs to critical domains: token efficiency and faithfulness. To address both constraints jointly, we present TEFM (Token-Efficient Faithful Modeling), a framework designed for structured data analysis in critical domains. TEFM achieves token efficiency by compressing lengthy structured observations into compact Behavioral Code tokens, dramatically reducing token consumption with minimal information loss. Moreover, TEFM enables faithful rationalization through a dual-fidelity objective that jointly optimizes code-level reconstruction and prediction-level fidelity, identifying minimal sufficient feature subsets grounded in input data. Comprehe…

arXiv Computational Linguisticsサイト内本文翻訳待ち:TEFM: Token-Efficient Faithful Modeling for Structured Data

翻訳待ち:Benchmarking Hybrid Deep Research Across Database Querying and Web Search

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.09410v1 Announce Type: new Abstract: While autonomous agents have made significant strides in "deep research" by iteratively navigating the open web to synthesize information, real-world problem-solving is rarely confined to a single environment. Complex analytical tasks inherently require agents to weave together evidence from both ambiguous unstructured text (e.g., the open web) and highly precise structured data (e.g., relational databases). However, existing benchmarks evaluate these modalities in isolation, failing to capture the critical "handoff" - the ability to preserve constraints when moving evidence between systems. We introduce HybridDeepResearch, to our knowledge the first deep-research benchmark that requires both web searc…

arXiv Computational Linguisticsサイト内本文翻訳待ち:Benchmarking Hybrid Deep Research Across Database Querying and Web Search

翻訳待ち:Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.09338v1 Announce Type: new Abstract: Speculative decoding is critical for accelerating LLM inference. However, the speedup is fragile: drafters are typically trained against a narrow distribution for a single target model, and their acceptance rate collapses under workload shifts. This is a striking inversion of modern LLM development, where target models are valued precisely for the broad generalization they acquire through large-scale pretraining. We argue that the natural remedy, pretraining, has been hard to apply to drafters because existing recipes are target-specific: the drafter consumes the target's hidden states and is distilled on the target's logits, so pretraining must be repeated for each target. We introduce Osprey, which i…

arXiv Computational Linguisticsサイト内本文翻訳待ち:Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding

翻訳待ち:GraphNOSE: A Graph Transformer in Olfaction

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.05694v1 Announce Type: new Abstract: Predicting olfactory qualities from molecular structure is an open problem in chemoinformatics. Although linear models can link molecular features to odor descriptors, they often fail when extrapolating to novel chemical scaffolds, extreme molecular weights, or complex odor mixtures. To address this, we introduce GraphNOSE, an open-source graph transformer framework that predicts multi-label odor descriptors from simplified molecular-input line-entry system (SMILES) strings for single molecules and binary mixtures. By integrating positional and structural encodings within a transformer-based graph architecture, GraphNOSE achieves strong performance with six times fewer parameters than standard graph ne…

arXiv Machine Learningサイト内本文翻訳待ち:GraphNOSE: A Graph Transformer in Olfaction

翻訳待ち:Connecting Score Matching, Maximum Likelihood, and Expectation-Maximization in Mixed Linear Regression

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.05688v1 Announce Type: new Abstract: We study variance-preserving diffusion of the response in mixed linear regression (MLR) with unknown mixing weights. Our analysis separates the statistical guarantees of score matching from the loss geometry and optimization signal at a fixed diffusion noise level. The KL divergence links the denoising score matching objective integrated over the diffusion path with the likelihood and a terminal discrepancy. Under mild regularity conditions and terminal schedule, the resulting estimator converges up to the ground truth parameters of MLR, and its scaled error converges to the Gaussian limit of the maximum-likelihood estimator. At a fixed scale of the diffusion noise level, we derive a decomposition link…

arXiv Machine Learningサイト内本文翻訳待ち:Connecting Score Matching, Maximum Likelihood, and Expectation-Maximization in Mixed Linear Regression

翻訳待ち:PAC-Private Autoregressive Generation: Calibrating Noise to Ensemble Disagreement

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.05676v1 Announce Type: new Abstract: Language models adapted on private text are often served through APIs, so privacy leakage occurs through generated outputs rather than exposed weights. Private prediction protects these releases. Methods such as PMixED incur privacy cost at each release and increasingly rely on the public model over long horizons. PAC privacy instead calibrates noise to output variability across possible secrets, adding less noise when predictions are stable. To our knowledge, PAC-private prediction has not previously been extended from classification to autoregressive generation. We construct $m=128$ overlapping worlds from the private corpus, with each record appearing in exactly $m/2$ worlds, and train one adapter p…

arXiv Machine Learningサイト内本文翻訳待ち:PAC-Private Autoregressive Generation: Calibrating Noise to Ensemble Disagreement

翻訳待ち:SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.05505v1 Announce Type: new Abstract: Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations of large language models (LLMs) typically examine review stages in isolation. We introduce SciLitBench, a multi-stage benchmark spanning title and abstract screening, full-text screening, and schema-guided data extraction, with 42,981 retrieved records, 1,012 full texts, and annotations for 888 included papers. Across 22 open-weight LLMs from six model families, explicit inclusion and exclusion criteria improve title and abstract screening $F_2$ by 28.8\%, while researcher-authored rationales improve full-text screening by 15\%. Data extraction reveals a different reliability regime: performance de…

arXiv AIサイト内本文翻訳待ち:SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews

翻訳待ち:AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.05446v1 Announce Type: new Abstract: We introduce AutoFyn, an agent harness inspired by the Expert Iteration algorithm, adapting a frozen model across many rounds by updating persistent state from verified reward signals rather than model weights. Each round begins from a fresh model session, and durable information is reintroduced only through explicit interfaces such as persistent memory files, reports, and repository state. Within a round, an orchestrator explores, plans and builds many alternative approaches with specialized agents, while a task-grounded verifier verifies the work and supplies an objective reward for measuring progress. This reward is distilled back into the persistent state, which updates the effective policy for the…

arXiv AIサイト内本文翻訳待ち:AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents

翻訳待ち:Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Learn how to deploy Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter open-weight model, on Amazon SageMaker HyperPod with vLLM. This walkthrough covers cluster provisioning, NVFP4 quantization, and an OpenAI-compatible endpoint with built-in reasoning, tool calling, and native MTP speculative decoding.

AWS Machine Learning Blogサイト内本文翻訳待ち:Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

翻訳待ち:[AINews] OpenAI reports Navier-Stokes singularity find in 88 hours using Astra-next, roughly 10,000 agents and 130B tokens (>$40M), a contender for second ever Millennium Prize awarded

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Overshadowing Cognition's $48B Series E, Mistral's $24B Series D, Meta's Muse agent, and GPT Image 2.5. The most jam packed, feel the AGI day in the history of AI.

Latent Spaceサイト内本文翻訳待ち:[AINews] OpenAI reports Navier-Stokes singularity find in 88 hours using Astra-next, roughly 10,000 agents and 130B tokens (>$40M), a contender for second ever Millennium Prize awarded

Mistralの新規資金調達、Sovereign AIへの架け橋に

MistralはオープンウェイトのAIスタートアップとして始まったが、欧州市場に対応するため、現在はSovereign AI(主権AI)へと戦略をシフトしている。今回の資金調達は、その移行を支える架け橋になると見られている。

AI Businessサイト内本文Mistralの新規資金調達、Sovereign AIへの架け橋に

Harbor AdaptersとHarbor-Index:大規模エージェント評価のためのインフラとキュレーション済みメタデータセット

本論文は、エージェント型ベンチマークのための統一評価インフラであるHarbor Adaptersを提案する。80以上のベンチマークを任意のエージェントで評価できるようにアダプタ化し、厳密なコードレビューとパリティ実験で検証した。さらに、8つのモデルを54のベンチマークで、Terminus-2および3種類のネイティブハーネスと組み合わせて大規模評価を実施。加えて、29のベンチマークから選抜した82の困難で多様な高品質タスクからなるHarbor-Indexも導入する。どのモデル・ハーネス構成も正解率30%を超えず、最強のGPT-5.5 with Codexでも28.0%にとどまり、評価資源はすべてオープンソースとして公開される。

arXiv AIサイト内本文Harbor AdaptersとHarbor-Index:大規模エージェント評価のためのインフラとキュレーション済みメタデータセット

Nous Research、Hermes Desktopにワンクリックのローカルモデルセットアップ機能を追加

Nous Researchは、Hermes Desktopでのローカルオープンウェイトモデルのセットアップをワンクリックに簡素化しました。アプリがハードウェアを読み取り、GPUに合うモデルを選択し、重みをダウンロードしてllama.cppを自動構成します。アカウント不要で利用可能。量子化は4ビット下限、推奨モデルは64K以上のコンテキストウィンドウを保証し、各モデルのメモリ適合状況を緑・黄・赤で表示します。

MarkTechPostサイト内本文Nous Research、Hermes Desktopにワンクリックのローカルモデルセットアップ機能を追加

プローブ汎化を部分空間選択とみなす:OOD詐欺検出に向けて

線形プローブは活性化から欺瞞などの概念を検出できるが、分布外への汎化が難しい。本研究はLlama-3.1-8B-Instructを用い、学習分布の活性化の少数の主成分だけを使う部分空間選択が有効であることを示し、LLM判定によるPC選択でオラクルとの性能差を縮小した。

arXiv Computational Linguisticsサイト内本文プローブ汎化を部分空間選択とみなす:OOD詐欺検出に向けて

Nvidia、アイドル状態のPCをつないで“個人AIデータセンター”にする無償ツール「PAIR」を発表

エヌビディアは、家庭ネットワーク上のアイドル状態にある互換PCを自動検出して接続し、ローカルAI推論やエージェント型ワークフローに活用できる無償のオープンソースソフトウェア「PAIR」を発表した。対応するのは GeForce RTX 20シリーズ以降、RTX Pro GPU、DGX Spark、Apple M4以降。ベータ版は本日からWindows、Linux、macOSで利用できる。

The Verge AIサイト内本文Nvidia、アイドル状態のPCをつないで“個人AIデータセンター”にする無償ツール「PAIR」を発表

エヌビディア、Hugging Faceを約130億ドルで買収へ

エヌビディアは、オープンソースAIモデルやデータセット、ツールのホスティングプラットフォームとして知られるHugging Faceを129.3億ドルで買収することで合意した。同プラットフォームは「AIのGitHub」とも呼ばれ、AI開発者同士のコラボレーションを支えている。クローズドAI大手が独自チップを開発する中、オープンソースAIエコシステムでの足場を固め、AIハードウェア市場での優位性を守る狙いとみられる。

The Verge AIサイト内本文エヌビディア、Hugging Faceを約130億ドルで買収へ

Omi Desktop

Omi Desktop は、オープンソースの AI ネックレス Omi の Mac 向けコンパニオンアプリで、見たり聞いたりしたことについて Mac に質問できます。

Product Hunt AIサイト内本文Omi Desktop

Perplexity が Mac 向けハイブリッド コンピューティングをリリース:クラウド エージェントがローカル モデルをオーケストレーション、デバイス上でゲート

Perplexity は Mac 向けハイブリッド コンピューティングを発表し、タスクをクラウドのフロンティア モデルとユーザーの Mac 上のローカル モデルに分割し、機密データがクラウドに送信される前に保護するデバイス上のプライバシー ゲートを備えました。この機能は Apple シリコン Mac の Pro、Max、Enterprise ユーザーが利用できます。Perplexity はまた、ゲートの背後にある分類器 PII-Tracer をオープンソース化し、ベンチマークで優れた性能を示しました。

MarkTechPostサイト内本文Perplexity が Mac 向けハイブリッド コンピューティングをリリース:クラウド エージェントがローカル モデルをオーケストレーション、デバイス上でゲート

Scientific Agent Skills: 研究エージェントのための手続き的知識ライブラリ

新しいarXiv論文は、科学分析における言語モデルエージェントの妥当性を向上させるために、ゲノミクス、ケモインフォマティクス、医用画像などを含む16の実践領域にわたる163の手続き的スキルを含むオープンソースライブラリ「Scientific Agent Skills」を紹介しています。各スキルはバージョン管理された人間が読める指示ファイルを中心としたディレクトリであり、エージェントはタスクに必要なときのみロードします。著者はタスクレベルの評価やホスト選択率を報告していません。

arXiv Computational Linguisticsサイト内本文Scientific Agent Skills: 研究エージェントのための手続き的知識ライブラリ

注意感度だけでは不十分:ファインチューニング下での注意レベルと行動レベルのインコンテキスト学習の乖離

本論文は、大規模言語モデルにおけるインコンテキスト学習(ICL)がファインチューニングによって損なわれる問題を扱い、注意パターンに基づく保存診断の信頼性に疑問を投げかけます。著者らは注意レベルのContext Sensitivity(ICS)と行動レベルのICL-GAPを定義し、Llama-2-7Bでの実験で、ICSを最大化する正則化によりICSがほぼ上限に達する一方、ICL-GAPはほぼゼロのままでMMLU精度が低下するというグッドハート的乖離を示しました。メカニズム分析では、注意が鋭くなりプレフィックス間でほぼ分離するものの、ラベルではなくフォーマットやデモ本体のトークンに向かうことが判明。ランダムラベル実験や構成的実験でも乖離が確認されました。結論として、注意レベルのICL指標は行動ギャップとの検証を経て初めて訓練目標として用いるべきです。

arXiv Machine Learningサイト内本文注意感度だけでは不十分:ファインチューニング下での注意レベルと行動レベルのインコンテキスト学習の乖離

REAL-Q: 動的勾配降下法によるエンドツーエンドLLM量子化

REAL-Qを提案。これは、動的ブロックワイズ勾配降下法とスライディングウィンドウ機構を用いてエンドツーエンドで整合した代理損失を直接最適化する新しい後訓練量子化(PTQ)手法であり、LLaMA-3.1およびQwen3モデルで最先端手法と比較してエンドツーエンドのKLダイバージェンスを最大約49%削減する。

arXiv Machine Learningサイト内本文REAL-Q: 動的勾配降下法によるエンドツーエンドLLM量子化

UI-Venus-2 テクニカルレポート

UI-Venus-2は、モバイル、ウェブ、デスクトップ環境で動作する汎用基盤GUIエージェントであり、統合されたクローズドループの推論・行動フレームワークを採用しています。環境(170以上の多言語モバイルアプリとネイティブデスクトップOS)、タスク(ディープリサーチパイプラインによる機能基盤の命令生成)、検証(視覚的キーポイントとマルチモデル投票を用いたトレースレベルおよびサンプルレベル評価)の3次元を拡張します。さらに、重大なアクションの制御された実行を保証する安全対策も統合しています。このオープンソースモデルは、実世界アプリケーション向けの汎用性が高く、検証可能で自己内省的なエージェントへの進歩を促進します。

arXiv AIサイト内本文UI-Venus-2 テクニカルレポート

命令チューニングされた小型言語モデルによる高齢者向け段階的金銭詐欺のインクリメンタルリスク評価

高齢者を標的とした金融詐欺は、電子メールやSMS、電話など複数ターンの会話を通じて行われることが増えており、リスクシグナルは各ターンで徐々に現れます。そのため、リソース制約のある環境で継続的にリスクを更新できるモデルが求められています。本研究では、累積ターンベースのリスク評価フレームワークを提案し、投資詐欺・慈善詐欺・テクニカルサポート詐欺のシナリオを含む多ターン対話データセットを構築しました。Phi-4、LLaMA-3.2、DeepSeek-R1、Qwen3の4つの小型言語モデルを微調整した結果、Phi-4とLLaMA-3.2がパラメータ規模に対して優れたターン認識リスク推定性能を示し、プライバシー保護とオンデバイス詐欺防止の可能性を示唆しています。

arXiv AIサイト内本文命令チューニングされた小型言語モデルによる高齢者向け段階的金銭詐欺のインクリメンタルリスク評価

I-CARE: テキストから画像へのモデルにおける、制御可能で多様かつ代表的なアンラーニング設定での干渉関連現象の分析

本論文は、生成的機械学習アンラーニングにおける干渉現象(保持されるべき意味的に関連する概念の意図しない性能低下)を体系的に研究するための手法I-CAREを提案する。新しいベンチマークやアルゴリズムを提案するのではなく、I-CAREはタスク、メトリクス、レポートテンプレートの形式的定義を提供し、さまざまなアンラーニング設定での干渉研究を標準化・再現可能にする。著者らは、最先端アルゴリズムと一般的なデータセットを用いてその実用性を実証し、オープンソースのソフトウェアとWebベースのグラフィカルインターフェースを提供する。

arXiv AIサイト内本文I-CARE: テキストから画像へのモデルにおける、制御可能で多様かつ代表的なアンラーニング設定での干渉関連現象の分析

Relaticle

Relaticleは、承認ゲート付きAIライティング機能を備えたオープンソースCRMで、AI生成コンテンツの公開前に人間のレビューを必須とします。

Product Hunt AIサイト内本文Relaticle

協力によって私たちはより強くなる

Databricksは外部セキュリティ研究者のMehmet Ince氏と協力し、オープンソースのPostGIS address_standardizer拡張機能におけるメモリ安全性の脆弱性を特定・緩和しました。この脆弱性はLakebase PostgresやNeonなどのマネージドPostgresプラットフォームのテナントがアクセス可能でしたが、DatabricksのマイクロVMアーキテクチャによりクロスカスタマーへの影響は防がれました。Databricksはこのサードパーティの脆弱性を自社の問題として迅速にパッチを展開し、最終的には修正が上流に貢献され、Postgresエコシステム全体のセキュリティが向上しました。

Databricks Blogサイト内本文協力によって私たちはより強くなる

Show HN: オープンソースのK8sネイティブAIプラットフォーム、分散マルチモデル推論用

shaideは、自社のKubernetesクラスター上で大規模にAIモデルを提供するセルフホスト型AIプラットフォームです。単一コマンドでインストールでき、ネットワーク境界内に完全に収まり、完全にエアギャップされた環境もサポートします。このプラットフォームはPulumiを使用してインフラをコードとして管理し、OpenAI互換APIを提供し、エージェント群向けに最適化されています。

Hacker News AIサイト内本文Show HN: オープンソースのK8sネイティブAIプラットフォーム、分散マルチモデル推論用

Foremerge: AIコーディングエージェント間の衝突をコーディング前に検出

Foremergeは、AIコーディングエージェントが編集前に意図を宣言することで、エージェント間の衝突を防ぐオープンソースの調整プロトコルです。Git上に構築され、共有データベースを使用して計画を意味的に比較し、エージェントがコーディングを開始する前に衝突を警告します。ファイルをロックしたり、モデルの判断に依存したりしません。

Hacker News AIサイト内本文Foremerge: AIコーディングエージェント間の衝突をコーディング前に検出

翻訳待ち:The Sequence Knowledge- Issue 924: The Distilled Models You Need to Know About

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:From DistilBERT and Gemini Flash to Gemma, Llama, Qwen, DeepSeek, Phi, Ministral, and PrismML’s Bonsai 27B.

TheSequenceサイト内本文翻訳待ち:The Sequence Knowledge- Issue 924: The Distilled Models You Need to Know About

翻訳待ち:5 Best Local LLMs You Can Run on a Mac mini in 2026

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Proprietary models are amazing! But sometimes what is of importance is configurability rather than raw power. This has led to the emergence of locally hosted models. The Mac mini has emerged as a surprisingly capable machine for running AI locally. With Apple Silicon, enough unified memory, and tools like Ollama and LM Studio, users can now run capable models entirely on-device. But […] The post 5 Best Local LLMs You Can Run on a Mac mini in 2026 appeared first on Analytics Vidhya.

Analytics Vidhyaサイト内本文翻訳待ち:5 Best Local LLMs You Can Run on a Mac mini in 2026

翻訳待ち:Improving Spatial-Temporal Reasoning in Video-Language Models with Structured Video Prompting

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.28666v1 Announce Type: new Abstract: Video-language models (VLMs) remain brittle on tasks that require tracking events over time and grounding answers in specific spatial regions. We propose that part of this limitation can be addressed through better organization of visual evidence at inference time. We introduce structured video prompting, a training-free inference-time method that augments the input video with lightweight spatial structure and temporal structure, providing explicit anchors for organizing evidence across space and time without changing model weights or decoding and without altering the question prompt in the main comparison. We evaluate this approach on two complementary video benchmarks and two open video-language mode…

arXiv Computer Visionサイト内本文翻訳待ち:Improving Spatial-Temporal Reasoning in Video-Language Models with Structured Video Prompting

翻訳待ち:Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.28623v1 Announce Type: new Abstract: Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating explicit chain-of-thought reasoning before answering. In language models it has been observed that this performance often comes with sycophancy, the tendency of a model to agree with the user over the evidence. However, for LMRMs no reliable method to measure sycophancy yet exists. We bridge this gap by introducing a benchmark and dataset for evaluating LMRM sycophancy when confronted with a wrong answer from a user. Our benchmark pairs four visually grounded datasets spanning mathematical, clinical, temporal, and demographic reasoning with five pressure conditions in single-turn and multi-turn sett…

arXiv Computational Linguisticsサイト内本文翻訳待ち:Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure

翻訳待ち:Gurukul AI: An Interactive AI-Driven Educational Platform for Indian Education System

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.28611v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) like ChatGPT and LLaMA have transformed AI-driven education, but these systems are predominantly trained on Western-centric data, making them ill-suited for regional curricula like India's. The Indian education system is linguistically diverse, exam-oriented, and structured around standardized syllabi, not addressed by existing datasets or tools. In this work, we curate a syllabus-aligned QA dataset based on NCERT (National Council of Educational Research and Training) textbooks for classes 9-12, capturing the content, context, and teaching style of Indian curricula. The final dataset, comprising 18,720 question-answer pairs across five subjects, is publi…

arXiv Computational Linguisticsサイト内本文翻訳待ち:Gurukul AI: An Interactive AI-Driven Educational Platform for Indian Education System

翻訳待ち:SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.28911v1 Announce Type: new Abstract: The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length. We show that uniform KV quantization on a fractional-bit grid does not degrade gracefully: under a prespecified multi-seed statistical protocol, Llama-3.1-8B-Instruct with an affine quantizer is statistically indistinguishable from FP16 KV down to 2.322 code bits/value and collapses at 2.0 bits - a quality cliff in (2.0, 2.322] that reappears in generation-time quantization and multi-turn dialogue and transfers to Mistral-7B. The cliff reframes importance-aware mixed precision: above it, eight model-internal importance indicators are statistically interc…

arXiv Machine Learningサイト内本文翻訳待ち:SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

翻訳待ち:The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.28859v1 Announce Type: new Abstract: Reasoning models do not stop when they know the answer. On DeepSeek-R1-Distill-Qwen-7B the chain of thought runs about twice as long as the model's own answer probability takes to settle, and how much of that excess is removable varies from problem to problem, so a global length penalty cannot take it out. We take it out by internalizing a causal interpretability finding into the weights. The mechanism is a halt vector: a difference-of-means direction at layer 18 of this model whose steering strength controls how long it thinks, while a replicated value axis does nothing. Installing that intervention in the weights is harder than it looks. Maximizing the scalar projection onto the direction corrupts th…

arXiv Machine Learningサイト内本文翻訳待ち:The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning

翻訳待ち:Curvature Cryptanalysis of Smooth Transformer Feed-Forward Networks

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.28843v1 Announce Type: new Abstract: We show that smooth two-layer feed-forward networks (FFNs) expose an additional structural model extraction channel under a chosen-input raw-output oracle at the FFN branch; consider transformer FFN branches with GELU or SiLU activations under chosen-input raw-output access, without access to parameters, gradients, or internal activations; exploit a second-order leakage channel in which projected input Hessians form different mixtures of the same hidden symmetric rank-one factors induced by the FFN input weights. We formalize resulting Hessian collection as a partially symmetric decomposition to establish conditions for local identifiability and stability to exploit vector-output stencil reuse to reduc…

arXiv Machine Learningサイト内本文翻訳待ち:Curvature Cryptanalysis of Smooth Transformer Feed-Forward Networks

翻訳待ち:A collective capability boundary in frontier large language models on guideline-conformant and case-specific oncology decision-making

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.28592v1 Announce Type: new Abstract: Large language models (LLMs) achieve high scores on medical knowledge examinations, yet real-world oncology is not a knowledge test--it is a sequence of guideline-pathway choices, escalation judgments, and commitments under uncertainty. Existing benchmarks largely measure factual recall, leaving open whether frontier LLMs share decision-path blind spots that combining models cannot fix. We built the Oncology Decision Boundary Benchmark (ODBB)--2,005 oncology decision points across NCCN guidelines and colorectal cancer cases--and evaluated nine frontier LLMs (four closed-source, five open-weight families) released between June 2025 and April 2026. A fully deterministic scorer (zero LLM inference) classi…

arXiv AIサイト内本文翻訳待ち:A collective capability boundary in frontier large language models on guideline-conformant and case-specific oncology decision-making

翻訳待ち:Expert-validated STEM QA

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.28591v1 Announce Type: new Abstract: Recent advancements in AI are helping scientists achieve breakthroughs in fields such as mathematics, medicine, and materials sciences. New evaluation datasets for AI models contribute to such advancement in AI. In the STEM domain, frontier models have consumed most of the available online data, creating the need for human-created datasets that codify the knowledge of leading experts in the domain. There are several STEM datasets available for the research community in this field. However, there are some gaps in these datasets, leaving room for improvement. Examples of gaps include (1) saturation in model performance on these datasets, leaving no head-room for meaningful evaluations, (2) skewed taxonom…

arXiv AIサイト内本文翻訳待ち:Expert-validated STEM QA

翻訳待ち:Google AI Releases TimesFM-3: A 330M Parameter Zero-Shot Foundation Model For Multivariate Time Series Forecasting

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Google Research has released TimesFM-3, a 330 million parameter time series foundation model that forecasts multiple related series in a single forward pass. Unlike every TimesFM checkpoint through 2.5, it is pretrained natively for multivariate forecasting, accepting multiple targets, past covariates, and past-future covariates with no task-specific fine-tuning. It takes the top average rank among pretrained foundation models on GIFT-Eval, fev-bench, and the TIME leaderboard. The weights, however, ship under a non-commercial, non-production license. The post Google AI Releases TimesFM-3: A 330M Parameter Zero-Shot Foundation Model For Multivariate Time Series Forecasting appeared first on MarkTechPost.

MarkTechPostサイト内本文翻訳待ち:Google AI Releases TimesFM-3: A 330M Parameter Zero-Shot Foundation Model For Multivariate Time Series Forecasting

翻訳待ち:Speed Up LLM Inference with DSpark Speculative Decoding

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Learn how DSpark speculative decoding can improve local LLM generation speed using the same GPU, with Qwen3-8B, llama.cpp, and CUDA.

KDnuggetsサイト内本文翻訳待ち:Speed Up LLM Inference with DSpark Speculative Decoding

翻訳待ち:Accelerating LLM Inference via Vector Index Based Output Embeddings

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.27460v1 Announce Type: new Abstract: Large output embedding matrices create a significant memory bandwidth bottleneck during autoregressive decoding, especially for compact LLMs with large multilingual vocabularies. We reformulate the output projection followed by top-k token selection as a maximum inner product search over token embeddings and replace the dense vocabulary projection with an HNSW-based vector index. The resulting output head retrieves only a small candidate set of high-scoring tokens and can be integrated into existing decoding pipelines by scattering retrieved logits into a sparse full-vocabulary tensor. On CPU inference with Gemma 3, Llama 3.2, and Qwen 3 models, our method substantially accelerates the output projectio…

arXiv Computational Linguisticsサイト内本文翻訳待ち:Accelerating LLM Inference via Vector Index Based Output Embeddings

翻訳待ち:Ollama's transparent pricing

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Ollama's Pro, Max, and Team plans now use industry-standard per-token pricing with usage included on every plan.

Ollama Blogサイト内本文翻訳待ち:Ollama's transparent pricing

翻訳待ち:90 days of attacks on AI infrastructure

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Wiz PricingGet a demo Get a demo Wiz Threat Research operates honeypots across AI and ML services including LiteLLM, Flowise, LangChain, Langflow, ChromaDB, Ollama, and others. Over 90 days of telemetry, we observed sus…

Hacker News AIサイト内本文翻訳待ち:90 days of attacks on AI infrastructure

その他の成長タグ

オープンソースモデル AI News | AI News Hub