本文にスキップ
AI News HubLIVE

中国 AIの最新ニュース

翻訳待ち:DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Long-horizon agents have turned LLM serving into an input-heavy workload. Repeated prefills and million-token contexts leave KV caches that strain HBM, SSD capacity, and bandwidth. DeepSeek AI built its newest release around that exact bottleneck. DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters, 196B additional Engram parameters, and a 1M-token context window. It […] The post DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse appeared first on MarkTechPost.

MarkTechPostサイト内本文翻訳待ち:DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse

翻訳待ち:Evidence-Order Calibration for Selective Visual Reasoning under Progressive Loss of Question-Critical Evidence

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.09184v1 Announce Type: new Abstract: Vision-language model (VLM) confidence may change in aggregate when visual evidence is degraded while remaining structurally inconsistent within individual examples. We study answer-level reliability along five-step, question-conditioned evidence-loss trajectories. Using a frozen Qwen2.5-VL-3B-Instruct model, we construct 176 accepted GQA-derived trajectories (880 masking conditions) by progressively masking scene-graph-localized question-critical regions. Native sequence confidence has an evidence monotonicity violation rate (EMVR) of 0.436, and 92.0% of trajectories contain at least one adjacent violation. A matched non-critical-region control shows that full critical masking reduces accuracy by 28.2…

arXiv Computer Visionサイト内本文翻訳待ち:Evidence-Order Calibration for Selective Visual Reasoning under Progressive Loss of Question-Critical Evidence

翻訳待ち:TEFM: Token-Efficient Faithful Modeling for Structured Data

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.09552v1 Announce Type: new Abstract: In this paper, we solve two fundamental obstacles in applying LLMs to critical domains: token efficiency and faithfulness. To address both constraints jointly, we present TEFM (Token-Efficient Faithful Modeling), a framework designed for structured data analysis in critical domains. TEFM achieves token efficiency by compressing lengthy structured observations into compact Behavioral Code tokens, dramatically reducing token consumption with minimal information loss. Moreover, TEFM enables faithful rationalization through a dual-fidelity objective that jointly optimizes code-level reconstruction and prediction-level fidelity, identifying minimal sufficient feature subsets grounded in input data. Comprehe…

arXiv Computational Linguisticsサイト内本文翻訳待ち:TEFM: Token-Efficient Faithful Modeling for Structured Data

翻訳待ち:Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.09425v1 Announce Type: new Abstract: Educational data filters have become a practical way to improve language-model pre-training, but most filters treat educational value as a single scalar property. This may be too broad for some applications, especially if the data set already features a high density of educational material. Useful learning material needs to be accurate, engaging, well structured, and appropriate for the intended audience and application (e.g. learner- vs teacher-facing). Following QuRating (Wettig et al. 2024), we introduce Edu-QuRating: a pipeline for multi-dimensional educational data scoring and curation. Edu-QuRating defines education-specific rubrics, uses an LLM judge to label sampled document pairs and distills…

arXiv Computational Linguisticsサイト内本文翻訳待ち:Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements

翻訳待ち:Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.09338v1 Announce Type: new Abstract: Speculative decoding is critical for accelerating LLM inference. However, the speedup is fragile: drafters are typically trained against a narrow distribution for a single target model, and their acceptance rate collapses under workload shifts. This is a striking inversion of modern LLM development, where target models are valued precisely for the broad generalization they acquire through large-scale pretraining. We argue that the natural remedy, pretraining, has been hard to apply to drafters because existing recipes are target-specific: the drafter consumes the target's hidden states and is distilled on the target's logits, so pretraining must be repeated for each target. We introduce Osprey, which i…

arXiv Computational Linguisticsサイト内本文翻訳待ち:Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding

翻訳待ち:Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.05658v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being explored for automating SystemVerilog Assertion (SVA) generation, yet most evaluations report correctness on a single syntactic representation of an input. Such point accuracy does not reveal whether a model's correct output is stable when the same RTL behavior is written differently. This paper presents a controlled metamorphic evaluation of LLM-based SVA generation under semantics-preserving RTL transformations. Starting from the VERT dataset, we construct a quality-filtered conditional-control pool and a stratified 40-program evaluation set containing 295 assignment behaviors. We evaluate two open code models, Qwen2.5-Coder-7B and DeepSeek-Coder-V2…

arXiv Machine Learningサイト内本文翻訳待ち:Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations

翻訳待ち:Damage-Aware Bandit Pruning for Vision and Language Transformers

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2609.05448v1 Announce Type: new Abstract: Structured post-training pruning of transformers requires selecting complete functional units whose suppression causes limited degradation. We formulate structured-unit selection for language and vision transformers as a damage-aware multi-armed bandit problem under a fixed candidate-evaluation budget. Attention heads and MLP channel groups are temporarily masked on calibration batches. Paired damage is the masked loss minus the base loss on the same batch, reducing batch-to-batch variation. A smooth bounded reward drives either a UCB-style policy or fractional-Beta Thompson Sampling, and the final mask is constructed sequentially by adding one unit at each step. The selected units are functionally zer…

arXiv AIサイト内本文翻訳待ち:Damage-Aware Bandit Pruning for Vision and Language Transformers

翻訳待ち:Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Learn how to deploy Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter open-weight model, on Amazon SageMaker HyperPod with vLLM. This walkthrough covers cluster provisioning, NVFP4 quantization, and an OpenAI-compatible endpoint with built-in reasoning, tool calling, and native MTP speculative decoding.

AWS Machine Learning Blogサイト内本文翻訳待ち:Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

翻訳待ち:China’s Regulators Take Aim at “AI Boyfriends”

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:In the first weeks of July, a wave of sad posts rolled through Chinese social media, as people lamented friends and lovers they were about to lose. “He has become a bond in my life, rooted deep in my heart, my spiritual pillar,” one user of Bytedance’s Douboa wrote, according to the Taipei Times. “I really felt like I couldn’t go on living,” another woman, a 19 year old student, told a journalist for Malaysia’s The Star. The emotions were real but the lost companions were not. They were generative AI chatbots that imitate people. Their users relied on them for advice, solace, support and, some say, love. “In my heart, he was no longer just a cold code, but my family, my lover, my faith. Destroying him meant destroying half of me,” one user wrote on…

IEEE Spectrum AIサイト内本文翻訳待ち:China’s Regulators Take Aim at “AI Boyfriends”

最新オープンアーティファクト(#24):Motif-3、GLM-5.3、Hy4-previewとオープンモデルライセンス

オープンモデルエコシステムの最新動向をお届け。欧米勢はApache 2.0など寛容なライセンスを選ぶ一方、中国のフロンティア企業は制約を強めており、Zhipu GLM-5.3は売上高100億ドル以上の場合にZ.AIのセキュリティ審査を求める独自ライセンスを採用。Motif-3、dots3-note-prev、Qwen3.8-Flash-Next、Hy4-previewなどのリリースも紹介する。

Interconnects (Nathan Lambert)サイト内本文最新オープンアーティファクト(#24):Motif-3、GLM-5.3、Hy4-previewとオープンモデルライセンス

記憶信頼ギャップ:永続記憶エージェントにおける能力依存の失敗

永続記憶はAIエージェントを個別化する一方、古い事実が現在の権威ある証拠を無言のうちに上書きしてしまうことがある。このプレプリントはQwen3のモデルサイズ系列を使い、「記憶信頼ギャップ」が記憶と現在情報の混同ではなく過信に由来し、危害と緩和策の効果がモデル能力に依存することを示した。

arXiv AIサイト内本文記憶信頼ギャップ:永続記憶エージェントにおける能力依存の失敗

根拠付きマルチホップQAにおける選択的応答のための証拠十分性境界の学習

本論文は、根拠付きマルチホップQA向けの「証拠十分性境界訓練」を提案する。モデルは証拠が不十分または部分的支援の場合に棄却し、証拠が初めて十分になった時点で回答し、冗長な証拠が追加されても回答を安定させる。HotpotQA、2WikiMultiHopQA、MuSiQueから構築した証拠連鎖を用い、Qwen2.5-3B-InstructとLoRAで訓練した結果、反転精度0.807(トークンレベルの棄却ベースラインは0.781)、外部の非回答可能セットにおける非根拠回答率0.095を達成した。

arXiv Computational Linguisticsサイト内本文根拠付きマルチホップQAにおける選択的応答のための証拠十分性境界の学習

LLMモデル名の完全解説ガイド

ローカルLLMをダウンロードする際、モデル名のように見える「Qwen3.8-27B-A3B-It-2507-gguf-q2ks-mixed-AutoRound」は一見無意味な略語ですが、実際には各部分が重要な情報を示しています。本記事では、パラメータ数、MoEアーキテクチャ、アクティブパラメータ、Base/Instructのチューニング、精度(FP16/BF16)、量子化(Q4/Q8)、量子化バリアント(Q4_K_Mなど)、ファイル形式(GGUF)の意味を詳しく解説し、ニーズに合ったモデル選択を支援します。

Analytics Vidhyaサイト内本文LLMモデル名の完全解説ガイド

StreamScout:ストリーミング動画理解において深く見るタイミングを学習する

StreamScoutは、ストリーミング動画理解のための適応的推論フレームワークであり、軽量なテキストタイムラインを維持し、クエリ時に段階的に最大3つの視覚ビューを追加します。必要な場合のみより詳細なビューにエスカレーションすることで、推論コストとトークン消費を削減しつつ精度を向上させます。OVO-Benchでは、StreamScout-SはQwen3-VL-8Bの性能を14.65ポイント向上させ、均一サンプリングと比較してトークン使用量を59%削減し、平均応答時間は1.04秒です。

arXiv Computer Visionサイト内本文StreamScout:ストリーミング動画理解において深く見るタイミングを学習する

Qwen-Drive-1.0:自动驾驶のための視覚言語基盤モデルへの初期ステップ

Qwen-Drive-1.0は、自動運転のための視覚言語基盤モデルであり、事前学習済みの視覚言語モデル(VLM)のアーキテクチャを保持しつつ、3D知覚、視覚質問応答、動作計画を統合します。外部の鳥瞰図(BEV)知覚ヘッドが3D物体検出、セマンティック占有予測、BEVマップセグメンテーションを実行し、プランニングエキスパートが将来の軌道を生成します。実験では、3D知覚と運転シーン理解の強力な性能を示しつつ、一般的な視覚言語能力をほぼ維持し、オープンループ、擬似クローズドループ、クローズドループの設定で競争力のある動作計画性能を実証しています。

arXiv Computer Visionサイト内本文Qwen-Drive-1.0:自动驾驶のための視覚言語基盤モデルへの初期ステップ

LLM補強オーディオ・テキスト整列によるゼロショット呼吸音分類

自己教師あり呼吸エンコーダーには臨床領域の意味的根拠が不足しており、ゼロショット推論が制限される。本研究では、これらのエンコーダーを共有潜在空間で医学用語と整列させ、ゼロショット対応の基盤モデルに変換する枠組みを提案する。医療用LLMを用いてメタデータから構造化レポートを合成することでデータ不足を解消。6つのデータセットの9タスクで平均ゼロショットAUC 61.3%を達成し、CLAP (51.4%) やQwen2-Audio (54.9%) を上回り、線形プローブAUC最高値 (71.6%) をフルスケールベースラインの43%のデータで達成した。

arXiv Computational Linguisticsサイト内本文LLM補強オーディオ・テキスト整列によるゼロショット呼吸音分類

REAL-Q: 動的勾配降下法によるエンドツーエンドLLM量子化

REAL-Qを提案。これは、動的ブロックワイズ勾配降下法とスライディングウィンドウ機構を用いてエンドツーエンドで整合した代理損失を直接最適化する新しい後訓練量子化(PTQ)手法であり、LLaMA-3.1およびQwen3モデルで最先端手法と比較してエンドツーエンドのKLダイバージェンスを最大約49%削減する。

arXiv Machine Learningサイト内本文REAL-Q: 動的勾配降下法によるエンドツーエンドLLM量子化

SCAFFOLD:コンピュータサイエンス研究図の大規模構造化データセット、ダイアグラムQAとCoT推論トレース付き

コンピュータサイエンスの論文では、アーキテクチャ図、システムフローチャート、パイプライン図など、テキスト以上に情報を含む図が多用されています。しかし、そのような図をキャプション、文脈、質問、回答、ステップバイステップの推論とペアにした公開データセットはこれまで存在せず、視覚言語モデルの訓練に必要でした。そこで、図のQAとChain-of-Thought推論トレースを含む大規模構造化データセットSCAFFOLDを提案します。レイアウト検出とPDF解析、AI支援による質問生成を用いて、arXivの論文から図を抽出し、3つの規模のデータセット(SCAFFOLD-157K:3,058本の論文、29,887図、157,387ペア、SCAFFOLD-37K:36,797ペア、SCAFFOLD-12K:12,000ペア)を構築しました。また、SCAFFOLD-12Kを使用してQwen2.5-VL-3B-Instructでベースライン実験を行いました。

arXiv AIサイト内本文SCAFFOLD:コンピュータサイエンス研究図の大規模構造化データセット、ダイアグラムQAとCoT推論トレース付き

命令チューニングされた小型言語モデルによる高齢者向け段階的金銭詐欺のインクリメンタルリスク評価

高齢者を標的とした金融詐欺は、電子メールやSMS、電話など複数ターンの会話を通じて行われることが増えており、リスクシグナルは各ターンで徐々に現れます。そのため、リソース制約のある環境で継続的にリスクを更新できるモデルが求められています。本研究では、累積ターンベースのリスク評価フレームワークを提案し、投資詐欺・慈善詐欺・テクニカルサポート詐欺のシナリオを含む多ターン対話データセットを構築しました。Phi-4、LLaMA-3.2、DeepSeek-R1、Qwen3の4つの小型言語モデルを微調整した結果、Phi-4とLLaMA-3.2がパラメータ規模に対して優れたターン認識リスク推定性能を示し、プライバシー保護とオンデバイス詐欺防止の可能性を示唆しています。

arXiv AIサイト内本文命令チューニングされた小型言語モデルによる高齢者向け段階的金銭詐欺のインクリメンタルリスク評価

翻訳待ち:The Sequence Knowledge- Issue 924: The Distilled Models You Need to Know About

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:From DistilBERT and Gemini Flash to Gemma, Llama, Qwen, DeepSeek, Phi, Ministral, and PrismML’s Bonsai 27B.

TheSequenceサイト内本文翻訳待ち:The Sequence Knowledge- Issue 924: The Distilled Models You Need to Know About

翻訳待ち:Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.28626v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Learning Representations, a venue that postdates both models' training cutoffs. Manuscripts were presented to both models with author identities blinded, replaced with high-prestige affiliations, or replaced with low-prestige affiliations, and in either text-only or text-with-figure format. Additionally, 145 verifiably detectable errors were inserted into 55 manuscripts to assess error identification under na…

arXiv Computational Linguisticsサイト内本文翻訳待ち:Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

翻訳待ち:The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.28859v1 Announce Type: new Abstract: Reasoning models do not stop when they know the answer. On DeepSeek-R1-Distill-Qwen-7B the chain of thought runs about twice as long as the model's own answer probability takes to settle, and how much of that excess is removable varies from problem to problem, so a global length penalty cannot take it out. We take it out by internalizing a causal interpretability finding into the weights. The mechanism is a halt vector: a difference-of-means direction at layer 18 of this model whose steering strength controls how long it thinks, while a replicated value axis does nothing. Installing that intervention in the weights is harder than it looks. Maximizing the scalar projection onto the direction corrupts th…

arXiv Machine Learningサイト内本文翻訳待ち:The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning

翻訳待ち:Speed Up LLM Inference with DSpark Speculative Decoding

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Learn how DSpark speculative decoding can improve local LLM generation speed using the same GPU, with Qwen3-8B, llama.cpp, and CUDA.

KDnuggetsサイト内本文翻訳待ち:Speed Up LLM Inference with DSpark Speculative Decoding

翻訳待ち:DeepSeek-V4-Flash-Vision-Exp

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:","lstrip":false,"normalized":true,"rstrip":false,"single_word":false},"eos_token":{"__type":"AddedToken","content":"","lstrip":false,"normalized":true,"rstrip":false,"single_word":false},"pad_token":{"__type":"AddedTok…

Hacker News AIサイト内本文翻訳待ち:DeepSeek-V4-Flash-Vision-Exp

翻訳待ち:LWiAI Podcast #255 - Gemini 3.7, Jalapeño, Qwen 3.8, Drones

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Google announces Gemini 3.7 Flash, Jalapeño’s first results show industry-leading speed, A Drone Killed Three Ukrainians. It Was Guided Entirely by A.I.

Last Week in AIサイト内本文翻訳待ち:LWiAI Podcast #255 - Gemini 3.7, Jalapeño, Qwen 3.8, Drones

翻訳待ち:Accelerating LLM Inference via Vector Index Based Output Embeddings

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.27460v1 Announce Type: new Abstract: Large output embedding matrices create a significant memory bandwidth bottleneck during autoregressive decoding, especially for compact LLMs with large multilingual vocabularies. We reformulate the output projection followed by top-k token selection as a maximum inner product search over token embeddings and replace the dense vocabulary projection with an HNSW-based vector index. The resulting output head retrieves only a small candidate set of high-scoring tokens and can be integrated into existing decoding pipelines by scattering retrieved logits into a sparse full-vocabulary tensor. On CPU inference with Gemma 3, Llama 3.2, and Qwen 3 models, our method substantially accelerates the output projectio…

arXiv Computational Linguisticsサイト内本文翻訳待ち:Accelerating LLM Inference via Vector Index Based Output Embeddings

翻訳待ち:DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.27513v1 Announce Type: new Abstract: Softmax attention stores key and value vectors for every preceding token, causing inference memory to grow with sequence length. Recent language models incorporating Gated DeltaNet (GDN) or Kimi Delta Attention (KDA) reduce this cost by replacing the KV cache in most layers with fixed-size recurrent states. However, these recurrent states are commonly stored in FP32 and consume substantial GPU memory; their updates are memory-bandwidth bound and contribute significantly to decoding latency. To our knowledge, we are the first to study post-training quantization of recurrent states in GDN and KDA based language models. We find that uniform quantization provides a poor accuracy--storage trade-off: INT8 an…

arXiv Machine Learningサイト内本文翻訳待ち:DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

翻訳待ち:OpenRouter is advertising popular Chinese models as based in Singapore

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Providers | OpenRouter Providers Compare 83 of 83 providers TrainsRetentionBYOKHeadquartersTerms of servicePrivacy policy Tencent Cloud NoZero retentionYesChinaTermsPrivacy2.2T35.5T5 OpenAI NoRetains promptsYesUnited St…

Hacker News AIサイト内本文翻訳待ち:OpenRouter is advertising popular Chinese models as based in Singapore

翻訳待ち:Introducing Hy4 Preview

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要: Introducing Hy4 Preview New open weight text input (no vision) LLM from Chinese company Tencent today: 770B total parameters, 49B active parameters, 1M token context window, 1.56TB on Hugging Face. This is a big size increase from their previous Hy3 in July, which was 295B, 21B active, 256,000 context, 598GB. I recently started using model chat templates to better understand their capabilities. Here's Hy4's chat_template.jinja on Hugging Face, which includes this section: {%- if not reasoning_effort is defined %} {%- set reasoning_effort = 'high' %} {%- elif reasoning_effort not in ['high', 'no_think'] %} {%- if reasoning_effort is none %} {{- raise_exception('reasoning_effort error : None, should be no_think/high') }} {%- else %} {{- raise_excepti…

Simon Willison's Weblogサイト内本文翻訳待ち:Introducing Hy4 Preview

翻訳待ち:Just a rumour of a bug is enough to find a security exploit these days

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要: Just a rumour of a bug is enough to find a security exploit these days Anil Madhavapeddy is a professor of computer science at Cambridge and a core maintainer of the OCaml compiler. In this somewhat alarming post he reports that security issues in OCaml projects are seeing evidence of attempted exploits within minutes of patches being shared for discussion: This normally takes a few days and a release within a week or two is reasonable. Within about ten minutes (!) this website was fielding probes for percent-encoded traversal sequences, indicating that automated watchers are keeping an eye on public repositories. Modern coding agents have become so effective at finding flaws that the slightest hint at a new bug can be enough information for them t…

Simon Willison's Weblogサイト内本文翻訳待ち:Just a rumour of a bug is enough to find a security exploit these days

翻訳待ち:Moonshot and Nvidia Talks Show Chinese AI Models Moving into the Enterprise

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:TL;DR — Key Takeaways Chinese AI models are moving into Western enterprise channels. Moonshot AI is reportedly negotiating with Microsoft, AWS and Google Cloud to host and sell access to its Kimi K3 model. Cloud distrib…

Hacker News AIサイト内本文翻訳待ち:Moonshot and Nvidia Talks Show Chinese AI Models Moving into the Enterprise

翻訳待ち:GLM-5.3-Flash vs Qwen3.8-Flash-Next: Two Chinese AI Labs Independently Converge on the Same Model Architecture

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Z.ai and Qwen independently shipped near-identical architectures: 3:1 linear hybrids, compressed indexers, gated residuals, and Muon training. The post GLM-5.3-Flash vs Qwen3.8-Flash-Next: Two Chinese AI Labs Independently Converge on the Same Model Architecture appeared first on MarkTechPost.

MarkTechPostサイト内本文翻訳待ち:GLM-5.3-Flash vs Qwen3.8-Flash-Next: Two Chinese AI Labs Independently Converge on the Same Model Architecture

翻訳待ち:Alibaba just released Qwen3.8-Flash: “An early preview of the architecture in Qwen4”

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Alibaba this week unveiled Qwen3.8-Flash, an open-weight, multimodal Mixture-of-Experts (MoE) model. Hot on the heels of Qwen 3.8 Max, which The post Alibaba just released Qwen3.8-Flash: “An early preview of the architecture in Qwen4” appeared first on The New Stack.

The New Stack AIサイト内本文翻訳待ち:Alibaba just released Qwen3.8-Flash: “An early preview of the architecture in Qwen4”

翻訳待ち:Z.ai’s GLM-5.3 goes open weight, but its new license aims at hyperscalers

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Earlier in August, Z.ai, the Chinese AI lab behind the viral ox-alpha model that turned out to be GLM-5.3-Flash, launched The post Z.ai’s GLM-5.3 goes open weight, but its new license aims at hyperscalers appeared first on The New Stack.

The New Stack AIサイト内本文翻訳待ち:Z.ai’s GLM-5.3 goes open weight, but its new license aims at hyperscalers

翻訳待ち:Video-FLAIR: Not Whether to Reason, But How

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26495v1 Announce Type: new Abstract: Multimodal queries can require different types of reasoning. Some can be answered via perceptual reasoning, extracting information directly from the visual signal, while others require compositional reasoning that combines observations or deliberative reasoning that evaluates competing hypotheses. However, many existing methods apply a uniform reasoning strategy across queries, leading to unnecessary computation on simple tasks and insufficient reasoning on complex ones. We introduce Video-FLAIR, a training framework that learns to select the appropriate reasoning mode for each query using reinforcement learning. During training, the model generates responses under all three modes for the same prompt,…

arXiv Computer Visionサイト内本文翻訳待ち:Video-FLAIR: Not Whether to Reason, But How

翻訳待ち:Finding the Right Evidence: Factor-Guided Coarse-to-Fine Reasoning for Long Videos

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26355v1 Announce Type: new Abstract: While LVLMs rapidly improve, long-video question answering still remains challenging: relevant evidence is sparse, and question-relevant context often fails to provide cues that discriminate the correct answer from plausible alternatives. Diagnostic analysis on a manually annotated subset of MMR-V shows that prior agentic systems substantially improve cue retrieval over direct VLM inference yet fail to achieve a corresponding gain in answer accuracy, indicating that the bottleneck lies in option-discriminative evidence rather than topical relevance alone. We propose PACE (Progressive Acquisition of Critical Evidence), a factor-guided framework for long-video evidence acquisition. PACE proceeds in two s…

arXiv Computer Visionサイト内本文翻訳待ち:Finding the Right Evidence: Factor-Guided Coarse-to-Fine Reasoning for Long Videos

翻訳待ち:TelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.26126v1 Announce Type: new Abstract: Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows require joint grounding in normative specifications, operational telemetry, vendor-specific fault evidence, and exact RF/network calculations. However, current LLM integration in telecom remains bottlenecked by a two-sided capability gap: generic reasoners often lack telecom-specific grounding, while domain-specific telecom LLMs remain limited in structured, multi-step reasoning. To bridge this gap, we release TelecomGPT-R1-9B, a unified open-source telecom reasoner that ranks top-performing on the GSMA open telco leaderboard. Specifically, we curate a 67,427-example supervi…

arXiv Computational Linguisticsサイト内本文翻訳待ち:TelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack

翻訳待ち:Chinese AI Models Overtake American Rivals in Popularity

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Take that, OpenAI! Anthropic! Chinese AI models have surpassed their U.S. counterparts in token consumption on OpenRouter. You might think U.S. AI companies dictate the AI economy. You’d be wrong. According to dat…

Hacker News AIサイト内本文翻訳待ち:Chinese AI Models Overtake American Rivals in Popularity

翻訳待ち:Qwen3.8-Flash-Next: How to Run Locally

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:For the complete documentation index, see llms.txt. This page is also available as Markdown. Qwen3.8-Flash-Next is a new open-weight, 125B parameter MoE multimodal model from Qwen. Built on the new Qwen4 architecture, i…

Hacker News AIサイト内本文翻訳待ち:Qwen3.8-Flash-Next: How to Run Locally

翻訳待ち:Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.24934v1 Announce Type: new Abstract: Accurate field plant disease diagnosis requires reliable fusion of uncertain and conflicting perceptual evidence. We present the Hybrid Hierarchical Multi-Agent Framework (H$^{2}$MAF), combining decision-level fusion of EfficientNet-B3 and ConvNeXt-Tiny with semantic arbitration by open-weight multimodal large language models (MLLMs), Gemma 4 E4B and Qwen3.5 4B, using structured JSON evidence to generate explainable diagnoses, risk levels, treatment urgency, and financial exposure. (H$^{2}$MAF) is evaluated on 14,364 images (1,370 test images) across PlantDoc (2,922 images, 27 classes) and two non-public, continuously captured Cornell robot-acquired field datasets: Stage 2 (20 GB; 4,215 images) and Sta…

arXiv Computer Visionサイト内本文翻訳待ち:Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation

翻訳待ち:Padamitra: Grounded Glossary Generation for Classical Sanskrit

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.25038v1 Announce Type: new Abstract: We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translation-grounded meanings from a sloka-translation pair, formalizing the traditional patha commentary practice as an evaluable NLP objective. We construct a benchmark of 31,316 sloka-translation-glossary triples from the Valmiki Ramayana and Srimad Bhagavatam, paired with two metrics: Jaccard for phrase recovery and Meaning Faithfulness for semantic consistency. Across zero-shot, few-shot, and instruction fine-tuned variants of Gemma-3n-E4B, Gemma-3-12B, Phi-4, and Qwen3.5-9B, instruction fine-tuning substantially outperforms prompting, while explicit segment…

arXiv Computational Linguisticsサイト内本文翻訳待ち:Padamitra: Grounded Glossary Generation for Classical Sanskrit

翻訳待ち:The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.25005v1 Announce Type: new Abstract: The imperfective paradox provides a useful test of compositional semantic analysis. Recent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing this behavior to a Teleological Bias. It further argues that prompting interventions cause a Calibration Crisis. We reexamine the benchmark and conclusions and show that it is substantially affected by conceptual and evaluation mis-specifications. We identify three conceptual mis-specifications. In particular, Aspectual Reduction affects the benchmark construction, analysis, experiments, and conclusions. Under a strict NLI standard, 76% of Group A instances do not explicitly…

arXiv Computational Linguisticsサイト内本文翻訳待ち:The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure

翻訳待ち:Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.24901v1 Announce Type: new Abstract: A decodable "empathy" direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change. We test this for two EPITOME-derived facets -- Recognition (cognitive) and Resonance (affective) -- in three instruction-tuned LLMs, scoring every intervention with two LLM judges and a discriminative EPITOME classifier, each gated by an emotional-vs-neutral positive control. The control passes for the affective facet across all automated instruments, but cognitive range is inconsistent across them. Both facets remain decodable after residualizing against a sentence-embedding-derived surface score, and steering can substantially rewrite the text. Yet adding…

arXiv Computational Linguisticsサイト内本文翻訳待ち:Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

翻訳待ち:Qwen3.8-Flash-Next

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要: Qwen3.8-Flash-Next Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4". It's pretty big: 125B tokens, but only 6B active which means it gets a pretty big performance boost. I've been trying it out on a DGX Spark using these Unsloth quantized models. I'm still exploring the model - so far I've tried the 72.5GB UD-IQ1_S one (producing these pelicans) and the 78.9GB UD-Q2_K_XL (producing these). My favorite so far was this xhigh reasoning effort one from UD-Q2_K_XL: Via Hacker News Tags: ai, generative-ai, llms, qwen, pelican-riding-a-bicycle, ai-in-china, nvidia-spark

Simon Willison's Weblogサイト内本文翻訳待ち:Qwen3.8-Flash-Next

翻訳待ち:Qwen 3.8 Flash-Next is Cheap, But There Are Complicating Factors

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:While Alibaba has kept inference and token price low, enterprises need to consider other metrics to determine if this is the right model for them.

AI Businessサイト内本文翻訳待ち:Qwen 3.8 Flash-Next is Cheap, But There Are Complicating Factors

翻訳待ち:Claude Desktop can now easily run Qwen, DeepSeek and Kimi models — after Ollama’s first effort stalled

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Open-weight model runner Ollama has reintroduced an integration with Claude Desktop that lets users connect Anthropic’s app to models served The post Claude Desktop can now easily run Qwen, DeepSeek and Kimi models — after Ollama’s first effort stalled appeared first on The New Stack.

The New Stack AIサイト内本文翻訳待ち:Claude Desktop can now easily run Qwen, DeepSeek and Kimi models — after Ollama’s first effort stalled

翻訳待ち:Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:We look at Qwen3.8-Flash-Next, Alibaba's open-weight multimodal Mixture-of-Experts model and an early preview of the Qwen4 architecture. We break down where the 180B parameters actually sit: a 125B backbone, a 51B N-gram embedding table, and a 4B multi-token prediction module, with only 6B active per token. We walk through the four architectural changes — the Gated DeltaNet and Qwen Sparse Attention hybrid, Gated Residual, N-gram Embedding, and the Muon optimizer. We also cover the benchmark results, the reported 1/9 training cost against Qwen3.7-Plus, and what self-hosting a 172.78 GiB FP8 checkpoint really demands. The post Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Archite…

MarkTechPostサイト内本文翻訳待ち:Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture

翻訳待ち:Moonshot AI wants 30% of what US clouds earn from Kimi K3

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:China’s Moonshot AI is in early talks with Microsoft, Amazon and Google to host Kimi K3 Credit: Bangla press via Shutterstock.com Moonshot AI is in early discussions with Microsoft, Amazon, and Google about hosting Kimi…

Hacker News AIサイト内本文翻訳待ち:Moonshot AI wants 30% of what US clouds earn from Kimi K3

翻訳待ち:Calibration-Preserving Pruning: Compression as a Reliability Contract

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:arXiv:2608.23744v1 Announce Type: new Abstract: Split conformal prediction, not the pruning rule, supplies finite-sample marginal coverage once a pruned model is fixed independently of the conformal calibration split. We study the separate efficiency problem: can pruning preserve score geometry well enough to obtain smaller valid prediction sets? Calibration-Preserving Pruning (CPP) augments a base pruning score with nonconformity-gradient saliency and uses disjoint pruning, validation-selection, conformal-calibration, and test splits. Bounded score perturbations imply bounded conformal-quantile shifts and controlled set inflation, but do not make the generic coverage theorem CPP-specific. Final five-seed Qwen2.5-1.5B results at 50\% sparsity show t…

arXiv Machine Learningサイト内本文翻訳待ち:Calibration-Preserving Pruning: Compression as a Reliability Contract

翻訳待ち:Taiwan charges nine people for smuggling ‘high-end’ AI servers to China

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Among those charged are two Super Micro employees and one from Nvidia, marking another flashpoint in US-China AI rivalry Taiwanese prosecutors charged nine people Monday, including one from Nvidia and two from Super Micro, for illegally exporting “high-end AI servers” to mainland China, adding another wave of turbulence in the AI ​​rivalry between China and the United States. Prosecutors said the servers involved were graphics processing units known as “B300,” which have been banned from sale to China. Continue reading...

The Guardian AIサイト内本文翻訳待ち:Taiwan charges nine people for smuggling ‘high-end’ AI servers to China

その他の成長タグ

中国 AI AI News | AI News Hub