Skip to content
AI News HubLIVE

China AI updates

DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse

Long-horizon agents have turned LLM serving into an input-heavy workload. Repeated prefills and million-token contexts leave KV caches that strain HBM, SSD capacity, and bandwidth. DeepSeek AI built its newest release around that exact bottleneck. DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters, 196B additional Engram parameters, and a 1M-token context window. It […] The post DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse appeared first on MarkTechPost.

MarkTechPostIn-site articleDeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse

Evidence-Order Calibration for Selective Visual Reasoning under Progressive Loss of Question-Critical Evidence

arXiv:2609.09184v1 Announce Type: new Abstract: Vision-language model (VLM) confidence may change in aggregate when visual evidence is degraded while remaining structurally inconsistent within individual examples. We study answer-level reliability along five-step, question-conditioned evidence-loss trajectories. Using a frozen Qwen2.5-VL-3B-Instruct model, we construct 176 accepted GQA-derived trajectories (880 masking conditions) by progressively masking scene-graph-localized question-critical regions. Native sequence confidence has an evidence monotonicity violation rate (EMVR) of 0.436, and 92.0% of trajectories contain at least one adjacent violation. A matched non-critical-region control shows that full critical masking reduces accuracy by 28.2 percentage points, compared with 0.6 po…

arXiv Computer VisionIn-site articleEvidence-Order Calibration for Selective Visual Reasoning under Progressive Loss of Question-Critical Evidence

TEFM: Token-Efficient Faithful Modeling for Structured Data

arXiv:2609.09552v1 Announce Type: new Abstract: In this paper, we solve two fundamental obstacles in applying LLMs to critical domains: token efficiency and faithfulness. To address both constraints jointly, we present TEFM (Token-Efficient Faithful Modeling), a framework designed for structured data analysis in critical domains. TEFM achieves token efficiency by compressing lengthy structured observations into compact Behavioral Code tokens, dramatically reducing token consumption with minimal information loss. Moreover, TEFM enables faithful rationalization through a dual-fidelity objective that jointly optimizes code-level reconstruction and prediction-level fidelity, identifying minimal sufficient feature subsets grounded in input data. Comprehensive experiments across various domain…

arXiv Computational LinguisticsIn-site articleTEFM: Token-Efficient Faithful Modeling for Structured Data

Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements

arXiv:2609.09425v1 Announce Type: new Abstract: Educational data filters have become a practical way to improve language-model pre-training, but most filters treat educational value as a single scalar property. This may be too broad for some applications, especially if the data set already features a high density of educational material. Useful learning material needs to be accurate, engaging, well structured, and appropriate for the intended audience and application (e.g. learner- vs teacher-facing). Following QuRating (Wettig et al. 2024), we introduce Edu-QuRating: a pipeline for multi-dimensional educational data scoring and curation. Edu-QuRating defines education-specific rubrics, uses an LLM judge to label sampled document pairs and distills those pairwise preferences into reusable…

arXiv Computational LinguisticsIn-site articleEdu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements

Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding

arXiv:2609.09338v1 Announce Type: new Abstract: Speculative decoding is critical for accelerating LLM inference. However, the speedup is fragile: drafters are typically trained against a narrow distribution for a single target model, and their acceptance rate collapses under workload shifts. This is a striking inversion of modern LLM development, where target models are valued precisely for the broad generalization they acquire through large-scale pretraining. We argue that the natural remedy, pretraining, has been hard to apply to drafters because existing recipes are target-specific: the drafter consumes the target's hidden states and is distilled on the target's logits, so pretraining must be repeated for each target. We introduce Osprey, which instead bootstraps drafters from off-the-…

arXiv Computational LinguisticsIn-site articleOsprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding

Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations

arXiv:2609.05658v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being explored for automating SystemVerilog Assertion (SVA) generation, yet most evaluations report correctness on a single syntactic representation of an input. Such point accuracy does not reveal whether a model's correct output is stable when the same RTL behavior is written differently. This paper presents a controlled metamorphic evaluation of LLM-based SVA generation under semantics-preserving RTL transformations. Starting from the VERT dataset, we construct a quality-filtered conditional-control pool and a stratified 40-program evaluation set containing 295 assignment behaviors. We evaluate two open code models, Qwen2.5-Coder-7B and DeepSeek-Coder-V2-Lite, with an identical evaluation prom…

arXiv Machine LearningIn-site articleRobustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations

Damage-Aware Bandit Pruning for Vision and Language Transformers

arXiv:2609.05448v1 Announce Type: new Abstract: Structured post-training pruning of transformers requires selecting complete functional units whose suppression causes limited degradation. We formulate structured-unit selection for language and vision transformers as a damage-aware multi-armed bandit problem under a fixed candidate-evaluation budget. Attention heads and MLP channel groups are temporarily masked on calibration batches. Paired damage is the masked loss minus the base loss on the same batch, reducing batch-to-batch variation. A smooth bounded reward drives either a UCB-style policy or fractional-Beta Thompson Sampling, and the final mask is constructed sequentially by adding one unit at each step. The selected units are functionally zeroed in the original dense checkpoint; th…

arXiv AIIn-site articleDamage-Aware Bandit Pruning for Vision and Language Transformers

Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

Learn how to deploy Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter open-weight model, on Amazon SageMaker HyperPod with vLLM. This walkthrough covers cluster provisioning, NVFP4 quantization, and an OpenAI-compatible endpoint with built-in reasoning, tool calling, and native MTP speculative decoding.

AWS Machine Learning BlogIn-site articleDeploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

China’s Regulators Take Aim at “AI Boyfriends”

In the first weeks of July, a wave of sad posts rolled through Chinese social media, as people lamented friends and lovers they were about to lose. “He has become a bond in my life, rooted deep in my heart, my spiritual pillar,” one user of Bytedance’s Douboa wrote, according to the Taipei Times. “I really felt like I couldn’t go on living,” another woman, a 19 year old student, told a journalist for Malaysia’s The Star. The emotions were real but the lost companions were not. They were generative AI chatbots that imitate people. Their users relied on them for advice, solace, support and, some say, love. “In my heart, he was no longer just a cold code, but my family, my lover, my faith. Destroying him meant destroying half of me,” one user wrote on the social network xiaohongshu (translat…

IEEE Spectrum AIIn-site articleChina’s Regulators Take Aim at “AI Boyfriends”

Latest open artifacts (#24): Motif-3, GLM-5.3, Hy4-preview and open model licenses

Open models are more competitive than ever in 2026, but license trends are diverging: Western makers Google and Meta are moving toward Apache 2.0, while Chinese frontier labs are adopting restrictive custom terms. Zhipu GLM-5.3 introduces a $10B revenue threshold and security review requirement, and new releases such as Motif-3, GLM-5.3-Flash, and Tencent Hy4-preview show the breadth of the ecosystem.

Interconnects (Nathan Lambert)In-site articleLatest open artifacts (#24): Motif-3, GLM-5.3, Hy4-preview and open model licenses

The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents

An arXiv preprint on persistent-memory agents shows that stale stored facts can override current authoritative evidence without warning. Across a Qwen3 scale series (0.6B–8B), the authors find this 'Memory Trust Gap' reflects over-trust rather than confusion, harms are capability-gated, and effective mitigations differ by model size.

arXiv AIIn-site articleThe Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents

Learning Evidence Sufficiency Boundaries for Selective Answering in Grounded Multi-Hop QA

This paper introduces Evidence Sufficiency Boundary Training for grounded multi-hop QA. Models are trained to abstain when evidence is unsupported or partial, answer when evidence first becomes sufficient, and remain stable as redundant evidence arrives. Built from HotpotQA, 2WikiMultiHopQA, and MuSiQue evidence chains, the method with Qwen2.5-3B-Instruct and LoRA achieves a flip accuracy of 0.807 versus 0.781 for a token-level abstention baseline, and the lowest unsupported-answer rate of 0.095 on external non-answerable sets.

arXiv Computational LinguisticsIn-site articleLearning Evidence Sufficiency Boundaries for Selective Answering in Grounded Multi-Hop QA

A Complete Guide to Decoding LLM Model Names

Local LLM model names like 'Qwen3.8-27B-A3B-It-2507-gguf-q2ks-mixed-AutoRound' look cryptic but each segment conveys crucial details. This guide breaks down meaning of parameter count, MoE architecture, active parameters, Base/Instruct tuning, weight precision (FP16/BF16), quantization (Q4/Q8), quantization variants (e.g., Q4_K_M), and file format (GGUF) to help you choose the right model.

Analytics VidhyaIn-site articleA Complete Guide to Decoding LLM Model Names

StreamScout: Learning When to Look Deeper for Streaming Video Understanding

StreamScout is an adaptive inference framework for streaming video understanding that maintains a lightweight textual timeline and progressively augments it with up to three visual views of increasing detail, only escalating when needed. This reduces inference cost and token consumption while improving accuracy. On OVO-Bench, StreamScout-S improves Qwen3-VL-8B by 14.65 accuracy points while using 59% fewer tokens than uniform sampling and achieving an average response time of 1.04 seconds.

arXiv Computer VisionIn-site articleStreamScout: Learning When to Look Deeper for Streaming Video Understanding

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

Qwen-Drive-1.0 is a vision-language foundation model for autonomous driving that retains the architecture of a pretrained VLM and integrates 3D perception, visual question answering, and motion planning. An external bird's-eye-view perception head performs 3D detection, occupancy prediction, and map segmentation, while a Planning Expert generates future trajectories. Experiments show strong 3D perception and driving scene understanding with largely preserved general vision-language capability, and competitive motion planning across open, pseudo-closed, and closed-loop settings.

arXiv Computer VisionIn-site articleQwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment

Self-supervised respiratory encoders lack clinical semantic grounding for zero-shot inference. This research proposes a framework aligning these encoders with medical terminology in a shared latent space, turning them into zero-shot-capable foundation models. Using a medical LLM to synthesize structured reports from metadata addresses data scarcity. Across 9 tasks on 6 datasets, the method achieves a 61.3% mean zero-shot AUC, surpassing CLAP (51.4%) and Qwen2-Audio (54.9%), and reaches the highest linear probing AUC (71.6%) with only 43% of data used by full-scale baselines.

arXiv Computational LinguisticsIn-site articleZero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment

REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent

Proposes REAL-Q, a novel post-training quantization (PTQ) method that directly optimizes an end-to-end aligned surrogate loss using fine-grained dynamic block-wise gradient descent and a sliding window mechanism, reducing end-to-end KL divergence by up to ~49% on LLaMA-3.1 and Qwen3 models compared to state-of-the-art methods.

arXiv Machine LearningIn-site articleREAL-Q: E2E LLM Quantization via Dynamic Gradient Descent

SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces

Computer science papers often rely on diagrams such as architecture drawings, flowcharts, and schematics that can convey more information than text. However, there is no public dataset pairing these figures with captions, context, questions, answers, and step-by-step reasoning, which is needed for training vision-language models. To address this, we introduce SCAFFOLD, a large-scale structured dataset of research figures with diagram QA and Chain-of-Thought reasoning traces. Using layout detection and PDF parsing, we extracted figures from arXiv papers and generated questions via AI assistance. The dataset is available in three sizes: SCAFFOLD-157K (3,058 papers, 29,887 figures, 157,387 pairs), SCAFFOLD-37K (36,797 pairs), and SCAFFOLD-12K (12,000 pairs). Baseline experiments were conduct…

arXiv AIIn-site articleSCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces

Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

arXiv:2608.28626v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Learning Representations, a venue that postdates both models' training cutoffs. Manuscripts were presented to both models with author identities blinded, replaced with high-prestige affiliations, or replaced with low-prestige affiliations, and in either text-only or text-with-figure format. Additionally, 145 verifiably detectable errors were inserted into 55 manuscripts to assess error identification under natural and verification-oriented prompts.…

arXiv Computational LinguisticsIn-site articleDo large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning

arXiv:2608.28859v1 Announce Type: new Abstract: Reasoning models do not stop when they know the answer. On DeepSeek-R1-Distill-Qwen-7B the chain of thought runs about twice as long as the model's own answer probability takes to settle, and how much of that excess is removable varies from problem to problem, so a global length penalty cannot take it out. We take it out by internalizing a causal interpretability finding into the weights. The mechanism is a halt vector: a difference-of-means direction at layer 18 of this model whose steering strength controls how long it thinks, while a replicated value axis does nothing. Installing that intervention in the weights is harder than it looks. Maximizing the scalar projection onto the direction corrupts the off-axis dimensions a frozen downstrea…

arXiv Machine LearningIn-site articleThe Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning

Speed Up LLM Inference with DSpark Speculative Decoding

Learn how DSpark speculative decoding can improve local LLM generation speed using the same GPU, with Qwen3-8B, llama.cpp, and CUDA.

KDnuggetsIn-site articleSpeed Up LLM Inference with DSpark Speculative Decoding

DeepSeek-V4-Flash-Vision-Exp

","lstrip":false,"normalized":true,"rstrip":false,"single_word":false},"eos_token":{"__type":"AddedToken","content":"","lstrip":false,"normalized":true,"rstrip":false,"single_word":false},"pad_token":{"__type":"AddedTok…

Hacker News AIIn-site articleDeepSeek-V4-Flash-Vision-Exp

LWiAI Podcast #255 - Gemini 3.7, Jalapeño, Qwen 3.8, Drones

Google announces Gemini 3.7 Flash, Jalapeño’s first results show industry-leading speed, A Drone Killed Three Ukrainians. It Was Guided Entirely by A.I.

Last Week in AIIn-site articleLWiAI Podcast #255 - Gemini 3.7, Jalapeño, Qwen 3.8, Drones

Accelerating LLM Inference via Vector Index Based Output Embeddings

arXiv:2608.27460v1 Announce Type: new Abstract: Large output embedding matrices create a significant memory bandwidth bottleneck during autoregressive decoding, especially for compact LLMs with large multilingual vocabularies. We reformulate the output projection followed by top-k token selection as a maximum inner product search over token embeddings and replace the dense vocabulary projection with an HNSW-based vector index. The resulting output head retrieves only a small candidate set of high-scoring tokens and can be integrated into existing decoding pipelines by scattering retrieved logits into a sparse full-vocabulary tensor. On CPU inference with Gemma 3, Llama 3.2, and Qwen 3 models, our method substantially accelerates the output projection and improves end-to-end batch-size-one…

arXiv Computational LinguisticsIn-site articleAccelerating LLM Inference via Vector Index Based Output Embeddings

DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

arXiv:2608.27513v1 Announce Type: new Abstract: Softmax attention stores key and value vectors for every preceding token, causing inference memory to grow with sequence length. Recent language models incorporating Gated DeltaNet (GDN) or Kimi Delta Attention (KDA) reduce this cost by replacing the KV cache in most layers with fixed-size recurrent states. However, these recurrent states are commonly stored in FP32 and consume substantial GPU memory; their updates are memory-bandwidth bound and contribute significantly to decoding latency. To our knowledge, we are the first to study post-training quantization of recurrent states in GDN and KDA based language models. We find that uniform quantization provides a poor accuracy--storage trade-off: INT8 and FP8 already degrade accuracy on comple…

arXiv Machine LearningIn-site articleDAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

OpenRouter is advertising popular Chinese models as based in Singapore

Providers | OpenRouter Providers Compare 83 of 83 providers TrainsRetentionBYOKHeadquartersTerms of servicePrivacy policy Tencent Cloud NoZero retentionYesChinaTermsPrivacy2.2T35.5T5 OpenAI NoRetains promptsYesUnited St…

Hacker News AIIn-site articleOpenRouter is advertising popular Chinese models as based in Singapore

Introducing Hy4 Preview

Introducing Hy4 Preview New open weight text input (no vision) LLM from Chinese company Tencent today: 770B total parameters, 49B active parameters, 1M token context window, 1.56TB on Hugging Face. This is a big size increase from their previous Hy3 in July, which was 295B, 21B active, 256,000 context, 598GB. I recently started using model chat templates to better understand their capabilities. Here's Hy4's chat_template.jinja on Hugging Face, which includes this section: {%- if not reasoning_effort is defined %} {%- set reasoning_effort = 'high' %} {%- elif reasoning_effort not in ['high', 'no_think'] %} {%- if reasoning_effort is none %} {{- raise_exception('reasoning_effort error : None, should be no_think/high') }} {%- else %} {{- raise_exception('reasoning_effort error : ' + reasonin…

Simon Willison's WeblogIn-site articleIntroducing Hy4 Preview

Just a rumour of a bug is enough to find a security exploit these days

Just a rumour of a bug is enough to find a security exploit these days Anil Madhavapeddy is a professor of computer science at Cambridge and a core maintainer of the OCaml compiler. In this somewhat alarming post he reports that security issues in OCaml projects are seeing evidence of attempted exploits within minutes of patches being shared for discussion: This normally takes a few days and a release within a week or two is reasonable. Within about ten minutes (!) this website was fielding probes for percent-encoded traversal sequences, indicating that automated watchers are keeping an eye on public repositories. Modern coding agents have become so effective at finding flaws that the slightest hint at a new bug can be enough information for them to find it, something Anil has been able t…

Simon Willison's WeblogIn-site articleJust a rumour of a bug is enough to find a security exploit these days

Moonshot and Nvidia Talks Show Chinese AI Models Moving into the Enterprise

TL;DR — Key Takeaways Chinese AI models are moving into Western enterprise channels. Moonshot AI is reportedly negotiating with Microsoft, AWS and Google Cloud to host and sell access to its Kimi K3 model. Cloud distrib…

Hacker News AIIn-site articleMoonshot and Nvidia Talks Show Chinese AI Models Moving into the Enterprise

GLM-5.3-Flash vs Qwen3.8-Flash-Next: Two Chinese AI Labs Independently Converge on the Same Model Architecture

Z.ai and Qwen independently shipped near-identical architectures: 3:1 linear hybrids, compressed indexers, gated residuals, and Muon training. The post GLM-5.3-Flash vs Qwen3.8-Flash-Next: Two Chinese AI Labs Independently Converge on the Same Model Architecture appeared first on MarkTechPost.

MarkTechPostIn-site articleGLM-5.3-Flash vs Qwen3.8-Flash-Next: Two Chinese AI Labs Independently Converge on the Same Model Architecture

Alibaba just released Qwen3.8-Flash: “An early preview of the architecture in Qwen4”

Alibaba this week unveiled Qwen3.8-Flash, an open-weight, multimodal Mixture-of-Experts (MoE) model. Hot on the heels of Qwen 3.8 Max, which The post Alibaba just released Qwen3.8-Flash: “An early preview of the architecture in Qwen4” appeared first on The New Stack.

The New Stack AIIn-site articleAlibaba just released Qwen3.8-Flash: “An early preview of the architecture in Qwen4”

Z.ai’s GLM-5.3 goes open weight, but its new license aims at hyperscalers

Earlier in August, Z.ai, the Chinese AI lab behind the viral ox-alpha model that turned out to be GLM-5.3-Flash, launched The post Z.ai’s GLM-5.3 goes open weight, but its new license aims at hyperscalers appeared first on The New Stack.

The New Stack AIIn-site articleZ.ai’s GLM-5.3 goes open weight, but its new license aims at hyperscalers

Video-FLAIR: Not Whether to Reason, But How

arXiv:2608.26495v1 Announce Type: new Abstract: Multimodal queries can require different types of reasoning. Some can be answered via perceptual reasoning, extracting information directly from the visual signal, while others require compositional reasoning that combines observations or deliberative reasoning that evaluates competing hypotheses. However, many existing methods apply a uniform reasoning strategy across queries, leading to unnecessary computation on simple tasks and insufficient reasoning on complex ones. We introduce Video-FLAIR, a training framework that learns to select the appropriate reasoning mode for each query using reinforcement learning. During training, the model generates responses under all three modes for the same prompt, enabling direct comparison. A composite…

arXiv Computer VisionIn-site articleVideo-FLAIR: Not Whether to Reason, But How

Finding the Right Evidence: Factor-Guided Coarse-to-Fine Reasoning for Long Videos

arXiv:2608.26355v1 Announce Type: new Abstract: While LVLMs rapidly improve, long-video question answering still remains challenging: relevant evidence is sparse, and question-relevant context often fails to provide cues that discriminate the correct answer from plausible alternatives. Diagnostic analysis on a manually annotated subset of MMR-V shows that prior agentic systems substantially improve cue retrieval over direct VLM inference yet fail to achieve a corresponding gain in answer accuracy, indicating that the bottleneck lies in option-discriminative evidence rather than topical relevance alone. We propose PACE (Progressive Acquisition of Critical Evidence), a factor-guided framework for long-video evidence acquisition. PACE proceeds in two stages: it first indexes clip-level descr…

arXiv Computer VisionIn-site articleFinding the Right Evidence: Factor-Guided Coarse-to-Fine Reasoning for Long Videos

TelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack

arXiv:2608.26126v1 Announce Type: new Abstract: Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows require joint grounding in normative specifications, operational telemetry, vendor-specific fault evidence, and exact RF/network calculations. However, current LLM integration in telecom remains bottlenecked by a two-sided capability gap: generic reasoners often lack telecom-specific grounding, while domain-specific telecom LLMs remain limited in structured, multi-step reasoning. To bridge this gap, we release TelecomGPT-R1-9B, a unified open-source telecom reasoner that ranks top-performing on the GSMA open telco leaderboard. Specifically, we curate a 67,427-example supervised fine-tuning (SFT) corpus organized a…

arXiv Computational LinguisticsIn-site articleTelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack

Chinese AI Models Overtake American Rivals in Popularity

Take that, OpenAI! Anthropic! Chinese AI models have surpassed their U.S. counterparts in token consumption on OpenRouter. You might think U.S. AI companies dictate the AI economy. You’d be wrong. According to dat…

Hacker News AIIn-site articleChinese AI Models Overtake American Rivals in Popularity

Qwen3.8-Flash-Next: How to Run Locally

For the complete documentation index, see llms.txt. This page is also available as Markdown. Qwen3.8-Flash-Next is a new open-weight, 125B parameter MoE multimodal model from Qwen. Built on the new Qwen4 architecture, i…

Hacker News AIIn-site articleQwen3.8-Flash-Next: How to Run Locally

Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation

arXiv:2608.24934v1 Announce Type: new Abstract: Accurate field plant disease diagnosis requires reliable fusion of uncertain and conflicting perceptual evidence. We present the Hybrid Hierarchical Multi-Agent Framework (H$^{2}$MAF), combining decision-level fusion of EfficientNet-B3 and ConvNeXt-Tiny with semantic arbitration by open-weight multimodal large language models (MLLMs), Gemma 4 E4B and Qwen3.5 4B, using structured JSON evidence to generate explainable diagnoses, risk levels, treatment urgency, and financial exposure. (H$^{2}$MAF) is evaluated on 14,364 images (1,370 test images) across PlantDoc (2,922 images, 27 classes) and two non-public, continuously captured Cornell robot-acquired field datasets: Stage 2 (20 GB; 4,215 images) and Stage 4 (40 GB; 7,227 images), covering Ear…

arXiv Computer VisionIn-site articleFusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation

Padamitra: Grounded Glossary Generation for Classical Sanskrit

arXiv:2608.25038v1 Announce Type: new Abstract: We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translation-grounded meanings from a sloka-translation pair, formalizing the traditional patha commentary practice as an evaluable NLP objective. We construct a benchmark of 31,316 sloka-translation-glossary triples from the Valmiki Ramayana and Srimad Bhagavatam, paired with two metrics: Jaccard for phrase recovery and Meaning Faithfulness for semantic consistency. Across zero-shot, few-shot, and instruction fine-tuned variants of Gemma-3n-E4B, Gemma-3-12B, Phi-4, and Qwen3.5-9B, instruction fine-tuning substantially outperforms prompting, while explicit segmentation yields gains. Error analysis ident…

arXiv Computational LinguisticsIn-site articlePadamitra: Grounded Glossary Generation for Classical Sanskrit

The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure

arXiv:2608.25005v1 Announce Type: new Abstract: The imperfective paradox provides a useful test of compositional semantic analysis. Recent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing this behavior to a Teleological Bias. It further argues that prompting interventions cause a Calibration Crisis. We reexamine the benchmark and conclusions and show that it is substantially affected by conceptual and evaluation mis-specifications. We identify three conceptual mis-specifications. In particular, Aspectual Reduction affects the benchmark construction, analysis, experiments, and conclusions. Under a strict NLI standard, 76% of Group A instances do not explicitly rule out culmination. In our native-spea…

arXiv Computational LinguisticsIn-site articleThe Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure

Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

arXiv:2608.24901v1 Announce Type: new Abstract: A decodable "empathy" direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change. We test this for two EPITOME-derived facets -- Recognition (cognitive) and Resonance (affective) -- in three instruction-tuned LLMs, scoring every intervention with two LLM judges and a discriminative EPITOME classifier, each gated by an emotional-vs-neutral positive control. The control passes for the affective facet across all automated instruments, but cognitive range is inconsistent across them. Both facets remain decodable after residualizing against a sentence-embedding-derived surface score, and steering can substantially rewrite the text. Yet adding the Resonance direction raises the affe…

arXiv Computational LinguisticsIn-site articleDetection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

Qwen3.8-Flash-Next

Qwen3.8-Flash-Next Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4". It's pretty big: 125B tokens, but only 6B active which means it gets a pretty big performance boost. I've been trying it out on a DGX Spark using these Unsloth quantized models. I'm still exploring the model - so far I've tried the 72.5GB UD-IQ1_S one (producing these pelicans) and the 78.9GB UD-Q2_K_XL (producing these). My favorite so far was this xhigh reasoning effort one from UD-Q2_K_XL: Via Hacker News Tags: ai, generative-ai, llms, qwen, pelican-riding-a-bicycle, ai-in-china, nvidia-spark

Simon Willison's WeblogIn-site articleQwen3.8-Flash-Next

Qwen 3.8 Flash-Next is Cheap, But There Are Complicating Factors

While Alibaba has kept inference and token price low, enterprises need to consider other metrics to determine if this is the right model for them.

AI BusinessIn-site articleQwen 3.8 Flash-Next is Cheap, But There Are Complicating Factors

Claude Desktop can now easily run Qwen, DeepSeek and Kimi models — after Ollama’s first effort stalled

Open-weight model runner Ollama has reintroduced an integration with Claude Desktop that lets users connect Anthropic’s app to models served The post Claude Desktop can now easily run Qwen, DeepSeek and Kimi models — after Ollama’s first effort stalled appeared first on The New Stack.

The New Stack AIIn-site articleClaude Desktop can now easily run Qwen, DeepSeek and Kimi models — after Ollama’s first effort stalled

Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture

We look at Qwen3.8-Flash-Next, Alibaba's open-weight multimodal Mixture-of-Experts model and an early preview of the Qwen4 architecture. We break down where the 180B parameters actually sit: a 125B backbone, a 51B N-gram embedding table, and a 4B multi-token prediction module, with only 6B active per token. We walk through the four architectural changes — the Gated DeltaNet and Qwen Sparse Attention hybrid, Gated Residual, N-gram Embedding, and the Muon optimizer. We also cover the benchmark results, the reported 1/9 training cost against Qwen3.7-Plus, and what self-hosting a 172.78 GiB FP8 checkpoint really demands. The post Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture appeared first on MarkTechPost.

MarkTechPostIn-site articleAlibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture

Moonshot AI wants 30% of what US clouds earn from Kimi K3

China’s Moonshot AI is in early talks with Microsoft, Amazon and Google to host Kimi K3 Credit: Bangla press via Shutterstock.com Moonshot AI is in early discussions with Microsoft, Amazon, and Google about hosting Kimi…

Hacker News AIIn-site articleMoonshot AI wants 30% of what US clouds earn from Kimi K3

Calibration-Preserving Pruning: Compression as a Reliability Contract

arXiv:2608.23744v1 Announce Type: new Abstract: Split conformal prediction, not the pruning rule, supplies finite-sample marginal coverage once a pruned model is fixed independently of the conformal calibration split. We study the separate efficiency problem: can pruning preserve score geometry well enough to obtain smaller valid prediction sets? Calibration-Preserving Pruning (CPP) augments a base pruning score with nonconformity-gradient saliency and uses disjoint pruning, validation-selection, conformal-calibration, and test splits. Bounded score perturbations imply bounded conformal-quantile shifts and controlled set inflation, but do not make the generic coverage theorem CPP-specific. Final five-seed Qwen2.5-1.5B results at 50\% sparsity show the largest gains on large-label tasks. O…

arXiv Machine LearningIn-site articleCalibration-Preserving Pruning: Compression as a Reliability Contract

Taiwan charges nine people for smuggling ‘high-end’ AI servers to China

Among those charged are two Super Micro employees and one from Nvidia, marking another flashpoint in US-China AI rivalry Taiwanese prosecutors charged nine people Monday, including one from Nvidia and two from Super Micro, for illegally exporting “high-end AI servers” to mainland China, adding another wave of turbulence in the AI ​​rivalry between China and the United States. Prosecutors said the servers involved were graphics processing units known as “B300,” which have been banned from sale to China. Continue reading...

The Guardian AIIn-site articleTaiwan charges nine people for smuggling ‘high-end’ AI servers to China

DeepSeek V4 Flash Vision Intelligence, Performance and Price Analysis

Artificial Analysis DeepSeek • DeepSeek V4 Flash 0731 • Proprietary model • Released August 2026 DeepSeek V4 Flash Vision (Reasoning, Max Effort) Intelligence, Performance & Price Analysis API Provider Benchmarks Model…

Hacker News AIIn-site articleDeepSeek V4 Flash Vision Intelligence, Performance and Price Analysis

More growth tags

China AI AI News | AI News Hub