AI News HubLIVE

GPU 基礎設施動態

待翻譯:Andreessen Horowitz raises $1.1B AI infrastructure fund

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Andreessen Horowitz today announced that it has raised a $1.1 billion fund to back artificial intelligence infrastructure startups. The Machine Age Fund will invest in companies that make data center equipment such as chips, memory and networking gear. Andreessen Horowitz also plans to prioritize providers of edge AI hardware. The venture capital firm listed smart […] The post Andreessen Horowitz raises $1.1B AI infrastructure fund appeared first on SiliconANGLE.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Andreessen Horowitz today announced that it has raised a $1.1 billion fund to back artificial intelligence infrastructure startups. The Machine Age Fund will invest in companies t…
站內正文

待翻譯:Moonshot and Nvidia Talks Show Chinese AI Models Moving into the Enterprise

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:TL;DR — Key Takeaways Chinese AI models are moving into Western enterprise channels. Moonshot AI is reportedly negotiating with Microsoft, AWS and Google Cloud to host and sell access to its Kimi K3 model. Cloud distrib…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • TL;DR — Key Takeaways Chinese AI models are moving into Western enterprise channels. Moonshot AI is reportedly negotiating with Microsoft, AWS and Google Cloud to host and sell ac…
站內正文

待翻譯:Vercel AI Open-Sources vgpu: A TypeScript WebGPU Library for AI Agent Shaders

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Vercel has open-sourced vgpu, the WebGPU library it built to ship the shaders on vercel.com. It treats .wgsl files as importable TypeScript modules, runs the same shader in the browser, in headless Node.js via Dawn, and in a deterministic CI mock, and ships a fullscreen effect in 25 KB gzipped. The post Vercel AI Open-Sources vgpu: A TypeScript WebGPU Library for AI Agent Shaders appeared first on MarkTechPost.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Vercel has open-sourced vgpu, the WebGPU library it built to ship the shaders on vercel.com. It treats .wgsl files as importable TypeScript modules, runs the same shader in the br…
站內正文

待翻譯:AI GPUs probably live longer than three years

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:People who think current AI use is unsustainable often rely on the claim that inference GPUs only last “three years at the most” under load1. The idea here is that once the AI bubble money drains away, current infrastru…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • People who think current AI use is unsustainable often rely on the claim that inference GPUs only last “three years at the most” under load1. The idea here is that once the AI bub…
站內正文

待翻譯:Labour rejects Zack Polanski’s call to ‘slam brakes’ on building AI datacentres

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Government says Green leader’s proposed moratorium on ‘energy-guzzling’ AI projects would be disaster for economy Labour has rejected calls to pause construction of major AI infrastructure in Britain after the Green party leader, Zack Polanski, said it was time to “slam the brakes on these energy-guzzling, water-guzzling datacentres”. The government hit back at the opposition party’s proposal of a “moratorium other than [for] local-scale datacentres for the local community”, saying it “would be a disaster for jobs and national security”. Continue reading...

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Government says Green leader’s proposed moratorium on ‘energy-guzzling’ AI projects would be disaster for economy Labour has rejected calls to pause construction of major AI infra…
站內正文

待翻譯:Prompt: The AI Infrastructure Boom Is Getting Bigger Than GPUs

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Nvidia’s record quarter shows AI demand is still surging as the infrastructure race expands into CPUs, networking, robotics and edge computing.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Nvidia’s record quarter shows AI demand is still surging as the infrastructure race expands into CPUs, networking, robotics and edge computing.
站內正文

待翻譯:It’s Nvidia’s world. We just live in it

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The artificial intelligence juggernaut kept cruising along this week thanks to big earnings results from Nvidia — and even Salesforce, the supposed epicenter of the SaaSpocalypse. Nvidia not only beat all expectations for revenue, CEO Jensen Huang (pictured) indicated it’s going to be capacity-constrained for awhile longer, which certainly indicates no diminution of demand. Likewise […] The post It’s Nvidia’s world. We just live in it appeared first on SiliconANGLE.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • The artificial intelligence juggernaut kept cruising along this week thanks to big earnings results from Nvidia — and even Salesforce, the supposed epicenter of the SaaSpocalypse.…
站內正文

待翻譯:Cross-Platform Benchmark of Neural 3D Reconstruction for Autonomous Laboratory Robots

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.26383v1 Announce Type: new Abstract: Autonomous robots performing laboratory tasks depend on 3D reconstruction pipelines that can turn raw camera streams into actionable object representations within the latency budget of a physical control loop. Neural 3D reconstruction methods have demonstrated high-quality view synthesis, but their real-time viability across the compute platforms on which laboratory robots actually run remains poorly characterized. In this work, we present a systematic compute-platform benchmark of neural 3D reconstruction methods, evaluating NeRF and 3D Gaussian Splatting training and rendering on GPU-enabled computing devices ranging from single-board computers to server-class nodes, and place Meta's SAM3D single-image reconstruction on the same axes to quantify its latency and fidelity gap relative to per-scene optimization. Our results show that Gaussian Splatting yields higher rendering quality than NeRF at greater GPU cost, and that onboard compute is insufficient for full per-scene optimization at interactive rates. Our preliminary assessment on SAM3D indicates that it delivers plausible object geometry within seconds, but with detail mismatches that can compromise downstream manipulation. Together, these findings motivate tiered pipelines in which lightweight feed-forward reconstruction sustains the real-time perception-and-tracking loop for laboratory robots, while heavier neural reconstruction is scheduled selectively on suitable compute.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.26383v1 Announce Type: new Abstract: Autonomous robots performing laboratory tasks depend on 3D reconstruction pipelines that can turn raw camera streams into actionabl…
站內正文

待翻譯:Nvidia reportedly acquires AI project hosting platform Hugging Face for $12.9B

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Nvidia Corp. has reportedly bought Hugging Face Inc., a startup with a popular platform for hosting open-source artificial intelligence projects. Rumors that an acquisition was in the cards first leaked on Monday. Business Insider broke the news that Hugging Face had received interest from multiple prospective buyers. On late Wednesday, The Information reported that Nvidia […] The post Nvidia reportedly acquires AI project hosting platform Hugging Face for $12.9B appeared first on SiliconANGLE.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Nvidia Corp. has reportedly bought Hugging Face Inc., a startup with a popular platform for hosting open-source artificial intelligence projects. Rumors that an acquisition was in…
站內正文

待翻譯:The AI storage stack gets an inference-era rethink

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Artificial intelligence is changing what storage and data management platforms look like. In collaboration with Super Micro Computer Inc. and Solidigm, DataDirect Networks Inc. has introduced DDN Enterprise AI HyperPOD, built on Nvidia Corp.’s AI Data Platform. The goal is to simplify the storage, scaling and deployment of AI inference for enterprise workloads. “Customers are […] The post The AI storage stack gets an inference-era rethink appeared first on SiliconANGLE.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Artificial intelligence is changing what storage and data management platforms look like. In collaboration with Super Micro Computer Inc. and Solidigm, DataDirect Networks Inc. ha…
站內正文

待翻譯:Pushing GPT-5.6 Luna from 0% to 56% on ARC-AGI-3 Public

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Research August 27, 2026 For long-horizon agents, model capability alone does not determine system capability. Infrastructure orchestration multiplies what models can do. INT21’s SwarmOS tested this hypothesis on the AR…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Research August 27, 2026 For long-horizon agents, model capability alone does not determine system capability. Infrastructure orchestration multiplies what models can do. INT21’s…
站內正文

待翻譯:AMD Jumps from ROCm 7.14 to ROCm 10.0 with Rocm.ai

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AMD Jumps From ROCm 7.14 To ROCm 10.0 With ROCm.AI Back in July ROCm 7.14 was announced as their new production release built atop TheRock build system and introducing Ryzen AI 400 series support. The versioning choice…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • AMD Jumps From ROCm 7.14 To ROCm 10.0 With ROCm.AI Back in July ROCm 7.14 was announced as their new production release built atop TheRock build system and introducing Ryzen AI 40…
站內正文

待翻譯:Jensen Huang says Nvidia achieved AGI, again — not that it matters

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:On Nvidia's earnings call Wednesday, CEO Jensen Huang casually announced the company had "achieved AGI," one of the tech industry's ultimate goals some of its biggest players have spent years chasing. Almost immediately, Huang dismissed the coveted milestone as "senseless." He's right. For the supposed finish line of the AI race, there is no consensus on what artificial general intelligence means, let alone how we'll know when we've actually got there, which makes achieving it equally arbitrary. Asked about OpenAI's pursuit of AGI, Huang said that when it comes to Nvidia, "for many tasks, we could say that we've already achieved AGI." … Read the full story at The Verge.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • On Nvidia's earnings call Wednesday, CEO Jensen Huang casually announced the company had "achieved AGI," one of the tech industry's ultimate goals some of its biggest players have…
站內正文

待翻譯:Deepgram deepens Amazon SageMaker AI observability with Enhanced Metrics

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Self-hosted speech AI carries an observability trade-off: the numbers that drive capacity planning and cost management stay locked inside the vendor container. Deepgram closes that gap on Amazon SageMaker AI with two capabilities that land billing, usage, and per-GPU metrics directly in your own Amazon CloudWatch account.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Self-hosted speech AI carries an observability trade-off: the numbers that drive capacity planning and cost management stay locked inside the vendor container. Deepgram closes tha…
站內正文

待翻譯:Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Serving automatic speech recognition (ASR) models at scale is costly when each request uses only a fraction of a GPU. Learn how NVIDIA CUDA Multi-Process Service (MPS) with NVIDIA Triton Inference Server on Amazon EC2 GPU instances cuts GPU infrastructure by 75% while holding sub-second latency at 92.1 requests per second per GPU.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Serving automatic speech recognition (ASR) models at scale is costly when each request uses only a fraction of a GPU. Learn how NVIDIA CUDA Multi-Process Service (MPS) with NVIDIA…
站內正文

待翻譯:Nvidia Targets Physical AI With New Jetson Edge Platform

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The Jetson Orin Nano 2 comes as the AI chip giant targets physical AI expansion.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • The Jetson Orin Nano 2 comes as the AI chip giant targets physical AI expansion.
站內正文

待翻譯:AWS, Nvidia Expand Partnership With 2 Million More GPUs

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The companies are expanding their collaboration beyond GPUs into CPUs, government AI infrastructure and robotics.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • The companies are expanding their collaboration beyond GPUs into CPUs, government AI infrastructure and robotics.
站內正文

待翻譯:GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:NVIDIA’s Gamescom announcements are revealing what’s next for GeForce NOW, with new ways to play, more supported devices and platforms, and even more big PC games headed to the cloud. New NVIDIA DLSS 4.5 technology controls give members more ways to fine-tune gameplay, while expanded support for new Steam devices, GOG single sign-on, Firefox browser […]

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • NVIDIA’s Gamescom announcements are revealing what’s next for GeForce NOW, with new ways to play, more supported devices and platforms, and even more big PC games headed to the cl…
站內正文

待翻譯:Nvidia NVLink Fusion Brings Nvhbm to Next-Generation AI Infrastructure

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AI factories must support increasingly large models and more complex reasoning workloads. To keep up with the insatiable compute demands of AI workloads, hyperscalers and AI-native companies are developing custom AI acc…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • AI factories must support increasingly large models and more complex reasoning workloads. To keep up with the insatiable compute demands of AI workloads, hyperscalers and AI-nativ…
站內正文

待翻譯:CRESSim-Neo: A Batched GPU Simulation Engine for Surgical Robotics and Robot Learning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.25192v1 Announce Type: new Abstract: We introduce CRESSim-Neo, a batched GPU simulation engine for surgical robotics and robot learning. CRESSim-Neo combines position-based simulation of rigid bodies, deformable tissues, fluids, and strands with batched rendering, surgery-specific sensing, and a GPU-resident data pipeline. The engine supports applications including tissue manipulation, fluid suction, suturing, cable-driven robots, and ultrasound image synthesis. Direct access to physics and rendering buffers enables GPU-resident robot learning and zero-copy PyTorch integration using DLPack. We demonstrate CRESSim-Neo across rigid-body, deformable-body, and fluid simulation tasks, including vision-based and surgical robot-learning scenarios. On an NVIDIA RTX 4090, the engine achieves up to 2.03 million environment steps per second for 8192 parallel CartPole environments, and scales to batched surgical scenarios involving tissue deformation, fluid interaction, and ultrasound sensing. Overall, CRESSim-Neo provides a unified and scalable platform for surgical simulation, synthetic data generation, and surgical robot learning.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.25192v1 Announce Type: new Abstract: We introduce CRESSim-Neo, a batched GPU simulation engine for surgical robotics and robot learning. CRESSim-Neo combines position-b…
站內正文

待翻譯:A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.24935v1 Announce Type: new Abstract: Accurate identification of early-stage apple fruitlet anatomical structures, including the calyx, fruitlet body, and peduncle, is essential for robotic thinning, crop-load management, and other precision orchard operations. This study presents a lightweight multimodal vision-language framework that adapts TinyCLIP for fine-grained fruitlet anatomy classification in complex orchard environments. A dataset of 600 high-resolution RGB images collected from Scilate and Scifresh apple orchards was converted into 224 x 224 image patches and annotated for three anatomical classes. Domain-specific language prompts, such as ``a photo of a class,'' were used to guide multimodal alignment between orchard imagery and horticultural structures. A sliding-window inference strategy with a stride of 112 pixels aggregates patch-level predictions into spatial heatmaps, enabling interpretable whole-image localization of fruitlet components relevant to robotic thinning. Patch-level evaluation on an NVIDIA T4 GPU achieved F1-scores of 0.95 for calyx, 0.98 for fruitlet, and 0.85 for peduncle, with a macro-F1 score of 0.93. Deployment-oriented optimization using ONNX and TensorRT enabled efficient inference on NVIDIA Jetson hardware, preserved accuracy under INT8 quantization, and supported model sizes of approximately 127-137 MB with millisecond-level patch inference. These results demonstrate that lightweight vision-language models can provide interpretable and edge-deployable perception for automated fruitlet analysis and future robotic thinning systems. The source code and implementation details are publicly available at https://github.com/WilliamBu1/A-Lightweight-Vision-Language-Model-for-Early-Stage-Fruitlet-Classification-in-Apple-Orchards.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.24935v1 Announce Type: new Abstract: Accurate identification of early-stage apple fruitlet anatomical structures, including the calyx, fruitlet body, and peduncle, is e…
站內正文

待翻譯:DataKernelBench: Can LLMs Optimize Database Queries on GPUs?

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.25061v1 Announce Type: new Abstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Existing LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement-heavy database-style operators untested. We introduce DataKernelBench, which translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that optimize either the core tensor-bounded snippet or the full query in CUDA or Triton through execution-guided repair. Across ten proprietary and open-weight models on TPC-H SF10 with an H100 GPU, the strongest full-query CUDA configuration achieves $2.11\times$ speedup over torch.compile at full pass rate. We find that higher-performing implementations commonly use kernel fusion and execution-strategy changes, stronger models benefit most from full-query specialization, and workload context matters more than hardware context. To handle data larger than GPU memory, we extend TorchPlan with Dask-cuDF for on-demand partition loading on TPC-H SF100 with four H100 GPUs, achieving $2.54\times$ speedup

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.25061v1 Announce Type: new Abstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Exis…
站內正文

待翻譯:MacroAgent: Regularity-Aware Macro Legalization with LLM-Agent-Designed Contour Algorithms

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.24946v1 Announce Type: new Abstract: Macros constitute a large part of the core area in modern very large-scale integration (VLSI) designs. Moreover, macro positions have a significant impact on the final quality of result (QoR), and macro legalization is typically the final step in determining the macro positions. However, existing approaches related to macro legalization either lack robustness or incur substantial computational costs or neglect the regularity between macros. To address these limitations, we introduce MacroAgent. The novel framework is a four-stage approach: clustering, contour generation, template matching, and inter-cluster refinement. We propose leveraging Large Language Models (LLMs) to discover multiple, effective heuristic regularity-aware contour algorithms. This framework successfully generates robust and effective algorithmic solutions for macro legalization. Compared with state-of-the-art macro legalization works, experimental results on TILOS and Chipyard benchmarks demonstrate a 2 to 8 fold improvement in layout regularity, a 3% to 5% reduction in routed wirelength with comparable congestion after global routing, and significantly better robustness with an acceptable runtime. Furthermore, end-to-end evaluation through Cadence Innovus place-and-route confirms that the regularity improvements translate into tangible PPA gains, including 2.9% lower routed wirelength and 68.3% TNS improvement over the DREAMPlace macro legalization baseline; it also achieves 1.8% lower routed wirelength when integrated into the Innovus macro placement flow.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.24946v1 Announce Type: new Abstract: Macros constitute a large part of the core area in modern very large-scale integration (VLSI) designs. Moreover, macro positions ha…
站內正文

待翻譯:FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.24945v1 Announce Type: new Abstract: Recent years have witnessed remarkable achievements of Large Language Models (LLMs) in multiple domains, while the excessive resource requirements of LLMs hinder the deployment on resource-constrained devices. Although model quantization stands out as an effective approach, conventional quantization approaches typically incur severe performance degradation due to uniform bit-width or simple heuristic sensitivity evaluation. In this paper, we propose a novel Fisher information-based Adaptive Mixed Precision Weight Quantization approach, i.e., FAMPWQ, which performs layer-adaptive weight quantization for effective LLM inference on commodity GPUs. First, we propose a system model with a novel Fisher information metric to measure the layer-wise sensitivity to quantization. Second, we propose a reinforcement learning-based bit-width allocator in FAMPWQ, which generates an adaptive bit-width allocation strategy based on the Fisher information sensitivity metric. Extensive experiments on 7 models and 5 benchmarks demonstrate that FAMPWQ significantly outperforms 7 baseline approaches in terms of PPL (up to 3.39 smaller), accuracy (up to 6.87% higher), and LLM-as-a-judge comparison (up to 76% win rate).

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.24945v1 Announce Type: new Abstract: Recent years have witnessed remarkable achievements of Large Language Models (LLMs) in multiple domains, while the excessive resour…
站內正文

待翻譯:ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.24938v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation. Yet low-latency MoE serving is increasingly challenging, because it spans two inference phases with fundamentally different bottlenecks: prefill is dominated by token-wise expert computation, whereas decode is constrained by memory traffic from the batch-wise activated expert set. However, existing training-free acceleration methods optimize only a single resource proxy, either the experts each token executes or the experts a batch activates, and either discard the excluded experts' contribution or leave it only implicitly approximated. In this paper, we propose ExFold, a unified training-free expert-folding framework for jointly accelerating MoE prefill and decode. ExFold casts both prefill and decode as one budgeted output-approximation problem: execute only a phase-specific constrained expert set while projecting the contribution of budget-excluded experts onto retained experts using calibrated scalar projectors. Motivated by the observation that many expert outputs are directionally aligned but differ in magnitude, ExFold calibrates a pairwise scalar-projector matrix on unlabeled data and uses it at inference time to fold excluded expert contributions into retained experts. Under this view, prefill acceleration becomes token-level Top-K folding, and decode acceleration becomes batch-level expert-pool folding. The two phases differ only in how retained experts are selected, while excluded contributions are recovered by one shared folding mechanism. We implement ExFold as a plug-and-play plugin in vLLM, with a lightweight expert-folding CUDA kernel, delivering up to 1.41x TTFT and 2.45x TPOT speedups while retaining about 99% of the original average quality.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.24938v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert act…
站內正文

待翻譯:Qwen3.8-Flash-Next

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:<p><strong><a href="https://qwen.ai/blog?id=qwen3.8-flash-next">Qwen3.8-Flash-Next</a></strong></p> Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4".</p> <p>It's pretty big: 125B tokens, but only 6B active which means it gets a pretty big performance boost.</p> <p>I've been trying it out on a DGX Spark using <a href="https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF">these Unsloth quantized models</a>. I'm still exploring the model - so far I've tried the 72.5GB UD-IQ1_S one (producing <a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff9c69ebdab90d8a45b8de4742cc7b840">these pelicans</a>) and the 78.9GB UD-Q2_K_XL (producing <a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F6ba7cbfc1a9336986703b41f7fccd73a">these</a>).</p> <p>My favorite so far was this xhigh reasoning effort one from UD-Q2_K_XL:</p> <p><img alt="Flat vector illustration: a white pelican with an orange beak and orange legs rides a red bicycle along a sandy path, a wicker basket on the handlebars holding a blue fish, with green rolling hills, a small tree and bushes, white clouds and a bright yellow sun in a blue sky behind it" src="https://static.simonwillison.net/static/2026-08-27/IMG_7667.png" /> <p><small></small>Via <a href="https://news.ycombinator.com/item?id=49448210">Hacker News</a></small></p> <p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/qwen">qwen</a>, <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle">pelican-riding-a-bicycle</a>, <a href="https://simonwillison.net/tags/ai-in-china">ai-in-china</a>, <a href="https://simonwillison.net/tags/nvidia-spark">nvidia-spark</a></p>

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • <p><strong><a href="https://qwen.ai/blog?id=qwen3.8-flash-next">Qwen3.8-Flash-Next</a></strong></p> Another open weights model from Qwen. This one is "a multimodal MoE model that…
站內正文

待翻譯:Nvidia is about to be a hundred-billion-dollar-a-quarter company

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Nvidia's predicting it will pull in $108 billion in revenue within just a few months. It wouldn't be the first company to rake in over $100 billion in quarterly revenue - Amazon, Apple, and Alphabet have repeatedly reached the milestone. Nvidia said in its latest earnings report that it brought in a record $96.2 billion in overall revenue in the past quarter, a jump of over $10 billion from the previous quarter. Its data center revenue alone more than doubled year-over-year to a record $89 billion, and the company's profits more than doubled to $59.7 billion. Nvidia's "edge computing" category, which includes its consumer gam … Read the full story at The Verge.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Nvidia's predicting it will pull in $108 billion in revenue within just a few months. It wouldn't be the first company to rake in over $100 billion in quarterly revenue - Amazon,…
站內正文

待翻譯:Meta's new MTIA 400 chip has a split personality: Training AI and serving ads

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Meta's new MTIA 400 chip has a split personality: Training AI and serving ads Faster than Blackwell, but still no replacement for AMD or Nvidia ... yet Tobias Mann Tobias Mann SYSTEMS EDITOR Published wed 26 Aug 2026 //…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Meta's new MTIA 400 chip has a split personality: Training AI and serving ads Faster than Blackwell, but still no replacement for AMD or Nvidia ... yet Tobias Mann Tobias Mann SYS…
站內正文

待翻譯:NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The next wave of AI is placing new demands on infrastructure. As AI agents and trillion-parameter workloads become mainstream, the performance of AI infrastructure depends not only on compute, but on how compute, memory, storage, networking and software are designed together as a unified system. To help hyperscalers and AI innovators build the next generation […]

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • The next wave of AI is placing new demands on infrastructure. As AI agents and trillion-parameter workloads become mainstream, the performance of AI infrastructure depends not onl…
站內正文

待翻譯:AMD, Supermicro and MinIO target the enterprise data pipeline bottleneck

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Despite rapid advances in artificial intelligence, the enterprise world is still dealing with a data pipeline problem. More than 80% of enterprise data is unstructured, and 99% of this data is dark to AI because there is no easy solution to query it, according to industry experts. Yet organizations are still trying to build AI […] The post AMD, Supermicro and MinIO target the enterprise data pipeline bottleneck appeared first on SiliconANGLE.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Despite rapid advances in artificial intelligence, the enterprise world is still dealing with a data pipeline problem. More than 80% of enterprise data is unstructured, and 99% of…
站內正文

待翻譯:Bring your own model with Amazon SageMaker AI: Script mode in SDK v3

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The SageMaker Python SDK v3 redesigns script mode with unified ModelTrainer and ModelBuilder classes. This post walks through two end-to-end examples, a scikit-learn Random Forest and a multi-GPU Stable Diffusion 3.5 LoRA fine-tune, showing how SourceCode syncs your local code into any container at runtime so you can iterate without rebuilding Docker images.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • The SageMaker Python SDK v3 redesigns script mode with unified ModelTrainer and ModelBuilder classes. This post walks through two end-to-end examples, a scikit-learn Random Forest…
站內正文

待翻譯:LangChain Announces Enterprise Agentic AI Platform Built with NVIDIA

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Build, deploy, and monitor production-grade AI agents at scale with LangChain's enterprise agentic AI platform integrated with NVIDIA.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Build, deploy, and monitor production-grade AI agents at scale with LangChain's enterprise agentic AI platform integrated with NVIDIA.
站內正文

待翻譯:Intel Crescent Island GPU Flexes 32 Xe3P Cores, 480GB LPDDR5X for Agentic AI

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Intel Crescent Island GPU Render - Image: Intel What kind of hardware do you need for AI processing? Well, every kind, because "AI processing" is a very broad term. Unlike a lot of specialized AI chips (e.g. d-Matrix Ra…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Intel Crescent Island GPU Render - Image: Intel What kind of hardware do you need for AI processing? Well, every kind, because "AI processing" is a very broad term. Unlike a lot o…
站內正文

待翻譯:Nvidia's Vera CPU outpaces AMD EPYC 9655P in Linux kernel compilation

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Photo: Steve A Johnson / Pexels Nvidia’s Vera CPU outpaces AMD EPYC 9655P in Linux kernel compilation at Hot Chips 2026 The chipmaker's new Vera CPU, Rubin GPU, and networking stack represent a coordinated bet that agen…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Photo: Steve A Johnson / Pexels Nvidia’s Vera CPU outpaces AMD EPYC 9655P in Linux kernel compilation at Hot Chips 2026 The chipmaker's new Vera CPU, Rubin GPU, and networking sta…
站內正文

待翻譯:Hope and concern swirl for Ohioans around ‘world’s largest datacenter’

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Piketon datacenter promises to generate thousands of jobs, but environmental groups voice concern over project On a winding road tucked away behind forests in the Appalachian foothills of southern Ohio is where OpenAI, Nvidia and Japanese investors are set to spend $500bn on one of the largest artificial intelligence datacenters on the planet. Last March, the energy secretary, Chris Wright, the commerce secretary, Howard Lutnick and a host of Japanese and other dignitaries briefly descended on Piketon to enthusiastically break ground on a project to build 8GW worth of AI computing power. Continue reading...

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Piketon datacenter promises to generate thousands of jobs, but environmental groups voice concern over project On a winding road tucked away behind forests in the Appalachian foot…
站內正文

待翻譯:Apple Updates Mini and Studio, AI Computers, OpenAI Jalapeño

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Apple Updates Mini and Studio, AI Computers, OpenAI Jalapeño Wednesday, August 26, 2026 Listen to Podcast Apple and OpenAI have two completely different hardware announcements; both represent pressure on Nvidia. Subscri…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Apple Updates Mini and Studio, AI Computers, OpenAI Jalapeño Wednesday, August 26, 2026 Listen to Podcast Apple and OpenAI have two completely different hardware announcements; bo…
站內正文

待翻譯:Nvidia Jetson Orin-guided Russian AI drone killed three civilians in Ukraine

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:5 Join the conversation Follow us Add us as a preferred source on Google A Russian Molniya drone carrying an Nvidia Jetson Orin module crashed and killed three civilians at a gas station in Zaporizhzhia last month after…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • 5 Join the conversation Follow us Add us as a preferred source on Google A Russian Molniya drone carrying an Nvidia Jetson Orin module crashed and killed three civilians at a gas…
站內正文

待翻譯:Perplexity's Portable Computer tackling local AI services market

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Now Perplexity is trying to get into the local AI action Amid talk of an Nvidia deal, the AI search biz is looking beyond the cloud Thomas Claburn Thomas Claburn AI AND SOFTWARE REPORTER Published wed 26 Aug 2026 // 00:…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Now Perplexity is trying to get into the local AI action Amid talk of an Nvidia deal, the AI search biz is looking beyond the cloud Thomas Claburn Thomas Claburn AI AND SOFTWARE R…
站內正文

待翻譯:Safety-aware Model Predictive Path Integral Control with Signal Temporal Logic

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:arXiv:2608.23972v1 Announce Type: new Abstract: Safety-aware motion planning remains a challenge in robotics, especially when missions are time-critical and are under complex specifications. In this paper, we propose safety-aware-stl-mppi, a computationally efficient sampling-based receding-horizon planning framework designed to promote satisfaction of constraints expressed in Signal Temporal Logic (STL). Our approach encodes discrete-time STL formulas into candidate time-varying control barrier functions (CBF), which are integrated into a model predictive path integral (MPPI) controller. Our method inherits the benefits of low computational cost from an efficiently parallelizable sampling based planner and utilizes CBF for constraints expressed in STL. We compare against several MPPI baselines using four artificial Mars Rover planning case studies with a diverse environment and cost setups, where we show our method consistently achieving high safety and efficiency. We show a quadcopter planning experiment with NVIDIA Isaac Lab.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • arXiv:2608.23972v1 Announce Type: new Abstract: Safety-aware motion planning remains a challenge in robotics, especially when missions are time-critical and are under complex spec…
站內正文

待翻譯:Perplexity AI launches Portable Computer on-device AI agent

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Perplexity AI Inc. today introduced Portable Computer, an artificial intelligence agent designed to run on desktops equipped with Nvidia Corp. silicon. The launch follows a report that Nvidia is weighing an investment in the startup that could value it at over $30 billion. Furthermore, Nvidia has reportedly floated the idea of licensing Perplexity’s technology and […] The post Perplexity AI launches Portable Computer on-device AI agent appeared first on SiliconANGLE.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Perplexity AI Inc. today introduced Portable Computer, an artificial intelligence agent designed to run on desktops equipped with Nvidia Corp. silicon. The launch follows a report…
站內正文

待翻譯:AI inference gets a new tier as context windows grow

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AI storage infrastructure is becoming a more consequential planning issue as organizations move from model training toward agentic AI. As agents reason, act and reassess, they build longer contexts and generate more data that they must access quickly during inference. Agentic AI is also changing the shape of the data problem. Interactions are growing longer […] The post AI inference gets a new tier as context windows grow appeared first on SiliconANGLE.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • AI storage infrastructure is becoming a more consequential planning issue as organizations move from model training toward agentic AI. As agents reason, act and reassess, they bui…
站內正文

待翻譯:Perplexity Ships Portable Computer on NVIDIA DGX Spark: Local Harness, OS-Enforced Sandbox, and Zero Per-Token Cost for Local Steps

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Perplexity releases Portable Computer, packaging local models, harness, sandbox, and connectors into one system running on NVIDIA DGX Spark. The post Perplexity Ships Portable Computer on NVIDIA DGX Spark: Local Harness, OS-Enforced Sandbox, and Zero Per-Token Cost for Local Steps appeared first on MarkTechPost.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Perplexity releases Portable Computer, packaging local models, harness, sandbox, and connectors into one system running on NVIDIA DGX Spark. The post Perplexity Ships Portable Com…
站內正文

待翻譯:Taiwan charges nine people for smuggling ‘high-end’ AI servers to China

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Among those charged are two Super Micro employees and one from Nvidia, marking another flashpoint in US-China AI rivalry Taiwanese prosecutors charged nine people Monday, including one from Nvidia and two from Super Micro, for illegally exporting “high-end AI servers” to mainland China, adding another wave of turbulence in the AI ​​rivalry between China and the United States. Prosecutors said the servers involved were graphics processing units known as “B300,” which have been banned from sale to China. Continue reading...

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Among those charged are two Super Micro employees and one from Nvidia, marking another flashpoint in US-China AI rivalry Taiwanese prosecutors charged nine people Monday, includin…
站內正文

待翻譯:Trump-backed Ken Paxton unveils ‘Texas first’ datacenter plan aimed at ‘negative impact’

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The Republican Senate candidate announced the plan with a graphic X flagged as ‘Made with AI’ US politics live – latest updates Ken Paxton, a Republican candidate for the US Senate in Texas, has unveiled plans to tackle the “negative impact” of datacenters amid a growing backlash across the political spectrum. Paxton, endorsed by Donald Trump but facing a close contest against Democrat James Talarico, posted a “Texas first” datacenter plan on X with a graphic that the social media network flagged as “Made with AI”. Continue reading...

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • The Republican Senate candidate announced the plan with a graphic X flagged as ‘Made with AI’ US politics live – latest updates Ken Paxton, a Republican candidate for the US Senat…
站內正文

待翻譯:Nvidia's Groq 3 LPX accelerator enters full production at 3,500 tokens/SEC

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Chipmaker Nvidia Corp. says its dedicated artificial intelligence inference accelerator Groq 3 LPX has now entered full production as it strives to maintain its dominance in the world of AI compute. The new chip, announ…

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Chipmaker Nvidia Corp. says its dedicated artificial intelligence inference accelerator Groq 3 LPX has now entered full production as it strives to maintain its dominance in the w…
站內正文

待翻譯:Perplexity’s Computer agent can now run locally — if you can afford it

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Perplexity, in partnership with Nvidia, has taken Computer, its agentic AI assistant, and brought it to the desktop in the The post Perplexity’s Computer agent can now run locally — if you can afford it appeared first on The New Stack.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Perplexity, in partnership with Nvidia, has taken Computer, its agentic AI assistant, and brought it to the desktop in the The post Perplexity’s Computer agent can now run locally…
站內正文

待翻譯:Leading Publishers Bring Blockbuster PC Games and Technology to NVIDIA RTX Spark

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:NVIDIA is bringing the next wave of RTX gaming to the Gamescom conference running this week in Cologne, Germany, with support for new games, anti-cheat technologies and increased visual quality. Electronic Arts, Embark and Ubisoft are among the latest game publishers and developers bringing their blockbuster titles to NVIDIA RTX Spark ahead of its launch […]

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • NVIDIA is bringing the next wave of RTX gaming to the Gamescom conference running this week in Cologne, Germany, with support for new games, anti-cheat technologies and increased…
站內正文

待翻譯:Nvidia doubles compute for entry-level edge robotics with Jetson Orin Nano 2

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Nvidia Corp. today announced the release of Jetson Orin Nano 2, a robotics computer “brain” for running artificial intelligence and frontier-level models at the edge. In the past months, foundational AI models have grown smaller and more efficient, adding numerous capabilities alongside language understanding, computer vision and audio processing. As more AI models compress in […] The post Nvidia doubles compute for entry-level edge robotics with Jetson Orin Nano 2 appeared first on SiliconANGLE.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Nvidia Corp. today announced the release of Jetson Orin Nano 2, a robotics computer “brain” for running artificial intelligence and frontier-level models at the edge. In the past…
站內正文

待翻譯:AI factories enter the execution era as Cisco and NVIDIA push rack-scale systems into production

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The artificial intelligence infrastructure market is crossing an important threshold. The conversation is shifting from acquiring graphics processing units to building complete AI factories that can generate tokens reliably, efficiently and at scale. This is the next bottleneck. GPUs may be the engine, but an AI factory is a system. Compute, networking, storage, cooling, software […] The post AI factories enter the execution era as Cisco and NVIDIA push rack-scale systems into production appeared first on SiliconANGLE.

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • The artificial intelligence infrastructure market is crossing an important threshold. The conversation is shifting from acquiring graphics processing units to building complete AI…
站內正文

更多增長標籤

GPU 基礎設施 AI News | AI News Hub