AI News HubLIVE

芯片动态

待翻译:Show HN: KinoPipe – FFmpeg as a service for AI agents (typed ops, no shell)

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:REC · YOUR AGENT IS EDITING FFmpeg as a service, built for agents. Typed operations your agent calls over MCP or REST. Trim, resize, compress, convert. A validated request in, a finished file out. No shell, ever. Start…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • REC · YOUR AGENT IS EDITING FFmpeg as a service, built for agents. Typed operations your agent calls over MCP or REST. Trim, resize, compress, convert. A validated request in, a f…
站内正文

待翻译:Nvidia Targets Physical AI With New Jetson Edge Platform

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The Jetson Orin Nano 2 comes as the AI chip giant targets physical AI expansion.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • The Jetson Orin Nano 2 comes as the AI chip giant targets physical AI expansion.
站内正文

待翻译:Hugging Face’s new robot is an adorable rollerskating duck

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Hugging Face's Pollen Robotics has launched its second cute AI robot, the Microduck, a one-eyed biped standing just under 10 inches tall. It's available to preorder now for $399 in cream, graphite, lavender, and sky blue, and Pollen Robotics says it plans to start shipping the little robot "before Christmas 2026." Video demos of the Microduck show it picking up socks and markers, kicking around a ball, and zipping around on tiny rollerskates. Pollen Robotics says it can also "react to its surroundings" and follow around a laser pointer, or users can control it with a game controller. The Microduck's software is open-source, like Hugging Fa … Read the full story at The Verge.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Hugging Face's Pollen Robotics has launched its second cute AI robot, the Microduck, a one-eyed biped standing just under 10 inches tall. It's available to preorder now for $399 i…
站内正文

待翻译:Is your really old iPhone still worth anything or just e-waste? How to tell

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:I found my old iPhone 4 in a drawer and wondered: Does it work? Is it worth anything? What should I do with it? I went looking for answers.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • I found my old iPhone 4 in a drawer and wondered: Does it work? Is it worth anything? What should I do with it? I went looking for answers.
站内正文

待翻译:AWS, Nvidia Expand Partnership With 2 Million More GPUs

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The companies are expanding their collaboration beyond GPUs into CPUs, government AI infrastructure and robotics.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • The companies are expanding their collaboration beyond GPUs into CPUs, government AI infrastructure and robotics.
站内正文

待翻译:GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:NVIDIA’s Gamescom announcements are revealing what’s next for GeForce NOW, with new ways to play, more supported devices and platforms, and even more big PC games headed to the cloud. New NVIDIA DLSS 4.5 technology controls give members more ways to fine-tune gameplay, while expanded support for new Steam devices, GOG single sign-on, Firefox browser […]

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • NVIDIA’s Gamescom announcements are revealing what’s next for GeForce NOW, with new ways to play, more supported devices and platforms, and even more big PC games headed to the cl…
站内正文

待翻译:Qwen3.8-Flash-Next: How to Run Locally

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:For the complete documentation index, see llms.txt. This page is also available as Markdown. Qwen3.8-Flash-Next is a new open-weight, 125B parameter MoE multimodal model from Qwen. Built on the new Qwen4 architecture, i…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • For the complete documentation index, see llms.txt. This page is also available as Markdown. Qwen3.8-Flash-Next is a new open-weight, 125B parameter MoE multimodal model from Qwen…
站内正文

待翻译:The Independent AI Coding Community for Cursor, Claude Code and LLMs

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The Independent AI Coding Community AI Tools Search & browse all AI tools AI Jobs International roles · opportunities Creative Studio Image · Video · Training AI Models Curated models, explained AI Skills Handy prompts,…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • The Independent AI Coding Community AI Tools Search & browse all AI tools AI Jobs International roles · opportunities Creative Studio Image · Video · Training AI Models Curated mo…
站内正文

待翻译:Give this skill to your AI to build decks for you

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Notifications You must be signed in to change notification settings Fork 0 Star 0 BranchesTags Open more actions menu Latest commit History 4 Commits 4 Commits Folders and files NameName Last commit message Last commit…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Notifications You must be signed in to change notification settings Fork 0 Star 0 BranchesTags Open more actions menu Latest commit History 4 Commits 4 Commits Folders and files N…
站内正文

待翻译:The Sequence Opinion #921: AI’s Sixth Layer Is Finance

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Jensen Huang’s five-layer cake explains how intelligence is manufactured. The missing layer explains how quickly - and by whom - it can scale.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Jensen Huang’s five-layer cake explains how intelligence is manufactured. The missing layer explains how quickly - and by whom - it can scale.
站内正文

待翻译:GLM 5.3 Flash faster and cheaper

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:GLM 5.3 Flash | Model APIs | RunInfra RunInfraby RightNow © 2026 RunInfra. All rights reserved. Join the communitySystem status Backed by Combinator AICPA Type II SOC 2 Ask AI about RunInfra Part of RightNow RunInfraby…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • GLM 5.3 Flash | Model APIs | RunInfra RunInfraby RightNow © 2026 RunInfra. All rights reserved. Join the communitySystem status Backed by Combinator AICPA Type II SOC 2 Ask AI abo…
站内正文

待翻译:The people, the money and the ownership behind China's leading AI companies

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:China & AI | WireScreen Briefings CareersProduct WWIRESCREEN · SPECIAL REPORT · CHINA AND ARTIFICIAL INTELLIGENCE WS-2026-034 · AUGUST 2026 Built and Owned The people, the money and the ownership behind China’s leading…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • China & AI | WireScreen Briefings CareersProduct WWIRESCREEN · SPECIAL REPORT · CHINA AND ARTIFICIAL INTELLIGENCE WS-2026-034 · AUGUST 2026 Built and Owned The people, the money a…
站内正文

待翻译:Nvidia NVLink Fusion Brings Nvhbm to Next-Generation AI Infrastructure

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:AI factories must support increasingly large models and more complex reasoning workloads. To keep up with the insatiable compute demands of AI workloads, hyperscalers and AI-native companies are developing custom AI acc…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • AI factories must support increasingly large models and more complex reasoning workloads. To keep up with the insatiable compute demands of AI workloads, hyperscalers and AI-nativ…
站内正文

待翻译:Google Research Introduces GlucoFM: A 0.72M-Parameter Dual-Stream Foundation Model for Continuous Glucose Monitoring

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Google Research and UNSW Sydney released GlucoFM, a self-supervised foundation model that splits a CGM trace into a slow physiological stream and a transient event stream instead of encoding it as one sequence. At 0.72M parameters it reached 58.8 task-averaged PR-AUC across 14 cohort–task evaluations, beating a 135M GluFormer and a 385M MOMENT. It remains a research prototype with no regulatory clearance. The post Google Research Introduces GlucoFM: A 0.72M-Parameter Dual-Stream Foundation Model for Continuous Glucose Monitoring appeared first on MarkTechPost.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Google Research and UNSW Sydney released GlucoFM, a self-supervised foundation model that splits a CGM trace into a slow physiological stream and a transient event stream instead…
站内正文

待翻译:CRESSim-Neo: A Batched GPU Simulation Engine for Surgical Robotics and Robot Learning

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25192v1 Announce Type: new Abstract: We introduce CRESSim-Neo, a batched GPU simulation engine for surgical robotics and robot learning. CRESSim-Neo combines position-based simulation of rigid bodies, deformable tissues, fluids, and strands with batched rendering, surgery-specific sensing, and a GPU-resident data pipeline. The engine supports applications including tissue manipulation, fluid suction, suturing, cable-driven robots, and ultrasound image synthesis. Direct access to physics and rendering buffers enables GPU-resident robot learning and zero-copy PyTorch integration using DLPack. We demonstrate CRESSim-Neo across rigid-body, deformable-body, and fluid simulation tasks, including vision-based and surgical robot-learning scenarios. On an NVIDIA RTX 4090, the engine achieves up to 2.03 million environment steps per second for 8192 parallel CartPole environments, and scales to batched surgical scenarios involving tissue deformation, fluid interaction, and ultrasound sensing. Overall, CRESSim-Neo provides a unified and scalable platform for surgical simulation, synthetic data generation, and surgical robot learning.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25192v1 Announce Type: new Abstract: We introduce CRESSim-Neo, a batched GPU simulation engine for surgical robotics and robot learning. CRESSim-Neo combines position-b…
站内正文

待翻译:A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24935v1 Announce Type: new Abstract: Accurate identification of early-stage apple fruitlet anatomical structures, including the calyx, fruitlet body, and peduncle, is essential for robotic thinning, crop-load management, and other precision orchard operations. This study presents a lightweight multimodal vision-language framework that adapts TinyCLIP for fine-grained fruitlet anatomy classification in complex orchard environments. A dataset of 600 high-resolution RGB images collected from Scilate and Scifresh apple orchards was converted into 224 x 224 image patches and annotated for three anatomical classes. Domain-specific language prompts, such as ``a photo of a class,'' were used to guide multimodal alignment between orchard imagery and horticultural structures. A sliding-window inference strategy with a stride of 112 pixels aggregates patch-level predictions into spatial heatmaps, enabling interpretable whole-image localization of fruitlet components relevant to robotic thinning. Patch-level evaluation on an NVIDIA T4 GPU achieved F1-scores of 0.95 for calyx, 0.98 for fruitlet, and 0.85 for peduncle, with a macro-F1 score of 0.93. Deployment-oriented optimization using ONNX and TensorRT enabled efficient inference on NVIDIA Jetson hardware, preserved accuracy under INT8 quantization, and supported model sizes of approximately 127-137 MB with millisecond-level patch inference. These results demonstrate that lightweight vision-language models can provide interpretable and edge-deployable perception for automated fruitlet analysis and future robotic thinning systems. The source code and implementation details are publicly available at https://github.com/WilliamBu1/A-Lightweight-Vision-Language-Model-for-Early-Stage-Fruitlet-Classification-in-Apple-Orchards.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.24935v1 Announce Type: new Abstract: Accurate identification of early-stage apple fruitlet anatomical structures, including the calyx, fruitlet body, and peduncle, is e…
站内正文

待翻译:DataKernelBench: Can LLMs Optimize Database Queries on GPUs?

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25061v1 Announce Type: new Abstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Existing LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement-heavy database-style operators untested. We introduce DataKernelBench, which translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that optimize either the core tensor-bounded snippet or the full query in CUDA or Triton through execution-guided repair. Across ten proprietary and open-weight models on TPC-H SF10 with an H100 GPU, the strongest full-query CUDA configuration achieves $2.11\times$ speedup over torch.compile at full pass rate. We find that higher-performing implementations commonly use kernel fusion and execution-strategy changes, stronger models benefit most from full-query specialization, and workload context matters more than hardware context. To handle data larger than GPU memory, we extend TorchPlan with Dask-cuDF for on-demand partition loading on TPC-H SF100 with four H100 GPUs, achieving $2.54\times$ speedup

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25061v1 Announce Type: new Abstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Exis…
站内正文

待翻译:MacroAgent: Regularity-Aware Macro Legalization with LLM-Agent-Designed Contour Algorithms

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24946v1 Announce Type: new Abstract: Macros constitute a large part of the core area in modern very large-scale integration (VLSI) designs. Moreover, macro positions have a significant impact on the final quality of result (QoR), and macro legalization is typically the final step in determining the macro positions. However, existing approaches related to macro legalization either lack robustness or incur substantial computational costs or neglect the regularity between macros. To address these limitations, we introduce MacroAgent. The novel framework is a four-stage approach: clustering, contour generation, template matching, and inter-cluster refinement. We propose leveraging Large Language Models (LLMs) to discover multiple, effective heuristic regularity-aware contour algorithms. This framework successfully generates robust and effective algorithmic solutions for macro legalization. Compared with state-of-the-art macro legalization works, experimental results on TILOS and Chipyard benchmarks demonstrate a 2 to 8 fold improvement in layout regularity, a 3% to 5% reduction in routed wirelength with comparable congestion after global routing, and significantly better robustness with an acceptable runtime. Furthermore, end-to-end evaluation through Cadence Innovus place-and-route confirms that the regularity improvements translate into tangible PPA gains, including 2.9% lower routed wirelength and 68.3% TNS improvement over the DREAMPlace macro legalization baseline; it also achieves 1.8% lower routed wirelength when integrated into the Innovus macro placement flow.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.24946v1 Announce Type: new Abstract: Macros constitute a large part of the core area in modern very large-scale integration (VLSI) designs. Moreover, macro positions ha…
站内正文

待翻译:FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24945v1 Announce Type: new Abstract: Recent years have witnessed remarkable achievements of Large Language Models (LLMs) in multiple domains, while the excessive resource requirements of LLMs hinder the deployment on resource-constrained devices. Although model quantization stands out as an effective approach, conventional quantization approaches typically incur severe performance degradation due to uniform bit-width or simple heuristic sensitivity evaluation. In this paper, we propose a novel Fisher information-based Adaptive Mixed Precision Weight Quantization approach, i.e., FAMPWQ, which performs layer-adaptive weight quantization for effective LLM inference on commodity GPUs. First, we propose a system model with a novel Fisher information metric to measure the layer-wise sensitivity to quantization. Second, we propose a reinforcement learning-based bit-width allocator in FAMPWQ, which generates an adaptive bit-width allocation strategy based on the Fisher information sensitivity metric. Extensive experiments on 7 models and 5 benchmarks demonstrate that FAMPWQ significantly outperforms 7 baseline approaches in terms of PPL (up to 3.39 smaller), accuracy (up to 6.87% higher), and LLM-as-a-judge comparison (up to 76% win rate).

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.24945v1 Announce Type: new Abstract: Recent years have witnessed remarkable achievements of Large Language Models (LLMs) in multiple domains, while the excessive resour…
站内正文

待翻译:ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24938v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation. Yet low-latency MoE serving is increasingly challenging, because it spans two inference phases with fundamentally different bottlenecks: prefill is dominated by token-wise expert computation, whereas decode is constrained by memory traffic from the batch-wise activated expert set. However, existing training-free acceleration methods optimize only a single resource proxy, either the experts each token executes or the experts a batch activates, and either discard the excluded experts' contribution or leave it only implicitly approximated. In this paper, we propose ExFold, a unified training-free expert-folding framework for jointly accelerating MoE prefill and decode. ExFold casts both prefill and decode as one budgeted output-approximation problem: execute only a phase-specific constrained expert set while projecting the contribution of budget-excluded experts onto retained experts using calibrated scalar projectors. Motivated by the observation that many expert outputs are directionally aligned but differ in magnitude, ExFold calibrates a pairwise scalar-projector matrix on unlabeled data and uses it at inference time to fold excluded expert contributions into retained experts. Under this view, prefill acceleration becomes token-level Top-K folding, and decode acceleration becomes batch-level expert-pool folding. The two phases differ only in how retained experts are selected, while excluded contributions are recovered by one shared folding mechanism. We implement ExFold as a plug-and-play plugin in vLLM, with a lightweight expert-folding CUDA kernel, delivering up to 1.41x TTFT and 2.45x TPOT speedups while retaining about 99% of the original average quality.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.24938v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert act…
站内正文

待翻译:S1: In-Context Learning for Robotics

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:0:00 / 0:00 Introducing S1: In-Context Learning for Robotics Unseen tasks10-minute horizonsOne video promptNo post-training 13-minute read Introduction The evolution of language modeling provides a blueprint for turning…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • 0:00 / 0:00 Introducing S1: In-Context Learning for Robotics Unseen tasks10-minute horizonsOne video promptNo post-training 13-minute read Introduction The evolution of language m…
站内正文

待翻译:Qwen3.8-Flash-Next

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:<p><strong><a href="https://qwen.ai/blog?id=qwen3.8-flash-next">Qwen3.8-Flash-Next</a></strong></p> Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4".</p> <p>It's pretty big: 125B tokens, but only 6B active which means it gets a pretty big performance boost.</p> <p>I've been trying it out on a DGX Spark using <a href="https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF">these Unsloth quantized models</a>. I'm still exploring the model - so far I've tried the 72.5GB UD-IQ1_S one (producing <a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff9c69ebdab90d8a45b8de4742cc7b840">these pelicans</a>) and the 78.9GB UD-Q2_K_XL (producing <a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F6ba7cbfc1a9336986703b41f7fccd73a">these</a>).</p> <p>My favorite so far was this xhigh reasoning effort one from UD-Q2_K_XL:</p> <p><img alt="Flat vector illustration: a white pelican with an orange beak and orange legs rides a red bicycle along a sandy path, a wicker basket on the handlebars holding a blue fish, with green rolling hills, a small tree and bushes, white clouds and a bright yellow sun in a blue sky behind it" src="https://static.simonwillison.net/static/2026-08-27/IMG_7667.png" /> <p><small></small>Via <a href="https://news.ycombinator.com/item?id=49448210">Hacker News</a></small></p> <p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/qwen">qwen</a>, <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle">pelican-riding-a-bicycle</a>, <a href="https://simonwillison.net/tags/ai-in-china">ai-in-china</a>, <a href="https://simonwillison.net/tags/nvidia-spark">nvidia-spark</a></p>

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • <p><strong><a href="https://qwen.ai/blog?id=qwen3.8-flash-next">Qwen3.8-Flash-Next</a></strong></p> Another open weights model from Qwen. This one is "a multimodal MoE model that…
站内正文

待翻译:Instinct.co Raises $350M

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:AI Assistant Instinct Hits $2.5 Billion Valuation In Weeks Amid VC Feeding Frenzy Editors' Pick VCs Are So Obsessed With This AI Assistant That Its Valuation Jumped Fivefold In Weeks A hot new AI agent called Instinct t…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • AI Assistant Instinct Hits $2.5 Billion Valuation In Weeks Amid VC Feeding Frenzy Editors' Pick VCs Are So Obsessed With This AI Assistant That Its Valuation Jumped Fivefold In We…
站内正文

待翻译:Deep Cogito raises $43M to develop self-improving AI models

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Artificial intelligence startup Deep Cogito Inc. today announced that it has raised $43 million in funding. TQ Ventures led the Series A round. It was joined by Benchmark, Nexus Venture Partners, Atreides Management, South Park Commons and Zscaler Inc., a publicly traded cybersecurity provider. The deal brings Deep Cogito’s total outside funding to more than […] The post Deep Cogito raises $43M to develop self-improving AI models appeared first on SiliconANGLE.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Artificial intelligence startup Deep Cogito Inc. today announced that it has raised $43 million in funding. TQ Ventures led the Series A round. It was joined by Benchmark, Nexus V…
站内正文

待翻译:Nvidia is about to be a hundred-billion-dollar-a-quarter company

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Nvidia's predicting it will pull in $108 billion in revenue within just a few months. It wouldn't be the first company to rake in over $100 billion in quarterly revenue - Amazon, Apple, and Alphabet have repeatedly reached the milestone. Nvidia said in its latest earnings report that it brought in a record $96.2 billion in overall revenue in the past quarter, a jump of over $10 billion from the previous quarter. Its data center revenue alone more than doubled year-over-year to a record $89 billion, and the company's profits more than doubled to $59.7 billion. Nvidia's "edge computing" category, which includes its consumer gam … Read the full story at The Verge.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Nvidia's predicting it will pull in $108 billion in revenue within just a few months. It wouldn't be the first company to rake in over $100 billion in quarterly revenue - Amazon,…
站内正文

待翻译:Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series — a 320B-total / 18B-active MoE with a 1,048,576-token context window, MIT-licensed weights on Hugging Face, and API pricing at $0.15/M input and $0.50/M output. It scores 84.3 on Terminal-Bench 2.1 and 63.4 on DeepSWE v1.1, using hybrid KDA linear plus NoPE sparse MLA attention to cut attention compute ~3× and KV cache 4.4× versus GLM-5.3. The post Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context appeared first on MarkTechPost.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series — a 320B-total / 18B-active MoE with a 1,048,576-token context window, MIT-licensed weight…
站内正文

待翻译:Meta's new MTIA 400 chip has a split personality: Training AI and serving ads

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Meta's new MTIA 400 chip has a split personality: Training AI and serving ads Faster than Blackwell, but still no replacement for AMD or Nvidia ... yet Tobias Mann Tobias Mann SYSTEMS EDITOR Published wed 26 Aug 2026 //…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Meta's new MTIA 400 chip has a split personality: Training AI and serving ads Faster than Blackwell, but still no replacement for AMD or Nvidia ... yet Tobias Mann Tobias Mann SYS…
站内正文

待翻译:NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The next wave of AI is placing new demands on infrastructure. As AI agents and trillion-parameter workloads become mainstream, the performance of AI infrastructure depends not only on compute, but on how compute, memory, storage, networking and software are designed together as a unified system. To help hyperscalers and AI innovators build the next generation […]

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • The next wave of AI is placing new demands on infrastructure. As AI agents and trillion-parameter workloads become mainstream, the performance of AI infrastructure depends not onl…
站内正文

待翻译:Gamescom highlights gaming boom amid AI concerns

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:https://p.dw.com/p/5JOAz AI is bringing significant challenges for the gaming industry, but it could also help significantly reduce costsImage: Political-Moments/IMAGO Earlier this month, gaming giant Electronic Arts wa…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • https://p.dw.com/p/5JOAz AI is bringing significant challenges for the gaming industry, but it could also help significantly reduce costsImage: Political-Moments/IMAGO Earlier thi…
站内正文

待翻译:Show HN: AI scientist builds an open-source Codex Micro from scratch for $40

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:TL;DR: AgentPad13 is an open-source take on the $230 Codex Micro and a test of whether an autonomous scientist can teach itself PCB routing. We wanted a more wallet-friendly, open-source Codex Micro, so we asked Marvin…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • TL;DR: AgentPad13 is an open-source take on the $230 Codex Micro and a test of whether an autonomous scientist can teach itself PCB routing. We wanted a more wallet-friendly, open…
站内正文

待翻译:AMD, Supermicro and MinIO target the enterprise data pipeline bottleneck

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Despite rapid advances in artificial intelligence, the enterprise world is still dealing with a data pipeline problem. More than 80% of enterprise data is unstructured, and 99% of this data is dark to AI because there is no easy solution to query it, according to industry experts. Yet organizations are still trying to build AI […] The post AMD, Supermicro and MinIO target the enterprise data pipeline bottleneck appeared first on SiliconANGLE.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Despite rapid advances in artificial intelligence, the enterprise world is still dealing with a data pipeline problem. More than 80% of enterprise data is unstructured, and 99% of…
站内正文

待翻译:Bring your own model with Amazon SageMaker AI: Script mode in SDK v3

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The SageMaker Python SDK v3 redesigns script mode with unified ModelTrainer and ModelBuilder classes. This post walks through two end-to-end examples, a scikit-learn Random Forest and a multi-GPU Stable Diffusion 3.5 LoRA fine-tune, showing how SourceCode syncs your local code into any container at runtime so you can iterate without rebuilding Docker images.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • The SageMaker Python SDK v3 redesigns script mode with unified ModelTrainer and ModelBuilder classes. This post walks through two end-to-end examples, a scikit-learn Random Forest…
站内正文

待翻译:Z.ai’s GLM-5.3 Flash is cheap, good, and served on Chinese chips

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Ox-alpha, the stealth model that quickly became the most popular model on OpenRouter in the last few days, is actually The post Z.ai’s GLM-5.3 Flash is cheap, good, and served on Chinese chips appeared first on The New Stack.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Ox-alpha, the stealth model that quickly became the most popular model on OpenRouter in the last few days, is actually The post Z.ai’s GLM-5.3 Flash is cheap, good, and served on…
站内正文

待翻译:The Importance of Reading (and Teaching) Cyberpunk in the Age of AI

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The Importance of Reading (and Teaching) Cyberpunk in the Age of AI - Reactor 0 Share Featured Essays Cyberpunk The Importance of Reading (and Teaching) Cyberpunk in the Age of AI Looking for answers — and finding hope…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • The Importance of Reading (and Teaching) Cyberpunk in the Age of AI - Reactor 0 Share Featured Essays Cyberpunk The Importance of Reading (and Teaching) Cyberpunk in the Age of AI…
站内正文

待翻译:LangChain Announces Enterprise Agentic AI Platform Built with NVIDIA

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Build, deploy, and monitor production-grade AI agents at scale with LangChain's enterprise agentic AI platform integrated with NVIDIA.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Build, deploy, and monitor production-grade AI agents at scale with LangChain's enterprise agentic AI platform integrated with NVIDIA.
站内正文

待翻译:Google Pixel 11 Review: Why the base model is still my favorite, even this year

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:In a year of iterative upgrades, Google is introducing smart features that make the base Pixel still the one to buy for most people.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • In a year of iterative upgrades, Google is introducing smart features that make the base Pixel still the one to buy for most people.
站内正文

待翻译:Show HN: I built a tool that finds people asking for what you sell

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Launch price: $49/month, yours until you cancel. It rises after the first 100 customers. ReachFastSign in Your next customers are already asking on ReachFast finds people asking for your products across X, Reddit, Linke…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Launch price: $49/month, yours until you cancel. It rises after the first 100 customers. ReachFastSign in Your next customers are already asking on ReachFast finds people asking f…
站内正文

待翻译:Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:We look at Qwen3.8-Flash-Next, Alibaba's open-weight multimodal Mixture-of-Experts model and an early preview of the Qwen4 architecture. We break down where the 180B parameters actually sit: a 125B backbone, a 51B N-gram embedding table, and a 4B multi-token prediction module, with only 6B active per token. We walk through the four architectural changes — the Gated DeltaNet and Qwen Sparse Attention hybrid, Gated Residual, N-gram Embedding, and the Muon optimizer. We also cover the benchmark results, the reported 1/9 training cost against Qwen3.7-Plus, and what self-hosting a 172.78 GiB FP8 checkpoint really demands. The post Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture appeared first on MarkTechPost.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • We look at Qwen3.8-Flash-Next, Alibaba's open-weight multimodal Mixture-of-Experts model and an early preview of the Qwen4 architecture. We break down where the 180B parameters ac…
站内正文

待翻译:🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Anima Anandkumar has spent two decades in AI, from classical math to deep learning and back. Now she's using it to model the physical world, from weather to fusion reactors.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Anima Anandkumar has spent two decades in AI, from classical math to deep learning and back. Now she's using it to model the physical world, from weather to fusion reactors.
站内正文

待翻译:Intel Crescent Island GPU Flexes 32 Xe3P Cores, 480GB LPDDR5X for Agentic AI

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Intel Crescent Island GPU Render - Image: Intel What kind of hardware do you need for AI processing? Well, every kind, because "AI processing" is a very broad term. Unlike a lot of specialized AI chips (e.g. d-Matrix Ra…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Intel Crescent Island GPU Render - Image: Intel What kind of hardware do you need for AI processing? Well, every kind, because "AI processing" is a very broad term. Unlike a lot o…
站内正文

待翻译:Nvidia's Vera CPU outpaces AMD EPYC 9655P in Linux kernel compilation

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Photo: Steve A Johnson / Pexels Nvidia’s Vera CPU outpaces AMD EPYC 9655P in Linux kernel compilation at Hot Chips 2026 The chipmaker's new Vera CPU, Rubin GPU, and networking stack represent a coordinated bet that agen…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Photo: Steve A Johnson / Pexels Nvidia’s Vera CPU outpaces AMD EPYC 9655P in Linux kernel compilation at Hot Chips 2026 The chipmaker's new Vera CPU, Rubin GPU, and networking sta…
站内正文

待翻译:Hope and concern swirl for Ohioans around ‘world’s largest datacenter’

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Piketon datacenter promises to generate thousands of jobs, but environmental groups voice concern over project On a winding road tucked away behind forests in the Appalachian foothills of southern Ohio is where OpenAI, Nvidia and Japanese investors are set to spend $500bn on one of the largest artificial intelligence datacenters on the planet. Last March, the energy secretary, Chris Wright, the commerce secretary, Howard Lutnick and a host of Japanese and other dignitaries briefly descended on Piketon to enthusiastically break ground on a project to build 8GW worth of AI computing power. Continue reading...

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Piketon datacenter promises to generate thousands of jobs, but environmental groups voice concern over project On a winding road tucked away behind forests in the Appalachian foot…
站内正文

待翻译:New Platform Peers Inside AI’s Black Box

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Prompt Claude, ChatGPT, Gemini, or any other popular large language model (LLM) with a question like “What is the best film ever made?” and the response will vary, and you (and most worryingly, the people who built the LLM) have little idea exactly how it came up with that specific answer. This mysterious behavior can be useful in some situations. But—as a recent incident where OpenAI could not explain why its advanced pre-release model hacked AI company Hugging Face highlighted—it can have negative and alarming consequences too. And when frontier AI models are writing code, generating results humans could not achieve alone, and performing other important tasks across society, the need to interpret AI ‘thinking’ and outputs has never been greater. Goodfire, an AI lab focused solely on this very problem, recently made its cutting-edge Silico platform, filled with tools to interpret the behavior of AI, generally available to the public. As part of this, the company recently announced a new grant program offering $1 million in free Silico usage for academic and nonprofit interpretability researchers. These efforts aim to democratize AI interpretability, placing techniques previously available to a clutch of elite labs into the hands of ambitious research teams and startups that want to build and understand their own models or adapt open-source models for different purposes. Mechanistic interpretability Founded in 2024 and based in San Francisco, Goodfire aims to provide the tools that build the next generation of safe and powerful AI by understanding the structures inside them instead of treating AI models as black boxes. “Treating models like black boxes isn’t inevitable, it’s a choice,” says Eric Ho, Goodfire co-founder and CEO. “With the right interpretability tools, we can see how models actually work.” The tools Ho refers to are built around a concept called mechanistic interpretability, which aims to understand what goes on inside an AI model when it carries out a task by interpreting the model’s weights, activations, and attention patterns, and mapping its neurons and the pathways between them. Mechanistic interpretability tools span the gamut. One approach is mapping a model’s activations in response to controlled prompts, and matching those activation patterns to a set of human-understandable concepts. Another tack is tracking changes in model weights before and after a specific training run in order to spot and understand what changed. Yet another option is changing specific model weights or activations and observing how that affects the model’s output. With Silico, uSilico combines a broad range of these tools, and provides a layer of AI agents to help users understand their model. Users describe what they want to investigate about their AI model in plain language, asking things like ‘Find out when and why my model is hallucinating.’ The platform then autonomously builds an experimental plan involving a host of tasks that can be performed using the various interpretability tools and techniques at its disposal. It then sends out agents to perform these tasks in parallel. Completion of these subtasks should add up to an answer to the original prompt, or at least insights that can be inspected and built upon. Ho says: “In a sense, Silico is like a microscope to peer inside an AI model to understand which parts are responsible for what behavior, and even edit those parts directly.” Understanding Alzheimer’s and AI These tools have already been used to make some impressive advances in a host of fields. In medicine, for instance, Prima Mente, a UK-based AI company, worked with Goodfire to understand its Pleiades epigenetic foundation model. The model performed well at its task of detecting Alzheimer’s disease from blood samples, but the company didn’t know why. “We reverse-engineered Pleiades and found it was using DNA fragment-length patterns to make its predictions—a signal humans hadn’t used to detect Alzheimer’s before,” recalls Ho. In other words, the team had discovered that Pleiades was using a completely new biomarker for the disease. “As far as we know, it’s the first significant finding in the natural sciences discovered purely by reverse-engineering a foundation model,” Ho adds. Elsewhere, Silico is being used to explore deep questions surrounding AI. Cameron Berg, Founder and Director of Reciprocal Research (a New York nonprofit research organization he founded to explore methods of gauging AI cognition), says that Silico almost fell out of the sky at the right time for him and his research. “Silico has been really helpful for operationalizing my research agenda and executing on it way faster than I would have expected,” he says. “ I feel like I have basically become the PI and my research scientists and research engineers are AI systems.” Berg sees general access to Silico and tools like it leading to greater trust in AI’s ability to conduct research tasks, which will accelerate the scientific process across the board. But beyond scientific research, the widespread release of Silico could signal a shift in how AI innovators build, debug, and deploy their models. “I think it’s a mistake to not understand the most consequential technology of our time, particularly given the emergent behavior we’re seeing from increasingly capable AI agents,” says Ho. “If we truly understand how AI models think, instead of discovering and trying to correct their behavior retroactively, we can design them intentionally and shape how models behave to be safer and more reliable.”

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Prompt Claude, ChatGPT, Gemini, or any other popular large language model (LLM) with a question like “What is the best film ever made?” and the response will vary, and you (and mo…
站内正文

待翻译:Moonshot AI wants 30% of what US clouds earn from Kimi K3

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:China’s Moonshot AI is in early talks with Microsoft, Amazon and Google to host Kimi K3 Credit: Bangla press via Shutterstock.com Moonshot AI is in early discussions with Microsoft, Amazon, and Google about hosting Kimi…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • China’s Moonshot AI is in early talks with Microsoft, Amazon and Google to host Kimi K3 Credit: Bangla press via Shutterstock.com Moonshot AI is in early discussions with Microsof…
站内正文

待翻译:Apple Updates Mini and Studio, AI Computers, OpenAI Jalapeño

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Apple Updates Mini and Studio, AI Computers, OpenAI Jalapeño Wednesday, August 26, 2026 Listen to Podcast Apple and OpenAI have two completely different hardware announcements; both represent pressure on Nvidia. Subscri…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Apple Updates Mini and Studio, AI Computers, OpenAI Jalapeño Wednesday, August 26, 2026 Listen to Podcast Apple and OpenAI have two completely different hardware announcements; bo…
站内正文

待翻译:When Smaller Models Win

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Even the best AI models can suck at chess The launch of ChatGPT had an interesting effect on the online chess discourse. Chess has already long been conquered by machines. As early as 1996 a computer (IBM’s Deep Blue) was able to beat the human world champion, grandmaster Garry Kasparov, in a game watched by […]

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Even the best AI models can suck at chess The launch of ChatGPT had an interesting effect on the online chess discourse. Chess has already long been conquered by machines. As earl…
站内正文

待翻译:Dataset: AI agent security failures, 1000 incidents classified

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:...\n **config_kwargs,\n )\n"," File \"/usr/local/lib/python3.14/site-packages/datasets/inspect.py\", line 291, in get_dataset_config_info\n raise SplitsNotFoundError(\"The split names could not be parsed from the datas…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • ...\n **config_kwargs,\n )\n"," File \"/usr/local/lib/python3.14/site-packages/datasets/inspect.py\", line 291, in get_dataset_config_info\n raise SplitsNotFoundError(\"The split…
站内正文

主题导航

芯片 — AI 话题新闻 | AI News Hub