AI News HubLIVE

Source Mix

  • SiliconANGLE AI12
  • Hacker News AI10
  • NVIDIA Blog7
  • MarkTechPost5
  • AI Business3
  • arXiv AI2
  • arXiv Machine Learning1
  • AWS Machine Learning Blog1

Topic Mix

  • Chips49
  • Agents27
  • Models10
  • Research10
  • Startups7
  • Robotics3
  • Tools1

Timeline

  • 2026-08-2512
  • 2026-08-178
  • 2026-08-217
  • 2026-08-246
  • 2026-08-184
  • 2026-08-204
  • 2026-08-193
  • 2026-08-162

Latest Updates

Nvidia Jetson Orin-guided Russian AI drone killed three civilians in Ukraine

5 Join the conversation Follow us Add us as a preferred source on Google A Russian Molniya drone carrying an Nvidia Jetson Orin module crashed and killed three civilians at a gas station in Zaporizhzhia last month after…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • 5 Join the conversation Follow us Add us as a preferred source on Google A Russian Molniya drone carrying an Nvidia Jetson Orin module crashed and killed three civilians at a gas…
In-site article

Perplexity's Portable Computer tackling local AI services market

Now Perplexity is trying to get into the local AI action Amid talk of an Nvidia deal, the AI search biz is looking beyond the cloud Thomas Claburn Thomas Claburn AI AND SOFTWARE REPORTER Published wed 26 Aug 2026 // 00:…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Now Perplexity is trying to get into the local AI action Amid talk of an Nvidia deal, the AI search biz is looking beyond the cloud Thomas Claburn Thomas Claburn AI AND SOFTWARE R…
In-site article

Perplexity AI launches Portable Computer on-device AI agent

Perplexity AI Inc. today introduced Portable Computer, an artificial intelligence agent designed to run on desktops equipped with Nvidia Corp. silicon. The launch follows a report that Nvidia is weighing an investment in the startup that could value it at over $30 billion. Furthermore, Nvidia has reportedly floated the idea of licensing Perplexity’s technology and […] The post Perplexity AI launches Portable Computer on-device AI agent appeared first on SiliconANGLE.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Perplexity AI Inc. today introduced Portable Computer, an artificial intelligence agent designed to run on desktops equipped with Nvidia Corp. silicon. The launch follows a report…
In-site article

Perplexity Ships Portable Computer on NVIDIA DGX Spark: Local Harness, OS-Enforced Sandbox, and Zero Per-Token Cost for Local Steps

Perplexity releases Portable Computer, packaging local models, harness, sandbox, and connectors into one system running on NVIDIA DGX Spark. The post Perplexity Ships Portable Computer on NVIDIA DGX Spark: Local Harness, OS-Enforced Sandbox, and Zero Per-Token Cost for Local Steps appeared first on MarkTechPost.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Perplexity releases Portable Computer, packaging local models, harness, sandbox, and connectors into one system running on NVIDIA DGX Spark. The post Perplexity Ships Portable Com…
In-site article

Taiwan charges nine people for smuggling ‘high-end’ AI servers to China

Among those charged are two Super Micro employees and one from Nvidia, marking another flashpoint in US-China AI rivalry Taiwanese prosecutors charged nine people Monday, including one from Nvidia and two from Super Micro, for illegally exporting “high-end AI servers” to mainland China, adding another wave of turbulence in the AI ​​rivalry between China and the United States. Prosecutors said the servers involved were graphics processing units known as “B300,” which have been banned from sale to China. Continue reading...

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Among those charged are two Super Micro employees and one from Nvidia, marking another flashpoint in US-China AI rivalry Taiwanese prosecutors charged nine people Monday, includin…
In-site article

Nvidia's Groq 3 LPX accelerator enters full production at 3,500 tokens/SEC

Chipmaker Nvidia Corp. says its dedicated artificial intelligence inference accelerator Groq 3 LPX has now entered full production as it strives to maintain its dominance in the world of AI compute. The new chip, announ…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Chipmaker Nvidia Corp. says its dedicated artificial intelligence inference accelerator Groq 3 LPX has now entered full production as it strives to maintain its dominance in the w…
In-site article

Perplexity’s Computer agent can now run locally — if you can afford it

Perplexity, in partnership with Nvidia, has taken Computer, its agentic AI assistant, and brought it to the desktop in the The post Perplexity’s Computer agent can now run locally — if you can afford it appeared first on The New Stack.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Perplexity, in partnership with Nvidia, has taken Computer, its agentic AI assistant, and brought it to the desktop in the The post Perplexity’s Computer agent can now run locally…
In-site article

Leading Publishers Bring Blockbuster PC Games and Technology to NVIDIA RTX Spark

NVIDIA is bringing the next wave of RTX gaming to the Gamescom conference running this week in Cologne, Germany, with support for new games, anti-cheat technologies and increased visual quality. Electronic Arts, Embark and Ubisoft are among the latest game publishers and developers bringing their blockbuster titles to NVIDIA RTX Spark ahead of its launch […]

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • NVIDIA is bringing the next wave of RTX gaming to the Gamescom conference running this week in Cologne, Germany, with support for new games, anti-cheat technologies and increased…
In-site article

Nvidia doubles compute for entry-level edge robotics with Jetson Orin Nano 2

Nvidia Corp. today announced the release of Jetson Orin Nano 2, a robotics computer “brain” for running artificial intelligence and frontier-level models at the edge. In the past months, foundational AI models have grown smaller and more efficient, adding numerous capabilities alongside language understanding, computer vision and audio processing. As more AI models compress in […] The post Nvidia doubles compute for entry-level edge robotics with Jetson Orin Nano 2 appeared first on SiliconANGLE.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Nvidia Corp. today announced the release of Jetson Orin Nano 2, a robotics computer “brain” for running artificial intelligence and frontier-level models at the edge. In the past…
In-site article

AI factories enter the execution era as Cisco and NVIDIA push rack-scale systems into production

The artificial intelligence infrastructure market is crossing an important threshold. The conversation is shifting from acquiring graphics processing units to building complete AI factories that can generate tokens reliably, efficiently and at scale. This is the next bottleneck. GPUs may be the engine, but an AI factory is a system. Compute, networking, storage, cooling, software […] The post AI factories enter the execution era as Cisco and NVIDIA push rack-scale systems into production appeared first on SiliconANGLE.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • The artificial intelligence infrastructure market is crossing an important threshold. The conversation is shifting from acquiring graphics processing units to building complete AI…
In-site article

A Go dependency wrote AGENTS.md mid-build and got Codex to hide the change

In a proof of concept published in mid-2026, NVIDIA's AI Red Team planted a malicious dependency in a Go project, let it execute during a normal build, and watched it write a new AGENTS.md file into the project director…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • In a proof of concept published in mid-2026, NVIDIA's AI Red Team planted a malicious dependency in a Go project, let it execute during a normal build, and watched it write a new…
In-site article

Nvidia and Cisco push the enterprise AI factory into the rack-scale era

The Cisco Secure AI Factory with Nvidia has been extended into the rack-scale era, offering enterprises greater full-stack operational capabilities as a result. Designed to give customers a framework for deploying artificial intelligence across their entire infrastructure, Secure AI Factory with Nvidia from Cisco Systems Inc. now integrates rack-to-fabric liquid cooling supporting systems beyond 200 […] The post Nvidia and Cisco push the enterprise AI factory into the rack-scale era appeared first on SiliconANGLE.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • The Cisco Secure AI Factory with Nvidia has been extended into the rack-scale era, offering enterprises greater full-stack operational capabilities as a result. Designed to give c…
In-site article

Cisco and Nvidia take AI factories from rack to runtime

AI factories are moving from ambitious plans toward production, but the path from graphics processing unit acquisition to usable systems remains a race against time. Neoclouds already have customers waiting for capacity, enterprises are looking to bring inference workloads closer to home and sovereign AI programs are being built now. Those distinct buyer motions are […] The post Cisco and Nvidia take AI factories from rack to runtime appeared first on SiliconANGLE.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • AI factories are moving from ambitious plans toward production, but the path from graphics processing unit acquisition to usable systems remains a race against time. Neoclouds alr…
In-site article

Robotics AI startup Generalist reportedly raises $200M

Generalist AI Inc., a startup that develops artificial intelligence software for robots, has reportedly raised $200 million in funding. Axios today cited a source as saying that 8VC led the investment. It was reportedly joined by a number of unnamed existing investors. Generalist’s previous $400 million round in June included the participation of Nvidia Corp., […] The post Robotics AI startup Generalist reportedly raises $200M appeared first on SiliconANGLE.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Generalist AI Inc., a startup that develops artificial intelligence software for robots, has reportedly raised $200 million in funding. Axios today cited a source as saying that 8…
In-site article

Nvidia reportedly eyes another investment in Perplexity AI at a $30B valuation

Nvidia Corp. is reportedly considering making another investment in the artificial intelligence search startup Perplexity AI Inc. A report by The Information says the chipmaker is holding talks with Perplexity over an investment that could push the startup’s valuation to more than $30 billion. That would represent a jump of more than 50% from the […] The post Nvidia reportedly eyes another investment in Perplexity AI at a $30B valuation appeared first on SiliconANGLE.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Nvidia Corp. is reportedly considering making another investment in the artificial intelligence search startup Perplexity AI Inc. A report by The Information says the chipmaker is…
In-site article

How XPUs Meet a World-Class AI Factory

To generate intelligence at scale, AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost per token, utilization and uptime. That requires AI infrastructure designed and built as a full factory, not a collection of individual accelerators. Hyperscalers and AI-native companies building custom XPUs must consider […]

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • To generate intelligence at scale, AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost per token, utilizat…
In-site article

With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents

The next era of AI inference won’t be defined by a single breakthrough chip, network or system. It’ll be defined by how every layer of the AI factory works together. That’s why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems. Announced today, the NVIDIA Vera Rubin rack-scale system NVIDIA Groq […]

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • The next era of AI inference won’t be defined by a single breakthrough chip, network or system. It’ll be defined by how every layer of the AI factory works together. That’s why NV…
In-site article

Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. Why? Consider what happens when an AI agent researches a company for an investment decision. The agent queries financial databases, searches news and filings, invokes a sub-agent to run peer comparisons and model valuations, then synthesizes everything into a […]

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. Why? Consider what happens when an AI agent researches a company for an inve…
In-site article

Best GPU Neoclouds 2026: CoreWeave, Nebius, Lambda, Crusoe, and Groq Ranked by Published Pricing and Contracted Power

The five largest GPU neoclouds now run on very different models. CoreWeave and Nebius report to the SEC; Lambda and Crusoe are private and heading toward IPOs; Groq rebuilt itself as an inference cloud after licensing its LPU technology to NVIDIA. This comparison checks each provider's live rate card, Q2 2026 financials, active and contracted gigawatts, anchor contracts, and SemiAnalysis ClusterMAX tier. Nebius posts the lowest H100 rate and the only published B300 price, Lambda has the cheapest B200, Crusoe is the only one with AMD on its card, and CoreWeave commands a 10–15% premium as the sole Platinum-rated provider. Figures verified August 21, 2026. The post Best GPU Neoclouds 2026: CoreWeave, Nebius, Lambda, Crusoe, and Groq Ranked by Published Pricing and Contracted Power appeared first on MarkTechPost.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • The five largest GPU neoclouds now run on very different models. CoreWeave and Nebius report to the SEC; Lambda and Crusoe are private and heading toward IPOs; Groq rebuilt itself…
In-site article

BF1: A Causal Dyadic Sparse-Attention Retrofit for Efficient Long-Context Transformers

arXiv:2608.20427v1 Announce Type: new Abstract: Dense causal attention remains expensive at long context even when implemented with highly optimized exact kernels. We study BF1, a deterministic block-aligned dyadic sparse-attention route that combines a small exact local neighborhood, a global first block, and logarithmically spaced historical blocks. The route is related to prior log-sparse and dilated attention patterns; our contribution is a correctness-gated pretrained-model retrofit, a matched topology-control study, and a systems characterization that connects per-layer sparsity to whole-model latency. For fixed block width, every converted layer uses O(n log n) selected token interactions and has O(log n) graph communication depth. On an NVIDIA RTX PRO 6000 Blackwell GPU, an optimized BF16 implementation crosses dense attention between 2K and 4K tokens and reaches a 10.91x per-layer prefill speedup at 32K. Retrofitting eight of 28 Qwen3-0.6B attention layers lowers warm whole-model time to first token by 7.7%, 11.3%, and 15.3% at 8K, 16K, and 32K, respectively, while the remaining dense layers keep the complete model asymptotically quadratic. Under a matched 1,000-step, 16.384M-token adaptation protocol, BF1 ranks first across three training seeds: mean report perplexity is 1.68639 versus 1.69154 for a matched static-random nonlocal graph, 1.69258 for dense continued training, and 1.81505 for equal-budget local sliding. At seed 1234, the packed-report paired interval places Dense-CT 0.3169-0.4055% above BF1 and static-random graph 17 0.2441-0.3642% above BF1. These results establish BF1 as a reproducible sparse operator and selective retrofit primitive with real long-context systems value. This paper evaluates numerical correctness, selected-interaction scaling, kernel performance, partial-model inference, and matched next-token language modeling.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • arXiv:2608.20427v1 Announce Type: new Abstract: Dense causal attention remains expensive at long context even when implemented with highly optimized exact kernels. We study BF1, a…
In-site article

Best GPU Neoclouds 2026: CoreWeave, Nebius, Lambda, Crusoe, and Groq Ranked by Published Pricing and Contracted Power

The five largest GPU neoclouds now run on very different models. CoreWeave and Nebius report to the SEC; Lambda and Crusoe are private and heading toward IPOs; Groq rebuilt itself as an inference cloud after licensing its LPU technology to NVIDIA. This comparison checks each provider's live rate card, Q2 2026 financials, active and contracted gigawatts, anchor contracts, and SemiAnalysis ClusterMAX tier. Nebius posts the lowest H100 rate and the only published B300 price, Lambda has the cheapest B200, Crusoe is the only one with AMD on its card, and CoreWeave commands a 10–15% premium as the sole Platinum-rated provider. Figures verified August 21, 2026. The post Best GPU Neoclouds 2026: CoreWeave, Nebius, Lambda, Crusoe, and Groq Ranked by Published Pricing and Contracted Power appeared first on MarkTechPost.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • The five largest GPU neoclouds now run on very different models. CoreWeave and Nebius report to the SEC; Lambda and Crusoe are private and heading toward IPOs; Groq rebuilt itself…
In-site article

Starcloud raises $250M to build AI data centers in orbit

Artificial intelligence hardware startup Starcloud Inc. today announced that it has raised $250 million in funding at a $2.3 billion valuation. Investment firm Manhattan West led the deal. It was joined by more than a dozen other backers including Nvidia Corp. and Cisco Investments. The cash infusion extends a Series A round that Starcloud announced in […] The post Starcloud raises $250M to build AI data centers in orbit appeared first on SiliconANGLE.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Artificial intelligence hardware startup Starcloud Inc. today announced that it has raised $250 million in funding at a $2.3 billion valuation. Investment firm Manhattan West led…
In-site article

Nvidia just showed that the harness, not the AI model, is now the real hero

Nvidia published some interesting new research on Friday suggesting it’s the harness, more than the underlying model, that is far more important when asking an AI to do long-horizon tasks. A harness is the software wrap…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Nvidia published some interesting new research on Friday suggesting it’s the harness, more than the underlying model, that is far more important when asking an AI to do long-horiz…
In-site article

Where Security Fits in an AI Agent Stack

As AI agents become more capable and operate over longer horizons, building security and trust into the applications they power becomes increasingly important. Drawing on work with NVIDIA OpenShell, agent developers, op…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • As AI agents become more capable and operate over longer horizons, building security and trust into the applications they power becomes increasingly important. Drawing on work wit…
In-site article

Nvidia's Switchyard router reshuffles AI models mid-task to cut task costs

Enterprises running always-on AI agents keep hitting the same tradeoff. Send every task to a frontier model and the bill climbs fast. Build custom routing logic to send easy tasks to cheaper models and that becomes its…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Enterprises running always-on AI agents keep hitting the same tradeoff. Send every task to a frontier model and the bill climbs fast. Build custom routing logic to send easy tasks…
In-site article

From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG

arXiv:2608.19535v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves language-model responses by grounding generation in external passages, which comes with overhead: retrieved context lengthens the prompt, increasing prefill work, KV-cache footprint, memory traffic, latency, and energy. Context compression offers a natural remedy by pruning retrieved text before generation. However, state-of-the-art context-compression methods are typically used with a fixed compression budget, or with the rate selected offline and then applied at inference time. This static view ignores both workload variation and the live state of the edge device. On an edge SoC, compression is not free: the compressor itself runs on the same SoC and consumes latency and energy that can offset any generation savings. This paper proposes a vision for telemetry-informed adaptive compression in edge RAG, grounded in experimental evidence. We characterize the compression tradeoff on the NVIDIA Jetson AGX Thor using Llama and Qwen generators, Natural Questions and HotpotQA datasets, and LLMLingua-2 compression. Our measurements show that generation dominates the RAG budget for larger models, reaching roughly 90% of per-query latency and 91% of GPU energy for 7B-8B generators. Exploring the impact of the compression rate reveals an adaptive operating region: mild compression can miss energy opportunities, and overly aggressive compression can hurt inference quality. Intermediate compression can reduce GPU energy by up to 53.2%, and SoC energy by up to 48.2%, with negligible quality loss. We argue for runtime policies that dynamically manage compression, guided by workload features and edge telemetry.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • arXiv:2608.19535v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves language-model responses by grounding generation in external passages, which comes wi…
In-site article

Nvidia’s SONIC Teaches Humanoids to Move

Nvidia’s model uses real-time human demonstrations and training data to give operators a one-stop shop for humanoid motion.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Nvidia’s model uses real-time human demonstrations and training data to give operators a one-stop shop for humanoid motion.
In-site article

Bring the Fire: Play Games on GeForce NOW With New Firefox Browser Support

It’s a new way into the cloud. GeForce NOW welcomes Firefox support to the cloud, opening up another way to jump into high-performance PC gaming straight from the browser, starting today. Whether on a school laptop or everyday PC, it’s now even easier to play supported PC games without downloading a dedicated app. Plus, discover […]

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • It’s a new way into the cloud. GeForce NOW welcomes Firefox support to the cloud, opening up another way to jump into high-performance PC gaming straight from the browser, startin…
In-site article

Blowing Off Steam: How Power-Flexible AI Factories Can Stabilize the Global Energy Grid

Editor’s note: This blog, originally published in March 2026, has been updated. At the half-time whistle of the UEFA EURO 2020 round of 16 football match between England and Germany, millions of viewers stepped away from their screens in the U.K. to do the same thing at the same time — turn on their kettles. […]

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Editor’s note: This blog, originally published in March 2026, has been updated. At the half-time whistle of the UEFA EURO 2020 round of 16 football match between England and Germa…
In-site article

Sanja Fidler’s world model startup Veeda AI raises $90M in seed funding

Veeda AI, a startup led by a team of former Nvidia Corp. researcher and renowned computer scientist Sanja Fidler, has taken its bow on the main stage after raising $90 million in a seed funding round today. The round, which was first reported by The Logic, was co-led by Khosla Ventures and Radical Ventures, is […] The post Sanja Fidler’s world model startup Veeda AI raises $90M in seed funding appeared first on SiliconANGLE.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Veeda AI, a startup led by a team of former Nvidia Corp. researcher and renowned computer scientist Sanja Fidler, has taken its bow on the main stage after raising $90 million in…
In-site article

Nvidia’s new financial strategy does not compute

“Compute is an asset class! Compute is an asset class!” I continue to insist as I slowly shrink down and turn into a corncob | Image: Cath Virginia / The Verge, Getty Images April - 1805 Napoleon is master of Europe Only the British fleet stands before him Compute is now an asset class I see it is once again time to talk financial innovation. Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR are all working with Nvidia to put together $500 billion in financing to turn compute into an asset class. "This is really the first time that technology chips have become an investable asset class," Nvidia CEO Jensen Huang said to CNBC. "These are revenue-generating assets now. They're productive, they're long-lived, they're fungible, they're flexible." "This is the very beginning, like what it was whe … Read the full story at The Verge.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • “Compute is an asset class! Compute is an asset class!” I continue to insist as I slowly shrink down and turn into a corncob | Image: Cath Virginia / The Verge, Getty Images April…
In-site article

KernelArc: A Multi-Agent Framework for GPU Kernel Optimization

arXiv:2608.17071v1 Announce Type: new Abstract: We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads. Strategy-specialized agents run in parallel and coordinate through conclusions-only shared memory, a deterministic benchmark guard, and read-only cross-agent state with plateau-triggered drafting. We evaluate \kernelarc{} on NVIDIA H100 and B200 GPUs using category-representative SOL-ExecBench workloads. The resulting implementations span custom BF16 GEMM, static cuBLASLt Expert-API configuration tables, fused mixture-of-experts backward, shape-gated decoder-layer fusion, native NVFP4 grouped-query attention, and paged prefill attention. At the public SOL-ExecBench leaderboard snapshot recorded on July~30, 2026, these submissions ranked first on representative L1, L2, Quantization, and FlashInfer tasks. The trajectories support the paper's central motivation: shared multi-agent search can broaden exploration and reach stronger incumbents within a fixed candidate budget, while the value of individual coordination features depends on the kernel and optimization stage.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • arXiv:2608.17071v1 Announce Type: new Abstract: We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads. Strategy-speci…
In-site article

NVIDIA Releases TensorRT Model Connect in Public Preview: Hugging Face Checkpoint to Native C++ Inference in Two Commands

NVIDIA has released TensorRT Model Connect (TRTMC) in public preview, an Apache-2.0 project that takes a supported Hugging Face or local checkpoint to end-to-end TensorRT inference in two commands, with no intermediate ONNX export. The build emits a versioned .bundle artifact that runs through native C++ task APIs, so inference executes without PyTorch in the runtime path. NVIDIA's July 29, 2026 GB300 snapshot covers 105 release profiles across 76 model families. The post NVIDIA Releases TensorRT Model Connect in Public Preview: Hugging Face Checkpoint to Native C++ Inference in Two Commands appeared first on MarkTechPost.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • NVIDIA has released TensorRT Model Connect (TRTMC) in public preview, an Apache-2.0 project that takes a supported Hugging Face or local checkpoint to end-to-end TensorRT inferenc…
In-site article

Nvidia to Back OpenAI Data Center With $105B Investment

The financing and infrastructure role puts the AI vendor at the center of another multibillion-dollar data center project.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • The financing and infrastructure role puts the AI vendor at the center of another multibillion-dollar data center project.
In-site article

ByteDance Seed and Tsinghua AIR Introduces CUDA Agent: A Large-Scale Agentic RL System for CUDA Kernel Generation

ByteDance Seed and Tsinghua AIR have released CUDA Agent, an agentic reinforcement learning system that trains a large language model to write GPU kernels that beat a compiler. The gap it targets is narrow but stubborn: frontier models already produce correct CUDA, they just produce slow CUDA. On KernelBench, the base model Seed1.6 passes 74.0% […] The post ByteDance Seed and Tsinghua AIR Introduces CUDA Agent: A Large-Scale Agentic RL System for CUDA Kernel Generation appeared first on MarkTechPost.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • ByteDance Seed and Tsinghua AIR have released CUDA Agent, an agentic reinforcement learning system that trains a large language model to write GPU kernels that beat a compiler. Th…
In-site article

How NVIDIA scales expertise with ChatGPT Work

NVIDIA teams use ChatGPT Work to reduce manual tasks, connect fast-moving signals, and scale successful workflows globally.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • NVIDIA teams use ChatGPT Work to reduce manual tasks, connect fast-moving signals, and scale successful workflows globally.
In-site article

AI cloud operator Groq raises $350M more in funding

Artificial intelligence startup Groq Inc. today announced that it has raised $350 million in funding. The Series A round was led by returning backer Disruptive. Grok stated that Nvidia Corp. plans to join the round later down the line, but didn’t specify how much the chip giant will invest. The cash infusion comes less than […] The post AI cloud operator Groq raises $350M more in funding appeared first on SiliconANGLE.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Artificial intelligence startup Groq Inc. today announced that it has raised $350 million in funding. The Series A round was led by returning backer Disruptive. Grok stated that N…
In-site article

NVIDIA Nemotron 3.5 Lightning now available in Amazon SageMaker JumpStart

NVIDIA Nemotron 3.5 Lightning, an open model built for high-volume agentic workloads, is now available in Amazon SageMaker JumpStart. This post shows how to deploy the 30B Mixture-of-Experts model (3B active), which delivers up to 4x higher throughput and up to 30% faster task completion for always-on agents.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • NVIDIA Nemotron 3.5 Lightning, an open model built for high-volume agentic workloads, is now available in Amazon SageMaker JumpStart. This post shows how to deploy the 30B Mixture…
In-site article

On theCUBE Pod: AI bubble debate heats up and neocloud earnings challenge doubters

The debate over whether we are in an artificial intelligence bubble took a new turn this week. Despite the ballooning AI spending, Dave Vellante (pictured, right), chief analyst for theCUBE Research, contends that any bursting point may be far off. Now that Nvidia Corp. Chief Executive Jensen Huang, has committed $500 billion to establish independent […] The post On theCUBE Pod: AI bubble debate heats up and neocloud earnings challenge doubters appeared first on SiliconANGLE.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • The debate over whether we are in an artificial intelligence bubble took a new turn this week. Despite the ballooning AI spending, Dave Vellante (pictured, right), chief analyst f…
In-site article

AI-enriched Linux 7.2 delivers cache-aware scheduling - here's everything new

The latest stable kernel also brings filesystem and I/O improvements, and substantial new support across AMD, Intel, Apple, Nvidia, USB4, and laptop hardware.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • The latest stable kernel also brings filesystem and I/O improvements, and substantial new support across AMD, Intel, Apple, Nvidia, USB4, and laptop hardware.
In-site article

Teaching Everyone to Fish for Tokens

Nvidia wants you building your own model, not buying from Anthropic/OpenAI.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Nvidia wants you building your own model, not buying from Anthropic/OpenAI.
In-site article

LG to Release Nvidia-Powered Humanoid in 2027

The robot is part of the companies’ expanded partnership to shift physical AI from concepts to real-world deployments.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • The robot is part of the companies’ expanded partnership to shift physical AI from concepts to real-world deployments.
In-site article

Securing the Infrastructure of Intelligence

AI factories are the defining infrastructure of the AI era—where compute transforms energy and data into intelligence that powers every business, industry and country. In the AI economy, compute is revenue. AI factories require a full stack of critical resources: advanced chips, packaging, memory, and networking – as well as land, power and shell. Just […]

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • AI factories are the defining infrastructure of the AI era—where compute transforms energy and data into intelligence that powers every business, industry and country. In the AI e…
In-site article

Nvidia AI Financing Is the $500B Risk Investors Aren't Watching

Nvidia AI Financing Is The $500 Billion Risk Investors Aren’t Watching ByJim Osman, Senior Contributor. Forbes contributors publish independent expert analyses and insights. Jim Osman is a finance expert with over 30 ye…

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Nvidia AI Financing Is The $500 Billion Risk Investors Aren’t Watching ByJim Osman, Senior Contributor. Forbes contributors publish independent expert analyses and insights. Jim O…
In-site article

Russian missile uses Nvidia AI chip to help target Ukraine

Ukrainian military intelligence found an Nvidia Jetson Orin module inside a downed Russian S-71 Monochrome cruise missile, suggesting the weapon may use AI. The chip was released after Nvidia exited Russia in 2022, indicating sanctions and export controls are being circumvented.

  • GUR recovered an Nvidia Jetson Orin NX 16GB module from a Russian S-71M cruise missile.
  • The module could enable autonomous target search and engagement, hinting at AI use.
In-site article

The CPU Comeback Is Upon Us

Agentic AI is driving a surge in CPU demand, as tool use, tokenization, and safety guardrails increasingly run on CPUs rather than GPUs. AWS has reportedly told engineers to conserve CPU cycles, Intel has sold out of server CPUs, and AMD, Arm, Qualcomm, and Nvidia are all pushing CPU-focused AI hardware.

  • Agentic AI workloads rely heavily on CPUs for tool calling, tokenization, and safety checks, creating a new bottleneck.
  • AWS has instructed engineers to minimize CPU cycle use as wait times for CPU server capacity explode.
In-site article

Did Nvidia’s Jensen Huang just make the AI buildout too big to fail?

Nvidia Corp. is no longer just selling technology. It is helping create a financial asset class around artificial intelligence compute. In our last Breaking Analysis, we argued that AI can be technologically transformative and still produce a capital bubble. Our thesis was simply that the bubble pops if deployable supply grows faster than monetizable demand – […] The post Did Nvidia’s Jensen Huang just make the AI buildout too big to fail? appeared first on SiliconANGLE.

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Nvidia Corp. is no longer just selling technology. It is helping create a financial asset class around artificial intelligence compute. In our last Breaking Analysis, we argued th…
In-site article

Show HN: I built a Claude Code plugin to query 10.6M earnings-call embeddings

FN2 is a Claude Code plugin that adds semantic search over 10.6 million earnings-call passages, plus real-time prices, SEC filings, and FRED macro data. Users can start, schedule, and read research agents from the terminal, and a demo shows how it summarized NVDA's Blackwell commentary across four quarters.

  • FN2 is a Claude Code plugin that enables semantic search across 10.6 million embedded earnings-call passages.
  • It also provides access to real-time prices, SEC filings, and FRED macro data, with agent scheduling and email delivery.
In-site article

Company Directory

NVIDIA AI News | AI News Hub