We are thrilled to announce that Groq will be among the first adopters of NVIDIA Groq 3 LPX, boosting inference token generation for NVIDIA Vera Rubin NVL72 which it will deploy to its purpose-built AI inference cloud.…
Llama 3 Now Available to Developers via GroqChat and GroqCloud™ Here’s what’s happened in the last 36 hours: April 18th, Noon: Meta releases versions of its latest Large Language Model (LLM), Llama 3. April 19th, Midnig…
Jonathan Ross, CEO and Founder of Groq, and Tareq Amin, CEO of Aramco Digital, a subsidiary of Saudi Aramco, proudly announced their partnership to build the largest AI inference data center in Saudi Arabia. The data ce…
What is Context Length? As businesses look to leverage Large Language Models (LLMs) for conversational AI, generative AI, and analytics, a crucial factor often gets overlooked: context length, also known as context wind…
This blog is adapted from an original post by Groq CEO and Founder, Jonathan Ross. If Large Language Models (LLMs) are the printing press of the Generative AI age, then what’s next? What will be the AI equivalent of the…
The rapidly evolving field of Artificial Intelligence (AI) has led to significant advancements in Machine Learning (ML), with "inference" emerging as a crucial concept. But what exactly is inference, and how does it wor…
Today, one year this week since we launched GroqCloud, we're celebrating the one million developers who are now on GroqCloud! In just a year this community of builders, makers, and innovators have shipped incredible app…
Overview Groq LPU™ AI Inference Technology Groq builds fast AI inference. Groq® LPU™ AI inference technology delivers exceptional AI compute speed, quality, and affordability at scale. Groq AI inference infrastructure,…
GroqCloud™ provides fast inference for complex AI solutions that require instant responsiveness. But what happens when your use cases expand and require features beyond speed? That’s where Batch Processing comes in –now…
You know Groq runs small models. But did you know we run large models including MoE uniquely well? Here’s why. The Evolution of Advanced Openly-Available LLMs There’s no argument that Artificial intelligence (AI) has ex…
Appointment supports Groq’s global expansion and growing AI inference demand Mountain View, Calif. — September 22, 2025 — Groq, the global pioneer in AI inference, today announced the appointment of Simon Edwards as Chi…
GroqCloud announces beta availability of remote Model Context Protocol (MCP) server integration, enabling faster, lower-cost AI applications with seamless tool connectivity and zero-code migration from OpenAI.
Groq announces two key updates for its GPT-OSS models: price reductions and prompt caching, aimed at improving cost efficiency and speed for AI inference. New pricing is effective immediately and retroactive to October 2025 invoices. Prompt caching offers up to 50% discount on cached tokens, lower latency, and higher rate limits with zero configuration.
Based on practical experience, this guide explains how to reliably integrate open-source LLMs into products. The core is a four-step loop: Read (only necessary context), Constrain (clear system and formatting rules), Act (structured outputs, function calls, or plain text), Explain (show users steps and citations). It covers common patterns (router, extractor, translator, etc.), safe shipping (testing, monitoring, fallbacks), and common pitfalls. The goal is to build invisible, reliable AI features that users depend on daily.
GroqCloud announces day zero support for OpenAI's GPT-OSS-Safeguard-20B, a new open-source safety-classification model running at over 1000 t/s. Key features include bring your own policy, configurable reasoning effort, full reasoning trace, prompt caching, and 128k token context window. Pricing matches the base GPT-OSS-20B model.
Groq announces MCP Connectors in beta on GroqCloud, starting with Google Workspace. These pre-built, Groq-hosted MCP servers enable AI agents to interact with Gmail, Drive, and Calendar via the Responses API without managing your own MCP server.
Groq has been recognized as a Cool Vendor in the 2025 Gartner AI Infrastructure report, highlighting its LPU chip's deterministic, low-latency inference that scales linearly. Over 2.5 million developers use Groq for up to 5x faster and cheaper performance than GPUs.
The article discusses U.S. leadership in AI compute, especially inference, and proposes an export policy that balances market flexibility with consortium coordination to maintain strategic advantage.
GroqCloud is expanding its AI inference infrastructure globally to meet the growing demand from real-time applications moving from experimentation to production. A new UK data center, in partnership with Equinix, brings deterministic, high-performance inference closer to European developers and enterprises. GroqCloud now has over 3.5 million developers and sustained increases in production traffic.
Groq’s LPU is purpose-built for inference, achieving ultra-low latency without sacrificing accuracy through TruePoint numerics, SRAM-based memory, static scheduling, and tensor parallelism. Kimi K2 runs at 40x performance on Groq, demonstrating the architecture’s efficiency.