Skip to content
AI News HubLIVE
Original source2 min read

Expanding our enterprise inference capacity with IBM Cloud and NVIDIA

Summary

Together AI is partnering with IBM and NVIDIA to launch a dedicated large-scale NVIDIA B300 GPU inference cluster on IBM Cloud, using Spectrum-X Ethernet networking, to scale enterprise-grade open-model inference.

Expanding our enterprise inference capacity with IBM Cloud and NVIDIA
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

All blog posts

Company

Published 10/6/2026

Expanding our enterprise inference capacity with IBM Cloud and NVIDIA

Authors

Kai Mak

Table of contents

40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production...

We’re excited to share that Together AI is working with IBM and NVIDIA to scale enterprise-grade AI inference, starting with a large cluster of NVIDIA B300 GPUs on IBM Cloud backed by NVIDIA Spectrum-X Ethernet networking. It's the first dedicated, large-scale inference cluster of its kind on IBM Cloud, and we're the first customer running on it.

What's actually happening

The infrastructure: A dedicated NVIDIA B300 GPU cluster, purpose-built for inference, running on IBM Cloud.

The model: We operate the inference layer, IBM provides the cloud, NVIDIA delivers the silicon and networking.

The trajectory: We're planning for the future as token demand goes parabolic

Why it matters

We started Together AI because we believe the future of AI shouldn't be owned by a handful of closed labs. That bet is paying off faster than even we expected: Hundreds of trillions of tokens served per month to over a million developers, and demand keeps climbing.

Enterprises and AI-native companies choose open models for two simple reasons: their sovereign data stays theirs and they get frontier-level performance at a fraction of closed-model cost. The bigger the usage, the better that math looks.

With this collaboration:

NVIDIA brings the silicon — B300 GPUs and Spectrum-X Ethernet networking engineered for high-throughput inference.

IBM brings decades of running mission-critical infrastructure for the world's largest enterprises.

Together AI brings the inference platform — the fastest, most efficient way to run open models in production, hardened by some of the most demanding AI workloads on the planet.

Put those together and the result is simple: enterprise-grade inference at massive scale – more production grade tokens, with the reliability, security and guardrails enterprises have come to expect.

Open-source AI has to run everywhere, at scale, as fast and reliably as anything closed. We have spent the last few years making sure it can and this collaboration is a big step toward exactly that.

Related articles

View All

View All

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • Together AI is the first customer on IBM Cloud's first dedicated large-scale NVIDIA B300 inference cluster.
  • The stack combines IBM Cloud, NVIDIA B300 GPUs with Spectrum-X Ethernet networking, and Together AI's inference platform.
  • The goal is to meet surging token demand while giving enterprises data sovereignty and frontier-level performance at a fraction of closed-model cost.

Highlights and analysis are generated automatically and may contain errors. Check the original source.