Skip to content
AI News HubLIVE
Public articles 50Collected articles 51Trust 88Refresh 5 min
Health HealthySource type OfficialFull-text rights Official full textLast ingested 2026-09-23ID together-ai-blogStatus Enabled

Official source; confirm reuse terms before enabling full body display.

Latest public articles

How to train your own Jev for $17

We just launched our own Jev-like classifier, together/Tev1-4B-experimental, on top of Qwen3.5 4B on Together’s serverless platform. In this blog post we’ll show you how to fine-tune your own version!

Together AI BlogIn-site articleHow to train your own Jev for $17

Canary rollouts: upgrade models in production without downtime

A hard model swap exposes every user at once, and rolling back means cold-starting the old deployment under pressure. Here's how staged traffic ramps, metric gates, and automatic rollback work on dedicated inference.

Together AI BlogIn-site articleCanary rollouts: upgrade models in production without downtime

Migrating from closed to open source models, Together

Moving from closed to open source models can take weeks, not years. A five-stage playbook: discover, evaluate, adapt, decide, and production.

Together AI BlogIn-site articleMigrating from closed to open source models, Together

Introducing preemptible compute: the same compute, half the price

Together GPU Clusters now supports preemptible compute: the same GPU capacity at a flat 50% of the on-demand rate, with a five-minute drain window.

Together AI BlogIn-site articleIntroducing preemptible compute: the same compute, half the price

To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

We ported ThunderKittens to NVIDIA's Vera Rubin NVL72 and rebuilt our NVFP4 GEMM around the new hardware, taking it from 42% of roofline to over 22 PFLOPS — competitive with cuBLAS and CuTe DSL. Here is what changed in the ISA and how we used it.

Together AI BlogIn-site articleTo Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

The Open Source AI Stack

All blog posts Inference Published 9/9/2026 The Open Source AI Stack Authors Hassan El Mghari Table of contents 40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production... A…

Together AI BlogIn-site articleThe Open Source AI Stack

GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing

We ran 900 DeepSWE rollouts on GLM-5.3 and GLM-5.3 Flash. Flash gives up 5.6 points of pass@1 at 17x lower cost, and only 2.6 points at pass@4.

Together AI BlogIn-site articleGLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing

GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

We ran 904 DeepSWE rollouts on GLM-5.3 and GPT-5.6 Sol. Sol leads pass@1 by 3.7 points; GLM-5.3 wins pass@4 at half the cost, and a GLM-first cascade hits 85.9%.

Together AI BlogIn-site articleGLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.

Together AI BlogIn-site articleGLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and GPT-5.6 Sol. Sol leads pass@1 by 10 points at 35x the cost; Pro wins pass@4, and a Pro-first cascade hits 83.0%.

Together AI BlogIn-site articleDeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

A/B test models in production

Shadow traffic proves a candidate is operationally sound. It can't tell you if users like it better. Run the split at the endpoint instead of in your app code.

Together AI BlogIn-site articleA/B test models in production

DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding

We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.

Together AI BlogIn-site articleDeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding

Kimi K3: The Complete Developer Guide

Kimi K3 is Moonshot AI's 2.8T-parameter open-weight model—the first in the 3-trillion-parameter class. Here's how it's architected, what it costs on Together AI, and how to call it with code examples for reasoning effort, streaming, vision, structured output, tools, and dynamic tool loading.

Together AI BlogIn-site articleKimi K3: The Complete Developer Guide

ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale

ThunderAgent is a program-aware scheduler for agentic inference. By treating each agent workflow as a schedulable program, it eliminates KV cache thrashing to deliver more than 2x single-node throughput and near-linear multi-node scaling.

Together AI BlogIn-site articleThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale

Configuring Dedicated Model Inference

A detailed guide on Together AI's Dedicated Model Inference architecture: endpoints, deployments, configs, and capacity-aware routing.

Together AI BlogIn-site articleConfiguring Dedicated Model Inference

Kimi K3 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

We ran 904 DeepSWE rollouts on Kimi K3 and GPT-5.6 Sol. Sol leads pass@1; Kimi K3 wins pass@4 at 2.8x the solves per dollar, and routing between them reaches ~85.6%.

Together AI BlogIn-site articleKimi K3 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

Kimi K3 vs Claude Fable 5 on DeepSWE: Cost and Coding

We ran 452 DeepSWE rollouts on Kimi K3 and Claude Fable 5. Fable leads pass@1 by 1.4 points; Kimi K3 wins pass@4 and delivers 2.8x the solves per dollar.

Together AI BlogIn-site articleKimi K3 vs Claude Fable 5 on DeepSWE: Cost and Coding

The production platform for open-weight AI inference

Together AI releases a major update to its inference platform, giving users control over performance, cost, and quality. New features include canary deployments, A/B testing, autoscaling, and a closed beta for custom training with reinforcement learning and fine-tuning.

Together AI BlogIn-site articleThe production platform for open-weight AI inference

Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community

Together AI and Y Combinator have partnered to provide YC startups with a dedicated GPU cluster, addressing the compute bottleneck. Startups can commit for weeks instead of two-year contracts, and manage GPUs directly through a self-service portal without YC involvement. The cluster is fully utilized and supports needs from single-node to large-scale expansion.

Together AI BlogIn-site articleTogether AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community

What does 99.9% uptime mean for inference?

This article breaks down the real meaning behind reliability numbers like 99%, 99.9%, and 99.99% uptime for AI inference services, explaining the failure domains each tier must survive and the architectural requirements. The authors from Together AI share their experience building reliable inference infrastructure and provide key questions to ask any provider before committing.

Together AI BlogIn-site articleWhat does 99.9% uptime mean for inference?

Together AI brings Thinking Machines Lab’s new model Inkling on day 0

Thinking Machines Lab released Inkling, a multimodal mixture-of-experts model for token-efficient reasoning, native multimodal understanding, and broad task versatility. Together AI makes it available on its inference platform with support for controllable reasoning effort, text/image/audio inputs, and a 1M context window.

Together AI BlogIn-site articleTogether AI brings Thinking Machines Lab’s new model Inkling on day 0

Open, convenient and predictable: Introducing Provisioned Throughput

Provisioned Throughput offers reserved inference capacity for frontier open models with token-based pricing and a 99% uptime SLA, reducing costs by up to 90% compared to proprietary APIs.

Together AI BlogIn-site articleOpen, convenient and predictable: Introducing Provisioned Throughput

Announcing our $800M Series C to accelerate the shift to open-source AI

Together AI raised $800M in Series C funding to accelerate the shift to open-source AI. The company argues that closed models' economics don't scale, and open models combined with full-stack optimization can achieve 6-20x cost reductions. Together AI has launched innovations like FlashAttention-4 and Together Megakernel, becoming one of the world's largest AI token producers.

Together AI BlogIn-site articleAnnouncing our $800M Series C to accelerate the shift to open-source AI

Together AI at ICML 2026: Frontier Research Across the Full Stack

Eight papers from Together AI accepted to ICML 2026, covering the full AI stack from agents to GPU kernels. These research works are integrated into the Together platform and already benefit production workloads.

Together AI BlogIn-site articleTogether AI at ICML 2026: Frontier Research Across the Full Stack

ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)

ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves under a third, but a few generated kernels beat any public implementation.

Together AI BlogIn-site articleParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)

All sources