How to train your own Jev for $17
We just launched our own Jev-like classifier, together/Tev1-4B-experimental, on top of Qwen3.5 4B on Together’s serverless platform. In this blog post we’ll show you how to fine-tune your own version!
Source profile
AI News Hub tracks Together AI Blog AI updates with visible source status, reuse boundaries, collection method, and published articles.
Official source; confirm reuse terms before enabling full body display.
We just launched our own Jev-like classifier, together/Tev1-4B-experimental, on top of Qwen3.5 4B on Together’s serverless platform. In this blog post we’ll show you how to fine-tune your own version!
A hard model swap exposes every user at once, and rolling back means cold-starting the old deployment under pressure. Here's how staged traffic ramps, metric gates, and automatic rollback work on dedicated inference.
Inside a global bank's shift to self-serve dedicated inference: how Together's DMI gave engineering teams direct control over scaling, models, and testing.
Moving from closed to open source models can take weeks, not years. A five-stage playbook: discover, evaluate, adapt, decide, and production.
Together Fine-Tuning adds the latest open-weight models, live experiment tracking, Expert LoRA, early stopping, tokenized dataset previews, pre-flight validation, and lower training prices on selected models.
Together GPU Clusters now supports preemptible compute: the same GPU capacity at a flat 50% of the on-demand rate, with a five-minute drain window.
We ported ThunderKittens to NVIDIA's Vera Rubin NVL72 and rebuilt our NVFP4 GEMM around the new hardware, taking it from 42% of roofline to over 22 PFLOPS — competitive with cuBLAS and CuTe DSL. Here is what changed in the ISA and how we used it.
All blog posts Inference Published 9/9/2026 The Open Source AI Stack Authors Hassan El Mghari Table of contents 40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production... A…
We ran 900 DeepSWE rollouts on GLM-5.3 and GLM-5.3 Flash. Flash gives up 5.6 points of pass@1 at 17x lower cost, and only 2.6 points at pass@4.
We ran 904 DeepSWE rollouts on GLM-5.3 and GPT-5.6 Sol. Sol leads pass@1 by 3.7 points; GLM-5.3 wins pass@4 at half the cost, and a GLM-first cascade hits 85.9%.
We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.
We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and GPT-5.6 Sol. Sol leads pass@1 by 10 points at 35x the cost; Pro wins pass@4, and a Pro-first cascade hits 83.0%.
We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.
Shadow traffic proves a candidate is operationally sound. It can't tell you if users like it better. Run the split at the endpoint instead of in your app code.
We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.
Kimi K3 is Moonshot AI's 2.8T-parameter open-weight model—the first in the 3-trillion-parameter class. Here's how it's architected, what it costs on Together AI, and how to call it with code examples for reasoning effort, streaming, vision, structured output, tools, and dynamic tool loading.
ThunderAgent is a program-aware scheduler for agentic inference. By treating each agent workflow as a schedulable program, it eliminates KV cache thrashing to deliver more than 2x single-node throughput and near-linear multi-node scaling.
Together AI and Moonshot AI form a strategic partnership, making Together AI the launch platform for Moonshot's open models, starting with Kimi K3, a 2.8-trillion-parameter sparse MoE model.
A detailed guide on Together AI's Dedicated Model Inference architecture: endpoints, deployments, configs, and capacity-aware routing.
We ran 904 DeepSWE rollouts on Kimi K3 and GPT-5.6 Sol. Sol leads pass@1; Kimi K3 wins pass@4 at 2.8x the solves per dollar, and routing between them reaches ~85.6%.
We ran 452 DeepSWE rollouts on Kimi K3 and Claude Fable 5. Fable leads pass@1 by 1.4 points; Kimi K3 wins pass@4 and delivers 2.8x the solves per dollar.
Together AI releases a major update to its inference platform, giving users control over performance, cost, and quality. New features include canary deployments, A/B testing, autoscaling, and a closed beta for custom training with reinforcement learning and fine-tuning.
Together AI and Y Combinator have partnered to provide YC startups with a dedicated GPU cluster, addressing the compute bottleneck. Startups can commit for weeks instead of two-year contracts, and manage GPUs directly through a self-service portal without YC involvement. The cluster is fully utilized and supports needs from single-node to large-scale expansion.
This article breaks down the real meaning behind reliability numbers like 99%, 99.9%, and 99.99% uptime for AI inference services, explaining the failure domains each tier must survive and the architectural requirements. The authors from Together AI share their experience building reliable inference infrastructure and provide key questions to ask any provider before committing.
See how Together AI is improving production GPU clusters with passive health checks, node repair, stronger Slurm reliability, OIDC, and startup scripts.
Thinking Machines Lab released Inkling, a multimodal mixture-of-experts model for token-efficient reasoning, native multimodal understanding, and broad task versatility. Together AI makes it available on its inference platform with support for controllable reasoning effort, text/image/audio inputs, and a 1M context window.
Provisioned Throughput offers reserved inference capacity for frontier open models with token-based pricing and a 99% uptime SLA, reducing costs by up to 90% compared to proprietary APIs.
Together AI raised $800M in Series C funding to accelerate the shift to open-source AI. The company argues that closed models' economics don't scale, and open models combined with full-stack optimization can achieve 6-20x cost reductions. Together AI has launched innovations like FlashAttention-4 and Together Megakernel, becoming one of the world's largest AI token producers.
Eight papers from Together AI accepted to ICML 2026, covering the full AI stack from agents to GPU kernels. These research works are integrated into the Together platform and already benefit production workloads.
ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves under a third, but a few generated kernels beat any public implementation.