跳到主要內容
AI News HubLIVE
站內改寫6 分鐘閱讀

待翻譯:Canary rollouts: upgrade models in production without downtime

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:A hard model swap exposes every user at once, and rolling back means cold-starting the old deployment under pressure. Here's how staged traffic ramps, metric gates, and automatic rollback work on dedicated inference.

待翻譯:Canary rollouts: upgrade models in production without downtime
回報錯誤

更正管道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

All blog posts Inference Published 9/22/2026 Canary rollouts: upgrade models in production without downtime Staged traffic ramps, metric gates, and automatic rollback on dedicated inference. Table of contents 40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production... Links in this article Dedicated Model Inference 101 Autoscaling A/B Testing 📚 Docs: ‍Start a rollout‍ ‍Gate rollouts with metrics‍ ‍CLI reference ‍API reference: Create a rollout ‍ Summary Rollouts move live traffic from your current model to a new checkpoint in gated steps. Health checks always run before traffic moves; on a canary you can also add metric gates (say p95 latency or error rate) that run after each step. If a gate trips, the rollout pauses at the canary share and you cancel it and run it in reverse. Below we run one for real: a Qwen2.5-7B → Qwen3.5-9B canary whose gate caught a 137% p95 regression at 10% of traffic; we canceled and reversed it with live requests served and none failed. Swapping models is common practice If you run a model in production, you already know the need to swap in a new checkpoint or a new model family: the open model ecosystem moves fast, and the candidate usually looks great in evals or promises better throughput. You want it in front of every user without hiccups, and a way to rollback if it disappoints. The usual options force a tradeoff: Hard swap: point the endpoint at the new model and every user is exposed at once. If p95 doubles, you find out from your dashboard or, worse, from your customers, and you roll back under pressure onto a cold-started old model. DIY staged cutover: a second deployment plus a script that nudges traffic percentages while you watch Grafana, remembering to scale the old deployment back up before you shift traffic back. Both of these options put a human in the loop as the safety mechanism. Rollouts move that mechanism into the platform: you describe the source, the target, the steps, and what "healthy" means, and the platform works against this plan at every stage. How rollouts work A rollout migrates traffic between two deployments on the same endpoint: a source (what's serving today) and a target (what you want to serve tomorrow). You pick one of three strategies: Canary: traffic moves in staged percentages you define (say 10% → 50% → 100%; the default ladder is 5% → 25% → 50% → 100%), with a wait period and optional metric checks between steps. Blue-green: one gated 0% → 100% cutover. You can think of this as a single-step canary. Rolling: an in-place, replica-by-replica swap that preserves total capacity. Best when capacity is constrained, especially for same-model config changes that don't need a traffic ramp. Here's what happens inside every canary step: We chose this ordering deliberately; each item prevents a class of incidents: The target scales up before any traffic moves. No capacity for the new deployment means no redirected requests: the rollout parks first. The health gate runs before the traffic shifts. Traffic only reaches replicas whose engine is loaded and answering, not merely started. A propagation wait sits between the shift and the drain. Routing caches converge before any source capacity is removed. The source drains after traffic has moved. Capacity leads traffic on the way up; traffic leads capacity on the way down. The wait period and the metric gate come before the step is recorded as complete. A step that regressed is never marked passed. Through the API or the console, a rollout is created in a PENDING state and does nothing until you explicitly start it (the CLI's rollout command creates and starts in one step). This two-step create/start is intentional because you can create the rollout, review it (or have a teammate review it), and start it when you're actually watching. Two states in the diagram above deserve a note: PAUSED means you pressed pause. The rollout holds exactly where it is and resumes from the same step. SYSTEM_PAUSED means the platform found something that went wrong, such as a failed metric gate, a capacity shortfall or missing metrics, and stopped to wait for human approval. It pauses, notifies you and waits; canceling is always your call. There is no FAILED end state that leaves traffic in limbo: a rollout ends COMPLETED (the target serves) or CANCELED (the split is frozen where it was, and you run the rollout in reverse to go back). Anatomy of a step The following is a breakdown of what happens in a single canary step, measured on the run at the end of this post (Qwen2.5-7B → Qwen3.5-9B on one H100 each). The propagation wait is what keeps stale global routing caches from sending requests to a shrinking source. The wait period is grown to the metric window plus ingestion lag. The cold start dominates the first step; later steps add replicas to a target that is already serving and warm. Choosing a strategy at a glance All three strategies run through the same engine and the same health gates; they differ in how traffic moves, how much extra capacity the overlap costs, and whether there is a wait window for a metric gate. CanaryBlue-greenRolling Traffic patternSteps through shares you choose (default 5% → 25% → 50% → 100%), each held for a wait windowOne cutover, 0% → 100%, once the target is healthyReplica by replica, traffic following the replica ratio Extra capacityNear source size; the target grows one step before the source drains that shareBoth deployments at full size until the source drainsSource's replica count at each step; one extra replica mid-step Typical durationOne cold start plus a wait per step (at least 390 s each with a metric gate)One cold start plus 30 s propagation; a few minutesOne cold start per replica; slowest on large deployments Metric gatesYes, after every stepNo (no wait window)No How to go backCancel freezes the current share, then run the rollout in reverseRun the rollout in reverse; --final-source-replicas 1 keeps the old model warm for an instant returnRun the rollout in reverse Best forMeasuring on live traffic before taking 100%The fastest switch, when you can briefly afford double capacitySame-model engine or config changes at a constant GPU footprint ‍ Creating a rollout Here's a three-step canary from a deployment serving your current model to one serving the candidate, with a latency regression gate. The CLI ships as tg in the together Python package (2.34.0 or newer). You pass the target deployment; the source is inferred when exactly one deployment is receiving traffic, otherwise pass --source: # 1. Create AND start the rollout in one command. # Intervals and windows are seconds with an "s" suffix ("600s", not "10m"). tg beta endpoints rollout $TARGET_DEPLOYMENT_ID \ --source $SOURCE_DEPLOYMENT_ID \ --canary \ --steps 10,50,100 \ --interval 600s \ --metric router_latency --metric-stat p95 \ --metric-max-regression 10 --metric-direction higher-is-worse \ --metric-window 300s # 2. Watch it move: pass the rollout ID printed under "Active Rollout", # or the endpoint ID for the endpoint summary tg beta endpoints get $ROLLOUT_ID tg beta endpoints get $ENDPOINT_ID # 3. Control it: pass the endpoint ID plus exactly one control flag tg beta endpoints rollout $ENDPOINT_ID --pause --reason "holding for review" tg beta endpoints rollout $ENDPOINT_ID --resume tg beta endpoints rollout $ENDPOINT_ID --promote tg beta endpoints rollout $ENDPOINT_ID --cancel --reason "latency regression on target" The source drains to zero replicas and stops when the rollout completes (--final-source-replicas defaults to 0), and the target lands with the source's replica count as its floor (--final-target-replicas). The CLI attaches one metric gate per rollout; for several rules use the console or the API. The same via the REST API, where create and start are separate calls and a rollout can carry several metric rules: # 1. Create the rollout. It comes back in state PENDING; save its "id" (rol_...) as $ROLLOUT_ID. curl -s -X POST \ "https://api.together.ai/v2/projects/$PROJECT_ID/endpoints/$ENDPOINT_ID/rollouts" \ -H "Authorization: Bearer $TOGETHER_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "sourceDeploymentId": "'$SOURCE_DEPLOYMENT_ID'", "targetDeploymentId": "'$TARGET_DEPLOYMENT_ID'", "canary": { "steps": [{"traffic": 10}, {"traffic": 50}, {"traffic": 100}], "stepInterval": "600s" }, "metrics": [{ "name": "router_latency", "stat": "METRIC_STAT_TYPE_PERCENTILE", "percentile": 95, "regressionCheck": { "direction": "REGRESSION_DIRECTION_HIGHER_IS_WORSE", "maxRegressionPercent": 10 }, "window": "300s" }] }' # 2. Start it: POST .../rollouts/$ROLLOUT_ID/start -d '{}' # 3. Watch it: GET .../rollouts/$ROLLOUT_ID A few things the API is strict about: percentile is an integer (95, not "p95"), enum values carry their full prefix (METRIC_STAT_TYPE_*, REGRESSION_DIRECTION_*, THRESHOLD_OPERATOR_*), durations are protobuf strings like "600s", and a metric name outside the catalog is rejected with a 400 that lists the supported names. Draining the source is the default, so there is nothing to pass for it. The regression check can be understood as: at each gate, compare the target's p95 router latency (the per-request duration measured at the router, in milliseconds) over the last 5 minutes against the source's. If the target is more than 10% worse, don't proceed. Python SDK The same rollout from Python, with the together package (2.34.0 or newer). Field names are snake_case here and camelCase on the wire; the SDK translates. from together import Together client = Together() # reads TOGETHER_API_KEY rollout = client.beta.endpoints.rollouts.create( endpoint_id=ENDPOINT_ID, project_id=PROJECT_ID, source_deployment_id=SOURCE_DEPLOYMENT_ID, target_deployment_id=TARGET_DEPLOYMENT_ID, canary={"steps": [{"traffic": 10}, {"traffic": 50}, {"traffic": 100}], "step_interval": "600s"}, metrics=[{ "name": "router_latency", "stat": "METRIC_STAT_TYPE_PERCENTILE", "percentile": 95, "regression_check": {"direction": "REGRESSION_DIRECTION_HIGHER_IS_WORSE", "max_regression_percent": 10}, "window": "300s", }], ) # state: PENDING client.beta.endpoints.rollouts.start(rollout.id, project_id=PROJECT_ID, endpoint_id=ENDPOINT_ID) # later: .retrieve / .pause / .resume / .promote / .cancel take the same (id, project_id, endpoint_id) Controlling a rollout Every rollout accepts the same four controls. An endpoint has at most one active rollout, so the CLI takes the endpoint ID and you rarely need the rollout ID. Each control returns as soon as it is accepted; poll tg beta endpoints get (or the GET endpoint) until the rollout reaches the state you expect. While a rollout is active, including while paused, the endpoint's traffic split is locked and its source and target cannot be stopped or deleted. Pause The rollout goes PAUSING, lets any step activity in flight finish, then holds at the current traffic split and replica counts as PAUSED. Both deployments keep serving. A pause can last for days; the platform never auto-resumes an operator pause. CLI tg beta endpoints rollout $ENDPOINT_ID --pause --reason "holding for review" REST POST …/rollouts/$ROLLOUT_ID/pause {"reason": "holding for review"} Resume Continues from the same step, for both PAUSED and SYSTEM_PAUSED. If a gate tripped, it re-evaluates against fresh data; the step is not skipped. CLI tg beta endpoints rollout $ENDPOINT_ID --resume REST POST …/rollouts/$ROLLOUT_ID/resume {} Promote Skips the remaining canary steps and runs the final 100% step in full: the target scales to its landing size, traffic shifts, the propagation wait and soak run, then the source drains. Skipped steps are recorded as SKIPPED. Not instantaneous: in a test run with a 10-minute step interval, a promote at step 0 still sat through the final step's full soak. CLI tg beta endpoints rollout $ENDPOINT_ID --promote [truncated for AI cost control]

展開要點與分析

文章情報

工程師進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級後設資料。
  • A hard model swap exposes every user at once, and rolling back means cold-starting the old deployment under pressure. Here's how staged traffic ramps, metric gates, and automatic…

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。