--interval in the CLI, or Step interval in the console) during which traffic stays at the current split level so the metric can settle and accumulate samples.
At the end of the window, the gate evaluates its rules against the target deployment and either advances the rollout or pauses it for review.
Configure a metric gate
Metric gates are canary-only. Blue-green and rolling strategies have no soak window, so a gate is rejected at create time.- CLI
- Console
Configure a gate by passing
--metric* flags to the rollout dep_target456 --canary command. Pass a supported metric plus either a threshold check or a regression check. For router_latency, set --metric-stat to avg or a percentile. For router_error_rate and inflight_requests, omit --metric-stat (the platform records it as AVG).--metric-stat picks the aggregation the gate applies over the lookback window: a percentile (p50, p90, p95, or p99) or avg, the mean over the window. It is required for router_latency and optional for the other metrics, which only aggregate one way.See the CLI reference for the full flag list.Threshold vs. regression checks
Each gate uses one of the following checks to evaluate the target at the end of each soak window:- Threshold check: Compares the target against a fixed SLO, set with
--metric-thresholdand--metric-operator(gt,gte,lt, orlte). For example, “p95 router latency must stay under 30 seconds”. - Regression check: Compares the target against the source, set with
--metric-max-regression(a percent) and--metric-direction(higher-is-worseorlower-is-worse). It fails when the target is worse than the source by more than the allowed amount. For example, “the new deployment must not be more than 50% slower than the old one”.
Supported metrics
The following metrics are supported (creating a rollout with any other metric name is rejected with a400 error). All of them are measured at the router.
For regression checks, higher values are worse for all three metrics.
Choose a lookback window
--metric-window (CLI) or Window (s) (console) sets the lookback period over which the gate samples the metric. The minimum is 60s. When omitted, it defaults to 300s (five minutes).
Use at least 300s for router_latency percentile gates. A short window such as 60s falls inside the metric pipeline’s ingestion lag (about 90 seconds) and will likely hold too few samples for a trustworthy percentile, causing the rollout to pause unnecessarily.
You don’t need to size the step interval to fit the window: at create time, the platform grows each step’s soak window to at least the metric window plus the ingestion lag, so the metric always has time to settle and accumulate samples before the target is evaluated.
Gate on more than one rule
The CLI configures a single-rule gate. To gate on several rules at once, add multiple gates in the console’s New rollout dialog, or create the rollout through the API and then start it manually. A step advances only when every rule passes. If any rule regresses or its metric is unavailable, the rollout pauses instead. You can attach both a threshold check and a regression check on the same metric. When you retrieve the rollout, each configured rule appears as its own row instatus.steps[].metrics, with its criteria and its own verdict.
How gate results appear on retrieve
Each metric row instatus.steps[].metrics or status.condition.metrics is a MetricResult object containing the rule’s name, check form, criteria, optional observed values, and a verdict (PASS, BREACHED, or UNAVAILABLE).
- Measured rows:
targetValueis set whenever the gate recorded an observation, including a legitimate0.sourceValueis set only on regression-check rows, because threshold checks never consult the source. - Unmeasured rules: When the gate couldn’t measure a rule (for example, no traffic reached the target), the rule still appears as a row with verdict
UNAVAILABLEand its criteria, but nosourceValueortargetValue. - Completed rollouts: A
COMPLETEDrollout never marks a step asFAILED. Finished steps arePASSED, and steps a promote jumped over areSKIPPED. For the full per-step state set, see Step status on retrieve.
What to do when a metric gate fails
If a metric gate fails (the target deployment doesn’t meet the gate’s threshold or regression criteria), the rollout pauses automatically. The traffic split that’s in place when the gate fails is maintained while the rollout is paused. Review the gate verdicts for each completed step by retrieving the rollout, or open the Rollout details drawer from the progress card or history table in the console. The pause reason (shown on the paused rollout and on itsrollout.system_paused event) starts with one of two prefixes:
metrics-regression:The gate confirmed the target is worse than the gate’s criteria allow. The rest of the reason names the failing rule and the observed values. Investigate why the target regressed before acting. If the target itself is at fault, cancel the rollout, then run it in reverse to move traffic back to the source. If the cause was external (like a traffic surge), resume instead.metrics-unavailable:The gate could not get a trustworthy reading, so it paused rather than advancing on incomplete data. Common causes are gaps in the metrics pipeline (the most common), too few samples in the lookback window, and zero ready replicas on the target or source. Most of these are transient—the platform confirms the reading before pausing and retries automatically. Resume the rollout manually if the pause persists. The paused step’s metrics include anUNAVAILABLErow for each rule the gate could not measure.
Next steps
Start a rollout
Create and run the canary rollout a gate protects.
Observability
Watch the same metrics on your endpoint’s dashboards.