Skip to main content
GitHub

Rollback

Rolling back a fix deployment — built, and gated off.

Built, and gated off

Rollback is built, and it cannot be requested: a deployment cannot be created, so no new deployment exists to roll back, and the rollback request is not available with an API key during the beta. Automatic rollback runs in the deployment worker, which does not run in production. This page describes the design.

Rollback is designed to protect your system from bad fixes.

Automatic Rollback

As designed, fixes are automatically rolled back when:

TriggerThresholdSpeed
Error rate increase>10% relativeInstant
P99 latency increase>2x baselineInstant
A/B test failsp < 0.05 (treatment worse)Instant

Rollback Speed

Target: under 500ms — a design target that has never been measured.

As designed, an SDK process with the fix runtime on learns about a rollback only by polling GET /v1/fixes/active on the API host, every 60 seconds by default. There is no push path, so a process can keep applying the old fix until its next successful poll.

Manual Rollback

Via Dashboard

There is no working rollback control in the dashboard in this release. As designed:

  1. Navigate to Healing -> Deployments
  2. Find the deployment
  3. Click "Rollback"
  4. Confirm

Via API

Not available with an API key during the beta. As designed, a rollback is a DELETE request on the deployment, and it requires the owner or admin role.

Deployment API

Not available with an API key during the beta. As designed, four operations manage the full deployment lifecycle:

OperationDescription
ListList all deployments
GetGet deployment detail
CreateCreate a new deployment — gated off in this release
Roll backRoll back a deployment (owner or admin)

Deployment Management

All deployment state transitions (ramping, graduating) are handled automatically by the system based on statistical tests. There are no separate pause, resume, or graduate endpoints.

Deployment States

A deployment row holds one of these statuses:

StateDescription
pendingDeployment created, not yet started
canaryCreated at its initial traffic percentage but not yet routed; the worker's next check moves it to running
runningLive: the fix is routed to its share of traffic
rampingTraffic percentage increasing through stages
testingStatistical test in progress
completedRollout finished at 100%
rolled_backDeployment reverted
failedUnrecoverable error during deployment

Rollback Events

The event below illustrates the design. The API does not record rollback events today: a rollback's reason is not stored, and there is no rollback-events endpoint.

{
  "event": "rollback",
  "deployment_id": "deploy-abc123",
  "fix_id": "fix-xyz789",
  "timestamp": "2024-01-15T10:30:00Z",
  "trigger": "automatic",
  "reason": "error_rate_exceeded",
  "metrics": {
    "baseline_error_rate": 0.10,
    "treatment_error_rate": 0.15,
    "increase_percentage": 50
  },
  "duration_ms": 234
}

Rollback History

There is no rollback history view or endpoint in this release. As designed:

TimeFixTriggerReason
10:30fix-abcAutomaticError rate +50%
09:15fix-xyzManualCustomer report
Yesterdayfix-123AutomaticLatency 2.5x

Post-Rollback Analysis

As designed, after a rollback:

  1. Alert sent to team
  2. Diagnosis triggered on new errors (diagnosis is held for the beta)
  3. Fix marked as failed
  4. Learning recorded for future (a knowledge store exists but nothing uses it)

Preventing Bad Deployments

Canary First

All fixes go through canary (5%) before wider rollout.

Gradual Ramp

5% -> 25% -> 50% -> 100%

Each stage requires passing a statistical A/B test.

Guardrails

Secondary metrics must not degrade even if primary improves.

Rollback Configuration

There is no rollback configuration today: as built, the create request accepts only fix_id and initial_traffic_percentage, and a deployment cannot be created in this release. As designed, rollback thresholds would be customizable:

{
  "deployment_config": {
    "rollback_thresholds": {
      "error_rate_increase": 0.05,
      "latency_increase_factor": 1.25
    },
    "rollback_delay_seconds": 0,
    "require_manual_for_graduated": true
  }
}

Recovery After Rollback

To retry a rolled-back fix:

  1. Analyze failure reason
  2. Modify fix configuration
  3. Create new fix version
  4. Deploy from canary

Rolled-back fixes cannot be directly re-deployed.

Next Steps