ROLLBACK AUTOMATION

When a deployment goes wrong, recovery is improvised. Engineers reconstruct the previous state from memory and backups while the outage clock runs.

Actors

  • Deployment Engineer
  • NOC Engineer
  • Change Manager

Systems / Vendors

  • Ansible
  • Git
  • ServiceNow

Business Question

"The deployment failed at 2am and it took us four hours to get back to where we were at midnight — why isn't 'undo' a button?"

What SPoG Does

  • Captures the last known good state automatically before every change.
  • Reverts failed deployments to that state in minutes, on demand or on health-check failure.
  • Logs every rollback with cause and evidence, feeding release quality improvement.

Outcome Metrics

−90%

Recovery time after failed change

−50%

Change-related downtime

2–4 wks

To first protected deployment