A rollback is not an undo. It is a second deploy, shipped under pressure, into a system the first deploy has already changed. The code goes back. The rows it wrote, the orders it sent and the machines it crashed stay where they are.
List every place a release writes, and mark which ones reverting the deploy actually undoes; usually it is 1 of 7. Commit reversible writes first and irreversible ones last, make every step idempotent, change schema by expand, migrate, contract, and never reuse a flag name. Knight Capital's rollback moved 7 healthy servers into the broken state.









Paulo Antunes published a short, honest postmortem this week called "The rollback that only rolled back half of it". He was testing a new screen against a factory's production ERP database: two writes, one transaction, and a ROLLBACK at the end "so the rehearsal wouldn't leave a trace." One table was InnoDB. The other dated from the MyISAM era. His next line is the whole article: "The rollback ran. No error. And half of it stayed."
It is a small bug in an old schema. It is also the most accurate description of rollback I have read in years, because it applies at every scale above a single table. Knight Capital lost over $460 million in 2012 partly because a rollback made things worse. CrowdStrike reverted a bad update 78 minutes after it started shipping in 2024, and the revert could not reach the machines already stuck in a boot loop. The mechanism is the same each time. Time only runs one way, and every write your system makes runs with it.
The rollback that reported success
MyISAM has no transactions. Antunes puts it plainly: BEGIN does nothing to a MyISAM table, COMMIT does nothing, and ROLLBACK does nothing either. The update lands the moment it runs. I reproduced his case on a stock MySQL 8.4 container to see exactly what the server says when this happens.
mysql> START TRANSACTION; mysql> UPDATE items SET code = NULL WHERE id = 1; -- InnoDB mysql> UPDATE volume SET qty = 0 WHERE id = 1; -- MyISAM mysql> ROLLBACK; Query OK, 0 rows affected, 1 warning (0.00 sec) mysql> SHOW WARNINGS; | Warning | 1196 | Some non-transactional changed tables couldn't be rolled back | mysql> SELECT code, qty FROM items JOIN volume USING (id); | A1 | 0 | ^ one table rolled back. the other kept the write.
So the server does tell you, strictly speaking. The MySQL manual documents it: roll back a transaction that touched a nontransactional table and "an ER_WARNING_NOT_COMPLETE_ROLLBACK warning occurs." But it arrives as a warning count on a statement that returned OK. Most application code checks for an exception, never gets one, and moves on. From the program's side, the rollback succeeded. That is the dangerous version of every failure in this article: the undo reports success and the state disagrees.
The usual advice is to turn warnings into errors, and Oracle's own Connector/Python documentation suggests the sql_mode setting for it. I tried both. Stock MySQL 8.4 already runs with STRICT_TRANS_TABLES, and adding STRICT_ALL_TABLES changed nothing: same warning 1196, same half-kept write. Connector/Python 26.7.0 with raise_on_warnings=True did raise error 1196 when I sent ROLLBACK as a statement through a cursor. The connection's own rollback() method, the call most code actually makes, returned normally. And in both cases the MyISAM write had already stayed. The exception only tells you afterwards.
Knight Capital rolled back into the fire
Antunes caught his half-rollback on one screen, in a rehearsal he ran himself. The same mismatch, between what the code says and what the state holds, went to the open market in 2012.
On August 1, 2012, Knight Capital's order router, SMARS, took 212 small retail orders and turned them into over 4 million executions in 154 stocks, more than 397 million shares, in about 45 minutes. The SEC's order is the most useful deployment document I know, because it reads like a postmortem written by someone with subpoena power.
The setup took years. Knight stopped using an old function called Power Peg in 2003 but never removed it. In 2012 the new Retail Liquidity Program code "repurposed a flag that was formerly used to activate the Power Peg code." The plan was to delete Power Peg so that the flag would now mean RLP. The new code went out to SMARS in stages, and "one of Knight's technicians did not copy the new code to one of the eight SMARS computer servers." Nobody checked. The SEC notes that Knight "did not have a second technician review this deployment" and "had no written procedures that required such a review."
So on the morning of August 1, seven servers read the flag as RLP and one read it as Power Peg. That was the deploy. Then came the rollback. While staff hunted for the cause, "Knight uninstalled the new RLP code from the seven servers where it had been deployed correctly. This action worsened the problem, causing additional incoming parent orders to activate the Power Peg code that was present on those servers."
Read that twice. The rollback reverted the code and left the flag. Incoming orders still carried the repurposed flag, the old code on the seven servers read it the old way, and those servers started doing what the eighth had been doing all morning. The revert was correct about the code and wrong about the system. It moved seven healthy servers into the broken state.
deploy the incident the "rollback" code, 7 srv new RLP new RLP old, Power Peg code, 8th old, Power Peg old, Power Peg old, Power Peg the flag set, means RLP set, means RLP set, means RLP --- --- --- routing 7 right, 1 wrong 8 wrong ^ the code moved back a version. the flag never moved at all.
Knight ended the morning net long about $3.5 billion across 80 stocks and net short about $3.15 billion across 74. It paid a $12 million civil penalty on top of the trading loss. The penalty was not for the bug. It was for having no controls that would have caught an eight-server deploy that reached seven. That is the same trap as a feature flag nobody retired: a flag is state, and state outlives the code that reads it.
What a revert can reach
Every change you ship writes to more than one place. The deploy tool knows about one of them. Here is the list I walk through before I believe any rollback plan, with an honest answer for each row.
| What changed | Does reverting the deploy undo it? |
|---|---|
| Application code | Yes. This is the one thing rollback was built for. |
| Config and feature flags | Only if they ship in the same artifact. Usually they live elsewhere and stay set. |
| Schema | Only with a written down-migration, and in MySQL a DDL statement commits the open transaction when it runs. |
| Rows written under the new code | No. They stay, in whatever shape the new code wrote them. |
| Messages already published | No. Consumers have them, and some have acted. |
| Calls to other companies' systems | No. Orders filled, payments captured and emails delivered belong to somebody else now. |
| Hosts the release crashed | No. A machine that cannot boot cannot download the fix. |
That last row is the CrowdStrike lesson. Its preliminary post-incident review says the faulty content update reached Windows hosts online between 04:09 and 05:27 UTC on July 19, 2024, and that "the defect in the content update was reverted" at 05:27. Hosts that came online after that were not affected. Hosts that had already taken the update were another matter, and CrowdStrike's own remediation guidance starts with "Reboot the host to give it an opportunity to download the reverted channel file," followed by a section headed "If the host crashes again on reboot." After that come bootable recovery drives, and a note that BitLocker-encrypted hosts may need a recovery key. The revert was fast. It could not reach machines that could no longer boot.
Count the rows. The deploy tool controls one of seven for certain and a second one sometimes. Everything else is state, and state only runs forward.
The case where rollback really is trivial
The strongest argument against all this is blue/green deployment, and it deserves a fair hearing. Google's SRE Workbook describes two environments, one serving traffic and one ready, where "rollback is a trivial reversal of the router change." That is true. It is also true that the workbook recommends smaller, more frequent releases precisely because they make any given release "cheaper and easier to roll back." Both are good advice. I use both.
Look at what the router reversal actually reverses. It moves traffic. If blue and green share a database, a queue, a payment provider or an email service, the router has nothing to say about any of them. For ten minutes blue wrote rows in blue's format, sent blue's messages and charged cards on blue's logic. Flip the router and green inherits all of it. Rollback is trivial only in the special case where the new version wrote nothing that outlives it. The same workbook, discussing a deploy that goes wrong, allows for the case where the process "doesn't provide us the option to roll back to a previously known good configuration." Their answer there is to patch forward. For state, that is the only answer there has ever been.
The other modern answer is GitOps: revert the commit and let the controller reconcile. That reverts the declaration. It does not reach what the environment did while the bad declaration was live. The Terraform AWS provider's documentation for S3 buckets is blunt about one case: with force_destroy set, destroying a bucket deletes its objects, and "these objects are not recoverable." Revert the commit that removed the bucket and Terraform will plan a new bucket with the same name. It will be empty.
I wrote in the anatomy of a production outage that the teams who recover fastest roll back to known good states. I still believe it, with one condition I would now put in bold. "Known good" has to describe the whole system, not just the binary. Knight's seven servers went back to known good code, and it was exactly the wrong move.
Where my own rollbacks stopped halfway
I do not need a trading firm to see this. This site runs a correction pipeline: when a fact check catches an error, I fix the article and redeploy. I treated the article body as the thing being corrected. It was one surface among many.
The same claim lives in the page title, the meta description, the image alt text, the summary that goes into the email digest, the carousel slides, the slide alt text and a folder of social drafts. None of them regenerate from the body. Once, the corrected body went live while the old figure stayed in the title, the meta description and the image alt text. No gate caught it; I found it by fetching the live page. Another time a retracted claim survived in the summary, a slide, its alt text and three social drafts after the body had dropped it. The checks passed both times, because the checks read the body. The first fix was a written rule: when a figure changes, search for the old value across every surface, then confirm it is gone from production as well as that the new one is there. It is Antunes's MyISAM table with more tables.
A rule depends on someone remembering it, so it became a gate. Every derived surface now carries a stamp naming the commit of the article it was last checked against, and the deploy refuses to ship when the article has moved on and a stamp has not. It fired this week. A correction made two days earlier had never reached one rendered page or two of its social drafts, which still stated the old timeline, and the deploy stopped until they matched.
And some of those surfaces cannot be reverted at all. On Bluesky there is no native edit. The tools that fake one delete the post and recreate it, which resets its likes, reposts and quotes and leaves no record that anything changed. A published post is a message already sent. The only honest correction is a new one, in the thread, where people can see it. That is forward-fixing, and it is the only kind of rollback a published claim gets.
Write the undo before the do
None of this says do not roll back. It says a rollback plan that lists only the code is a plan for the one layer that was never the problem. What works is deciding, before the deploy, which parts can go backward and which can only go forward, and ordering the work around that.
Classify every write. For each change, name where it lands and whether reverting the deploy undoes it. Use the table above. Anything in the lower five rows needs its own plan, and that plan is a forward fix you write in advance.
Put the reversible writes first. This is Antunes's fix, and it generalises. Commit the part that can be undone, then do the part that cannot. If the run dies in between, you land in a half-finished state you chose ahead of time. Make every step idempotent so the recovery is "run it again" rather than a hand-written repair.
Change schema in three moves, not one. Danilo Sato's parallel change pattern, also called expand and contract, splits a breaking change "into three distinct phases: expand, migrate, and contract." Add the new shape alongside the old. Move every reader and writer across. Remove the old shape only when nothing uses it. Between the first and last step, both versions of the code can run against the same data, so rolling the code back stops being a bet on the schema. A deploy that fails during the expand phase needs no database rollback at all: the old code does not know the new column exists.
Retire the flag, not just the code. Knight's failure needed an old function left callable and a flag given a new meaning. Delete dead code in its own release, and never reuse a flag name. Give the new behaviour a new flag.
Find the tables that cannot roll back, and know what that search misses. On MySQL this finds the obvious ones:
mysql> SELECT table_schema, table_name, engine
-> FROM information_schema.tables
-> WHERE engine <> 'InnoDB'
-> AND table_schema NOT IN ('mysql','sys','information_schema','performance_schema');
I tried to break that recommendation before printing it, and it broke. I set up a transaction touching only InnoDB tables: an UPDATE, then an ALTER TABLE, then ROLLBACK. The update survived. The query above returned nothing, because every table was InnoDB. The manual explains why: DDL statements "implicitly end any transaction active in the current session, as if you had done a COMMIT before executing the statement." So the query is a screen for one cause, not proof of atomicity. Also search your migration and job code for DDL inside a transaction, and treat any routine that writes to two stores, such as a database and a queue or a database and a payment API, as having no transaction at all. A saga does not change that. Its compensating step is a new write that tries to cancel an old one, and it can fail like any other write. The protocol that can span two stores, two-phase commit, asks every participant to promise before anyone acts, and nothing runs it for you when a request writes a row and then publishes to a queue. That last case is one AI coding agents write readily, and an agent that reverts its own commit has reverted the commit, nothing else. For more on why a database is a harder boundary than it looks, see your database is already an API.
Rehearse the rollback against production-shaped state. A revert tested on an empty staging database proves the binary installs. It says nothing about the rows, flags and queues the new version will have touched by the time you need it. Run the new version long enough to write real state, then roll back and check every row of the table above.
Before Your Next Deploy
Five checks. Each one fits in an afternoon, and none needs a new tool.
- Pick your next release and list every place it writes: code, config, flags, schema, rows, messages, external calls. Mark which ones reverting the deploy actually undoes.
- For each write that cannot be undone, write the forward fix now, before you ship, and put it in the runbook next to the rollback command.
- Reorder the job so reversible writes commit first and irreversible ones run last, and make every step safe to run twice.
- Run the engine query against production, then search migrations and jobs for DDL inside a transaction and for any routine that writes to two stores.
- Rehearse one rollback on a copy that the new version has already written real state into, and check every row of the list afterwards.
The Bottom Line
Rollback earned its reputation in the era when a deploy meant swapping a binary. Deploys now write rows, flip flags, publish messages and call other companies' systems, and a revert reaches almost none of that. Knight's rollback reverted the code correctly and made the system worse. CrowdStrike's revert was fast and still could not reach machines that no longer booted. Treat every rollback as a deploy of its own. Decide before you ship which writes can go backward, put those first, and write the forward fix for everything else while you are calm enough to think.
"A rollback plan that lists only the code is a plan for the one layer that was never the problem."
Sources
- The rollback that only rolled back half of it — Paulo Antunes's postmortem of a transaction spanning InnoDB and MyISAM tables in a factory ERP, where ROLLBACK undid one write and silently kept the other.
- In the Matter of Knight Capital Americas LLC, Release No. 34-70694 — SEC order on Knight Capital's August 1, 2012 trading incident: code missing from one of eight servers, a repurposed flag, and a rollback that worsened the problem. $12 million penalty.
- Falcon Content Update Preliminary Post Incident Report — CrowdStrike's preliminary review of the July 19, 2024 content update, which reached Windows hosts between 04:09 and 05:27 UTC and was reverted at 05:27.
- Falcon Content Update Remediation and Guidance Hub — CrowdStrike's recovery steps for hosts affected on July 19, 2024: reboot to fetch the reverted file, then bootable recovery drives, with BitLocker recovery keys for encrypted hosts.
- Parallel Change — Danilo Sato on expand and contract: splitting a breaking interface or schema change into expand, migrate and contract phases.
- Canarying Releases (The Site Reliability Workbook, Chapter 16) — Google's SRE Workbook on canary releases, smaller release artifacts that are cheaper to roll back, and blue/green deployment where rollback is a reversal of the router change.
- MySQL 8.4 Reference Manual: START TRANSACTION, COMMIT, and ROLLBACK Statements — Documents that changes to nontransactional tables cannot be rolled back and that ROLLBACK raises ER_WARNING_NOT_COMPLETE_ROLLBACK when a transaction touched one.
- MySQL 8.4 Reference Manual: Statements That Cause an Implicit Commit — Lists the statements, including DDL such as ALTER TABLE, that implicitly end any active transaction as if a COMMIT had been issued.
- MySQL Connector/Python Developer Guide: Connector/Python Connection Arguments — Lists get_warnings and raise_on_warnings, both False by default, and suggests the sql_mode setting for turning warnings into errors.
- Terraform AWS Provider: aws_s3_bucket resource — Documents force_destroy: when a bucket is destroyed, all its objects are deleted and are not recoverable.
Disagree? Have a War Story?
I read every reply. If you've seen this pattern play out differently, or have a counter-example that breaks my argument, I want to hear it.
Send a Reply →