Cutting flaky end-to-end tests by 80%
The problem
The commerce end-to-end suite failed often enough that engineers re-ran it by reflex. A red build no longer meant a broken product, so real regressions slipped through and deploys slowed down.
What I did
- Audited eight months of recorded failures and classified each one by root cause instead of by test name.
- Drove the migration of the suite from Cypress to Playwright, using the move to fix waiting and isolation problems at the source.
- Coordinated fixes with the teams that owned the underlying causes, since many failures came from shared services, not the tests.
Result
The flaky-test rate dropped by 80% and the team trusted CI again, which raised deployment confidence for every commerce release.
What I'd tell another team
Treat flakiness as a concurrency problem. Most "random" failures fall into a handful of buckets: shared test data, timing assumptions, environment dependencies, and real race conditions in the product. Count failures per bucket and fix the biggest bucket first. Quarantine a flaky test only with a named owner and a deadline, or quarantine becomes deletion.