Testing Automation, CI Sharding & Flaky Test Engineering
Writing tests is only half the battle; ensuring that thousands of tests run reliably, fast, and deterministically within CI/CD pipelines is the hallmark of senior engineering leadership. Flaky tests erode team trust in CI, leading to ignored failures and blocked releases.
In this lesson, we explore automated test pipeline orchestration, test sharding and parallelization across multiple CI machines, diagnosing and eliminating flaky tests, trace artifact collection, and enforcing code coverage gates.
┌────────────────────────────────────────────────────────────┐
│ CI Test Sharding Architecture (Playwright) │
├────────────────────────────────────────────────────────────┤
│ GitHub Actions Matrix Runner: │
│ ├── Job 1 (Shard 1/4): Runs Tests 001 - 050 on Worker 1 │
│ ├── Job 2 (Shard 2/4): Runs Tests 051 - 100 on Worker 2 │
│ ├── Job 3 (Shard 3/4): Runs Tests 101 - 150 on Worker 3 │
│ └── Job 4 (Shard 4/4): Runs Tests 151 - 200 on Worker 4 │
├────────────────────────────────────────────────────────────┤
│ Result: 200 E2E tests finish in 3 minutes instead of 12 min│
└────────────────────────────────────────────────────────────┘
1. Test Sharding & Parallel Execution in CI
When an application's end-to-end test suite grows to hundreds of scenarios, running them sequentially on a single CI machine creates agonizing 30-minute build bottlenecks.
Test Sharding splits the test suite evenly across multiple parallel runner containers:
# .github/workflows/e2e.yml
name: Playwright E2E Tests
on: [push, pull_request]
jobs:
e2e:
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
shardIndex: [1, 2, 3, 4]
shardTotal: [4]
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
- run: npm ci
- run: npx playwright install --with-deps
# Run only the assigned slice of the test suite!
- run: npx playwright test --shard=${{ matrix.shardIndex }}/${{ matrix.shardTotal }}
# Upload failure traces & video artifacts
- uses: actions/upload-artifact@v4
if: always()
with:
name: playwright-report-shard-${{ matrix.shardIndex }}
path: playwright-report/
retention-days: 7
2. Root Causes of Flaky Tests & How to Eliminate Them
A Flaky Test is a test that sometimes passes and sometimes fails on the exact same commit without code changes.
The 4 Major Causes of Flakiness:
- Hardcoded
sleep()/ Timeouts (setTimeout(1000)):- Problem: CI machines are slower than developer laptops; an arbitrary 1000ms delay will randomly fail under CI CPU load.
- Fix: Use auto-waiting assertions:
await expect(page.getByText('Success')).toBeVisible().
- Shared State Between Test Cases:
- Problem: Test B fails only when run after Test A because database records or localStorage cookies leaked.
- Fix: Isolate browser contexts per test (
browser.newContext()) and use ephemeral database transactions.
- Asynchronous Network Race Conditions:
- Problem: Clicking a button before the component has finished attaching its event handler or loading initial data.
- Fix: Wait for network idle or specific API response interception before triggering user events.
- Timezone & Locale Assumptions:
- Problem: Tests formatting dates using the local system timezone (
UTC-5on dev machine vsUTCon GitHub Actions). - Fix: Pin the timezone explicitly in Playwright config:
use: { timezoneId: 'UTC', locale: 'en-US' }.
- Problem: Tests formatting dates using the local system timezone (
3. Playwright Trace Collection & Failure Artifacts
When an E2E test fails inside a remote Linux CI runner, debugging from text logs alone is painful. Configure Playwright to record full Execution Traces on failure:
// playwright.config.ts
import { defineConfig } from "@playwright/test";
export default defineConfig({
use: {
// Record video, screenshots, and DOM snapshots ONLY on test failure
trace: "retain-on-failure",
screenshot: "only-on-failure",
video: "retain-on-failure",
},
retries: process.env.CI ? 2 : 0, // Auto-retry flaky tests in CI
workers: process.env.CI ? 4 : undefined,
});
The resulting trace.zip file can be inspected using npx playwright show-trace trace.zip, giving you a full time-travel debugger with DOM snapshots, network requests, console logs, and mouse coordinate timelines.
4. Enforcing Code Coverage Gates in Vitest
Configure strict coverage thresholds to prevent untested features from merging into the main branch:
// vitest.config.ts
import { defineConfig } from "vitest/config";
export default defineConfig({
test: {
coverage: {
provider: "v8",
reporter: ["text", "json", "html"],
thresholds: {
lines: 80,
functions: 85,
branches: 75,
statements: 80,
},
exclude: ["**/*.d.ts", "**/*.config.ts", "**/dist/**"],
},
},
});
Summary
- CI Test Sharding distributes test suites across parallel runner machines, reducing test run times linearly.
- Flaky tests stem from arbitrary sleep timers, shared global state, network race conditions, and timezone discrepancies.
- Playwright's auto-waiting assertions eliminate timing race conditions.
- Playwright Traces (
retain-on-failure) capture interactive time-travel recordings for instant remote CI debugging. - Coverage gates in Vitest enforce minimum quality thresholds across lines, functions, and branches.
Best Practices
- Never Use
page.waitForTimeout()orsleep(): Always await explicit DOM state assertions (toBeVisible(),toHaveText()). - Isolate Test State Completely: Ensure every E2E test runs with a fresh browser context and isolated test user.
- Quarantine Unstable Flaky Tests: Mark flaky tests with
.fixme()until fixed to preserve developer confidence in CI. - Enforce Coverage Thresholds on Business Logic: Set strict 80%+ coverage gates on pricing, authentication, and validation domains.