← Back to writing

Every engineer pays the CI tax on every push. That's what makes this the highest-leverage change I've made — it isn't one team's problem solved once, it's time returned to everybody, permanently, on every commit for as long as the project lives.

We were at 30–40 minutes. We're at 8.

Full pipeline before
30–40 min
Full pipeline now
8 min
Elixir test step before
20.5 min
Elixir test step now
3.5 min

The twenty-minute password

Most of it was one thing, and it was almost funny once we found it.

Every test that needed a user created one through a shared data case. That case ran the real registration path, which ran the real password hashing — with the production work factor. Password hashing is designed to be slow; that's the entire point of it. We were paying that cost thousands of times per run, deliberately, in an environment where the passwords were password123 and nobody was attacking anything.

That alone was roughly twenty minutes of the run.

The fix is a two-line configuration change that reduces the work factor in the test environment only. Production hashing is untouched, and it's worth being explicit about that, because this is exactly the kind of change that looks like a security regression in a diff if you don't read which config file it landed in.

Disk that didn't need to be disk

Test data is disposable by definition. It exists for the length of one transaction and then gets rolled back or dropped. Writing it to a real filesystem buys you durability you will never once use.

Postgres moved onto tmpfs. Faster, obviously — but it also removed a whole class of intermittent I/O failures that had been quietly filed under "flaky CI" for years. That's a recurring theme in this work: the thing you think is a performance problem and the thing you think is a reliability problem are frequently the same problem.

Two speed fixes that were really correctness fixes

Background jobs were being killed mid-flight by the test harness. Supervised async work would start, the test would finish, and the harness would tear everything down before the job completed.

So the suite was fast and lying. Tests were passing without the work they were nominally testing having actually run. Fixing the timeouts made CI slower in one narrow place and made every result after it worth believing. I'd have taken that trade at any price; getting it alongside a speedup was luck.

The second one: runner allocation. Each pull request was claiming four runner slots. With several PRs open, everything queued behind everything else, and queue time is what "slow CI" actually feels like to a human being — nobody experiences pipeline duration, they experience waiting. Dropping to two slots per PR meant one PR's jobs land on one worker and the queue drains.

Verifying it

The temptation with a change like this is to ship it the moment the number goes green. The number going green is precisely when you should be most suspicious — a test suite can be made arbitrarily fast by making it test less.

2,178 tests, three consecutive green runs, before merge. Same tests, same assertions, same coverage. Only the waiting is gone.

What I'd tell someone starting this

Profile before you optimise, obviously. But more usefully: look for the thing that runs thousands of times, not the thing that takes the longest once. Everybody's instinct is to attack the slowest single step. The twenty-minute win here was hiding in a function that took forty milliseconds.

And treat a slow pipeline as a correctness problem before you treat it as a performance one. A loop long enough that people stop waiting for it is a loop people start ignoring, and a test suite people ignore is worse than no test suite at all — because it teaches the whole team that red doesn't mean anything.