Work
Selected engineering work. Client names are omitted deliberately; longer write-ups are linked where they exist.
CI: 40 minutes to 8
Cut the pipeline from 30–40 minutes to 8 across a 2,178-test Elixir suite. Most of it was one pathology — every test paying full production password-hashing cost — plus Postgres on tmpfs and a fix to runner allocation.
Two of the changes were correctness fixes in disguise: background jobs were being killed mid-flight by the test harness, so the suite had been fast and lying.
Read the write-up →
Elixir, GitHub Actions, PostgreSQL
Observability From Zero
There were no metrics, no alerts and no dashboards — production being down was something we learned from a customer. Built the monitoring platform end to end across the fleet: Prometheus and Grafana, blackbox uptime probes, alert routing, and a consolidated inventory so no machine runs somewhere we forgot about.
All of it defined in git and shipped through the same reviewed pipeline as product code. Each product got its own view, because a dashboard everybody owns is a dashboard nobody opens.
Read the write-up →
Prometheus, Grafana, Azure, GitOps
A Deploy Pipeline With No Password In It
One path to production for everything we run, holding no long-lived credentials. CI proves its identity with a token that lives for minutes; OpenBao validates it and returns only the secrets that job needs, which then expire.
The same path deploys a Worker at the edge and a container on a dedicated server — same review, same audit trail, same one-click rollback. Four applications now run on it.
Read the write-up →
GitHub Actions, OIDC, OpenBao, Cloudflare
Bento — Draft Zero
A product I build under Draft Zero: paste one self-contained HTML document, get a clean sandboxed URL. No build step, no hosting to configure.
It landed in a niche I hadn't planned for — AI tooling is good at producing a single HTML file and bad at editing a WordPress theme, so "generate a page and share it" had nowhere to land. Since deployed for internal use at my employer behind a Cloudflare Access perimeter, where it also hosts build artifacts.
Cloudflare Workers, Cloudflare Access, MCP
A Black Box Recorder for Containers
Container crashes were the least debuggable failure we had — the evidence died with the process, so every investigation started from zero. Wrote a watchdog that installs as a system service on every host and captures the full picture at the moment of exit, posting it to the team channel within seconds.
It immediately surfaced problems that had been happening invisibly for months: workloads killed under memory pressure mid-job, restart loops nobody had seen, and jobs that had outgrown their machines.
Read the write-up →
Go, Docker, systemd
Cloud Migration and IaC Foundation
Migrated the production estate — 38 VMs, 46 disks, AI deployments and DNS — into a new Azure subscription single-handedly, with near-zero post-migration issues.
Structured the destination from scratch rather than lifting the old mess across: a tagging taxonomy that made cost analysis possible for the first time, and a three-region topology with clean environment segregation. Used the migration to bootstrap infrastructure as code. Cost dashboards built for capacity planning turned up ~$8.4k/year of right-sizing savings as a side effect.
Azure, OpenTofu, Terragrunt, Cloudflare
Also
Security moved inside engineering. Secret, dependency, static and container scanning as CI gates — report-only first, blocking second, because a scanner that blocks on day one is a scanner everyone learns to ignore. Supply-chain controls against package poisoning: version cooldowns, blocked install scripts, locked dependencies.
An end-to-end test suite, finally. A years-open gap. Not a script somebody runs occasionally but a system with its own dashboard, running on every merge and again nightly on ephemeral infrastructure, with flakiness tracked over time.
Shipping to production without an engineer. A Cloudflare Worker in front of the marketing domain routes bespoke pages from git while everything else passes through to the WordPress origin. A team that doesn't write code now ships to the real site, reviewed and revertible, with no engineer in the loop.
Ingestion pipeline, 5× faster. Cut a production data pipeline from ~14 hours to ~3 by profiling and restructuring query patterns and the MySQL configuration behind them.
Federated SSO for enterprise customers. Single sign-on between our platform and a Big Four consultancy's identity infrastructure. Difficult structurally rather than protocol-wise: a split frontend/backend deployment meant the standard flows didn't apply, so it needed a custom token exchange that leaked nothing at any boundary.
Frontend architecture. Set the layout, theming, state and component conventions the team still builds on, then codified them so they get applied without supervision.