One Week of Looking. Two Weeks of Truth.
There is an email we get a dozen times a year: the product works, but nothing ships anymore. The team is good, customers are paying, the roadmap exists — and yet the last three releases slipped and the newest engineer has been “almost done” with their first ticket for six weeks. Roughly a third of our work is rescue-and-scale, and nearly all of it starts there.
The Pattern
Rescue projects share a skeleton: a capable team, a real deadline, and an architecture nobody ever wrote down. Not from carelessness — it was built under delivery pressure, one honest decision at a time, and the decisions were never read back together. The 2021 diagram is a lie; the real one lives across four README files, a whiteboard photo in Slack, and the head of the person who left in March.
Leadership feels four symptoms and files them as four separate problems:
- Velocity cliff. Throughput drops quarter over quarter and nobody can name the cause. Usually it is shared mutable code: everything touches everything, so every ticket costs a survey.
- Incident rate. The same three subsystems keep appearing in postmortems under three different-looking root causes — one coupling problem.
- Deploy fear. Releases are batched and done on quiet days with two people awake. Friday deploys are banned — not by policy, by trauma.
- Onboarding time. Time-to-first-merged-PR stretches from days to months. If a new hire cannot trace a request without a guided tour, neither can anyone else.
Four symptoms, one disease: nobody has looked at the whole system at once since it was small enough to fit in one head.
Inside the Audit Week
The audit is one week, fixed price — long enough to be evidence-based, short enough that nobody manages the engagement instead of their product:
- Day 1 — Repository and history. Tree, build graph, git log: churn, hot spots, commit size trends, bus factor per directory.
git log --numstat— archaeology, not judgment. - Day 2 — Runtime and deploys. We sit with the pipeline: how a commit becomes production, how long, who approves, what rolls back, what happens at 02:00.
- Day 3 — Data model. Schema, migrations, constraints, orphan rules, fields that mean two things in two services — the defects you cannot refactor away cheaply, so they get a day.
- Day 4 — Tests and CI. Not the coverage number — what the suite asserts, which tests are load-bearing, which are decorative, which fail on Tuesdays — plus pipeline duration and flake rate.
- Day 5 — Interviews and findings. Four to six structured thirty-minute conversations with the people who carry the system, then the severity-ranked write-up.
What we deliberately do not do in a week: no pull requests, no migrations, no refactors, no tooling swaps, no 80-page deck, no opinions about individuals. Estimating a full remediation plan to the story point takes longer than the audit: that is a guess wearing a spreadsheet.
Refusing to look at the thing before rebuilding it is not bold — it is skipping the only cheap step available.
What We Examine
Seven lenses:
- Architecture and coupling. Module boundaries versus runtime boundaries, who imports whom, where the cycles are, which “services” are one deployable wearing a costume.
- Data model integrity. Constraints enforced in the schema versus in three layers of application code — or nowhere. Nullable-everything columns, orphan rows, dual-write paths, stale migrations.
- Test coverage quality. Coverage percentage is a vanity metric — 80% on getters proves nothing about the payment path. We read for what breaks when it changes: contract tests, integration depth, whether it catches last quarter’s bug class.
- Pipeline and release risk. Build duration, flake rate, artifact reproducibility, rollback path, migration safety, environment drift, and every manual gate only one person can open.
- Security posture. Dependency age and known CVEs, secret handling, authz at the boundary versus buried in handlers, upload paths, the “temporary” admin endpoint. What an attacker with five minutes finds.
- Performance and observability. Real traces over folklore: p50/p95 where it matters, log volume versus signal, and whether anyone can answer “slow for everyone or just this tenant” without SSHing.
- Team velocity signals. Git history and incident records: PR size and review latency, rework rate, hot spots owned by fewer than two people, MTTR, how many of the last ten incidents share a subsystem.
Our capabilities page lists this as an engagement type; the honest description is narrower — find the two or three constraints that explain ninety percent of the pain.
The Report
One findings document, structured to survive being forwarded to someone who was not in the room. Every finding carries a severity, evidence, and an effort-versus-impact position — in that order, because severity without evidence is an opinion.
| Severity | Definition | Example finding | Typical response window |
|---|---|---|---|
| P0 | Data loss, security exposure, or an outage path with no workaround. Already biting you. | Payment retries write duplicate ledger rows: the idempotency key is checked after the write. | Before the next release. |
| P1 | Structural defect that compounds weekly: blocks a roadmap item or taxes every deploy. | Catalog and checkout share one mutable module, so no change ships without re-testing both. | Next quarter, scheduled into the train. |
| P2 | Waste and friction. Nobody is bleeding, but capacity leaks: slow builds, flaky tests, dead paths. | CI takes 14 minutes with a 40% re-run rate, so engineers batch changes — and batched changes merge badly. | Batched opportunistically. |
Evidence means the artifact: the failing query, the commit hash, the trace, the EXPLAIN output. A finding without evidence in a week is a hypothesis, and it lands in an appendix labelled as one.
Remediation is ordered by effort against impact — a two-by-two, not a ranked list, because “do this first” is meaningless if it needs a quarter of platform work to unlock. The front of the report is a one-page executive summary, because that is the only page most of the room will read; a report nobody reads is an expensive PDF.
Three P0s and a sequenced plan beat forty observations. If the findings run longer than a few pages, we failed to prioritize.
Why We Almost Never Recommend the Rewrite
About one audit in five opens with the client half-committed to a rebuild. The rewrite trap has three teeth:
- Dual running. The new system must serve production while the old one still does — two bug queues, two deploy paths, two on-call playbooks — for far longer than the plan says. The plan says six months. It is never six months.
- The feature parity myth. Parity is measured against the current product, so the new system spends its whole budget arriving where the old one has already left — while the old one accrues the undocumented edge cases production taught you — five years of “why is this column not null”.
- The second system effect. Freed from constraints, the team redesigns for the architecture they wish they had. Elegant, general, late, and optimized for problems you do not have.
The alternative is the strangler-fig pattern: a seam in front of the legacy path, one capability moved across it at a time, old code deleted once traffic has provably moved. Each step ships through your existing release train behind a feature flag — dark first, then ramped, reversible in one commit. You get the win without the cutover weekend, and the team learns the new shape on real traffic instead of a staging database that lies.
# remediation-sequence.plan — 12 weeks, one release train
Weeks 1-2 [P0] break the dual-write path; backfill behind flag
flag: ledger.dual_write ramp 0% → 5% → 25% → 100%
Weeks 3-4 [P1] introduce seam module; route read traffic only
flag: catalog.seam_read ramp 1% → 10% → 100%
Weeks 5-6 [P1] contract tests around the 4 payment boundaries
Weeks 7-8 [P2] cut CI 14m → 4m; gate on changed files only
Weeks 9-10 [P1] migrate write traffic; delete the old path's flag
Weeks 11-12 [P2] tracing on checkout; p95 budget enforced in CI
Rules: every step merges to trunk, ships dark, then ramps.
every step is revertible in one commit.
no step requires a cutover window or a freeze.Sequencing is the craft of a rescue — see the Atlas rebuild: an 18ms IPC boundary and a 42MB footprint, from a desktop rebuild that could have become a two-year rewrite. Want that on your codebase? Start here.
When a Rewrite Genuinely Is Right
We are not ideologically against rebuilding. The honest list of cases where we recommend it is short, and none of them are “the code is ugly”:
- Unsupported runtime. The interpreter, framework or OS major version is end-of-life and the upgrade path genuinely does not exist — not that nobody tried, but that the vendor abandoned it.
- Licensing or hosting dead end. The cost curve or the compliance story makes the current substrate untenable inside the planning horizon. You are leaving either way; leave on purpose.
- Data model that cannot express the domain. If the schema cannot represent a first-class business concept however you contort it, every feature becomes a hack. That one is worth paying for.
- Security defect with no viable patch. A flaw in the component’s fundamental design, with no mitigation you can ship and verify. Rare, non-negotiable.
Notice what is missing: boredom, resume-driven development, a new framework everyone likes. If that is the real motivation, the report will say so.
Cost, Timeline, and What Happens After
The audit is fixed-price and quoted before it starts: a week of senior time, no change orders, no discovery before the discovery. What follows:
| Stage | Duration | What you get | Who from your side |
|---|---|---|---|
| Audit | 1 week, fixed price | Repo, runtime, data, test and security findings with evidence attached. | Engineering lead (present, not parked) plus a domain person. |
| Findings and plan | Within 3 working days | Severity-ranked report, one-page executive summary, effort-vs-impact map. | Roadmap owner — someone who can say yes. |
| Remediation increments | 6–12 weeks, in your release train | Flagged, sequenced changes shipped by your team; we pair, review, unblock. | 2–4 of your engineers; we never become a shadow team. |
| 90-day outcome | Measured at day 90 | Baseline vs. current on deploy frequency, lead time, change failure rate, MTTR. | Engineering lead; numbers go to the roadmap sign-off. |
Those roles are not optional — without the engineering lead in the room and a roadmap owner who can say yes, findings become a PDF nobody acts on. We still need no production credentials: just history, one deploy watched end to end, and permission to ask uncomfortable questions. Our process covers what happens after the report.
What “Good” Looks Like at 90 Days
Ninety days is long enough to move the numbers and short enough to stay attributable to the work rather than a lucky quarter. We baseline these in the audit week and re-measure on day 90:
- Deploy frequency — from 2–3 per week to daily or better; small batched changes ship without ceremony.
- Lead time for changes — commit to production in hours, not sprints; the queue is fear, not compute.
- Change failure rate — down; seams and contract tests catch the break class that used to reach customers.
- MTTR — down hardest: observability and rollback were part of the fix, not an afterthought.
And the one that is on no dashboard: the team stops being afraid of their own release button. Deploys stop being events. Nobody schedules them for a quiet afternoon with two people watching a graph, and the newest engineer ships in week two instead of month four. Nobody quotes a metric for it; you notice it before you chart it.
That is the cheapest way to find out whether you need a sequence of changes or a conversation about a rebuild — one of those answers is usually right.
One Week of Looking Beats a Year of Guessing.
If your release button has become a loaded weapon, bring it to a principal engineer — a fixed-price audit week usually settles whether you need a sequence of changes or a rebuild, in writing, with evidence.