[ Engineering · Performance ]

A Performance Budget Nobody Enforces Is a Wish.

$ lighthouse-ci assert --assertions.largest-contentful-paint
  route /           LCP 940ms p75   ✓ budget 1000ms
  route /services    LCP 1120ms     ✗ budget 1000ms
  route /           INP 164ms      ✓ budget 200ms
  route /           CLS 0.00       ✓ budget 0.10
bundle 187.4kb / 190kb  fonts 96kb / 120kb  PASS
exit 1 — block merge, comment posted on PR #2841

Every team we audit has a performance budget. It lives in a Notion page written during a planning offsite, it says something like “LCP under 2.5s”, and it has not been read since — let alone enforced. That document is not a budget. It is a wish with a unit of measurement attached.

Why Performance Budgets Rot

The failure mode is always the same shape. Someone writes the budget, nobody wires it into the pipeline, and the first real test comes six months later during a launch — when p75 has quietly climbed past 3.5s and the growth team is describing the conversion dip as “seasonal”. By then the retrofit costs an order of magnitude more than a red build would have.

Budgets rot because they are treated as documentation instead of machinery. Documentation ages silently. Machinery either holds or breaks, and broken machinery gets noticed the same day it breaks.

Yours has rotted if any of these are true:

  • The budget lives outside the repo — a doc, a slide, a Slack thread — so changing it needs a conversation, not a reviewed pull request.
  • Nobody can name the last build that failed on performance. Zero failures in a year does not mean zero regressions; it means zero enforcement.
  • The number was copied from Google’s public thresholds rather than derived from your own baseline. Generic thresholds produce generic compliance.
A budget that only exists in a document is a wish with a unit of measurement. The unit does no work — only the failing build does work.

The Actual Numbers

What we enforce, and why each number sits where it does — lab assertions on a pull request, expressed as what a median run must beat:

  • LCP ≤ 1.0s at p75 on simulated 4G — far tighter than the public 2.5s “good” threshold.
  • INP ≤ 200ms — at the boundary of “good”, because interaction latency degrades nonlinearly on mid-tier Android.
  • CLS ≤ 0.1 in the field; CLS = 0 asserted in the lab on marketing routes.

The LCP number looks absurd until you account for the gap between a lab machine and a real user. Our lab box has a warm cache and a controlled network; the field p75 user is a three-year-old Android on a train. Lab median 1.0s → field p75 around 2.1s, comfortably green. Lab median drifting to 1.6s → field p75 past 2.5s and a yellow flag in Search Console within a reporting cycle. The lab budget is a leading indicator; set it where the lagging field number still has headroom.

On CLS we split hairs deliberately. Field budget is 0.1, the threshold real users are judged against. Lab budget on marketing routes is 0 — every shift on a landing page is ours: an unsized image, a font swap, a banner injected before the hero. If a machine can load the page with zero shift, a shift in review is a defect, not an environmental fact.

Budgets differ by route class, and conflating the two is where programs go wrong. A marketing route is LCP-bound: tiny JS budget, no hydration, hard cap on render-blocking CSS, hero image preloaded and dimensioned. An app shell is INP-bound: it may ship more initial JavaScript because it needs a runtime, but it gets a hard main-thread budget — no task over 50ms during interaction, no layout thrash in the hot path, lazy routes that stay lazy.

Why p75 field data rather than a lab median, given we just spent a paragraph defending lab numbers? Because the median hides the population you are losing. A regression that slows 30% of sessions — a legacy browser path, a device tier — may not move the median at all while moving p75 sharply. p75 is where you start paying for the tail; the median is a vanity metric with a dashboard.

Lab Versus Field

These are not competing sources of truth; they answer different questions. Mixing them up is how teams declare victory while their Search Console graph burns.

Lighthouse in CI is a controlled experiment: fixed device, fixed throttling, fixed Chrome version. It is excellent at exactly one thing — detecting when this change made the page slower than the last build: a regression detector with a stopwatch, not a measurement of your users.

A green lab run does not guarantee a green field p75. The lab never runs your tag manager’s consent flow, never hits a cold CDN edge where your traffic actually lives, and — critically — exercises almost no real interaction, so INP in the lab is close to fiction. Field INP comes from users tapping things: a select with a 400ms handler, a beacon fired on click, a modal re-rendering a list. Lighthouse reproduces none of it.

What each is good for:

  • Lab (Lighthouse CI): every pull request, ~60–90 seconds, blocks regressions in bytes, layout, and paint timing.
  • Field (CrUX API / RUM): daily or weekly, tells you what users experienced over a rolling 28 days. Never a PR gate — it is delayed, aggregated, and only exists once you have traffic.

Our split: lab blocks merges, field alerts. If field p75 degrades two weeks running while lab stays green, the lab model is wrong — usually a missing third-party or device tier — and the fix belongs in the model, not the number. That distinction is most of what our performance engineering work comes down to.

Building the Gate

Lighthouse CI gives you assertions natively. This config runs on two routes per pull request — home page and the heaviest marketing route — three runs each, median aggregation:

{
  "ci": {
    "collect": {
      "numberOfRuns": 3,
      "startServerCommand": "npm run preview",
      "url": [
        "http://localhost:4173/",
        "http://localhost:4173/services"
      ]
    },
    "assert": {
      "assertions": {
        "categories:performance": ["error", { "minScore": 0.92 }],
        "largest-contentful-paint": ["error", { "maxNumericValue": 1000 }],
        "cumulative-layout-shift": ["error", { "maxNumericValue": 0.01 }],
        "total-blocking-time": ["error", { "maxNumericValue": 150 }],
        "unsized-images": "error",
        "font-display": "error",
        "render-blocking-resources": ["warn", { "maxLength": 1 }]
      }
    },
    "upload": { "target": "temporary-public-storage" }
  }
}

Three things earn their keep. numberOfRuns: 3 takes the median, which kills most single-run noise. render-blocking-resources is a warning, not an error — one acceptable stylesheet is not worth blocking a deploy. And unsized-images as an error makes CLS = 0 enforceable rather than aspirational.

A full browser run is the slowest check in the pipeline, so it should not be the first line of defence. Two static gates run in well under two seconds and catch most regressions before Lighthouse even launches:

// scripts/budget-check.mjs — runs before any browser does
const budgets = { js: 190_000, css: 60_000, fonts: 120_000, images: 400_000 };
const actual  = await measure('dist/assets');
const fonts   = await countFontFiles('dist/assets/fonts');

for (const [name, bytes] of Object.entries(actual)) {
  if (bytes > budgets[name] && routeClass !== 'app-shell') {
    fail(`${name}: ${fmt(bytes)} > budget ${fmt(budgets[name])}`);
  }
}
if (fonts.files > 6 || fonts.woff2Total > budgets.fonts) {
  fail(`font budget blown — ${fonts.files} files, ${fmt(fonts.woff2Total)}`);
}

The font budget is the one people skip and then regret. Two families, one weight axis, subset per script, ≤ 120kb total — a failed font check tells you in a second what Lighthouse buries inside a composite score. Byte budgets also fail with a diff: “hero.webp grew 340kb” is actionable, a 0.89 score is not.

Gate the cause, not the score

A composite performance score is an average of averages — it can stay green while LCP regresses, because CLS improved. Assert on the individual metric audits with explicit numeric ceilings, and treat the category score as a floor, never as the budget.

Keeping It Non-Flaky

A flaky performance gate gets disabled within a month — not cynicism, the observed lifecycle of every pipeline we inherit. The variance sources are known and mostly tractable:

  • Noisy neighbours. A parallel job saturates the shared runner. Pin the runner class, or run perf checks on a dedicated one.
  • Cold caches. First run after a build misses HTTP cache and service worker state. Warm it with a throwaway run, or assert consistently cold — never mix.
  • Third-party scripts. An ad or analytics vendor can deploy on their side and move your numbers without a single line of your code changing. Block non-essential third parties in lab runs; gate vendor bytes separately.
  • Extension and telemetry drift. Pin the Chrome channel and the Lighthouse version — an auto-bumped browser is a silent budget change.

Mechanically: median of three runs, numeric tolerances rather than booleans, two-tier severity. Start every assertion at warn for two weeks, collect the pass rate per route, and promote it to error only once it has passed ≥ 98% of runs on main. Warn first, block later — you earn the right to block with data instead of hope.

Quarantine the route, never the budget

When one legacy route cannot meet the global ceiling, add a per-route override that names the route and carries an expiry date. Weakening the global number because one page is bad is how budgets rot — the exception silently becomes the standard.

The same applies to a page under construction: assert against its own recorded baseline with a delta tolerance (“no more than 10% worse than main”), then promote it to the global budget once it lands. It is the kind of pipeline we take on when a team needs someone to build the gate properly rather than inherit ignored warnings.

What Changes After Six Months

Numbers from one engagement — a marketing site plus a logged-in dashboard, 400 sessions a day before the work, 3,000 after. Six months of enforced budgets, no heroics, no rewrite:

MetricBeforeAfter 6 monthsCI cost per PR
LCP (field p75)2.9s1.4s—
INP (field p75)340ms164ms—
CLS (field p75)0.190.02—
Bounce, marketing routes62%47%—
Perf-related rollbacks / quarter40—
Median PR wall time added0s+92s (Lighthouse) + 1.4s (byte gates)92s
False-positive block raten/a~1.5% of runs—

The honest headline is the second-to-last row: 92 seconds per pull request, plus a second and a half of static checks. At 15 merges a day that is under 25 minutes of CI time — the cost of one round of review — in exchange for four rollbacks a quarter turning into none. About one run in sixty needs a re-run, which is why the gate comments on the PR instead of failing silently, and why quarantining a route beats arguing with a red build at 6pm.

The softer change matters more: performance stopped being a discussion. Nobody negotiates a hero image size in review anymore, because the size negotiates first, in the pipeline, before a human is involved.

When Not to Gate

Enforcement has a blast radius — three places where it does more harm than good.

Early-stage prototypes. If product-market fit is still open, a 92-second perf gate on every PR is tax on learning. Ship the checks report-only so the numbers accumulate without blocking anyone, then promote them to errors once the team settles into a normal merge cadence. The data is useful later; the gate is not.

Third-party-dominated pages. When 60% of main-thread time belongs to a consent platform, a tag manager and two chat widgets, gating your own bundle is theatre — your code is not the constraint. Budget third-party bytes and blocking time on their own, renegotiate with the vendors, then bring back per-metric assertions. Gating your own 40kb while a 900kb embed runs free teaches the wrong lesson about where the cost is.

One-off campaign microsites. A page with a two-week lifespan and a fixed launch date gets one manual Lighthouse pass at delivery, not a permanent CI job. The gate protects assets that keep changing for years; a microsite is never touched after week three, and the pipeline you leave behind for it is one more thing to delete later.

The pattern underneath: enforce where regression is possible. Anything edited repeatedly deserves a failing build; anything else deserves a checkpoint. Confusing the two is how teams end up with red pipelines nobody reads — a wish, wearing a stricter outfit.

Put It to Work

Your Budget Is Only as Real as Its Last Failing Build.

If your pipeline reports green while Search Console reports yellow, that gap is the whole problem. Bring it to a principal engineer — thirty minutes is usually enough to say whether it is a measurement problem or an enforcement problem.