← All insights Program Health

The KPIs that keep a program honest

A status color hides. A number cannot. The measurement system that keeps a large ERP, CRM, or service management program honest from kickoff through hypercare, and the go-live metrics that tell you the truth before the org does.

I have watched a lot of programs report themselves healthy right up until the week they were not. The status was green. The milestones were "on track." The steering deck had a comforting row of amber-trending-green. Then a data conversion missed its window, or the first live payroll ran wrong, or a division quietly went back to spreadsheets, and everyone in the room discovered the program had been red for 2 quarters. Nobody lied. The reporting just drifted, the way reporting does when it is a color instead of a number.

There is research behind that drift. A study in MIT Sloan Management Review found that status reports on troubled IT projects are biased about 60% of the time, and the bias runs roughly two to one toward optimism. Worse, the executives receiving those reports mostly could not tell. A green light told them nothing, and they had no way to know it told them nothing.

Green tells you nothing, and you cannot tell that it tells you nothingA study in MIT Sloan Management Review found status reports on troubled IT projects are biased about 60% of the time, and the bias runs roughly two to one toward optimism. The executives receiving those reports mostly could not detect it. That is the case for measuring instead of coloring.

This is the case for measuring instead of coloring. Not more reporting. Better numbers, chosen in advance, produced on a cadence, shown as trends, and owned by name. A KPI framework earns its place by doing more than filling a dashboard. It is the instrument that lets a program leader, a sponsor, or an independent set of eyes see the truth early enough to still do something about it.

None of what follows is specific to one platform. Whether you are standing up Workday, SAP S/4HANA, Oracle Fusion, Salesforce, or ServiceNow, the physics are the same. The system changes. The numbers that predict success do not.

Why the color lies, and the number does not

Start with why this matters at the scale most of these programs run.

A McKinsey study with the University of Oxford, drawn from more than 5,400 large IT projects, found that on average they run 45% over budget and deliver 56% less value than predicted. Read that second number twice. More than half the value, gone, on the average large project. And these were the ordinary ones, not the disasters. The disasters, the "black swans" that blow past 200% overrun, were a separate 17% on top.

45%Average budget overrun across more than 5,400 large IT projects. These were the ordinary ones, not the disasters.McKinsey with University of Oxford
56%Less value delivered than predicted, on the same average project. Read that number twice before you accept a business case.McKinsey with University of Oxford
17%Were black swans on top of that, blowing past 200% overrun. The tail is not rare enough to plan around ignoring.McKinsey with University of Oxford

The pattern has not aged out. Gartner projects that by 2027, more than 70% of recently implemented ERP initiatives will fail to fully meet their original business goals. Note the precise wording. Not "fail" in the sense of collapse. Fail to meet the goals that justified the spend. Panorama Consulting's 2026 ERP Report found more than a quarter of implementations went over budget and nearly a quarter over schedule, with the leading schedule cause being organizational, governance gaps, change resistance, delayed sign-offs, and unplanned change orders. Not the software.

70%Of recently implemented ERP initiatives will fail to fully meet their original business goals by 2027. Not collapse. Fail to meet the goals that justified the spend.Gartner projection
25%Went over budget, and nearly a quarter over schedule, in the 2026 ERP survey.Panorama Consulting
GovernanceThe leading schedule cause: governance gaps, change resistance, delayed sign-offs, unplanned change orders. Every one of those is measurable before it becomes a headline.Panorama Consulting

That last point is the whole argument for a real KPI system. These programs do not fail on technology. They fail on decisions made late, data trusted too early, adoption assumed rather than measured, and vendor performance nobody was holding to a number. Every one of those is measurable before it becomes a headline. A good framework measures them on purpose.

Five views, each answering a different question

I run program health across five views. Each one answers a different question, each has a small set of KPIs, and each KPI has a formula, a red-flag threshold, an owner, and a cadence. The point of five views is that no single number can hide a program. A green schedule cannot cover a red budget. Strong adoption cannot excuse dirty data. You have to look from all five directions at once.

KPIHow it is calculatedRed flagOwnerCadence
01  Delivery and governanceis the program moving, and deciding?
Schedule Performance IndexEarned work divided by planned work< 0.90Program leadWeekly
Milestone hit rateAgainst baseline dates, not against dates quietly reset last month2 consecutive missesProgram leadWeekly
Decision agingCount open past 14 days, plus average days to decide> 14 days openPMOWeekly
02  Financialis spend buying real progress?
Cost Performance IndexEarned value over actual cost< 0.90Finance leadMonthly
Burn against accepted deliverablesBudget spent against work formally accepted, not against the calendarburn leads by 10 ptsFinance leadMonthly
03  People and adoptionis the organization actually coming with you?
Training completion by roleWeighted to go-live-critical roles, never a flattering all-user average< 90% critical rolesChange leadWeekly
Super-user coverageNamed, trained super users per site or departmentany site at zeroChange leadBiweekly
Wave readinessAgainst gate criteria at set intervals before each cutoveramber at final gateChange leadPer gate
Utilization and proficiencyPost go-live. Completion measures exposure, this measures capabilityflat or fallingChange leadPost go-live
04  Data and qualityis what you are moving fit to run a business on?
Duplicate rate, priority objectsOn the objects that carry compliance exposure, worker and pay first> 2% at final mockBusiness data ownerPer mock
Rehearsal varianceAt least three full mock loads. Finance reconciles control totals to the dollar> 2% final mockBusiness data ownerPer mock
UAT exit qualityDefects found in user testing are the expensive kind. These are exit criteria, not aspirationsany open criticalTest leadPer gate
05  SI performanceis your integrator earning the invoice?
First-pass acceptanceShare of deliverables accepted without reworkfalling trendVendor managerMonthly
Change-order velocitySI-originated changes per month, in count and dollars. Watch the acceleration, not the levelacceleratingVendor managerMonthly
Key-person turnoverAgainst the roles named in the contractany named roleVendor managerMonthly
Billed against acceptedWhen billing runs ahead of acceptance, you are financing the integrator's optimismbilled > acceptedVendor managerMonthly
Two thresholds worth knowing where they come from. The 0.90 lines on both performance indices and the 20% rule (past roughly one fifth complete, cumulative CPI rarely recovers) come from decades of earned value research on large government programs. The SI thresholds do not come from a study, because nobody benchmarks first-pass acceptance or consultant turnover on implementations. That absence is the argument for instrumenting it yourself, not a gap in the research. Treat those four as a starting position and tighten them to your context.

Delivery and governance asks whether the program is moving and deciding. The workhorse is the Schedule Performance Index from earned value, earned work divided by planned work, with 0.90 as the line you do not want to cross. Alongside it: milestone hit rate against baseline dates, not against the dates you quietly reset last month. And the one I trust most, decision aging. Count the decisions open past 14 days and track the average days to decide. The Standish Group's work on decision latency makes the case that the interval matters more than the elegance of the decision. Programs that decide in days beat programs that decide in weeks. An aged-decision list, reviewed oldest first, predicts trouble earlier than any schedule metric I know, because a program that cannot decide is a program that has already begun to slip. It just has not shown up in the plan yet.

10%Of "active" worker records on a recent conversion were duplicates or carried termination dates contradicting payroll history. That is compliance exposure riding into the new system.Field, anonymized
400 → 90"Critical" legacy reports collapsed to 90 once checked against actual run logs. Nobody asked for the missing 310 again.Field, anonymized
3Full mock loads, minimum. By the final one you want under 2% error overall, with Finance reconciling control totals to the dollar.Standing rule

Financial asks whether spend is buying real progress. The Cost Performance Index, earned value over actual cost, again with 0.90 as the warning line. There is a hard lesson buried in the earned value research from decades of large government programs. Once a program passes roughly 20% complete, the cumulative CPI rarely recovers. If you are over budget at the one-fifth mark, the honest base case is that you finish over budget, and the intervention has to happen now, not at a future point where the math has already closed. The number I watch even more closely is burn against accepted deliverables. Not budget spent against the calendar. Budget spent against work the client has formally accepted. When burn leads acceptance by ten points or more, you are paying for progress that has not been proven. That gap is the financial watermelon.

The twenty percent ruleDecades of earned value research on large government programs found that once a program passes roughly 20% complete, the cumulative Cost Performance Index rarely recovers. If you are over budget at the one fifth mark, the honest base case is that you finish over budget. The intervention has to happen now, not at a future point where the math has already closed.

People and adoption asks whether the organization is actually coming with you. This is the view programs fund last and regret first. Prosci's benchmarking is blunt about the stakes: programs with excellent change management run on or ahead of schedule about five times as often as programs with poor change management, and hold budget about one and a half times as often. All of that upside sits in the one workstream that gets raided for budget when things get tight. Yet Panorama's data shows fewer than a quarter of organizations put an intense focus on change management. The KPIs here are training completion by role, weighted toward the go-live-critical roles rather than a flattering all-user average, super-user coverage per site, and wave readiness against gate criteria at set intervals before each cutover. Track completion for what it is, though. It measures exposure, not capability. Utilization and proficiency after go-live are the numbers that tell you whether the training took.

5xPrograms with excellent change management run on or ahead of schedule about five times as often as programs with poor change management.Prosci benchmarking
1.5xThe same programs hold budget roughly one and a half times as often. All of that upside sits in the workstream that gets raided when things get tight.Prosci benchmarking
<25%Of organizations put an intense focus on change management, which is why the upside above stays theoretical on most programs.Panorama Consulting

Data and quality asks whether what you are moving is fit to run a business on. Two field numbers I have seen hold across programs: on a recent conversion, about 10% of the "active" worker records were duplicates or carried termination dates that contradicted payroll history. Call that a spreadsheet problem and you have mislabelled it. That is payroll-tax and compliance exposure riding into the new system. On another, 400 "critical" legacy reports collapsed to 90 once we checked them against the actual run logs, and nobody asked for the missing 310 again. The KPIs are duplicate rate on priority objects, migration rehearsal variance, and UAT exit quality. My rule on rehearsals: run at least three full mock loads, and by the final one you want under 2% error overall with Finance reconciling control totals to the dollar. Data problems found in user testing are the expensive kind. These thresholds are exit criteria, not aspirations.

SI performance asks whether your integrator is earning the invoice. This is the view that never appears on the vendor's own status slide, which is exactly why it belongs on yours. Be honest about the evidence here: there is no published industry benchmark for first-pass deliverable acceptance, rework rates, or consultant turnover on implementations. No analyst tracks it. That absence is the argument, not a gap in the research. Because no one benchmarks it, the client has to instrument it. Track first-pass acceptance, the share of deliverables accepted without rework. Track change-order velocity, SI-originated changes per month in count and dollars, watching for the acceleration that means requirements gaps are being monetized. Track key-person turnover against the named roles in the contract. And track billed against accepted, because when billing runs ahead of acceptance, you are financing the SI's optimism. The thresholds I use come from field judgment across many programs, not from a study. Treat them as a starting position and tighten them to your context.

The Program KPI Wall Chart: 29 KPIs across six views with red-flag thresholds, delivery and governance, financial, people and adoption, data and quality, SI performance, and go-live and hypercare
The whole instrument on one page. 29 KPIs, six views, every red flag. Print it and score your program against it.

Most steering committees will not hold 29 numbers in their heads

Twenty-nine KPIs is the full instrument. Most steering committees will not hold that many numbers in their head, and they do not need to. If I had to put five on one slide and make them the spine of every meeting, it would be these:

  • Schedule Performance Index. Is the program moving.
  • Burn against accepted deliverables. Is spend buying proven progress.
  • Decision aging. Is the program deciding, or quietly stalling.
  • Migration rehearsal variance. Is the data fit to go live on.
  • First-pass SI acceptance. Is the integrator earning the invoice.

Five numbers, one from each view. There is no place for a watermelon to hide across all five at once.

The other question executives ask is which numbers matter right now. The honest answer is that the instrument changes by phase: the KPIs that run a design phase are not the ones that run a cutover, and each phase has a gate that ends it.

Which KPIs matter right now: the measurement system by phase from Phase Zero through Design, Build, Test and Data, Cutover and Go-Live, to Hypercare and BAU, with the gate that ends each phase
The measurement system by phase. Each station has its KPIs, and each phase has the gate that ends it.

Go-live and hypercare: when the KPIs change shape

Everything above measures whether you will get to go-live in one piece. The day you cut over, the questions change, and so do the metrics. Delivery KPIs go quiet. Operational KPIs take their place, and they move by the hour instead of by the week. This is the phase where a program's real health becomes visible faster than any status deck can keep up with. It is also where I have spent some of the most useful hours of my career, running a command center for a $25M Workday big-bang across HCM, Finance, Supply Chain, and Payroll.

The mental shift is this. Before go-live, you are asking "are we ready." After go-live, you are asking "is it working, right now, for real people doing real work." The command center is where you answer that. It runs on a daily scorecard, triaged several times a day in the first week and daily after that, on one queue with one set of priorities and one version of the truth.

Incident volume and aging. The spine of the command center is the ticket picture, pulled from your service management platform, ideally the one you already run the enterprise on. Total volume by severity, resolved, open, and then the aging buckets that matter most: open past 24 hours, past 72 hours, past 7 days. The 7-day band is the one exit gates watch. Break it down by site and by functional area so a single struggling location or module shows up as a shape, not a surprise. On the big-bang I mentioned, the command center opened first for time and scheduling, then expanded to HR, payroll, finance, and supply chain as each went live, and the ticket tags by functional area told us where to move people days before anyone raised a hand.

Before go-live you watchAfter go-live it becomes
Schedule and milestone performanceTicket volume and its trend, by functional area, against the demobilization curve
Training completionUtilization and proficiency in the system. Completion measured exposure. This measures capability
Rehearsal varianceTransaction throughput against the pre-go-live baseline. Far below normal means users found a way around you
Defect counts at the gateTime to resolve by severity, and the reopen rate, which is the tell for a superficial fix
Readiness against gate criteriaWhether the business ran a real cycle clean: payroll twice, a period closed on time and reconciled
The instrument changes shape at go-live, and most programs forget to change it. They keep reporting delivery metrics into a period where delivery is over and stabilization is the only question. Hand off to business as usual on the proof that a real cycle ran clean, not on ticket counts alone. A compensating control that runs on heroics is deferred failure with a nicer name.

Resolution quality, not raw speed. Mean time to resolve is the obvious one, total resolution time over tickets resolved. The one people forget is reopen rate, reopened tickets over resolved tickets. Below 5% is a healthy signal that fixes are sticking. A rising reopen rate during hypercare is worse news than raw volume, because it means you are closing tickets that reopen, and the backlog you think is shrinking is about to come back. Set your severity SLAs tighter for hypercare than for steady state, and separate time-to-acknowledge from time-to-resolve so a fast-triaging command center is not judged as if every fix were instant.

Adoption, measured for real. At go-live, adoption stops being a training statistic and becomes a login and usage fact. Active users, web and mobile, by role and by site. Not "logged in once," but performing real work in the system. Prosci is refreshingly honest that there is no universal benchmark for a "good" adoption rate, so be skeptical of anyone who quotes you a magic 95%. What matters is the trend and the outliers. A department far below the others is a change-management signal, not a software defect, and it is cheaper to fix in week one than in month three. Watch transaction throughput against the pre-go-live baseline too. If volume is far below normal, users have found a way around the system. They are working around you.

Transaction success, by process, agnostic to the platform. This is where go-live either proves itself or does not, and the metrics translate across every system because they follow the business process, not the software.

  • Procure to pay. The three-way match exception rate, and the touchless or straight-through rate on invoices. Best-in-class AP teams run exception rates near 9% against an industry average around 22%, and Ardent Partners pegs best-in-class touchless processing at roughly half of all invoices while the average buyer sits closer to a quarter. At go-live, watch the clean-match rate trend. A sharp drop below the legacy baseline usually means a match-tolerance or master-data problem, not a user problem.
  • Record to report. The first month-end close on the published calendar. Expect it to run long. APQC's benchmarking puts the median monthly close at about 6.4 calendar days, with top performers under 4.8 and laggards past 10, so a first close that overruns is normal. The success metric lives across the first three closes, in a declining pain curve, rather than in the first one.
  • Order to cash. Order-entry accuracy, invoice generation success, and cash-application match rate. Early breakage here hits revenue and customers faster than anywhere else, so it gets watched first.
  • Payroll. The least forgiving event in the entire program. The two numbers that matter are clean net-pay match rate across the parallel runs, where you want as close to every employee inside tolerance as you can get across at least two consecutive clean cycles, and the on-time, accurate first live payroll. A 99% accuracy target across all pay elements is the working standard most teams hold, and any single employee's net pay moving more than about 10% period over period earns a manual review before the run releases. There is no metric on the program more visible to more people than the first paycheck being right.

Defect leakage and data integrity. Escaped defects, defects found in production over total defects, is the cleanest read on how good your testing actually was. Under 5% leakage is the working target, and a spike in high-severity escapes into hypercare tells you the problem was in UAT scope or test environments, not in your support team. Underneath it all, keep reconciling. Post-migration opening balances must equal legacy closing balances, to the penny on financials. Suspense and clearing accounts used during migration have to return to zero, and climbing unposted or suspense transactions are the earliest hard signal of a mapping or interface defect. Watch them daily, not at month-end, when it is too late to do anything but explain.

Exit on conditions, not on the calendar. The most important discipline in hypercare is knowing when it ends, and the answer is never a date. Hypercare ends when the operating model can absorb issues through normal support, not when the calendar says the 4 weeks are up. Define the exit criteria before you go live and make them measurable: severity-1 volume back to the steady-state baseline, no release-related incidents older than 7 days, several consecutive clean days with no new critical issues, the first close completed, and, the one people skip, the ongoing support team demonstrating it can resolve within SLA on its own. Gate the handoff to business-as-usual on that last proof, not on ticket counts alone. A compensating control that runs on heroics is just deferred failure with a nicer name.

When does hypercare end: a stabilization curve of severity-1 volume decaying to the steady-state baseline with the exit gate at the intersection and four exit conditions
Hypercare ends where the curve meets the baseline and four conditions hold, never on a calendar date.

The operating rhythm that makes the numbers work

Metrics without a meeting are a spreadsheet nobody opens. The framework earns its keep in a 60-minute steering meeting run on a few unbreakable rules.

Trends beat snapshots. Every number is shown against the prior two meetings, never as a single point-in-time color. Decisions are logged with an age and reviewed oldest first, because the aged-decision list is the meeting, not a footnote to it. Every red or amber arrives with a written ask, the specific decision or resource that turns it, and a red without an ask goes back to its owner. The SI presents in the same format, their deliverable acceptance and change-order numbers sitting right next to yours. And no slide restates status. If it is not a number, a decision, or a risk, it is not in the room.

The five-number steering slide, a worked example with sample data: SPI, burn vs accepted, decision aging in red with its written ask, rehearsal variance, and first-pass acceptance, each shown against the prior two meetings
A worked example with sample data. Five numbers, trends against the prior two meetings, and one red arriving with its ask.

Each view has a question that opens the conversation, and the question is the fastest way to find the watermelon. For delivery: which decisions older than 14 days sit on the critical path, and who closes them this week. For financial: show me burn against accepted deliverables, not against the calendar. For adoption: which go-live-critical roles are not at 100% training, by site. For data: what was the variance on the last full mock load, and who dispositioned each item. For the SI: what share of deliverables passed first time, and what has been billed that we have not accepted. Every one of those questions is designed to be un-colorable. You either have the number or you do not, and not having it is itself the finding.

Anything you cannot produce a number for within 48 hours is itself a finding.

A program that measures cannot be surprised the way the colour-coded ones are

A program that measures cannot be surprised the way the color-coded ones are. The red shows up while there is still budget, time, and scope to move it, not the week before go-live when all three have run out. The truth arrives early enough to be useful, which is the only kind of truth that matters on a program betting tens of millions on getting this right.

The systems will keep changing. The vendors will keep shipping. And a status deck will still, quarter after quarter, tell a room full of executives that everything is fine. Build the instrument that lets you check. Measure the program, not the color, and you will know which side of the Gartner number you are on long before go-live tells you for free.

Get the framework: The ERP Program KPI Command Deck, the full 29-KPI catalog with formulas, red-flag thresholds, owners, and the 60-minute meeting agenda. Free, no form, along with the rest of the field guides in the toolkit.

Related reading: The steering committee that earns its hour · The watermelon status report · Running a go-live command center that works · Hypercare. Related case study: An end-to-end enterprise Workday program, where this command-center scorecard ran a multi-wave go-live.

Let's talk

Put independent eyes on your program.

If you're betting tens of millions on an ERP program, a candid second opinion is the cheapest insurance you'll buy.

Field notes

Get the next lesson in your inbox.

One hard-won program lesson at a time. No cadence promises, no spam.