Every number on the Monday dashboard is green. Calls are up on last month. Tickets closed is the highest it has been since March. Pipeline created has a pleasant upward slope. The team looks at it for four minutes, nods, and moves to the next agenda item.
Meanwhile revenue is flat, the same three clients are unhappy for the same reasons they were unhappy in February, and nobody in the room can say whether the business is in better shape than a quarter ago.
This is not a tooling problem. The dashboard is working as designed, pulling accurate data from correctly configured systems. The numbers on it were simply never capable of telling you whether things were getting better. Most KPI sets are assembled out of whatever the systems already count, which is a very different exercise from deciding what result you want and working out how you would know if you were getting it.
Activity counts and progress measures are not the same thing
An activity count tells you how much of something happened. A progress measure tells you whether the thing you wanted is more true than it was last time you looked. They sit side by side on the same dashboard, which makes it easy to miss that only one is doing any work.
Take a sales team. "Calls made" is clean, reliable and easily collected. It also goes up when a rep has a slow week and burns through a list of people who will never buy, and in a quarter where nothing closes. The count is real and tells you almost nothing about whether the team is getting better at selling.
Now a different measure: the proportion of first conversations where the prospect has a problem you actually solve and a reason to solve it this year. That moves when targeting improves, when discovery questions get sharper, when marketing sends better leads, and not when someone pads their call log. It is harder to collect, because a person has to make a judgement after each conversation and record it honestly, and that difficulty is the whole point.
Delivery is where this gets expensive. Most services and implementation teams measure projects delivered, utilisation, or hours billed, because the timesheet system hands those over willingly. A project can be delivered on time, fully billed, at ninety percent utilisation, and the client can still be quietly deciding not to renew because what they bought is sitting unused. The useful measure is harder: how many clients are using what you built sixty days after handover. That moves when scoping improves and when you stop selling work that does not fit, and it sits still when the team simply works longer hours.
Support is the clearest case. Tickets closed is the default support metric almost everywhere and close to useless alone, because closing a ticket is entirely within the agent's control: they can close it because the problem is solved, because the customer stopped replying, or because the queue was long on Friday.
Ask instead whether the issue stayed fixed, counting the proportion that generate no repeat contact from the same customer about the same problem within thirty days. That drops when the team clears the queue by closing tickets nobody resolved. You have moved from the department's output to the customer's outcome, and those two diverge more often than anyone likes to admit.
Why everyone reaches for activity counts anyway
The people picking these measures are not lazy. Activity counts win for three good reasons, and understanding them is the only way to resist them.
First, the data is already there. Your CRM counts calls whether you ask or not. Building a measure from existing data takes an afternoon. Building one that requires a recorded judgement takes a process change, training, and a few months of people getting it wrong.
Second, they are always available, never having a month with no data, which matters when something has to go on screen every week.
Third, and this is the real reason, nobody argues about them. "We made 412 calls" is a fact that implicates no one. "Only 22 percent of our first conversations were with a qualified buyer" is also a fact, but it leads straight to an uncomfortable conversation about lead quality and whether two people are calling the wrong list. Measures that start arguments get quietly deprioritised.
There is a structural reason too. In a business of twenty or fifty people, the person building the dashboard is doing it between other jobs, with no training in measure design. They ask "what can we measure?" when the useful question is "what result are we trying to produce?". Stacey Barr, the Australian performance measurement specialist behind the PuMP methodology, has spent years making this argument: the common failure in KPI development is not the choice of metric, it is skipping the thinking about what the result actually is.
Start by writing down the result, in plain language
Before you choose any measure, write a sentence describing the result you want, in words a new employee would understand, with no numbers in it. Then rewrite it until it describes something you could observe.
Most goals fail on the first attempt. "Improve customer experience" is not observable, and two people can read it and picture different things. "Become the market leader in our region" hides three arguments about what leadership means.
The way through is to ask what would be visibly different if the result were achieved. Not what you would do, what would be different. If customer experience improved, customers might contact you less about problems, stay longer, or stop splitting their work between you and someone else. Each is observable, and each points at a different measure, which is why you choose before you measure.
Work one through. "Improve onboarding" is a direction, not a result. Ask what better onboarding looks like from the client's side and you land closer to: new clients use the product for real work without us holding their hand. Sharpen it once more. What does "real work" mean here? Perhaps they have run a full month end cycle in the system, or three or more of their people log in weekly.
Whatever it is, write it down. You now have a statement that is true or false for any given client, so it can be counted across clients, so you have a measure without having gone looking for one. That is the correct order, and almost nobody follows it.
Testing a measure before you commit to it
Put a candidate through a few questions before it goes near a dashboard.
The first is whether the number moves when the result improves and holds still when it does not. Test it against recent history: take a month where the result genuinely improved and ask whether this measure would have shown it, then take a month where nothing improved but everyone was busy and ask whether it would have stayed flat. If it rises in both, it is an activity count wearing a KPI badge.
The second is how someone would game it. Any measure can be gamed by someone determined enough, so the useful version of the question is about effort and visibility. Tickets closed can be gamed in seconds with no trace. Repeat contacts within thirty days can only be gamed by talking the customer out of contacting you, which is hard and tends to get noticed. Prefer measures where the cheapest way to move the number is to do the work, and be honest that if you attach a bonus, someone will find the gap between what you asked for and what you meant.
The third is who would have to behave differently for the number to move. If you cannot name a person or a team, the measure is not connected to anything you control and will function as weather commentary. If the answer is "everyone", it is too broad to act on. Qualified conversations move because specific reps prospect differently. Market share moves because of a dozen things, most outside your influence this quarter.
The fourth separates measures from decoration: what decision would a bad reading trigger? If the honest answer is "we would discuss it", the measure is not earning its place. If it is "we would stop taking that category of work", or "we would change the qualification criteria and retrain the team", the measure is load bearing. Decide that response while nobody is defending a bad month.
Leading, lagging, and the comfortable lie in between
The leading and lagging distinction is taught everywhere and misapplied almost everywhere. The textbook version is fine: a lagging measure tells you what happened, a leading one gives early warning of what is coming.
The problem is the application. Almost every KPI set in a growing business leans lagging, because lagging measures are what the finance system produces. Revenue, margin, churn, cash collected. You need them, and they all describe decisions made months ago.
So teams go hunting for leading measures and reach for the nearest activity count. Calls made gets labelled leading, and so does proposals sent. The label sticks because the activity happens before the revenue, and that feels like enough.
It is not. A leading measure has to predict the result, not merely precede it. Calls made only predicts revenue if call volume is your actual constraint, and usually it is not. If conversion from first conversation to opportunity is poor, more calls produce more of the same while the so called leading indicator points cheerfully upwards.
The test is whether you have a real reason to believe the causal link holds in your business, not in general. If qualified conversations rose last quarter, did closed revenue rise this quarter? If you have never checked, you do not have a leading measure, you have a hypothesis you decided to treat as fact. This is one of the quieter ways plans come apart: the standard reference on execution failure remains Sull, Homkes and Sull's 2015 Harvard Business Review article "Why Strategy Execution Unravels, and What to Do About It", which finds that plans break down in execution rather than in strategy formulation.
Keep the lagging measures, because they tell you the truth. Add one or two genuine leading measures per goal, chosen because you have a defensible reason to think they predict the outcome, then check after two quarters and drop what did not hold.
Targets pulled out of the air ruin good measures
A well designed measure can be undermined in one step by an arbitrary target. Someone says "let's make it 20 percent", everyone nods, and a number that could have told you something becomes one that generates excuses.
Arbitrary targets do two kinds of damage. Wildly optimistic ones stop being taken seriously within six weeks, and the measure goes the same way. Accidentally easy ones get hit in month two and give you nothing for the rest of the year. There is a subtler problem too: when a target is set before anyone knows the current value, you cannot tell whether missing it means performance is poor or the target was fiction, so every review becomes a debate about the target rather than the work.
The fix is unglamorous. Measure first, set the target second. Run the measure long enough to establish a baseline, at least three months and longer if your business has seasonal swing. You want the normal range and how much it bounces, because without that you will react to noise.
Then set a target you can defend, meaning you can say where the number came from: what the process is capable of once the known bottleneck is fixed, what the financial plan requires, or the level a comparable part of the business already achieves. "It felt about right" is not an argument. A new measure will produce no target in its first quarter, and a baseline with no target beats a target with no baseline.
How many KPIs you should actually have
Far fewer than you have now. For a business of ten to a hundred people, five to ten measures at company level is plenty, with each team owning a small number beneath that.
The argument is about attention, not tidiness. A dashboard with thirty numbers gets scanned, not read. Nobody holds thirty trends in their head, so people look at whatever is red or concerns them personally and the rest becomes wallpaper. A year later half the charts are broken and nobody has noticed.
Good measures also cost something to produce, a cost worth paying for six and impossible across thirty, which is why large measure sets drift back towards whatever the systems count automatically.
Every team does need visibility of its own numbers, which argues for layered measures rather than one sprawling dashboard. The company view holds the handful that would make leadership change course. Team views can hold operational detail, including activity counts, which are fine for managing capacity as long as nobody mistakes them for progress.
The link between team measures and company goals is where this usually breaks. Research on OKR adoption in 2026 found that around 65% of teams say their OKRs are not clearly linked to company goals. Without that link, team measures multiply without constraint, because nothing above them is deciding what matters.
Nobody ever deletes a KPI
Ask any operations lead when they last removed a measure from the dashboard. The answer is usually never, or once during a rebuild when everything went at the same time.
Measures accumulate because adding one is easy and removing one feels like a judgement on whoever added it. There is always a plausible reason to keep a number: we might need it, someone looks at it. So the dashboard grows, the signal degrades, and eventually the whole thing gets replaced.
The fix is a standing question in your quarterly review, applied to every measure without exception: has this number changed a decision in the last two quarters? Not has it been interesting. Has a decision been made differently because of what it showed.
If the answer is no, retire the measure, or rebuild it if the result it pointed at still matters. What you may not do is leave it alone and agree to watch it more carefully next quarter. That is how you got here.
A measure with no owner and no forum is decoration
You can do all of the above correctly and still change nothing if two things are missing: a named owner, and a forum where a bad number has to produce a decision.
Ownership means a person, not a team and not a role. "Sales owns conversion rate" means nobody owns it. "Priya owns qualified conversations" means a specific person is expected to know why the number moved, what she is doing about it, and what she needs from others. The owner rarely controls every input, but has to be accountable for understanding and acting.
The forum matters as much. A measure that only appears in a monthly report is one nobody answers for. It needs a regular meeting where the owner presents the number, with the standing expectation that a bad reading produces a decision, not an explanation.
That distinction is the hard part. Explanations are comfortable and often true: two deals slipped, someone was on leave, it is seasonal. The explanation can be correct and still be a way of avoiding the decision. Accept it, then ask what changes as a result. Sometimes the answer is nothing, because the variance really was noise, but that has to be a decision someone makes rather than a place the conversation quietly stops.
This gap shows up in the adoption data. The same 2026 OKR benchmarking found 71% of companies using OKRs say they have not yet mastered the process, which says less about the framework than about the distance between setting measures and running them. Keeping each measure attached to a named owner, the goal it serves and the meeting where it gets reviewed is the connection Empiraa GPS is built to hold. Platform or spreadsheet, the requirement is the same: a name against every measure and a recurring slot where somebody speaks to it.
Running one goal all the way through
Take a services firm of about thirty people whose priority this year is reducing client churn.
Write the result in plain language first, with no numbers. The initial draft, "clients stay with us longer", is directionally right and not observable at a point where you could act on it. Rewrite it around what would visibly change: clients keep engaging us for new work after their first project finishes, instead of going quiet.
That is specific enough to measure. The candidate: the proportion of clients whose first project completed in a quarter who commission a second piece of work within six months. Run it through the tests.
- Does it move with the result? Yes, and not because anyone worked more hours.
- Can it be gamed? Somewhat. A delivery lead could push a token follow up engagement, so the definition should require a second project of meaningful size.
- Who has to behave differently? The delivery leads running first projects, and whoever decides how handover is done.
- What does a bad reading trigger? A review of the last five first projects, looking at how scope was set and what happened in the four weeks after handover. Then the baseline. Pulling the last four quarters of first projects gives you the current rate without waiting six months. Say it sits in the low thirties and bounces by several points each quarter. That bounce matters: it tells you a three point improvement means nothing.
Only now does the target get set, with an argument attached. If the financial plan needs a certain amount of repeat revenue, work back to the rate that delivers it. If one team already runs well above the firm average, that is evidence the higher rate is achievable. Either way the target survives being questioned.
Finally the plumbing: one named owner, a slot in the monthly leadership meeting, and an agreed response if the number comes in low. That is the whole system for one goal, and it took a couple of hours rather than a software project.
The point
The test for any measure on your dashboard is simple and slightly brutal. If the number goes up, does that mean the business is better off? If you have to think about it, or the honest answer is "it depends what caused it", you are looking at activity.
Pick one measure and run it through these questions. Most will survive as context and fail as a KPI, and the work from there is to go back to the result you wanted and derive something better. That is slower than adding a chart, and it is the only version that ends with a number capable of changing what anyone does on Monday.


