The thing that kills small consulting practices is almost never a shortage of work. It is six engagements running at once, each one slightly over scope, none of them individually alarming, and a principal who has not looked properly at the margin on any of them since March.
Consultants are unusually good at the work and unusually bad at the operating system around the work. This is not a character flaw. It is a structural consequence of the model. Every hour spent on practice management is an hour not billed, so the practice management gets done at nine on a Sunday night, badly, or not at all.
The result is a business that looks healthy from the outside, with a full calendar and good client relationships, while quietly losing money on two of its six engagements and having no mechanism that would reveal which two.
This article is about the operating layer. What actually goes wrong when you run several client engagements concurrently, how to see it early, and what to change.
The two things that erode margin
Almost all margin loss in a small consulting practice comes from one of two sources, and they behave very differently.
The first is scope creep, which is well understood and still constant. Project Management Institute research has consistently found that around half of projects experience scope creep, with meaningful cost overruns attached. In consulting specifically, the mechanism is usually not a dramatic renegotiation. It is a sequence of small yeses.
The client asks whether you could also look at the sales data while you are in there. You could, and it takes two hours, and saying no over two hours would be graceless. Three weeks later they ask for an additional workshop with a team you had not planned to include. Then the board wants a version of the deck. Then someone wants a follow-up conversation to walk through the recommendations again.
None of those requests is unreasonable. Each one, on its own, is the kind of thing that builds the relationship. Collectively they can add twenty or thirty per cent to the delivery effort on a fixed-fee engagement, and because they arrived one at a time, nobody ever had a conversation about the cumulative effect.
The second source is quieter and worse: utilisation drift. Industry benchmarking from Deltek's professional services research has shown billable utilisation across the sector declining in recent years, from levels above seventy per cent earlier in the decade into the high sixties. For a small practice the driver of that decline is usually not idle time between projects. It is the growing volume of unbilled work attached to projects that are running.
Status calls that were not scoped. Rework because a stakeholder was not consulted early enough. Reading a document that a client sent at eleven at night. Preparing a summary because someone missed the session. Chasing a client for an input that has been outstanding for two weeks. None of this appears on a timesheet in most small practices, because most small practices do not keep timesheets on fixed-fee work.
That is the trap. If you bill a fixed fee, you have no natural mechanism that counts the hours, which means you have no way of knowing that engagement four is now taking twice what you priced it at.
Why concurrency makes both of these worse
Running one engagement at a time, most consultants would notice a problem quickly. Running six, the signal gets buried.
The main reason is that context switching has a cost that is real and largely invisible. Moving between six client contexts in a week means six separate sets of names, histories, politics and half-finished thinking to reload. The reload time does not feel like work, so it does not get counted, but it is a substantial share of the day.
The second reason is that concurrency lets you compensate. When one engagement runs hot, you borrow time from the others. Nobody complains immediately, because the other engagements are not visibly starved yet, they are just slightly slower. Two months later, three engagements are behind and it is no longer possible to identify which one caused it.
The third is that averaging hides the loser. If the practice is profitable overall, the two engagements running at a loss are subsidised by the four that are fine, and the aggregate number looks acceptable. You cannot fix what the average is concealing, and small practices tend to look at the aggregate because it is the number the accountant produces.
Seeing the problem early
The fix here is not complicated but it is unpopular, because it involves tracking time on fixed-fee work.
The objection is understandable. Time tracking feels like a return to the hourly model you left, and it adds admin to a day that has no room in it. But the purpose is different. You are not tracking time to bill it. You are tracking it to price the next engagement properly and to notice when this one has gone wrong.
The lightweight version is enough. You do not need six-minute increments and a code hierarchy. You need a rough daily note of how many hours went to each client, kept honestly, taking about two minutes. At the end of a month you divide the fee by the hours and you have an effective rate per engagement.
That single number does more work than any other metric in a small practice. It tells you which clients are actually profitable rather than which ones feel pleasant. It frequently produces an unwelcome result, because the client who is easiest to work with is sometimes the one absorbing the most unbilled time, and the demanding client with clear requirements is sometimes the most profitable.
It also gives you a defensible basis for repricing at renewal. "Our rates have increased" is a weak conversation. "This engagement has consistently required about forty per cent more input than we scoped, here is what that looked like, and here is what the next twelve months should be priced at" is a different one, and it usually goes better than consultants expect.
The threshold worth setting in advance is a percentage. Decide that any engagement running more than twenty-five per cent over its scoped effort triggers a conversation with the client. Writing it down beforehand is what makes it happen, because in the moment there is always a reason to leave it another fortnight.
Handling scope requests without damaging the relationship
The advice to simply say no is useless, and consultants know it, which is why they ignore it. Client relationships in a small practice are the entire asset. Refusing a small request to protect a margin you have not measured is a bad trade.
What works better is a two-tier response that most experienced consultants arrive at eventually.
Small requests that take under an hour and sit near the existing work get absorbed without comment. They cost little, they build goodwill, and tracking them creates more friction than they are worth. This is a deliberate investment, not an accident.
Anything larger gets named as additional scope in the moment, and named neutrally. Not as a complaint or a negotiation, but as information: this sits outside what we scoped, it is a genuinely good idea, here is roughly what it would take, do you want to add it now or hold it for the next phase.
The critical part is doing it in the moment rather than at the end. Raising scope at the point of request is a normal professional conversation. Raising it three months later, when the work has already been done and you are trying to explain why the invoice is larger or the timeline slipped, is a dispute.
Most clients respond well to this, and the reason is worth understanding. A client who asks for something extra is usually not trying to extract free work. They genuinely do not know what it costs, because the internal shape of your delivery is invisible to them. Telling them is helpful rather than difficult. The consultants who struggle with this conversation are usually the ones who have avoided it long enough that the accumulated gap is now too large to raise casually.
There is one further habit that prevents most of it: writing down what is out of scope, not just what is in. Statements of work almost always list deliverables and almost never list exclusions. Two or three lines naming the obvious adjacent things you are not doing removes the ambiguity that scope creep grows in.
How the pricing model changes the problem
The shape of the margin problem depends heavily on how you charge, and it is worth being clear about which version you are dealing with.
On time and materials, scope creep is largely self-correcting. Extra work produces extra invoice. The risk shifts to the client, which is why clients dislike the model, and the constraint on your practice becomes capacity rather than margin. The problem in this model is different: you are capped at hours multiplied by rate, and growth requires either raising the rate or hiring, both of which are slow.
On fixed fee, all the scope risk sits with you. This is the model where the failures described above do the most damage, because the revenue is locked at the start and every additional hour comes directly out of margin without any signal that it is happening. It is also the model clients most often prefer, and the one most consultants default to for exactly that reason.
On retainer, the failure is subtler and takes longer to notice. A monthly retainer priced against an expected level of demand will drift as the client gets comfortable, and the drift is gradual enough that no single month feels wrong. Retainers that have been running unchanged for two years are almost always underpriced relative to what they now involve, and the longer they run the harder the conversation becomes.
Outcome or performance-based arrangements have grown in popularity among independent consultants, typically structured as a smaller implementation fee plus a bonus tied to a defined milestone. They can work well where the outcome is genuinely measurable and largely within your influence. They go badly where the outcome depends on client execution you do not control, which describes most strategy work. If you use them, the milestone needs to be something you can actually move, and the base fee needs to cover your costs on its own.
The practical point across all four is that the pricing model determines which number you need to watch. On fixed fee, watch effective hourly rate. On retainer, watch effective rate over time rather than in a single month, because the drift only shows up across a trend. On time and materials, watch utilisation. Watching the wrong number for your model is how practices stay confident while margin erodes.
The fractional model changes the maths again
A growing share of independent consultants now work on fractional arrangements, holding an ongoing part-time executive role inside a client business rather than delivering discrete projects. Market analysis through 2026 has consistently described this as one of the faster-growing segments of the independent consulting market, though the size estimates published vary widely and most come from firms with a commercial interest in the category, so they are better read as a direction than a measurement.
What is clear from the practitioner side is that fractional work has a different operating profile to project work, and the differences matter for anyone running both.
The scope boundary is harder to hold. A project has deliverables. A fractional role has a remit, which is inherently elastic, and the elasticity is the reason clients like it. Two days a week as a fractional operations lead can quietly become three without any conversation, because there is no deliverable list to point at.
The concurrency limit is lower than it looks. Consultants moving into fractional work often assume they can hold two or three roles alongside project work, on the basis that each one is only a couple of days. In practice each fractional role carries a full context load, including internal politics, team relationships and meetings that do not respect your other commitments. Most people find the sustainable ceiling is lower than the arithmetic suggests.
The upside is revenue stability, which is the reason the model has grown. Predictable monthly income from two or three fractional engagements removes the feast and famine cycle that makes small practice financial planning so difficult.
If you run a mix, the thing to protect is the boundary on the fractional days. They should be specific days, communicated, and defended, because the alternative is that fractional work expands to fill the week and the project work, which is usually the higher-margin half, gets squeezed into evenings.
A cadence for the practice, not just the projects
Each engagement has its own rhythm, set by the client. What most small practices lack is a rhythm for the practice as a whole, which is where the cross-engagement problems become visible.
Thirty minutes a week is enough, and it is the highest-value half hour in the calendar. Go through every live engagement and answer three things for each. Where is it against plan. What is blocked and who is blocking it. Is the effort tracking to what we priced.
That last question is the one that changes behaviour, because it forces an actual look at the hours rather than an impression of them.
A monthly view sits on top of this and asks different questions. Which engagements are ending in the next sixty days, and what replaces them. What is the effective rate across the portfolio and which engagements are dragging it down. Which clients have asked for something that should have become a proposal and did not.
The pipeline question in particular gets neglected in exactly the wrong months. Small practices tend to do business development when they are quiet, which produces a cycle where you are busiest and least visible at the same time, then quiet and scrambling ninety days later. Keeping a small amount of consistent pipeline work in the calendar during busy periods is what smooths that, and it only survives if it has a slot rather than depending on spare capacity that never appears.
For practices tracking several engagements at once, this is exactly the kind of thing Empiraa GPS is used for, keeping the plan for each client engagement, the owners and the review rhythm in one place rather than spread across separate documents per client. Whatever the mechanism, the requirement is the same: one view where all engagements are visible together, because the cross-engagement problems are invisible from inside any single one.
Client reporting that costs less and lands better
A significant share of unbilled time in small practices goes into reporting, and much of it is wasted because the format is wrong.
The common pattern is a long written update, produced monthly, which takes several hours to assemble and which the client skims. It is expensive to make and low value to receive, which is the worst combination available.
Shorter and more frequent generally works better. A brief note every fortnight covering what moved, what is blocked and what is needed from the client is faster to write and more useful to read than a comprehensive monthly document. It also surfaces blockers while they can still be cleared, rather than reporting them after they have caused a delay.
The section clients value most is almost always the one naming what you need from them. Consultants underuse it, partly out of politeness and partly because it feels like passing the problem back. It is not. Most delivery delays in consulting are caused by inputs a client owes and has forgotten, and a standing line in every update that names those outstanding items resolves more of them than chasing does.
It also protects you. When an engagement runs long, the record of when you asked for something and how many times you repeated it is the difference between a shared understanding and an argument.
Where to start
If you are running several engagements and suspect the margin is not what you think, do three things over the next month.
Track hours roughly on every live engagement, however crudely, and calculate an effective rate for each at month end. Expect at least one surprise.
Add an exclusions section to your next two proposals, naming the adjacent work you are not doing. It takes ten minutes and it prevents the ambiguity most creep grows in.
Put thirty minutes in the calendar weekly to look across all engagements at once, and hold it even when the week is busy, because the weeks it gets cancelled are exactly the weeks it would have caught something.
None of this is the work you got into consulting to do. It is the difference between a practice that grows and one that is busy for three years and no more profitable at the end of it.


