Home/Blog/AI SDRs: What to Automate in Outbound and What to Keep Human

AI SDRs: What to Automate in Outbound and What to Keep Human

Sales rep reviewing an automated outbound sequence on screen

Two years of "replace your SDR team with AI" marketing has produced a predictable outcome. A lot of teams bought an AI SDR, ran it for a quarter, got worse results than the humans it was meant to replace, and concluded the category is nonsense. Both the original claim and the backlash are wrong, and the confusion is expensive. There is real value in automating parts of outbound. There is also a set of tasks where automation reliably makes results worse, and the line between them is not where the marketing put it. This is an attempt to draw that line properly, for a team of five to fifty rather than an enterprise with a dedicated revenue operations function. The framing is not "should we use AI in outbound", because you already do, in your email tool, your data provider and your CRM. The useful question is which specific tasks in the outbound process should be handed over, and what happens to the ones that should not be.

Why the "AI SDR" framing caused the problem

The reason so many pilots disappointed is that the product was sold as a replacement for a role rather than a tool for a set of tasks. An SDR does a lot of different things. They build lists, research accounts, find contact details, write first-touch messages, follow up, handle objections in early replies, qualify, book meetings, and pass context to an account executive. Those tasks have wildly different characteristics. Some are pattern-matching at volume, which is exactly what current systems are good at. Others require judgment about an ambiguous human situation, which is exactly what they are not. Sold as a role replacement, the product has to attempt all of them, and the weakest link determines the outcome. A system that researches accounts brilliantly and writes openers that make prospects wince produces worse results than a mediocre human doing both adequately, because the visible output is the email. Industry commentary through 2025 and 2026 has increasingly landed on the hybrid position: automation handles research, enrichment and volume work while people handle the conversation. That is a reasonable conclusion but it is stated too vaguely to act on. What follows is a more specific split. What automation genuinely does better Start with the tasks where handing over is straightforwardly a good idea. Finding and enriching companies that match a defined profile is the clearest case. This is high-volume, rule-based work with objectively checkable output. A person doing it manually is slower, more inconsistent and more likely to give up on a difficult record. There is no upside to a human building a list by hand, and the time saved is substantial: analyses of how sales reps spend their week have consistently found that direct selling occupies well under half of it, with research and administration taking large shares. Contact discovery and verification is the same category. Finding a likely email format, checking whether the address resolves, identifying the right person by title, keeping the record current as people change roles. All of it is mechanical, all of it is checkable, and none of it benefits from human judgment. Account research summarisation is the third clear win, with a caveat. Pulling together what a company does, its recent announcements, its technology footprint, its headcount trend and its likely priorities into a readable brief is genuinely useful and genuinely fast to automate. The caveat is that the brief should be read by a person before it informs a message, because summarisation errors are confident and the cost of referencing something inaccurate in an email is high. Follow-up scheduling and sequence logistics belong in the automation column without argument. Nobody should be manually tracking who is due for touch three. Finally, the administrative residue of selling. Logging activity, updating stage, writing call notes, keeping the pipeline current. This is where a large amount of a small team's selling time disappears, and it is the least controversial thing to hand over. It is also, oddly, the last thing many teams automate, because it feels less exciting than automating the outreach itself. What automation reliably makes worse Now the other side, which gets less attention because it is harder to sell. First-touch message writing is the most common mistake. It looks like a writing task with a template, which is why it is automated first, and it is actually a judgment task about what will matter to a specific person in a specific week. Automated first-touch email fails in a recognisable way: it opens with manufactured enthusiasm, references a pain point the prospect may not have, and pitches capability rather than relevance. Recipients recognise the pattern immediately, and recognition produces either deletion or a spam complaint. The complaint outcome is the part that has changed the calculation materially. Mailbox providers now enforce complaint rate thresholds around 0.3 per cent for high-volume senders, with non-compliant bulk mail rejected rather than filtered. An automated system sending many multiples of a human's volume with a slightly higher complaint rate per message does not just perform worse, it damages the sending domain that all your outbound depends on. Volume used to be the compensation for poor targeting. It is now the amplifier for it. Second, reply handling. An early reply is the highest-value moment in the entire outbound process and it is where ambiguity is greatest. "Not right now" means five different things depending on who said it and what else is in the message. Automating this stage saves a small amount of time and costs a meaningful proportion of the opportunities that were genuinely available. Third, qualification judgment. A system can check whether a company meets stated criteria. It cannot tell that the person who replied is enthusiastic but has no budget authority and is about to leave, which is the kind of thing an experienced rep picks up from tone in two sentences. Fourth, and most important for a small team, deciding what to say. The strategic content of outbound, which segment to target, what problem to lead with, which proof matters, is your commercial judgment. Handing that to a system that optimises for reply rate on a small sample produces local optimisation towards whatever gets clicks, which is not the same as whatever gets customers.

The hybrid split in practice

Put together, a workable division for a small team looks like this. Automation owns list building, enrichment, verification, research briefs, sequence logistics, follow-up timing and CRM hygiene. People own the segmentation decision, the message, the first-touch send in low-volume high-value segments, all reply handling, and qualification. The interesting middle ground is first-touch at volume. If you are selling something with a small deal size to a broad market, automated first-touch is defensible because the economics of human-written outreach do not work and the cost of a weaker message is tolerable. If your deal size is meaningful and your addressable market is a few thousand companies, automating the first touch is close to indefensible, because you are trading the quality of a scarce resource for volume you do not need. The test is straightforward. Divide your target market size by your team's realistic sending capacity. If you could reach your entire addressable market by hand in a quarter, do not automate the message. If reaching it by hand would take four years, automation of first-touch is a reasonable trade and you should focus your quality effort on the reply stage instead.

What to fix before automating anything

The most reliable predictor of a failed automation project is automating on top of a broken foundation, and there are three specific things worth checking. The first is whether your ICP is actually defined. Not aspirationally described, defined: firmographic criteria specific enough that two people would build the same list. If it is vague, automation will produce a large volume of near-misses, and near-misses generate complaints. The second is whether your messaging has been validated by a human. If you do not have a message that gets replies when a person sends it carefully to a well-matched prospect, automating it will not create one. It will produce the same non-performing message at scale, which is worse than the status quo because of what it does to your domain. The third is data quality. B2B contact data degrades continuously, with published decay estimates varying widely by field and methodology but all pointing the same direction. Automation applied to a stale list produces bounces at scale, and bounce rate is one of the thresholds now enforced strictly. There is an unglamorous version of this advice: the highest-return automation available to most small teams is not in outreach at all, it is in the administrative layer underneath it. Removing manual CRM updates, manual list building and manual research gives your existing people back a meaningful share of their week, which they then spend on the conversation work that automation cannot do. That is the actual mechanism by which these tools produce returns. This is broadly the design of Empiraa Signal, which handles the finding, enrichment and sequencing side while keeping the message and the conversation with the person selling. ANI, the AI assistant across the product, does the summarising and drafting work rather than the deciding. You can review the full set of Signal features and Signal pricing before deciding which parts of outbound to automate.

Where the research step goes wrong even when it should work

Account research is in the automation column, but it fails often enough in practice that it is worth examining why, because the failure is instructive about the whole category. The problem is that a research brief is only useful if it changes what you say, and most automated briefs do not. They return a competent summary of what a company does, which the rep could have got from the homepage, alongside a list of recent events with no interpretation. What a rep actually needs is the answer to a narrower question: given what this company just did, what is likely to be difficult for them right now in the area I sell into. That is an inference rather than a summary, and it depends on knowledge of your own market that a general-purpose system does not have. A brief that says a company posted four operations roles is a summary. A brief that says four operations hires in a quarter usually means their current process is about to break at the handover step, because that is what happened with the last six customers who did the same thing, requires your accumulated pattern knowledge. The workable version is to write the inference rules yourself and have the system apply them. Decide in advance what each signal type implies for your product, write that down as a short set of if-then statements, and use automation to detect the signal and attach your interpretation. This is considerably less impressive than an AI that figures out relevance on its own and it works far more reliably, because the commercial judgment stays with the person who has it. It also compounds. Each time a signal-based hypothesis turns out to be wrong, you update the rule, and the briefs get better. Teams that rely on a system to generate relevance from scratch have nowhere to put that learning.

The governance question small teams skip

Enterprise discussions of AI in sales spend a lot of time on governance, and small teams tend to dismiss it as bureaucracy for companies with compliance departments. There are two parts of it that matter regardless of size. The first is knowing what is being sent in your name. An automated system that generates message variations will produce output nobody has read. Most of it will be fine. Some proportion will make a claim about your product that is not true, misstate pricing, or address a prospect in a register that does not match how you want the business to sound. At small scale you will hear about it from a prospect, which is an expensive way to find out. The mitigation is not approval on every message, which defeats the purpose. It is sampling. Read twenty generated messages a week, chosen at random rather than selected by the system, and read them as a recipient rather than as the person who built the sequence. This takes fifteen minutes and it catches the drift that otherwise accumulates unnoticed over a quarter. The second is data handling. Enrichment tools and outreach platforms process personal information about people who never agreed to be in your database, and the obligations attached to that vary by jurisdiction and are not optional. Australian businesses have obligations under the Privacy Act, and any team selling into Europe or the United Kingdom is dealing with a materially stricter regime. The practical minimum is knowing which tools hold contact data, having a way to delete a person's record on request, and honouring opt-outs across every system rather than just the one where the request arrived. Small teams frequently fail the last of those, because an unsubscribe in the sequencing tool does not remove the contact from the list in the data provider, and the person gets contacted again next quarter from a different sequence.

Deciding when a pilot has failed

A structural weakness of small-team automation projects is that nobody defines failure in advance, so the pilot continues on the strength of the effort already invested. Write the failure condition down before starting. It should be a specific number on a specific metric by a specific date, chosen so that hitting it means stopping. Something like: if this has not produced eight qualified meetings in ten weeks at a complaint rate under 0.1 per cent, we turn it off and go back to the previous process. The reason to commit in advance is that mid-pilot, every number has a story attached. Volume was lower than expected because of a warm-up delay. Reply rate was down because of a holiday period. The messaging was still being tuned in weeks two and three. All of these may be true and none of them can be assessed fairly by the person who chose the tool. Ten to twelve weeks is usually the right window for outbound automation. Shorter than that and you are measuring warm-up. Much longer and the sunk cost has become large enough to distort the decision, and you will keep something mediocre because turning it off would mean the quarter was wasted.

Measuring whether it worked

Finally, a warning about evaluation, because a lot of automation pilots are judged on the wrong number and either killed unfairly or kept too long. Volume metrics are not evidence. An automated system will produce more sends, more opens and often more clicks than the humans it replaced, and none of that establishes anything. Judge on meetings held with qualified prospects, and on the same definition you used before, over a period long enough to be meaningful. Watch the negative metrics with equal attention. Bounce rate, unsubscribe rate and complaint rate where visible. A pilot that produced more meetings and doubled the complaint rate has not succeeded, it has borrowed against future delivery. And measure the time returned, because that is often the real benefit and it goes uncounted. If automating enrichment and CRM updates gave two reps six hours a week each, that is the return, and it shows up as improved output somewhere other than the automated process itself.

Frequently asked questions
Do AI SDRs actually work?

It depends entirely on which tasks are being automated. Automation performs well on list building, data enrichment, contact verification, research summarisation, sequence logistics and CRM administration, all of which are high-volume and objectively checkable. It performs poorly on writing first-touch messages in considered markets, handling replies, and exercising qualification judgment. Products sold as complete replacements for the SDR role have to attempt all of these, and results are determined by the weakest component, which is usually the message. Framed as a set of task-level decisions rather than a role replacement, the technology is useful.

What should stay human in outbound?

The segmentation decision, the core message, reply handling and qualification. Reply handling matters most: an early reply is ambiguous, high-value, and the point where a small amount of human judgment changes the outcome disproportionately. In markets with meaningful deal sizes and a limited number of target accounts, first-touch writing should also stay human, because there is no volume advantage worth trading message quality for.

Why do AI outbound pilots fail?

Most commonly because the foundation underneath them was not ready. An undefined ideal customer profile produces high volumes of near-miss targeting. Messaging that has never worked when a person sent it will not work at scale. A stale contact list produces bounces at scale. Automation multiplies whatever it is given, so a process with a quality problem gets a larger quality problem. Deliverability consequences are the second common cause, since complaint and bounce thresholds are now enforced with outright rejection rather than spam filtering.

Does automating outbound hurt email deliverability?

It can, and the mechanism is volume interacting with complaint rate. Mailbox providers including Google, Yahoo and Microsoft enforce a spam complaint rate threshold around 0.3 per cent for high-volume senders, and Microsoft's requirements effective May 2025 reject non-compliant bulk mail rather than filtering it. An automated system sending several times a human's volume at a similar or slightly worse complaint rate per message can cross that threshold and damage the reputation of the sending domain, which reduces delivery for all subsequent mail including well-targeted messages.

What is the first thing a small sales team should automate?

The administrative layer rather than the outreach. Contact enrichment, list building, verification and CRM updates consume a large share of a rep's week, are mechanical, and carry no downside risk if automated. Handing those over returns hours to people who can then spend them on conversations, which is the part automation handles least well. Automating the message first is the more common choice and the one that most often produces a worse result than the process it replaced.

Ash Brown

Ash Brown

Founder & CEO of Empiraa

Published 13 September 2026

Ready to fix the part of your business that feels messy?

Whether you're trying to execute strategy, grow pipeline, or connect the way your team works, Empiraa gives you a clearer system to run from.

GPS for strategy execution. Signal for sales growth.