How to solve AI outreach slop

Mohamed ChahinSeptember 30, 202620 min read

AI personalization done correctly took a customer's outbound campaign from a ~4.5% to a 12% reply rate. Here is the step-by-step on how to achieve the same.

When AI outreach actually delivers on its promise, you go from a ~4.5% to a 12% reply rate.

That's what a customer achieved on their very first campaign running with Twain's Ultra mode.

In this article, you'll learn exactly how you can achieve the same. With or without Twain. Everyone is wondering if they should buy or build anyway, so why not make that decision easier for you?

Big fat disclaimer. We'll never be able to fix missing PMF. If you don't have PMF, this won't help you. No matter if you use Twain or build it yourself.

This customer has PMF. Same goes for all the other customers reporting success with Ultra, the newest version of Twain (we gave a handful early access).

A customer reports a 12% reply rate on their first Ultra campaign, roughly 2.5x every other campaign they tested.

What problem is AI outreach trying to solve?

Let's start with the obvious non-obvious thing: why does it sound so smart to "use AI to personalize outreach," but when we all tried it, it didn't feel right? And somehow we can't really explain why we were excited in the first place. At least it's harder to do.

We have to start with the basics of why people reply to outreach, positive or negative.

  • Reason 1: it's relevant. Timing is right on 2-3% of your TAM.
  • Reason 2: there is effort behind your outreach.

The primary purpose of personalization is for the 97-98% of your TAM. The people you won't catch with great timing, but that you could catch by showing effort.

We want to focus on these 97-98%. The ones we would love to build a relationship with, create a high-quality touchpoint, and get them to associate a positive emotion with our brand.

The question that will decide how they feel about your outreach is: "Is there effort behind this outreach?"

Prospects reply out of reciprocity. Effort works as inference: visible investment predicts the conversation will be worth their time.

Let's make this more tangible with examples.

Why does gifting work? Gifting is a high-effort signal because it's expensive. At the top of the podium, we probably have an invite to the Super Bowl, and at the bottom is an Instagram ad. Why? The invite costs more than $20k. The ad costs $0.005 per impression.

No one will ask, "How much time did they spend on sending me the invite?" The financial investment is enough. That's the effort signal. Even when they're not buying, people will accept your invite or thank you for the gift.

You can do the same with time spent. When you invest 30 minutes to get to know someone (or their online profile) and write a thoughtful outreach message, you'll get a reply appreciating your effort. That's reciprocity. And that's a high-quality touchpoint with your ICP!

That's why data is backing the fact that personalization in outreach gets better results. Personalization that signaled effort.

Now, if the data says personalization improves campaign performance, why hasn't AI personalization got us to these results?

The reason why this failed is because we were all greedy and wanted to signal effort while spending $0 on this. Why would we, since content creation was always free, right?

We wanted to fake effort.

With all the AI slop, inbox fatigue, and AI SDR hate, I think it's safe to say this failed.

Why did we get AI slop instead of AI efficiency?

There are three parts to the answer: the data, the AI, and the user.

The data

Before AI, the degree of personalization we saw in email was industry, company size, company name, and prospect name.

Template with merge fields

Hey Prospect name,

We often hear from companies in {INDUSTRY} that productivity slows down when they hit {COMPANY_SIZE} FTE. Is that something you're experiencing at {COMPANY_NAME}?

The problem here starts with the data. In some databases, Twain is labeled as "IT services," team size "11-20," and the full legal entity name, "Twain Technology GmbH," sits under company name. So that outreach would read to me:

The same template with the database values merged in

Hey Mohamed,

We often hear from companies in IT services that productivity slows down when they hit 11-20 FTE. Is that something you're experiencing at Twain Technology GmbH?

How do you feel when reading this? Does it signal effort?

These labels existed before AI, and data providers did not fix them when AI came along. It's too expensive.

So people started feeding wrong and outdated data to an LLM and generated personalized messages that did nothing more than templates with merged fields did. With the difference that, next to failed field merges, prospects now got to enjoy AI hallucinations.

Which led to the now infamous "LLM injections."

A prompt hidden in a LinkedIn profile turns a recruiting email into a flan recipe.

Which Twain, of course, stops.

Twain flags the injected instructions as a warning on the lead instead of following them.

The AI

So, humans put in effort when you signal effort. What's happening with hallucinations is that we ask a prospect to put in effort after we failed to write their name correctly, failed to understand what their company does, or where they work.

Why is it so hard to get this right? Why has AI been so impactful in coding, customer success, or creative work, but not in sales outreach?

You've read a shit ton of content on that. I did too. But I've never felt like someone was explaining the actual problem. So here it is.

When you write code, vibe-code a landing page, generate an image, even a video, you're working with AI to create one artifact. You write a prompt and you get one asset out of it. You can look at that asset and decide if you'd want to change it, leave it, or just use it as it is.

The problem with AI in outreach is that a list of 1,000 prospects equals 1,000 artifacts. A 3-step sequence? Make it 3,000 artifacts.

You've written one big ass prompt, and you think it's airtight. Now, would you bet money that it is? Because it has to hold 3,000 times, one-shotted. You can't review it, and you can't fix it later.

That's the problem, and it's a problem we found not even state-of-the-art AI models can solve.

I can tell you right now: if you need to create a single banger outreach sequence and you have the time to fact-check everything, don't use Twain. Go to Claude, put Fable on max, and see magic happen. But the moment you try that on a list, Claude will betray you.

The user

Startups post-PMF typically have better products at better prices than the incumbents they're competing with. So they actually provide a better value proposition. But how do they break into enterprise accounts? Historically, email outreach has been one of the most effective ways to directly target decision makers in enterprise organizations.

Today, thanks to AI slop, successful startups are afraid to use AI. They care about their brand and won't risk having one of their outreach messages fall for an LLM injection and go viral on LinkedIn.

It's word for word what the VP of Demand Generation of one of the scale-ups that use Twain told me:

"The reason why we use Twain is because my worst nightmare is for one of our outreach emails to go viral on LinkedIn."

Now they found a solution, but the majority of companies with great products have not (yet). This is sad because you know what it has led to?

  • Great products care about their brand, so they don't use AI outreach.
  • Bad products don't care about their brand, so they do.
  • Your inbox fills up with bad outreach pitching bad products, and you get fatigued.

I've seen it a million times. It always starts with how you look at your own product. It's the single most common trait in our customer base: they care about their brand. Twain is a premium solution, and you have to care about your brand, your ICP, and your product to invest in Twain.

Remember the LLM injection screenshot from earlier? It is incredibly rare to catch these. But it's incredibly important to do so, because whoever puts that in their LinkedIn profile feels a certain way about outreach. So we invested an unreasonable amount of effort to spot them. That's one of the things you need to consider when building it in-house: invest the time to catch LLM injections.

No matter how great the message is, there's no point in reaching out to people you shouldn't be reaching out to. Yes, Twain actually stops you from reaching out to people that you shouldn't.

Qualification warnings on a lead who left their company and sits outside the target persona.

There are more irrelevant messages in my inbox than there are badly written ones.

Now, we've established the problem with AI in outreach at scale:

  • GTM data is historically wrong, incomplete, and outdated.
  • No LLM can follow your prompts at scale (especially not with bad data).
  • Great products (that should be using AI) are not using it because of the risk.

So if we manage to resolve these problems, we should be able to actually achieve what we wanted to achieve with AI in the first place: signal effort. The right amount of effort.

What the right amount of effort costs

What do I mean by "the right amount of effort"? Well, your economics have to work. Big ACV, bigger budget for acquiring that customer. Small ACV, small budget.

Translated: you'll throw human SDR time at enterprise accounts, even AEs, to make sure every touchpoint is thoughtful. You have maybe 1,000 of these accounts, tops, in your entire TAM. But there's a big belly of 50,000 accounts in your TAM that might convert with little effort.

For the longest time, little effort in email outreach meant a template plus nurturing every other quarter. In recent times, it meant AI slop. Cost per account is close to $0.

Would you spend $100 creating a touchpoint for them? No. $1? Yes.

So what can you get for $1?

I have to admit, sometimes I need to process that just a year ago you might have spent $0.01 to generate a sequence with Twain. The value proposition we could deliver at that price point was marginal.

Today, you might spend $1. So 100x that. Which might 2.5x your reply rates immediately. You might also get replies that you could've only dreamed of. Here's another customer reporting back after getting early access to Ultra.

A customer's reply from the exact profile they wanted at a 25,000-person company.

So how did we figure out how to solve this?

Solve the problems step by step

There's a simple but hard answer to the problem. You need to write like a human would.

You're now thinking, "ehhh, duh?!" It's more complicated than that. Writing like a human means a lot more than removing em dashes.

LLMs today take your input tokens and output your outreach sequence in one go. Humans don't do that. Humans are intuitive creatures. We intuitively understand rules about outreach without having to ever learn them explicitly. We make a lot of micro decisions along the way before we hit send.

Quick thought experiment: how many times have you written an email and then made edits to it to see "how the sentence feels"? We don't send the first draft. And when we use AI for manual outreach, we always get to review and make edits. Remember the problem with AI at scale? It's 3,000 artifacts, not one. That's why reviewing and iterating doesn't work at scale.

So how do we get AI to mimic human intuition? 3,000 times in a row.

What does human intuition look like in outreach? When AI would congratulate a prospect on a funding round, the AI would not question that the funding is six months old. It only understands that this is a typical email opener from its training data, and if that data point is available, it will be biased to use it. AI won't try to understand how much time has passed since the round was announced and how it has changed the organization since. But humans do exactly that. An entry-level SDR intuitively understands not to congratulate a raise that's six months old.

But for an AI, that's hard because of training data biases. Meaning the AI is lazy and biased toward choosing simple, high-frequency outreach templates ("Congrats on the new job!") that it saw millions of times in its training data. It will completely ignore better, more recent details because safe, generic openers are mathematically easier for it to generate.

Now here's what most people miss when they talk about AI outreach slop. It's not only em dashes that destroy your effort signal, it's decision-making biases.

So while you might care about your own personal tone of voice too much, what you should actually care about are AI biases. All types of biases. Your prospects don't know your personal tone of voice. They won't know if an email "sounds like you." What they will know is if it's overusing the word "real." They'll know if the signal mentioned is recent, if it's actually relevant to what they're working on, and if the social proof mentioned makes sense.

Getting AI not to mention words like "real," "effort," or "caught my attention" too much is a hard problem to solve. Getting AI to mimic human intuition felt a lot harder, and importantly, costs a lot more.

We believe Ultra is the closest thing on the market to human intuition. It feels incredibly intelligent. Let me walk you exactly through how we did this (this is the part you want to steal).

Relevance at scale used to require human effort. We made it require computing effort instead.

So how do we get the AI to put in the effort? The effort can't be about how AI works; it must work like humans work.

How do humans approach outreach? Intuitively, we take one decision at a time. Meaning when we start, we don't know yet exactly what we will do. We decide what to write based on research. We decide when to write based on qualification. We decide what to lead with based on what we learn about you, and the CTA based on what we know would be our best shot at having a conversation.

So far, templates and AI alike would write the same to everyone. So the way we solved it is by starting with research on contact and account level. And then, based on what we find, we qualify, we strategize, and only then we write.

Below, I use product screenshots for the most important steps. Showing you exactly what we've built feels the closest I can get to helping you build this yourself. It's important that you don't mix up steps and that you don't try to cut corners.

How these examples were made. Everything that follows comes from running Ultra on real people from the GTM world (all friends of mine), researched from their public activity. Nothing is edited or human reviewed. The closing will show a message for someone I "accidentally" put into a cold outbound campaign who turns out to have an existing relationship with Twain. You will see how Twain catches this and changes course to make the message relevant and save me from embarrassing myself.

Research

We start with real-time deep research to verify data. One of the core value propositions, if not the one that Twain stands for the most, is real-time data accuracy. Our users don't have to trust databases, and they don't have to build messy tables to clean and enrich data that ends up being outdated anyway. We first find truth.

I can't comment on the cost here since we have bulk agreements in place, but something that might help reduce cost is scoping out exactly what the minimal data necessary would look like to confirm real-time information on your ICP. Meaning, you might not need to search everywhere. Searching in one specific place or for one specific data point is for sure cheaper. When testing solutions, always make sure that it's real-time. Data gets outdated real fast.

Contact and account research, each signal with its source and age. The Relationship line already says this lead is a Twain advisor.

Qualification

This is where AI and humans typically fail at scale. Humans can't go through the entire list by hand and update who moved jobs, who isn't in a decision-making position, or what account doesn't actually fit your ICP. And while AI could, it struggles to access, process, and sort accurate data at scale.

So after we run research, we understand who we're reaching out to on account and contact level. Now here is where we've stumbled on something that feels magical. No matter your intention, Ultra will stop you from pitching a bad-fit lead. We've given Ultra the freedom to decide what angle makes sense. And when there is no direct angle, there is no direct pitch. This is where the experience becomes unique, because AI will typically try to follow your instructions, and humans won't put in the effort to decide how to qualify each and every prospect at scale.

It's equal to spending 20 minutes per prospect going back and forth with Claude Fable max just to qualify them. By the way, for 1,000 prospects that's 333 hours, or 13.89 days. No one would do that. But if we would, surely we wouldn't approach all 1,000 the same way. But that's exactly what we're doing with AI outreach slop.

Twain on Ultra can now make better decisions at scale because we mimic what 13.89 days of effort look like.

So when building the qualification layer, give freedom and thinking time. Models are highly biased to "complete the task." Any hint will nudge them towards pitching what your overall intention is. The way to approach it is to find reasons why you shouldn't be reaching out to them. Why today is not the right timing. What I'm trying to say here is that your AI will try to make you happy by giving you what you asked for, even when it's not in your best interest. Don't let it fool you.

A persona mismatch: this lead runs partnerships, not SDR tooling, so the pilot pitch is off the table.

Strategy

Here's where it all comes together. We take outreach so seriously that we came up with a messaging strategy "product" that documents all decisions and ends with a priority list of what to write.

So after we fix biases caused by outdated data and biases caused by not questioning the fit or intention, we now go to the hardest part: making sense of all data points and deciding what to write and what not.

Who knows your prospect best? They do. Writing something that makes sense to them (and ideally only them) is actually a hard task.

We already touched on the date math behind signals. Here's another bias that human intuition solves implicitly but AI fails to. Signals don't concern everyone in the organization. A small startup CEO has her priorities aligned with the company's priorities. A COO probably too. But a director of engineering at a Fortune 500 company is not tasked with the organization's 2027 product roadmap.

That's why Twain starts with understanding internal alignment.

Level alignment: what this lead actually owns, and which part of the offer speaks to that.

After Twain understands the alignment, we go deep on the prospect's actual profile and preferences. See, with more and more decision makers sharing their opinions online, it makes sense to understand sentiment before you reach out. Especially when you're selling something as timely as AI.

When building this, make sure you QA signals from social activity. I've been seeing a lot of posts on outreach to someone mentioning "Hey Mohamed, I've seen your comment on Mohamed's post on deep research." Does that signal effort? You need to process all signals, and comments or posts are no different. That's why it's important to separate these tasks. When using social listening, process every input into qualification and reasoning, and convert it based on that reasoning. Don't let it be injected into your message writing.

Risks, channel preference, and language and tone for a CRO who publicly judges outreach.

Another bias that, in all fairness, happens to humans just as much as it happens to AI is that social proof mentioned in your outreach is just not relevant to me. Sure, social proof is better than no social proof, but you know what's a true sign of effort? Social proof that's relevant to me. So that I (as the prospect) don't have to spend my valuable (and very scarce) cognitive energy to understand if you're credible or not.

Now, with verifying social proof before writing anything, we give Twain time to think about what makes sense to mention.

Verified social proof: one proof point cleared for this lead, with a limit on how to use it.

Reading about verified social proof, you might argue, "We have a list of social proof, case studies, and reference customers which we store in our Claude instance and constantly update." The problem with that approach is that whenever that list gets created and used, you don't have all the context. Per campaign, per account, per contact. Only the moment you're building that campaign do you have all the information. That's why you need to build another layer of social proof verification within the campaign.

Timing and date math

We all constantly talk about timing. When 2-3% of your TAM is in the market for your solution, it's fair to assume that timing is mostly off. And when timing is off for 97-98% of your TAM, the signals you'll find are not going to be hot and fresh buying signals.

When you're working on this, you won't be able to fix that. Instead, what we've found works better is understanding timing and date math. That's what a human understands implicitly (remember the example with six-month-old funding). We've built this for every signal. And every prospect. That's how we mimic intuition when it comes to understanding signal timing. It's one of the easiest giveaways of AI outreach slop.

Like I mentioned earlier, you need fresh data to achieve the same. But what you might not need is to research as heavily and broadly as we do. Know your ICP extremely well, and you might be able to reduce cost and latency by narrowing your searches.

Timing and date math: every signal dated, and only the recent ones left as the angle.

Now, having worked through level alignment, risks, language, and signals, the next thing we have to do, and what we would do as humans writing outreach, is to decide what data point to ignore (irrelevant, don't write), what data point to use as context (relevant, but don't write), and what data point to mention explicitly (relevant, write).

This solves the issue of models being overwhelmed by information, biased by certain signal types (funding announcement), and being able to consider date math.

Three buckets

We call it three buckets.

Three buckets: what to ignore, what to keep as context, and what to mention.

What we get from all of this is a sequence strategy. This strategy now controls the writing. It's the result of real-time deep research, qualification on accurate data, and an unreasonable amount of thinking to avoid all typical biases.

Real effort. Just not invested by a human, but by Twain.

The sequence strategy: one warm note from Mohamed instead of a cold sequence, because the lead is already a Twain advisor.

Writing

The thing that surprised me most is a side effect of what we've built. In Ultra, a user does not have to give any writing instructions. Twain will make every decision on opener, body, social proof, and CTA based on what is best for that prospect.

Before we built this feature, the key to succeeding with Twain was creating great frameworks. It was an actual blocker for new users to get value out of Twain. You've probably felt that with any AI: restricting it as much as you can to cover edge cases, build fallbacks, and catch hallucinations. Humans don't need any of that. They have intuition.

So seeing that Ultra doesn't need that either feels amazing. It feels like a success. It feels like the right step to solving AI outreach slop.

The message Twain wrote for the advisor I put into the cold campaign by accident: no cold pitch, one personal note about the MCP overlap.

The ultimate test: would you compliment AI outreach?

If someone tells you they're getting compliments left and right for 100% AI-generated outreach, would you believe them? I wouldn't, because most of what I've seen is AI slop.

People use AI just like they use software, not understanding that you can't achieve better results than with a template without investing more effort. You only get an edge out of AI if you're willing to do things differently.

Very few companies invest $1 per prospect in an outreach campaign. Our customers now do. And what they get from it is immediately visible. Here's another customer sharing screenshots of prospects complimenting their BDR team.

The reason why this is happening is because they're standing out. And prospects appreciate that. So even when it's not the right timing, they manage to have more conversations with their ICP. And that's all that matters.

How many times does your team get complimented for their outreach?

A customer's first report on their Ultra campaign: replies complimenting the outreach.

The catch

This will not work forever. It will break when it isn't rare anymore. When this level of AI intelligence gets 100x cheaper, 100x more teams will use it. Which will destroy your effort signal.

Effort works in the context of what everyone else is doing. And you can't fake it. I wish I could tell you this is your forever silver bullet, but the truth is you and I have to continue innovating to stand out. I have to come up with the next version of Twain that will be an order of magnitude better than everything else, and you have to continue figuring out how to stand out from the competition. How to signal effort.

I truly hope this helped you answer some questions. Questions like: what 98% of your TAM would reply to, why AI has had an impact in coding but not in sales outreach, and how to build your own tool that conveys effort. I have intentionally not touched on subject lines (I don't think there's a core issue here), lead magnets (hard to provide general advice), or email infrastructure and deliverability (not my expertise).

I also hope it inspired you to drop templates for good. A template will never signal effort. You might argue there's targeting effort behind templated outreach, but that doesn't convert if the message can't convert it.

The message is the entire surface of what your prospect sees. It's the surface of your product's first impression. To your prospect, it is your product. So let's ship a better one.

Either build the decision-making chain yourself or try Ultra on Twain. Twain's Ultra mode is natively available in Claude through the Twain MCP server and in Clay.

If you're building this and would like my opinion, book a call with me.