Apollo's AI told me my email was making a classic cold email mistake.
The subject line was "Sales Ops." 2 words. The body opened with the contact's first name, then described what I do. I work inside sales teams as a fractional RevOps lead. I spend time understanding how reps are measured, what resources they've actually got, and whether the process is working. The email ended with a question. Are you open to a cold call from me tomorrow afternoon?
01Apollo told me to soften it
I asked Apollo to review it. Apollo told me the opener was a classic cold email mistake because it used the first name and then talked about me instead of the prospect. It called the cold call ask "presumptuous." It suggested softening to "Would it make sense to swap notes?" It rewrote the subject line to "Your RevOps setup — quick question." That em dash is Apollo's, not mine. Apollo writes everything with em dashes.
Every one of those suggestions made the email worse. The cold call ask is the line that makes the email memorable, because nobody asks for a cold call in a cold email. "Would it make sense to swap notes" is the exact phrase every SDR tool on the planet has been A/B testing since 2022. Apollo was right that the subject line was vague. It just reached for "quick question" instead of something shorter and sharper, because "quick question" is what the sequence library holds.
I reverted. Tightened the draft further. Kept the cold call ask.
02What it costs you to accept the rewrite
If I had clicked accept on Apollo's rewrite, I would have sent a worse email and felt productive about it. The interface would have shown me a green checkmark. The dashboard would have logged the assist. I would have moved on to the next draft a little faster than the last one, and the next draft would have started a little closer to Apollo's defaults, because that's how taste works. Nothing about a single edit feels like drift. It shows up 6 months later, when every email in your sent folder sounds like every other email everyone else is sending.
The reason I caught it was a document, a voice system I've been building for months, with rules about em dashes, about sentences that knock down one claim just to set up the next, and the specific phrases that read as AI-generated to anyone who knows what to look for. My gut never fired. I checked Apollo's suggestions against a reference that lives outside the chat window, and without that reference Apollo would have won every round.
You're moving faster through work that reflects whatever you already thought going in.
You're moving faster through work that reflects whatever you already thought going in. The AI is amplifying the taste and framing you already had. If you'd had a better answer you'd have written a better prompt. You didn't, so the model's guess becomes your answer, and the model's guess is downstream of whatever the training set was biased toward, which is usually the mean.
That swap, speed for averaging, is one most people take without noticing it happened. The speed is real. What goes away is variance, which was the only part that sounded like you.
03The rubric that looked done
I think about this more when I'm building things than when I'm writing.
A few months ago I was designing the call classification system that filters voicemails out of call logs so managers only review real conversations. My first instinct was to write the scoring rubric. 12 criteria. Did the rep identify the decision maker. Did they handle the primary objection. Did they establish next steps. 9 more like that. I wrote it out in a chat window and the model agreed it was a reasonable rubric. It looked done.
What stopped me was reading the actual transcripts. A rep logs 35 calls in a week. 31 of them are 3 seconds long because nobody picked up. 2 are voicemail greetings the rep sat through before the beep. One is an automated phone menu. One is a 90-second conversation that ended with "not interested." Running a 12-criteria rubric against that data produces scores that look authoritative and mean nothing.
The model agreed with my rubric because I gave it my version of the problem, and my version of the problem was wrong. The model wasn't checking my premise. It was extending it. If I'd shipped what we agreed on, I would have called it progress, and the dashboard would have agreed with me.
Nobody warns you about this part. The model agrees with whatever version of the problem you hand it. Whatever you assumed walking in is what you'll be assuming walking out, except now it has a rubric attached that makes it feel rigorous.
04I'm not outside this either
The AI drafting this post with me got stuck on a different example for 3 messages before I pulled it back. It was a dramatic bug catch from another project, vivid and specific, and the model wanted to open the post with it. I kept saying no and the model kept explaining why it was a good opener. I had to name the tunnel vision before it stepped back. You're reading the cleaned-up version. I'm not outside this. Neither is anyone using these tools.
Whatever you bring to the chat is what comes back, smoothed and confident, a little closer to the average of everyone else who asked something similar. The output looks like progress because you have something now that you didn't have before. But what you have is downstream of the assumption you walked in with, and that's the part the model can't see.
A second model reading the first model's output is still inside the loop. The check has to be something the model can't reach.
A good assumption survives this fine, a little flatter than it went in. A wrong one just comes back looking finished.
A second model reading the first model's output is still inside the loop. The check has to be something the model can't reach. A document with rules in it. A transcript you actually opened. A friend who tells you the email sounds like Apollo wrote it. The reps who say the dashboard doesn't reflect the work.
056 months from now your sent folder
Most people aren't going to build that check. They're going to use the tools the way the tools are designed to be used and accept every suggestion that looks reasonable. That adds up to 6 months of work that all sounds like each other. They'll feel faster, maybe even like they're getting better. And they will be, in the narrow sense that they're shipping more output per hour than they were a year ago.
But the output will reflect whoever they were a year ago, processed through a model that pulled them toward the mean, with no mechanism for noticing the pull. The thing they actually wanted to get better at, the judgment underneath the output, will be exactly where it was. Possibly worse, because they stopped practicing it.
My check is the document I keep open while I work. Yours might be something else. An earlier version of this post closed on one of the exact patterns it names as a tell, a tidy 2-sentence reframe that sounded smart enough to ship. The document flagged it. I'd already signed off.
Most of what I build works this way now: the check learns a rule, the work goes back through it. This post included, so the version you just read depends on when you read it.