Apollo's AI told me my email was making a classic cold email mistake.

The subject line was "Sales Ops." Two words. The body opened with the contact's first name, then described what I do. I work inside sales teams as a fractional RevOps lead. I spend time understanding how reps are measured, what resources they've actually got, and whether the process is working. The email ended with a question. Are you open to a cold call from me tomorrow afternoon?

01Apollo told me to soften it

I asked Apollo to review it. Apollo told me the opener was a classic cold email mistake because it used the first name and then talked about me instead of the prospect. It called the cold call ask "presumptuous." It suggested softening to "Would it make sense to swap notes?" It rewrote the subject line to "Your RevOps setup — quick question." That em dash is Apollo's, not mine. Apollo writes everything with em dashes.

Every one of those suggestions made the email worse. The cold call ask is the line that makes the email memorable. Nobody asks for a cold call in a cold email. That's the whole point. "Would it make sense to swap notes" is the exact phrase every SDR tool on the planet has been A/B testing since 2022. The subject line was vague, which Apollo was right about, but the fix wasn't adding "quick question." The fix was something shorter and sharper that didn't sound like it came out of a sequence library.

I reverted. Tightened the draft further. Kept the cold call ask.

02What it costs you to accept the rewrite

Here's the part that should bother you. If I had clicked accept on Apollo's rewrite, I would have sent a worse email and felt productive about it. The interface would have shown me a green checkmark. The dashboard would have logged the assist. I would have moved on to the next draft a little faster than the last one, and the next draft would have started a little closer to Apollo's defaults, because that's how taste works. You don't notice the drift in any single edit. You notice it six months later when every email in your sent folder sounds like every other email everyone else is sending.

The reason I caught it wasn't instinct. It was a document. I have a voice system I've been building for months, with rules about em dashes and parallel negation pairs and the specific phrases that read as AI-generated to anyone who knows what to look for. I wasn't catching Apollo with my gut. I was catching it against a reference that lives outside the chat window. Without that reference, Apollo would have won every round.

You're not getting better. You're moving faster through work that reflects whatever you already thought going in.

You're not getting better. You're moving faster through work that reflects whatever you already thought going in. The AI is amplifying the taste and framing you already had. If you'd had a better answer you'd have written a better prompt. You didn't, so the model's guess becomes your answer, and the model's guess is downstream of whatever the training set was biased toward, which is usually the mean.

This is the trade most people accept without noticing they're accepting it. Speed in exchange for averaging. The output goes up, the variance goes down, and the variance is the part that was yours.

35 → 8logged calls vs real conversations after filtering voicemails
12scoring criteria the model agreed looked rigorous before I read a single transcript

03The rubric that looked done

I think about this more when I'm building things than when I'm writing.

A few months ago I was designing the call classification system that filters voicemails out of call logs so managers only review real conversations. My first instinct was to write the scoring rubric. Twelve criteria. Did the rep identify the decision maker. Did they handle the primary objection. Did they establish next steps. I wrote it out in a chat window and the model agreed it was a reasonable rubric. It looked done.

What stopped me was reading the actual transcripts. A rep logs 35 calls in a week. Thirty-one of them are three seconds long because nobody picked up. Two are voicemail greetings the rep sat through before the beep. One is a menu tree. One is a 90-second conversation that ended with "not interested." Running a 12-criteria rubric against that data produces scores that look authoritative and mean nothing.

The model agreed with my rubric because I gave it my version of the problem, and my version of the problem was wrong. The model wasn't checking my premise. It was extending it. If I'd shipped what we agreed on, the dashboard would have looked great and measured nothing, and I would have called it progress.

That's the thing nobody warns you about. The model agrees with the version of the problem you bring. It does not bring its own. Whatever you assumed walking in is what you'll be assuming walking out, except now there's a rubric attached to it and the rubric makes the assumption feel rigorous.

04I'm not outside this either

A small honest thing before I keep going. The AI drafting this post with me got stuck on a different example for three messages before I pulled it back. It was a dramatic bug catch from another project, vivid and specific, and the model wanted to open the post with it. I kept saying no and the model kept explaining why it was a good opener. I had to name the tunnel vision before it stepped back. You're reading the cleaned-up version. I'm not outside this. Nobody using these tools is outside this.

So here's the thing I want you to sit with. Every prompt you send is a draft of you. Whatever you brought to the chat is what comes back, smoothed and confident and a little closer to the average of everyone else who asked something similar. The output looks like progress because you have something now that you didn't have before. But the something is downstream of the assumption you walked in with, and the assumption is the thing the model can't see.

Every prompt you send is a draft of you. Whatever you brought to the chat is what comes back, smoothed and confident and a little closer to the average of everyone else who asked something similar.

If your assumption was good, your output is good and a little flatter than it would have been. If your assumption was wrong, the model just made the wrong assumption look finished.

The check is not another AI. The check is something the model can't reach. A document with rules in it. A transcript you actually opened. A friend who tells you the email sounds like Apollo wrote it. The reps who say the dashboard doesn't reflect the work. The check has to live outside the loop or it isn't a check.

05Six months from now your sent folder

Most people aren't going to build that check. They're going to use the tools the way the tools are designed to be used, accept the suggestions that look reasonable, and produce six months of work that all sounds like each other. They'll feel faster. They'll feel more productive. They'll feel like they're getting better. And they will be, in the narrow sense that they're shipping more output per hour than they were a year ago.

But the output will reflect whoever they were a year ago, processed through a model that pulled them toward the mean, with no mechanism for noticing the pull. The thing they actually wanted to get better at, the judgment underneath the output, will be exactly where it was. Possibly worse, because they stopped practicing it.

Instinct is the name we give to the thing we can't explain. Most of the time it's not instinct, it's a system we haven't bothered to make legible. Mine is a document I keep open while I work. Yours might be something else. The question isn't whether you have one.

The question is whether you'd notice if you didn't.