The demo on my site takes a visitor's website, reads it, then goes out and finds companies that would plausibly buy from them. The first version read it with a model that could run web searches, scored each company 1 to 5, returned up to 5, and cost me about 2 cents a run.
I pointed it at my own site over and over while I was building it, because that's the only way to tune a prompt. The companies it kept handing back were companies I could have named from memory. Funded, hiring, publishing, already sitting in somebody's sequence. Nothing was wrong with the output. The fit reasons were accurate and the companies were real.
I started paying attention to what it never returned.
I build and fix sales systems for small B2B teams, and outbound is part of that, so I run Apollo myself, including the MCP connector that lets me prospect by typing requests into Claude instead of clicking filters. The list it hands me is narrower than I assumed. It took me longer than it should have to notice, because nothing about the output looks wrong.
01Somebody already wrote the easy version of this
Sachin Jha published the version of this argument I'd have written first in June. "Everyone has the same ZoomInfo list. Same job titles. Same revenue filters." He's right, and he gets to the consequence fast: reply rates sit around 2% and nobody looks surprised anymore.
His prescription is intent data. Stop targeting profiles, target moments. Funding rounds, job postings, headcount changes, and by his numbers, stacking 2 or more of those signals gets you to about 15%.
That prescription is a product. A handful of vendors sell it, to everyone in the category, at the same time. If the disease is that we all query the same inventory with the same filters, the treatment is that we all subscribe to the same signal feed and set the same thresholds on it. All that does is move the crowd to a newer signal.
To his credit, Jha half-says this himself. He warns that standard intent feeds are commoditizing and steers readers toward public records, permits and procurement notices, the things that appear before commercial platforms pick them up. I think that warning is the actual finding of his piece. What he's describing is data nobody has packaged yet, which has a shorter shelf life than anyone selling it will admit.
Search this topic and you get vendor blogs, nearly all of them published this year, nearly all arguing that AI has made targeting more precise. I could not find a page anywhere making the opposite case.
02What actually happens when you type the prompt
AI prospecting names 2 different things, and they fail in different ways.
The first is a database with a chat interface on it. Apollo's MCP connector is a set of filter tools: search people, search organizations, enrich, create contacts, add to a sequence. The documentation lists 50 or so. When I type "find me people near here who look like my customers," no model is inventing companies. It's picking filter values on my behalf and reading rows out of a database that, by Apollo's own marketing, holds 230M+ contacts and serves 600K+ companies.
Convergence comes from 3 places there, and only the third one is worth arguing about.
Shared inventory is the boring one. One database, one set of records, everybody pulling from it. Default sort is nearly as boring: whatever the platform ranks first is what the model reads first, and in my own use it doesn't go deep. It takes the top of the list and stops, the way I would if I were in a hurry, except it's in a hurry every time.
Then there's the prompt. A vague request has to become a specific query, and the specific query it becomes is the most ordinary one available. Apollo's own documentation shows you what that looks like, in their sample prompt: "Find VP-level marketing leaders at SaaS companies with 200-1000 employees in the Bay Area." Title, industry, headcount band, geography. 4 filters, and the vendor picked them as the canonical example because they are the canonical example. My vague prompt collapses to roughly that. So does yours. A person building the same list by hand would at least have brought a hunch along, some unreasonable conviction about a vertical, a company somebody mentioned at a bad conference. The collapse strips that out on the way in, before any filter you set ever gets applied.
Apollo also ships a feature called Lookalikes. You paste in your best customers and it returns companies with matching firmographics. Apollo's own description says it helps you "expand your prospect list by doubling down on what already works." The name is accurate. It's a product for finding more of what you already have, sold to people whose problem is usually that they need something they don't have yet.
The reasonable objection is that good operators customize, and they do. The pitch of the connector is that you don't have to.
One other thing on that product page, which doesn't help my argument at all: there's a line at the bottom saying the page exists for promotional purposes and results are not guaranteed. I noticed it while I was pulling the contact number and it stuck with me.
03The machine picks the famous one when nothing else separates them
The second kind of AI prospecting sends a model out to the open web to find companies. That's what my demo does. There's no fixed inventory to pull from, so the failure mode is different, and the sales world hasn't picked this one up yet. The research lives in venues no operator has a reason to read.
Start with the finding that changed how I read my own output. In June, 2 researchers ran experiments on brand recommendations across GPT-4o-mini, Claude Sonnet, and Gemini 3 Flash, using skincare products, and found what they call a conditional monopoly. When every product carried identical specifications, the well-known brand got recommended 100% of the time, which is a number I had to read twice. The effect also collapsed the moment a competitor had any real edge. A rating advantage of 0.1 of a star was enough to break it.
Read that second half and the paper looks reassuring. The bias only shows up when the model has nothing to separate the options with.
Now think about what a prospect list is. Whether a company will buy from you has no star rating. There's no spec sheet, no review average, no benchmark. Put 2 logistics companies side by side, same metro, same headcount, same stack. On the one dimension you care about they are identical to anything reading the public internet, because the thing that separates them lives in a conversation nobody has had yet. So in prospecting the tie condition is the standing condition, every query, every row, which means the tiebreak does nearly all of the work, and the tiebreak is fame. That's the part I can't get out of my head. The paper's reassuring finding, that a small quality difference dissolves the monopoly, is a promise a model can only keep in a category where quality gets published, and prospecting has never been one of those.
A March audit of brand and culture preferences in AI assistants named the consequence. The authors write that "recommendation sets define the effective consideration space available to users and, increasingly, to autonomous agents," and their top-preferred brands turned up in 71 to 76% of relevant answers. Consideration space is the phrase worth stealing. A company that doesn't surface never gets a score at all. There's no row for it to sit at the bottom of, no filter you can loosen until it appears, and nothing in the output to tell you it was ever a candidate. The same authors flag amplified market concentration as a systemic risk, which in a paper means a line in the discussion section and on a Monday morning means your list.
The lists aren't even stable across tools. A June study ran 250 category queries through GPT-5.2, Gemini 3 Flash, and Perplexity's sonar-pro, and all 3 picked the same top brand in 104 of them, 41.6%. In consulting it was 22%. So which companies you see depends partly on which model your vendor happened to wire in, and 2 competitors on different tools may be working different pools without either of them knowing. That cuts against the tidy version of my argument, and I'll come back to it.
04The companies with the most internet
My first note on this said the companies surfacing in AI results were there because somebody was spending on GEO. That was wrong, and I'm glad I checked before writing it. As of early 2026, no major platform sells placement in generative answers. Nobody bought their way into your list.
Visibility gets earned instead, per that same source, through schema markup, content structured so an answer can be lifted straight out of it, density of named entities, mentions on review sites and in press and on Reddit, and freshness, with recently updated pages appearing roughly 4.3x more often than stale ones.
Part of what the score is measuring is how much marketing a company does, and it hands that back to you labeled as fit.
Every item on that list describes a company with a marketing department, or one that hired somebody to act like a marketing department. Not one of them says anything about whether that company has the problem you fix, the budget to fix it, or an executive who has been complaining about it in meetings.
That's what I mean when I call a company legible: it has produced enough public text about itself to be retrieved reliably. Legibility is a marketing outcome. The tools measure it and hand it back to you labeled as fit.
The loop closes on itself. A company markets, which makes it legible to the tools, and the tools surface it to everybody at once. So it hears from all of us, every week, which makes it a worse prospect this year than it was last year. Nothing in a fit score knows that.
A fit score assembled out of public signal can only reward public signal. The companies scoring highest are the ones with the largest legible surface. Part of what the score is measuring is how much marketing a company does, and it hands that back to you labeled as fit. That's true of my demo, of Apollo's scoring, and of anything else that reads the internet and returns a number.
05The thing I built had this problem
The demo used to source the way everything else does. It ran a live web search, a live web search returns legible companies, and so it handed back the same names for the same reason I have spent this piece describing. I had most of a draft written before I noticed my own demo belonged on the wrong side of the argument.
So I rebuilt the middle of it. It does not search for companies any more. It reads the visitor's own pages, works out what a company with their customers' problem would be hiring right now, and matches that against an index of which companies have which roles open, read off the hiring software those companies run themselves. A company turns up because it did something, not because somebody wrote about it. It still runs the ordinary search alongside as a control and prints both lists next to each other, which is the only part of this I can hand you as a check rather than a claim: on the runs I measured this week, the overlap was 0.
The bias that replaces it is the substrate's. Greenhouse, Ashby and Lever skew hard toward funded technology companies, and of 12 reference domains I tested, 2 ran a board I could read at all. So it reaches the unglamorous end of the market thinly, and healthcare staffing is worse served now than it was before I started.
What it is for is still narrower than a prospecting workflow. It shows a visitor what output looks like when a machine has read their site first, with a source link on every claim and the arithmetic behind every score printed underneath it. A run takes about 40 seconds and costs me about 9 cents, which is several times the search version and buys the one thing the search version could not do: I am holding the page the quote came from, so a quote is checked by exact match instead of by trusting that a hostname implies a sentence.
The discipline I'm describing here is a layer a person adds after any tool returns anything, mine included. I don't have a way to automate that layer. If I did, I'd be selling it.
There's a fact here I went back and forth on including. Every client I've closed at Outblox came through somebody who already knew me. None of them came off a list, including mine. I sell outbound systems and my own pipeline has never been machine-surfaced, which either undercuts everything above or is the most direct evidence for it I've got. I lean toward the second one, which is convenient for me, so weigh it accordingly.
06Contacting somebody without a score
The standard advice when a list stops working is to tighten your ICP. Add a signal. Get more specific about who you're chasing. I've given that advice. It doesn't touch this, and it took me a while to see why. Your ICP is a filter, and a filter runs on whatever got retrieved. Get the profile perfectly right and you've described a precise subset of the machine-legible slice. You've made the profile sharper without changing where the names came from.
What moves the slice is contacting a company the machine can't see, which means a person deciding something with no score behind it.
Nothing confirms you were right for weeks. There's no fit score to screenshot for your manager and no dashboard that turns green.
Pull 20 companies nobody's tool surfaced, start working them, and nothing confirms you were right for weeks. There's no fit score to screenshot for your manager and no dashboard that turns green. If it fails, nothing tells you, and you've spent a month finding out. Everyone I know who works this way works on conviction, and conviction is the least reportable thing in a sales org, which is most of why the ground is still open.
I'm not making a case for the phone here. The question is where the names come from before anybody picks a channel, and for a lot of teams right now the answer is a ranked list they didn't rank.
What I'm doing about it is unglamorous, and it stopped being about position. Reading deeper into the list was my first answer and it was still the list: taking number 150 instead of number 3 accepts somebody else's order and walks further down it. What replaced it is a filter for absence. How many pages are in their sitemap, whether they keep a blog, whether there is marketing tooling in their own page source. Every input is measured on the company's own site and every input is a published-marketing output, so a company that scores low is quiet rather than small. I sort down on it and cap how many of the loudest can make a list. It never touches the fit score. Those companies are harder to research, and there's less on their sites to personalize with.
There's a second-order version of this I can't prove. If every tool surfaces the same accounts, those domains receive the most machine-written mail, and their filters harden first. Does that make the most-contacted slice also the slice where your email is least likely to arrive? I don't know. I've never seen anyone measure it, and I'd want the data before I said it out loud.
07Where I could be wrong
I've built this argument out of research that mostly wasn't done on prospect lists, so here's where it's thin.
The strongest counter is a 2024 study that measured popularity bias in an LLM recommender against traditional collaborative filtering and found the LLM less biased, with no mitigation applied at all. The same paper tested instructing the model to avoid popular items. It worked, and bias went strongly negative. Accuracy fell hard across every model they tried. So the effect is partly correctable, and correcting it costs hit rate. That trade is what I'm actually recommending, and it should get named as a trade.
The skincare study is skincare. Consumer products with published star ratings, not 120-person logistics companies. I think the tie condition carries over and I've argued why. I haven't shown it.
The number that would settle this doesn't exist. Nobody has published the overlap between 2 competitors' prospect lists. Not the vendors, and not me. People have been writing "everyone has the same list" for a year, myself now included, and none of us has measured it. An early title for this piece had a specific count of companies in it. I invented that count, noticed, and cut it. The right denominator for a ratio like that would be the machine-legible portion of a market, and I've never seen anyone size that against the whole.
I've been here before with a smaller number. I once pulled 35 logged calls off an activity report and 8 of them turned out to be real conversations, with the rest voicemail and unanswered rings. I trust that instinct more than I trust any figure above, and it's only worth anything if I point it at my own arguments too.
The change I'm making to my own list is small enough to be embarrassing. When a query comes back, I'm taking a set off the bottom along with the top, tagging which half each company came from, and leaving it alone for a quarter so I can compare replies. If the bottom half is dead, I'll have spent 3 months emailing companies nobody else wanted, for a reason that turned out to be wrong. I don't know yet. The tagging takes about a minute.