The speech recogniser gave the words it got wrong an average confidence of 0.864. Its average across everything it heard was 0.948. I measured that in August, on recorded meetings I was turning into a searchable record, to see whether the low scores would flag the misheard names for me. They didn't. A wrong name and a right one looked nearly the same to the machine that wrote them both down.
I thought about those 2 numbers when I read HubSpot's Fall 2026 release, which announces a self-updating CRM. The release page says it captures "every call, email, and meeting" automatically, and follows that with a promise: "No manual entry."
In April I wrote about a Validity survey of 602 CRM users and stakeholders. The company's release put one of its findings in a single line: "37% of staff regularly fabricate data to tell leaders what they want to hear." A model asked to fill those same fields has no reason to leave one empty, unless somebody builds the reason in.
I build and fix sales systems for small B2B teams, and a lot of that work starts as CRM cleanup. Most of this release is good news for the records I clean up. One setting in it deserves a slower read.
01The boring fix shipped
The April piece ended on a list of boring fixes for made-up CRM data. 2 of them were "Call recordings that auto-log so reps don't have to manually summarize every conversation" and "Email sync that captures activity without someone copy-pasting a note." HubSpot now ships both as the headline. "Every call, email, and meeting gets captured automatically," the release page says. "The CRM uses it to keep customer data current."
I'm not going to argue with any of that. A transcript of the call is a better record than whatever a rep types from memory at 6pm, and a rep who never has to choose between logging and selling has one less reason to end up in that 37%. If the release stopped at capture, this would be a short post telling you to go turn it on.
02Recording a call and deciding what it meant
Those are 2 different jobs. A machine can record that a call happened on Tuesday, who was on it and what was said. The buyer's budget, their timeline and who signs the contract are conclusions somebody has to draw from what was said, and that second job is the one people were making up answers to.
HubSpot's product page for the feature, called Deal Progression, describes the second job in plain words: "Budget, timeline, decision-makers, objections: if it came up in a conversation, it self-updates in the CRM. Reps review and apply with one click."
A person reviewing every low-confidence word would have been pointed at the wrong part of the record.
The release page says the same feature will "update your CRM automatically." So does it write the budget in, or does it ask first? HubSpot's help article, which says an admin has to be enrolled in a public beta to use the feature, describes both. The main path is a suggestion: reps are told to "Review the recommendations for accuracy, then accept or reject the value update." In the upper right of that screen is a button called "Approve all." Setup has a step where an admin can "turn on auto-updates," and once a sales methodology is saved, the article says, "methodology properties update based on buyer conversations, reducing manual data entry."
As I read it, the careful version is what a rep meets first and the automatic versions are things somebody chooses, which is the right way round and HubSpot should get credit for it.
The phrase I read twice is "if it came up in a conversation." Plenty of first calls never get to budget. On the 3 HubSpot pages I read about this feature, I found what happens when the subject comes up. I didn't find a sentence about what goes in the field when it doesn't.
That's the gap my 2 numbers came out of. The recogniser never left a blank where it couldn't hear. It wrote down a word and attached a score, and the score on its mistakes sat within a tenth of the score on everything else. The words I knew were wrong also scored under 0.5 less often than the rest of the transcripts did. A person reviewing every low-confidence word would have been pointed at the wrong part of the record.
What made that record worth trusting was letting it say it didn't know. A voice in those meetings gets a person's name only when the sound of the voice and what was said both point to the same person, and otherwise it stays "Voice 4." A looser version of that rule measured about 93% right on the borderline cases, and I turned it down, because a wrong name is worse than a number. Under the stricter rule a lot of what was said still sits under "Voice 4."
A CRM field has the same choice to make. "Budget: not discussed" is a true entry, and the first thing I'd want to know about any self-updating CRM is whether it's allowed to write that.
03A measure of how full the boxes are
Another part of the release is called Context Home. Thomas Randall at Info-Tech Research Group wrote the release up on September 25, and he got to the objection before I did. By his account, Deal Progression recommends "updates to fields such as budget, timeline, stakeholders, and objections," and "Context Home then scores the completeness of the customer context available to HubSpot and identifies gaps that could reduce AI effectiveness." He tells buyers to find out "whether organizations can distinguish between technically complete records and information that is actually accurate, current, and appropriate for automated decisions. In other words, determine how trustworthy the trustworthiness score is."
I couldn't find the word "score" on HubSpot's own Context Home page, so the scoring is Randall's description. What HubSpot's page does say is "The more complete your context, the smarter and more effective HubSpot becomes." It prints 3 results for customers with "high-quality context" and defines the term in a footnote: "High-quality context means actively populated CRM fields, logged activity, connected integrations, and configured team settings." Populated is a lower bar than accurate, since a field with the wrong budget in it still counts.
The meeting record I built would grade badly on populated fields, because a lot of what was said in it has no name attached on purpose.
Put that next to the Validity line. People fabricated to tell leaders what they wanted to hear. My argument then was that a required field is how a leader says what that is: something in every box. A measure built on populated fields asks for the same thing, and this time the request goes to software that never runs short of time to fill a box. The meeting record I built would grade badly on populated fields, because a lot of what was said in it has no name attached on purpose.
The release page has a line that reads differently after the footnote: "Customers that use AI with high-quality context make 19.3x more connected calls on average." If the definition is the same one, that says the teams who fill in their fields, log their activity, connect their integrations and configure their settings make far more connected calls. I believe it. My guess is those teams were out-calling everybody before the AI showed up.
04What I'd ask before turning on auto-updates
Capture goes on the first day. Calls, emails and meetings logged without a person typing is the fix, and nothing below is a reason to wait on it.
Randall's list for buyers has a line I'd start from: "Determine which AI agent recommendations require human review and which low-risk actions may be executed automatically." I'd sort them like this: a field the system only had to record can update itself, and a field it had to conclude waits for a person. A transcript can still mishear a name, but the recording is there to check it against.
So the model had to name its clue from a short fixed list and quote the words, and code checked that the quote was really in the passage.
For every field on the auto-update list, I'd want to see the sentence the value came from. On the meeting record, a small model was proposing name changes in the transcripts, and every wrong one rested on a clue the passage didn't show. So the model had to name its clue from a short fixed list and quote the words, and code checked that the quote was really in the passage. Precision went from about 88% to about 96%, for about a cent a document. Over 4 rounds of that work the prompt changed twice and the code changed 9 times.
A suggested budget with the buyer's own sentence attached can be checked at a glance. Without it, "review and apply with one click" asks a rep who's short on time to approve a stack of values on trust, and "Approve all" makes the stack one click too.
Then the undo. In a system I built that writes into a client's CRM, 2 rules held up. A contact update that a second data source couldn't corroborate didn't get written. Every write was logged with its before and after values, so a whole run could be reversed. If auto-updates fill in a night's worth of budgets and a manager doubts them on Monday, somebody needs the list of what changed and a way to take it back as a set. I'd ask HubSpot's team to show me both before the setting went on.
05Where I could be wrong
A speech recogniser is a different machine from a language model reading a finished transcript. My 2 numbers show that one system's confidence didn't separate its mistakes from everything else it wrote. They don't show HubSpot's will behave the same way, so treat them as a reason to ask the question and nothing more.
The strongest case against this piece is the one I conceded at the top. A model reading the real call will beat a rep typing from memory on most fields, most days. If that 37% falls because nobody has to type anything, that's a win and I'll say so here. What I'm worried about is narrower: the fields where the call didn't say, filled in anyway and counted as complete.
Everything I've said about HubSpot comes from 4 of its own pages and one analyst's note, all read on October 5. A product page tells you what the vendor says, and I haven't shown you what the feature does on a live deal. The help article puts the feature behind a public beta, so some of this may change. The main path I found, suggest first and let a person accept or reject, is the right one, and the auto-update setting may turn out to be more careful than its name.
If I were switching this on for a team this month, I'd take the capture, leave auto-updates off for anything a person would have had to conclude, and count how often reps reject a suggestion. If that count is 0, I'd assume nobody is reading them.