You get off yet another Zoom call, and the meeting summary arrives about ninety seconds after the call ends. Five bullets, a couple of action items, and you skim it and think, yeeeaaaahhhhhhh, that’s about right. Then you close it.

Now think about the last time you opened the transcript underneath and read it to check whether the summary got things right. Not skimmed. Checked. I’ve been doing this a while, and it seems my answer is almost never.

In April a federal judge in California was handed one of these transcripts as evidence. She checked.

The recording nobody started

Two sets of lawyers logged into a Teams call, the free version, for a meeting the court had ordered them to have. That log-in triggered a recording by a tool called Otter. Nobody in the meeting started it, and nobody in the meeting knew it was running.

Afterwards both sides got an email offering them the transcript. One side paid for it and filed it in federal court. Judge Jennifer Thurston sealed it and wouldn’t consider it at all.

“The Court will not consider the transcript or any arguments based upon those ‘transcripts’ for the simple reason that no party has made any effort to demonstrate its accuracy or trustworthiness.”

Camarillo Hospitality LLC v. G6 Hospitality LLC, E.D. Cal., order of 24 April 2026

A court wouldn’t take the record of a meeting until somebody could show it was accurate. Most of us forward these to our teams every day without opening the transcript once, and that’s the bit I keep chewing on.

So, one question. Is that summary any good, and would you be able to tell?

Being fair about what these tools genuinely do

Microsoft’s own economists ran a randomized controlled trial across 66 companies and 7,137 knowledge workers. They allocated the licenses at random and measured the outcomes from the software’s own telemetry, rather than asking anyone how it felt.

Among people who used it regularly, email time fell by roughly 3.6 hours a week, about 31 percent. Focus time, meaning the long uninterrupted stretches where real work happens, went up nearly four hours. Out-of-hours email dropped.

That’s a good result and worth the effort (or lack of?). Keep in mind that three of the four researchers worked for Microsoft and the paper hasn’t been peer reviewed, which would normally make me take the entire thing with a large grain of salt. Here, it cuts the other way, because of what the same trial found next.

Meeting time didn’t significantly change. The headline feature of the product is the meeting recap, and the recap gave nobody their afternoon back. Six of the companies went the other way, with meeting time rising about 13 percent.

The authors’ own explanation is the honest one, and I think they’re right: when meetings get cheaper to sit through, people don’t reclaim the time, they hold more meetings.

The same trial turned up something I didn’t expect. People with the tool were measurably more likely to join more than five minutes late and to leave more than ten minutes early. They didn’t attend fewer meetings. They attended less of the meetings they were in, because the recap would be there.

So on the meetings half of your week you didn’t buy time. You bought the summary.

Hearing the room is harder than it looks

Before anything gets summarized, something has to hear the room, and that turns out to be the hard part.

An international research competition called NOTSOFAR-1 runs on 315 real recorded meetings, with all the crosstalk and the poor microphones and the person dialed in from a car. After a full international effort the winning system still got about 22 percent of the words wrong when the audio came from a single device in the room. A proper multi-microphone setup brought that down to about 11 percent. So the laptop sitting on the conference table is doing a lot of the damage, which is a cheaper problem to fix than most of what follows.

One specific failure is worth knowing about. When these transcription models hit a long stretch of silence they don’t always just leave a gap, they sometimes generate text that was never spoken, and researchers found the fabrication tracked closely with how much non-speech the audio contained.

Consider what a meeting actually consists of. Somebody hunting for the right slide, somebody unmuting, nobody talking while the screen share loads. Meetings are largely made of the thing that makes these models invent.

(There’s a separate problem with how well these systems handle different accents and dialects, which is a bigger story than I can do justice to here.)

The failure is omission, not fabrication

Fabrication or hallucination is the failure everyone worries about, but it’s not the problem you really think it is. What these summaries mostly do is leave things out.

In one study researchers took real meeting transcripts, had models summarize them, then had human annotators go through the results line by line. Ninety-seven percent of the summaries from the model tested were missing information. That number needs care, though, because it came from an earlier generation of model on a sample of 35 meetings, and today’s frontier models are better, so I wouldn’t quote it as a live figure for anything you’re using this week. What matters isn’t the number, it’s the shape. When these systems fail, they mostly fail by leaving something out rather than by making something up, and that distinction turns out to matter enormously, because the two failures are not equally visible. The researchers had a formal category for it, and the name is the part that stopped me.

Omission of significant decisions or actions.

I had this wrong for about two years. I assumed the risk was that a summary would tell me something false.

It isn’t lying to you. Almost everything in it did happen, exactly as written. You were right that it reads accurately, because it is accurate. What’s wrong with it isn’t in there to notice.

And you can’t really proofread something that isn’t there. You can read that summary ten times, as carefully as you like, and you’ll never find the decision that didn’t make it in. There’s nothing on the page to catch your eye.

Nothing else catches it either

If a person can’t spot the problem by reading, surely the software checks itself. That’s what I assumed. The industry grades these summaries automatically, with scoring tools comparing a summary against its source transcript.

Researchers tested whether those automatic scores track what human judges think. They don’t, at least not particularly. In roughly a third of the comparisons the scores actively masked the errors, because material that’s absent can’t drag a score down.

So run the chain forward. The transcription can mishear the room, the summary mostly fails by leaving things out, and the automatic scoring doesn’t reliably catch what got left out. Nothing anywhere in that pipeline flags it, which leaves exactly one check standing, and it’s you.

And we don’t do it

Researchers at Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers about their real use of these tools at work, asking them to say honestly when they’d applied any genuine scrutiny to the output. By their own account it came to about 59 percent.

So roughly four times in ten, with no real scrutiny. When the researchers looked at why, one of the most common reasons people gave was that the task felt too trivial to be worth checking, and the example named in the paper, for exactly that reason, is a meeting minutes summary.

The same paper goes further. The more confidence someone had in the tool, the less critical thinking they applied; the more confidence they had in their own ability at the task, the more they applied. Out of 319 people, 23 ever checked a source the tool had cited.

I’ve argued for a while that the thing to worry about isn’t artificial intelligence getting too clever. It’s how quickly we hand decisions to systems that aren’t clever at all, and then lose the ability to overrule them. If your system produces a record nobody can practically check, you haven’t automated anything. You’ve abdicated.

I swear abdication is going to become my buzzword, because I see it everywhere, and honestly, sometimes I find myself doing it as well. Little bits and pieces I convince myself aren’t important enough to check by hand. I’ve even said it in a previous article and video that AIs are great at tasks, not at replacing people, so if I’m just having it replace a task, then it should be fine, right?

The same missing check is running on your inbox

This is the part the audience got to before I did. Someone watching a Copilot tutorial left this:

“If I email my staff with Copilot, what if they just respond with Copilot and we have a back and forth with AI’s chatting on our behalf.”

That’s a joke. It’s also a measured finding.

A randomized trial in JAMA Network Open studied 52 physicians and 10,679 replies, testing what happens when software drafts your response to a message and you edit it. The drafted replies came out about 18 percent longer. Time spent reading messages rose nearly 22 percent. Time spent replying didn’t significantly change. More words, more reading, the same amount of time.

That’s clinicians and patient messages, not your inbox, and I won’t pretend it’s the same job. It’s still the cleanest randomized measurement anyone has of what happens when a machine drafts your replies.

Sidebar: There’s a video that makes me laugh so hard from YouTuber Angela Collier, called “ the malicious optimism of AI-first companies.” (link below) in this video, she talks about how the CEO of Zoom said in a public forum that eventually Zoom participants can be replaced by AI avatars that have been developed from each person’s respective personalities, emails, calendar, and all of that. The avatars will actually have Zoom meetings and make decisions on your behalf. The absurdity of this still blows my mind. Of course, the part that makes me laugh is the thought of an actual Zoom call with no human participants, because at that point, why do you need Zoom?

Ok, back to whatever my own point was.

What people report versus what gets measured

A gap runs underneath all of this. Ask people how much time they saved and you get one number. Measure it and you get a much smaller one.

In a UK Government Digital Service trial of the same product, participants reported saving 26 minutes a day. The measured randomized trial found about 12 minutes a week at intent to treat.

And the figure quoted most often, “64 percent less time on email,” literally doesn’t say that. What the source says is that 64 percent of users said the tool helps them spend less time processing email. That’s a directional opinion with no magnitude attached, from 297 volunteers in an early access program, surveyed by the company selling the product.

One more thing from that courtroom

Nobody in that meeting started the recording. It happened because of how an account had been connected to a calendar.

About a dozen US states require everybody on a call to consent to being recorded, not just one participant, and when people are in different states the strictest rule governs. A class action is running against Otter on precisely this question. Nothing has been decided, the hearing is still ahead, and those are allegations rather than findings.

You don’t need a court to supply the useful part, though. It’s the same missing step, one stage earlier. A decision got made about recording a private conversation, and no human made it.

And it’s probably worth being aware that just because it might be illegal in your jurisdiction, somebody in your Zoom call probably has invited one or two note-takers. They might be recording the call on their iPhone. It’s going to happen, so it pays to learn to be smart because clutching your pearls isn’t going to get everyone to magically change their behavior

The small thing

Sometime this week, pick one meeting where something was decided (Note: this originally said “something actually got decided”, which just seemed too snarky, so I decided to soften the language, but if you’re still reading this and that resonates with you, leave a comment, because meetings already sucked. Now, with five or six recorders and five or six participants, it’s just become a real drag to sit through some of these calls, right?). Open the transcript underneath the summary, search it for that decision, and see whether it made it into the bullets. It takes about four minutes, and you only have to do it once.

If it comes back clean, that’s genuinely useful to know. If one of these has quietly dropped something that mattered, I’d like to hear about it.

In short: these tools give back real time on email and none at all on meetings, so on the meetings half of your week what you actually bought was the summary. Summaries mostly fail by leaving things out, omission is invisible on the page, the automatic scoring doesn’t catch it either, and the only check left standing is a person who skips it because the task doesn’t look like it matters.

Sources: Camarillo Hospitality LLC v. G6 Hospitality LLC, E.D. Cal. (24 April 2026). Dillon, Jaffe, Immorlica & Stanton, NBER Working Paper 33795. NOTSOFAR-1 / CHiME-8. Koenecke et al., ACM FAccT 2024. Kirstein et al., Findings of EMNLP 2024. Tang et al., TofuEval, NAACL 2024. Lee et al., CHI 2025. Tai-Seale et al., JAMA Network Open (April 2024). UK Government Digital Service M365 Copilot trial, 2025. Microsoft Work Trend Index Special Report, November 2023. I’m not a lawyer, and the court material here is evidence about accuracy and consent, not legal advice.