Real information in, invented information mixed in, one confident answer out. The interface never shows you how it got there.We built it to sound sure of everything. The best thing it can learn to say is “I’m not sure.”The latest argument my team and I are going back and forth on is whether we are being too honest with our users.One of the AI products we are building has a panel on the side shows you what it’s doing as you’re working through a problem. As you talk to the platform, the building blocks come together with, “Here’s what I heard you say,” “Here’s the thread I’m pulling,” “Here’s where this is headed.”What worries the folks who argue against it is that it’s too much for the user to absorb. It’s a fair argument. Nobody wants to watch the sausage get made, and real products bury the process on purpose, on the theory that an answer looks smarter when you can’t see where it came from.That’s an entirely backwards way of thinking and it’s one of the most important things we’re getting wrong about AI right now.Right now, what you actually get from Claude or ChatGPT is a play-by-play that tells you nothing. “Analyzing your request” or “Working on it.” Cool. You never see the one thing that matters, how it got from your question to its answer. We’ve all quietly agreed this smoothness is what good looks like.A confident reply sliding out of a black box, the Mechanical Turk in reverse: this time there’s really nothing behind the curtain but the machine, and it’s guessing.That’s the whole problem.We should not automatically trust what you can’t see and you cannot catch a mistake we were never allowed to see.Sunnie S. Y. Kim calls this “trust calibration,” where people work out how much to trust an AI from what it shows them about how it reached an answer. Show them nothing, and they’ve got nothing to go on but confidence, which is exactly the wrong thing to trust.So far, we’ve built ours to do the opposite. It shows you the actual reasoning as it comes together, laid out where you can see it, argue with it, and stop it when it’s heading somewhere dumb.The people who think it’s too much aren’t wrong that it’s more. That’s the point.We trust people who show their work and say “here’s what I’m sure about and here’s what I’m guessing.” Then we turn around and build AI that does neither, and act surprised when nobody can tell the good answers from the confident garbage.Haochen Guo and Petr Polak point out the obvious here, that showing the reasoning only builds the right kind of trust when the machine also tells you what it’s unsure about. Transparency without “here’s what I’m not sure about” just gives you more to read and no idea what to trust. Otherwise known as “explainability.”The way you get there is by letting people see how the thing works and, crucially, having it tell you when it’s unsure. Kim, along with Q. Vera Liao, Mihaela Vorvoreanu, Stephanie Ballard, Jennifer Vaughan actually tested this and when an AI that said “I’m not sure, but…” got people to stop blindly agreeing with it and, this is the part that matters most, catch more of its mistakes.Going back to my team who argue against showcasing the sausage being made, they’re correct that more is not automatically better. Herman Saksono, Vivien Morris, Andrea G. Parker, and Krzysztof Z. Gajos found that the harder an explanation was to parse, the more people just deferred to the AI.Dumping the entire contents of the machine’s head on the table helps nobody. The trick is showing the parts a person actually needs to make the call, and being honest about the parts you’re guessing at.The harnessLeft alone, the thing inside will happily wreck the place. The harness is whatever keeps it from doing that.All of the above sounds fine and good in a bubble, but the catch is that a machine won’t do any of this on its own.Left alone, a frontier model is a Gremlin after midnight. That same brain behind ChatGPT or Claude will cheerfully hand you a confident answer to a question it has no business answering, whatever direction the conversation takes, because it was built to produce something plausible and keep you moving.This isn’t just a hunch of mine. Hiroshi Okumura dug into context-dependent suppression, testing big models, they held back from overconfident judgment almost every time in a careful, academic setting, and almost never once you asked them for practical advice. When he asked for a concrete recommendation, caution basically vanished. In one test, one response out of two hundred kept it.The helpfulness is the problem. The model wants to give you something usable, so it stops hedging exactly when the stakes go up.Miao Xiong and colleagues found that when you simply ask these models how confident they are, they come back overconfident almost across the board, parroting the sound of certainty rather than measuring anything real.Ask it to be honest, and it’ll perform honesty as smoothly as it performs everything else, which is its own kind of problem.Engineers call this a harness. You take one of these giant commodity models and build a structure around it, create rules about what it can and cannot do, tools it’s allowed to reach for, checkpoints it has to hit, guardrails that keep it from wandering off.This is a real, named approach. Traian Rebedea and his team at NVIDIA built one of the better-known versions of it, a toolkit that sits between the user and the model and enforces rules on what it’s allowed to say and do, at the moment it runs, without touching the underlying model at all.You cannot just reach into the giant brain and rewire it. You need to build a structure around it and make the structure do the discipline the model won’t.The model still does the thinking. The harness decides what good behavior looks like and holds it there.Harness is the right word for it. It isn’t a cage, and it isn’t a leash you use on something weak. You put a harness on something strong enough to hurt you or itself if it goes wherever it wants. A climber wears one over a long drop. A stunt driver straps into one before the crash. A parachutist trusts one at ten thousand feet.The point is to create accountability, to keep the power and lose the part where it drags you off a cliff sounding completely sure of itself.Showing the work and admitting doubt aren’t features we decided to add because they’re unique or human. They’re behaviors the harness forces. The reasoning shows up on the side because the structure makes it, not because the model volunteered it.There’s a second reason to force the work into the open. When the machine shows each step, you can teach it. You can put your finger on the exact place the logic bent and fix that, instead of throwing out the whole answer and hoping the next one lands closer.Jonathan Uesato and his colleagues found that giving a model feedback on each step of its reasoning, rather than just grading the final answer, made the reasoning far more reliable, and that only works because a person can see the steps well enough to correct them.A black box gives you one blunt tool, thumbs up or thumbs down, but a visible process gives you a red pen.The machine says “I’m not sure about this part” because there’s a rule that says it has to, instead of smoothing over the gap and hoping you don’t notice. Asking a model to be honest doesn’t work. It just costs you more tokens and gets you the same confident guess in a humbler voice.The honesty has to be built in, forced by the structure, or it isn’t really there.Knowing what it doesn’t knowThe machine has to pick an answer before it knows whether the answer is right. Nothing about how sure it sounds tells you which way it guessed.The harness gives you a place to make one specific decision that almost nobody makes, but let’s talk about what the machine should do when it isn’t sure.The default answer, the one baked into every out-of-the-box model, is answer anyway. Confidently. It’s Cliff Clavin at the end of the bar in Cheers: “the average human being only uses 17% of his brain. Boy, you realize what that means? We don’t use a full, uh… 64%.” A fact he’s completely sure of and completely wrong about, except the machine says it in a nice font and you have no bartender rolling their eyes to tip you off.On a recent episode of Pod Save America, Casey Newton and Tommy Vietor played a clip, originally from the Instagram account husk.irl, of someone asking a chatbot which month is spelled with an X. It answered instantly that the answer was December, “right in the middle, like a little holiday surprise.” Asked to double-check, it switched to October, then spelled October out and noted there’s no X in it, then landed on February.Three confident answers, each one wrong, each delivered with the exact certainty it would have used if it were right. Nobody built that machine to lie. It was built to always have an answer, and having an answer and being right turn out to be different things.One member of my team said these models will give you something almost regardless. Sometimes it’s right, sometimes it’s close, and sometimes you’re left staring at a tidy little answer trying to reverse-engineer how it got there.The box gets filled either way and whether it should have been is your problem to sort out later.As a result of this, we built a rule into the harness that the machine only offers an answer when it clears a bar for how sure it actually is. When it doesn’t clear the bar, it isn’t allowed to guess. It has to stop and come back to you and ask, “Here’s what I’ve got, here’s where I’m thin, tell me more before I keep going.”It turns out we’d reinvented something with a name: selective prediction, an idea that goes back to Ran El-Yaniv and Yair Wiener, who described letting a model stay quiet on the cases where it’s shaky instead of forcing an answer for every box.When an AI answers everything, people fall into automation bias. They rubber-stamp the confident wrong ones right alongside the good ones, because from the outside the two look identical.Zana Buçinca, Maja Barbara Malaya and Krzysztof Gajos showed you can push back on this by building in what they call cognitive forcing functions, small moments that make a person stop and actually think instead of nodding along. A machine that stops and asks is one of those moments, engineered right into the product.There’s a version of this everyone’s already seen. In Slumdog Millionaire, the whole tension is that a kid from the slums answers with total certainty, and nobody can tell from the outside whether he knows the answers or is the luckiest guesser alive.That’s the problem with a confident machine in one image.Confidence tells you nothing about whether the knowledge underneath it is real.The only thing that helps is the show’s other ritual, the one where they stop you before you commit and ask, “is that your final answer?” That pause, that forced second look, is worth more than any amount of certainty in the voice.The move we’re after is closer to the “Miracle on the Hudson” than to a game show. When both engines went out over the Hudson, Captain Chesley “Sully” Sullenberger didn’t pretend that he could make it back to LaGuardia. When he worked out, fast, that he couldn’t, he said so out loud and safely landed the plane in the river instead.Knowing the edge of what’s possible, and owning it in real time, is exactly what we want the machine to do. It’s what we trust Sully for. A machine that says “I can’t do this part reliably, let’s talk it through” is doing the same thing, on a much smaller stage.The person on my team was right about why it matters, beyond just being “more honest.” He said it could be the thing that sets us apart. Think about that for a second. We’ve gotten so used to AI that sounds certain no matter what, that a machine willing to say “I don’t know yet” reads as a competitive advantage.That’s how far the bar has dropped.Teaching goes both waysYou don’t get to say you taught it. Like Eliza at the ball, it has to show, out loud, that it actually learned.That red pen is the whole game, and it cuts both ways.Start with the teaching, because it’s the part people miss. When a model shows its reasoning, you’re not just watching for mistakes, you’re able to correct the specific one and make the next answer better.Hunter Lightman and his colleagues at OpenAI compared grading a model only on its final answer against grading each step of its reasoning, and the step-by-step version won clearly, in part because it “specifies the exact location of any errors.”The reason step-level feedback works is almost aggressively obvious once you say it out loud. You can only teach what you can see. Long Ouyang and his colleagues, also at OpenAI, showed how far that goes. A smaller model taught with steady human feedback was preferred by people over a version more than a hundred times its size that hadn’t been. The teaching mattered more than the horsepower, but you only get to teach if the thing is willing to show you what it did and let you tell it where it went wrong.What should maybe make anyone building this way a little nervous is that if we make the machine show its process, we’ve quietly agreed to something on our end too. We now have to be right about how we’re checking it.Showing the work cuts both directions.The user watches the machine, and the machine’s process is now watching us, daring us to actually do the verification we’re claiming to do.We are actually worse at that than we think. Matthias Hümmer and colleagues, tracking people solving problems with AI over six months, found the gap widen as problems got harder. On the toughest ones, people stayed highly confident the AI’s answer was right, over ninety percent sure, while fewer than half of those answers actually were. The confidence didn’t track the correctness at all.This is where Pygmalion is instructive, and I mean the George Bernard Shaw play, not just the musical it turned into. Henry Higgins claims he can pass a flower girl off as a duchess, but the claim is worthless on its own. Eliza has to stand in front of the room and prove it, “the rain in Spain,” the whole performance, live, where anyone can hear a slip.Eliza has to show she learned. We should want the same thing from a machine we’re claiming has gotten better.A model that hides its process can tell you it’s improving, tell you it’s sure, tell you it double-checked, and you have no way to call the bluff. A model that shows its work is making a claim you can actually test, and holding still while you test it.That’s the only version of “trust me” that means anything, from a machine or from the person who built it.Everyone’s building without a harnessEnchant the broom, put your feet up, and walk away. The flood is what happens when nobody built the part that knows when to stop.Even if you never touch an AI product, this matters. The tools are everywhere now, doing real work, and almost nobody has built a harness around them.Walk into any company right now and you’ll find people who’ve never designed anything shipping designs, and people who’ve never written a line of code shipping code. And they’re trusting what comes out.In 2025, KPMG found that 58% of employees admit to relying on AI output without checking whether it’s accurate, and 57% say they’ve already made mistakes because of it. Only 41% work somewhere with any policy governing how AI gets used.The machine is answering everything, and most of the people leaning on it have no bar for when to believe it.It’s the Sorcerer’s Apprentice, the Mickey Mouse sequence in Fantasia. He enchants the broom to haul his water, it works beautifully, and he puts his feet up. Then the broom doesn’t stop. It floods the whole place, because the one thing he never built was the part that knows when to quit.Paul Boag put the workplace version well when he said a lot of what gets generated is “a small mountain of plausible looking rubbish that nobody has the expertise to spot.” It looks finished, it sounds right, and it arrives faster than anyone can check it.Therein lies the trap! The quicker it comes and the more done it looks, the less anyone thinks to stop and verify it.Our immediate reflex could be to slam the brakes, route everything back through the one team that knows what “good” looks like. That never works though. They’ll bypass the team every time, because the AI is instant and the team is slow.Boag’s answer is the org-level version of the harness. You stop trying to inspect every output and you build the structure that shapes the outputs before they exist, the rules, the guardrails, the standards for what good means and how to ask the machine for it. His line for it is sharp in that the difference between useful output and confident nonsense sits almost entirely in the brief.While my team built a harness around a model, a company has to build a harness around its people using models. It’s the same logic either way.Build quality into the structure up front. Make the right way the easy way, and make the work show itself as it happens. Then you’re not stuck inspecting every answer at the end, hoping to catch the bad ones before they ship.This is urgent, not down the line, but now. It’s not a risk that’s coming, it’s the way most companies already work. Every one of them that skips the harness is quietly betting that a quick glance counts as checking and that a confident answer is close enough. Most of the time, it is close enough. That’s exactly what makes it dangerous.As Owen Rust put it, human mistakes in this kind of work are more common but far less costly, because a person moving slowly tends to catch themselves, while an AI won’t flag its own error before handing it to you. So it slips through, looking finished, until the day it’s wrong in a way that costs you.By then it already shipped, wearing your name.Pay attention to the man behind the curtainPull the curtain all the way back and the confident voice was always just a nervous little operation working the levers, hoping you wouldn’t look.Let’s return to the folks on my team who thought we were showing users too much. As I’ve mentioned, they were right that it’s more work and that it’s slower. Building a machine that shows its reasoning, flags what it doesn’t know, and stops to ask instead of guessing is harder than building one that just answers.The challenge here, though, is that a polished black box is easier to build and easier to sell. It looks smarter. It moves faster. It may even win…for a while.Think about what that black box actually is, though. In The Wizard of Oz, the terrifying floating head, the booming voice, the smoke and fire, all of it falls apart the second Toto pulls back the curtain and shows you a nervous man working the levers. The wizard’s whole trick was keeping you from seeing how the effects were made.“Pay no attention to the man behind the curtain” is the most honest thing the Wizard ever says, because what lay behind the curtain was the entire product.That’s the choice we’re making with AI and most people are making it without noticing. You can build the Wizard, the confident voice, the hidden machinery, the answer that looks like it came from nowhere or you can build the thing that pulls its own curtain back and shows you the levers, and trusts you to keep using it anyway.The second one is the only version worth building, because the first one only works until it’s wrong.And it will be wrong.While this may seem pessimistic, it’s simply how these machines are built. A team of researchers from OpenAI and Georgia Tech showed that models hallucinate because the way we train and score them rewards confident guessing over admitting uncertainty.Basically, a confident wrong answer scores better than an honest “I don’t know,” so we’ve spent years teaching these machines that bluffing pays and honesty doesn’t.And it isn’t getting better as the models get smarter. If anything, it’s getting worse. Scott M. Graffius, analyzing 2025 hallucination data, found that the newer reasoning-focused models hallucinate more than the ones they replaced, some of them making things up on a third to a half of open-ended factual questions, more than double the rate of the models from a year earlier. We’re building machines that are more capable and more confidently wrong at the same time.So it will be wrong, confidently, at the worst possible time.The machine that showed you its work is the one you had a chance to catch. The machine that hid everything behind a curtain is the one that takes you down with it, still insisting it was sure.We’re going to keep showing our users the sausage being made. Yes, it’s more and it’s messier. Some people will say it’s too much, but the first time the thing says “I’m not sure about this part, let’s talk it through,” and saves someone from shipping a confident mistake with their name on it, that will settle the whole argument.References and further readingOn trust, transparency, and showing the workSunnie S. Y. Kim, Q. Vera Liao, Mihaela Vorvoreanu, Stephanie Ballard, and Jennifer Wortman Vaughan, “‘I’m Not Sure, But…’: Examining the Impact of Large Language Models’ Uncertainty Expression on User Reliance and Trust” (FAccT 2024). When an AI flags its own uncertainty, people stop blindly agreeing and catch more of its mistakes.Haochen Guo and Petr Polak, on why showing the reasoning only builds the right kind of trust when the machine also says what it’s unsure about.Herman Saksono, Vivien Morris, Andrea G. Parker, and Krzysztof Z. Gajos, the harder an explanation is to parse, the more people defer to the AI instead of thinking it through.On what the models actually doHiroshi Okumura, “When Helpfulness Overrides Causal Caution,” a model’s caution nearly vanishes the moment you ask it for a practical recommendation instead of an academic judgment.Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi, “Can LLMs Express Their Uncertainty?”, ask a model how confident it is and it comes back overconfident, performing certainty rather than measuring it.Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang, “Evaluating Large Language Models for Accuracy Incentivizes Hallucinations” (Nature), models hallucinate because training and scoring reward confident guessing over admitting uncertainty.Scott M. Graffius, an analysis of 2025 data showing newer reasoning models hallucinate more than the ones they replaced, up to a third to half of open-ended factual questions.Casey Newton and Tommy Vietor, “AI Apocalypse Now,” Pod Save America, a chatbot confidently gives three different wrong answers to a simple question, each with total certainty.On the harnessTraian Rebedea, Razvan Dinu, Makesh Narsimhan Sreedhar, Christopher Parisien, and Jonathan Cohen, “NeMo Guardrails,” a structure that sits between user and model and enforces rules at runtime without touching the model itself.Ran El-Yaniv and Yair Wiener, “On the Foundations of Noise-free Selective Classification,” the foundational work on letting a model abstain on the cases where it’s shaky instead of answering everything.Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos, “To Trust or to Think,” cognitive forcing functions, small design moments that make a person stop and think instead of rubber-stamping.On teaching and accountabilityJonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins, “Solving Math Word Problems with Process- and Outcome-Based Feedback,” feedback on each reasoning step, not just the final answer, makes a model far more reliable, and only works because a person can see the steps.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe, “Let’s Verify Step by Step,” step-level feedback beats grading the final answer, in part because it pinpoints the exact location of an error.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and colleagues, “Training Language Models to Follow Instructions with Human Feedback” (InstructGPT), a much smaller model taught with human feedback was preferred over one a hundred times its size.Matthias Hümmer, Franziska Durner, Theophile Shyiramunda, and Michelle J. Cummings-Koether, a six-month study finding that people stay highly confident in AI answers on hard problems even as their actual accuracy drops below half.On the organizational stakesKPMG, 58% of employees rely on AI output without checking it, 57% have made mistakes because of it, and only 41% work somewhere with any AI policy.Paul Boag, the org-level harness, protecting quality by shaping the brief up front instead of inspecting every output.Owen Rust, on why AI errors cost more than human ones, a person moving slowly tends to catch themselves; the machine won’t.AI is lying to us, and nobody seems to care was originally published in UX Collective on Medium, where people are continuing the conversation by highlighting and responding to this story.
AI is lying to us, and nobody seems to care