I handed a UX review over to AI. Here’s what happened.

AI is a powerful partner for UX reviews, not a replacement for the expert, and the accuracy comes from the expertise you feed it, not the model.

I got a task that isn’t complicated in itself: I had to write a UX expert review for the search, results, and product detail pages of a portal. I had a set of questions about the target audience, the goals, and the USP, an industry research document, plus access to Google Analytics and Clarity.

My starting hypothesis (one you read more and more everywhere) was that AI is already more than good enough for this kind of thing. You barely have to work with it; the designer just has to look over the result, and on topics like this it basically replaces the human. Since a task like this happened to come along, I figured I’d check whether that’s really true.

The test

I worked with two models in parallel: Claude Fable 5 and ChatGPT-5.6 Sol, both on paid subscriptions, of course. I gave each of them everything I had as input: the brief, the questions and answers, the research material, the GA exports, the Clarity heatmaps, attention maps and scroll maps, plus the necessary URLs. The task was precisely defined: produce a UX expert review document, with a specific goal.

Importantly, I didn’t work with a single prompt. I broke the task down in a structured, multi-step way, and I asked the models to back up their findings with the GA and Clarity data where possible, rather than just assuming. In other words, I wasn’t testing the “throw everything into one prompt” approach, but what you can get out of a deliberate workflow.

And naturally, I also did the analysis by hand in parallel (like an animal), so I’d have something to compare against.

What I found

Let me say this up front. This is an experience report, not a controlled experiment. One project, one expert (myself), and my own analysis isn’t an absolute yardstick either. Read the following with that in mind.

Both models put together an analysis, but they were full of irrelevant, and even non-existent, problems. There were useful observations too, but even those I could only really use at the level of wording and structure. Claude was noticeably more accurate than ChatGPT, at least in phrasing and structure, though this is my subjective impression, not a controlled measurement. One difference was objective, namely that Claude could produce the result in a single document while ChatGPT only managed it in part, but that’s more a question of output format than of analytical quality.

I’ll admit it: there were problems I hadn’t spotted myself, and that was useful. But there were more that the AI, in turn, didn’t catch.

This actually lines up quite well with a 2025 study that tested GPT-4o against Nielsen’s usability heuristics:

“GPT-4o found roughly a fifth (21%) of the usability problems identified by the experts, while also raising several new, partly false problems. So it’s an excellent first-pass UX review tool, but on its own it doesn’t yet replace expert evaluation.”Can GPT-4o Evaluate Usability Like Human Experts? (2025)

On a project somewhat more complex than these Nielsen-style usability heuristics, I found essentially the same thing.

How the material finally came together

In the end I put the analysis together with Claude. I fed it my own observations, and in an ongoing conversation we shaped what to cut and what to add. It reviewed the text for grammar, consistency and similar aspects, and suggested changes. That part was incredibly useful.

Get Kolozsi István (kolboid)’s stories in your inbox

Join Medium for free to get updates from this writer.

Remember me for faster sign in

I didn’t measure the time precisely, but by feel: if I’d done everything on my own, it would have taken roughly 15% longer overall.

So AI is a great companion and helping partner that lets me speed up and sharpen my work, but right now it can’t replace me. It’s still a strong partner, and together we’re better than ever.

Press enter or click to view image in full size

You might ask whether a browser walk-through would have been better. These days there are ways and tools to have the models walk through the site live in a browser, rather than only working from static data and URLs. I can’t know for sure, since I didn’t run it, but I strongly suspect it wouldn’t have led to better results either, because the core limitation isn’t access, it’s judgment, meaning understanding the context and the goals, weighing severity, separating real problems from noise. A live walk-through changes little about that.

What I will actually try next is different. I’ll build my own agent skill for this task, one that encodes the established usability heuristics, the severity weighting shaped by professional experience, and the expert lens a UX practitioner brings to an interface, and then refine it project by project. The point isn’t for it to decide instead of me, but to turn my own expert judgment into a reusable, consistent form. That’s what I’ll look at next time.

This actually lines up with what you see in the research on the topic: Baymard’s 95%-accuracy AI heuristic evaluations weren’t achieved through better prompting, but by grounding the model in their own researched UX knowledge base. Granted, in their case this is narrowly about usability heuristics, which is a much simpler story than a complex, context-dependent review like this one, but the principle is the same. The accuracy doesn’t come from the model, it comes from the expert framework you feed into it. That’s exactly why it’s worth encoding your own standard into a skill.

And in the long run?

I’m well aware that sooner or later AI will replace our work, that’s essentially not in question. The only question is when it will happen.j

I’ve read that there are signs of a kind of slowdown, a plateau. According to a late-2024 Reuters piece, several leading AI researchers, including Ilya Sutskever, say that simply scaling up model pretraining has reached its limits. Further progress increasingly comes from new approaches, such as “o1”-style models that reason step by step, rather than from sheer size. In other words, development hasn’t stopped, but the emphasis is shifting from size toward smarter methods.

Of course, all sorts of other developments and research are going on in the background, so there’s no question that AI will gain total ground in the long run.

I recently stumbled on a page suggesting we still have a few good years left, especially in research and complex design, where we use various AI models everywhere to speed up, sharpen and support the work.

See when AI will replace your job!

And while we’re at it, I switch between models depending on which one I use for what, so don’t forget that every LLM is good at different things. Check out this 2026 LLM Compass!

Press enter or click to view image in full size

So what should you do tomorrow?

Start: use AI as a first-pass partner, feed it your real data, work in steps, and codify your own heuristics and severity into a reusable prompt or skill.

Stop: shipping its output as-is, expecting one prompt to replace a workflow, and handing over your judgment. Severity, prioritization and business context stay with you.

And stop panicking that it will take your job tomorrow. It won’t, but the designers who learn to direct it will outpace the ones who don’t.