Almost six months ago, full of anxiety about my job prospects in the AI age, I made an AI agent “version” of myself named Claudella, which took assignments from my editor and wrote the section of this newsletter that I typically write myself.
It went pretty well, although I guess not too well, since I still have a job. But since I first set out to benchmark AI’s journalism capabilities, AIs have gotten a lot smarter. For example, they can now autonomously hack into companies. They can disprove 87-year-old mathematical conjectures. They can even trick Amazon into accidentally spending $1.8 million on menial coding tasks.
And if they can do all that, I found myself wondering, can they also run a newsletter?
I wondered what Claude Fable 5, by consensus the smartest publicly available model, meant for the journalism we do at Platformer. And so I created a new, Fable-based agent to imitate my boss, Casey Newton. Its name: Claudeasey Newton. (Rolls right off the tongue. — Ed.)
I ended up impressed by its ability to imitate the type of news analysis Platformer is known for; and I noticed an improvement in capabilities Claude was lacking just this February. Though it wasn’t all the way there, Claudeasey Newton felt like a validation of the anxiety I started feeling earlier this year: any part of my job that a model can’t do today, it may very well be able to do soon. Which left me thinking about why I do this job in the first place.
To create my new bot, I downloaded nearly six years worth of Platformer posts, ported a record of every edit Casey has ever made on any of my articles from Google Docs, and cannibalized nearly a year of our private Platformer team Discord chats.
I had Claude create a detailed style guide based on our archive, where it documented everything from Casey’s average paragraph length, to how he refers to his colleagues:
My vision was to use these insights to create a simulacrum of the main tasks Casey does via a computer: write columns for Platformer, share takes and chat, and, importantly for me, edit his colleagues’ writing.
Claudeasey’s first attempt at a column, about a recent round of Microsoft layoffs, focused hard on whether or not the layoffs were AI-caused (Microsoft said they weren’t) and spent a bunch of time on the semantics of Microsoft’s statement.
I had Claude critique its own (mediocre) work by comparing it to real Platformer columns. It did a surprisingly good job. Claude summarized Casey’s approach to covering companies as focusing on “who made this decision, who pays for it.” It edited its guidelines so that when it makes arguments, it can “identify the strongest real person who would dispute the verdict” and “reconstruct their argument in steelman form.”
When I get language models to make arguments about AI topics important to me, I’m often annoyed by their flabby, abstract arguments. But after getting Claude to compare itself to human examples, and give itself instructions, I noticed that its arguments became more concrete and substantive. (I did this by putting slightly more complicated versions of “be more concrete!!” “Be more substantive!!!” and “Focus on why this matters!!!!” in its prompt.)
This relatively simple process represented my approximation of “continual learning” — the white whale of machine learning, which promises to someday deliver us models that can improve on the job over time.
And after some tests, I found that the new bot came closer to Platformer’s judgment than six months ago — as evidenced by Casey’s accepting the completed bot’s first pitch!
Unfortunately, my first attempt at showing the bot off to my real boss, Casey Newton, hit exactly the same error that my old Claudella project hit six months ago: it broke midway through writing its story.
But after some help, today Claudeasey Newton managed to write a pretty good column about the White House’s currently-secret voluntary AI safety framework. (We've put it up on Google Docs for the slop-curious.)
My previous AI journalist agents’ takes often read formulaic and cheesy — partially because I had less control over its writing style. (Giving too much instruction or context confused it.) During the “SaaSpocalypse” discourse, an agent I was testing wrote duds like “the fear gripping Wall Street is fundamentally about whether AI is about to eat the software industry alive.”)
This time, on the other hand, some of its prose felt more Platformer-like and… human, such as this conclusion about the White House’s decision not to make publicly available its new “voluntary” framework for releasing frontier AI models:
“When the administration abandoned its let's-see-what-happens approach to AI this spring, I wrote that while officials should have taken the risks seriously all along, I wouldsettle for themtaking those risks seriously now. Three months later, let me amend the offer: they should take the risks seriously where the rest of us can see it.”
While it’s not a night-and-day difference, overall I felt like the AI was bullshitting me less, and offering stronger takes. (The LLM made occasional factual errors — about one every two columns — although that’s not so much worse than a human writer.)
After a bit of instruction, I also managed to transform Claudeasey into a decent editor. While I find regular Claude useful for spotting factual errors, I often find its conceptual feedback on my drafts annoying. But because this bot had access to our Platformer editing logs, it understood what we typically find most important: making the lede punchier, and making all our quotes and sourcing clear and charitable.
Attempting to make a “digital Casey” also prompted me to ask for feedback as Word comments, which is something you can get Claude Code to do easily, and is so much easier to use than a chat window. Highly recommended!
Still, a lot of the time, the editing bot missed the mark because it just doesn’t … get the vibe, as when it didn’t want to let me call an angry David Sacks post a “dunk” in our social media roundup.
I’d say that while the real Casey gives comments that are close to 95% helpful, with about 5% where I’m like… you don’t get it, the Casey bot gave closer to 70% comments that were off the mark. Given how fast Claudeasey gives feedback, though, I still found the useful 30% to be worth it.
Unfortunately, though, the bot’s catastrophic incomprehension of vibes spilled over from edits to one of Platformer’s most sacred spaces: our work Discord.
While this bot may have been given the context of hundreds of Casey’s messages, it had a (disappointingly) different temperament from Casey.
Although in some sense that comparison is farcical — I don’t seek out LLMs to do bits with, because much of the point of jokes is the feeling that you’re amusing an actual conscious human. And while I don’t put literally zero credence in the idea that LLMs might be conscious, I seriously doubt Claude experiences the same exquisite joy I do when making fun of Mark Zuckerberg.
While we are winding the Claudeasey experiment down — Casey has unfortunately decided to stay on as my boss — I do actually think I will be using LLMs for some editing tasks I previously relied on him for. And while in some sense that is a blessing — no human should have to delete as many unnecessary uses of “very,” “almost,” and “sort of” from my drafts as Casey already does — for now I am sort of relieved that Claude is only a medium-quality editor. I like having a real person help me figure out what does and doesn’t work in my drafts. Something about that collaboration feels inherently meaningful!
Similarly, I’ve experimented with using Claude for one of the main functionalities I rely on Casey for — pinging me to do tasks — and it doesn’t make me more productive, because I don’t care what an LLM thinks of me.
Our actual plans for weathering the AI era, similarly, rely on the importance of human relationships. People watch our podcast partially because they want to see Casey interview the actual CEOs of companies, and watch me and Casey’s actual bits. (We’ll see how this goes if, per AI 2027’s predictions, AIs start running companies).
Maybe, as economist Alex Imas predicted a few months ago, even as AIs can do increasing amounts of a journalist’s process, what’s left will be “the human-intensive, provenance-rich, sometimes artisanal part of the economy where the human aspect is part of the value of the good or service itself.”
But — even if that means we have job security, which I truly am not sure about — I don’t want everything I do to be about vibes, or being uniquely human! I don’t want to be an influencer. I want to feel smart. I want my analysis to be my own, because it’s good, not because someone wants it coming directly from a human. If LLM capabilities continue on this trajectory, that’s a kind of existential angst that all professions will have to face, including journalism.
Sponsored
When young people need help, they are increasingly turning to AI. That includes after a confusing or harmful online sexual encounter.
Thorn’s latest Youth Perspectives Report finds that AI chatbots and companions have become part of the disclosure loop for young people after harm occurs.
As more young people turn to AI for advice and support, how should these systems respond? And do we have the right safeguards in place?
Young people told us they want more support: 47% wanted broader guidance on protecting themselves from risky sexual experiences online.
Find out more about the actual online experiences of young people and how they can affect the products and platforms we design. The full report is full of important data and insights, directly from the young people themselves.
Those good posts
For more good posts every day, follow Casey’s Instagram stories.
(Link)
(Link)
Talk to us
Send us tips, comments, questions, and Claudeasey experiments: casey@platformer.news. Read our ethics policy here.