Rachel Feltman: For Scientific American’s Science Quickly, I’m Rachel Feltman. Today we’ll be hearing from a few folks, one of whom is a young man named Jakob Jordan.
Jakob Jordan: Hi, my name is Jakob.
Feltman: Jakob has autism and apraxia, which is a disorder of the brain and nervous system that keeps people from carrying out physical actions at their own command. In Jakob’s case, it can be difficult or even impossible to speak his thoughts aloud.
On supporting science journalism
If you're enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.
So it might surprise you to learn that Jakob just starred in an opera called Sensorium Ex.
[CLIP from a song from the Sensorium Ex album entitled “What I Am.”]
Feltman: In a moment, you’ll hear Jakob using a custom-built AI speech synthesis device that made this opera sing—one that was specifically designed to sound like him and, crucially, to give him expressive control over his own voice.
Jordan: Hearing my voice produce my desired words feels so exhilarating. My mouth is constantly spouting off repetitive loops that I get stuck in. It is not my intended speech, so it can be frustrating when I can’t express my intentions. Sensorium Ex gave me the gift of hearing my trapped thoughts come to life. It’s like my voice was released from prison after a wrongful conviction. I’m so grateful to the Sensorium Ex team for making dreams that felt impossible come true.
Feltman: That sound bite was processed with the AI tech that Jakob was referring to, but later on we’ll hear how it contrasts with his more traditional text-to-speech device.
Along with Jakob, today we’ll be chatting with New York University professor Luke DuBois, who developed Jakob’s AI speech synthesis device, and composer Paola Prestini, who wrote the opera this tech was designed for. Their work together raises a radical question about AI: What if we built a tool not to replace human creativity but to amplify it?
Feltman: Paola, how did the idea of this opera first arise?
Paola Prestini: It started with a commission from an opera company that then didn’t end up participating in the final run. But they wanted me to do something in relation, actually, to blindness, and at that point I was very, you know, determined to only do something that I had lived experience in or collaborate with someone who had. And I was already working with this amazing poet Brenda Shaughnessy, and Brenda has lived experience as a mother to a nonverbal and nonambulatory child. And so I went back to the opera company, and I was like, “Listen, you know, this is something that we really feel we could, a story we could tell.”
And from there we refined it, and eventually my company VisionIntoArt really took over because we learned in taking this kind of intersectional approach to disability that we really had to create the whole framework, uh, you know, not just the opera.
Feltman: I’m really curious what the reaction was to your proposal to work with a different group.
Prestini: So I think what I can say with certainty is that a work like Sensorium Ex doesn’t really exist in the opera field. Even though the word “opera” means “to work” and historically was really the place of innovation, I think in contemporary society, it doesn’t exist that way anymore. It’s really about, you know, preserving the canon. And when it is done to commission new work, it generally doesn’t treat the stage as an R & D [research and development] lab, which is what we were essentially doing. And so I wanna be graceful in saying that, you know, my desire for the piece was to honor what was in the libretto—and the libretto was asking questions that I didn’t know how to answer without collaborating with someone, you know, like Luke or like Jakob. And to honor, also, the fact that the characters were written with certain disabilities, I wanted to really make sure that the context for disability would thrive. And so what I was being asked to do was not just create an opera but create a framework for success for the performers.
And so that’s when I learned, well, I have to be my own opera company, you know? Like, I have to collaborate with folks who will trust me, and I will make mistakes, but I’ll do so in the most empathic way and the most intelligent way I could, which was, of course, then collaborating with experts and collaborating with folks who I could learn from and then translate that into a system.
Feltman: Well, and that’s a great segue into Luke. What did you think of this technological challenge that was being set for you?
Luke DuBois: I mean, it’s a good one, right? If you think about an opera, an opera is a great excuse to do crazy things with voices.
Feltman: Mm-hmm.
DuBois: And speech synthesis technology, right?—so, you know, having a computer synthesize a voice—has had a really interesting history, and one of the things that we focused on in this project was looking at the things that a speech synthesizer typically is not designed to do.
So speech synthesizers were originally designed as a disability aid to help people who are visually impaired hear books. So if you go back to Ray Kurzweil, who sort of developed this stuff in the 1970s, it was married to an optical character recognition system so that a blind, low-vision individual could visit a library and hear a book, right?
And so if you think about philosophically what that means is that means the synthesized voice is the computer’s voice. So you don’t really care if it kind of speaks in a monotone, or it’s always the same voice, or it’s kind of never changing, or it just takes a string of text and just-flat out says it with no expressivity.
But at a certain point, starting in the 1990s really, we started building these augmentative and alternative communication devices, these things called AAC devices, where you’ve got a speech synthesizer acting as a proxy for a person like Jakob, and now it’s not a computer’s voice, it’s a person’s voice.
And so things like personalization and expressivity become really important. But they didn’t really change the voices at first, right? And so the challenge that we were trying to rise to here is: Can we make a system that Jakob and everybody else who might use this thing really feels like it sounds like them, and they can control it, just in the way that you and I are speaking now?
And that is one of the better kind of fun design challenges we’ve ever been faced with, right? I work at a thing at NYU called The Ability Project, right? And we are a design shop around technology and the social model of disability. We do all kinds of things, and this has been an incredibly exciting collaboration ’cause it ticks all the boxes of “nothing about us without us,” sort of, working with community with lived experience. It’s technologically innovative, but it is social-model technologically innovative; I am not fixing anyone, I am fixing the stuff they use. And that’s what makes it really fun, and it’s just been really great to hang out with everybody that’s been working on this project. I work at an engineering school, so it’s really great to escape once in a while and do some art.
Feltman: Yeah. Could you walk me through what aspects of the system are, are really unique and what sort of R & D went into making that possible?
DuBois: Yeah. So the first thing we do at The Ability Project is we do design research. Like, we don’t dive in and start trying to fix shit. We try to, like, think through: What are some, you know, ways in which we can do some stakeholder equity around the problem?
So we talked to lots of people who use these things, and the thing that lots of AAC users come back with is these three problems. The first problem we didn’t really fix, which is: It’s a pain in the butt to enter what you want to say, so that’s our next challenge. But the two things we thought, like, “Oh, we can go after this,” is: “This thing does not sound like I think I should sound, and while I’m talking, I have no agency on the delivery.” You typically enter in the full thing you want to say, and then it just edits it. So you say, “I want a cheeseburger,” and then it says, “I want a cheeseburger.” You can maybe stop it, but you can’t, like, make it fast or slow or nuance it or whatever. So one of the best pieces of stakeholder information we got was a line that people want the right to be able to inject emotional affect into their delivery. Right? You wanna be able to say, “I’m fine,” to your romantic partner in a way that it is very obvious that you are not fine. Right?
Feltman: Yeah.
DuBois: So then it is no longer just in the space of opera, it becomes like a communication civil rights issue. Like, you know, everyone should have the right to be sarcastic. Everyone should have the right to sound overjoyed. Everyone should have the right to sound hangry or upset or happy or whatever. So how do you get there? And so what we did was two things. We wrote some tech so that we can sample someone’s speech, right? Disordered or otherwise. So we can take what Jakob says normally, take that recording, and do something called an inference model. So the key word is “infer.” And so what we are doing from that recording is we are inferring the architecture of the voice that made that sound, right? So the computer will extrapolate from the recording 401 parameters that represent all the stuff that make your voice you. And it’s really your voice, not your language, so it’s not, like, your choice of words or your accent, but it is, like, the width of your shoulders and the depth of your nasal cavity and the tightness of your tongue against your teeth when you make an S sound and all that stuff, right?
So you get all those numbers, and then I can run normal language through it, and it will synthesize sound that sounds like that person. It does it all on a local computer, does it all on what we call local metal. It’s not using a cloud service. It’s not using something where there’s, like, a third-party privacy issue around your data. It’s all there. If you delete it, everything goes away. And then we took some lidar sensors. So these are the things on your automobile that deploy airbags in a crash, right? A modern car has a bunch of things that are 3,000 times a second being like, “Is there anything close to me? Is there anything close to me? Is there anything close to me?” And if something is close to you and it’s moving toward you, you know something’s about to hit you, so you deploy an airbag. So those things are about $6. And so we got a bunch of them, and we can put them on a platform or a tray or for a wheelchair user, we can attach them to a wheelchair, whatever. And then we can map them to different parameters of the speech as it’s being delivered. What Jakob can do is: he can start the playback and then speed it up and slow it down and change its pitch. So instead of it being like, [monotone] “Hi, Rachel, how are you?” It can be, “Hi, Rachel. [Beat.] How are you?” And you practice it just like any other musical instrument; you get good at it, right? Or just, like, you practice with your voice, and then you can do these really expressive things.
Feltman: Jakob still uses a more traditional text-to-speech device for everyday life, including his conversation with me. But the folks at Sensorium AI were happy to render his answers after the fact so we could share his words in the voice he prefers. Let’s hear more from Jakob about his experience getting involved in the opera world. For this next question, he’s joined by Julie Sando.
Julie Sando: I’m Jakob’s communication partner. He communicates through text-based communication, so he types out every word on a keyboard here. He’s prepared a lot of his responses in advance ’cause it is a slower process.
Feltman: Thanks so much for coming on to chat with us, Jakob. And, and Julie, thank you for being here to help. Listeners, Jakob is going to answer my next question with his standard text-based communication software so you can hear what it sounds like. But then we’ll switch back to an answer rendered by Sensorium AI to remind you what he sounds like.
So Jakob, how did you first become aware of and involved with the project?
Jordan: We have a social media page called Cards by Jakob where we share about my business and life. A casting director reached out and offered me two auditions, one of which was the opera. To my surprise, I got both. It sure has given me the acting bug. I absolutely feel most alive when I am onstage.
Feltman: Yeah, I totally get that. I, I love performing as well, and there’s really nothing like, like doing it live. I’m curious, what was your initial reaction to, to the idea of being in an opera?
Jordan: I was shocked. I honestly thought it must be some sort of charity that was giving us opportunities in some small-scale performance where parents are the majority of the audience. I couldn’t have been more wrong. I was blown away to witness the depth of skill I was surrounded by at that first workshop. I am still convinced this must be a dream. I mean, I am someone who can’t control my voice, and I have a lead role in an opera. Pinch me.
Feltman: Yeah. And Paola, I’m curious: What felt important for you, for the technology to have, in terms of the actual operatic performance?
Prestini: Yeah. I mean, Luke and I talked about this early on, which is that we wanted the ability, you know, for it to kind of take the best aspects of a synthesizer and be able to create, you know, not just the kind of actual tonality of the voice, which is what’s so interesting on a sociological level, but to also make it a musical instrument.
And in the opera, basically, there’s a moment where Jakob’s character, Kitsune, manifests their escape. And it’s this beautiful musical moment where, you know, Jakob essentially DJs his own voice and is able to kind of layer and create expressivity and, you know, be a musical instrument, if you will.
[CLIP from a song from the Sensorium Ex album entitled “What I Am.”]
Prestini: And I think that’s something that, you know, I think we will play with continuously as we build the instrument. And I, I have to say, you know, imagine building an instrument and then being told you have, like, six weeks to learn how to use that instrument, and that’s essentially what we did to Jakob because, you know, the funding for this came through very late, and Luke is not just doing Sensorium AI, he leads an entire department. And so we were, you know, building as quickly as we could, but essentially this came together, like, a month and a half before we went onstage.
And so each time we’re able to do it, there’s a refinement process and a learning of the instrument, which is really wonderful. But for me, it was really, again, this question of saying, “How could we build something that is not just essential to the story, but that’s actually potentially something that can contribute positively to Jakob’s life and to anyone who uses it?” And I think we’re in the step in the right direction.
Feltman: Yeah. And what applications are, are you hoping to see for this beyond the opera?
Prestini: Yeah. I mean, so I would say that, you know, from a kind of equity perspective that this question of really not understanding what it takes to create a thriving ecosystem that needs all the voices that we have in society is addressed through: one, seeing the opera; two, using the open-source Codex that we’re building to help folks understand what the steps we took that, of course, are not the only steps you can take, but for folks, you know, who are afraid of the work or of putting disability onstage, that they have a tool.
But yeah, my hope is that the stage [looks] like the life we live in. And I think when we complain in the arts about our work not having audiences, it’s like, we also have to take our part in building work that’s really representative of the world we wanna see. And so that means really creating a thriving ecosystem that if, you know, somebody says no to you, you don’t just stop and say, “Okay, well, I’ll just do something else.” But it’s like, well, no, what can I do in my power to really make sure that this is seen.
And composers, like, what we do is we create from the abstract, right? Like, we create something out of nothing, and so we are equipped in many ways to be able to do more than just create sound, right? But to create worlds. And so I feel like I probably got really far away from your question, but my hope is that, um, things like this become more everyday. You know, that art continue[s] to ask really hard questions so that, you know, things that don’t exist do exist.
Feltman: Luke, anything to add to that?
DuBois: Yeah, I mean, all of that is totally the plan. We are continuing to develop the synthesizer, and what we’re trying to do is get it to a stage where we can test run it with a lot of people. And a lot of that is about figuring out a way to get it running on, you know, the tablet computers that, that folks like Julie and Jakob use to drive a speech synthesizer and stuff like that. So we’re still working on it.
Another fun thing about collaborating with Paola on this as an opera is, you know, you get people from, like, all over the place who write you once they hear about it, and they want in on it. So I have, you know, hundreds of e-mails from people who are sort of waiting to try this thing.
Prestini: This idea that performance becomes the lab for innovation, I think is what we need to get back to as artists, you know? That we’re not building things to be perfect. We’re building things to really try to create what doesn’t exist.
When we built this project, we built a really large disability advisory board, and one of our friends on it is Jacob Nossell, who’s an activist from Copenhagen, and he really coined the phrase that we use with this project, which is, you know, “To have a voice is everything.” And so I feel like, you know, everything we’ve done is in service of that message.
Feltman: Yeah. Luke, I’m curious, here we cover so many stories about AI and, you know, about companies that sort of imagine that generative AI can supplant creative humans and this use is so much more helpful and interesting, and I’m just curious if you have any thoughts on, you know, what, this project can maybe tell us or teach us about AI and its use cases?
DuBois: Good question. I run the design program at the Tandon School of Engineering, so we’re the little art and design shop within this broader school, and we are often engaging with students, and to some degree our colleagues, around this very question about, like, you know: How do you do AI in a way so that AI is not just a mechanism to give people with wealth access to skill without giving people with skill access to wealth, right?
Feltman: Great way to put it. [Laughs.]
DuBois: That’s the bad trap, right?
If you think about it, we’ve had the technology to completely replace human performance for decades now—you don’t need AI to do it. There are these things called cartoons, right? And there are things like film and whatever. People see live musicians despite the fact that there are records out there, right? So why? They see that because they wanna see people doing things in real time where they believe in themselves, and they’re collaborating and all that stuff. And that really gives me hope, you know, against this idea that AI is this thing that’s gonna replace creative labor. I don’t think that’s true. every civilization uses the maximum level of technology available to make art. That’s a pretty normal thing. And hopefully what it does is it just adds into the stack.
So it just ends up becoming this thing that integrates, and I think AI will end up being something like that. It’ll be a set of tools to help people do some interesting things. Sensorium AI is really fun to talk about in the context of AI and ethics because it’s a use of machine learning that augments human interaction and human agency instead of replacing it, right? So if anything, it’s a labor-creating tool, right? It gives Jakob more opportunities to talk. And so I think flipping the question on its head to say, you know, if we were to situate AI into a thing that, like, foregrounds human creativity, this is a good example of it.
Feltman: And Jakob, the idea of working with AI? Is it, you know, something you had any preexisting feelings about?
Jordan: It is an honor to receive the benefits of AI and to be on the good side of the revolution. The world needs more Lukes and Paolas who are advancing the industry for the greater good.
Feltman: How do you think that your experience compares to and contrasts with that of your average performer onstage?
Jordan: It’s hard to say, but perhaps the level of gratitude is enhanced by my inability to produce purposeful speech while starring in a groundbreaking opera. I had the opportunity to ask Donnie Wahlberg for advice, and his answer showed me that we may be more alike than different. He said he comes back to gratitude in every performance. I have taken that advice to heart.
Feltman: What are you hoping that viewers will take away from your performance?
Jordan: I hope that people see me and remember to make dreams that feel absurdly unrealistic. I want them to connect more deeply with their own voices that might be hidden away, like a false prisoner held captive by walls that we don’t know we have the power to knock down.
Feltman: And how has Sensorium AI changed the way that you interact with the world?
Jordan: I have a completely altered view of myself now that I see what is possible. Before Sensorium Ex, I lived a small life without knowing how small it was. My existence is now wide open and full of confidence. I will show people a fuller version of myself with access to the Sensorium AI. Nothing brings me greater joy than freeing more voices, so I will continue interacting with as many people as I can.
Feltman: Well, Jakob, you mentioned that you have gotten the performance bug. Do you have any, uh, future plans past the opera?
Jordan: I am honored to have a role in a feature film that has yet to film. I signed an NDA, so I wish I could share more. But stay tuned to my Facebook page at Cards by Jakob.
DuBois: I would be delighted if the outcome of this is that Jakob becomes a movie star.
Prestini: Yes. I love that.
Feltman: Thank you all so much for coming on to chat with us. This has been great.
That’s all for today’s episode. We’ll be back on Monday with our weekly news roundup.
Science Quickly is produced by me, Rachel Feltman, along with Fonda Mwangi, Sushmita Pathak and Jeff DelViscio. This episode was edited by Alex Sugiura. Marielle Issa and Aaron Shattuck fact-check our show. Our theme music was composed by Dominic Smith. Subscribe to Scientific American for more up-to-date and in-depth science news.
For Scientific American, this is Rachel Feltman. Have a great weekend!