How Dheemanth Reddy Taught a Speech Model to Laugh, Whisper, and Sound Angry

A sentence can be pronounced clearly and still feel lifeless. The words may be accurate, but the delivery carries no sense of mood or intent. Dheemanth Reddy built his work around that gap, focusing on the difference between generated speech that is understandable and speech that feels recognizably human. His goal was to give synthetic voices more of the expressive range people rely on every day.
Reddy develops voice artificial intelligence models that can communicate emotion and sound natural across languages. He has concentrated on qualities that often disappear beneath polished demonstrations. Many systems can read a script smoothly, but their voices remain flat when the meaning calls for warmth, tension, hesitation, or humor. Reddy saw that limitation as a technical problem that needed to be addressed inside the model.
"Speech carries meaning in the way it is delivered," Reddy says. "Getting the words right is only part of the job. The voice also has to reflect what the speaker is trying to communicate."
That principle shaped Maya 1, the open-weights speech model Reddy built with 21 distinct emotions. The model could produce laughter, whispers, anger, and other forms of expression that gave generated speech more variation. Reddy treated those capabilities as part of the system's core design. He wanted the model to respond to emotional context instead of applying the same delivery to every sentence.
Human voices shift constantly during conversation. People soften a phrase when they want to reassure someone, pause before giving a difficult answer, or change emphasis when a point matters. Those choices influence how a listener interprets the words. Reddy wanted Maya 1 to reflect more of that range without reducing emotion to a simple change in volume.
"A whisper changes how the listener receives the sentence," he says. "The model has to learn that delivery is connected to meaning."
Building that kind of speech required more than refining the model's output. Reddy also had to work around the uneven availability of high-quality speech data. English has received far more attention and investment than many other languages, leaving gaps in the material needed to train systems that sound natural to native listeners. Those shortages become even more significant when the goal includes emotional delivery and regional variation.
Reddy gathered real speech across individual dialects when strong datasets were not available. The work required close attention to the way pronunciation and pacing change within a language. It also reinforced his view that multilingual development must begin with the people who actually speak each language. A model trained around limited examples can produce technically correct words while still sounding unfamiliar to the listener.
"You cannot assume one version of a language represents everyone who speaks it," he says. "The model has to learn from real voices and real differences in how people communicate."
Maya 1 went on to receive more than 330,000 downloads and has been used by developers and companies worldwide. Its open-weights release gave people outside Reddy's team the ability to work with the model directly. They could test it in their own applications and evaluate how it handled speech that had not been selected for a controlled demonstration. That broader access fit Reddy's preference for outside scrutiny.
That work has also put Reddy in roles on both sides of technical evaluation. Maya 1 ranks among the world's top open-weights speech models and is the only India-based model on the Artificial Analysis / Speech Arena leaderboard. Reddy has separately served as a judge at the Cerebral Valley AI hackathon, where his role shifted from submitting his own work for scrutiny to assessing the work of other builders.
He wanted the model to be judged through use rather than a small collection of ideal samples. Carefully chosen demonstrations can conceal weaknesses, while an open release exposes a system to unfamiliar prompts and new settings. Reddy viewed that process as part of technical development because it allowed other users to identify strengths and limitations that the original team might not have anticipated.
"Once people can use the model themselves, you get a more honest picture," he says. "They try things you did not plan for, and that shows you where the work still needs to improve."
The download count and worldwide use gave Maya 1 a public reach beyond Maya Research. Those facts did not explain every reason people chose the model, but they showed that developers and companies were willing to incorporate it into their own work. For Reddy, that reach created more opportunities for the model to be tested outside the environment where it had been built.
His larger goal was to move synthetic speech closer to real conversation. Emotion mattered because a conversational system had to respond in a way that fit what had already been said. The model also needed to preserve context and recognize what a speaker meant before producing a response. Voice quality alone could not carry an interaction if the system failed to follow the exchange.
Reddy believes many speech systems still operate like script readers, even when the audio sounds polished. The voice may be fluent, but the interaction can break down when the system has to interpret a continuing conversation. His work connects expressive delivery with understanding, because both influence whether a response feels appropriate to the moment.
"People hear the words and the voice together," he says. "The model has to understand that those two things are connected."
Reddy's interest in voice technology began with people who struggled to use English-first interfaces, including members of his family in Vempalli, Andhra Pradesh. That experience shaped the problem he chose to pursue, while the development of Maya 1 required years of technical training and experimentation. He later earned a master's degree in Computing, Entrepreneurship, and Innovation from New York University, where he studied deep learning under Yann LeCun.
Reddy now serves as co-founder and CEO of Maya Research, which is backed by South Park Commons. He also received an Emergent Ventures grant for his work on voice AI that sounds native. His work remains focused on building models that communicate with greater emotional and linguistic accuracy.
Maya 1 showed how expressive speech could be approached through concrete model capabilities and made available for outside use. For Reddy, the achievement was the creation of a system that could vary its delivery in ways tied to human communication. It gave him a technical foundation for continuing to study how machines can speak with greater emotional awareness.
"When the delivery changes, the meaning can change with it," Reddy says. "Voice AI has to learn that relationship."
© Copyright IBTimes 2025. All rights reserved.

























