

Thousands of listeners heard competing Hindi voice samples without being told which company had produced them. Their choices shaped a public leaderboard based on blind comparison. Dheemanth Reddy’s Maya 2 reached the number one position for Hindi on Voice Arena through more than 15,000 human votes, placing it ahead of models from ElevenLabs and Cartesia. The result gave Reddy an outside measure of the quality he had been working to achieve.
Reddy builds voice artificial intelligence models designed to sound natural in the languages people use every day. He believes listeners can recognize details that a product description cannot capture. A voice may pronounce every word clearly and still feel unfamiliar because its rhythm does not match the language as people actually speak it. Blind voting placed that judgment with the audience hearing the samples.
“You do not have to explain natural speech to someone who has spoken the language their whole life,” Reddy says. “They hear the voice and know whether it feels right.”
That evaluation method suited the way Reddy approaches technical proof. Voice Arena removed brand recognition from the comparison and allowed each sample to stand on what listeners heard. The voters did not know which model Reddy had created, which meant his name and company could not influence their choices. Each response became part of a larger record created by people outside his team.
The ranking also addressed a question that has guided much of his work: What does it mean for a machine to sound native? For Reddy, basic intelligibility is only the starting point. He considers how a sentence moves, how the delivery affects meaning, and whether the voice sounds as though it belongs in the conversation. Those qualities become especially important when a model is expected to interact with someone instead of reading a prepared script.
“Native listeners can hear when the words are right but the voice still feels unfamiliar,” he says. “That difference matters.”
Reddy believes the voice AI field has devoted too much attention to English demonstrations that perform well under limited conditions. Success in one language does not automatically carry into another. A model built around English patterns may struggle elsewhere because each language has its own cadence and expressive habits. Maya 2 was developed around the expectation that Hindi listeners should receive the same level of care.
This focus separates Reddy’s work from the practice of treating additional languages as later extensions. His goal is to build speech that reflects how people communicate within the language itself. That requires close attention to the sound of the voice and the model’s ability to understand what the speaker means. A system cannot sustain a convincing exchange when the voice feels unnatural or the conversation loses its direction.
“The language cannot feel like an added feature,” Reddy says. “It has to feel like the model was built to speak with that person from the beginning.”
Voice Arena gave Reddy a public way to test that standard. Listeners were not responding to the model’s size or the reputation of the organization behind it. They heard the available samples and selected the ones they preferred. More than 15,000 blind votes contributed to Maya 2’s Hindi ranking, giving the result a basis beyond Reddy’s own assessment.
He does not present the ranking as proof that every problem in multilingual voice has been solved. The result identifies one area in which the model performed strongly under public comparison. It also supports his belief that listening tests should play an important role when naturalness is being evaluated. The people who speak the language bring knowledge that cannot be replaced by a company’s description of its own model.
“Voice is experienced by hearing it,” he says. “The people listening should have a real role in deciding whether the model works.”
Maya 2 followed Reddy’s earlier work on Maya 1, an open-weights speech model made available for developers to use directly. The later model faced a different form of outside judgment through blind listener preference. It had to earn votes without depending on Reddy’s explanation of what the audience was supposed to hear.
Reddy earned a master’s degree in Computing, Entrepreneurship, and Innovation from New York University. He now serves as co-founder and CEO of Maya Research, where he builds models and directs the technical work behind them. He was also awarded an Emergent Ventures grant for developing voice AI that sounds native.
As co-founder and CEO of Maya Research, Reddy is building that work within a company backed by South Park Commons. His earlier Maya 1 model is one of the world’s top open-weights speech models and the only model from India on the Artificial Analysis / Speech Arena leaderboard. His technical experience has also led to an evaluative role outside Maya Research, including serving as a judge at the Cerebral Valley AI hackathon.
The significance of Maya 2’s Hindi ranking lies in the standard it applies. Reddy wants multilingual models to be evaluated within the languages they claim to speak well. A successful English demonstration says little about whether a Hindi voice will sound familiar to someone who hears the language at home. Each language deserves direct testing by listeners who can recognize its details.
Reddy expects people to use speech for more of the tasks that currently require typing or moving through menus. Those interactions will depend on whether the system understands the speaker and responds in a voice that feels appropriate. Poor native-language quality can create distance before the conversation has properly begun. A model that sounds natural is more likely to help the user focus on the task rather than on the interface's limitations.
Reddy sees the Hindi leaderboard as one public checkpoint in a much longer technical effort. It showed that Maya 2 could compete through blind listening and placed native speakers at the center of the evaluation. For him, that is where judgment of a voice model belongs.
“People should be able to hear the voice and trust their own judgment,” Reddy says. “That is the standard the model has to meet.”