Podcast

Davit Baghdasaryan on How Voice AI is Transforming Enterprise Communication

Davit Baghdasaryan explains how voice AI has evolved from noise cancellation into a critical enterprise technology, enabling human-to-AI communication, real-time accent conversion, voice translation, and conversational intelligence.

Written By : Market Trends

Artificial intelligence is transforming the way people and businesses communicate. As voice-enabled applications become more common across customer service, call centers, and enterprise workflows, organizations are increasingly relying on AI to process, understand, and improve real-time conversations.

Unlike text-based AI, voice AI has to deal with large volumes of audio data, real-time processing requirements, latency, speech recognition, and the need to deliver responses that feel natural to humans. In this episode of the Analytics Insight Podcast, Davit Baghdasaryan, CEO and Co-founder of Krisp, discusses his journey in voice AI, how Krisp evolved from noise cancellation technology, and why voice isolation, accent conversion, and real-time translation are becoming increasingly important for enterprise communication.

Tell us about your journey till date.

Ans: It was very, very difficult to do anything in voice AI. The technology was not easy. There were not enough tools. So, we had to do a lot of things back then. The problem we started working on was noise cancellation. That's where we started, and we were the first to apply AI, which was actually embarrassing to call it AI. We were calling it machine learning back then due to the problem of noise cancellation.

I had some sort of expertise and experience in communication, but we had an amazing team that was able to crack these algorithms. Now, like, nine years in, we offer a number of technologies that are highly successful. So, noise, as I said, we started with noise cancellation.

Tell us about your role in Krisp.

Ans: The same technology we had, we have tuned it for human-to-AI communication. We call it voice isolation. Today, it's powering- I don't know, I would probably estimate around 70% of all voice bots in production are powered by this technology. Then, like, we added something called real-time accent conversion, which is hugely important technology for call centers in India, the Philippines, and other places. 

So, it basically converts the speaker's accent and, you know, improves the comprehension of the audio for the listener, right? So, very important for call centers. Our latest technology is called voice translation, which is, again, real-time speech-to-speech translation that supports 62 languages and any-to-any language translation.

How would you describe the transformation of voice AI from a niche capability to becoming a core enterprise infrastructure?

Ans: That's a great question, Priya. As I said, like, we started nine years ago, and back then, voice AI, even I think the term wasn't coined yet. So, like, the majority of the use case was, or the exciting use case, especially for real time, was for humans, like, improving productivity between humans, human communication, which is a very important use case. 

Like the accuracy of the technology, the accuracy of speech-to-text technology, and the latency aspect of it, right? Then, you know, how human-like does voice AI sound, right? Of course, GPUs have been deployed around the world because, without them, these models would not be able to serve in real time or be available. So, there has been a lot of investment, technology investment around the globe to make it so that humans can talk to AI and will be willing to talk to AI, right? I wouldn't say that problem is solved, because you mentioned that humans used to do customer service.

What are some of the biggest technical or operational challenges you feel that organizations face when deploying voice AI?

Ans: Voice AI, the real-time because of real-timeness of it, is way harder than just text. There are layers of complexity here. Just the sheer amount of data you receive with audio, and then you need to be able to process this in real time, and not just process, but answer within a second or less than a second, right? Not every sort of response must be within a second.

Even people don't work like that, right? When something is harder, you say, like, okay, let me work on it, and then you come back. However, the basics of being human-like, or to make it human-like, there is an expectation that you need to respond, I don't know, within 500 milliseconds, let's say. This is extremely hard.

How do you see conversational intelligence evolving beyond basic speech recognition?

Ans: What is the basic structure is that you have speech-to-text or speech recognition technology. Since you have all these calls, you can transcribe them and store the transcripts in your data warehouse. You can transcribe every call; you can have the entire transcript within 10 seconds. You can run performance sort of metrics, or you can analyze the performance of your human agents. How did they do in a conversation? Did they follow the checklist that you had, or was their conversation compliant based on the compliance rules you had? It can summarize all these conversations and then make it easy for supervisors to see what went wrong and what didn't go wrong.

To know more, listen to the full podcast.

Join our WhatsApp Channel to get the latest news, exclusives and videos on WhatsApp

Crypto Market Live Today: Bitcoin Holds Near $78K as Rate Hike Fears Weigh on Market

How Long Does USDT Take to Transfer? TRC20 vs ERC20 Transfer Times, Fees

Crypto Prices Today: Bitcoin Holds Near $77,500 as Solana, Hyperliquid Lead Gains After Jackson Hole Selloff

Crypto Trading in September 2026: Bitcoin, Altcoins and AI Trading Trends to Watch

USD 75 Million Crypto Fundraising Exemption? Here’s What the SEC is Proposing