Rime, an AI voice technology company developing speech models for enterprise applications, has raised $24 million in funding led by M13, with participation from Corazon Capital, Unusual Ventures, Cadenza Ventures and Twilio Ventures. The investment will support development of voice interaction models designed to power natural, real-time conversations with AI across complex and regulated environments.
Building the Voice Layer for AI
Rime develops speech infrastructure designed to make voice a primary interface for interacting with artificial intelligence. The company’s models already support voice applications used by Fortune 500 companies, combining advanced AI research with linguistic expertise to create more natural and reliable conversational experiences.
"The design patterns that will define the speech-to-speech era of AI haven't been built yet. Rime is doing the foundational work by combining frontier AI with deep linguistic expertise to create the speech infrastructure the next generation of AI products will rely on. Lily, Ares and Brooke are building Rime with the rare combination of world-class AI research and deep linguistic expertise to deploy real time voice people can trust in complex, regulated environments." - Morgan Blumberg, Partner, M13
Advancing Conversational Intelligence
The funding will accelerate Rime’s work on AI models built specifically for speech-to-speech interaction. As AI increasingly performs tasks independently in the background, Rime sees voice becoming a key interface between people and intelligent systems.
Rafael Valle has also joined Rime as Chief Scientist. Valle previously led audio understanding at Meta Superintelligence Labs and worked with NVIDIA’s Applied Deep Learning Research audio team.
Expanding Enterprise Voice AI
Rime plans to continue developing the underlying technology needed for more expressive, responsive and human-like AI conversations. The focus combines machine learning, speech technology and linguistics to improve how enterprise AI systems communicate with users in real time.
