Tanay Kothari’s path to building one of the fastest-rising voice AI companies did not begin with a transcription app. It began in Delhi with a fascination for conversational computing, followed by years of building software, studying artificial intelligence at Stanford and eventually trying to rethink one of computing’s oldest interfaces: the keyboard.

That idea has now become Wispr, the San Francisco-based company behind AI dictation product Wispr Flow. In August 2026, Wispr raised $280 million in a Series B led by Menlo Ventures at a $2 billion valuation, almost three times the roughly $700 million valuation reported after its November 2025 financing. The latest round brought the company’s total capital raised to about $361 million.

For a business still best known for converting speech into text, that valuation reflects a much larger bet. Wispr and its investors believe voice could become a primary way people write, communicate with AI and ultimately tell computers what to do.

From Delhi to Stanford

Kothari grew up in Delhi and later moved to the United States to study at Stanford. Wispr’s official biography says he earned a bachelor’s degree in computer science and a master’s degree in artificial intelligence, taught Stanford’s deep-learning course alongside Andrew Ng and published research through the Stanford AI Lab. Before Wispr, he built several products, including AI-personalization startup FeatherX, which was acquired by Cerebra Technologies.

His interest in voice interfaces goes much further back. Kothari has said watching Iron Man as a child left him fascinated not primarily by Tony Stark, but by JARVIS, the fictional computer that could be spoken to almost like another person.

At Stanford, he met Sahaj Garg during their first year as undergraduates. Kothari has described the pair as close friends who later roomed together and spent late nights building projects across AI and other emerging technologies.

Garg brought a strong research background of his own. Wispr says he worked at Stanford’s AI Lab, published generative-model research and later became an early employee and AI team lead at Luminous Computing, a startup working on photonic hardware for artificial intelligence. Today he is Wispr’s co-founder and chief technology officer.

Wispr did not start as a dictation company

The company Kothari and Garg started in 2021 looked very different from Wispr Flow.

Early Wispr was a neurotechnology project. The founders experimented with wearable hardware designed to translate subtle physiological signals into computer input, pursuing the broader goal of communicating with machines without relying on conventional keyboards and screens.

The concept was ambitious, but hardware brought difficult engineering and scaling constraints. As speech recognition and large language models improved, the founders moved toward software.

That pivot eventually became Wispr Flow.

What Wispr Flow actually does

Flow is designed to work anywhere a user would normally type. On supported Macs, Windows PCs, iPhones and Android devices, a user activates Flow, speaks naturally and receives formatted text inside the application they are using.

But the product is not intended to behave like old-fashioned dictation software that simply converts every spoken word literally.

Flow removes filler words, adds punctuation, formats lists, understands verbal corrections and uses context to improve names, jargon and technical vocabulary. If someone says, “Let’s meet at five, actually six,” Flow is designed to write the corrected time instead of transcribing the entire false start.

Wispr says Flow supports more than 100 languages.

For developers, the software can also use nearby code context. Wispr says Flow can recognize variables, function names and filenames from supported editors such as VS Code, Cursor and Windsurf, improving transcription of technical prompts and programming terminology.

That is an important distinction. Wispr is not merely competing on raw speech recognition. It is trying to produce text that requires little or no correction after the user stops speaking.

Growth became difficult for investors to ignore

By August 2026, Wispr had moved well beyond a niche dictation experiment.

Menlo Ventures said Wispr’s revenue had grown more than 30-fold year over year, with Flow being used in 162 countries and more than 100 languages. Reuters reported that the product was being used across more than 10,000 enterprises. Wispr says users have generated more than 60 billion words through its software.

The company has also published unusually strong engagement claims. Its media materials say that after six months, an average user generates about 72% of their typed characters through Flow, across nearly 70 apps and websites.

Earlier reporting showed how quickly those habits were forming. In November 2025, Wispr said an average user who had been using the product for three months was already creating more than half their characters through Flow. At the time, the startup was valued around $700 million after a $25 million round led by Notable Capital.

Less than nine months later, that valuation became $2 billion.

The $280 million Series B

Menlo Ventures led Wispr’s $280 million Series B announced on August 17, 2026.

Existing investors included Notable Capital, NEA, Neo Ventures, 8VC and MVP Ventures. New participants included Acrew, Forerunner, Goodwater, Peak XV, Together Fund and PLUS Capital, alongside other investors.

The financing lifted total funding to approximately $361 million.

The size of the round matters because Wispr is increasingly competing not just with other startups, but with companies that control operating systems, AI models and workplace software.

Voice and dictation competitors include startups such as Willow, Monologue, Aqua and Superwhisper, while much larger companies including Apple, Google, Microsoft and major AI labs are also improving speech and conversational interfaces.

Wispr’s strategy is to remain useful regardless of which underlying application wins. Flow can be used to dictate an email, write in a document, prompt an AI assistant or speak into a coding environment.

Canto gives Wispr control of the speech layer

Alongside the funding, Wispr previewed Canto, its first proprietary speech-recognition model.

The reason for building it was straightforward: laboratory speech-recognition benchmarks often use relatively clean recordings, while real users dictate in cars, streets, offices, noisy rooms and with accents the system may encounter less frequently.

Wispr says Canto was trained for those more difficult conditions. According to Kothari, in the company’s hardest tests involving background noise, wind, music or strong accents, word-error rates fell from more than 30% to roughly 5% to 10%. Wispr also expects Canto to reduce the number of dictations requiring editing by roughly 30% to 35%.

Those are company-reported results, not independent benchmark findings, so they should not be treated as universally verified performance numbers.

Still, Canto is strategically important. Owning more of the speech-recognition stack gives Wispr greater control over accuracy, latency and product differentiation instead of relying entirely on third-party models.

Dictation is supposed to be only phase one

Wispr’s ambitions extend considerably beyond speech-to-text.

Garg has described a three-stage strategy. The first is reliable voice input. The second is voice to action, where speaking causes software to perform tasks rather than merely insert words. The third is making voice broadly available through future wearables and hardware.

The company has also established the Wispr Advanced Interfaces Lab, which is researching systems that combine voice with vision, memory, application context and generative user interfaces. Its aim is an interface layer capable of understanding what a user wants and routing that intent to the appropriate model, service or tool.

Wispr has already moved slightly beyond dictation with a meeting notetaker and experimental voice commands. The eventual vision is closer to a software interface that can act on spoken intentions.

The $2 billion question

The opportunity is large, but so are the risks.

Speech recognition is becoming cheaper. AI models are becoming better at understanding messy language. Apple and Google own device platforms. Microsoft owns major workplace software. OpenAI and Anthropic are developing increasingly capable agents.

Wispr must also maintain extremely low latency and high accuracy while processing enormous amounts of voice data. Privacy is particularly important because dictated material can contain emails, confidential company information and personal messages.

The biggest challenge may ultimately be behavioral.

People have spent decades typing. A voice interface becomes valuable only when it is reliable enough that speaking feels easier than reaching for the keyboard.

That is why Wispr’s $2 billion valuation is not simply a bet on speech recognition.

It is a bet that Kothari and Garg can change a computing habit.

They began with futuristic hardware intended to rethink human-computer interaction, retreated to a narrower problem that users could adopt immediately, and turned that product into a rapidly growing business.

Now, with $361 million raised, a $2 billion valuation and its own speech model, Wispr has to prove that Flow is not the destination.

It has to prove that dictation is the doorway to a much larger interface for AI.

Reader questions

Frequently asked questions

Who is Tanay Kothari?

Tanay Kothari is the Indian-origin co-founder and CEO of Wispr, the company behind AI voice-dictation product Wispr Flow. He studied computer science and artificial intelligence at Stanford.

What is Wispr Flow?

Wispr Flow is an AI-powered dictation product that converts natural speech into formatted text across applications, handling punctuation, filler words, verbal corrections and contextual terminology.

How much is Wispr worth?

Wispr was valued at $2 billion in its August 2026 Series B, up from roughly $700 million in November 2025.

How much funding has Wispr raised?

Wispr has raised about $361 million in total funding after its $280 million Series B.

Who invested in Wispr?

Investors include Menlo Ventures, Notable Capital, NEA, 8VC, Peak XV, Acrew, Forerunner, Goodwater, Together Fund, PLUS Capital and others.

What is Canto?

Canto is Wispr’s proprietary speech-recognition model, designed to improve transcription accuracy in difficult real-world conditions such as noise, wind, music and varied accents.

What does Wispr want to build beyond dictation?

Wispr plans to move from voice-to-text toward voice-driven actions, multimodal interfaces combining voice with context and vision, and potentially future voice-first hardware.


Corrections and updates

Nexuswild welcomes factual corrections. Email [email protected] with evidence and the article URL.