Whistle and Phonon 2: The Smallest Speech-to-Text Models
6 min read

Two very small speech-to-text models were announced this week: Whistle from Cactus and Phonon 2 from Fermion Research. Speech recognition models, the ones that turn audio into text, usually run to several gigabytes. Whistle is 16.9 MB and Phonon 2 is a 164 MB download. Both run on your own device without the cloud.
This article summarizes the relevant sections of the AI Search channel's weekly AI news video on YouTube and extends them with the measurements, install commands and limitations from the two projects' own pages; the video link is at the end. This is a news roundup, and I have not tested the models. Numbers come from the benchmark tables on the project pages, and the developers ran the measurements themselves.
Why should a speech model be small?
When audio is processed in the cloud, latency and privacy become problems; every recording goes to a server. An on-device model solves both: no network delay, and the audio never leaves the device. For devices with limited memory, such as phones, watches, robots, smart homes and microcontrollers, model size is decisive. So these two releases try different routes on size and speed.
What is Whistle?
Whistle is a speech recognition model Cactus Compute released on October 2, 2026. It is a single 16.9 MB file; it runs on a CPU with no extra dependencies and loads into the C++ engine used by Cactus's small language model Needle.
Highlights:
- Languages: English, German, French, Spanish, Italian, Dutch and Polish. The language is detected automatically unless you name it.
- Input: 16 kHz mono audio, up to 30 seconds in one pass.
- Word timestamps: start, end and probability for every word.
- Speech embedding: encoder output for every 80 ms frame, without producing a transcript.
- Keyword biasing: you can boost names and terms (for example "Siobhan", "Krzysztof") with a list.
- Silence: if loudness is below a threshold, it returns an empty transcript without entering the decoder.
- Deployment: prebuilt engines for seventeen targets, from macOS and Linux to Android, iOS, watchOS, Windows on ARM, RISC-V and the browser. In the demo on the page the model runs in a browser tab and the audio never leaves the device.
Whistle measurements
Cactus's benchmark table (10 seconds of audio, Apple M4 Pro CPU; each model on its official runtime):
| Whistle | Whisper base | Moonshine tiny v2 | |
|---|---|---|---|
| Size | 16.9 MB | 145.3 MB | 41.9 MB |
| Time to first token | 11.1 ms | 73.2 ms | 22.8 ms |
| Decode speed | 1,319 tokens/s | 266 tokens/s | 262 tokens/s |
On word error rate (WER, lower is better) Whistle is ahead on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average. Whisper base is ahead on TED-LIUM, AMI and the MLS average. So Whistle does not win every benchmark, but it is remarkable at 8.6 times smaller. The video's "much faster than Whisper base" fits the measurements: about 5x on decoding and about 6.6x on time to first token.
Installing Whistle
You can try it with a single command in Python:
pip install cactus-needleimport needle
print(needle.transcribe("clip.wav")["text"])word_timestamps=True gives word times, keywords=[...] boosts names and language="de" forces the language. Weights are on Hugging Face and the engine is on GitHub (Cactus-Compute/needle3).
What is Phonon 2?
Phonon 2 is Fermion Research's open-weight, English-only speech recognition model. The company says it is the most accurate open speech recognition model under 900 MB. The download is 164 MB. It is derived from NVIDIA's Parakeet TDT 0.6B v3, and the weights are released under the same CC-BY-4.0 license. Each encoder weight is stored as one of five learned levels at about 2.1 bits.
Phonon 2 measurements
Average word error rate across the seven English sets of the Open ASR Leaderboard:
| Model | Download | Average WER |
|---|---|---|
| Parakeet TDT 0.6B v3 (teacher) | 2,508 MB | 4.96% |
| Phonon 2 | 164 MB | 5.21% |
| Parakeet Redux | 178 MB | 5.69% |
| Canary 180M Flash | 737 MB | 5.69% |
| Whisper large-v3-turbo | 1,618 MB | 6.58% |
So a file 15 times smaller comes very close to the teacher's accuracy. Phonon 2 beats the teacher on meeting recordings (AMI) and parliamentary speech (VoxPopuli).
Speed: on an Apple M5 MacBook Air on the GPU (MLX) it runs at 174 times real time; an hour of audio becomes text in about 20 seconds. According to the page, that is 1.7 times faster than the fastest other runtime measured on the same machine. On eight CPU cores an hour of audio takes 25 seconds; on a single H100, a full day of audio takes 13 seconds in batches of 128.
Installing Phonon 2
pip install fermion-research
phonon transcribe meeting.wav
phonon serve --port 8010phonon serve opens an OpenAI-compatible endpoint and phonon listen transcribes the microphone live on a Mac. There are Docker containers (CPU and CUDA) and FermionResearch/Phonon-2 on Hugging Face. The same lab also offers Detta, a Mac dictation app that runs Phonon 2 in any text field.
How to choose between them
| Whistle | Phonon 2 | |
|---|---|---|
| Size | 16.9 MB | 164 MB |
| Languages | 7 languages, auto-detected | English only |
| Input | Up to 30 s | Long recordings (hours) |
| Strength | Very constrained devices such as microcontrollers, watches, browsers | Long English recordings such as meetings, parliament, podcasts |
| Output | Text, word times, embeddings | Text, word times (--json) |
The two models do not do the same job: Whistle is multilingual and aimed at short clips on very small devices; Phonon 2 stresses accuracy and speed for English and long recordings. Do not decide before trying your own voice and accent; the benchmark tables partly rest on clean recordings. The video occasionally calls Phonon 2 "text-to-speech," but the task described is speech to text.
For an example of a voice-driven app, see What is Talkamble; for the same week's voice model news, see the ElevenLabs v4 article.
Frequently Asked Questions
Does Whistle need a GPU?
No. According to Cactus it runs on a CPU only, with no extra dependencies.
Which languages does Whistle support?
English, German, French, Spanish, Italian, Dutch and Polish; the language is detected automatically.
Is Whistle more accurate than Whisper base?
On some benchmarks (LibriSpeech, SPGISpeech, Earnings-22, FLEURS average), yes; on TED-LIUM, AMI and the MLS average Whisper base is better. Whistle is 8.6 times smaller and faster.
Does Phonon 2 know Turkish?
No, it was released for English only.
How fast is Phonon 2?
On an Apple M5 MacBook Air an hour of audio becomes text in about 20 seconds; on eight CPU cores, 25 seconds.
Source
- AI Search, Gemini 4, GPT 6.1, Dots, Claude Sonnet 5.5, Ideogram 4.5, Flux 3: AI NEWS (YouTube, October 4, 2026): the Whistle and Phonon 2 sections.
- Cactus Compute, Whistle: Speech to Text in 16.9 MB (October 2, 2026): features, measurements, install.
- Fermion Research, Introducing Phonon-2: accuracy, speed and install.
Related Posts
What Is OpenAI Dots? Open-Source Open Dots Alternatives
OpenAI Dots is a persistent agent with its own computer that keeps working when you are away. Who gets it, its safety controls, open-source Open Dots options.
What Is Claude Sonnet 5.5? Benchmarks, Cost, Free Plan
Claude Sonnet 5.5 is now the free plan's default model. It scores 70.6% on Terminal-Bench 4.0 at $2 / $10 per million tokens. How it compares.
What Is GPT-6.1 Sol? Price, Performance, Who Can Use It
GPT-6.1 Sol offers near-Astra intelligence at a fifth of the price: $2 / $10. New plan tiers, Ultrafast speed and who can use it, explained.