İçeriğe geç / Skip to content / Zum Inhalt

Whistle and Phonon 2: The Smallest Speech-to-Text Models

Ahmet Balaman

6 min read

Vibe CodingWhistlePhonon 2Speech to TextWhisperOpen SourceAI
Whistle and Phonon 2: The Smallest Speech-to-Text Models

Two very small speech-to-text models were announced this week: Whistle from Cactus and Phonon 2 from Fermion Research. Speech recognition models, the ones that turn audio into text, usually run to several gigabytes. Whistle is 16.9 MB and Phonon 2 is a 164 MB download. Both run on your own device without the cloud.

This article summarizes the relevant sections of the AI Search channel's weekly AI news video on YouTube and extends them with the measurements, install commands and limitations from the two projects' own pages; the video link is at the end. This is a news roundup, and I have not tested the models. Numbers come from the benchmark tables on the project pages, and the developers ran the measurements themselves.

Why should a speech model be small?

When audio is processed in the cloud, latency and privacy become problems; every recording goes to a server. An on-device model solves both: no network delay, and the audio never leaves the device. For devices with limited memory, such as phones, watches, robots, smart homes and microcontrollers, model size is decisive. So these two releases try different routes on size and speed.

What is Whistle?

Whistle is a speech recognition model Cactus Compute released on October 2, 2026. It is a single 16.9 MB file; it runs on a CPU with no extra dependencies and loads into the C++ engine used by Cactus's small language model Needle.

Highlights:

  • Languages: English, German, French, Spanish, Italian, Dutch and Polish. The language is detected automatically unless you name it.
  • Input: 16 kHz mono audio, up to 30 seconds in one pass.
  • Word timestamps: start, end and probability for every word.
  • Speech embedding: encoder output for every 80 ms frame, without producing a transcript.
  • Keyword biasing: you can boost names and terms (for example "Siobhan", "Krzysztof") with a list.
  • Silence: if loudness is below a threshold, it returns an empty transcript without entering the decoder.
  • Deployment: prebuilt engines for seventeen targets, from macOS and Linux to Android, iOS, watchOS, Windows on ARM, RISC-V and the browser. In the demo on the page the model runs in a browser tab and the audio never leaves the device.

Whistle measurements

Cactus's benchmark table (10 seconds of audio, Apple M4 Pro CPU; each model on its official runtime):

Whistle Whisper base Moonshine tiny v2
Size 16.9 MB 145.3 MB 41.9 MB
Time to first token 11.1 ms 73.2 ms 22.8 ms
Decode speed 1,319 tokens/s 266 tokens/s 262 tokens/s

On word error rate (WER, lower is better) Whistle is ahead on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average. Whisper base is ahead on TED-LIUM, AMI and the MLS average. So Whistle does not win every benchmark, but it is remarkable at 8.6 times smaller. The video's "much faster than Whisper base" fits the measurements: about 5x on decoding and about 6.6x on time to first token.

Installing Whistle

You can try it with a single command in Python:

pip install cactus-needle
import needle
print(needle.transcribe("clip.wav")["text"])

word_timestamps=True gives word times, keywords=[...] boosts names and language="de" forces the language. Weights are on Hugging Face and the engine is on GitHub (Cactus-Compute/needle3).

What is Phonon 2?

Phonon 2 is Fermion Research's open-weight, English-only speech recognition model. The company says it is the most accurate open speech recognition model under 900 MB. The download is 164 MB. It is derived from NVIDIA's Parakeet TDT 0.6B v3, and the weights are released under the same CC-BY-4.0 license. Each encoder weight is stored as one of five learned levels at about 2.1 bits.

Phonon 2 measurements

Average word error rate across the seven English sets of the Open ASR Leaderboard:

Model Download Average WER
Parakeet TDT 0.6B v3 (teacher) 2,508 MB 4.96%
Phonon 2 164 MB 5.21%
Parakeet Redux 178 MB 5.69%
Canary 180M Flash 737 MB 5.69%
Whisper large-v3-turbo 1,618 MB 6.58%

So a file 15 times smaller comes very close to the teacher's accuracy. Phonon 2 beats the teacher on meeting recordings (AMI) and parliamentary speech (VoxPopuli).

Speed: on an Apple M5 MacBook Air on the GPU (MLX) it runs at 174 times real time; an hour of audio becomes text in about 20 seconds. According to the page, that is 1.7 times faster than the fastest other runtime measured on the same machine. On eight CPU cores an hour of audio takes 25 seconds; on a single H100, a full day of audio takes 13 seconds in batches of 128.

Installing Phonon 2

pip install fermion-research
phonon transcribe meeting.wav
phonon serve --port 8010

phonon serve opens an OpenAI-compatible endpoint and phonon listen transcribes the microphone live on a Mac. There are Docker containers (CPU and CUDA) and FermionResearch/Phonon-2 on Hugging Face. The same lab also offers Detta, a Mac dictation app that runs Phonon 2 in any text field.

How to choose between them

Whistle Phonon 2
Size 16.9 MB 164 MB
Languages 7 languages, auto-detected English only
Input Up to 30 s Long recordings (hours)
Strength Very constrained devices such as microcontrollers, watches, browsers Long English recordings such as meetings, parliament, podcasts
Output Text, word times, embeddings Text, word times (--json)

The two models do not do the same job: Whistle is multilingual and aimed at short clips on very small devices; Phonon 2 stresses accuracy and speed for English and long recordings. Do not decide before trying your own voice and accent; the benchmark tables partly rest on clean recordings. The video occasionally calls Phonon 2 "text-to-speech," but the task described is speech to text.

For an example of a voice-driven app, see What is Talkamble; for the same week's voice model news, see the ElevenLabs v4 article.

Frequently Asked Questions

Does Whistle need a GPU?

No. According to Cactus it runs on a CPU only, with no extra dependencies.

Which languages does Whistle support?

English, German, French, Spanish, Italian, Dutch and Polish; the language is detected automatically.

Is Whistle more accurate than Whisper base?

On some benchmarks (LibriSpeech, SPGISpeech, Earnings-22, FLEURS average), yes; on TED-LIUM, AMI and the MLS average Whisper base is better. Whistle is 8.6 times smaller and faster.

Does Phonon 2 know Turkish?

No, it was released for English only.

How fast is Phonon 2?

On an Apple M5 MacBook Air an hour of audio becomes text in about 20 seconds; on eight CPU cores, 25 seconds.

Source

Comments