TL;DR
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
Cactus Compute released Whistle, a 16.9 MB speech-recognition model designed to run locally on a CPU without dependencies. The company reports support for seven languages and an 11.1 ms time to first token on an Apple M4 Pro, while its word-error results vary against Whisper and Moonshine across benchmarks.
Cactus Compute released Whistle on October 2, describing it as a 16.9 MB speech-recognition model that runs on a CPU without dependencies. The company says it can transcribe up to 30 seconds of audio in English, German, French, Spanish, Italian, Dutch and Polish, with processing kept on the device—an approach aimed at speech features in phones, wearables, vehicles, robots and other systems with limited resources.
Whistle is built to load into the same C++ engine and container as Cactus Compute’s Needle model. The company says both models use the same quantisation, allowing developers to run speech recognition alongside Needle in a single binary. Its browser demonstration downloads the model on first use; Cactus says subsequent audio is processed locally and does not leave the device.
The model offers three outputs: a transcript, word-level start and end timestamps with probabilities, and speech embeddings that represent the audio without generating a transcript. It accepts 16 kHz mono audio in clips up to 30 seconds. Language detection is automatic unless a language is specified. The decoder supports keyword biasing, which the company says can raise the likelihood of terms supplied by an application.
Cactus reports an 11.1 ms time to first token for a 10-second clip on an Apple M4 Pro CPU. In the company’s comparison, Whistle took 5.9 ms for five seconds of audio and 36.3 ms for 30 seconds. Those measurements cover audio input through the first generated token; the company says Whisper’s processing time remained flat across clip lengths because it pads input to 30 seconds.
Local Speech on Smaller Devices
A model with a reported size of 16.9 MB could make speech input easier to add to devices where storage, connectivity or cloud processing is a constraint. Local processing can also keep audio on the device, which may be useful for products handling private conversations. Those are potential advantages of the design, not proof of performance across every phone, microcontroller or embedded system.
Cactus’s figures suggest Whistle is designed for low-latency applications: the company measured a short time to first token and reports a high decode rate on its test CPU. But its benchmark table does not show a universal accuracy lead. Whistle outperformed Whisper base on several listed datasets, while Whisper base scored better on TED-LIUM and the AMI and MLS comparisons. Moonshine tiny v2 is listed as English-only, limiting direct language coverage comparisons.
The release also connects speech recognition to Cactus’s existing model stack. Sharing an engine with Needle may reduce integration work for developers already using that system, though the announcement does not provide independent testing of the combined deployment or quantify memory use beyond the model file’s size.
As an affiliate, we earn on qualifying purchases.
How Whistle Processes Audio
Cactus describes a pipeline that converts 16 kHz audio into log-mel features, then reduces those frames through a convolutional stem before encoding them. For a 30-second clip, the company says the audio becomes 3,000 initial frames and then 375 frames, with each later frame representing about 80 milliseconds. An eight-block encoder processes the audio before a decoder generates text.
The decoder uses five-beam search, according to the report, and its transcript is capped at 320 tokens. Cactus says the encoder runs all eight blocks at every model depth, while a setting can choose the number of decoder layers. A silence check measures the clip before decoding; audio below the engine’s threshold returns an empty transcript without running beam search.
The company compares Whistle with Whisper base, listed at 145.3 MB, and Moonshine tiny v2, listed at 41.9 MB. Its report says the three systems ran in their official runtimes with default settings, including five beams for Whistle. The reported word-error-rate comparisons use different datasets, and the company notes that some results are unavailable because the other model developers did not publish them. It also flags that Whisper’s AMI result uses a different subset from the AMI figures for the other systems.
““It is one 16.9 MB file, runs on the CPU with no dependencies, and loads into the same C++ engine as Needle.””
— Cactus Compute
low latency speech recognition device
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Benchmark Limits and Device Performance
The performance and accuracy figures come from Cactus Compute’s own report; the supplied material does not include independent replication or detailed benchmark files. The exact hardware, runtime settings and measurement definitions matter when comparing speech models, and results from an Apple M4 Pro CPU do not establish speeds on phones, microcontrollers or other processors.
Word-error rates differ by dataset, and missing results prevent a complete comparison. The report also identifies a mismatch between the AMI subset used for Whisper and those used for the other systems. It does not provide a single combined accuracy score or establish how well the model performs in noisy conditions, with accents, or on speech outside its seven listed languages.
It is also not clear from the announcement how much total memory Whistle needs while running, what its licensing terms are, or which specific devices have been tested. The stated 16.9 MB refers to the model file, not necessarily the full memory footprint of an application using it.
privacy-focused speech transcription device
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Developer Access and Independent Tests
Cactus Compute presents Whistle as available through a browser sandbox and says it is intended for mobile, wearable, robotics, smart-home, automotive and microcontroller use. Developers will be able to assess its practical fit by testing the model on their target hardware and checking language accuracy, runtime memory and integration requirements.
No further release dates or independent benchmark results are specified in the announcement. The next useful evidence would be reproducible tests across a wider range of devices, with consistent datasets and clear reporting of accuracy, latency and total memory use.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Whistle?
Whistle is a speech-recognition model released by Cactus Compute. The company says it transcribes audio, returns word-level timestamps and can produce speech embeddings.
Which languages does it support?
Cactus lists English, German, French, Spanish, Italian, Dutch and Polish. The language is detected automatically unless the user specifies it.
Does Whistle send recordings to the cloud?
Cactus says processing in its demonstration happens on the device and that audio does not leave it. That claim describes the company’s stated design; the supplied announcement does not include an independent privacy audit.
How does Whistle compare with Whisper?
In Cactus’s tests, Whistle was smaller than Whisper base and scored better on several listed word-error benchmarks. Whisper base led on TED-LIUM, AMI and the MLS average, and the AMI results use different subsets, so the comparison is not a consistent win on every dataset.
What does the 16.9 MB figure measure?
It is the stated size of the Whistle model file. The announcement does not specify the full runtime memory footprint of an application using it.
Source: hn
Columbus Day / Indigenous Peoples' Day Picks
long weekend sales
As an affiliate, we earn on qualifying purchases.
