HyperAIHyperAI

Command Palette

Search for a command to run...

Cactus-Compute Releases Whistle, a 16.9 MB On-Device Speech Recognition Model

Cactus-Compute has released Whistle, an open-weight speech recognition model optimized for edge devices and embedded systems. Weighing just 16.9 megabytes, the model runs entirely on CPU architecture through a shared C++ inference engine designed for Needle, eliminating external dependencies and enabling deployment across smartphones, wearables, automotive systems, and microcontrollers. The entire processing pipeline operates locally, ensuring audio never leaves the user device. Whistle supports real-time transcription for seven languages: English, German, French, Spanish, Italian, Dutch, and Polish. It accepts up to 30 seconds of 16 kHz mono audio per pass, automatically detecting language or allowing manual override. Beyond standard transcription, the engine outputs precise word-level timestamps with probability scores and extracts speech embeddings at 80-millisecond intervals. A built-in silence threshold prevents unnecessary processing during quiet intervals, while an Aho-Corasick automaton allows users to bias recognition toward specific keywords during decoding. Architecturally, Whistle utilizes a convolutional front end that processes audio into 80 log-mel bins, feeding into an encoder composed of eight Simple Attention blocks with Monarch Hadamard multilayer perceptrons. The decoder employs laddered attention layers with gated cross-attention mechanisms, enabling efficient beam search across five parallel hypotheses. The model supports dynamic depth selection, allowing users to load lighter or heavier configurations by adjusting audio depth parameters at runtime without retraining. Independent benchmarking demonstrates Whistles competitive edge in efficiency and accuracy. On an Apple M4 Pro CPU, the model generates its first transcription token in approximately 11 milliseconds for a ten-second clip, with decoding latency scaling linearly to 36 milliseconds for full thirty-second inputs. In word error rate testing across LibriSpeech, SPGISpeech, Earnings-22, and FLEURS datasets, Whistle outperformed OpenAI Whisper base model. The performance gap is particularly notable given Whistles significantly smaller footprint, operating at less than one-ninth the size of Whisper 145.3 megabyte checkpoint while maintaining multilingual capabilities. Deployment is streamlined through a unified C API and Python bindings, compatible with 17 operating environments including macOS, Linux, Android, iOS, Windows ARM, RISC-V, MIPS, and browser-based WASI targets. The needle_load runtime seamlessly handles both text and speech models from a single binary, returning structured JSON outputs with transcription, language identification, and processing metadata. Weights and source code are publicly available on Hugging Face and GitHub, underscoring a growing industry shift toward lightweight, privacy-first AI inference capable of running reliably without cloud connectivity.

Related Links