Whistle: Speech to Text in 16.9 MB

Whistle packages speech-to-text into a 16.9 MB on-device model that runs on CPU with no dependencies, delivers first tokens in about 11 ms, and supports seven languages while keeping audio local. For CIOs, the business value is lower cloud inference cost, better privacy and compliance, and much faster voice experiences for edge and embedded products such as mobile devices, wearables, robots, smart home systems, automotive platforms, and microcontrollers. Strategically, this points IT organizations toward more offline-first, edge-native voice workflows and tighter integration between speech, transcription, and downstream automation in a single deployment path.

Hacker News3 min read
Read full article
Whistle: Speech to Text in 16.9 MB

Read the full story at Hacker News →