· 5 min
Whisper on the Neural Engine: what on-device inference actually costs
I built a dictation tool for myself in a day. The interesting part was not the app — it was discovering that the on-device model beat the larger one on speed and accuracy, and that the naming told me the opposite.
on-device AIspeech recognitionmacOSarchitecture