Have a listen for yourself: https://www.youtube.com/watch?v=92BQg2oozBg
The pipeline can be mostly run locally on an M-series Mac with at least 16 GB of RAM.
The translation stage is the main quality constraint but local models are getting better and better. Could see a specialized 30B model be more than enough here.
Did this on a whim so I could listen to the Kimi founder being interviewed. What stood out to me is how fast the software and even the models themselves seem to be getting commoditized. Far faster than I ever anticipated.
Built using Kimi Code, thought the model needed some guidance, so not a 1-shot effort quite yet.