Please add japanese!

#3
by CYISNOTHERE - opened

Pretty please

+1

I was so excited to try this model, only to find it dramatically cuts down on Whisper languages down to just 8. Japanese in particular!

nyra labs org
β€’
edited Aug 5

Generally transcribing japanese should still work too... just pass the original whisper japanese language flag..... it was not further finetuned on japanese so maybe quality slightly degraded but should still be usable....let me know how that goes... otherwise i will have to see about what kind of datasets i can get my hands on and how one could sensibly define verbatim in japanese in written form....

+1 for fine tuning Japanese. I've used this model which is pretty good "efwkjn/whisper-ja-1.5B". But I understand that CrisperWhisper2.0 is a lot faster at transcription?

The latest commit has some fixes for CJK text that was displayed wrong in some cases. I don't speak and read Japanese, but if some Japanese speakers could test CrisperWhisper 2.0 Large on some Japanese text with and without disfluencies I would love to know results ;)

Sign up or log in to comment