Streaming models for speech-to-text, with VAD. Plus models include speaker segmentation and verification for cascaded diarization.
-
futo-org/asr4all-l
Automatic Speech Recognition • 97M • Updated • 169 • 1 -
futo-org/asr4all-m
Automatic Speech Recognition • 55.8M • Updated • 147 • 1 -
futo-org/asr4all-s
Automatic Speech Recognition • 28.2M • Updated • 378 • 1 -
futo-org/asr4all-l-plus
Automatic Speech Recognition • 0.1B • Updated • 25