Artificial Intelligence
Speech Recognition
Turning spoken audio into text, now usually with a single end-to-end neural network.
Classical speech systems chained together an acoustic model, a pronunciation dictionary and a language model. Modern systems replace the chain with one network that maps audio features directly to characters or tokens.
The remaining difficulty is not the clean case but the messy one: accents, background noise, overlapping speakers and domain vocabulary. Evaluation uses word error rate, which punishes insertions, deletions and substitutions equally even though they rarely matter equally.
Also in Artificial Intelligence
JOIN NOW
Begin the first module
It is free, it is the real curriculum, and if it is not for you, you have lost nothing but an evening.
Join any time · Build AI skills at your pace