Meta, formerly known as Facebook, has developed a unique AI language model called Massively Multilingual Speech (MMS) that sets itself apart from ChatGPT clones. MMS has the ability to recognize over 4,000 spoken languages and produce speech in more than 1,100 languages. In an effort to promote language diversity and encourage further research, Meta has decided to open-source MMS, making the models and code available to the research community. By sharing their work, Meta hopes to contribute to the preservation of the world’s incredible language diversity.
Read: Facebook fined $1.3 billion over data transfers
Creating speech recognition and text-to-speech models usually involves training them on large amounts of audio data accompanied by transcription labels. However, for languages that are not widely used in industrialized nations and are at risk of disappearing, such data is scarce. Meta tackled this challenge by taking an unconventional approach. They utilized audio recordings of translated religious texts, such as the Bible, which have been widely studied for text-based language translation research. These recordings allowed Meta to expand the available languages of their model to over 4,000.
At first glance, this approach may raise concerns about bias towards Christian worldviews. However, Meta assures that the model remains unbiased. The use of a connectionist temporal classification (CTC) approach, along with the analysis of religious content, prevents the model from producing more religious language. Furthermore, despite the majority of the recordings being read by male speakers, the model performs equally well with female and male voices.
Meta trained their models using wav2vec 2.0, a self-supervised speech representation learning model that can leverage unlabelled data. By combining unconventional data sources with a self-supervised speech model, Meta achieved impressive results. Their Massively Multilingual Speech models outperformed existing models like OpenAI’s Whisper, exhibiting half the word error rate while covering 11 times more languages.
Meta acknowledges that their models are not perfect and there is a risk of mistranscribing certain words or phrases, which could lead to offensive or inaccurate language. They emphasize the importance of collaboration across the AI community for responsible development of AI technologies.
With the release of MMS for open-source research, Meta aims to counter the trend of technology contributing to the dwindling of languages, often leaving only the top 100 or fewer languages supported by big tech companies. They envision a world where assistive technology, text-to-speech systems, and even virtual reality and augmented reality technologies enable people to speak and learn in their native languages, thus encouraging the preservation of linguistic diversity. Meta believes that technology can have a positive impact by empowering individuals to access information and use technology in their preferred language.



