Developing for Indic languages | Gemma and Navarasa
May 14, 2024
46,140
637
52
1.49%
Search the Record
IndexedEvery word spoken in this episode is indexed. Type any phrase to jump straight to the moment it was said.
Type any word or phrase that may have been spoken. Click a result to seek the player to that exact moment.
Try a name, a topic, or a quoted line
Google Episodes Around May 14, 2024
See what was published immediately before and after this episode.
1:47Now PlayingDeveloping for Indic languages | Gemma and Navarasa
YouTube Description
as posted by the channelWhile many early large language models were predominantly trained on English language data, the field is rapidly evolving. Newer models are increasingly being trained on multilingual datasets, and there's a growing focus on developing models specifically for the world’s languages. However, challenges remain in ensuring equitable representation and performance across diverse languages, particularly those with less available data and computational resources.
Gemma, Google's family of open models, is designed to address these challenges by enabling the development of projects in non-Germanic languages. Its tokenizer and large token vocabulary make it particularly well-suited for handling diverse languages. Watch how developers in India used Gemma to create Navarasa — a fine-tuned Gemma model for Indic languages.
Watch the full keynote
To watch this keynote with American Sign Language (ASL) interpretation, please click here
#GoogleIO #GoogleIO2024
Subscribe to our Channel
Find us on X
Watch us on TikTok
Follow us on Instagram
Join us on Facebook
Guests & Subjects Covered
Sentinel Indexing in Progress
Metadata and chapters are available. Claim extraction for this episode is pending.
All video content is delivered via YouTube embedded players in accordance with the YouTube Terms of Service. Sentinel provides research tools that promote discovery and accountability across political media.









