South Korea · Open Source · Hong Kong · Nvidia · China · Llama · Fortune Technology
Several firms are working to build models for what’s deemed “low resource languages,” or those that don’t have a large
Compiled by KHAO Editorial — aggregated from 1 source + 2 references discovered via search. See llms.txt for citation guidance.
◌ Single Source
Indosat, Indonesia’s second-largest telecoms company, is building Sahabat AI, an open-source large language model that focuses on Indonesian languages like Bahasa.
Key facts
- South Korea has gone further, staging a state-sponsored elimination tournament, which local media has dubbed the “AI Squid Game”, to pick national champions for homegrown foundation models, backed
- Ting estimates that the company uses between 500 million and 1 billion tokens to train its models, compared to the trillions used for English-language models, and at a cost of roughly $250,000
- Together, these efforts grew the corpus of Cantonese data from 100 million tokens to more than 500 million
- Pak-Sun Ting will be speaking at the Fortune Leaders Forum, held in Macau on Sep. 8
Summary
The AI boom is being driven by two languages: English and Mandarin Chinese. That leaves countless other languages and dialects, even those spoken by tens of millions of people, by the wayside. “The whole AI revolution is in English and Mandarin,” says Pak-Sun Ting, CEO of Votee AI, a Hong Kong-based startup that’s striving to build AI models for Cantonese and other neglected tongues. Votee is one of a growing number of companies that are trying to tackle AI’s neglect of other languages. “It’s more than culture.