Featured Story

Adaption and AI Singapore Pilot to Improve Model Training Efficiency in Southeast Asia

·

·

4min read

Adaption and AI Singapore Pilot to Improve Model Training Efficiency in Southeast Asia

Why frontier AI models underperform in Southeast Asian languages

Southeast Asia is home to more than 700 million people, over 1,200 languages, and 11 countries. Frontier language models, trained overwhelmingly on Western, English-centric data, perform inconsistently on Southeast Asian languages and often miss the region's cultural context.

Adaption collaborated with AI Singapore (AISG) to enhance dataset quality and expand training dataset size via localization across five low-resource Southeast Asian languages. Adaption's AutoScientist was able to save 56 engineering hours’ worth of manual optimization of model training hyperparameters for AISG.

AISG is a national initiative driven by the Ministry of Digital Development and Information (MDDI), the Infocomm Media Development Authority (IMDA), and the National Research Foundation, Singapore (NRF). AISG brings together Singapore-based research institutions and the broader ecosystem of AI start-ups and companies to support use-inspired research, grow knowledge, create tools, and develop the talent to power Singapore’s AI efforts.

Among AISG's flagship products is SEA-LION (Southeast Asian Languages In One Network), Southeast Asia's first family of open-source large language models, built to understand the region's languages, cultures, and context. SEA-LION is trained to cover 11 national Southeast Asian languages and a number of major regional dialects. Since its first release in December 2023, SEA-LION has seen over 900,000 model downloads and more than 6.5 million API calls, across an ecosystem spanning over 70 community partners.

Frontier language models, trained overwhelmingly on Western, English-centric data, perform inconsistently on Southeast Asian languages and often miss the region's cultural context.

Targeting this gap means AISG’s team has to address two connected problems at once: the quality and coverage of the training data, and the speed of the cycle that turns that data into an evaluated model.

The data problem. Southeast Asian training data is scarce relative to demand, quality is inconsistent across sources, and much of what exists misses the local dialects, nuance, and cultural context that separate a merely multilingual model from a genuinely regional one.

The training problem. Techniques proven in English text don't automatically transfer to Southeast Asian languages, and the field moves fast enough that confirming what works, and what needs adjusting, is a constant race against the clock.

How Adaptive Data expanded SEA-LION’s training set to 1.75 million samples

Adaption partnered with AISG for a pilot using Adaptive Data, Adaption's data enhancement and localization engine, to overcome both gaps at once: raise the quality of existing training data, and extend coverage into low-resource SEA languages, on a timeline that a fully manual approach couldn't match.

Based on a sample of 1,050,000 rows from AISG's SEA-Instruct-2602, spanning domains including writing, math, science, and education, across seven Southeast Asian languages, Adaption’s data engine delivered the following within one month:

  • Enhanced the 1 million data samples, applying Adaptive Data to improve the overall quality score through reasoning traces in the language of the sample.
  • Generated approximately 750,000 new samples by further localizing the existing 1 million samples, adding additional native-quality coverage in five additional languages and regional variants: Singapore Tamil, Singapore Malay, Burmese, Khmer, and Lao.

What data enhancement changes in a Thai training sample

Enhancement here means more than cleanup. Take a Thai prompt asking for "10 words describing the benefits of renewable energy." The instruction is ambiguous in the original: it could mean a single ten-word sentence or ten separate sentences, and the original completion resolves it with one vague line. The enhanced version specifies the structure explicitly, a numbered list of ten sentences in Thai, and the completion delivers ten distinct, substantive points. Ambiguous quantity instructions produce unpredictable outputs. Fixing them creates consistent, measurable training signals.

Why localizing training data is not the same as translating it

Localization goes further because moving a sample between languages is not translation. A Thai prompt asking for an asteroid joke relies on Thai wordplay that has no equivalent in Tamil. Rather than translating the joke and losing it, Adaptive Data reoriented the sample to Singapore Tamil and produced a two-asteroid dialogue in colloquial register, using the natural spoken markers of Singapore Tamil casual speech rather than textbook Tamil. The result teaches the model to adapt creative tasks to cultural and linguistic context, not just to swap vocabulary.

Blog section illustration

How long manual data localization takes by comparison

The scale of what Adaption automates becomes clearer against the manual alternative: localizing and validating training data by hand takes significant time and specialized expertise. For example, three evaluation sets would typically take two visiting scholars two full working weeks to localize and validate: SEA-IFEval (105 rows), SEA-MTBench (58 x 2 rows), and a 1,000-row translation set. With Adaptive Data, 750,000 localized samples across five languages were processed in under a month.

What AutoScientist changes for enterprises

By automating the parts of model training that once required specialized engineering skills, Adaption’s platform opens the research loop to more of AISG’s team and shortens the path from idea to evaluated model.

What the team no longer builds by hand: data pipelines, failure handling as datasets grow, and hyperparameter exploration. Adaptive Data enriches and localizes seed data automatically, while AutoScientist abstracts the experimentation that previously demanded specialized tuning expertise.

Where the team focuses instead: the judgment calls that drive the science. AISG's team sources seed data, designs evaluation criteria, reviews intermediate and final outcomes, and adjusts the training recipe based on what they learn.

What runs in the background: AutoScientist manages the infrastructure behind long-running training and evaluation jobs and adds data to mitigate issues like catastrophic forgetting or lack of diversity, without researchers monitoring each run.

Build Sovereign AI with Adaption

The pattern behind this project is not specific to national AI programs. It applies to any organization with proprietary data and requirements that general-purpose models don't meet, whether those requirements are a language, a regulatory environment, or a domain vocabulary.

Adaption provides enterprises, early-stage startups, communities, and sovereign initiatives with ownership of their intelligence, building continual learning systems that let organizations shape, train, and own their AI. Three pillars make this possible: Adaptive Data for shaping training data, Adaptive Intelligence for models built for any industry or language, and Adaptive Interfaces for reimagining how people interact with AI.

The full case study includes side-by-side samples of enhanced and localized data in Thai and Tamil, with commentary on why each change matters for dataset quality. Read the full case study.

Author

Sudip Roy, Co-founder and Darius Liu

Date