How researchers adapted Dolma for better Thai language models

Imported from official source

AI Open Source Classified by Officially

Thai researchers adapted Ai2’s open Dolma toolkit to build Mangosteen, a 47-billion-token Thai corpus that filters low-quality web data while maintaining or improving model performance and strengthening Thai cultural knowledge.

This is an extract. The publication continues at the source.

Read the original at the source: https://allenai.org/blog/thai-llm-dolma

Officially imported this from Allen Institute for AI’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.

Provenance

Organization
Allen Institute for AI — imported from official source
Official source
https://allenai.org/rss.xml RSS
Imported
September 20, 2026 19:52
Versions
1 recorded
Identity
https://allenai.org/blog/thai-llm-dolma

Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.