South Korea opens 1.56 trillion tokens of AI contest data to public

South Korea will publicly release 35.44 million AI training items totaling 1.56 trillion tokens, enough to train 70B-80B parameter models. The data, collected from five consortia including Naver Cloud and LG AI Research, includes multimodal and red-teaming sets.

Categorized in: AI News IT and Development
Published on: Aug 27, 2026
South Korea opens 1.56 trillion tokens of AI contest data to public

South Korea's Ministry of Science and ICT said Thursday it will publicly release 35.44 million items of AI training data, totaling 1.56 trillion tokens, collected during a government contest to build proprietary foundation models. The datasets, gathered by teams led by Naver Cloud, Upstage AI, SK Telecom, NC AI, and LG AI Research, are enough to train models with roughly 70 billion to 80 billion parameters.

The data will be posted on the government's AI development support hub website. It includes large-scale pretraining datasets, multimodal content such as video and audio, and red-teaming data designed to test model safety.

What's in the release

The information was collected during the first round of the contest, which wrapped up earlier this year. The ministry's statement describes the release as a way to make the contest's work broadly available to developers and researchers.

For teams working on large language models, the scale matters. The 1.56 trillion token count places the dataset in the range used for serious foundation model pretraining, not just fine-tuning experiments. The inclusion of multimodal and safety-testing data also means the release covers more than text-only training.

Who benefits

The public release gives smaller teams and independent developers access to training material that would otherwise be expensive to assemble. The data was originally collected by five separate consortia, each with different approaches to model development. That variety could make the combined dataset more useful than a single-source corpus.

Developers working on Generative AI and LLM projects will likely find the pretraining data most relevant. Those focused on deployment and safety evaluation can draw on the red-teaming sets.

Why this matters for IT and development

For AI for IT & Development teams, the release removes a common bottleneck: access to large, cleaned training corpora. A team that previously lacked the resources to gather pretraining data at this scale can now experiment with model architectures that require it.

The practical takeaway is straightforward. If your work involves training or fine-tuning models in the 70B-80B parameter range, this dataset is worth evaluating when it goes live on the ministry's hub. The combination of pretraining, multimodal, and safety data in one release is unusual, and the fact that it comes from five competing consortia means the material has already been through real-world use in a national contest.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)