South Korea's Ministry of Science and ICT said Thursday it will publicly release 35.44 million items of AI training data, totaling 1.56 trillion tokens, collected during a government contest to build proprietary foundation models. The datasets, gathered by teams led by Naver Cloud, Upstage AI, SK Telecom, NC AI, and LG AI Research, are enough to train models with roughly 70 billion to 80 billion parameters.
The data will be posted on the government's AI development support hub website. It includes large-scale pretraining datasets, multimodal content such as video and audio, and red-teaming data designed to test model safety.
What's in the release
The information was collected during the first round of the contest, which wrapped up earlier this year. The ministry's statement describes the release as a way to make the contest's work broadly available to developers and researchers.
For teams working on large language models, the scale matters. The 1.56 trillion token count places the dataset in the range used for serious foundation model pretraining, not just fine-tuning experiments. The inclusion of multimodal and safety-testing data also means the release covers more than text-only training.
Who benefits
The public release gives smaller teams and independent developers access to training material that would otherwise be expensive to assemble. The data was originally collected by five separate consortia, each with different approaches to model development. That variety could make the combined dataset more useful than a single-source corpus.
Developers working on Generative AI and LLM projects will likely find the pretraining data most relevant. Those focused on deployment and safety evaluation can draw on the red-teaming sets.
Why this matters for IT and development
For AI for IT & Development teams, the release removes a common bottleneck: access to large, cleaned training corpora. A team that previously lacked the resources to gather pretraining data at this scale can now experiment with model architectures that require it.
The practical takeaway is straightforward. If your work involves training or fine-tuning models in the 70B-80B parameter range, this dataset is worth evaluating when it goes live on the ministry's hub. The combination of pretraining, multimodal, and safety data in one release is unusual, and the fact that it comes from five competing consortia means the material has already been through real-world use in a national contest.
Your membership also unlocks: