(17 Sep 2026) The National Library of Korea is releasing 38.33 million AI learning datasets, processed from the library’s national document collection, to the public for the first time.
The library announced that, starting September 17, it will make available via the “Gongyoo Seojae” platform a total of 3,974 text datasets, 21,485 image datasets—including tables, illustrations, advertisements, and photos—plus metadata and 38.3 million character-level datasets.
The datasets include modern magazines from the 1930s–1940s whose copyright issues have been resolved, government publications and textbooks from the 1940s–1960s, material published by the library, and open access (OA) academic papers licensed for AI training. Users on the Gongyoo Seojae platform can freely view, download, and use these materials.
The text datasets have been refined to be easily recognized by AI, so that they can be used immediately in research and service development. The character-level datasets include various forms of character data, which can be utilized for developing optical character recognition (OCR) and natural language processing technologies—areas where it is difficult for small and venture companies to secure large-scale learning data, according to the library.
Find out more here.




