Lifestyle

Taiwan Sovereign AI Training Corpus Opens Public Call for Content: License Terms and the State of Its Hakka-Language Data

On September 15, 2026, Taiwan's Ministry of Digital Affairs (moda) announced a public call for content for the Taiwan Sovereign AI Training Corpus, inviting publishers and writers to license their works for AI training. Drawing on three moda press releases and the license-terms page of the corpus website, this article covers the scope of the call, the royalty-free terms, the real limits of the "opt-out" option, and what the official documents leave unsaid. Checked September 17, 2026.

About 15 min read

Original illustration: three books on the left; an arrow to a stamped licensing document; another arrow to a large square holding nine small squares, representing books entering the corpus
Image: Mokaair (© Mokaair)

On September 15, 2026, Taiwan's Ministry of Digital Affairs (moda) formally launched a public call for content for the Taiwan Sovereign AI Training Corpus (臺灣主權AI訓練語料庫), inviting publishers, writers and cultural institutions to license their published works to the corpus for AI training. Minister Yi-Jing Lin (林宜敬), in his own capacity as an author, was the first to donate two of his own books, The Happy Ghost Island (幸福的鬼島) and Roving Rebels and Innovators (流寇與創新者); a number of publishers, e-book platforms and writers also signed licensing consent forms at the press conference. See the next section for the list.

This article was checked on September 17, 2026, against the full text of three moda press releases and the license-terms page of the Taiwan Sovereign AI Training Corpus website. This site has not tested anything itself, and does not judge for readers whether they should license or submit their work. The scale figures, dates and clause language cited below all reflect what these four official documents stated as of the day they were checked, and may change afterward.

The September 15 Press Conference: Who Signed What

moda held a "Public Content Call Launch Press Conference" the same day. Its release states that those attending in person to sign licensing consent forms "included publishers and e-book platforms such as Showwe Information Co., Ltd. (秀威資訊), INK (印刻文學), Readmoo (讀墨), foodNEXT (食力) and Cunext Group (巨思文化), as well as writers such as Lan Yifeng (藍弋丰) and Lee Kuei-shien (李魁賢, represented by family)" — this is a list of examples, not a complete roster. Writers Chen Mingzhong (陳明忠), Lian Mingwei (連明偉), Zhu Guozhen (朱國珍) and Zhu Hezhi (朱和之) are separately described as having "also contributed works in response." The official account describes these two groups separately and does not say whether the second group was present in person.

Minister Yi-Jing Lin said that how an AI model understands concepts such as democracy and checks on power is shaped by the social and cultural context embedded in its training corpus, and noted that the Chinese-language training data used by major international large language models is, at present, still mostly Simplified Chinese content. This is the minister's own characterization of the current situation; this site has not independently verified whether the claim itself is accurate.

moda emphasizes that the content call follows three principles — voluntary participation, explicit licensing, and the ability to opt out — and states that a complete process for requesting withdrawal is in place. The official release does not, however, explain the URL, form, or processing time for that process, nor does it set a deadline or a quota for this call; it only states that "moda sincerely invites authors, publishers, cultural institutions, corporate bodies and content providers across all fields to take part."

How the Corpus Grew to Its Current Scale

The Taiwan Sovereign AI Training Corpus did not first appear on September 15. moda published a launch announcement on December 24, 2025, whose text states that "[the corpus] currently has more than 200 government agencies participating, with over 2,000 datasets and more than 600 million tokens uploaded," covering fields such as language, culture, education, biology and geography. The official wording used is "currently," without giving a reference date for any of these three figures; the page itself carries a creation date of December 24, 2025 and a last-updated date of March 19, 2026.

On July 24, 2026, moda issued a press release announcing that the corpus had partnered with the Hakka Affairs Council to formally add Hakka-language resources, publications, cultural and historical archives, and research reports, adding about 20 million tokens of high-quality training data for the first time; moda did not state the actual go-live date. The release positions Hakka as "one of our nation's national languages." The same release states that the corpus has accumulated more than 1.5 billion tokens of data since going live, but moda did not attach any cutoff point to this cumulative figure.

By September 15, moda stated that the corpus had reached about 2.2 billion tokens "as of the end of August this year," and called the launch of the public content call a milestone marking the corpus's "entry into a new stage of development." Of the three announcements, only this one gives a reference date for its scale figure: the December release says "currently," the July release says "since going live," and neither carries a cutoff point. The three figures therefore rest on different bases for counting; this article does not add them together, subtract one from another, or convert them into a single same-day figure.

Source: moda press releases dated December 24, 2025, July 24, 2026 and September 15, 2026. Checked on September 17, 2026.
Release DateOfficial Scale StatedSource of DataVerification Basis
December 24, 2025"Currently" over 200 agencies, 2,000+ datasets, 600 million+ tokens (no reference date given)Datasets uploaded by government agenciesmoda launch announcement
July 24, 2026"Since going live," over 1.5 billion tokens accumulated (no reference date given)About 20 million tokens of new Hakka-language data from the Hakka Affairs Councilmoda Hakka-language data press release
September 15, 2026"As of the end of August this year," about 2.2 billion tokensLaunch of the public content call; current partners are publishers and e-book platformsmoda public content call press release

Who Can Contribute Data, What Counts, and Who Sits in the Middle

moda explains that, at this stage, it is partnering with publishers and e-book platforms that hold rich corpus resources, and is proceeding on a royalty-free licensing basis; authors who wish to contribute their own works may have their publisher help upload them to the corpus. The official release does not mention whether authors may apply to upload their works on their own. The priority content categories are four: publications licensed with consent, publication synopses, publication excerpts, and classics and creative works whose copyright protection has expired and that can be activated as public-domain resources.

The license terms cast the whole process in three roles: the "Licensor" is the rights holder who provides the data; the "Licensee" is the recipient who uses the data under the terms; and the "Corpus Platform" is the third party that obtains the data from the Licensor and, on the Licensor's instructions, bridges it to the Licensee. The terms state explicitly that the Corpus Platform "is not a party to any of the licensing relationships defined in this License," but that it may, with the Licensor's consent, evaluate and select recipients and/or validity periods of the data on the Licensor's behalf.

The license terms define "Corpus Data" more broadly than just books: it includes, but is not limited to, works usable for AI language and multimodal learning, such as text, music, artwork, photographs, graphics, audiovisual content and sound recordings, as well as compilation works resulting from the creative selection and arrangement of data. The September 15 release does not state how many books or tokens this call aims to collect before stopping, nor does it say whether a second round of calls will follow.

Four-panel diagram: who can contribute data, what content can be contributed, the licensing conditions, and how completed training results are unaffected by opting out
Source: moda press release dated September 15, 2026 and the Taiwan Sovereign AI Training Corpus License. Checked on September 17, 2026. · Image: Mokaair (© Mokaair)

What the License Terms Actually Do to a Contributor's Rights

The Taiwan Sovereign AI Training Corpus License – Version 1 (臺灣主權AI訓練語料授權條款-第1版), published on the corpus website, was, according to moda's December 24, 2025 launch announcement, jointly introduced by moda and the Taiwan Intellectual Property Office (TIPO), Ministry of Economic Affairs. The license terms call the individual or entity with legal rights who provides the data the "Licensor," and the party that receives and uses the data the "Licensee." What the Licensor grants the Licensee is the right to reproduce, adapt, compile and exercise other necessary rights under copyright and related rights, so that the data can be lawfully used for AI training; "this license is non-transferable and non-sublicensable," and its validity period may be specified by the Licensor or its representative as a set number of years or in perpetuity.

The terms also draw the real boundary of the "opt-out" option. Models, weights, code, documentation or other output that the Licensee produces through training belong to the Licensee or the model operator if they meet the requirements for copyright protection; and "even if the original corpus data is subsequently withdrawn from use, this does not affect training results already completed, including but not limited to the resulting models, weights, and code, documentation or other forms of output." The terms do not say whether copies of data already downloaded may continue to be used for training afterward. A separate proviso states that, unless the results are substantially similar to the original data and negatively affect the market or value of the original data, and exceed a reasonable scope of learning, the Licensor agrees not to treat the AI training results as an adaptation or compilation of its own work, and agrees not to impose any restrictions on, or bring litigation over, the subsequent use, modification, redistribution or computation of the results.

Unless exempted by law or otherwise agreed by the Licensor, models or tools that the Licensee trains using the data must carry identifying information provided by the Corpus Platform — the terms give examples such as the dataset name, version number, data provider, year of release, and the dataset's application page or official website — and must note that it was released under this license. The terms state explicitly, however, that the Licensee "is not required to attribute each individual component within the Corpus Data" item by item, meaning there is no need to list the source of every single book. Where the output takes the form of code or documentation, the attribution obligation can be simplified; the example the terms give is that when such output is publicly performed or displayed, it should be labeled "AI-generated output." The metadata on the license page shows it was last modified on September 3, 2026, before the September 15 press conference.

What the Terms Don't Say, and How to Check for Yourself

The disclaimer in the terms is stated plainly: this is a "royalty-free basis," under which the Licensor asserts only the copyright and related rights it holds in the data itself, provides no warranty of any kind for any other matter, and bears no responsibility for direct or indirect losses arising from use of the data. The terms separately list ten legally or ethically controversial activities for which the Licensee alone bears responsibility, including violating international or local laws, harming minors, generating false information, using the data to track personal information, and violating the lawful rights of others. The terms state that the Licensor — the individual or entity that provided the data — is not, merely by providing it, to be deemed to have expressed consent, permission or approval for such activities.

As of the September 17, 2026 check of the four official documents cited in this article, this article found no URL, form or processing time for the withdrawal process, nor any figure for how many works this content call has received so far, nor which agency or vendor operates and maintains the corpus. Compensation, however, is not a matter the documents stay silent on — they do address it: the press release states the call is "carried out on a royalty-free licensing basis," and the license terms describe a "royalty-free basis"; neither is a fee or a revenue share.

Publishers or authors interested in contributing data can check the Taiwan Sovereign AI Training Corpus's official website directly, or call the inquiry line published in the press release, 0800-023-300. The terms also state that both the Traditional Chinese and English versions are official texts with equal legal effect; where the two language versions differ in wording, this article follows only the Chinese version.

Frequently asked questions

Does this content call have a deadline?

The press release doesn't set one. moda's September 15, 2026 release does not give a deadline or a quota for this public content call, and as of the September 17, 2026 check, the other three official documents cited in this article showed no deadline of any kind either.

Is there a fee or revenue share for contributing data?

What the official documents state is royalty-free licensing. moda's press release explains that, at this stage, it is partnering with publishers and e-book platforms "on a royalty-free licensing basis," and the Taiwan Sovereign AI Training Corpus License is likewise established on a "royalty-free basis." As of the September 17, 2026 check, the four official documents cited in this article showed no fee, royalty, or revenue-sharing mechanism of any kind.

If I change my mind, can I withdraw data I've already licensed?

moda states that a process for requesting withdrawal is in place, but as of the September 17, 2026 check, the four official documents cited in this article showed no URL, form, or applicant-eligibility details for that process. More importantly, the license terms state explicitly that "even if the original corpus data is subsequently withdrawn from use, this does not affect training results already completed," so models, weights and output already trained will not be clawed back because of a withdrawal. Whether copies of data already downloaded can continue to be used for training afterward is not addressed in the terms.

Was the Hakka-language data added as part of this content call?

No, it's a separate batch. The Hakka-language data was added through a partnership between moda and the Hakka Affairs Council; the July 24, 2026 press release states it went live formally "recently," adding about 20 million tokens for the first time. moda did not give an exact go-live date — July 24 is only the release's publication date. The September 15 call is a separate effort aimed at publishers and writers; the two have different sources and different timing.

How much data has the corpus accumulated in total right now?

According to moda's September 15, 2026 press release, the corpus had reached about 2.2 billion tokens "as of the end of August this year"; the same release also states that the corpus went live at the end of 2025 and initially drew mainly on data from central and local government agencies. The scale figures in the other two releases are given as "currently" and "since going live" respectively, neither with a reference date, so this article does not combine the three figures into a single "current" number.

Can authors submit their own work to the corpus directly?

The path the official documents describe runs through a publisher. moda's press release explains that authors who wish to contribute their own works "may have their publisher help upload them to the corpus." As of the September 17, 2026 check, the four official documents cited in this article gave no indication of whether authors can bypass their publisher and apply to upload their work themselves.

Latest travel guides

Sources

Lifestyle