Lifestyle
Taiwan Sovereign AI Training Corpus Opens Public Call for Content: License Terms and the State of Its Hakka-Language Data
On September 15, 2026, Taiwan's Ministry of Digital Affairs (moda) announced a public call for content for the Taiwan Sovereign AI Training Corpus, inviting publishers and writers to license their works for AI training. Drawing on three moda press releases and the license-terms page of the corpus website, this article covers the scope of the call, the royalty-free terms, the real limits of the "opt-out" option, and what the official documents leave unsaid. Checked September 17, 2026.
About 15 min read

On September 15, 2026, Taiwan's Ministry of Digital Affairs (moda) formally launched a public call for content for the Taiwan Sovereign AI Training Corpus (臺灣主權AI訓練語料庫), inviting publishers, writers and cultural institutions to license their published works to the corpus for AI training. Minister Yi-Jing Lin (林宜敬), in his own capacity as an author, was the first to donate two of his own books, The Happy Ghost Island (幸福的鬼島) and Roving Rebels and Innovators (流寇與創新者); a number of publishers, e-book platforms and writers also signed licensing consent forms at the press conference. See the next section for the list.
This article was checked on September 17, 2026, against the full text of three moda press releases and the license-terms page of the Taiwan Sovereign AI Training Corpus website. This site has not tested anything itself, and does not judge for readers whether they should license or submit their work. The scale figures, dates and clause language cited below all reflect what these four official documents stated as of the day they were checked, and may change afterward.
The September 15 Press Conference: Who Signed What
moda held a "Public Content Call Launch Press Conference" the same day. Its release states that those attending in person to sign licensing consent forms "included publishers and e-book platforms such as Showwe Information Co., Ltd. (秀威資訊), INK (印刻文學), Readmoo (讀墨), foodNEXT (食力) and Cunext Group (巨思文化), as well as writers such as Lan Yifeng (藍弋丰) and Lee Kuei-shien (李魁賢, represented by family)" — this is a list of examples, not a complete roster. Writers Chen Mingzhong (陳明忠), Lian Mingwei (連明偉), Zhu Guozhen (朱國珍) and Zhu Hezhi (朱和之) are separately described as having "also contributed works in response." The official account describes these two groups separately and does not say whether the second group was present in person.
Minister Yi-Jing Lin said that how an AI model understands concepts such as democracy and checks on power is shaped by the social and cultural context embedded in its training corpus, and noted that the Chinese-language training data used by major international large language models is, at present, still mostly Simplified Chinese content. This is the minister's own characterization of the current situation; this site has not independently verified whether the claim itself is accurate.
moda emphasizes that the content call follows three principles — voluntary participation, explicit licensing, and the ability to opt out — and states that a complete process for requesting withdrawal is in place. The official release does not, however, explain the URL, form, or processing time for that process, nor does it set a deadline or a quota for this call; it only states that "moda sincerely invites authors, publishers, cultural institutions, corporate bodies and content providers across all fields to take part."
How the Corpus Grew to Its Current Scale
The Taiwan Sovereign AI Training Corpus did not first appear on September 15. moda published a launch announcement on December 24, 2025, whose text states that "[the corpus] currently has more than 200 government agencies participating, with over 2,000 datasets and more than 600 million tokens uploaded," covering fields such as language, culture, education, biology and geography. The official wording used is "currently," without giving a reference date for any of these three figures; the page itself carries a creation date of December 24, 2025 and a last-updated date of March 19, 2026.
On July 24, 2026, moda issued a press release announcing that the corpus had partnered with the Hakka Affairs Council to formally add Hakka-language resources, publications, cultural and historical archives, and research reports, adding about 20 million tokens of high-quality training data for the first time; moda did not state the actual go-live date. The release positions Hakka as "one of our nation's national languages." The same release states that the corpus has accumulated more than 1.5 billion tokens of data since going live, but moda did not attach any cutoff point to this cumulative figure.
By September 15, moda stated that the corpus had reached about 2.2 billion tokens "as of the end of August this year," and called the launch of the public content call a milestone marking the corpus's "entry into a new stage of development." Of the three announcements, only this one gives a reference date for its scale figure: the December release says "currently," the July release says "since going live," and neither carries a cutoff point. The three figures therefore rest on different bases for counting; this article does not add them together, subtract one from another, or convert them into a single same-day figure.
| Release Date | Official Scale Stated | Source of Data | Verification Basis |
|---|---|---|---|
| December 24, 2025 | "Currently" over 200 agencies, 2,000+ datasets, 600 million+ tokens (no reference date given) | Datasets uploaded by government agencies | moda launch announcement |
| July 24, 2026 | "Since going live," over 1.5 billion tokens accumulated (no reference date given) | About 20 million tokens of new Hakka-language data from the Hakka Affairs Council | moda Hakka-language data press release |
| September 15, 2026 | "As of the end of August this year," about 2.2 billion tokens | Launch of the public content call; current partners are publishers and e-book platforms | moda public content call press release |
Who Can Contribute Data, What Counts, and Who Sits in the Middle
moda explains that, at this stage, it is partnering with publishers and e-book platforms that hold rich corpus resources, and is proceeding on a royalty-free licensing basis; authors who wish to contribute their own works may have their publisher help upload them to the corpus. The official release does not mention whether authors may apply to upload their works on their own. The priority content categories are four: publications licensed with consent, publication synopses, publication excerpts, and classics and creative works whose copyright protection has expired and that can be activated as public-domain resources.
The license terms cast the whole process in three roles: the "Licensor" is the rights holder who provides the data; the "Licensee" is the recipient who uses the data under the terms; and the "Corpus Platform" is the third party that obtains the data from the Licensor and, on the Licensor's instructions, bridges it to the Licensee. The terms state explicitly that the Corpus Platform "is not a party to any of the licensing relationships defined in this License," but that it may, with the Licensor's consent, evaluate and select recipients and/or validity periods of the data on the Licensor's behalf.
The license terms define "Corpus Data" more broadly than just books: it includes, but is not limited to, works usable for AI language and multimodal learning, such as text, music, artwork, photographs, graphics, audiovisual content and sound recordings, as well as compilation works resulting from the creative selection and arrangement of data. The September 15 release does not state how many books or tokens this call aims to collect before stopping, nor does it say whether a second round of calls will follow.
What the License Terms Actually Do to a Contributor's Rights
The Taiwan Sovereign AI Training Corpus License – Version 1 (臺灣主權AI訓練語料授權條款-第1版), published on the corpus website, was, according to moda's December 24, 2025 launch announcement, jointly introduced by moda and the Taiwan Intellectual Property Office (TIPO), Ministry of Economic Affairs. The license terms call the individual or entity with legal rights who provides the data the "Licensor," and the party that receives and uses the data the "Licensee." What the Licensor grants the Licensee is the right to reproduce, adapt, compile and exercise other necessary rights under copyright and related rights, so that the data can be lawfully used for AI training; "this license is non-transferable and non-sublicensable," and its validity period may be specified by the Licensor or its representative as a set number of years or in perpetuity.
The terms also draw the real boundary of the "opt-out" option. Models, weights, code, documentation or other output that the Licensee produces through training belong to the Licensee or the model operator if they meet the requirements for copyright protection; and "even if the original corpus data is subsequently withdrawn from use, this does not affect training results already completed, including but not limited to the resulting models, weights, and code, documentation or other forms of output." The terms do not say whether copies of data already downloaded may continue to be used for training afterward. A separate proviso states that, unless the results are substantially similar to the original data and negatively affect the market or value of the original data, and exceed a reasonable scope of learning, the Licensor agrees not to treat the AI training results as an adaptation or compilation of its own work, and agrees not to impose any restrictions on, or bring litigation over, the subsequent use, modification, redistribution or computation of the results.
Unless exempted by law or otherwise agreed by the Licensor, models or tools that the Licensee trains using the data must carry identifying information provided by the Corpus Platform — the terms give examples such as the dataset name, version number, data provider, year of release, and the dataset's application page or official website — and must note that it was released under this license. The terms state explicitly, however, that the Licensee "is not required to attribute each individual component within the Corpus Data" item by item, meaning there is no need to list the source of every single book. Where the output takes the form of code or documentation, the attribution obligation can be simplified; the example the terms give is that when such output is publicly performed or displayed, it should be labeled "AI-generated output." The metadata on the license page shows it was last modified on September 3, 2026, before the September 15 press conference.
What the Terms Don't Say, and How to Check for Yourself
The disclaimer in the terms is stated plainly: this is a "royalty-free basis," under which the Licensor asserts only the copyright and related rights it holds in the data itself, provides no warranty of any kind for any other matter, and bears no responsibility for direct or indirect losses arising from use of the data. The terms separately list ten legally or ethically controversial activities for which the Licensee alone bears responsibility, including violating international or local laws, harming minors, generating false information, using the data to track personal information, and violating the lawful rights of others. The terms state that the Licensor — the individual or entity that provided the data — is not, merely by providing it, to be deemed to have expressed consent, permission or approval for such activities.
As of the September 17, 2026 check of the four official documents cited in this article, this article found no URL, form or processing time for the withdrawal process, nor any figure for how many works this content call has received so far, nor which agency or vendor operates and maintains the corpus. Compensation, however, is not a matter the documents stay silent on — they do address it: the press release states the call is "carried out on a royalty-free licensing basis," and the license terms describe a "royalty-free basis"; neither is a fee or a revenue share.
Publishers or authors interested in contributing data can check the Taiwan Sovereign AI Training Corpus's official website directly, or call the inquiry line published in the press release, 0800-023-300. The terms also state that both the Traditional Chinese and English versions are official texts with equal legal effect; where the two language versions differ in wording, this article follows only the Chinese version.
Frequently asked questions
Does this content call have a deadline?
The press release doesn't set one. moda's September 15, 2026 release does not give a deadline or a quota for this public content call, and as of the September 17, 2026 check, the other three official documents cited in this article showed no deadline of any kind either.
Is there a fee or revenue share for contributing data?
What the official documents state is royalty-free licensing. moda's press release explains that, at this stage, it is partnering with publishers and e-book platforms "on a royalty-free licensing basis," and the Taiwan Sovereign AI Training Corpus License is likewise established on a "royalty-free basis." As of the September 17, 2026 check, the four official documents cited in this article showed no fee, royalty, or revenue-sharing mechanism of any kind.
If I change my mind, can I withdraw data I've already licensed?
moda states that a process for requesting withdrawal is in place, but as of the September 17, 2026 check, the four official documents cited in this article showed no URL, form, or applicant-eligibility details for that process. More importantly, the license terms state explicitly that "even if the original corpus data is subsequently withdrawn from use, this does not affect training results already completed," so models, weights and output already trained will not be clawed back because of a withdrawal. Whether copies of data already downloaded can continue to be used for training afterward is not addressed in the terms.
Was the Hakka-language data added as part of this content call?
No, it's a separate batch. The Hakka-language data was added through a partnership between moda and the Hakka Affairs Council; the July 24, 2026 press release states it went live formally "recently," adding about 20 million tokens for the first time. moda did not give an exact go-live date — July 24 is only the release's publication date. The September 15 call is a separate effort aimed at publishers and writers; the two have different sources and different timing.
How much data has the corpus accumulated in total right now?
According to moda's September 15, 2026 press release, the corpus had reached about 2.2 billion tokens "as of the end of August this year"; the same release also states that the corpus went live at the end of 2025 and initially drew mainly on data from central and local government agencies. The scale figures in the other two releases are given as "currently" and "since going live" respectively, neither with a reference date, so this article does not combine the three figures into a single "current" number.
Can authors submit their own work to the corpus directly?
The path the official documents describe runs through a publisher. moda's press release explains that authors who wish to contribute their own works "may have their publisher help upload them to the corpus." As of the September 17, 2026 check, the four official documents cited in this article gave no indication of whether authors can bypass their publisher and apply to upload their work themselves.
2026 Tech News Roundup: Key Points on Hardware, Platforms, Telecom, and Regulation2026 Tech News Roundup: Key Points on Hardware, Platforms, Telecom, and RegulationThis site's 2026 tech-news explainers in five groups — hardware, platforms, computing infrastructure, Taiwan policy, EU regulation: iPhone Duo and September hardware, M6/M5 Ultra, Snapdragon 8 Elite Gen 6, Project Zenith, Pixel Drop, App Store subscriptions, WordPress and Synology patches, NVIDIA, 6G, the Taiwan–Matsu cables, the sovereign AI corpus, the Cyber Resilience Act, the KIDS Act and Apple's EU terms. No purchase advice; vendor claims attributed; official sources and check dates.Read the full article
moda Hosts International 6G Spectrum Symposium: What Was Discussed, What Remains Undecidedmoda Hosts International 6G Spectrum Symposium: What Was Discussed, What Remains UndecidedOn September 10, 2026, Taiwan's Ministry of Digital Affairs (moda) held the "Spectrum for Tomorrow" International Symposium, inviting ITU-R Study Group 5's chairman and spectrum-policy experts from several countries to discuss 6G's spectrum demand from land-space integration and AI. Using moda's press release, program page, and the Radio Frequency Supply Plan (amended February 2025), this article lays out what was discussed and what the official documents had not yet said as of the check date.Read the full article
Lifestyle
moda Hosts International 6G Spectrum Symposium: What Was Discussed, What Remains Undecided
On September 10, 2026, Taiwan's Ministry of Digital Affairs (moda) held the "Spectrum for Tomorrow" International Symposium, inviting ITU-R Study Group 5's chairman and spectrum-policy experts from several countries to discuss 6G's spectrum demand from land-space integration and AI. Using moda's press release, program page, and the Radio Frequency Supply Plan (amended February 2025), this article lays out what was discussed and what the official documents had not yet said as of the check date.
Lifestyle
Taiwan–Matsu No. 4 Submarine Cable and Matsu’s Microwave Stations: moda Says Completion Is Coming Soon, and the State of Matsu’s Communications Backups
On June 23, 2026, moda said the Taiwan–Matsu No. 4 cable and microwave stations in Matsu’s four townships, built with Forward-looking Infrastructure subsidies, are due to be completed soon; once active, there will be 3 cables and at least 2 microwave systems per township. This article traces the cable outages and backups of the first half of 2026 from official notices, plus the official cable list as of September 17, 2026. This site did no on-site testing and gives no travel or carrier advice.
Lifestyle
Cloudflare launches Traces in public beta: site operators can follow every step a request takes through the platform on one timeline
On October 2, 2026, Cloudflare announced the public beta of Cloudflare Traces. Website operators and developers using Cloudflare can see a request pass through security rules, caching, routing and the origin server on a single timeline, making it easier to find why a request was blocked or slowed. New pricing takes effect on December 1, 2026. Information comes from the official Cloudflare blog.
Lifestyle
Cloudflare launches Web Search API via AI Gateway, requiring search partners to follow its crawler rules
On October 2, 2026, Cloudflare announced a Web Search API that lets AI agents query live web information through AI Gateway. The first partners are Ceramic.ai, Exa and Linkup. Cloudflare says these partners' crawlers must meet its Verified bots requirements and cite sources. This matters both to developers building AI applications and to website owners whose content may be crawled.
Articles that cite this one
Latest travel guides

GuideTokyo
Where to Stay in Tokyo: Comparing Shinjuku, Ueno, Tokyo Station, Shibuya, Asakusa, Ikebukuro, and Ginza, Plus Airport Access, Accommodation Tax, and Luggage Delivery
Where should you stay in Tokyo? Compare Shinjuku, Ueno, Tokyo Station, Shibuya, Asakusa, Ikebukuro, and Ginza by the same criteria: access from Narita and Haneda, transit routes, nearby attractions, neighborhood character, and who each area suits. Includes a comparison table, a Yamanote Line diagram, Tokyo’s accommodation tax as verified in 2026/9 (changing to 3% in 2027/4), and Airport TA-Q-BIN luggage shipping rules.
- Budget
- Hotels

GuideTokyo
How to Choose Tokyo Transit Passes: Are Suica, Welcome Suica, the Tokyo Subway Ticket, and the JR Pass Worth It?
On a first Tokyo trip, start with an IC card and pay per ride (Welcome Suica has no deposit and is valid for 28 days). If you take four or more subway rides in a day, add a 72-hour Tokyo Subway Ticket for 2,000 yen; a JR Pass is never worthwhile if you stay in Tokyo and do not go to Kansai. See what TOURIST PASMO, Suica on iPhone, and the Tokyo Metro day pass do and do not cover, with a decision chart. Prices verified in September 2026.
- Transport
- Budget

GuideTokyo
Tokyo Disneyland and DisneySea Guide: Ticket Prices, Fantasy Springs, Disney Premier Access (DPA), Standby Pass, and Which Park to Choose for Your First Visit
Tokyo Disney one-day Passport prices vary: most weekdays in 9/2026 cost ¥9,900 and weekends ¥10,900. At 14:00 daily, tickets go on sale for the same date two months later. Free Priority Pass is no longer on the official service list; only paid Disney Premier Access (¥1,000–3,500 per person per use) shortens waits. Covers hours, the 25th anniversary, Standby Pass, Entry Request, Fantasy Springs access and first-visit park choice; checked on the official site in 9/2026.
- Itineraries
- Family
Sources
- moda Launches Public Call for Content for Sovereign AI Corpus, Rallying Writers and Publishers to Take Part · Checked:
- "Taiwan Sovereign AI Training Corpus" Goes Live! Ministry of Digital Affairs Partners with 200 Agencies to Build a Local Data Resource · Checked:
- moda Adds 20 Million Tokens of Hakka-Language Data to the "Taiwan Sovereign AI Training Corpus," Strengthening the Local AI Language Data Foundation · Checked:
- Taiwan Sovereign AI Training Corpus License – Version 1 · Checked: