NüshuRescue: Reviving the Endangered Nüshu Language with AI - ACL Anthology ACL Anthology
AboutAnnouncementsCommunication channelsRelated workCopyrightCreditsVolunteerDevelopmentFeedbackUsingAuthor directoryCiting papersLinks in the AnthologyData accessAll FAQsDetailsAnthology identifiersNamesORCID iDsDOIsVerified authorsContributionsSubmissionsCorrectionsMaintain author pagesAttachmentsGitHub
NüshuRescue: Reviving the Endangered Nüshu Language with AIIvory Yang, Weicheng Ma, Soroush VosoughiCorrect Metadata for Use this form to create a GitHub issue with structured data describing the correction. You will need a GitHub account. Once you create that issue, the correction will be reviewed by a staff member.⚠️ Mobile Users: Submitting this form to create a new issue will only work with github.com, not the GitHub Mobile app.Important: The Anthology treat PDFs as authoritative. Please use this form only to correct data that is out of line with the PDF. See our corrections guidelines if you need to change the PDF.Title Adjust the title. Retain tags such as <fixed-case>. Authors Adjust author names and order to match the PDF.Add AuthorAbstract Correct abstract if needed. Retain XML formatting tags such as <tex-math>. You may use <b>...</b> for bold, <i>...</i> for italic, <u>...</u> for underline, <sc>...</sc> for small-caps, <tt>...<tt> for typewriter text, <url>...</url> for URLs, <a href=...> for hyperlinks, and <par/> for paragraph breaks. Verification against PDF Ensure that the new title/authors match the snapshot below. (If there is no snapshot or it is too small, consult the PDF.)Authors concatenated from the text boxes above: ALL author names match the snapshot above—including middle initials, hyphens, and accents.Create GitHub issue for staff reviewAbstractThe preservation and revitalization of endangered and extinct languages is a meaningful endeavor, conserving cultural heritage while enriching fields like linguistics and anthropology. However, these languages are typically low-resource, making their reconstruction labor-intensive and costly. This challenge is exemplified by Nüshu, a rare script historically used by Yao women in China for self-expression within a patriarchal society. To address this challenge, we introduce NüshuRescue, an AI-driven framework designed to train large language models (LLMs) on endangered languages with minimal data. NüshuRescue automates evaluation and expands target corpora to accelerate linguistic revitalization. As a foundational component, we developed NCGold, a 500-sentence Nüshu-Chinese parallel corpus, the first publicly available dataset of its kind. Leveraging GPT-4-Turbo, with no prior exposure to Nüshu and only 35 short examples from NCGold, NüshuRescue achieved 48.69% translation accuracy on 50 withheld sentences and generated NCSilver, a set of 98 newly translated modern Chinese sentences of varying lengths. In addition, we developed FastText-based and Seq2Seq models to further support research on Nüshu. NüshuRescue provides a versatile and scalable tool for the revitalization of endangered languages, minimizing the need for extensive human input. All datasets and code have been made publicly available at https://github.com/ivoryayang/NushuRescue.Anthology ID:2025.coling-main.468Volume:Proceedings of the 31st International Conference on Computational LinguisticsMonth:JanuaryYear:2025Address:Abu Dhabi, UAEEditors:Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, Steven SchockaertVenue:COLINGSIG:Publisher:Association for Computational LinguisticsNote:Pages:7020–7034Language:URL:https://aclanthology.org/2025.coling-main.468/DOI:Bibkey:yang-etal-2025-nushurescueCite (ACL):Ivory Yang, Weicheng Ma, and Soroush Vosoughi. 2025. NüshuRescue: Reviving the Endangered Nüshu Language with AI. In Proceedings of the 31st International Conference on Computational Linguistics, pages 7020–7034, Abu Dhabi, UAE. Association for Computational Linguistics.Cite (Informal):NüshuRescue: Reviving the Endangered Nüshu Language with AI (Yang et al., COLING 2025)Copy Citation:BibTeX Markdown MODS XML Endnote More options…PDF:https://aclanthology.org/2025.coling-main.468.pdfPDF Cite Search
Fix dataExport citationBibTeXMODS XMLEndnotePreformatted@inproceedings{yang-etal-2025-nushurescue, title = {{N}{\"u}shu{R}escue: Reviving the Endangered N{\"u}shu Language with {AI}}, author = "Yang, Ivory and Ma, Weicheng and Vosoughi, Soroush", editor = "Rambow, Owen and Wanner, Leo and Apidianaki, Marianna and Al-Khalifa, Hend and Eugenio, Barbara Di and Schockaert, Steven", booktitle = "Proceedings of the 31st International Conference on Computational Linguistics", month = jan, year = "2025", address = "Abu Dhabi, UAE", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2025.coling-main.468/", pages = "7020--7034", abstract = {The preservation and revitalization of endangered and extinct languages is a meaningful endeavor, conserving cultural heritage while enriching fields like linguistics and anthropology. However, these languages are typically low-resource, making their reconstruction labor-intensive and costly. This challenge is exemplified by N{\"u}shu, a rare script historically used by Yao women in China for self-expression within a patriarchal society. To address this challenge, we introduce N{\"u}shuRescue, an AI-driven framework designed to train large language models (LLMs) on endangered languages with minimal data. N{\"u}shuRescue automates evaluation and expands target corpora to accelerate linguistic revitalization. As a foundational component, we developed NCGold, a 500-sentence N{\"u}shu-Chinese parallel corpus, the first publicly available dataset of its kind. Leveraging GPT-4-Turbo, with no prior exposure to N{\"u}shu and only 35 short examples from NCGold, N{\"u}shuRescue achieved 48.69{\%} translation accuracy on 50 withheld sentences and generated NCSilver, a set of 98 newly translated modern Chinese sentences of varying lengths. In addition, we developed FastText-based and Seq2Seq models to further support research on N{\"u}shu. N{\"u}shuRescue provides a versatile and scalable tool for the revitalization of endangered languages, minimizing the need for extensive human input. All datasets and code have been made publicly available at \url{https://github.com/ivoryayang/NushuRescue}.} }Download as File Copy to Clipboard<?xml version="1.0" encoding="UTF-8"?> <modsCollection xmlns="http://www.loc.gov/mods/v3"> <mods ID="yang-etal-2025-nushurescue"> <titleInfo> <title>NüshuRescue: Reviving the Endangered Nüshu Language with AI</title> </titleInfo> <name type="personal"> <namePart type="given">Ivory</namePart> <namePart type="family">Yang</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Weicheng</namePart> <namePart type="family">Ma</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Soroush</namePart> <namePart type="family">Vosoughi</namePart> <role> <roleTerm authority="marcrelator" type="text">author</roleTerm> </role> </name> <originInfo> <dateIssued>2025-01</dateIssued> </originInfo> <typeOfResource>text</typeOfResource> <relatedItem type="host"> <titleInfo> <title>Proceedings of the 31st International Conference on Computational Linguistics</title> </titleInfo> <name type="personal"> <namePart type="given">Owen</namePart> <namePart type="family">Rambow</namePart> <role> <roleTerm authority="marcrelator" type="text">editor</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Leo</namePart> <namePart type="family">Wanner</namePart> <role> <roleTerm authority="marcrelator" type="text">editor</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Marianna</namePart> <namePart type="family">Apidianaki</namePart> <role> <roleTerm authority="marcrelator" type="text">editor</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Hend</namePart> <namePart type="family">Al-Khalifa</namePart> <role> <roleTerm authority="marcrelator" type="text">editor</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Barbara</namePart> <namePart type="given">Di</namePart> <namePart type="family">Eugenio</namePart> <role> <roleTerm authority="marcrelator" type="text">editor</roleTerm> </role> </name> <name type="personal"> <namePart type="given">Steven</namePart> <namePart type="family">Schockaert</namePart> <role> <roleTerm authority="marcrelator" type="text">editor</roleTerm> </role> </name> <originInfo> <publisher>Association for Computational Linguistics</publisher> <place> <placeTerm type="text">Abu Dhabi, UAE</placeTerm> </place> </originInfo> <genre authority="marcgt">conference publication</genre> </relatedItem> <abstract>The preservation and revitalization of endangered and extinct languages is a meaningful endeavor, conserving cultural heritage while enriching fields like linguistics and anthropology. However, these languages are typically low-resource, making their reconstruction labor-intensive and costly. This challenge is exemplified by Nüshu, a rare script historically used by Yao women in China for self-expression within a patriarchal society. To address this challenge, we introduce NüshuRescue, an AI-driven framework designed to train large language models (LLMs) on endangered languages with minimal data. NüshuRescue automates evaluation and expands target corpora to accelerate linguistic revitalization. As a foundational component, we developed NCGold, a 500-sentence Nüshu-Chinese parallel corpus, the first publicly available dataset of its kind. Leveraging GPT-4-Turbo, with no prior exposure to Nüshu and only 35 short examples from NCGold, NüshuRescue achieved 48.69% translation accuracy on 50 withheld sentences and generated NCSilver, a set of 98 newly translated modern Chinese sentences of varying lengths. In addition, we developed FastText-based and Seq2Seq models to further support research on Nüshu. NüshuRescue provides a versatile and scalable tool for the revitalization of endangered languages, minimizing the need for extensive human input. All datasets and code have been made publicly available at https://github.com/ivoryayang/NushuRescue.</abstract> <identifier type="citekey">yang-etal-2025-nushurescue</identifier> <location> <url>https://aclanthology.org/2025.coling-main.468/</url> </location> <part> <date>2025-01</date> <extent unit="page"> <start>7020</start> <end>7034</end> </extent> </part> </mods> </modsCollection> Download as File Copy to Clipboard%0 Conference Proceedings %T NüshuRescue: Reviving the Endangered Nüshu Language with AI %A Yang, Ivory %A Ma, Weicheng %A Vosoughi, Soroush %Y Rambow, Owen %Y Wanner, Leo %Y Apidianaki, Marianna %Y Al-Khalifa, Hend %Y Eugenio, Barbara Di %Y Schockaert, Steven %S Proceedings of the 31st International Conference on Computational Linguistics %D 2025 %8 January %I Association for Computational Linguistics %C Abu Dhabi, UAE %F yang-etal-2025-nushurescue %X The preservation and revitalization of endangered and extinct languages is a meaningful endeavor, conserving cultural heritage while enriching fields like linguistics and anthropology. However, these languages are typically low-resource, making their reconstruction labor-intensive and costly. This challenge is exemplified by Nüshu, a rare script historically used by Yao women in China for self-expression within a patriarchal society. To address this challenge, we introduce NüshuRescue, an AI-driven framework designed to train large language models (LLMs) on endangered languages with minimal data. NüshuRescue automates evaluation and expands target corpora to accelerate linguistic revitalization. As a foundational component, we developed NCGold, a 500-sentence Nüshu-Chinese parallel corpus, the first publicly available dataset of its kind. Leveraging GPT-4-Turbo, with no prior exposure to Nüshu and only 35 short examples from NCGold, NüshuRescue achieved 48.69% translation accuracy on 50 withheld sentences and generated NCSilver, a set of 98 newly translated modern Chinese sentences of varying lengths. In addition, we developed FastText-based and Seq2Seq models to further support research on Nüshu. NüshuRescue provides a versatile and scalable tool for the revitalization of endangered languages, minimizing the need for extensive human input. All datasets and code have been made publicly available at https://github.com/ivoryayang/NushuRescue. %U https://aclanthology.org/2025.coling-main.468/ %P 7020-7034Download as File Copy to ClipboardMarkdown (Informal)[NüshuRescue: Reviving the Endangered Nüshu Language with AI](https://aclanthology.org/2025.coling-main.468/) (Yang et al., COLING 2025)NüshuRescue: Reviving the Endangered Nüshu Language with AI (Yang et al., COLING 2025)ACLIvory Yang, Weicheng Ma, and Soroush Vosoughi. 2025. NüshuRescue: Reviving the Endangered Nüshu Language with AI. In Proceedings of the 31st International Conference on Computational Linguistics, pages 7020–7034, Abu Dhabi, UAE. Association for Computational Linguistics.Copy Markdown to Clipboard Copy ACL to Clipboard ACL materials are Copyright © 1963–2026 ACL; other materials are copyrighted by their respective copyright holders. Materials prior to 2016 here are licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 3.0 International License. Permission is granted to make copies for the purposes of teaching and research. Materials published in or after 2016 are licensed on a Creative Commons Attribution 4.0 International License.The ACL Anthology is managed and built by the ACL Anthology team of volunteers.Site last built on 13 September 2026 at 19:42 UTC with commit d6430d2. |
The preservation and revitalization of endangered and extinct languages is presented as a significant endeavor that conserves cultural heritage while advancing fields such as linguistics and anthropology. A major obstacle in this work is that these languages are typically low-resource, which renders their reconstruction labor-intensive and expensive, a challenge exemplified by the Nüshu script, a rare writing system historically utilized by Yao women in China for self-expression within a patriarchal structure. To address this scarcity of data, the authors introduce NüshuRescue, an artificial intelligence driven framework specifically designed to train large language models (LLMs) using minimal linguistic data, aiming to accelerate the process of linguistic revitalization through automated evaluation and corpus expansion.
As a foundational element of this framework, the researchers developed NCGold, which constitutes a 500-sentence parallel corpus between Nüshu and Chinese, establishing it as the first publicly available dataset of this nature. The NüshuRescue methodology then leveraged this corpus and advanced models to perform translation tasks. Specifically, using GPT-4-Turbo, and despite having no prior exposure to Nüshu, the system was tested with only 35 short examples from NCGold. This process resulted in a translation accuracy of 48.69 percent when tested on 50 withheld sentences, and the system further generated NCSilver, which comprised a set of 98 newly translated modern Chinese sentences of varying lengths. Furthermore, the research incorporated the development of FastText-based and Seq2Seq models to provide additional support for ongoing research concerning the Nüshu language. Ultimately, NüshuRescue is positioned as a versatile and scalable instrument for endangered language revitalization, designed to minimize the reliance on extensive human input. All associated datasets and the corresponding code have been made publicly accessible through a designated repository. |