The First Sunaŵi Workshop
AI & Indigenous Language Revitalization Workshop
June 8–10, 2026; Loyola Marymount University, Los Angeles
What it is
Sunaŵi brought together people from very different backgrounds (Indigenous community members, computer scientists, AI experts, linguists, archivists, language teachers, nonprofit leaders, education experts, students, industry representatives, and junior and senior faculty) to surface shared concerns and move toward defining what responsible AI looks like in the context of Indigenous language revitalization. It went a long way toward breaking down some of the barriers that prevent us all from working more closely together.
This wiki is the workshop's living output. See the people who took part.
Program
Three days of talks, panels, and small-group work. Talks and a tutorial came from researchers, practitioners, and community members sharing their own experience. Panels put those perspectives in direct conversation with each other. The breakouts were where the room worked in small groups toward something concrete: the first mapped the concerns people brought with them, the second drafted recommendations, and the final morning was given to presenting that work back to the room. This wiki grew out of that material.
Each day ran roughly 9:00 to 5:00, with talks and panels in the mornings and breakouts in the afternoons. Times below are from the program as scheduled.
Day 1 — Monday, June 8
| Time | Session | |
|---|---|---|
| 9:00 | Opening remarks | Jared Coleman |
| 9:15 | A view on AI & Indigenous Languages from the Advancing Indigenous Language Technologies Working Group | Amy Fountain |
| 9:45 | Whose Data, Whose Voice? Exploring Indigenous Sovereignty from a Choctaw Perspective | Jacqueline Brixey |
| 10:30 | Tutorial: LLM-Assisted Rule-Based Machine Translation | Jared Coleman |
| 11:00 | Beyond Data Scarcity: Challenges and Opportunities for Indigenous Language NLP | Emily Prud'hommeaux |
| 11:30 | Panel 1: Concerns & Hopes | Helema Andrews, Jacqueline Brixey, Citlali Arvizu, Jay Prakash Thakur, Kiana Maillet |
| 2:00 | Breakout 1: Concern mapping | |
| 4:00 | Plenary cross-pollination |
Day 2 — Tuesday, June 9
| Time | Session | |
|---|---|---|
| 9:00 | Manuscripts, Microfiche, Typewriters, Databases and LLMs: Language and Technology at the Myaamia Center | Hunter Lockwood |
| 9:30 | AI for Endangered Language Documentation and Preservation | Antonios Anastasopoulos |
| 10:00 | Agentic AI Systems for Developing Language Education Tools | Khalil Iskarous |
| 11:00 | The AmericasNLP Shared Tasks: Results, Artifacts, and Future Directions | Abteen Ebrahimi |
| 11:30 | Between the Academy and the Community: Are Our Contributions Doing What We Think They Are? | Belu Ticona |
| 12:00 | Panel 2: Community & Capacity Building | Kristine Hildebrandt, Ben Ankiel, Melissa Floca, Daniel Bögre Udell |
| 2:30 | Breakout 2: Drafting recommendations | |
| 4:30 | Plenary cross-pollination |
Day 3 — Wednesday, June 10
| Time | Session | |
|---|---|---|
| 9:00 | Final breakout presentations | |
| 10:45 | Panel 3: Future of the Field | Raina Heaton, Khalil Iskarous, Luis Barragan, Tiffany Wright |
| 11:45 | Closing remarks |
Talk abstracts
In program order. The tutorial had no abstract.
A view on AI & Indigenous Languages from the Advancing Indigenous Language Technologies Working Group
This talk will focus on issues and opportunities arising from language work with communities here in the southwest, especially focusing on language policy and data sovereignty. We will discuss language ideology, policy and practical matters.
Whose Data, Whose Voice? Exploring Indigenous Sovereignty from a Choctaw Perspective
Data sovereignty is a topic that impacts all people in this era of big data and artificial intelligence. In short, it is the ability of a country to control and access the data that is generated in its territory. In this talk, I will explore what data sovereignty means from an Indigenous perspective, specifically from my viewpoint as a citizen of the Choctaw Nation of Oklahoma. I will show examples of policies that CNO has enacted. Additionally, I will present a framework for evaluating commercially available Large Language Models (LLMs) and explain how their unauthorized use of Indigenous data is undermining Indigenous language revitalization efforts.
Beyond Data Scarcity: Challenges and Opportunities for Indigenous Language NLP
The majority of the world's languages are underserved by both big tech companies and the NLP research community. Although the marginalization of these languages is driven to some extent by a general lack of training resources, the quality and content of the data that is available, rather than the overall quantity of data, play an important role. In this talk, I will discuss the impact of transcription quality, domain mismatch, and data modality on NLP task performance in endangered, indigenous, and under-resourced languages.
Manuscripts, Microfiche, Typewriters, Databases and LLMs: Language and Technology at the Myaamia Center
Miami-Illinois is an Algonquian language spoken in the southern Great Lakes region. Documentation of the language dates back to the late 17th century, and continued into the mid- to late-20th century with the death of the last generation of semispeakers and rememberers. Since the 1980s, there has been a concerted effort to revitalize the language using archival sources. Reclamation efforts originally relied on photocopies and microfiche of handwritten and typewritten notes, but serious digitization work began in the early 2000s. This work was supercharged after the development of ILDA (Indigenous Languages Digital Archive), allowing us to store, analyze, and share data in a more robust way, on a platform that was designed with the needs of Indigenous language communities in mind. Since 2022, as technology continues to evolve, we have explored the use of LLMs/AI, with a focus on practical applications for language documentation and revitalization. In this talk, I will situate the Myaamia Center's investigation of LLMs within our broader history of archives-based language revitalization and use of technology; discuss our current efforts to explore the potential of LLMs; and reflect on the ethical, privacy, and sovereignty considerations that guide our approach.
AI for Endangered Language Documentation and Preservation
In this talk I will cover three case studies from my lab's work, that center on projects around building language technologies for underserved languages in three different contexts in different continents. In particular, I will cover how we built educational tools for Mapuzugun in Chile, how we collected data to build the first translation and speech recognition system for Bemba in Zambia, and how we employ AI-human feedback loops to build resources for under-resourced endangered Greek varieties like Griko.
Agentic AI Systems for Developing Language Education Tools
Communities often need fun and engaging learning tools to help their children learn their endangered language, when most engaging materials are in some other dominant language. This talk will describe work of this sort performed with member of the Ladin community in Italy 8 years ago, and how modern agentic AI systems could transform the development of such tools.
The AmericasNLP Shared Tasks: Results, Artifacts, and Future Directions
Natural language processing tools have great potential in aiding language documentation and revitalization efforts, however, for many languages the quality of these systems is far below the minimum threshold needed for real-world effectiveness. In this talk I will present the AmericasNLP Shared Tasks, hosted as part of the AmericasNLP workshop since its inception in 2021, and how these competitions can provide, in addition to pure research, an additional avenue towards addressing the weaknesses of current language technologies available for data-scarce languages.
Between the Academy and the Community: Are Our Contributions Doing What We Think They Are?
How do academic language technology initiatives actually impact the people they claim to serve? From rural speakers to urban descendants, the reality of who those people are varies widely. In this talk, Belu draws on her indigenous roots and a South American perspective to reflect on how researchers engage with "communities," and whether that engagement aligns with a researcher's positionality, particularly when that researcher shares roots with the people being studied.
How it came to be
It goes all the way back to why I chose computer science as a major in the first place: I wanted to create a "Rosetta Stone" (the language-learning software) for my tribe's critically endangered language, Owens Valley Paiute. My research ended up going in a very different direction (distributed computing), but I never lost sight of that original motivation. I helped create online tools (like an online version of Glenn Nelson's Owens Valley Paiute dictionary) to help make the language more accessible to my community, but it wasn't until ChatGPT exploded in 2022 that I saw, for the first time, an opportunity to do research in this area as well.
I started working with linguists, archivists, and other scholars in language documentation, preservation, and revitalization. Until then, my only real exposure to the field was through my own tribe's language program, and most of that was interacting with historical materials. As many of the workshop participants can attest, the traditional language documentation field was quite extractive. Linguists would come into communities, gather data from speakers, produce potentially valuable resources (dictionaries, grammars, etc.), and then publish them behind academic paywalls or in ways that were otherwise inaccessible to the communities they were about.
When I started studying Owens Valley Paiute seriously, I came across a grammar that I never knew existed and found it locked behind UC San Diego's academic paywall. As a Ph.D. student at USC, I was able to get access to it, but it remains largely inaccessible to the rest of my community. While reading the first few pages, I was stunned to learn that Ida Stewart (my great-great-grandmother) was the source of most of the data in it. I had no idea that this resource existed until I was in graduate school, and none of the family members I talked to knew about it either. This was my experience coming into this work.
I learned, however, that the field has come a long way toward making documentation and preservation a much more community-centered practice. In fact, in every conversation I've observed or taken part in with documentation scholars and experts, "community-centered" is at the top of everyone's mind.
This is not my experience with AI research, which I think is still largely in an extractive phase of its lifespan (and hopefully it is just a phase). So the idea behind this workshop was to bring together people from very different backgrounds (Indigenous communities, language programs, linguistics, archival work, computer science, and AI) to surface shared concerns and move toward defining what responsible AI looks like in this space.
— Jared Coleman, organizer