The First Sunaŵi Workshop

AI & Indigenous Language Revitalization Workshop

June 8–10, 2026; Loyola Marymount University, Los Angeles

What it is

Sunaŵi brought together people from very different backgrounds (Indigenous community members, computer scientists, AI experts, linguists, archivists, language teachers, nonprofit leaders, education experts, students, industry representatives, and junior and senior faculty) to surface shared concerns and move toward defining what responsible AI looks like in the context of Indigenous language revitalization. It went a long way toward breaking down some of the barriers that prevent us all from working more closely together.

This wiki is the workshop's living output. See the people who took part.

Program

Three days of talks, panels, and small-group work. Talks and a tutorial came from researchers, practitioners, and community members sharing their own experience. Panels put those perspectives in direct conversation with each other. The breakouts were where the room worked in small groups toward something concrete: the first mapped the concerns people brought with them, the second drafted recommendations, and the final morning was given to presenting that work back to the room. This wiki grew out of that material.

Each day ran roughly 9:00 to 5:00, with talks and panels in the mornings and breakouts in the afternoons. Times below are from the program as scheduled.

Day 1 — Monday, June 8

TimeSession
9:00Opening remarksJared Coleman
9:15A view on AI & Indigenous Languages from the Advancing Indigenous Language Technologies Working GroupAmy Fountain
9:45Whose Data, Whose Voice? Exploring Indigenous Sovereignty from a Choctaw PerspectiveJacqueline Brixey
10:30Tutorial: LLM-Assisted Rule-Based Machine TranslationJared Coleman
11:00Beyond Data Scarcity: Challenges and Opportunities for Indigenous Language NLPEmily Prud'hommeaux
11:30Panel 1: Concerns & HopesHelema Andrews, Jacqueline Brixey, Citlali Arvizu, Jay Prakash Thakur, Kiana Maillet
2:00Breakout 1: Concern mapping
4:00Plenary cross-pollination

Day 2 — Tuesday, June 9

TimeSession
9:00Manuscripts, Microfiche, Typewriters, Databases and LLMs: Language and Technology at the Myaamia CenterHunter Lockwood
9:30AI for Endangered Language Documentation and PreservationAntonios Anastasopoulos
10:00Agentic AI Systems for Developing Language Education ToolsKhalil Iskarous
11:00The AmericasNLP Shared Tasks: Results, Artifacts, and Future DirectionsAbteen Ebrahimi
11:30Between the Academy and the Community: Are Our Contributions Doing What We Think They Are?Belu Ticona
12:00Panel 2: Community & Capacity BuildingKristine Hildebrandt, Ben Ankiel, Melissa Floca, Daniel Bögre Udell
2:30Breakout 2: Drafting recommendations
4:30Plenary cross-pollination

Day 3 — Wednesday, June 10

TimeSession
9:00Final breakout presentations
10:45Panel 3: Future of the FieldRaina Heaton, Khalil Iskarous, Luis Barragan, Tiffany Wright
11:45Closing remarks

Talk abstracts

In program order. The tutorial had no abstract.

A view on AI & Indigenous Languages from the Advancing Indigenous Language Technologies Working Group

Amy Fountain

This talk will focus on issues and opportunities arising from language work with communities here in the southwest, especially focusing on language policy and data sovereignty. We will discuss language ideology, policy and practical matters.

Whose Data, Whose Voice? Exploring Indigenous Sovereignty from a Choctaw Perspective

Jacqueline Brixey

Data sovereignty is a topic that impacts all people in this era of big data and artificial intelligence. In short, it is the ability of a country to control and access the data that is generated in its territory. In this talk, I will explore what data sovereignty means from an Indigenous perspective, specifically from my viewpoint as a citizen of the Choctaw Nation of Oklahoma. I will show examples of policies that CNO has enacted. Additionally, I will present a framework for evaluating commercially available Large Language Models (LLMs) and explain how their unauthorized use of Indigenous data is undermining Indigenous language revitalization efforts.

Beyond Data Scarcity: Challenges and Opportunities for Indigenous Language NLP

Emily Prud'hommeaux

The majority of the world's languages are underserved by both big tech companies and the NLP research community. Although the marginalization of these languages is driven to some extent by a general lack of training resources, the quality and content of the data that is available, rather than the overall quantity of data, play an important role. In this talk, I will discuss the impact of transcription quality, domain mismatch, and data modality on NLP task performance in endangered, indigenous, and under-resourced languages.

Manuscripts, Microfiche, Typewriters, Databases and LLMs: Language and Technology at the Myaamia Center

Hunter Lockwood

Miami-Illinois is an Algonquian language spoken in the southern Great Lakes region. Documentation of the language dates back to the late 17th century, and continued into the mid- to late-20th century with the death of the last generation of semispeakers and rememberers. Since the 1980s, there has been a concerted effort to revitalize the language using archival sources. Reclamation efforts originally relied on photocopies and microfiche of handwritten and typewritten notes, but serious digitization work began in the early 2000s. This work was supercharged after the development of ILDA (Indigenous Languages Digital Archive), allowing us to store, analyze, and share data in a more robust way, on a platform that was designed with the needs of Indigenous language communities in mind. Since 2022, as technology continues to evolve, we have explored the use of LLMs/AI, with a focus on practical applications for language documentation and revitalization. In this talk, I will situate the Myaamia Center's investigation of LLMs within our broader history of archives-based language revitalization and use of technology; discuss our current efforts to explore the potential of LLMs; and reflect on the ethical, privacy, and sovereignty considerations that guide our approach.

AI for Endangered Language Documentation and Preservation

Antonios Anastasopoulos

In this talk I will cover three case studies from my lab's work, that center on projects around building language technologies for underserved languages in three different contexts in different continents. In particular, I will cover how we built educational tools for Mapuzugun in Chile, how we collected data to build the first translation and speech recognition system for Bemba in Zambia, and how we employ AI-human feedback loops to build resources for under-resourced endangered Greek varieties like Griko.

Agentic AI Systems for Developing Language Education Tools

Khalil Iskarous

Communities often need fun and engaging learning tools to help their children learn their endangered language, when most engaging materials are in some other dominant language. This talk will describe work of this sort performed with member of the Ladin community in Italy 8 years ago, and how modern agentic AI systems could transform the development of such tools.

The AmericasNLP Shared Tasks: Results, Artifacts, and Future Directions

Abteen Ebrahimi

Natural language processing tools have great potential in aiding language documentation and revitalization efforts, however, for many languages the quality of these systems is far below the minimum threshold needed for real-world effectiveness. In this talk I will present the AmericasNLP Shared Tasks, hosted as part of the AmericasNLP workshop since its inception in 2021, and how these competitions can provide, in addition to pure research, an additional avenue towards addressing the weaknesses of current language technologies available for data-scarce languages.

Between the Academy and the Community: Are Our Contributions Doing What We Think They Are?

Belu Ticona

How do academic language technology initiatives actually impact the people they claim to serve? From rural speakers to urban descendants, the reality of who those people are varies widely. In this talk, Belu draws on her indigenous roots and a South American perspective to reflect on how researchers engage with "communities," and whether that engagement aligns with a researcher's positionality, particularly when that researcher shares roots with the people being studied.

How it came to be

It goes all the way back to why I chose computer science as a major in the first place: I wanted to create a "Rosetta Stone" (the language-learning software) for my tribe's critically endangered language, Owens Valley Paiute. My research ended up going in a very different direction (distributed computing), but I never lost sight of that original motivation. I helped create online tools (like an online version of Glenn Nelson's Owens Valley Paiute dictionary) to help make the language more accessible to my community, but it wasn't until ChatGPT exploded in 2022 that I saw, for the first time, an opportunity to do research in this area as well.

I started working with linguists, archivists, and other scholars in language documentation, preservation, and revitalization. Until then, my only real exposure to the field was through my own tribe's language program, and most of that was interacting with historical materials. As many of the workshop participants can attest, the traditional language documentation field was quite extractive. Linguists would come into communities, gather data from speakers, produce potentially valuable resources (dictionaries, grammars, etc.), and then publish them behind academic paywalls or in ways that were otherwise inaccessible to the communities they were about.

When I started studying Owens Valley Paiute seriously, I came across a grammar that I never knew existed and found it locked behind UC San Diego's academic paywall. As a Ph.D. student at USC, I was able to get access to it, but it remains largely inaccessible to the rest of my community. While reading the first few pages, I was stunned to learn that Ida Stewart (my great-great-grandmother) was the source of most of the data in it. I had no idea that this resource existed until I was in graduate school, and none of the family members I talked to knew about it either. This was my experience coming into this work.

I learned, however, that the field has come a long way toward making documentation and preservation a much more community-centered practice. In fact, in every conversation I've observed or taken part in with documentation scholars and experts, "community-centered" is at the top of everyone's mind.

This is not my experience with AI research, which I think is still largely in an extractive phase of its lifespan (and hopefully it is just a phase). So the idea behind this workshop was to bring together people from very different backgrounds (Indigenous communities, language programs, linguistics, archival work, computer science, and AI) to surface shared concerns and move toward defining what responsible AI looks like in this space.

— Jared Coleman, organizer

Created · Updated
Supported By the National Science Foundation Award 2542375.