For Indigenous Peoples of the Apya Yala (or Abya Yala), particularly in the Kara and Kichwa citizens of the Pan-Andean-Amazonian biocultural region, data is not merely a knowledge or information resource, it is the extension of Khipu Panaka (Indigenous data authority), treading the data lifecycle, genealogical and relational memory held within customary law and collective responsibility. This perspective paper presents the Kara-Kichwa Data Sovereignty Framework, a living legal-ethical instrument developed through autopoietic Indigenous storytelling, rights to story and place, and Indigenous-informed scope review to engage with external Indigenous data frameworks, counteracting intellectual gentrification and the systemic invisibility of Andean-Amazonian Indigenous Peoples within global digital transformation. The framework codifies five customary pillars, Kamachy (self-determination, community owns data about itself), Aylu-laktapak kamachy (collective authority and polygovernance), Tantanakuy (collective deliberation and relational accountability), Wilay-panka-tantay (physical custody of data and knowledge confidentiality), and Sumak kawsay (biocultural ethics and intergenerational
Indigenous peoples across Turtle Island (North America) face disproportionate rates of disappearance and murder, a "genocide" rooted in settler-colonial violence and systemic erasure. Technology plays a crucial role in the Missing and Murdered Indigenous Relatives (MMIR) crisis: perpetuating harm and impeding investigations, yet enabling advocacy and resistance. Communities utilize technologies such as AMBER alerts, news websites, social media groups, and campaigns (like #MMIW, #MMIWR, #NoMoreStolenSisters, and #NoMoreStolenDaughters) to mobilize searches, amplify awareness, and honor missing relatives. Yet, little research in HCI has critically examined technology's role in shaping the MMIR crisis by centering community voices. Through a large-scale study, we analyze 140 webpages to identify systemic, technological, and institutional barriers that hinder communities' efforts, while highlighting socio-technical actions that foster healing and safety. Finally, we amplify Indigenous voices by providing a dataset of stories that resist epistemic erasure, along with recommendations for HCI researchers to support Indigenous-led initiatives with cultural sensitivity, accountability, and
This paper focuses on the essential global issue of protecting and transmitting indigenous knowledge. It reveals the challenges in this area and proposes a sustainable supply chain framework for indigenous knowledge. The paper reviews existing technological solutions and identifies technical challenges and gaps. It then introduces cutting-edge technologies to protect and disseminate indigenous knowledge more effectively. The paper also discusses how the proposed framework can address real-world challenges in protecting and transmitting indigenous knowledge, and explores future research applications of the proposed solutions. Finally, it addresses open issues and provides a detailed analysis, offering promising research directions for the protection and transmission of indigenous knowledge worldwide.
In this paper, we examine the intersections of indigeneity and media representation in shaping perceptions of indigenous communities in Bangladesh. Using a mixed-methods approach, we combine quantitative analysis of media data with qualitative insights from focus group discussions (FGD). First, we identify a total of 4,893 indigenous-related articles from our initial dataset of 2.2 million newspaper articles, using a combination of keyword-based filtering and LLM, achieving 77% accuracy and an F1-score of 81.9\%. From manually inspecting 3 prominent Bangla newspapers, we identify 15 genres that we use as our topics for semi-supervised topic modeling using CorEx. Results show indigenous news articles have higher representation of culture and entertainment (19%, 10% higher than general news articles), and a disproportionate focus on conflict and protest (9%, 7% higher than general news). On the other hand, sentiment analysis reveals that 57% of articles on indigenous topics carry a negative tone, compared to 27% for non-indigenous related news. Drawing from communication studies, we further analyze framing, priming, and agenda-setting (frequency of themes) to support the case for dis
Indigenous narratives have long preserved observations of celestial phenomena, offering insights that resonate with modern astrophysical research. Deeply embedded in cultural traditions, these stories describe events such as stellar variability, supernovae, eclipses, and planetary alignments. Indigenous communities have also used the stars for navigation and calendar systems, reflecting a sophisticated understanding of celestial patterns. These narratives not only complement historical records but offer human-centric perspectives on stellar life cycles that often parallel modern models. Drawing from traditions from diverse Indigenous communities - including Aboriginal Australians, Pueblo peoples, Inuit, and Polynesians - this paper highlights the profound connections between Indigenous knowledge and stellar astrophysics. Integrating these traditions acknowledges their value, safeguards their relevance in the Space Age, and fosters mutual learning through cross-cultural collaboration. We explore how such narratives can enrich education, public engagement, and even inspire scientific hypotheses, bridging cultural heritage and modern astrophysics.
The paper focuses on the marginalization of indigenous language communities in the face of rapid technological advancements. We highlight the cultural richness of these languages and the risk they face of being overlooked in the realm of Natural Language Processing (NLP). We aim to bridge the gap between these communities and researchers, emphasizing the need for inclusive technological advancements that respect indigenous community perspectives. We show the NLP progress of indigenous Latin American languages and the survey that covers the status of indigenous languages in Latin America, their representation in NLP, and the challenges and innovations required for their preservation and development. The paper contributes to the current literature in understanding the need and progress of NLP for indigenous communities of Latin America, specifically low-resource and indigenous communities in general.
Across the Northern Hemisphere, Indigenous hunters developed arrows capable of skipping across the water surface to strike waterfowl. Archaeological and ethnographic records reveal remarkably similar projectile designs spanning millennia and geographically distant cultures, suggesting a convergent technological solution. Despite extensive study of water-entry dynamics, the physical principles underlying this behaviour remain poorly understood. Here we show that successful water-skipping arises from a small set of coupled geometric and dynamical parameters that define a bounded operational regime separating rebound, plunging, and overshoot. Using a combination of controlled experiments, hydrodynamic modeling, and historical reconstruction, we demonstrate that reconstructed arrow designs from independent cultures consistently fall within this predicted regime. These results demonstrate that Indigenous technologies were effectively tuned to satisfy the hydrodynamic constraints governing controlled skipping, providing evidence of convergent optimization in human-engineered systems. More broadly, our results suggest that material culture encodes physical knowledge that formal science is
Speech foundation models struggle with low-resource Pacific Indigenous languages because of severe data scarcity. Furthermore, full fine-tuning risks catastrophic forgetting. To address this gap, we present an empirical study adapting models to real-world Pacific datasets. We investigate the impact of data volume, adaptation strategies, and representational drift on speech foundation models for various Pacific languages. Additionally, we analyze a continual learning framework for sequential language acquisition. Empirical results across three distinct Pacific Indigenous languages demonstrate that adapting to these linguistically distant languages induces severe internal representational drift. Consequently, these models face a strict plasticity and stability dilemma. While LoRA adapts well initially, it suffers from catastrophic forgetting during sequential learning. Ultimately, this study highlights the urgent need for robust adaptation strategies tailored to underrepresented languages.
As part of the mission of the International Astronomical Union Centre for the Protection of the Dark and Quiet Sky from Satellite Constellation Interference (IAU-CPS) Policy Hub to consider national and international regulations about the usage and sustainability in outer space, we also included discussion specific to the rights of Indigenous peoples with respect to outer space under the context of the United Nations Declaration for the Rights of Indigenous Peoples (UNDRIP). In this work, we review how some of the articles of UNDRIP require various actors in the use and exploitation of outer space including satellite companies, nation states, and professional/academic astronomy to consult and support Indigenous peoples/nations and respect Indigenous sovereignties. This work is concluded with recommendations for consulting and collaborating with Indigenous peoples and recommendations for moving from the traditional colonial exploitation of outer space and building an anti-colonial future in relationship with outer space.
Low-resource indigenous languages often lack the parallel corpora required for effective neural machine translation (NMT). Synthetic data generation offers a practical strategy for mitigating this limitation in data-scarce settings. In this work, we augment curated parallel datasets for indigenous languages of the Americas with synthetic sentence pairs generated using a high-capacity multilingual translation model. We fine-tune a multilingual mBART model on curated-only and synthetically augmented data and evaluate translation quality using chrF++, the primary metric used in recent AmericasNLP shared tasks for agglutinative languages. We further apply language-specific preprocessing, including orthographic normalization and noise-aware filtering, to reduce corpus artifacts. Experiments on Guarani-Spanish and Quechua-Spanish translation show consistent chrF++ improvements from synthetic data augmentation, while diagnostic experiments on Aymara highlight the limitations of generic preprocessing for highly agglutinative languages.
Since 2022 we have been exploring application areas and technologies in which Artificial Intelligence (AI) and modern Natural Language Processing (NLP), such as Large Language Models (LLMs), can be employed to foster the usage and facilitate the documentation of Indigenous languages which are in danger of disappearing. We start by discussing the decreasing diversity of languages in the world and how working with Indigenous languages poses unique ethical challenges for AI and NLP. To address those challenges, we propose an alternative development AI cycle based on community engagement and usage. Then, we report encouraging results in the development of high-quality machine learning translators for Indigenous languages by fine-tuning state-of-the-art (SOTA) translators with tiny amounts of data and discuss how to avoid some common pitfalls in the process. We also present prototypes we have built in projects done in 2023 and 2024 with Indigenous communities in Brazil, aimed at facilitating writing, and discuss the development of Indigenous Language Models (ILMs) as a replicable and scalable way to create spell-checkers, next-word predictors, and similar tools. Finally, we discuss how
The task of recognizing the age-separated faces of an individual, Age-Invariant Face Recognition (AIFR), has received considerable research efforts in Europe, America, and Asia, compared to Africa. Thus, AIFR research efforts have often under-represented/misrepresented the African ethnicity with non-indigenous Africans. This work developed an AIFR system for indigenous African faces to reduce the misrepresentation of African ethnicity in facial image analysis research. We adopted a pre-trained deep learning model (VGGFace) for AIFR on a dataset of 5,000 indigenous African faces (FAGE\_v2) collected for this study. FAGE\_v2 was curated via Internet image searches of 500 individuals evenly distributed across 10 African countries. VGGFace was trained on FAGE\_v2 to obtain the best accuracy of 81.80\%. We also performed experiments on an African-American subset of the CACD dataset and obtained the best accuracy of 91.5\%. The results show a significant difference in the recognition accuracies of indigenous versus non-indigenous Africans.
In this work, the authors describe efforts aimed at Indigenizing a second-year linear algebra course at a small liberal arts university in Manitoba, Canada. This is done through an assignment, part hands-on and part written work, that explores the connection between Indigenous beadwork and linear algebra. Our collaboration was perhaps unconventional: Sarah, the first author, is a mathematics professor; while Cathy, the second author, is an associate professor in art history. However, we both had similar goals of putting theory into practice and making positive changes to student learning outcomes in a culturally appropriate way. We situate our work in the context of the current scholarly literature, adding to the important ongoing dialogue on Indigenization of course content and reflecting on the process and outcomes. This transformation of the course curriculum represented an applied approach to immerse Indigenous knowledge and pedagogy into a mathematics classroom. We hope that it may serve as an example of how other educators, particularly in science, technology, engineering, and mathematics (STEM), can integrate Indigenous knowledge-centered pedagogy into their classroom.
Data mining reproduces colonialism, and Indigenous voices are being left out of the development of technology that relies on data, such as artificial intelligence. This research stresses the need for the inclusion of Indigenous Data Sovereignty and centers on the importance of Indigenous rights over their own data. Inclusion is necessary in order to integrate Indigenous knowledge into the design, development, and implementation of data-reliant technology. To support this hypothesis and address the problem, the CARE Principles for Indigenous Data Governance (Collective Benefit, Authority to Control, Responsibility, and Ethics) are applied. We cover how the colonial practices of data mining do not align with Indigenous convictions. The included case studies highlight connections to Indigenous rights in relation to the protection of data and environmental ecosystems, thus establishing how data governance can serve both the people and the Earth. By applying the CARE Principles to the issues that arise from data mining and neocolonialism, our goal is to provide a framework that can be used in technological development. The theory is that this could reflect outwards to promote data sover
Argentina has a large yet little-known Indigenous linguistic diversity, encompassing at least 40 different languages. The majority of these languages are at risk of disappearing, resulting in a significant loss of world heritage and cultural knowledge. Currently, unified information on speakers and computational tools is lacking for these languages. In this work, we present a systematization of the Indigenous languages spoken in Argentina, classifying them into seven language families: Mapuche, Tupí-Guaraní, Guaycurú, Quechua, Mataco-Mataguaya, Aymara, and Chon. For each one, we present an estimation of the national Indigenous population size, based on the most recent Argentinian census. We discuss potential reasons why the census questionnaire design may underestimate the actual number of speakers. We also provide a concise survey of computational resources available for these languages, whether or not they were specifically developed for Argentinian varieties.
In this paper, we offer an overview of indigenous languages, identifying the causes of their devaluation and the need for legislation on language rights. We review the technologies used to revitalize these languages, finding that when they come from outside, they often have the opposite effect to what they seek; however, when developed from within communities, they become powerful instruments of expression. We propose that the inclusion of Indigenous knowledge in large language models (LLMs) will enrich the technological landscape, but must be done in a participatory environment that encourages the exchange of knowledge.
In this position paper, we first discuss the uptake of speculative design as a method for Indigenous HCI. Then, we outline how a key assumption about temporality threatens to undermine the usefulness of speculative design in this context. Finally, we briefly sketch out a possible alternative understanding of speculative design, based on the concept of "spiraling time," which could be better suited for Indigenous HCI.
This paper presents the winning submission of the RaaVa team to the AmericasNLP 2025 Shared Task 3 on Automatic Evaluation Metrics for Machine Translation (MT) into Indigenous Languages of America, where our system ranked first overall based on average Pearson correlation with the human annotations. We introduce Feature-Union Scorer (FUSE) for Evaluation, FUSE integrates Ridge regression and Gradient Boosting to model translation quality. In addition to FUSE, we explore five alternative approaches leveraging different combinations of linguistic similarity features and learning paradigms. FUSE Score highlights the effectiveness of combining lexical, phonetic, semantic, and fuzzy token similarity with learning-based modeling to improve MT evaluation for morphologically rich and low-resource languages. MT into Indigenous languages poses unique challenges due to polysynthesis, complex morphology, and non-standardized orthography. Conventional automatic metrics such as BLEU, TER, and ChrF often fail to capture deeper aspects like semantic adequacy and fluency. Our proposed framework, formerly referred to as FUSE, incorporates multilingual sentence embeddings and phonological encodings t
Indigenous languages are historically under-served by Natural Language Processing (NLP) technologies, but this is changing for some languages with the recent scaling of large multilingual models and an increased focus by the NLP community on endangered languages. This position paper explores ethical considerations in building NLP technologies for Indigenous languages, based on the premise that such projects should primarily serve Indigenous communities. We report on interviews with 17 researchers working in or with Aboriginal and/or Torres Strait Islander communities on language technology projects in Australia. Drawing on insights from the interviews, we recommend practices for NLP researchers to increase attention to the process of engagements with Indigenous communities, rather than focusing only on decontextualised artefacts.
Commercial endeavours have already compromised our relationship with space. The Artemis Accords are creating a framework that will commercialize the Moon and further impact that relation. To confront that impact, a number of organizations have begun to develop new principles of sustainability in space, many of which are borne out of the capitalist and colonial frameworks that have harmed water, nature, peoples and more on Earth. Indigenous methodologies and ways of knowing offer different paths for living in relationship with space and the Moon. While Indigenous knowledges are not homogeneous, there are lessons we can use from some of common methods. In this talk we will review some Indigenous methodologies, including the concept of kinship and discuss how kinship can inform our actions both on Earth and in space.