The Indo-European Cognate Relationships (IE-CoR) dataset is an open-access relational dataset showing how related, inherited words ('cognates') pattern across 160 languages of the Indo-European family. IE-CoR is intended as a benchmark dataset for computational research into the evolution of the Indo-European languages. It is structured around 170 reference meanings in core lexicon, and contains 25731 lexeme entries, analysed into 4981 cognate sets. Novel, dedicated structures are used to code all known cases of horizontal transfer. All 13 main documented clades of Indo-European, and their main subclades, are well represented. Time calibration data for each language are also included, as are relevant geographical and social metadata. Data collection was performed by an expert consortium of 89 linguists drawing on 355 cited sources. The dataset is extendable to further languages and meanings and follows the Cross-Linguistic Data Format (CLDF) protocols for linguistic data. It is designed to be interoperable with other cross-linguistic datasets and catalogues, and provides a reference framework for similar initiatives for other language families.
暂无摘要(点击查看详情)
暂无摘要(点击查看详情)
The Indo-European languages are among the most widely spoken in the world, yet their early diversification remains contentious1-5. It is widely accepted that the spread of this language family across Europe from the 5th millennium BP correlates with the expansion and diversification of steppe-related genetic ancestry from the onset of the Bronze Age6,7. However, multiple steppe-derived populations co-existed in Europe during this period, and it remains unclear how these populations diverged and which provided the demographic channels for the ancestral forms of the Italic, Celtic, Greek, and Armenian languages8,9. To investigate the ancestral histories of Indo-European-speaking groups in Southern Europe, we sequenced genomes from 314 ancient individuals from the Mediterranean and surrounding regions, spanning from 5,200 BP to 2,100 BP, and co-analysed these with published genome data. We additionally conducted strontium isotope analyses on 224 of these individuals. We find a deep east-west divide of steppe ancestry in Southern Europe during the Bronze Age. Specifically, we show that the arrival of steppe ancestry in Spain, France, and Italy was mediated by Bell Beaker (BB) populations of Western Europe, likely contributing to the emergence of the Italic and Celtic languages. In contrast, Armenian and Greek populations acquired steppe ancestry directly from Yamnaya groups of Eastern Europe. These results are consistent with the linguistic Italo-Celtic10,11 and Graeco-Armenian1,12,13 hypotheses accounting for the origins of most Mediterranean Indo-European languages of Classical Antiquity. Our findings thus align with specific linguistic divergence models for the Indo-European language family while contradicting others. This underlines the power of ancient DNA in uncovering prehistoric diversifications of human populations and language communities.
The aim of this study was to investigate the relationship among Lithuanian, Latvian, Indian, and some other populations through a genome-wide data analysis of single nucleotide polymorphisms (SNPs). Limited data of Baltic populations were mostly compared with geographically closer modern and ancient populations in the past, but no previous investigation has explored their genetic relationships with distant populations, like the ones of India, in detail. To address this, we collected and merged genome-wide SNP data from diverse publicly available sources to create a comprehensive dataset with a substantial sample size especially from Lithuanians and Latvians. Principal component analysis (PCA) and admixture analysis methods were employed to assess the genetic structure and relationship among the populations under investigation. Additionally, we estimated an effective population size (Ne) and divergence time to shed light on potential past events between the Baltic and Indian populations. To gain a broader perspective, we also incorporated ancient and modern populations from different continents into our analyses. Our findings revealed that the Balts, unsurprisingly, have a closer genetic affinity with individuals from Indian population who speak Indo-European languages, compared to other Indian linguistic groups (such as speakers of Dravidian, Austroasiatic, and Sino-Tibetan languages). However, when compared to other populations from the European continent, which also speak Indo-European and some Uralic languages, the Balts did not exhibit a stronger resemblance to Indo-European-speaking Indians. In conclusion, this study provides an overview of the genetic relationship and structure of the populations investigated, along with insights into their divergence times.
The Yamnaya archaeological complex appeared around 3300 BC across the steppes north of the Black and Caspian Seas, and by 3000 BC it reached its maximal extent, ranging from Hungary in the west to Kazakhstan in the east. To localize Yamnaya origins among the preceding Eneolithic people, we assembled ancient DNA from 435 individuals, demonstrating three genetic clines. A Caucasus-lower Volga (CLV) cline suffused with Caucasus hunter-gatherer1 ancestry extended between a Caucasus Neolithic southern end and a northern end at Berezhnovka along the lower Volga river. Bidirectional gene flow created intermediate populations, such as the north Caucasus Maikop people, and those at Remontnoye on the steppe. The Volga cline was formed as CLV people mixed with upriver populations of Eastern hunter-gatherer2 ancestry, creating hypervariable groups, including one at Khvalynsk. The Dnipro cline was formed when CLV people moved west, mixing with people with Ukraine Neolithic hunter-gatherer ancestry3 along the Dnipro and Don rivers to establish Serednii Stih groups, from whom Yamnaya ancestors formed around 4000 BC and grew rapidly after 3750-3350 BC. The CLV people contributed around four-fifths of the ancestry of the Yamnaya and, entering Anatolia, probably from the east, at least one-tenth of the ancestry of Bronze Age central Anatolians, who spoke Hittite4,5. We therefore propose that the final unity of the speakers of 'proto-Indo-Anatolian', the language ancestral to both Anatolian and Indo-European people, occurred in CLV people some time between 4400 BC and 4000 BC.
暂无摘要(点击查看详情)
This paper exhorts communication specialists to look beyond English language knowledge by providing evidence to disrupt the unsubstantiated belief that there are few assessment and intervention resources for supporting multilingual children's speech. The Multilingual Children's Speech website https://www.csu.edu.au/research/multilingual-speech/home has curated 1337 (mostly free) resources for supporting multilingual children's speech acquisition, assessment, and intervention in 131 of the world's languages and dialects (86 languages). Specifically, there are 658 speech acquisition studies in 55 languages, 423 speech assessment resources in 77 languages, and 178 speech intervention resources in 21 languages. This free website includes links to assessment tools, intervention manuals, journal articles, books, chapters, theses, and video recordings for 16 of the top 20 most spoken languages in the world and many minority languages, Indigenous languages (e.g. Māori, Samoan, Sesotho, Setswana, Warlpiri, isiXhosa, Zapotec, isiZulu) and languages and dialects impacted by colonisation and slavery (e.g. African American English, Fiji English, Jamaican Creole, Tok Pisin). Only 17.95% of the resources are about English, with 51.68% about 39 other Indo-European languages, and 30.37% about 46 languages belonging to 15 non-Indo-European language families. Previous analyses of curated knowledge about children's development in psychology and linguistics have found a WEIRD bias 'Western, Educated, Industrialized, Rich, and Democratic societies'; however, only 29.07% of the languages included on the Multilingual Children's Speech website are WEIRD. While only 1.23% of the 7000 world languages are represented on the website, these assessment and intervention resources will continue to grow due to ongoing work of multilingual communication specialists across the globe.
The decline of polygyny and the rise of monogamy among complex, stratified societies characterized by high wealth inequality present a long-standing puzzle in anthropology. Competing explanations suggest that monogamy either i) reduces reproductive inequality, fosters cooperation, and enhances success in intergroup competition, or ii) mitigates conflict over heritable wealth-especially land-under conditions of high social stratification and ecological constraints. Both frameworks are influenced by the history of Indo-European societies, where monogamy has long been normative and closely associated with land inheritance and state formation. However, normative monogamy is also found in many societies beyond this historical context. To evaluate these competing hypotheses, we formalized their causal structure and applied Bayesian phylogenetic multilevel models to a global sample of 186 societies. Our results show that monogamy is strongly associated with land privatization and, in some regions, with ecological or demographic proxies for land scarcity-supporting the view that competition over heritable wealth promotes monogamy globally. In contrast, support for the intergroup competition model is inconsistent: While it may explain monogamy in some language families, these dynamics do not extend to most societies in our sample. Our findings suggest that monogamy arose repeatedly under similar socioecological conditions and cannot be fully explained by theories based primarily on Indo-European history.
ObjectivesSpeech characteristics of children with cleft palate (CP) have primarily been studied in Indo-European languages. Vietnamese has a unique phonological system that differs from English and other Indo-European languages. The speech characteristics of Vietnamese children with CP, however, have not been thoroughly explored. This study aimed to investigate the speech and resonance characteristics of Vietnamese children with repaired CP who attended a speech evaluation clinic in Hanoi, Vietnam.MethodSeventy-two monolingual Vietnamese children with repaired CP, aged 3 to 12 years, participated in the study. Resonance, cleft-related errors, phonological errors, and speech production accuracy were examined. In particular, cleft-related errors and phonological errors were analyzed separately for word-initial and word-final positions. A database of 63 typically developing Vietnamese children was used to compare speech production accuracy.ResultsBoth language-common and language-specific characteristics were identified. Most Vietnamese children with repaired CP in this study demonstrated persistent abnormal resonance and lower articulation skills than their typically developing peers. Nasal air emission and cleft-related speech errors primarily occurred in the word-initial position, whereas the most frequent type of errors in the word-final position were developmental errors. This language-specific characteristic was likely due to the Vietnamese phonotactic constraints.ConclusionsLanguage-common and language-specific speech characteristics in Vietnamese children with CP may improve our understanding of CP speech. These error patterns are especially beneficial for Vietnamese speech therapists worldwide when assessing and treating children with CP who speak Vietnamese.
Quantitative phylogenetics in historical linguistics has relied almost entirely on lexical cognate data. This study asks a different question: how much genealogical signal can be recovered from structural features extracted from annotated corpora, and whether it survives at deep time depths. We compute 29 structural features-including Shannon entropies of dependency direction and of dependency-relation distributions, relation-specific directionality ratios, dependency-distance measures, and constructional ratios-across 25 Transeurasian languages from the five proposed groups (Turkic, Mongolic, Tungusic, Japonic, and Koreanic) and three outgroups (Chinese, Vietnamese, and Hindi), 28 languages in all. Most of the Tungusic and Mongolic languages have no running-text corpus, so we built new Universal Dependencies treebanks for them by glossing example sentences from reference grammars; thirteen are used here. Each feature was tested for phylogenetic signal (Pagel's λ and Blomberg's K, with FDR correction) under four competing reference topologies, and the features that passed were used for tree inference (Bayesian inference in MrBayes, with Neighbor-Joining as a check). The same pipeline was first run on Indo-European in a companion study, where it recovers only individual subgroups and does not resolve a stable tree. At the depth proposed for the Transeurasian family (a Proto-Transeurasian root of about 9000 years before present), the structural signal was not enough to reconstruct the family's internal relationships. The signal tests favoured a flat three-way division of the major branches (7 strict/20 relaxed features) over any nested hypothesis (≤2 strict features each), and the strongest signal lay in core word-order parameters (e.g., object direction, λ = 1.00, K = 6.06). But both Bayesian and distance-based inference returned near-complete polytomies: although the chains converged (ASDSF < 0.01), no branch reached a posterior probability above 0.75, and none of the three multi-language branches (Turkic, Mongolic, or Tungusic) was recovered. The outgroup test made the reason clear: Hindi, which is Indo-European but SOV, grouped with the head-final Transeurasian languages rather than with the other two (head-initial) outgroups, so the features are tracking typological similarity, not shared descent, at this depth. The study contributes 13 new treebanks for poorly documented languages, a reproducible framework for testing how much genealogical signal structural features carry, and direct evidence that, at Transeurasian time depths, this signal reflects typology rather than genealogy.
A critical gap remains in understanding how the architectural philosophies of different large language models (LLMs) perform in specialized clinical domains conducted in non-Indo-European languages, particularly regarding the effectiveness of Chain-of-Thought (CoT) reasoning models compared with those optimized for local languages. The purpose of this study was to compare the performance of 6 LLMs with distinct architectural philosophies (CoT-reasoning, general-purpose, and Korean-optimized) on the prosthodontics section of the Korean Dental Licensing Examination (KDLE) and to contextualize their performance with published human averages. A total of 161 Korean-language, text-only multiple-choice questions from the prosthodontics section of the KDLE (2020-2024) were presented to 6 LLMs (ChatGPT-o1, ChatGPT-4o, DeepSeek-R1, DeepSeek-V3, Gemini 1.5 Flash, and CLOVA X). Each test set was posed 6 times. The questions were further classified into 5 domains: diagnosis and treatment planning, mandibular movements and occlusal relationships, removable complete denture, removable partial denture, and fixed prosthodontics. Performance was measured by percentage accuracy and analyzed using the Cochran Q and post hoc McNemar tests (α=.05). LLM scores were contextually benchmarked against the average performance of human examinees. Significant performance differences were observed among the models (P<.001). The CoT-based model, ChatGPT-o1, achieved the highest overall accuracy (80.54%); the total human average (79.51%) fell within this LLM's 95% confidence interval. ChatGPT-4o (71.84%) and DeepSeek-R1 (70.19%) also demonstrated consistent passing-level performance. The Korean-language-optimized model, CLOVA X, obtained the lowest score (34.37%). The performance ranking of the models within individual domains generally mirrored the overall ranking. LLMs with CoT-reasoning architectures can achieve passing-level accuracy on non-English dental licensing examinations at a level contextually comparable to the human average, but performance varied significantly by architecture, and localized language optimization did not ensure domain expertise.
The history of the Albanian people has long been debated, as they first appear in historical records in the eleventh century CE and their language is not closely related to any surviving Indo-European branches. Here, to reconstruct their history, we analysed over 6,000 ancient West Eurasian genomes and 74 newly sequenced present-day ethnic Albanians. Using a range of population genetics methods, including an enhanced protocol to detect identity-by-descent segments between ancient and present-day individuals, we detect continuity of West Balkan Late Bronze and Iron Age ancestry in Early Medieval Albania, to a greater degree than in neighbouring Balkan regions. We find that present-day Albanians predominantly descend from this remnant palaeo-Balkan group, which by at least 800-900 CE already exhibited a genetic profile suggesting that they are ancestral to many modern Albanians. In addition, we observe geographically structured admixture with Medieval East European-related groups, averaging 10-20% across present-day Albanians. Our findings provide insight into the demographic processes shaping Albanian ancestry and help locate the origin area of the Albanian language.
When describing faces, people often struggle with verbalizing facial features. Free descriptions seem to focus predominantly on aspects of faces that are inferred, for example, psychological traits, age, attractiveness, and so on, whereas facial features themselves are often described in a limited and imprecise fashion. However, existing research relies heavily on large, industrialized societies of the West, so it is unclear if the peculiar characteristics of facial language are universal or limited to a narrow set of languages. As part of the ongoing effort to diversify cognitive sciences, we investigate this under-researched area through a verbal description task targeting a selection of 51 facial features in a small-scale nonindustrialized group of Maniq speakers (Austroasiatic) and an industrialized group of Polish speakers (Indo-European). Our results suggest facial appearance is poorly coded in both Maniq and Polish. In addition, while the major types of physical facial descriptors used across both languages are similar, Maniq speakers show a more uniform focus on directly observable aspects of faces, whereas Polish speakers are additionally inclined to infer socially relevant meaning from faces, even when not prompted to do so. All in all, although facial appearance is generally difficult to put into words, culture-specific factors shape the language of faces in distinct ways.
Down syndrome (DS) is associated with persistent language impairments that extend beyond early childhood, yet evidence from agglutinative languages remains limited. While morphosyntactic weaknesses have been well-documented in Indo-European languages, less is known about how such difficulties are manifested in Turkish, a language in which grammatical relations are primarily marked through morphology. In addition, short-term memory (STM) limitations, particularly in verbal domains, are characteristic of DS and may contribute to language outcomes. This study examined the interaction between morphosyntax and STM in Turkish-speaking children and adolescents with DS. A cross-sectional observational design was employed, including 12 monolingual Turkish-speaking participants with DS (aged 6;7-15;11) and 10 TD peers matched on nonverbal mental age. Participants completed standardized assessments of syntax and morphology, spontaneous language sampling, and STM tasks assessing verbal and visual memory. Children with DS performed significantly below controls on syntactic comprehension and production as well as morphological measures, with larger effects observed for syntax. Noun morphology was less accurate than verb morphology, likely reflecting increased morphophonological complexity. Regression analyses indicated that auditory digit span predicted sentence comprehension, whereas nonword repetition predicted morphological production indexed by mean length of utterance in morphemes. Substantial inter-individual variability was observed within the DS group. These findings suggest that morphosyntactic outcomes in Turkish-speaking children with DS are closely linked to verbal STM capacities and vary considerably across individuals, underscoring the importance of integrated assessment and individualized intervention planning. Future research with larger samples is warranted to confirm and extend these preliminary findings. Findings should be interpreted cautiously due to the limited sample size and are presented as preliminary descriptive evidence. This study provides initial data on Turkish-speaking individuals with Down syndrome.
Long-distance migrations across Eurasia have led to extensive admixture among ethnolinguistically diverse Altaic speakers and their Indo-European neighbors; however, the population histories of the Xibe and Daur in northern China remain poorly understood. We address this gap by analyzing genome-wide SNP data from 324 individuals to characterize genetic structure, admixture, and migration. Both groups cluster with East Asian populations, particularly Altaic-speaking groups, consistent with a shared origin of the geographically distinct Xibe and Daur populations. We identify divergent post-migration histories, in which northwestern Chinese Xibe show distinct profiles and stronger genetic ties to the Daur, while northeastern Xibe have substantial contributions from indigenous ancient Northeast Asian (ANA) and Sino-Tibetan speakers. Admixture modeling indicates that Daur and the majority of Xibe populations are best described as two-way mixtures with Han and ANA groups; a distinct West Eurasian/Steppe signal persists in specific Shandong subgroups, reflecting historical trans-Eurasian cultural interactions and gene flow. Notably, Xinjiang Xibe share a high-ANA genetic profile, whereas other Xibe groups possess additional Han-related ancestry. Shared-segment analyses support a northeastern origin for the Xinjiang Xibe, where subsequent westward migration served as a complementary event that preserved this ancestral diversity, whereas the Daur retains strong ties to ANA, especially those from Inner Mongolia. Together, these results clarify the layered demographic histories of Xibe and Daur, refine models of population structure in northern East Asia, and illustrate how repeated movements and regional interactions generate present-day genomic diversity at the crossroads between North China and Siberia.
Estimating evolutionary rates and divergence times for hepatitis B virus (HBV) has long been complicated by conflicting calibration approaches and extensive rate variation. To unlock the full potential of ancient and modern HBV genomic data, we develop a Bayesian mixed-effects molecular clock model that accounts for various sources of rate variation including time-dependent rate decay. Our analyses reveal a pronounced decline in evolutionary rates over time, reconciling HBV divergence estimates with human migration events across both deep and more recent timescales. We show that HBV spread into Europe through both Neolithic farming expansions and later steppe migrations, paralleling patterns proposed for Indo-European language origins. Phylogeographic reconstructions suggest that the Neolithic-associated lineage dispersed at approximately 1 km/year, consistent with archaeological estimates, while genotype D expanded during the Bronze Age at an almost threefold higher rate, plausibly driven by technological innovations underlying steppe expansions. Historical overlap between these lineages facilitated recombination, giving rise to genotype E, which has become a dominant HBV genotype in Africa. These findings demonstrate that ancient viral genomes, when analyzed with models capturing complex rate dynamics, provide a powerful lens on human prehistory and the processes shaping pathogen diversity.
Human language exhibits both universal characteristics and remarkable diversity. While neuroimaging has identified a shared fronto-temporo-parietal network across languages, causal inferences require techniques such as direct electrical stimulation (DES), which is performed during awake craniotomy. However, it is unclear which native languages have been mapped using DES. We conducted an analysis of which native languages were mapped in monolingual patients undergoing DES during awake craniotomy. Data was extracted from an open-access systematic review dataset of language errors from patients during awake craniotomy. An additional literature review and author communication were performed to identify the language type. Descriptive statistics and chi-squared analysis were used. The analysis identified 893 stimulation sites across 73 studies, involving nine languages: French, English, Chinese, Japanese, Italian, Dutch, Spanish, German, and Persian. These languages represent 27.2% of native languages spoken globally. The predominant mapped language family was Indo-European (670 stimulation sites; 75%), followed by Sino-Tibetan (186 stimulation sites; 20.8%) and Japonic (37 stimulation sites; 4.1%). There was a significant difference between the frequencies of the total number of global native speakers and stimulation sites (x2 = 2071.33, df = 28, p <0.00001). We calculated a 'stimulation site per million native speakers' metric and identified that the top mapped languages per population size were French (4.62), Dutch (0.99), and English (0.65). This study underscores a significant bias in DES-based language mapping literature, with under-representation of languages spoken by most of the global population. Standardized, multi-language testing paradigms are crucial for addressing these disparities and advancing inclusive cognitive neuroscience.
This study aims to present a comprehensive maternal genetic perspective on the population history of coastal northwestern India, especially Gujarat. The region is important due to its paleoanthropological significance, its role in the Indus Valley civilization, and its location within routes of Indo-European and Dravidian speaking populations. Previous studies have often used limited samples, underscoring the need for more extensive mitochondrial DNA analyses. We sequenced and analyzed complete mitochondrial genomes from 168 individuals in Gujarat. To place these within a broader phylogenetic framework, we included an additional 529 complete mitogenomes representing East-West Eurasian and South Asian lineages. Phylogenetic relationships and chronological expansions were examined, and Bayesian analysis was used to assess changes in maternal effective population size over time. Our findings show that the majority of (76%) of maternal lineages in Gujarat are from South Asia. East Eurasian and West Eurasian contributions are comparatively low, at 0.6% and 21%, respectively. Looking at the events of the last few millennia, we note that only 19% of West Eurasian lineages appear to have entered the region within the last 5000 years. Most Eurasian haplogroups represent early founding lineages, while 81% of West Eurasian lineages predate the Steppe migration. The data suggest that western India retains largely indigenous maternal ancestry, with little evidence for major maternal migration or replacement over the last 40,000 years. West Eurasian lineages entered in several small waves rather than a single large influx, challenging the classic Indo-Aryan migration model.