Let $G$ be a graph with vertex set $V=V(G)$. A double Roman dominating function on a graph $G$ is a function $f : V \to \{0,1,2,3\}$ satisfying the conditions that if $f(v) = 0$, then vertex $v$ must have at least two neighbors in $V_2$ or one neighbor in $V_3$, if $f(v) = 1$, then vertex $v$ must have at least one neighbor in $V_2 \cup V_3$. The weight of a double Roman dominating function $f$ is the sum $f(V) = \sum_{v \in V} f(v)$, and the double Roman domination number $γ_{dR}(G)$ is the minimum weight of a double Roman dominating function on $G$. A double Italian dominating function on a graph $G$ is a function $f : V \to \{0,1,2,3\}$ satisfying the condition that for every vertex $u \in V$, if $f(u) \in \{0,1\}$, then $\sum_{v \in N[u]} f(v) \ge 3$. The double Roman domination number $γ_{dI}(G)$ is the minimum weight of a double Italian dominating function on $G$. Mojdeh and Volkmann [D.A. Mojdeh and L. Volkmann, Roman {3}-domination (double Italian domination), Discrete Appl. Math. 283 (2020), 555--564] proved that $γ_{dI}(T) = γ_{dR}(T)$ for any tree $T$. However, we find that there is a minor issue in the proof. In this paper, we first prove that $γ_{dI}(T) eq γ_{dR}(T)$.
While Italian is a high-resource language, there are few Italian-native benchmarks to evaluate generative Large Language Models (LLMs) in this language. This work presents three new benchmarks: Invalsi MATE to evaluate models performance on mathematical understanding in Italian, Invalsi ITA to evaluate language understanding in Italian and Olimpiadi MATE for more complex mathematical understanding. The first two benchmarks are based on the Invalsi tests, which are administered to students of age between 6 and 18 within the Italian school system and have been validated by several experts in teaching and pedagogy, the third one comes from the Italian high school math Olympics. We evaluate 10 powerful language models on these benchmarks and find that they are bound by 71% accuracy on Invasli MATE, achieved by Llama 3.1 70b instruct and by 88% on Invalsi ITA. For both Invalsi MATE and Invalsi ITA we compare LLMs with the average performance of Italian students to show that Llama 3.1 is the only one to outperform them on Invalsi MATE while most models do so on Invalsi ITA, we then show that Olimpiadi MATE is more challenging than Invalsi MATE and the highest accuracy, achieved by Llama
We describe Evalita-LLM, a new benchmark designed to evaluate Large Language Models (LLMs) on Italian tasks. The distinguishing and innovative features of Evalita-LLM are the following: (i) all tasks are native Italian, avoiding issues of translating from Italian and potential cultural biases; (ii) in addition to well established multiple-choice tasks, the benchmark includes generative tasks, enabling more natural interaction with LLMs; (iii) all tasks are evaluated against multiple prompts, this way mitigating the model sensitivity to specific prompts and allowing a fairer and objective evaluation. We propose an iterative methodology, where candidate tasks and candidate prompts are validated against a set of LLMs used for development. We report experimental results from the benchmark's development phase, and provide performance statistics for several state-of-the-art LLMs.
Spoken language datasets are vital for advancing linguistic research, Natural Language Processing, and speech technology. However, resources dedicated to Italian, a linguistically rich and diverse Romance language, remain underexplored compared to major languages like English or Mandarin. This survey provides a comprehensive analysis of 66 spoken Italian datasets, highlighting their characteristics, methodologies, and applications. The datasets are categorized by speech type, source and context, and demographic and linguistic features, with a focus on their utility in fields such as Automatic Speech Recognition, emotion detection, and education. Challenges related to dataset scarcity, representativeness, and accessibility are discussed alongside recommendations for enhancing dataset creation and utilization. The full dataset inventory is publicly accessible via GitHub and archived on Zenodo, serving as a valuable resource for researchers and developers. By addressing current gaps and proposing future directions, this work aims to support the advancement of Italian speech technologies and linguistic research.
Adapting models to a language that was only partially present in the pre-training data requires fine-tuning, which is expensive in terms of both data and computational resources. As an alternative to fine-tuning, we explore the potential of activation steering-based techniques to enhance model performance on Italian tasks. Through our experiments we show that Italian steering (i) can be successfully applied to different models, (ii) achieves performances comparable to, or even better than, fine-tuned models for Italian, and (iii) yields higher quality and consistency in Italian generations. We also discuss the utility of steering and fine-tuning in the contemporary LLM landscape where models are anyway getting high Italian performances even if not explicitly trained in this language.
The development of domain-specific language models has significantly advanced natural language processing applications in various specialized fields, particularly in biomedicine. However, the focus has largely been on English-language models, leaving a gap for less-resourced languages such as Italian. This paper introduces Igea, the first decoder-only language model designed explicitly for biomedical text generation in Italian. Built on the Minerva model and continually pretrained on a diverse corpus of Italian medical texts, Igea is available in three model sizes: 350 million, 1 billion, and 3 billion parameters. The models aim to balance computational efficiency and performance, addressing the challenges of managing the peculiarities of medical terminology in Italian. We evaluate Igea using a mix of in-domain biomedical corpora and general-purpose benchmarks, highlighting its efficacy and retention of general knowledge even after the domain-specific training. This paper discusses the model's development and evaluation, providing a foundation for future advancements in Italian biomedical NLP.
Background and aim: Considering the scope of the application of artificial intelligence beyond the field of computer science, one of the concerns of researchers is to provide quality explanations about the functioning of algorithms based on artificial intelligence and the data extracted from it. The purpose of the present study is to validate the Italian version of system causability scale (I-SCS) to measure the quality of explanations provided in a xAI. Method: For this purpose, the English version, initially provided in 2020 in coordination with the main developer, was utilized. The forward-backward translation method was applied to ensure accuracy. Finally, these nine steps were completed by calculating the content validity index/ratio and conducting cognitive interviews with representative end users. Results: The original version of the questionnaire consisted of 10 questions. However, based on the obtained indexes (CVR below 0.49), one question (Question 8) was entirely removed. After completing the aforementioned steps, the Italian version contained 9 questions. The representative sample of Italian end users fully comprehended the meaning and content of the questions in the I
The extraction of pharmacological knowledge from regulatory documents has become a key focus in biomedical natural language processing, with applications ranging from adverse event monitoring to AI-assisted clinical decision support. However, research in this field has predominantly relied on English-language corpora such as DrugBank, leaving a significant gap in resources tailored to other healthcare systems. To address this limitation, we introduce DART (Drug Annotation from Regulatory Texts), the first structured corpus of Italian Summaries of Product Characteristics derived from the official repository of the Italian Medicines Agency (AIFA). The dataset was built through a reproducible pipeline encompassing web-scale document retrieval, semantic segmentation of regulatory sections, and clinical summarization using a few-shot-tuned large language model with low-temperature decoding. DART provides structured information on key pharmacological domains such as indications, adverse drug reactions, and drug-drug interactions. To validate its utility, we implemented an LLM-based drug interaction checker that leverages the dataset to infer clinically meaningful interactions. Experiment
Telegram has grown into a significant platform for news and information sharing, favored for its anonymity and minimal moderation. This openness, however, makes it vulnerable to misinformation and conspiracy theories. In this study, we explore the dynamics of conspiratorial narrative dissemination within Telegram, focusing on Italian and English landscapes. In particular, we leverage the mechanism of message forwarding within Telegram and collect two extensive datasets through snowball strategy. We adopt a network-based approach and build the Italian and English Telegram networks to reveal their respective communities. By employing topic modeling, we uncover distinct narratives and dynamics of misinformation spread. Results highlight differences between Italian and English conspiracy landscapes, with Italian discourse involving assorted conspiracy theories and alternative news sources intertwined with legitimate news sources, whereas English discourse is characterized by a more focused approach on specific narratives such as QAnon and political conspiracies. Finally, we show that our methodology exhibits robustness across initial seed selections, suggesting broader applicability. T
In recent years Large Language Models (LLMs) have increased the state of the art on several natural language processing tasks. However, their accessibility is often limited to paid API services, posing challenges for researchers in conducting extensive investigations. On the other hand, while some open-source models have been proposed by the community, they are typically English-centric or multilingual without a specific adaptation for the Italian language. In an effort to democratize the available and open resources for the Italian language, in this paper we introduce Camoscio: a language model specifically tuned to follow users' prompts in Italian. Specifically, we finetuned the smallest variant of LLaMA (7b) with LoRA on a corpus of instruction prompts translated to Italian via ChatGPT. Results indicate that the model's zero-shot performance on various downstream tasks in Italian competes favorably with existing models specifically finetuned for those tasks. All the artifacts (code, dataset, model) are released to the community at the following url: https://github.com/teelinsan/camoscio
Consider a finite simple digraph $D$ with vertex set $V(D)$. An Italian dominating function (IDF) on $D$ is a function $f:V(D)\rightarrow\{0,1,2\}$ satisfying every vertex $u$ with $f(u)=0$ has an in-neighbor $v$ with $f(v)=2$ or two in-neighbors $w$ and $z$ with $f(w)=f(z)=1$. A total Italian dominating function (TIDF) on $D$ is an IDF $f$ such that the subdigraph $D[\{ u\, |\, f(u)\ge 1\}]$ contains no isolated vertices. The weight $ω(f)$ of a TIDF $f$ on $D$ is $\sum_{u\in V(D)}f(u)$. The total Italian domination number of $D$ is $γ_{tI}(D)=\min\{ ω(f)\, |\, \mbox{$f$ is a TIDF on $D$}\}$. In this paper, we present bounds on $γ_{tI}(D)$, and investigate the relationship between several different domination parameters. In particular, we give the total Italian domination number of the Cartesian products $P_2\Box P_n$ and $P_3\Box P_n$, where $P_n$ represents a dipath with $n$ vertices.
The present project endeavors to enrich the linguistic resources available for Italian by constructing a Universal Dependencies treebank for the KIParla corpus (Mauri et al., 2019, Ballarè et al., 2020), an existing and well known resource for spoken Italian.
A discussion on the readiness of Italian universities to address gender-related issues from a regional standpoint is proposed. A statistical analysis is conducted on data of all scholars enrolled in Italian universities from 2000 to 2023 to investigate why the glass ceiling of the full professor position remains so challenging to break in almost all scientific fields and across all regions of Italy.
An \textit{Italian dominating function} on a digraph $D$ with vertex set $V(D)$ is defined as a function $f : V(D) \rightarrow \{0, 1, 2\}$ such that every vertex $v \in V(D)$ with $f(v) = 0$ has at least two in-neighbors assigned $1$ under $f$ or one in-neighbor $w$ with $f(w) = 2$. The \textit{weight} of an Italian dominating function $f$ is the value $ω(f) = f(V(D)) = \sum_{u \in V(D)} f(u)$. The \textit{Italian domination number} of a digraph $D$, denoted by $γ_I(D)$, is the minimum taken over the weights of all Italian dominating functions on $D$. The \textit{Italian bondage number} of a digraph $D$, denoted by $b_I(D)$, is the minimum number of arcs of $A(D)$ whose removal in $D$ results in a digraph $D'$ with $γ_I(D') > γ_I(D)$. The \textit{Italian reinforcement number} of a digraph $D$, denoted by $r_I(D)$, is the minimum number of extra arcs whose addition to $D$ results in a digraph $D'$ with $γ_I(D') < γ_I(D)$. In this paper, we initiate the study of Italian bondage and reinforcement numbers in digraphs and present some bounds for $b_I(D)$ and $r_I(D)$. We also determine the Italian bondage and reinforcement numbers of some classes of digraphs.
This paper presents Fauno, the first and largest open-source Italian conversational Large Language Model (LLM). Our goal with Fauno is to democratize the study of LLMs in Italian, demonstrating that obtaining a fine-tuned conversational bot with a single GPU is possible. In addition, we release a collection of datasets for conversational AI in Italian. The datasets on which we fine-tuned Fauno include various topics such as general question answering, computer science, and medical questions. We release our code and datasets on \url{https://github.com/RSTLess-research/Fauno-Italian-LLM}
Recent large-scale Spoken Language Understanding datasets focus predominantly on English and do not account for language-specific phenomena such as particular phonemes or words in different lects. We introduce ITALIC, the first large-scale speech dataset designed for intent classification in Italian. The dataset comprises 16,521 crowdsourced audio samples recorded by 70 speakers from various Italian regions and annotated with intent labels and additional metadata. We explore the versatility of ITALIC by evaluating current state-of-the-art speech and text models. Results on intent classification suggest that increasing scale and running language adaptation yield better speech models, monolingual text models outscore multilingual ones, and that speech recognition on ITALIC is more challenging than on existing Italian benchmarks. We release both the dataset and the annotation scheme to streamline the development of new Italian SLU models and language-specific datasets.
The recent introduction of Transformers language representation models allowed great improvements in many natural language processing (NLP) tasks. However, if on one hand the performances achieved by this kind of architectures are surprising, on the other their usability is limited by the high number of parameters which constitute their network, resulting in high computational and memory demands. In this work we present BERTino, a DistilBERT model which proposes to be the first lightweight alternative to the BERT architecture specific for the Italian language. We evaluated BERTino on the Italian ISDT, Italian ParTUT, Italian WikiNER and multiclass classification tasks, obtaining F1 scores comparable to those obtained by a BERTBASE with a remarkable improvement in training and inference speed.
This paper presents an overview of the state of scientific research in physics in the Italian peninsula for the first thirty years of XIX century. In doing so, we focus geographically on the Kingdom of Sardinia and the Lombardo - Veneto territories, because of the role played in the development of Italian physics and the political future of the peninsula; we line out a quantitative analysis of the scientific community using basic tool of graph theory; we somehow try to avoid the major personalities of the century (e.g. Volta, Avogadro, Lagrange) to highlight the development of the local communities. The language of the paper is Italian.
Online medical forums have long served as vital platforms where patients seek professional healthcare advice, generating vast amounts of valuable knowledge. However, the informal nature and linguistic complexity of forum interactions pose significant challenges for automated question answering systems, especially when dealing with non-English languages. We present two comprehensive Italian medical benchmarks: \textbf{IMB-QA}, containing 782,644 patient-doctor conversations from 77 medical categories, and \textbf{IMB-MCQA}, comprising 25,862 multiple-choice questions from medical specialty examinations. We demonstrate how Large Language Models (LLMs) can be leveraged to improve the clarity and consistency of medical forum data while retaining their original meaning and conversational style, and compare a variety of LLM architectures on both open and multiple-choice question answering tasks. Our experiments with Retrieval Augmented Generation (RAG) and domain-specific fine-tuning reveal that specialized adaptation strategies can outperform larger, general-purpose models in medical question answering tasks. These findings suggest that effective medical AI systems may benefit more from
In 1929, Enrico Fermi wrote "Problemi attuali della fisica" ("Contemporary Problems of Physics"), a short article in Italian published in the magazine "Annali dell'istruzione media". The magazine was sponsored by the Italian Ministry of Schooling and Education, and it was addressed to teachers and principals of middle and high schools. This short text written by Fermi had been forgotten for a long time. However, in recent years it has been republished in Italy; moreover, the "Annali dell'istruzione media" have been digitized and made freely available online. It is unclear why Fermi wrote an article of a clearly divulgative nature for a non-technical journal, but is still interesting to to read his words and appreciate his skills as a great scientific communicator. In this work we include the transcription of the original article in Italian, and we also propose an English translation to make the text available world-wide and accessible to a broader public.