Background: Infant cry acoustics provide a promising window into early neurodevelopment and may serve as scalable biomarkers for neurodevelopmental disorders. However, conventional microphone-based recordings are highly susceptible to environmental noise and raise privacy concerns in real-world clinical settings. Chest-surface accelerometers may offer a robust alternative by capturing vibrations directly from the larynx. Methods: We evaluated the validity of a chest-mounted accelerometer (ACC) for infant cry analysis by comparing acoustic features derived from ACC and simultaneously recorded microphone (MIC) signals during routine vaccination visits. The final sample included 85 infants (41 at 4 months; 44 at 12 months) from a diverse pediatric population. Seven vocal measures were extracted from both modalities, including fundamental frequency (F0), jitter, shimmer, cepstral peak prominence (CPP), and harmonics-to-noise ratio (HNR). Agreement and consistency between modalities was assessed using intraclass correlation coefficients (ICCs). Results: F0 demonstrated excellent agreement between ACC and MIC recordings (ICC > 0.94). Jitter measures also showed good-to-excellent agree
Deep learning has advanced pathological voice detection rapidly, yet rare laryngeal diseases remain underexplored due to data scarcity. Recurrent Respiratory Papillomatosis (RRP) exemplifies this gap: an HPV-induced disease of the larynx in which patients oscillate between recurrence and post-surgical remission over the years. RRP demands continuous voice monitoring that existing cross-sectional corpora cannot support. We introduce the first longitudinal voice dataset for RRP, comprising recordings from 26 patients with up to ten years of follow-up. Each session pairs sustained vowels with sentence-level utterances, which are annotated by otolaryngologists and confirmed synchronously with laryngoscopy. Building on this resource, we establish a systematic benchmark spanning handcrafted features, end-to-end deep networks, self-supervised pretrained models, and recent audio large language models, all evaluated under session-level cross-validation with patient-level audit. Per-subject longitudinal analyses further confirm that the cross-sectional discriminative signal reflects laryngoscopic disease state rather than stable speaker attributes. This work lays a foundation for rare longit
Recent advances in foundation models have enabled conversational agents that aim for sustained companionship rather than mere task completion. Yet most still remain unable to support natural, long-term companion-like interactions, resulting in experiences that feel episodic and inauthentic. We argue that current agents overlooked cross-temporal modeling of agents' social behaviors and internal emotions: generated behaviors rarely influence an agent's emotional state, and emotional states seldom shape subsequent behaviors. We present Cross-Temporal Emotion Modeling (CTEM), a framework that links long-term behavioral history to moment-to-moment emotional expression. CTEM establishes a closed loop where past experiences update an evolving emotional state; this state conditions immediate interactions; and user feedback continually revises both memory and emotional state, enabling reflection and anticipation. We instantiate CTEM as Auri, a companion agent on an instant-messaging platform, and report a 21-day in-the-wild study showing that CTEM shows improvements in perceived naturalness, coherence, and emotional harmony.
Articulatory-to-acoustic inversion strongly depends on the type of data used. While most previous studies rely on EMA, which is limited by the number of sensors and restricted to accessible articulators, we propose an approach aiming at a complete inversion of the vocal tract, from the glottis to the lips. To this end, we used approximately 3.5 hours of RT-MRI data from a single speaker. The innovation of our approach lies in the use of articulator contours automatically extracted from MRI images, rather than relying on the raw images themselves. By focusing on these contours, the model prioritizes the essential geometric dynamics of the vocal tract while discarding redundant pixel-level information. These contours, alongside denoised audio, were then processed using a Bi-LSTM architecture. Two experiments were conducted: (1) the analysis of the impact of the audio embedding, for which three types of embeddings were evaluated as input to the model (MFCCs, LCCs, and HuBERT), and (2) the study of the influence of the dataset size, which we varied from 10 minutes to 3.5 hours. Evaluation was performed on the test data using RMSE, median error, as well as Tract Variables, to which we a
This study examines the implications of Russia's full-scale invasion of Ukraine for the international mobility of Ukrainian scholars. The dataset, drawn from the CWTS in-house Scopus database, includes Ukrainian scholars who were internationally mobile between 2020 and 2023. The analysis focuses on scholars affiliated with universities and the National Academy of Sciences of Ukraine (NASU) prior to moving abroad. The findings reveal an increase in the number of internationally mobile scholars in 2022-2023, driven primarily by rising mobility from universities. For NASU-affiliated scholars, Russia was the top destination country in 2020-2021 but fell to fourth place in 2022-2023, overtaken by Germany, China, and Poland. For university-affiliated scholars, Poland, Germany, and Russia consistently ranked as the top three destination countries across both periods. Statistical tests indicate no significant difference in mean Field-Weighted Citation Impact (FNCI) between scholars who were internationally mobile in 2020-2021 and those mobile in 2022-2023. However, the share of internationally mobile scholars with articles among the top 10% most cited globally increased among those previou
This study examines the role of foreign co-affiliations in shaping the research performance of Ukrainian universities and research institutes of the National Academy of Sciences of Ukraine (NASU) before and during Russias full-scale invasion. In 2023, the share of articles with foreign co-affiliations was higher for NASU (17.1 percent) than for universities (10.6 percent), reflecting NASUs research-oriented profile and strong focus on physical sciences & engineering. In 2022 and 2023, this share increased across all disciplines in both types of institutions. The largest shares of articles with foreign co-affiliations involved institutions in advanced science systems such as Germany and China, as well as neighboring Poland, Czechia, and Slovakia. Articles with foreign co-affiliations showed citation impact comparable to internationally co-authored articles and substantially outperformed purely domestic publications.On the one hand, articles with foreign co-affiliations ensure research continuity and enhance the citation visibility of a resource-constrained and war-affected national science system. On the other hand, they distort the measurement of the national and institutional
The analysis of case-control point pattern data is an important problem in spatial epidemiology. The spatial variation of cases if often compared to that of a set of controls to assess spatial risk variation as well as the detection of risk factors and exposure to putative pollution sources using spatial regression models. The intensities of the point patterns of cases and controls are estimated using log-Gaussian Cox models, so that fixed and spatial random effects can be included. Bayesian inference is conducted via the integrated Nested Laplace approximation (INLA) method using the inlabru R package. In this way, potential risk factors can be assessed by including them as fixed effects while residual spatial variation is considered as a Gaussian process with Matérn covariance. In addition, exposure to pollution sources is modeled using different smooth terms. The proposed methods have been applied to the Chorley-Ribble dataset, that records the locations of lung and larynx cancer cases as well as the location of an disused old incinerator in the area of Lancashire (England, United Kingdom). Taking the locations of lung cancer as controls, the spatial variation of both types of c
Accurate simulations of the flow in the human airway are essential for advancing diagnostic methods. Many existing computational studies rely on simplified geometries or turbulence models, limiting their simulation's ability to resolve flow features such shear-layer instabilities or secondary vortices. In this study, direct numerical simulations were performed for inspiratory flow through a detailed airway model which covers the nasal mask region to the 6th bronchial bifurcation. Simulations were conducted at two physiologically relevant \textsc{Reynolds} numbers with respect to the pharyngeal diameter, i.e., at Re_p=400 (resting) and Re_p=1200 (elevated breathing). These values characterize resting and moderately elevated breathing conditions. A lattice-Boltzmann method was employed to directly simulate the flow, i.e., no turbulence model was used. The flow field was examined across four anatomical regions: 1) the nasal cavity, 2) the naso- and oropharynx, 3) the laryngopharynx and larynx, and 4) the trachea and carinal bifurcation. The total pressure loss increased from 9.76 Pa at Re_p=400 to 41.93 Pa at Re_p=1200. The nasal cavity accounted for the majority of this loss for both
This study explores funding, authorship patterns, and citation impact of articles funded by the Ministry of Education and Science of Ukraine (MESU), the National Academy of Sciences of Ukraine (NASU), and the National Research Foundation of Ukraine (NRFU). The analysis focuses on articles published in Scopus-indexed journals between 2020 and 2023. The findings show that the share of articles funded by these agencies increased from 8.6% in 2020-2021 to 11.9% in 2022-2023. Foreign co-funding as well as international co-authorship and co-affiliations are consistently associated with higher citation impact. In particular, foreign co-affiliations are associated with higher field-normalised citation impact (FNCI) for MESU-funded articles in 2022-2023, exceeding that of articles jointly funded by MESU and foreign agencies. NASU funding is associated with only modest differences in citation impact relative to unfunded articles. These effects are small and not consistently significant across authorship patterns and become less pronounced in 2022-2023, as the citation impact of unfunded articles partially converges with that of funded articles. While the results should be interpreted as aver
This study aimed to explore the relationship between access models, authorship patterns, and citation impact in Ukrainian research output from 2020 to 2023. The focus was on scholars affiliated with the National Academy of Sciences of Ukraine (NASU) and universities. Findings highlight that open access (OA) articles constituted the majority of publications by Ukrainian scholars during this period. This percentage reached 75.4% for NASU and 85.8% for universities. In both cases, the increase was driven by Gold OA and Hybrid Gold OA, the latter benefiting in part from Elsevier's waivers. Diamond OA prevailed for NASU, while Gold OA was dominant for universities. The effects of Russia's full-scale invasion of Ukraine included (1) a decline in the share of articles in foreign journals for both NASU and universities, (2) a decrease in Gold OA in foreign journals and an increase in Gold OA in Ukrainian journals for universities, and (3) a rise in internationally co-authored Gold OA articles in foreign journals for both entities. Despite waivers for Gold OA provided by major publishers and an increase in Gold OA articles in Elsevier and Springer journals, MDPI and Aluna Publishing House r
This study explores the effects of Russia's full-scale invasion of Ukraine on the international collaboration of Ukrainian scholars. First and foremost, Ukrainian scholars deserve respect for continuing to publish despite life-threatening conditions, mental strain, shelling and blackouts. In 2022-2023, universities gained more from international collaboration than the NASU. The percentage of internationally co-authored articles remained unchanged for the NASU, while it increased for universities. In 2023, 40.8% of articles published by the NASU and 32,2% of articles published by universities were internationally co-authored. However, these figures are still much lower than in developed countries (60-70%). The citation impact of internationally co-authored articles remained statistically unchanged for the NASU but increased for universities. The highest share of internationally co-authored articles published by the NASU in both periods was in the physical sciences and engineering. However, the citation impact of these articles declined in 2022-2023, nearly erasing their previous citation advantage over university publications. Universities consistently outperformed the NASU in the c
The graphical operation of insplitting is key to understanding conjugacy of shifts of finite type (SFTs) in both one and two dimensions. In this paper, we consider two approaches to studying 2-dimensional SFTs: textile systems and rank-2 graphs. Nasu's textile systems describe all two-sided 2D SFTs up to conjugacy, whereas the 2-graphs (higher-rank graphs of rank 2) introduced by Kumjian and Pask yield associated C*-algebras. Both models have a naturally-associated notion of insplitting. We show that these notions do not coincide, raising the question of whether insplitting a 2-graph induces a conjugacy of the associated one-sided 2-dimensional SFTs. Our first main result shows how to reconstruct 2-graph insplitting using textile-system insplits and inversions, and consequently proves that 2-graph insplitting induces a conjugacy of dynamical systems. We also present several other facets of the relationship between 2-graph insplitting and textile-system insplitting. Incorporating an insplit of the bottom graph of the textile system turns out to be key to this relationship. By articulating the connection between operator-algebraic and dynamical notions of insplitting in two dimension
We propose ARTI-6, a compact six-dimensional articulatory speech encoding framework derived from real-time MRI data that captures crucial vocal tract regions including the velum, tongue root, and larynx. ARTI-6 consists of three components: (1) a six-dimensional articulatory feature set representing key regions of the vocal tract; (2) an articulatory inversion model, which predicts articulatory features from speech acoustics leveraging speech foundation models, achieving a prediction correlation of 0.87; and (3) an articulatory synthesis model, which reconstructs intelligible speech directly from articulatory features, showing that even a low-dimensional representation can generate natural-sounding speech. Together, ARTI-6 provides an interpretable, computationally efficient, and physiologically grounded framework for advancing articulatory inversion, synthesis, and broader speech technology applications. The source code and speech samples are publicly available.
Laryngeal cancer imaging research lacks standardised public datasets to enable reproducible deep learning (DL) model development. We present LaryngealCT, a curated benchmark of 1,029 computed tomography (CT) scans aggregated from six collections from The Cancer Imaging Archive (TCIA). Uniform 1 mm isotropic volumes of interest encompassing the larynx were extracted using a weakly supervised parameter search framework validated by clinical experts. Six 3D DL architectures (custom 3D CNN, ResNet18,50,101, DenseNet121 and MedicalNet-pretrained ResNet50) were benchmarked on (i) early (Tis,T1,T2) vs. advanced (T3,T4) and (ii) T4 vs. non-T4 classification tasks. On the independent test set, the 3D CNN achieved the strongest overall performance across global and per-class metrics (Accuracy 0.854, F1-macro 0.841) in early vs. advanced classification. In the T4 task, AU-ROC values exceeded 0.82 for most models, but sensitivity for T4 disease remained limited (less than or equal to 0.412), with ResNet101 showing the most promising calibrated T4 recall (0.706. Model explainability assessed using GradCAMpp with thyroid cartilage overlays for T4 classification task revealed anatomically plausib
This letter reports the design, construction, and experimental validation of a novel hand-held robot for in-office laser surgery of the vocal folds. In-office endoscopic laser surgery is an emerging trend in Laryngology: It promises to deliver the same patient outcomes of traditional surgical treatment (i.e., in the operating room), at a fraction of the cost. Unfortunately, office procedures can be challenging to perform; the optical fibers used for laser delivery can only emit light forward in a line-of-sight fashion, which severely limits anatomical access. The robot we present in this letter aims to overcome these challenges. The end effector of the robot is a steerable laser fiber, created through the combination of a thin optical fiber (0.225 mm) with a tendon-actuated Nickel-Titanium notched sheath that provides bending. This device can be seamlessly used with most commercially available endoscopes, as it is sufficiently small (1.1 mm) to pass through a working channel. To control the fiber, we propose a compact actuation unit that can be mounted on top of the endoscope handle, so that, during a procedure, the operating physician can operate both the endoscope and the steerab
Introduction Speech is an integral component of human communication, requiring the coordinated efforts of various organs to produce sound (Titze & Alipour, 2006). The glottis region, a key player in voice production, assumes a crucial role in this intricate process. As air, emanating from the lungs in a confined space, interacts with the vocal folds (VFs) within the human body, it gives rise to the creation of voice (Alipour & Vigmostad, 2012). Understanding the mechanical intricacies of this process is very important. Studying VFs in vivo situations is hard work. However, the orientation, shape and size of VFs fibers have been extracted with synchrotron X-ray microtomography. (Bailly et al., 2018) The investigation of mechanical properties of both human and animal VFs has been carried out through various methodologies in the literature. The mechanical properties of VFs have been studied using the uniaxial extension test (Alipour & Vigmostad, 2012) assuming a linear behavior, while the nonlinearity and anisotropy of VFs has been determined using a multiscale method as in Miri et al. (2013). Pipette aspiration has also been used to extract in vivo elastic properties of V
This paper studies changes in articulatory configurations across genders and periods using an inversion from acoustic to articulatory parameters. From a diachronic corpus based on French media archives spanning 60 years from 1955 to 2015, automatic transcription and forced alignment allowed extracting the central frame of each vowel. More than one million frames were obtained from over a thousand speakers across gender and age categories. Their formants were used from these vocalic frames to fit the parameters of Maeda's articulatory model. Evaluations of the quality of these processes are provided. We focus here on two parameters of Maeda's model linked to total vocal tract length: the relative position of the larynx (higher for females) and the lips protrusion (more protruded for males). Implications for voice quality across genders are discussed. The effect across periods seems gender independent; thus, the assertion that females lowered their pitch with time is not supported.
Patients who have had their entire larynx removed, including the vocal folds, owing to throat cancer may experience difficulties in speaking. In such cases, electrolarynx devices are often prescribed to produce speech, which is commonly referred to as electrolaryngeal speech (EL speech). However, the quality and intelligibility of EL speech are poor. To address this problem, EL voice conversion (ELVC) is a method used to improve the intelligibility and quality of EL speech. In this paper, we propose a novel ELVC system that incorporates cross-domain features, specifically spectral features and self-supervised learning (SSL) embeddings. The experimental results show that applying cross-domain features can notably improve the conversion performance for the ELVC task compared with utilizing only traditional spectral features.
This article deals with large-eddy simulations of 3D incompressible laryngeal flow followed by acoustic simulations of human phonation of five cardinal english vowels /u, i, \textipa{A}, o, æ/. The flow and aeroacoustic simulations were performed in OpenFOAM and in-house code openCFS, respectively. Given the large variety of scales in the flow and acoustics, the simulation is separated into two steps: (1) computing the flow in the larynx using the finite volume method on a fine 2.2M grid followed by (2) computing the sound sources separately and wave propagation to the radiation zone around the mouth using the finite element method on a coarse 33k acoustic grid. The numerical results showed that the anisotropic minimum dissipation model, which is not well known since it is not available in common CFD software, predicted stronger sound pressure levels at higher harmonics and especially at first two formants than the wall-adapting local eddy-viscosity model. We implemented the model as a new open library in OpenFOAM and deployed the model on turbulent flow in the larynx with positive impact on the quality of simulated vowels. Numerical simulations are in very good agreement with posi
Context. Cassiopeia A occupies an important place among supernova remnants (SNRs) in low-frequency radio astronomy. The analysis of its continuum spectrum from low frequency observations reveals the evolution of the SNR absorption properties over time and suggests a method for probing unshocked ejecta and the SNR interaction with the circumstellar medium (CSM). Aims. In this paper we present low-frequency measurements of the integrated spectrum of Cassiopeia A to find the typical values of free-free absorption parameters towards this SNR in the middle of 2023. We also add new results to track its slowly evolving and decreasing integrated flux density. Methods. We used the New Extension in Nançay Upgrading LOFAR (NenuFAR) and the Ukrainian Radio Interferometer of NASU (URAN-2, Poltava) for measuring the continuum spectrum of Cassiopeia A within the frequency range of 8-66 MHz. The radio flux density of Cassiopeia A has been obtained on June-July, 2023 with two sub-arrays for each radio telescope, used as a two-element correlation interferometer. Results. We measured magnitudes of emission measure, electron temperature and an average number of charges of the ions for both internal an