Some children are slower than average to acquire early spoken language milestones. Little is known about how this heterogeneity in expressive language impacts later childhood language and reading outcomes in children with dyslexia. We investigated oral language and reading differences in 63 children with dyslexia aged 8-12 years, comparing those with (35%) and without (65%) a prior history of delayed two-word phrase production (beyond 24 months of age). All children subsequently acquired fluent spoken language. Language outcome measures included standardized tests of current sound awareness, rapid picture naming, memory for words, and sentence repetition. Current single word and paragraph reading abilities were measured using standardized timed and untimed tests. Results indicated that both groups showed comparable single word and paragraph reading performance in middle childhood. Those with early spoken language delay had poorer sentence repetition and memory for words than those who used two-word phrases by 24-months. Our findings point to distinct language contributions to dyslexia in children with and without a history of early spoken language acquisition difficulties, underscoring the importance of considering early language milestones in this population.
Postoperative nausea and vomiting are common complications after anesthesia. However, vomiting represents a clinically distinct and objectively measurable endpoint. This study aimed to develop and internally validate predictive models for postoperative vomiting within 24 hours using structured perioperative data and unstructured clinical text, while introducing a structured framework that separates feature construction from interpretability using large language models (LLMs). We analyzed 33,460 anesthesia records from a single center (2019-2022). Two temporally defined prediction tasks were constructed to reflect real-world clinical decision-making and prevent information leakage: a preoperative model using variables available before anesthesia induction, and a perioperative model using variables available up to the end of surgery. Structured data were modeled using machine learning algorithms (logistic regression, Extreme Gradient Boosting, Light Gradient Boosting Machine [LightGBM]). Unstructured clinical text was incorporated through a deterministic, concept-driven preprocessing pipeline, where LLMs were used solely for normalization (temperature=0) without feature generation, followed by rule-based concept mapping and feature encoding. Post hoc interpretability was further supported using an LLM-based Question Answering Chain module. Model performance was evaluated using receiver operating characteristic-area under the curve (AUC), precision-recall AUC, calibration metrics, and threshold-based operating characteristics. Classification thresholds were selected using the Youden J statistic, and all metrics were reported with 95% CIs derived from bootstrap resampling. Decision curve analysis was performed to assess clinical utility. A total of 33,460 surgical procedures were included, of which 3607 (10.8%) experienced postoperative vomiting within 24 hours. In the preoperative task, LightGBM achieved an AUC of 0.729 (95% CI 0.706-0.749), compared with 0.610 (95% CI 0.588-0.632) for the Apfel score. In the end-of-surgery task, LightGBM achieved an AUC of 0.735 (95% CI 0.714-0.757). At the Youden-optimal threshold, the negative predictive value exceeded 0.95 across all models. Decision curve analysis demonstrated positive net benefit across clinically relevant threshold probabilities. Incorporating text-derived features provided modest improvements, while LLM-based explanation modules generated structured, natural-language explanations intended to enhance interpretability without substantially improving predictive performance. Machine learning models can effectively predict postoperative vomiting within 24 hours using perioperative data. The proposed framework demonstrates that LLMs can be integrated in a controlled and reproducible manner-restricted to deterministic normalization and post hoc reasoning-to generate natural-language explanations intended to enhance the interpretability of model predictions, without introducing information leakage or altering predictive modeling. As no formal clinician-based evaluation was conducted, this interpretability benefit cannot yet be objectively confirmed, and the generated explanations should be regarded as a useful interpretability aid to be validated in future clinician-centered studies. External, multicenter validation is required before broader clinical applicability can be assumed.
As global teacher mobility increases, transnational educators often encounter acculturative stress and challenges to their professional identity. However, the psychological processes through which adaptation unfolds over time remain underexplored. Drawing on Ecological Systems Theory and the Transactional Model of Stress and Coping, this study examines intercultural adaptation as a dynamic process shaped by changing environmental demands, cognitive appraisals, and coping responses. Using a longitudinal phenomenological case study design, the study followed an Asian English language teacher over a three-year period as she transitioned into a self-initiated expatriate (SIE) role in Türkiye. Data were collected through semi-structured interviews, reflective journals, email correspondence, and portfolio documents across three developmental phases. Reflexive Thematic Analysis (RTA) informed by Interpretative Phenomenological Analysis (IPA), guided the analysis. The findings indicate that adaptation was experienced as an ongoing reconstruction of professional identity rather than a linear process of adjustment. Initially, heightened visibility and macrosystem pressures were associated with anxiety and emotion-focused coping. Over time, participation in new educational contexts and increased host-language engagement supported more problem-focused coping strategies and a growing sense of agency. In the final phase, institutional demands linked to the participant's perceived foreign identity and classroom challenges, contributed to acculturative stress and identity tension. However, these experiences were accompanied by the development of a more flexible transnational professional identity rather than assimilation into dominant norms. The findings suggest that, in this case, acculturative stress was shaped not only by individual characteristics but also by the interaction between ecological conditions, cognitive appraisals, coping processes, and evolving professional identities. The study contributes to psychological research by illustrating how identity reconstruction may emerge through repeated cycles of appraisal and coping across changing sociocultural contexts, with implications for language teacher education and the support of transnational teachers.
The aging Canadian population has led to an increase in Canada's use of home and community care services. In this context, the healthcare system relies heavily on personal support workers who work in various healthcare settings. A significant proportion of these workers are immigrants, and many live in linguistic minorities. Despite their pivotal role, the experiences and working conditions of these personal support workers are underrepresented in the scientific literature, specifically in intersectional analyses. This scoping review aims to examine the work experiences, health conditions, and well-being of immigrant personal support workers in Canada who work in a minority language setting. Studies addressing the work experiences, health conditions, and well-being of immigrant personal support workers in Canada who work in a minority language setting, and those published in English or French. There will have no restrictions on publication date. This review will follow the JBI recommendations for scoping reviews and incorporate the PRISMA-ScR checklist. We will develop an adapted search strategy for five relevant databases (Medline [Ovide], Web of Science, Embase, CINAHL, and Google Scholar). Two independent reviewers will select full-text articles based on pre-established inclusion criteria and extracted relevant data. The results will be presented in accordance with the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-analyses Extension for Scoping Reviews) guidelines. This scoping review contributes to expanding knowledge about the professional, health, and social realities of immigrant PSWs in Canada's linguistic minority communities.
Large language models (LLMs) have recently demonstrated exceptional capabilities, offering the potential to streamline workflows and enhance efficiency across diverse tasks. However, their application in domain-specific areas, such as radiotherapy treatment planning, remains challenging due to the lack of publicly available, specialized knowledge required for these tasks. To investigate the critical architectural and functional components necessary to adapt an off-the-shelf large language model into a clinically viable planning agent for inverse treatment planning in intensity-modulated radiation therapy (IMRT). Twenty head-and-neck (HN) cancer patients who received IMRT at our institution were retrospectively collected under IRB approval. The LLM agent was implemented to directly interact with the clinical treatment planning system (TPS) to iteratively extract intermediate plan states and propose new constraint values to guide inverse optimization. Its decision-making was informed by real-time plan evaluations and prior optimization outcomes, enabling adaptive refinement of planning strategies across iterations. The agent was equipped with three core capabilities: comprehension of clinical planning objectives, contextual understanding of the optimization environment, and arithmetic proficiency for quantitative reasoning. Treatment planning was conducted in a reference-free inference setting, wherein the LLM operated without prior exposure to manually generated treatment plans and without any fine-tuning or task-specific training. LLM-generated treatment plans were compared with clinically approved plans created by certified dosimetrists. Key dosimetric endpoints were evaluated and statistically analyzed. Paired comparisons between LLM-generated and clinical plans were performed using the Wilcoxon signed-rank test. LLM-generated plans achieved comparable organ-at-risk (OAR) sparing relative to clinical plans, while demonstrating improved hot spot control (Dmax: 106.5% vs. 108.8%, p < 0.05) and superior conformity (conformity index: 1.18 vs. 1.39, p < 0.05 for boost PTV; 1.82 vs. 1.88, p = 0.47 for primary PTV). This study demonstrates the feasibility of a reference-free, LLM-driven workflow for automated IMRT treatment planning in a commercial TPS. The proposed approach provides a generalizable solution that could reduce planning variability and support broader adoption of AI-based planning strategies.
Recent advances in large language models (LLMs) such as GPT-3/4 have spurred the development of artificial intelligence (AI) chatbots and advisory tools in medicine. These systems are posited to assist or augment physician-patient communication, potentially improving empathy, clarity, and responsiveness. However, their actual impact on communication outcomes remains uncertain. This study aimed to systematically review and meta-analyze peer-reviewed studies (2020-2025) evaluating how LLM-based interventions affect physician-patient communication, including empathy, clarity, trust, and patient understanding. Following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 guidelines, we searched PubMed/MEDLINE, Embase, Scopus, and Web of Science for studies published from 2020 to 2025 examining LLM or chatbot applications in clinical communication contexts. Eligible designs included randomized, observational, cross-sectional, and qualitative studies. Two reviewers (WHP and SR) independently screened titles or abstracts, assessed full texts, and extracted data on study design, population, LLM type, communication measures, and outcomes. We conducted a qualitative synthesis and random-effects meta-analysis, reporting pooled standardized mean differences or odds ratios with 95% CIs. From 312 records, 10 studies were included, all quantitative and predominantly cross-sectional. Populations ranged from patients with chronic conditions to health care professionals and laypersons. Outcomes assessed included empathy (8 studies), clarity or information quality (6 studies), satisfaction or usefulness (4 studies), and trust perceptions (2 studies). In 6 direct comparisons of AI- versus physician-generated responses, LLMs were rated significantly higher in empathy in 5 studies. One large study found that chatbot replies were judged empathetic in 45.1% of cases versus 4.6% for physician replies (odds ratio approximately 9.8, P<.001). Similarly, ChatGPT-4 answers scored higher in empathy on a 5-point scale than human-written responses (mean 4.18 vs 2.70, P<.001). One neurology study showed higher empathy scores (Consultation and Relational Empathy Scale +1.38, P<.01) for ChatGPT answers. Only 1 study found no significant empathy difference. LLM content was also longer and more information-rich, improving patient-perceived clarity and understanding. On the other hand, GPT-4 simplified pathology reports, increasing patient comprehension scores (7.98 vs 5.23/10, P<.001) and reducing consultation time by 70%. However, AI replies were sometimes less concise or less readable for low-literacy patients. In pooled analyses (k=4 studies; total evaluations N=2604), LLM assistance showed a large positive effect on empathy (standardized mean difference 1.02, 95% CI 0.44-1.60; random-effects model). Patient satisfaction results were mixed. No study directly assessed long-term trust. Current evidence suggests that LLM-based chatbots can enhance physician-patient communication by producing more empathetic, detailed, and understandable responses. These improvements may positively influence patient experience and engagement. However, LLMs may also generate overly lengthy or occasionally inaccurate advice, emphasizing the need for physician oversight. While meta-analytic findings are promising, robust randomized controlled trials, real-world and longitudinal studies are needed to confirm benefits, assess trust outcomes, and define optimal clinical integration strategies.
Coaching is a promising approach to professional development that emphasizes collaboration, goal setting, and skill building through a supportive, non-competitive relationship. Grounded in an individualized approach, coaching engages individuals in defining and pursuing goals that promote professional growth and positive health behavior change. While existing literature largely positions clinicians as recipients of coaching, far less attention has been given to how physicians might apply these principles in direct interactions with patients and families. This perspective explores how clinician-delivered coaching in pediatric primary care - particularly during well-child visits - can strengthen early relational health and language development. Core coaching competencies including trust-building, collaborative goal setting, reflective dialogue, and strengths-based feedback, offer a framework for reimagining the clinician's role as a coach. Integrating these strategies into routine care has the potential to enhance parent engagement, reinforce family-centered practice, and extend the preventative impact of primary care. Evidence from physician-led coaching in professional development and peer mentoring suggests that physicians can successfully adopt coaching roles, supporting the application of these principles within pediatric primary care.
Large language model (LLM)-powered virtual standardized patients (VSPs) enable scalable clinical skills practice, but the validity of AI-generated scores relative to faculty ratings remains unclear. To assess agreement between LLM-generated and faculty ratings of history-taking and communication performance, and to examine the influence of rater and case heterogeneity. In this cross-sectional study, 92 fourth-year medical students completed one of three 15-minute voice-based VSP cases (fever, diarrhea, cough). Ten blinded faculty raters scored performance (0-100 total; 0-50 domains). AI scores were generated by DeepSeek-V3 using a calibrated prompt. Agreement was evaluated using mixed-effects models, intraclass correlation coefficients (ICC[2,1]), Spearman correlations, mean absolute error (MAE), Bland-Altman analysis, and variance partition coefficients (VPC). Median total scores were similar for AI and faculty (93.0 [IQR 6.0] vs 94.0 [IQR 4.0]). Rater variability accounted for 37% of residual variance in faculty total scores (VPC = 0.37). AI total scores were positively associated with faculty total scores (β = 0.37, 95% CI 0.26-0.48, P<.001; Spearman ρ = 0.50, 95% CI 0.34-0.65). Absolute agreement was moderate (ICC[2,1] = 0.51, 95% CI 0.34-0.65), with MAE of 3.11 points. Mixed-effects Bland-Altman analysis showed a non-significant mean bias (1.26 points) and 95% limits of agreement from -4.95 to 7.48 (width = 12.43 points), with proportional bias (β_mean = -0.55, P<.001). Agreement was stronger for information gathering (β = 0.46, ρ = 0.49, ICC = 0.54, VPC = 0.23) than for communication (β = 0.27, ρ = 0.28, ICC = 0.29, VPC = 0.52). A sensitivity analysis in the lowest quartile showed attenuated but consistent agreement (ICC = 0.38). LLM-based scoring in a VSP showed moderate agreement with faculty ratings, performing better for information gathering than for communication. Due to rater and case heterogeneity, ceiling effects, and proportional bias, it is suitable for formative use and enhanced sampling in programmatic assessment, but not for independent high-stakes summative decisions.
Computational models of biochemical networks provide frameworks for predicting how molecular cues guide cell decisions. These models are typically limited by the time-intensive manual curation required to extract network mechanisms from incomplete literature. Here, we test whether general-purpose large language models (LLMs) can generate accurate models of signaling and metabolic networks. We find that general-purpose LLMs generate 24-65% of the reactions of literature-curated signaling networks for cardiomyocyte hypertrophy, myofibroblast activation, and mechanosignaling. Further, logic-based models based on these networks predict responses to perturbations with accuracies of 6-33%. In the context of metabolic modeling, LLMs are able to generate 64-91% of the reactions within the core Escherichia coli metabolic network and demonstrate highly variable accuracies in predicting substrate utilization. Current general-purpose LLMs generate biochemical networks with moderate accuracy, and this study provides a pipeline and benchmarks to guide future improvements.
Large language models (LLMs) have achieved impressive performance across various domains. However, the substantial hardware resources required for their training present a significant barrier to efficiency and scalability. To mitigate this challenge, low-precision training techniques have been widely adopted, leading to notable advancements in training efficiency. Despite these gains, low-precision training involves several components, such as weights, activations, and gradients, each of which can be represented in different numerical formats. The resulting diversity has created a fragmented landscape in low-precision training research, making it difficult for researchers to gain a unified overview of the field. This survey provides a comprehensive review of existing low-precision training methods. To systematically organize these approaches, we categorize them into three primary groups based on their underlying numerical formats, which is a key factor influencing hardware compatibility, computational efficiency, and ease of reference for readers. The categories are (1) fixed-point and integer-based methods, (2) floating-point-based methods, and (3) customized format-based methods. Additionally, we discuss quantization-aware training approaches, which share key similarities with low-precision training during forward propagation. Beyond efficiency, we examine robustness and deployment reliability under low precision. Finally, we highlight several promising research directions to advance this field. A collection of papers discussed in this survey is provided in Awesome-Low-Precision-Training.
Despite clinical advances, breast cancer screening adherence remains stagnant in Japan (<50%) compared with the United States (>70%). Understanding distinct cross-cultural barriers is essential; however, traditional methodologies often fail to capture visceral, real-world individual experiences and hidden deterrents to screening. This study aims to characterize and compare cross-cultural informatics profiles of barriers to breast cancer screening across Japanese-language and English-language social media discourse. We developed an automated natural language processing pipeline on a cloud-based informatics platform to analyze 46,823 screening-related posts (30,027 in Japanese and 16,796 in English) from X (formerly Twitter) collected in 2025, derived from an initial 76,955 posts after noise exclusion. The methodology integrated large language model-assisted sentiment polarity scoring with strict negation-handling, co-occurrence network topology analysis, and advanced distributional visualizations, including raincloud and ridgeline plots. Subgroup comparisons (prescreening vs postscreening and ultrasound with vs without mammography) were evaluated. Among the 46,823 screening-related posts, overall sentiment distributions showed no practically meaningful cross-cultural divergence (Cohen d=0.049); however, domain-specific analyses revealed sharp disparities in barrier prevalence. Within the English-language cohort containing 16,796 posts, discourse exhibited a concentrated, moderate negative sentiment regarding systemic barriers, featuring "Cost" as the primary barrier (n=1239, 7.4%), while "Pain" ranked considerably lower (n=842, 5.0%). In contrast, the Japanese-language cohort containing 30,027 posts was heavily bottlenecked by psychosomatic barriers, governed by a tightly interconnected network of "Pain," "Fear," and "Appointment." "Pain" emerged as the overwhelmingly dominant barrier (n=5788, 19.3%). In the English-language cohort, "Dense" breasts emerged as a prominent clinical topic due to elevated public awareness, distinct from the financial narrative. The Japanese subgroup analyses identified a temporal transition from anticipatory psychological anxiety ("Fear," 739/3805, 19.4%) and logistical concerns ("Appointment," 1556/3805, 40.9%) before screening to a strong persistence of the discomfort memory of "Pain" (1160/7419, 15.6%). Furthermore, sentiment scores for mammography were significantly more negative than those for ultrasound alone (P<.001, Cohen d=0.264), and pain-related descriptors for ultrasound spiked from 5.0% (106/2110, ultrasound alone) to 18.2% (733/4017) when performed concurrently with mammography. Despite comparable overall emotional equilibrium, a fundamental dichotomy emerged. The English-language discourse predominantly reflects US-specific systemic financial burdens, whereas the Japanese experience is characterized by emotional volatility transitioning from anticipatory anxiety to a tightly interconnected "Pain-Fear-Appointment" network. Physical discomfort of mammography dominates the screening narrative, overshadowing concurrent painless modalities like ultrasound. Improving adherence in Japan requires individualized pain-mitigating compression protocols and optimized clinical workflows to decouple mammographic discomfort from supplemental screening, thereby preventing pain-associated defensive avoidance and reducing logistical hurdles to improve equitable access.
Understanding how multiple languages are represented in the multilingual brain is critical to theories of multilingual assimilation and accommodation. Whether a second language (L2) recruits native-like networks (assimilation) or engages additional neural resources (accommodation) depends partly on the linguistic distance between L1 and L2. However, it remains unclear whether multilinguals can flexibly modulate assimilation and accommodation mechanisms when dealing with different L2 systems. Using 7 T-fMRI, we investigated brain activation patterns during a covert picture-naming task involving one-back phonological retrieval in 15 native Chinese speakers highly proficient in Japanese and English. Whole-brain univariate analysis revealed a common fronto-parieto-occipital network across all three languages (Chinese, Japanese, and English), supporting a shared core language system. However, the Japanese task elicited greater activation in the left lingual and middle occipital gyri, suggesting increased implicit visual-lexical demands. Representational Similarity Analysis (RSA) showed greater pattern similarity between Chinese and English than between Chinese and Japanese, indicating differential neural representation across L2 languages. VOI-based RSA showed a trend toward correlations between Japanese proficiency and representational similarity in the executive regions (pars triangularis and anterior cingulate gyrus), although these correlations did not reach statistical significance. These findings demonstrated that the trilingual brain relied on both assimilation and accommodation, with the level of each modulated by differences in linguistic structure and language experience. Language with a distinct writing system (English) seemed to be assimilated into the dominant L1 Chinese neural framework, whereas orthographically related language (Japanese) rather drove neural accommodation, probably due to script-driven interference. This study highlighted multi-level neural adaptation processes underlying multilingual lexical processing and provided a neurocognitive basis for tailor-made educational strategies considering each student's linguistic background.
Despite the growing number of bilingual and multilingual adults, most existing reading difficulty assessments focus primarily on individuals' first language (L1) and often overlook their literacy experiences across languages. To address this gap, this study developed the Bilingual Adult Reading History Questionnaire (Bi-ARHQ) and evaluated its psychometric properties and factor structure in 593 Hong Kong Chinese-English bilingual adults. All participants completed the self-reported Bi-ARHQ, which assesses reading, writing, and learning histories in L1 Chinese and second language (L2) English. Subsets of participants also completed Chinese (N = 456) and English word reading tasks (N = 358). Exploratory and confirmatory factor analyses supported a three-factor structure: L1 literacy difficulty (25 items), L1 literacy ability (15 items), and L2 English literacy competency (38 items). The L1 literacy difficulty and L1 literacy ability subscales significantly predicted Chinese word reading, whereas the L2 English literacy competency subscale significantly predicted English word reading. These findings indicate that the Bi-ARHQ is a reliable and valid tool for screening and identifying reading and writing difficulties in bilingual adults, particularly for Chinese-English readers.
Perspectives of families from historically marginalized groups regarding pediatric oncology clinical trial participation are not well-represented in the literature. To describe clinician- and parent-perceived facilitators and barriers to clinical trial participation. This single-center cross-sectional study with an explanatory sequential mixed-methods design enrolled parents of Black and Hispanic children with cancer as well as pediatric oncology clinicians from a large pediatric cancer center in Boston, Massachusetts. Parent participants completed single-time point surveys, and a subset, purposively sampled based on self-identified race and ethnicity, language, and household material hardship (HMH; ie, food, housing, transportation, or utility insecurity), completed semistructured interviews from September to December 2021. Clinicians completed semistructured interviews from February to March 2022. Data were analyzed from April 2022 to October 2025. Key factors influencing clinical trial participation in pediatric oncology among parents from historically marginalized groups. Quantitative data were summarized descriptively. Interview transcripts were analyzed using thematic analysis and integrated along key domains. A total of 60 parents completed the questionnaire; self-identified race and ethnicity included 5 Hispanic Black (8%), 10 Hispanic White (17%), 21 Hispanic other (35%), 21 non-Hispanic Black (35%), and 3 non-Hispanic White (5%) parents; most were mothers (51 [85%]). Twenty parents participated in interviews. Fifteen clinicians (10 [67%] female participants; 10 [67%] with ≥10 years caring for children with cancer) were interviewed, including 12 (80%) attendings and 3 (20%) advanced practice practitioners; most identified as non-Hispanic White (14 [93%]). Most families experienced HMH (44 [73%]) and reported high trust in their oncology team (mean [SD] score, 4.63 [0.65] of 5.00). Qualitatively, parents and clinicians aligned in identifying altruism and trustworthiness as facilitators to trial participation, while the informed consent discussion, non-English language preference, trial materials, and study requirements were participation barriers. Unlike clinicians, parents did not identify HMH or the experimental nature of trials as significant barriers to participation. Parents identified the desire for representation as a facilitator to participation, and clinicians identified gatekeeping as a barrier. In this cross-sectional study of pediatric oncology families from historically marginalized groups and clinicians, clinician- and parent-perceived barriers identified opportunities to increase equitable trial participation. Next steps include standardization of trial eligibility screening and systematic HMH screening and support to reduce gatekeeping.
In the article by Lim et al. entitled "The Influence of Language Dominance, Type of Language, and Narrative Task on Speech Disfluencies in Typically Fluent Bilingual English-Mandarin Children" [Folia Phoniatr Logop. 2026, https://doi.org/10.1159/000550426], there is an error where the author Chia Yi Lin's name was incorrectly displayed as Chia Y. Lin.The original article has been updated.
To provide an updated bibliometric overview of the research landscape, hotspots, and emerging trends of acupuncture for primary headaches (PH). Publications related to acupuncture for PH were retrieved from the Web of Science Core Collection (WoSCC) for the period from January 1, 2005, to December 31, 2025. Only English-language articles and reviews were included. After manual screening based on title, abstract, publication year, document type, language, and topic relevance, 286 eligible publications were included. Bibliometric analyses and visualizations were performed using CiteSpace, VOSviewer, Bibliometrix, and Microsoft Excel. The annual number of publications showed an overall upward trend, particularly after 2018. China was the leading contributor in terms of publication output and collaborative activity, followed by Germany, Italy, and the United States. Collaboration networks were identifiable at the country, institution, and author levels, although they remained concentrated among a limited number of leading countries and core research groups. The intellectual base of the field was mainly shaped by randomized controlled trials, systematic reviews, and evidence syntheses. Thematic analyses indicated that research was primarily centered on migraine, prophylaxis, efficacy, and headache-related clinical outcomes. Tension-type headache was also represented, whereas cluster headache and other less-studied PH subtypes received comparatively limited attention. Mechanism-related themes, including functional connectivity, neuroimaging, and trigeminocervical mechanisms, have emerged in recent years but remain less developed than the clinical literature. Research on acupuncture for PH has become increasingly productive, structured, and clinically evidence-oriented over the past 2 decades. However, the field remains characterized by thematic concentration, selective collaboration, and relatively limited mechanistic depth. Future studies should place greater emphasis on underrepresented PH subtypes, broader interdisciplinary collaboration, and more rigorous translational and mechanistic investigation.
The evolution of large models has witnessed the emergence of In-Context Learning (ICL) capabilities. In Natural Language Processing (NLP), numerous studies have demonstrated the effectiveness of ICL. Inspired by the success of Large Language Models (LLMs), researchers have developed Large Multimodal Models (LMMs) with ICL capabilities. However, explorations of demonstration configuration for multimodal ICL remain preliminary. Additionally, the controllability of In-Context Examples (ICEs) provides an efficient and cost-effective means to observe and analyze the inference characteristics of LMMs under varying inputs. This paper conducts a comprehensive external and internal investigation of multimodal in-context learning on the image captioning task. Externally, we explore demonstration configuration strategies through three dimensions: shot number, image retrieval, and caption assignment. We employ multiple metrics to systematically and thoroughly evaluate and summarize key findings. Internally, we analyze typical LMM attention characteristics and develop attention-based metrics to quantify model behaviors. We also conduct auxiliary experiments to explore the feasibility of attention-driven model acceleration and compression. We further compare performance variations between LMMs with identical model design and pretraining strategies and explain the differences from the angles of pre-training data features. Our study reveals both how ICEs configuration strategies impact model performance through external experiments and characteristic typical patterns through internal inspection, providing dual perspectives for understanding multimodal ICL in LMMs. Our method of combining external and internal analysis to investigate large models, along with our newly proposed metrics, can be applied to broader research areas.
Timely recognition and response to acute coronary syndrome (ACS) by emergency medical services is critical to reducing delays and improving outcomes. This study examines whether emergency medical services identification, care, and times (the interval from emergency medical services call to hospital arrival) differ by culturally and linguistically diverse (CALD) background among patients with ACS. We conducted a retrospective cohort study using ambulance data linked with the Department of Health's hospital data sets (January 2015-June 2019). The CALD group comprised individuals born in non-English-speaking countries or who preferred to speak a language other than English (LOTE). Based on preferred language, the CALD group was stratified as CALD-LOTE versus CALD-English. Of the 28 557 ACS cases, 30.2% were CALD immigrants: 9.5% CALD-LOTE and 20.7% CALD-English. Chest pain was the most common chief complaint (69.1%) but was lowest among CALD-LOTE patients (59.4%) (CALD-English=68.6%, Australian born=70.8%). Time-critical ambulance dispatch (lights and sirens) was not statistically different across groups, but paramedic ACS identification was lower in CALD-LOTE than in Australian-born patients (57.2% versus 69.0%; adjusted odds ratio=0.84 [95% CI, 0.76-0.92]). CALD-LOTE and CALD-English patients had higher rates of direct transfer to percutaneous coronary intervention-capable hospitals compared with Australian-born patients with ACS. Emergency medical services times were not statistically different in metropolitan areas but were marginally longer in rural areas among CALD patients than among non-CALD patients (67.5 versus 63.7 minutes; P=0.02). Despite lower ACS identification by paramedics, CALD-LOTE patients were more often transported to percutaneous coronary intervention-capable hospitals. Given the implications of missed early identification of ACS for outcomes, further research is needed to understand this variation and to inform tailored strategies for improvement.
Cone-beam computed tomography (CBCT) is considered the gold standard for 3-dimensional superimposition in assessing tooth movement during orthodontic treatment. However, given current guidelines emphasizing the reduction of radiation exposure, this study aimed to evaluate the accuracy, reliability, and reproducibility of a digital model superimposition technique that does not require CBCT. This retrospective study included 10 adult patients who underwent comprehensive orthodontic treatment without extractions. All patients had full-head CBCT scans and corresponding maxillary and mandibular digital dental models. Two evaluators assessed tooth movement using a new superimposition technique based on Standard Triangle Language files, using the palatal rugae as a superimposition reference for the maxillary arch and the transfer of the mandibular position through the interarch occlusal relationship. The accuracy of this method was compared with that of a CBCT-based reference workflow involving the integration of scanned dental models into CBCT images. The mean variation in reference landmark positions across the 3 spatial planes was similar between the methods for both arches. Interevaluator differences (Δ) ranged from -0.07 to 0.06 mm, which were not considered clinically significant. Student t tests showed no statistically significant differences between evaluators for any axis in either method (P >0.05). Intraclass correlation coefficients demonstrated high agreement between the 2 methods, ranging from 0.88 to 0.98. The 3-dimensional digital model superimposition technique using the palatal rugae for the maxillary arch and the interarch occlusal relationship for the mandibular arch is a reliable and reproducible alternative to CBCT-based methods for evaluating orthodontic tooth movement in nongrowing patients treated without extractions.
10-30% of the world's population suffers from insomnia. One-quarter of those who take hypnotic medication become dependent on it. Given the wide variety of available drugs, switching them is complex. The existing recommendations do not cover all drug classes. There is a need for practical, evidence-based recommendations that take account of specific aspects of care both in Europe as a whole and regionally. This narrative review is based on publications from May 1953 to February 2026 that were retrieved by a search in six English- and German-language databases, including guidelines, meta-analyses, systematic reviews, randomized controlled trials, and observational studies. Following the literature review, new recommendations were developed in a consensus process (consensus conference followed by the Delphi method) to supplement the current insomnia guideline. Treatment can be discontinued abruptly after short-term use (1-2 weeks) or if the drug in question does not produce withdrawal symptoms, e.g., daridorexant (an orexin receptor antagonist), melatonin, antihistamines, and phytotherapeutic drugs. Benzodiazepine receptor agonists, antidepressants, antipsychotics, and gabapentinoids should be tapered off gradually after medium- or long-term use (> 2 weeks), with a weekly dose reduction of 10-25%. Various gradual tapering methods can be used, depending on the drug, dose, duration of use, treatment regimen, and intended subsequent treatment. A longer taper is advisable after use at high doses (≥ 2/3 of the maximum dose) or over the long term (≥ 1 year). Withdrawal symptoms from benzodiazepine receptor agonists can be alleviated with the overlapping administration of daridorexant or eszopiclone. Behavioral therapy is recommended as an accompanying nonpharmacological treatment (relative risk [RR]: 1.68; 95% confidence interval: [1.19; 2.39]). A drug-specific approach with evidence-based and practice-oriented protocols should be used when hypnotic medications are switched or discontinued.