Metabolic reprogramming is a substantial obstacle for anticancer drug screening, as targeted therapeutics often lose efficiency due to the dynamic adaption of cancer cells. Glutamine metabolism in cancer profoundly impacts tumor initiation, progression and metastasis. The existing agents are compromised by resistance and off-target toxicity. In this study, a real-time NMR tracking method for intracellular glutamine metabolic flux was established. This method enables comprehensive profiling of nitrogen metabolism and serves as a valuable tool for characterizing specific cancer metabolic phenotypes and screening drugs against targeted cancer cells. Applying this approach to traditional Chinese medicine (TCM) discovery, we identified Astragalus membranaceus as a potent regulator of glutamine metabolism. Through virtual screening via molecular docking, 12 potential compounds from Astragalus membranaceus were initially flagged as candidate binders toward the allosteric pocket of glutaminase 1 (GLS1). Crucially, subsequent in vitro recombinant human GLS1 enzyme activity assays successfully ruled out computational false positives and demonstrated that Compound 2 (quercetin) acts as the exclusive, direct enzymatic inhibitor among the tested monomers, capable of effectively suppressing GLS1 activity. Overall, this work provides a robust platform for real-time metabolic profiling of glutamine metabolism and drug screening at the living cell level, and offers new insights into the mechanisms of TCMs in anticancer therapy.
BeanGPT is a domain-specific retrieval augmented generation system designed to support research and breeding decisions in common bean (Phaseolus vulgaris L.) by transforming natural language questions into citation-backed, verifiable answers. The platform integrates a large, curated corpus of legume-focused peer-reviewed literature with structured multi-year agronomic trial records collected across diverse environments, climate projections extending to 2090 under multiple emission scenarios, and standardized cultivar nomenclature to resolve naming inconsistencies across datasets and publications. BeanGPT combines semantic retrieval from a vector database with intent-based query routing and structured parameter extraction to direct questions to genetics, field performance analytics, or climate modules. To reduce errors that commonly occur in general-purpose language models, BeanGPT incorporates a genomic index that enables constant time membership lookup of gene and protein identifiers against authoritative resources, ensuring that molecular entities are either validated or clearly flagged as literature-derived. The system is implemented with a streaming web interface and an asynchronous backend that supports concurrent users and can generate interactive visualizations through automated Plotly code generation. Beta testing demonstrated strong retrieval relevance, low response latency, reliable gene verification, and high citation precision, indicating that domain-grounded RAG can improve accuracy and usability for Phaseolus vulgaris research.
Helicobacter (H.) pylori is characterized by a high degree of genomic diversity, with regional differences in virulence determinants. This study aims to explore genomic composition, phylogeography and accessory-gene relationships of Iraqi H. pylori isolates in a global contextualized dataset. A total of 198 H. Pylori genomes were reviewed, including 41 isolates sequenced from gastric samples of patients undergoing diagnostic endoscopy at Al-Yarmouk Teaching Hospital, Baghdad from June 2024 to February 2025. Illumina MiSeq was used to sequence genomes, which were quality-filtered and assembled using SPAdes. Prokka was used to perform annotation and Roary to infer pan-genome structure. FastTree was used to reconstruct core-genome phylogeny. Anatomical micro-niche (corpus vs. other gastric sites) were explored with pan-genome-wide association study (pan-GWAS) with Pyseer (linear mixed model, kinship based on the core alignment). The H. Pylori pan-genome showed 955 core gene families and 9,400 accessory genes. Isolates from Iraq were polyphyletic, mixing with European and Middle Eastern lineages. In the primary Pyseer linear mixed-model analysis, Benjamini-Hochberg correction across 3,259 valid lrt-pvalue tests identified 151 FDR-significant associations; after excluding rows with problematic Pyseer diagnostic notes, 70 unflagged loci remained significant. The strongest unflagged positive association was group_2036, whereas a co-occurring block including cagS, cagT, and virB4_1 was strongly depleted in corpus-derived isolates. These signals implicate accessory-genome variation in gastric micro-niche adaptation while also underscoring the need to interpret flagged Pyseer rows cautiously. The genomic variation of squamous H. Pylori isolates in Iraq corresponds to the global recombination trends and to the regional admixture. The corpus sampling-related accessory-gene cluster implies possible micro-niche adaptation. These preliminary results highlight the necessity of large, stratified Middle Eastern cohorts and long-read sequencing to dispel functional genetic constructions of tissue tropism and virulence.
Many therapists and counselors are not well informed about how political stress affects clients' lives, especially in its nonviolent forms. Capitalizing on Israel's ongoing judicial reform/overhaul, we compared the effects of this political stressor on depression and anxiety with those of childhood and adult stress. A nationally representative sample of Israeli-Jewish adults (N = 1,202) was recruited in August 2023, at the peak of the judicial reform/overhaul, via an online platform. Participants indicated their support (28%), opposition (52%), or deliberation about ("don't-knowers"; 20%) the reform/overhaul and completed measures of depression/anxiety, negative and positive affect, and childhood and adult stressful events. Structural equation modeling and logistic regression analyses were conducted to predict continuous and discrete levels of the outcomes. Opposition contributed 5.5%, 2%, 1.3%, 2.2%, and 2% to the variance of continuous negative and positive affect, depression, and anxiety and was associated with 61% and 104% increased risk for binary (flagged) depression and anxiety compared with support. Deliberation contributed 1.4%, .5%, .9%, 1.7%, and 1.5% to these continuous outcomes and was associated with 59% and 107% increased risk for flagged depression and anxiety. Deliberation, but not opposition, predicted an increased risk for flagged anxious depression. Except for many adult stressful life events, the effect of political stress on depression and anxiety was stronger than, or comparable to, the effects of childhood and adult stress. Political stress, even in its nonviolent form, is a clinical risk factor that should be addressed by therapists and counselors. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
To evaluate the effectiveness of an intelligent clinical decision support system (CDSS) for neonatal hypoglycemia management in mother-infant rooming-in settings, and to dissect the differential hypoglycemia risk conferred by individual high-risk factors and their specific combinations under standardized surveillance. A multidisciplinary team developed a knowledge-driven CDSS grounded in national expert consensus, integrating automated maternal-neonatal risk identification, dynamic tiered monitoring reminders, and structured stratified management recommendations. Effectiveness was assessed using a pre-post self-controlled analysis (historical control: January-March 2024, n = 522; CDSS-implemented: April-June 2024, n = 417) and a concurrent parallel controlled analysis (non-CDSS wards: n = 389; CDSS wards: n = 352). Neonatal hypoglycemia was defined as blood glucose <2.2 mmol/L. Risk factor combination patterns were explored among 6,667 system-flagged high-risk neonates. CDSS implementation significantly reduced hypoglycemia incidence in both the pre-post (5.76% vs. 11.88%, P < 0.05) and parallel (5.40% vs. 9.25%, P < 0.05) analyses. Under CDSS-managed surveillance, the overall hypoglycemia incidence in the high-risk cohort was 6.3%. Marked heterogeneity was observed: preterm birth (15.4%) and low birth weight (25.0%) carried the highest independent risks, while risk escalated non-linearly with specific factor combinations, reaching 18.8% in neonates with five concurrent factors. Serial monitoring demonstrated a sharp decline in hypoglycemia incidence from 6.1% at first measurement to ≤0.4% thereafter. The intelligent CDSS effectively reduces neonatal hypoglycemia in rooming-in settings. Hypoglycemia risk depends more on the specific types and combinations of high-risk factors than on their quantity alone, providing evidence for precise risk stratification. This closed-loop, guideline-driven workflow enhances clinical standardization and patient safety. Future multicenter studies incorporating machine learning and long-term neurodevelopmental follow-up are warranted.
Response to neoadjuvant PD-1 inhibitor plus chemotherapy in locally advanced esophagogastric adenocarcinoma varies widely, and tumor biomarkers alone do not fully explain the variation. We developed an integrated Body Composition and Immunonutritional Signature (BCIS) from pretreatment CT body composition and routine blood markers, and tested whether it predicts pathological response, immune-related adverse events (irAEs), and survival. From four tertiary hospitals in Hebei Province, China, we enrolled 720 patients with histologically confirmed gastric or gastroesophageal junction (Siewert II/III) adenocarcinoma treated between 2019 and 2023 with neoadjuvant PD-1 inhibitor plus SOX or XELOX followed by D2 gastrectomy. BCIS was a 0-10 additive score from ten prespecified adverse host features. Analyses included logistic regression, restricted cubic spline (RCS) modeling, multivariable Cox regression, DeLong tests for nested AUC comparisons, decision-curve analysis, eleven sensitivity scenarios, and collinearity checks by Spearman correlation and variance inflation factors. BCIS classified 243 (33.8%) patients as favorable, 303 (42.1%) as intermediate, and 174 (24.2%) as unfavorable. Major pathological response (MPR) rates dropped sharply across strata (64.2%, 39.6%, 21.8%; P-trend<0.001). After adjusting for age, sex, cT/cN, PD-L1 CPS, MMR status, EGJ origin, and regimen, every BCIS point cut the odds of MPR by roughly 40% (adjusted OR = 0.61, 95% CI 0.55-0.68, P<0.001) and pCR by close to half (adjusted OR = 0.55, 95% CI 0.46-0.66, P<0.001). RCS modeling flagged CRP and CAR as nonlinear; BCIS itself was strictly linear. Adding BCIS to a clinical baseline lifted overall AUC for MPR from 0.597 to 0.710 (DeLong P<0.001) and external AUC from 0.588 to 0.704. At 42.5 months' median follow-up, each BCIS point raised the hazard of progression by 53% (adjusted HR = 1.53, 95% CI 1.44-1.63) and of death by 38% (adjusted HR = 1.38, 95% CI 1.29-1.47; both P<0.001). Direction of effect held across every prespecified subgroup (P-interaction>0.10 throughout) and all eleven sensitivity scenarios. BCIS is independently associated with pathological response and survival under neoadjuvant PD-1-based immunochemotherapy and adds discrimination beyond tumor-centered biomarkers and clinical staging. Because every component comes from routine pretreatment workup at no extra cost, BCIS could feasibly inform prehabilitation, toxicity surveillance, and shared decision-making.
Health-related quality of life is a key secondary end point in stroke trials. Differential item functioning (DIF) occurs when individuals with the same underlying health-related quality of life interpret and respond differently to questionnaire items, potentially biasing treatment comparisons. This study evaluates DIF in the patient-reported 5-level EuroQOL questionnaire among patients with acute ischemic stroke across age, sex, and treatment groups. Data were from the AcT trial (Alteplase Compared to Tenecteplase), a registry-based randomized comparison of alteplase and tenecteplase conducted at 22 stroke centers across Canada (December 2019-January 2022). Patients with acute ischemic stroke presenting within 4.5 hours of symptom onset and eligible for thrombolysis completed the 5-level EuroQOL questionnaire at 90 days poststroke. DIF was assessed using multigroup graded response models with the Wald-based sweep procedure, which accounts for between-group differences in latent trait distributions. We quantified effect sizes using signed weighted area between curves (sWABC); |sWABC| <0.10=negligible. Of 1577 patients enrolled in the trial, 1264 survived to 90 days with complete 5-level EuroQOL questionnaire data (51.2% tenecteplase; 46.5% female; 30.1% aged ≥80). Omnibus testing revealed significant DIF only for age (χ2=86.9, P<0.001); neither sex (χ2=31.7, P=0.063) nor treatment (χ2=22.4, P=0.379) showed evidence of DIF. Four items flagged for age-related DIF: self-care, usual activities, pain/discomfort, and anxiety/depression. However, only self-care (sWABC=-0.46) and usual activities (sWABC=-0.34) showed moderate effects, while pain/discomfort (sWABC=-0.002) and anxiety/depression (sWABC=0.09) were negligible. Importantly, factor scores from models with and without DIF adjustment correlated (correlation coefficient=0.98). The 5-level EuroQOL questionnaire appears to function equivalently across sex and treatment groups in this stroke population. Age-related DIF, though statistically detectable in physical functioning items, had little practical consequence for individual scores, supporting the instrument's use for health-related quality of life comparisons in stroke trials. URL: https://www.clinicaltrials.gov; Unique identifier: NCT03889249.
Long-term changes in forest management are documented across reports, plans, and scientific papers written for different purposes and with changing vocabularies. This makes it difficult to show how a documentary record was converted into a temporal claim. FORM-TRACE is a formula-based workflow that records corpus decisions, extraction quality, domain terms, and calculations before interpretation. We demonstrate it with Harvard Forest and New England documents dated 1908-2026. Of 257 PDFs inspected, 215 met the analytical criteria; 201 were extracted and scored, 14 were flagged as unreadable, scanned, or corrupted, and 42 methods-support references were kept outside the scored corpus. The method provides: • A corpus manifest and extraction log that expose inclusion, exclusion, and coverage gaps; • A keyword-domain matrix and five numbered equations that produce document- and period-level indicators; • Saved score tables, plot data, and validation records that allow independent checking without redistributing copyrighted PDFs. FORM-TRACE measures documented attention rather than management performance and keeps interpretation separate from scoring.
The extent and characteristics of Large Language Model (LLM) utilization in arthroplasty literature remain undefined. In this study, we aimed to quantify the extent of LLM utilization in manuscripts across major arthroplasty journals. Additionally, we sought to assess temporal trends in LLM utilization, as well as associations with author productivity, geographic origin, and citation impact. A cross-sectional analysis of 3,352 original research articles from six arthroplasty journals from the era before advanced Large Language Models (LLMs) (Pre-AI) (2018 to 2022) and the era after advanced LLMs (post-AI) (2023 to 2025) was performed. The text was processed using a detection algorithm. Journal-specific thresholds for significant AI involvement were established (mean + two standard deviations of pre-AI scores). Author productivity, primary language of affiliate country, and citation counts were analyzed. In the post-AI era, five of six journals demonstrated significantly higher odds of AI involvement (P < 0.05). The proportion of AI-flagged articles rose from less than 4.2% (2018 to 2021) to 20.4% in 2025, which was a notable nonlinear increase compared to 2024. The first authors in the 90th percentile of the dataset for authorship demonstrated significantly greater odds of exceeding AI thresholds compared to authors who only had one publication in the dataset (odds ratio (OR) = 1.83, P < 0.001). Conversely, non-English affiliate country authors (P = 0.436) and zero-citation-count articles (P = 0.882) did not have higher odds. Unsurprisingly, detectable AI assistance in arthroplasty research has increased significantly since the public release of LLMs. The AI tools are disproportionately utilized by high-productivity authors but not non-English-speaking country authors, suggesting adoption is driven by research efficiency and scalability rather than language barriers. Higher rates of LLM use were not found in zero citation articles.
Multivariable regression tables are common in observational clinical research, but their coefficients are often over-interpreted. A model built to estimate the effect of 1 exposure may also report coefficients for age, sex, comorbidities, behaviours, and other adjustment variables. These additional rows are frequently read as independent risk factors, even when the analysis was not designed to estimate their effects. This is the Table 2 fallacy. The problem is not the use of adjustment, but the interpretation of adjustment terms as if each were a separate causal estimate. In this commentary, we use a directed acyclic graph and a single worked example to show why the coefficient for the exposure of interest can answer the intended clinical question, while coefficients for adjustment variables may not represent clinically actionable effects. We also show how similar errors arise when interaction terms are interpreted as causal. Authors should specify the target estimand (the causal effect the analysis is designed to estimate) and exposure, choose adjustment variables from the assumed causal structure, interpret only the exposure coefficient, and fit a separate model for each further question. We set out red flags, recurring pitfalls, and questions to ask of any risk-factor table. A regression table is not a menu of modifiable risks. An association is clinically actionable only when the study was designed to support that interpretation.
PSMA PET has been proposed for guiding targeted prostate biopsies. As the rate of clinically significant prostate cancer (csPCa) is positively correlated with PSA density (PSAd), a potential application for PSMA-guided biopsies may be men with negative MRI-first (PI-RADS-1-2) but increased PSAd. The PRIMARY scoring system, using the intraprostatic PSMA-pattern was developed and tested using [68Ga]Ga-PSMA-11 and proposed to be universal across PSMA-ligands, however no data on [18F]PSMA-1007 for biopsy guidance exists. Thus, the aim of this study was to evaluate the PRIMARY score for prostate biopsy guidance using [18F]PSMA-1007 in a prospective clinical setup. 91 patients were prospectively enrolled for [18F]PSMA-1007 PET/MRI, 86 of these for biopsy guidance after negative MRI or non-csPCa biopsies from PI-RADS 3-5 lesions and persistent suspicion of csPCa. Transperineal MRI/ultrasound fusion target biopsies were performed predominantly from PRIMARY 3-5 lesions. Nine patients with PRIMARY 1-2 did not undergo target biopsies. CsPCa was defined as ISUP grade ≥ 2. In a per patient analysis, 23.3% (20/86) were diagnosed with csPCa. In line with previous studies, the rate of csPCa increased with higher PRIMARY score in a per lesion analysis, constituting 0% (0/1), 0% (0/8), 7.2% (7/97), 60% (15/25), and 60% (6/10) for PRIMARY 1, 2, 3, 4, and 5 lesions, respectively. The 7.2% PRIMARY-3 lesions containing csPCa exhibited common characteristics such as anterior location or lack of well-defined separation from the peripheral zone. We evaluated the PRIMARY score for use with [¹⁸F]PSMA-1007 and identified benign features and potential "red flags" for PRIMARY-3 lesions. We found csPCa in 23.3% of patients with negative MRI or non-csPCa biopsies from PI-RADS 3-5 lesions and persistent suspicion of csPCa.
Whereas correlates of cyberbullying have been studied extensively, there has been comparatively less work examining predictors of the different ways in which bystanders to cyberbullying might respond. Adopting a person-situation interaction approach, this study investigated the extent to which the Big Five personality traits and severity of cyberbullying interactively predict the likelihood of different cyberbystander behaviors. Adults in the U.S. (N = 303) took part in an online survey in which they were presented with a series of nine simulated social media interactions in the form of screenshots that involved exchanges between two social media users. Each screenshot depicted one of three distinct levels of cyberbullying severity: none, low severity, and high severity. For each screenshot, participants were asked to report the likelihood that they would respond in a range of ways as a bystander, including remaining a passive observer, confronting a bully, reinforcing a bully, supporting a victim, and flagging or reporting a post. Participants then completed a measure of the Big Five personality traits. Regardless of cyberbullying severity, participants were significantly more likely to indicate that they would remain a passive observer in response to the depicted social media interactions than any other cyberbystander behavior. Informative two-way interactions did, however, emerge between cyberbullying severity and cyberbystander behavior and between these variables and Big Five traits across a series of mixed effects models. A significant three-way interaction emerged for agreeableness, such that participants higher in agreeableness reported a greater likelihood of bystander action in response to high severity cyberbullying than those with moderate or lower levels of agreeableness. This research offers support for the predictive value of both individual differences in the Big Five personality traits and cyberbullying severity for understanding diverse forms of cyberbystander behavior.
The COVID-19 pandemic highlighted the ability of epidemics to evolve through the emergence of successive strains of greater infectiousness, which prompted the insight that hyper-exponential growth (HEG) can arise in the development of an epidemic. The phenomenon of HEG has intrigued many researchers, because of some radical differences from exponential growth such as the finite-time singularity. However, the actual mechanism of HEG was usually hidden or needed to be added phenomenologically. Thus, there was little insight into how constraints would be triggered. In this study, we explore an SIR model with evolving parameters leading to a discrete sequence of variants. This allows us to consider the HEG phenomenon in greater depth and with greater mathematical rigour. The model yields a mechanistic description of what happens at the collapse of HEG. The model also yields closed expressions for important features, such as the critical time to the singularity, in terms of basic demographic parameters. Our analysis flags a few important issues needing further research, such as the stochastic character of the emergence of variants. Greater understanding of the HEG process will yield dividends in other fields, since modern societies exhibit HEG at several levels, such as human population growth, economic indicators and technological innovation. For this reason, more research on HEG remains imperative, especially on HEG in the presence of limited resources.
ICU-to-ward transfers are high-risk transitions marked by information loss and burdensome handoff preparation. We developed PAUSE-Agents, a clinician-in-the-loop multi-agent LLM pipeline that drafts source-attributed handoff briefs from structured ICU data and clinical notes using the clinician-developed ICU-PAUSE template. Mirroring ICU team structure, PAUSE-Agents routes each record through a scribe extractor, 6 role-specialized agents, explicit conflict surfacing, and deterministic safety checks before synthesis, producing an editable first draft rather than an autonomous note. In a single-center medical ICU cohort, 5 physicians completed 100 reviews of 84 agent-drafted briefs. Among adjudicable claims, 98.8% were verified and 1.2% were incorrect; 88% of briefs had no pertinent omission, and mean PDSQI-9 quality was 4.20/5. PAUSE-Agents surfaced 118 conflict warnings and 421 safety flags, making documentation inconsistencies visible before handoff. An o4-mini PDSQI-9 judge showed limited case-level discrimination but supported aggregate monitoring. We release PAUSE-Agents and its clinician evaluation application.
Samara Oblast is an ecological transition zone in Russia, yet its spotted fever group Rickettsia (SFGR) diversity and prevalence in Dermacentor ticks remain unexplored. This study characterized SFGR species circulating in Dermacentor reticulatus and Dermacentor marginatus across 15 districts of Samara Oblast. We collected 681 adult Dermacentor ticks via vegetation flagging during 2023-2025. SFGR screening was performed via a commercial qPCR kit. Genospecies were identified via Sanger sequencing of the partial gltA gene and validated by means of ompB analysis, while tick species were confirmed using a cox1 gene fragment. The overall SFGR prevalence was 33.5% (228/681). Infection rates were significantly higher in D. marginatus (44.7%) than in D. reticulatus (27.5%). Rickettsia raoultii predominated (94.4%), while Rickettsia slovaca was relatively rare (4.6%) and significantly associated with D. marginatus. Notably, partial ompB sequencing revealed two distinct R. raoultii putative genotypes circulating in the region: one with a 9 bp deletion and one without. No coinfections were detected. Unexpectedly, a single D. reticulatus tick tested positive for Rickettsia felis, confirmed by means of gltA and ompB, marking the first molecular detection of R. felis in Russia and globally in this tick species. Samara Oblast represents a high-prevalence sympatric zone for Dermacentor-borne SFGR. The molecular detection of R. felis in D. reticulatus suggests a potential association that requires confirmation through further studies.
Hypospadias is a common congenital malformation requiring surgery. Caregivers face substantial perioperative information needs, and large language models (LLMs) offer a potential health education channel, but their performance in pediatric urology and the relation between citation accuracy and clinical content safety lack systematic evaluation. This study aimed to evaluate 5 LLMs (ChatGPT-4o, Gemini-2.5-Pro, OpenEvidence, Zhipu Qingyan, and DeepSeek) for pediatric hypospadias perioperative consultation, and to characterize the dimensions clinicians and caregivers prioritize. A noninterventional cross-sectional study was conducted at a tertiary hospital in April 2025. From a 40-item bank, 10 high-priority questions were selected by an independent caregiver screening cohort (N=34, cohort A) and classified into 3 risk levels. Twenty-three pediatric urology experts (6 dimensions) and 36 primary caregivers (cohort B; 4 dimensions) evaluated responses by double-blind forced-ranking (reverse-scored, 5=best). Friedman tests with Kendall W assessed overall differences; paired Wilcoxon tests with Bonferroni correction (adjusted α=.005) and rank-biserial r with Hodges-Lehmann 95% CIs were used post hoc. Reference authenticity was independently verified by 2 reviewers (XH and WH) using a 5-category scheme (V/PV/F/G/NR [V: Verifiable, PV: Partially Verifiable, F: Fabricated, G: Guideline-Based, Nonspecific, and NR: No References]), with consensus after canonical-source reverification (Cohen κ=0.702 preadjudication). An 8-reviewer clinical safety audit (7 senior specialists plus 1 European Association of Urology [EAU]-anchored intermediate-title clinician) applied a 4-level severity scheme (None/Mild/Moderate/Severe). Models differed significantly (caregiver: χ²4=77.5, W=0.538, P<.001; expert: χ²4=62.2, W=0.676, P<.001). Gemini-2.5-Pro ranked first (expert median 5.0, IQR 3.0-5.0; caregiver 4.0, IQR 3.0-5.0). DeepSeek ranked second (4.0 both), with superior Empathy versus ChatGPT-4o (r=-0.343; P<.001). OpenEvidence scored lowest (2.0 both), despite high citation accuracy. Expert-caregiver agreement was strong (Spearman ρ=0.89; P=.04). Citation accuracy diverged sharply: OpenEvidence was fully verifiable (V=100%, F=0%), whereas DeepSeek and Zhipu Qingyan showed the highest fabrication (F of 33% and 24%, respectively); Gemini-2.5-Pro fabricated none but used nonspecific guideline citations (G=85%). The safety audit yielded 78 flags, including 9 Severe-level flags across 5 question-model combinations; OpenEvidence carried the largest Severe burden (5 of 9) and the highest severity-weighted score, whereas Gemini-2.5-Pro had the lowest. Bibliographic accuracy and clinical safety were dissociable, and the ranking held under poststratification weighting. High citation accuracy does not guarantee clinical safety. In the first dual-perspective evaluation, no model was uniformly best. Gemini-2.5-Pro was most comprehensive but relied on nonspecific guidelines. DeepSeek scored highest on caregiver-rated Empathy, yet it had the highest fabrication rate. OpenEvidence produced the most verifiable citations but carried the heaviest Severe-flag burden. These dimension-level priorities, the dissociation between citation quality and safety, and the portable evaluation framework can inform future pediatric medical-artificial intelligence (AI) development. For perioperative use, AI should follow a tiered human-machine collaboration model with mandatory clinician oversight in high-risk scenarios.
Spinal metastases may progress to debilitating pain, spinal instability, and neurological deficits. Timely referral is essential, yet delays are common because patients often first present to non-spine clinicians where red flags rarely expedite referral and guidelines primarily target spine specialists. We aimed to develop a staging-based referral tool to support non-spine clinicians in recognizing progression and guiding referral urgency. We defined the Spinal Metastasis Staging (SMS) system as four stages: SMS I, asymptomatic; SMS II, inflammatory pain; SMS III, mechanical pain and/or spinal instability; and SMS IV, neurological deficits and/or high-grade spinal cord compression. Stages were translated into a referral algorithm organized by urgency and presented as a pocket map. The tool was refined through regional and international multidisciplinary expert panels, and feasibility was evaluated in an international survey. Panels endorsed the four-stage SMS system and referral algorithm. Among all survey respondents (n = 120), high acceptability was reported. Among non-spine clinicians (n = 32), 94% found the tool easy to understand, 91% considered the format suitable for clinical use, and 91% anticipated improved referrals. Overall, 88% would use the tool at least occasionally, including 55% who would use it frequently or always. The SMS staging system and referral tool (link) was rated feasible by expert panels and survey respondents. However, only 32 of 120 survey respondents (27%) were non-spine clinicians, so findings in this group are preliminary and may overstate acceptance. The tool should be considered provisional: prospective studies are needed to validate effects on referral and patient outcomes.
Psychiatric disorders represent a major burden for patients with epilepsy (PwE). This study examined how demographic, epilepsy-related, and clinical risk factors contribute to the onset of mood and psychotic disorders in epilepsy. In this retrospective cohort, we intersected a hospital structured database with results from direct text mining in all available documentation from each patient to include PwE without prior psychiatric history. Logistic regression evaluated age, sex, epilepsy type, treatment resistance, histories of status epilepticus and febrile seizures, neurological and neurodevelopmental comorbidities as risk factors of psychotic and mood disorders. A five-fold cross-validated logistic regressor tested their predictive performance. Pairwise comparisons examined psychiatric outcomes across monotherapy treatments. We screened 4269 individuals ; 1709 were excluded for lack of an established epilepsy diagnosis or anti-epileptic medication ; 696 already had a psychiatric history at the time of first antiseizure medication. In 1864 PwE without prior psychiatric history, psychiatric disorders occurred in 25% of cases (15.8% mood, 2.9% psychosis, 3.4% mood and psychosis). Risks for mood disorders included female sex, focal epilepsy, drug resistance, a history of febrile seizures, neurological comorbidities, and neurodevelopmental disorders. Risk factors for psychotic disorders included drug resistance, a history of febrile seizures and neurodevelopmental disorders. Supervised analyses significantly predicted outcomes for mood disorders (AUC = 0.70, p < 0.001) and psychosis (AUC = 0.61, p < 0.001). Pairwise treatment comparisons found no significant differences. Identification of these risk factors to psychosis and mood disorder onset provides clinicians with red flags that will facilitate timely psychiatric referral and intervention, as well as integrated care between neurologists and psychiatrists.
Background: A substantial proportion of clinical data is stored in unstructured text, limiting its utility for evidence generation. Small language models (SLMs) can extract information but face hallucination risks and privacy concerns. Locally deployable SLMs are needed for secure and reliable healthcare use. Methods: We designed an end-to-end pipeline integrating clinical text extraction with stroke outcome prediction. Records of 1,398 patients screened for ischemic stroke were reviewed, with 1,166 included. A Llama 3 8B SLM was fine-tuned using low-rank adaptation and 4-bit quantization for local feasibility, guided by few-shot prompting. Multi-tiered validation-rule-based checks, retrieval-augmented generation, cosine similarity flagging, and human-in-the-loop review-was implemented. Structured data from 767 patients were used to train models predicting poor outcomes at 3 months. Results: Baseline extraction accuracy was 64.9% (95% CI, 62.0% to 67.8%), improving to 86.0% after fully automated multi-tiered validation, and further to 97.0% (95% CI, 95.7% to 98.3%) following human-in-the-loop review. Template-based variables achieved F1 > 0.90 (95% CI, 0.88 to 0.96). Narrative extraction reached F1 = 0.87 (95% CI, 0.84 to 0.90). NIHSS scores were extracted with a mean absolute error of 0.853 (95% CI, 0.791 to 0.915). TabPFN achieved an AUROC of 0.816 (95% CI, 0.784 to 0.847) with good calibration, confirming reliable risk stratification. Conclusion: This study demonstrates a privacy-preserving, efficient pipeline for clinical text processing. By combining an SLM with multi-tiered validation and predictive modeling, it offers a proof-of-concept solution with potential for broader deployment to transform unstructured records into structured data suitable for stroke outcome research and decision-support modeling.
ObjectiveSymptom-based screening alone is insufficient for identifying perinatal mental health risk in general hospital settings. Although Edinburgh Postnatal Depression Scale (EPDS)-centred screening is embedded in Australian policy, psychosocial adversity and somatic illness may contribute to psychiatric morbidity that is not captured by symptom screening alone. This perspective proposes a consultation-liaison framework that augments, but does not replace, universal screening.Intimate partner violence and hyperemesis gravidarum are associated with substantial psychiatric morbidity that may not be fully reflected by EPDS scores, yet are inconsistently captured by symptom-based screening.ConclusionThe Triangulated Perinatal Risk Assessment (TPRA) framework is a hypothesis-generating consultation-liaison framework that builds on established psychosocial risk-assessment approaches, including the Antenatal Risk Questionnaire and Psychosocial Risk Assessment Model. TPRA operates after initial risk flagging and integrates symptom assessment, psychosocial risk, and somatic burden to support CL-stage formulation and care allocation. Where perinatal liaison nursing capacity exists, TPRA may provide a pragmatic structure for organising information already gathered in routine care.TPRA may be relatively low-burden and potentially sustainable within existing consultation-liaison structures, but its feasibility, workload, referral burden, false-positive burden, and clinical impact remain unevaluated and require prospective study.