共找到 20 条结果
This study aimed to assess whether multimodal large language models (LLMs) can distinguish cholesteatoma from non-cholesteatomatous chronic otitis media (COM) on representative key-image temporal bone high-resolution computed tomography (HRCT) and to evaluate the short-interval reproducibility of their outputs. This retrospective, single-center study (2019-2024) included 101 patients (48 with cholesteatoma, 53 with non-cholesteatomatous COM) who underwent surgical treatment. The reference standard was intraoperative diagnosis with histopathological confirmation for cholesteatoma and surgical documentation for COM. For each case, six anonymized representative HRCT images reflecting standard diagnostic criteria were selected by consensus between a 4th-year radiology resident and a board-certified head and neck radiologist, both unaware of the diagnosis. Subsequently, the same images were analyzed by the GPT-5 and Gemini 2.5 Pro LLMs through their official web interfaces, utilizing structured prompts and a zero-shot approach. These evaluations were conducted in two distinct sessions (S1 and S2) with a 1-week interval. The primary endpoint was accurate binary classification. Accuracy, sensitivity, specificity, positive predictive value, and negative predictive value were calculated with 95% confidence intervals [(CIs); Wilson method] agreement with the reference standard and between sessions and models was assessed with the Cohen kappa (κ) coefficient; and differences in classification were assessed with the McNemar test. The radiologist achieved an accuracy of 96.0% (95% CI: 90.3-98.4) with almost perfect agreement with the reference standard (κ: 0.921). In S1, GPT-5 and Gemini 2.5 Pro achieved accuracies of 43.6% and 49.5%, and in S2, 46.5% and 47.5%, respectively. Both models combined high sensitivity (83.3%-97.9%) with low specificity (1.9%-13.2%), and balanced accuracy ranged from 0.45 to 0.52. Between-session reproducibility was fair for GPT-5 (κ: 0.360) and moderate for Gemini 2.5 Pro (κ: 0.485), and inter-model agreement was slight at both sessions (κ: 0.035 at S1 and κ: 0.086 at S2). Accuracy did not differ significantly between the two models (P = 0.211). In this single-center study, GPT-5 and Gemini 2.5 Pro, in the versions evaluated, combined high sensitivity with low specificity and showed only fair-to-moderate between-session reproducibility and exhibited slight inter-model agreement on temporal bone key-image HRCT. These findings do not support their use as independent second readers, and broader generalization to other multimodal LLMs would require the evaluation of additional models. The evaluated models lacked the spatial precision and consistency required for the accurate assessment of complex middle ear structures. This finding underscores the necessity for verification by a radiologist and continuous monitoring.
Trauma-related hemorrhage remains a leading cause of preventable death, requiring rapid diagnosis and timely intervention. Interventional Radiology (IR) plays a central role in the management of non-compressible bleeding, especially in solid organ injuries and pelvic trauma. This article outlines three key recommendations for integrating IR into trauma care. First, IR must be embedded in trauma teams with 24/7 availability at Level I trauma centers and structured access at Level II and III centers. Second, whole-body contrast-enhanced CT should be performed in hemodynamically stable or initially unstable but responsive patients, with immediate embolization when active extravasation or pseudoaneurysm is identified. Third, standardized embolization protocols and immediate access to essential materials-such as coils, plugs, liquid embolics, and stentgrafts-are critical for effective bleeding control. These recommendations are supported by current European guidelines and selected observational studies. To implement this guidance, trauma centers should develop IR-inclusive algorithms, define access pathways, and maintain trauma-ready IR inventories. Close collaboration between radiologists, surgeons, and emergency teams is essential to optimize patient outcomes and ensure timely intervention. KEY POINTS: Interventional radiologists should be fully integrated into trauma teams with 24/7 availability in Level I centers. Whole-body contrast-enhanced CT should be performed in stable trauma patients, followed by immediate embolization when active bleeding is detected. Standardized protocols and materials must be in place to ensure rapid and effective embolization in trauma-related hemorrhage.
To assess the diagnostic value of portable handheld digital radiography (DR) in a beagle model of thoracic trauma. Twenty-seven beagles were randomly assigned to three experimental groups: the pneumothorax group (induced by intrapleural air injection at 50 mL/kg), the pleural effusion group (induced by intrapleural normal saline injection at 30 mL/kg), and the rib fracture group (created by surgical transection of the 6th rib). All animals underwent three imaging examinations in a randomized order: conventional chest X-ray, mobile DR, and portable handheld DR. Detection rates (using surgical outcomes as the gold standard), image quality, and total examination time (from positioning to image acquisition) were compared among the three modalities. A total of 21 animals completed the full protocol (7 per group). There was no significant difference in detection rates among the three examination methods (P > 0.050). The image quality of both mobile DR and portable handheld DR was significantly superior to that of conventional chest X-ray (P = 0.021). Examination times for mobile DR (8.37 ± 0.80 minutes) and portable handheld DR (7.07 ± 0.67 minutes) were significantly shorter than for conventional chest X-ray (10.40 ± 0.96 minutes) (P < 0.001). Furthermore, portable handheld DR had a significantly shorter examination time than mobile DR (P < 0.05). Portable handheld DR provides detection rates comparable with mobile DR and conventional chest X-ray for thoracic trauma, with the advantages of superior image quality over conventional chest X-ray and the shortest examination time. Its user-friendly operation and high portability make it a valuable tool for emergency imaging in austere environments. Based on diagnostic results in a beagle model of thoracic trauma, this study demonstrates that portable handheld DR can provide reliable methodological support for real-time imaging and rapid triage in harsh environments such as field and post-disaster settings.
Multimodal large language models (LLMs) offer emerging capabilities in medical image interpretation; however, their efficacy in orthopedic oncology remains unverified. This study aimed to evaluate and benchmark the performance of five contemporary LLMs-ChatGPT 5.2, Gemini 3 Flash, MedGemma 4B, Claude Sonnet 4.6, and DeepSeek-VL2-in detecting and characterizing bone lesions on plain radiographs without task-specific fine-tuning. A retrospective analysis was conducted using 3,746 anonymized images from the Bone Tumor X-ray Radiograph Dataset (BTXRD), comprising normal, benign, and malignant cases. Reference standard annotations were provided directly by the BTXRD dataset. Models were evaluated on two tasks: lesion detection and lesion characterization. Diagnostic performance metrics, including accuracy, precision, sensitivity, specificity, and Cohen's kappa, were calculated and compared with reference-standard annotations. ChatGPT 5.2 demonstrated the highest overall accuracy (0.803) and specificity among the models (0.916) for lesion detection, although its sensitivity (0.689) was comparatively low. MedGemma 4B showed relatively low performance, with an overall accuracy of 0.677. Claude Sonnet 4.6 and Gemini 3 Flash had the highest sensitivities among the models (0.991 and 0.972, respectively) but low specificities (0.038 and 0.201, respectively), resulting in excessive false positives. In the characterization task, ChatGPT 5.2 consistently achieved the highest performance among the models, with an accuracy of 0.758 and a weighted F1 score of 0.745. DeepSeek-VL2 achieved high specificity but very low sensitivity for malignancy (0.714 and 0.022, respectively). Gemini 3 Flash provided high sensitivity for malignancy (0.711) but low overall accuracy. Multimodal LLMs demonstrated heterogeneous performance in the evaluation of bone lesions on plain radiographs, with substantial differences across models and tasks. Although some models achieved high accuracy in lesion detection and overall classification, performance was inconsistent across tasks, particularly in identifying malignant lesions and balancing sensitivity and specificity. These findings suggest that, despite their potential, current multimodal LLMs are not yet sufficiently reliable for diagnostic use in orthopedic oncology and should be considered investigational until further development and validation. Multimodal LLMs currently lack the diagnostic reliability required for bone lesion assessment, often exhibiting excessive false positives or failing to detect malignancy. Although generalist models show promise, expert radiologist oversight remains essential to ensure patient safety and oncologic accuracy in musculoskeletal imaging.
This study aimed to compare the diagnostic performance of contrast-enhanced mammography (CEM) and dynamic contrast-enhanced breast magnetic resonance imaging (MRI) in the preoperative evaluation of invasive lobular carcinoma (ILC) in terms of lesion size measurement, morphological characteristics, and enhancement kinetics. In this retrospective single-center study, 62 lesions from 46 patients with histopathologically confirmed ILC who underwent both CEM and MRI between February 2021 and December 2024 were analyzed. Lesion size was measured based on the maximum diameter of the index lesion on both modalities. Morphological patterns, including architectural distortion, mass and non-mass enhancement (NME) patterns, and lesion conspicuity, were analyzed according to the American College of Radiology Breast Imaging Reporting and Data System v2025 lexicon. Enhancement kinetics on CEM were categorized as persistent, plateau, or washout based on visual assessment of early- and delayed-phase images and compared with MRI kinetic curve types. The mean age of the patients was 50.6 ± 10.1 years, and multifocal or bilateral disease was observed in 35% of cases. Lesions presented as a mass in 71%, NME in 50%, and a small mass in 37.1% of cases. Both CEM and MRI demonstrated comparable accuracy in measuring the size of the index lesion (intraclass correlation coefficient range: 0.995-0.997; P < 0.001). A high level of agreement was found between CEM and MRI in kinetic curve patterns (Kappa: 0.764; P < 0.001). The sensitivity of CEM was 94.1% for type 1, 89.2% for type 2, and 62.5% for type 3 enhancement patterns. Architectural distortion was significantly more frequent in type 2 and type 3 patterns. High conspicuity was the most common finding on CEM, whereas low conspicuity was predominantly observed in type 1 lesions. MRI was superior for the evaluation of extra-mammary extension, soft tissue involvement, and axillary lymph nodes. CEM demonstrates diagnostic performance comparable with MRI in assessing lesion size, morphology, and enhancement kinetics in ILC. It represents a valuable alternative imaging modality in cases where MRI is contraindicated or limited in availability. However, MRI remains the gold standard for investigating NME and extra-mammary disease extension. CEM provides diagnostic accuracy comparable with MRI in assessing lesion size, morphology, and additional lesions in ILC. Its high sensitivity, shorter examination time, and cost-effectiveness make it a valuable alternative imaging modality, particularly in cases where MRI is contraindicated or limited in availability. Furthermore, the integration of CEM into clinical practice may contribute to improved preoperative assessment.
Preoperative risk estimation of occult high-volume central lymph node metastasis (CLNM) in clinically node-negative (cN0) papillary thyroid carcinoma (PTC) remains challenging. We aimed to develop and validate a practical multimodal model integrating conventional ultrasound and radiomics for individualized preoperative risk stratification. In this retrospective, two-center, multicohort study, 814 patients with cN0 PTC were included and assigned to a development cohort (n = 470), a temporal validation cohort (n = 202), and an external validation cohort (n = 142). Preoperative clinical variables and grayscale ultrasound radiomics features were evaluated. Least absolute shrinkage and selection operator regression was used for feature selection. Four candidate machine-learning algorithms were compared in the development cohort, and the selected classifier was used to construct clinical, radiomics, and fusion models. Model performance was assessed in terms of discrimination, calibration, and clinical utility. SHapley Additive exPlanations (SHAPs) analysis and a web-based calculator were used to improve interpretability and applicability. Three clinical predictors and 20 radiomics features were retained. Random forest was selected for final model construction. In the temporal validation cohort, the clinical, radiomics, and fusion models achieved area under the curve (AUC) values of 0.703, 0.755, and 0.845, respectively; in the external validation cohort, the corresponding AUC values were 0.676, 0.730, and 0.817. For the fusion model, sensitivity, specificity, and accuracy were 0.837, 0.695, and 0.755 in the temporal validation cohort and 0.769, 0.761, and 0.764 in the external validation cohort, respectively. SHAPs analysis identified radiomics score and maximum tumor diameter as the major contributors to model output. A practical multimodal model integrating conventional ultrasound and radiomics showed favorable performance for preoperative risk estimation of occult high-volume CLNM in cN0 PTC and may support individualized risk stratification, pending further prospective validation. This multimodal model may refine preoperative risk estimation for patients with cN0 PTC who harbor occult high-volume CLNM, thereby providing an adjunctive tool for individualized perioperative risk assessment.
Prostate cancer (PCa) is the second most common cancer and cause of cancer deaths among American men. Existing risk prediction methods have limited accuracy and reproducibility, resulting in difficulty in predicting treatment outcomes. We demonstrate the development and external validation of an automated multimodal artificial intelligence (AI) algorithm using biparametric magnetic resonance imaging (bpMRI) and clinical covariates for predicting biochemical recurrence (BCR) after radical prostatectomy (RP) in patients with PCa. The development cohort included 80% of patients from center 1 (n = 240) who underwent prostate MRI prior to RP between January 2008 and December 2018, with a minimum of 2 years of follow-up after RP. The test cohort included the remaining 20% of center 1 patients (n = 71) and an external validation cohort from center 2 (n = 168). Center 2 patients included those who underwent prostate MRI and RP between January 2015 and January 2024, with a minimum of 2 years of follow-up. Clinical comparisons were made using the Cancer of the Prostate Risk Assessment Postsurgical (center 1) and International Society of Urological Pathology Gleason Grade Group (ISUP GGG) scoring systems from post-RP pathology (center 2). The models developed were as follows: clinical (M0), automated clinical (M1), radiomics (M2), and a multimodal model (M3). Clinical variables (M0) included prostate-specific antigen (PSA), age, primary Gleason, and ISUP GGG. Automated clinical variables (M1 and M3) included PSA and age. Radiomic features (M2 and M3) were extracted from bpMRI using a lesion detection AI model. Accuracy, sensitivity, specificity, and area under the receiver operating characteristic curve (AUC) were calculated, and log-rank tests compared BCR-free survival to assess the models' ability to discriminate relative to clinical standards. Intermediate-risk groups were also assessed. The multimodal model (M3) had the highest AUC across test sets (combined: 0.71; center 1: 0.70; center 2: 0.75). This was the only model that significantly differentiated BCR-free survival outcomes in intermediate-risk groups across both centers (P < 0.05). This automated multimodal model leveraging radiomics and clinical covariates can predict BCR after RP, approaching clinical gold standards, and may enhance imaging-based prognostication following further validation. Given that this model demonstrated the potential to outperform pre-surgical and post-surgical clinical gold standards in an external cohort's intermediate-risk patient subgroup (for whom it is more challenging to predict disease trajectory), this model may contribute to enhanced personalized care in PCa after further validation.
To assess myocardial and hepatic tissue remodeling in patients with sickle cell anemia (SCA) using T1 mapping and extracellular volume (ECV) measurements, and to evaluate diastolic dysfunction (DD) using magnetic resonance imaging (MRI). This prospective study enrolled 32 patients with SCA (19 women; mean age: 37 years) and 12 healthy controls (8 men; mean age: 31 years). Myocardial T1, T2, and T2* values and hepatic T1 and T2* values were measured in both groups. Myocardial and hepatic ECV measurements were performed in the patient group. DD was evaluated using transmitral flow (TMF) curves and left atrial volume index (LAVI). In patients with SCA, the mean myocardial T1 value was 1,030.91 ms, significantly higher than in controls, and the mean ECV was 31.7% ± 3.1%, higher than published reference ranges. Moreover, LAVI was significantly higher in patients than in controls (46.48 ± 15.35 vs. 30.58 ± 5.0 mL/m²; P < 0.001). Diastolic function assessment revealed findings suggestive of a restrictive filling pattern in 10 individuals, who had an average age of 33 years. Within the SCA cohort, patients with TMF-derived early/late ratio >2 (n = 10) had significantly higher LAVI than the remaining patients (54.05 ± 10.12 vs. 43.0 ± 16.26 mL/m²; P = 0.027). Myocardial ECV was significantly higher in patients with findings suggestive of a restrictive filling pattern than in the remaining patients (34.14% ± 2.39% vs. 30.62% ± 2.67%; P = 0.001). The mean myocardial T2* value was 39.25 ± 5.9 ms, and no cardiac iron accumulation was detected. The mean hepatic T2* value was 12.20 ms, with iron accumulation observed in 16 patients. Iron accumulation can have an impact on T1 values, and measurements in patients without iron accumulation revealed a mean liver T1 value (626.5 ms) that was significantly higher than that in controls (P < 0.001). These patients also had an ECV (40.4%) that was higher than published reference ranges (P < 0.001). Increased myocardial and hepatic T1 and ECV values were observed in patients with SCA, suggesting expansion of the extracellular space, largely consistent with diffuse interstitial fibrosis in this clinical context. Diastolic function assessment revealed a restrictive filling pattern in a substantial proportion of patients, supporting a possible association between SCA and this pattern. Myocardial ECV was significantly higher in patients with findings suggestive of a restrictive filling pattern than in the remaining patients, suggesting an association between diffuse interstitial fibrosis and advanced DD patterns. MRI may be a valuable modality for evaluating myocardial and hepatic fibrosis and assessing cardiac function. MRI with mapping and ECV quantification may contribute to a non-invasive, quantitative assessment of myocardial and hepatic fibrosis-related tissue remodeling in patients with SCA. Moreover, TMF analysis may provide valuable insights into diastolic function. Together, these methods could contribute to the clinical evaluation and management of SCA and other diseases characterized by fibrosis-related remodeling.
Non-mass enhancement (NME) in breast magnetic resonance imaging (MRI) is a diagnostically challenging entity due to overlapping benign and malignant features, observer variability, and high false-positive rates. This study evaluated the diagnostic performance of three-dimensional (3D) volumetric radiomics and multiple machine learning (ML) algorithms, using early (2nd) and late (7th) post-contrast phases to differentiate benign from malignant NMEs. A total of 110 NMEs (86 benign and 24 malignant) from 108 patients were analyzed. Radiological features were recorded. Radiomics features were extracted from manual 3D segmentations using LIFEx software. Multivariate logistic regression (LR) and supervised ML algorithms-LR, support vector machine, random forest, and gradient boosting-were applied. The methodological quality was assessed using the Multicenter Evaluation of Radiomics in Clinical Studies framework. A total of 54 lesions were histopathologically confirmed, and 56 were confirmed by follow-up. Among the 56 lesions evaluated by follow-up, 34 remained stable (≥ 24 months), whereas 22 showed regression (≥ 6 months). Distribution, internal enhancement, size, and laterality differed significantly between benign and malignant NMEs (P < 0.05). Radiomics analysis extracted 123 features, of which 92 on the early and 80 on the late post-contrast images were significant for benign-malignant differentiation (P < 0.05). In the early phase, combining all radiomics features increased specificity from 78% to 98% and accuracy from 82% to 93%. ML models further improved performance, achieving specificity up to 99% and area under the curve (AUC) values exceeding 0.91. Similar improvements were observed on the late phase, with accuracies up to 91% and AUC values up to 0.93. Volumetric 3D radiomics, combined with ML, using early and late post-contrast phases improves diagnostic accuracy and specificity for NME on breast MRI. Integrating 3D radiomics and ML into breast MRI evaluation supports more accurate decision-making in NME and may reduce unnecessary biopsies.
To conduct a systematic review and meta-analysis to indirectly compare the diagnostic performance of contrast-enhanced ultrasonography (CEUS) and magnetic resonance imaging (MRI) in the evaluation of indeterminate testicular lesions and to explore their potential roles in guiding clinical decision-making. PubMed, Web of Science, and Scopus were searched in August 2025. Eligible studies were prospective or retrospective cohorts evaluating CEUS or MRI in small, impalpable, or incidentally detected testicular lesions, with histopathology or clinical/radiological follow-up as reference standard. Two reviewers independently performed selection, data extraction, and Quality Assessment of Diagnostic Accuracy Studies-2 tool assessment. Random-effects meta-analyses and meta-regression were conducted in R (R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria). This review was registered in the Prospective Register of Systematic Reviews (CRD420251113920). A total of 12 studies including 912 patients with 920 lesions (279 malignant, 641 benign) were analysed. CEUS showed a sensitivity of 90% (95% confidence interval [CI] 80-95%) and specificity of 74% (95% CI 47-91%), while MRI achieved a sensitivity of 94% (95% CI 88-97%) and specificity of 84% (95% CI 72-91%). CEUS provided a higher negative predictive value (NPV; 96% vs 88%), whereas MRI demonstrated superior positive predictive value (PPV; 83% vs 65%) and accuracy (91% vs 80%). Differences in PPV and accuracy significantly favoured MRI (P = 0.025 and P = 0.031). Limitations included protocol heterogeneity, operator dependence, and small sample sizes. Both CEUS and MRI demonstrated a high sensitivity and NPV, supporting their role in excluding malignancy. CEUS, with accessibility, low cost, and particularly high NPV, may serve as the first-line adjunct, while MRI is best reserved for equivocal or higher-risk cases. No funding was received for this study.
To update the pooled diagnostic accuracy and safety of cone-beam computed tomography (CBCT)-guided percutaneous transthoracic needle biopsy (PTNB) for pulmonary lesions. PubMed, Embase, and Cochrane CENTRAL were searched through April 2026 for studies of CBCT-guided PTNB including ≥ 10 patients. Sensitivity and specificity were pooled using a bivariate random-effects model. Complication rates (pneumothorax and pulmonary hemorrhage, encompassing both clinically significant hemoptysis and radiographic perilesional change) were pooled using generalized linear mixed models. Heterogeneity and publication bias were assessed. The review followed the Preferred Reporting Items for Systematic Reviews and Meta-Analyses of Diagnostic Test Accuracy Studies, used the Quality Assessment of Diagnostic Accuracy Studies 2, and was prospectively registered in PROSPERO (CRD420261370382). Twenty-four studies were eligible (k: 22 with extractable 2 × 2 data; n = 3,681; 2010-2026). Pooled sensitivity was 92.2% [95% confidence interval (CI): 90.4-93.7] and pooled specificity was 96.7% (95% CI: 94.3-98.2). The summary receiver operating characteristic area under the curve was 0.974. Pneumothorax incidence was 18.4% [95% CI: 15.5-21.7; prediction interval (PI) 8.3-35.8%]. Hemorrhage incidence (k: 19) pooled at 6.1% (95% CI: 3.5-10.6); high heterogeneity (I2: 90.9%) reflected variation in outcome definitions. CBCT-guided PTNB demonstrates high diagnostic accuracy for pulmonary lesions. The pooled specificity should be interpreted as a plausible upper-bound estimate because of universal differential verification bias across the included studies. The wide pneumothorax PI (8-36%) indicates that complication rates may vary substantially across centers despite the precise pooled estimate.
To identify magnetic resonance imaging (MRI)-based morphologic features of the distal tibiofibular syndesmosis and talocrural joint associated with ankle sprains and to determine which parameters are specifically linked to an increased risk of ligament tears after sprains. This retrospective study included ankle MRI examinations performed between January 2022 and November 2025. Two analytic datasets were constructed: Dataset 1 compared patients with ankle sprains and MRI-confirmed ligament tears to healthy controls with completely normal ankle MRIs; Dataset 2 compared controls, patients with sprains but no ligament tears, and patients with sprains and ligament tears. Standardized morphometric measurements of the distal tibiofibular syndesmosis and talocrural joint were obtained, including absolute and ratio-based parameters. Group comparisons were performed using appropriate univariable tests with multiple-comparison correction. Multivariable logistic regression was used to identify independent predictors of ligament tears. Inter-reader reliability was assessed using intraclass correlation coefficients and Bland-Altman analysis. In Dataset 1, compared with healthy controls, tibiofibular clear space and the fibular notch depth-to-tibial thickness ratio were significantly higher in patients with ankle sprains and ligament tears, whereas the lateral malleolar height-to-talar articular width ratio was significantly lower (all P < 0.05). Multivariable analysis demonstrated that tibiofibular clear space and the fibular notch depth-to-tibial thickness ratio were independent predictors. In Dataset 2, tibiofibular clear space, fibular notch depth-to-tibial thickness ratio, and lateral malleolar height-to-talar articular width ratio differed significantly across three groups (all P < 0.05). Notably, only the lateral malleolar height-to-talar articular width ratio independently differentiated patients with sprains and ligament tears from those without tears (P = 0.025). Model discrimination was moderate to good (area under the curve: 0.699 and 0.785). Specific MRI-based morphologic features are associated with both ankle sprain susceptibility and an increased risk of ligament tears. Among these, the lateral malleolar height-to-talar articular width ratio appears to be a morphometric parameter associated with an increased likelihood of ligament tears. MRI-based morphometric assessment, particularly the lateral malleolar height-to-talar articular width ratio, may help identify patients with ankle sprains who are at an increased risk for ligament tears and support more individualized clinical management.
Early hematoma expansion (HE) after intracerebral hemorrhage (ICH) critically affects patient outcomes. Identifying HE risk using emergency non-contrast computed tomography (NCCT) features and clinical data is important for early risk stratification and evaluation. This study analyzed factors associated with early HE and evaluated the value of NCCT features combined with clinical data for HE identification. A total of 571 patients with spontaneous supratentorial ICH admitted between January 2022 and December 2025 were enrolled. Five NCCT signs were evaluated at baseline. Four signs significantly associated with HE (black hole, swirl, blend, and satellite signs) were summed to construct a primary comprehensive sign score (0-4); a 5-sign score additionally including the island sign was examined in a sensitivity analysis. Patients were stratified by HE status, sign score quantiles, sex, and hematoma morphology, then randomly assigned to training (n = 391) and validation (n = 180) sets. Univariate and multivariate logistic regression identified independent HE factors and constructed a combined prediction model. Model performance was evaluated using receiver operating characteristic curves, calibration curves, and decision curve analysis (DCA). HE was defined as an absolute volume increase > 6 mL or a relative increase > 33% on follow-up CT compared with baseline CT. Among 571 patients, 292 (51.1%) experienced early HE. Multivariate analysis showed that the 4-sign score [odds ratio (OR): 1.850, 95% confidence interval (CI): 1.505-2.274, P < 0.001], irregular hematoma morphology (OR: 2.071, 95% CI: 1.262-3.399, P = 0.004), time from onset to initial CT (OR: 0.845, 95% CI: 0.750-0.953, P = 0.006), male sex (OR: 1.500, 95% CI: 1.028-2.189, P = 0.036), and baseline CT attenuation (OR: 0.964, 95% CI: 0.934-0.995, P = 0.024) were independently associated with HE. The primary combined model achieved an area under the curve (AUC) of 0.731 (95% CI: 0.682-0.781) in the training set and 0.716 (95% CI: 0.641-0.791) in the validation set. The 5-sign sensitivity model showed similar discrimination (AUC: 0.722 and 0.704, respectively), with no significant difference (P = 0.188 and P = 0.283). Calibration and DCA were broadly comparable between models. A 4-sign comprehensive imaging score, irregular hematoma morphology, time from onset to initial CT, male sex, and baseline CT attenuation were independently associated with early HE in patients with ICH. A model based on NCCT features combined with clinical data may provide a practical adjunct for early HE risk stratification in patients with ICH.
This study aimed to evaluate the utility of intratumoral and peritumoral radiomics derived from multi-parametric magnetic resonance imaging for predicting microsatellite instability (MSI) in endometrial cancer (EC). We retrospectively analyzed 161 patients with pathologically confirmed EC, assigning them to a training (n = 112) and a test cohort (n = 49) at a 7:3 ratio, and collected their full clinical and imaging data. We manually delineated regions of interest on axial T2-weighted imaging (T2WI), diffusion-weighted imaging (DWI), and dynamic contrast-enhanced (CE) T1-weighted imaging sequences, expanding them by 3, 5, and 7 mm to define peritumoral regions. We then extracted and selected radiomic features from both intratumoral and peritumoral areas. Using six machine learning algorithms, we developed separate radiomics models and assessed their performance via the area under the receiver operating characteristic curve (AUC). To complete the evaluation, we compared statistical differences using the DeLong test and generated calibration curves to verify predictive accuracy. Clinical characteristics exhibited no significant correlation with MSI status. In single-sequence analysis, the CE model demonstrated the highest performance (training AUC: 0.912; test AUC: 0.856). The multi-parametric model (T2WI + DWI + CE) outperformed single-sequence models, achieving AUCs of 0.934 and 0.914 in the training and test cohorts, respectively. Peritumoral radiomics independently showed robust predictive value; specifically, the model derived from 3 mm peritumoral features (across DWI) yielded a training AUC of 0.898 and a test AUC of 0.790. Notably, the integration of intratumoral and peritumoral features maximized predictive accuracy. The final model, combining optimal intratumoral features (T2WI + DWI + CE) with 3 mm peritumoral DWI features, achieved the best performance with a remarkable AUC of 0.998 in the test cohort. Peritumoral radiomics possesses strong independent predictive value and significantly enhances the performance of intratumoral models. The integration of intratumoral and peritumoral radiomics offers a precise, non-invasive method for preoperative MSI prediction, serving as a valuable tool to facilitate clinical decision-making. This study establishes a reliable, non-invasive radiomics approach for the preoperative assessment of MSI status.
Patients with superior vena cava (SVC) obstruction (SVCO) are referred for stenting to alleviate symptoms such as plethora and dyspnea. Accurate visualization and measurement of the central veins are essential for appropriate stent selection and optimal placement. This is critical to avoid extension into the right atrium, which can lead to arrhythmias or stent migration. However, in cases of severe SVCO, the sinoatrial junction (SAJ) often cannot be opacified on angiograms. Altered flow dynamics cause suboptimal perfusion and distention, complicating precise measurements and challenging stent planning-especially when the tumor encroaches near or beyond the SAJ. We integrated intravascular ultrasound (IVUS) into our procedural workflow and evaluated its utility in a series of 27 cases involving malignancy-related SVCO between November 2023 and March 2025. We measured the distance from the midpoint of the stenosis to the SAJ (stenosis-to-SAJ distance) and identified cases requiring IVUS for accurate assessment of the caudal landing zone. Parametric testing revealed a stronger correlation between stenosis-to-SAJ distance measurements on IVUS and computed tomography (CT) than between digital subtraction angiography and CT. Statistical analysis determined that a stenosis-to-SAJ distance of ≤ 40 mm was significantly associated with the need for IVUS (P = 0.008), whereas the association at ≤ 50 mm was not statistically significant (P = 0.069). Stenting becomes particularly challenging when tumors invade the SAJ. Our findings suggest that IVUS provides valuable visualization and measurement, particularly in cases with a stenosis-to-SAJ distance of ≤ 40 mm, making it a useful adjunct for safe and effective SVC stent placement. Visualization of venous anatomy and the exact extent of SVC stenosis is difficult in cases of severe obstruction, especially when the tumor encroaches upon the SAJ, making stent selection and deployment challenging. The use of IVUS facilitates visualization of the precise extent of the stenosis and delineates the location of the SAJ. A stenosis-to-SAJ distance of ≤ 40 mm significantly benefits from the concomitant use of IVUS for accurate stent placement.
Hepatic alveolar echinococcosis is a tumor-mimicking parasitic disease characterized by infiltrative hepatic involvement and substantial radiologic overlap with malignancy. Predominantly solid morphology, pseudo-capsular margins, and deceptive enhancement patterns frequently contribute to misclassification as primary or secondary liver cancer. Intralesional calcifications on non-contrast computed tomography (CT) and the absence of true arterial phase hyperenhancement on CT and magnetic resonance imaging represent important diagnostic clues. Diffusion-weighted imaging may further aid characterization, although overlap with malignant lesions persists. Accurate interpretation requires systematic multimodality imaging assessment integrated with clinical and serologic context. Increased awareness of characteristic imaging features and common pitfalls is essential to reduce diagnostic delay and prevent inappropriate oncologic management.
This study aimed to evaluate the prognostic value of apparent diffusion coefficient (ADC) measurements and a predefined magnetic resonance imaging (MRI)-based radiomics score for predicting early recurrence and disease-free survival (DFS) in patients with hepatocellular carcinoma (HCC) undergoing curative surgical resection. This retrospective study included 200 patients who underwent curative-intent resection for HCC between January 2021 and January 2025, with follow-up data reviewed through January 2026. Preoperative multiparametric MRI examinations, including DWI, ADC maps, and dynamic contrast-enhanced sequences, were reviewed. The predefined radiomics score was derived from a validated institutional MRI-based radiomics pipeline using ADC maps and contrast-enhanced MR images. In the present cohort, the score was analyzed together with region-of-interest (ROI)-based ADC measurements as a fixed imaging biomarker, without any additional feature selection, model training, or model retraining. Early recurrence was defined as recurrence within 12 months after surgery. Univariable and multivariable logistic regression analyses were performed to identify predictors of recurrence. Prognostic performance was evaluated using receiver operating characteristic analysis, and model discrimination was compared using the DeLong test. The DFS was assessed using Kaplan-Meier and Cox proportional hazards analyses. Early recurrence occurred in 66 of 200 patients (33.0%). In multivariable logistic regression analysis, tumor size [odds ratio (OR): 1.55, 95% confidence interval (CI): 1.27-1.89, P < 0.001), ADC (OR: 0.29, 95% CI: 0.18-0.46, P < 0.001), and radiomics score (OR: 4.74, 95% CI: 2.85-7.89, P < 0.001) were independent predictors of recurrence. The radiomics score alone achieved an area under the curve (AUC) of 0.767, whereas ADC achieved an AUC of 0.707. Combining the ADC and radiomics score significantly improved predictive performance (AUC: 0.837), outperforming both ADC alone (P < 0.001) and the radiomics score alone (P = 0.008). The full multivariable model demonstrated the highest discriminatory performance (AUC: 0.892). For survival analysis, larger tumor size, multiple lesions, lower ADC values, and higher radiomics scores were independently associated with shorter DFS. Patients classified as high-risk according to the optimal radiomics score cut-off exhibited significantly worse DFS than low-risk patients (log-rank P < 0.001). Measurements of the ADC and a predefined MRI-based radiomics score provided complementary prognostic information in patients with HCC. The combined ADC-radiomics score improved prediction of early recurrence and DFS compared with either parameter alone. These findings support further validation of quantitative MRI biomarkers for non-invasive risk stratification after curative HCC resection. The combined assessment of ROI-based ADC measurements and a predefined MRI-based radiomics score may help refine postoperative surveillance strategies after curative HCC resection, pending validation in independent cohorts.
Randomized controlled trials (RCTs) comparing hydrogel-coated coils (HGCs) with bare platinum coils (BPCs) have yielded heterogeneous results, and the clinical relevance of longitudinal angiographic assessment remains uncertain. Following the publication of the HYBRID trial, which emphasized occlusion trajectory rather than static end points, a post-HYBRID updated meta-analysis is warranted to compare angiographic durability and safety outcomes between these coil types. A systematic literature search of MEDLINE, Web of Science, and Scopus was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses guidelines. RCTs comparing HGCs with BPCs for the treatment of intracranial aneurysms were included. Six RCTs comprising 2,486 patients and 2,513 treated intracranial aneurysms were included in the analysis. There was no statistically significant difference between HGCs and BPCs in immediate complete occlusion [risk ratio (RR): 0.87; 95% confidence interval (CI), 0.70-1.09; I²: 63.0%] or immediate adequate occlusion (RR: 0.96; 95% CI, 0.85-1.07; I²: 31.0%). Immediate residual aneurysm was significantly more frequent with HGCs than with BPCs (RR: 1.12; 95% CI, 1.01-1.24; I²: 0.0%; P = 0.041). At the last available angiographic follow-up, complete occlusion, adequate occlusion, and residual neck rates remained comparable between HGCs and BPCs. However, residual aneurysm at follow-up was significantly less frequent with HGCs than with BPCs (RR: 0.75; 95% CI, 0.68-0.83; I²: 0.0%; P = 0.006). HGCs were also associated with a significantly lower rate of major recurrence than BPCs (RR: 0.71; 95% CI, 0.54-0.94; P = 0.024). Safety and clinical outcomes were similar between treatment groups. Despite comparable safety and clinical outcomes between HGCs and BPCs, HGCs demonstrated superior angiographic durability, as reflected by lower rates of residual aneurysm and major recurrence. HGCs may offer improved long-term aneurysm durability compared with BPCs by reducing residual aneurysm and major recurrence without compromising safety or clinical outcomes. These findings support consideration of HGCs when durable occlusion is a priority in endovascular treatment planning for intracranial aneurysms.
To develop and evaluate a glass-box, agentic-style radiology pipeline that separates perception from reasoning for auditable multiclass diagnosis on cine cardiac magnetic resonance imaging (MRI), and to quantify accuracy, robustness across decoding temperatures, and fidelity/safety of generated narrative explanations. Using the labeled Automated Cardiac Diagnosis Challenge training cohort (n = 100; five diagnostic classes), cine bSSFP images were segmented at end-diastole (ED) and end-systole (ES) with a pretrained nnU-Net, and 17 clinically interpretable biomarkers were extracted. A large language model (LLM) (GPT-OSS-120B) queried prompts under three different prompt strategies (V1-V3) with majority-vote self-consistency after a stratified split into prompt development (n = 20) and independent evaluation (n = 80). Temperatures (T = 0.1, 1.0, and 2.0) were tested for stability. A decoupled narrative module generated radiologist-style reports. Narratives underwent radiologist audit for numeric fidelity and clinical safety. Machine learning algorithms [Random Forest, Support Vector Machine (SVM), Logistic Regression, Decision Tree] were trained on the same biomarker set for benchmarking. Automated segmentation showed high agreement with reference masks [Dice at ED: right ventricle [RV] cavity 0.984 ± 0.004, left ventricle (LV)] myocardium 0.965 ± 0.009, LV cavity 0.989 ± 0.003; ES: RV cavity 0.979 ± 0.013, LV myocardium 0.975 ± 0.009, LV cavity 0.985 ± 0.005). The hierarchical veto-logic strategy (V3) achieved an accuracy of 0.925 (95% confidence interval: 0.863-0.975) and a macro-F1 of 0.924, remaining stable across temperatures, outperforming V2 (accuracy 0.787-0.800) and V1 (0.562-0.600). Reproducibility was highest for V3 at T = 0.1 (Fleiss' kappa: 0.969) with a low failure rate (0.83%). Narrative generation produced 97.5% valid reports with 100% numeric fidelity and audited safety ≥ 97.5%. Performance was comparable to supervised models (Random Forest accuracy 0.938; SVM/Logistic Regression accuracy 0.925). In this single-dataset internal evaluation, a glass-box workflow combining automated segmentation-derived biomarkers with an LLM enables robust multiclass cardiac MRI diagnosis while producing numerically faithful, safety-audited narratives, supporting auditability and governance for radiology artificial intelligence (AI). External multicenter validation is needed to confirm generalizability. A glass-box, biomarker-driven agentic-style workflow enables auditable cine cardiac MRI classification with numerically grounded explanations, addressing interpretability and stability barriers that limit translation of radiology AI into routine practice.
To investigate the accuracy of hematoma volume (V), surface area (S), and surface regularity (SR) measured on dual-energy computed tomography angiography (DECTA), using non-contrast computed tomography (NCCT) as the reference standard. A total of 129 patients with spontaneous intracerebral hemorrhage (sICH) who underwent both NCCT and DECTA scans were retrospectively studied. Patients were stratified by V: Group 1 (< 30 mL), Group 2 (30-60 mL), and Group 3 (> 60 mL); and by morphology: regular, irregular, and lobular. DECTA data were post-processed to generate conventional computed tomography angiography (CTA) and 60 keV virtual monoenergetic imaging (VMI). These, along with NCCT images, were imported into 3D Slicer software to obtain V and S; SR was subsequently calculated. Deviation percentages (ΔV, ΔS, ΔSR) of conventional CTA and 60 keV VMI relative to NCCT were calculated. We assessed correlations using simple linear regression, compared different modalities with unpaired t-tests or Mann-Whitney U tests, evaluated agreement via Bland-Altman analysis, and compared deviation percentages across groups using one-way analysis of variance or Kruskal-Wallis H tests. Inter-reader consistency was assessed using the intraclass correlation coefficient (ICC). The average V was 28.86±24.15 mL. All ICCs were excellent (0.991-0.999). Strong correlations were found for V and S between conventional CTA/60 keV VMI and NCCT (r: 0.987-0.996). Bland-Altman analysis for V showed mean biases of 0.21 mL (conventional CTA) and -0.04 mL (60 keV VMI) against NCCT, with 95% limits of agreement of -4.72 to 5.14 mL and -4.51 to 4.43 mL, respectively. However, SR values from both conventional CTA and 60 keV VMI were significantly lower than those from NCCT (all P < 0.001). In volume-based stratification, no significant differences in ΔS (conventional CTA) or ΔV, ΔS, and ΔSR (60 keV VMI) were found among groups (all P ≥ 0.410). For conventional CTA, ΔV was significantly smaller in Group 3 (> 60 mL) than in Group 1 (3.07% vs. 5.57%, adjusted P = 0.047), whereas ΔSR was larger (13.26% vs. 9.88%, adjusted P = 0.040). Morphology-based stratification revealed no significant differences in ΔV, ΔS, or ΔSR across groups for either modality (all P values ≥ 0.085). Hematoma volume measurements from DECTA show good agreement with NCCT measurements, suggesting potential utility for follow-up assessment in certain clinical scenarios. However, DECTA-derived SR measurements are significantly lower than those from NCCT. Volume measurement accuracy of conventional CTA was higher for large hematomas (> 60 mL). Accurate hematoma measurement is crucial for prognosis and management in sICH. This study indicates that V measurements from DECTA are comparable to those from NCCT in certain clinical scenarios, offering potential added value when CTA is clinically indicated.