This study examines the effects of Artificial Intelligence (AI) capabilities and entrepreneurial competency on Small and Medium Enterprise (SME) performance through the mediating role of strategic intelligence in SMEs operating in Malaysia and Saudi Arabia. Grounded in Resource Orchestration Theory and Core Competency Development Theory, the study explores how AI integration skills, smart decision-making abilities, innovativeness, and risk management competencies contribute to enhanced organizational performance in an increasingly digital business environment. A cross-sectional quantitative research design was adopted. Data were collected from 320 SMEs in Malaysia and Saudi Arabia using a structured questionnaire. The measurement model and structural relationships were analyzed using Partial Least Squares Structural Equation Modeling (PLS-SEM). The study examined direct relationships between AI capabilities, entrepreneurial competency, strategic intelligence, and SME performance, as well as the mediating effect of strategic intelligence. The findings reveal that both AI capabilities and entrepreneurial competency have a significant positive impact on strategic intelligence and SME performance. Furthermore, strategic intelligence was found to significantly mediate the relationship between AI capabilities, entrepreneurial competency, and SME performance. These results indicate that SMEs with stronger AI-driven capabilities and entrepreneurial competencies are more likely to develop higher strategic intelligence, which in turn enhances overall organizational performance. The study contributes to the growing body of literature on AI-driven organizational capabilities by providing empirical evidence from developing economies, specifically Malaysia and Saudi Arabia. The results highlight the importance of integrating AI competencies with entrepreneurial skill development to improve strategic decision-making and business competitiveness. Practically, the findings suggest that SME managers and policymakers should prioritize AI adoption and entrepreneurial capability development as key strategies for improving innovation, resilience, and long-term performance in SMEs.
Artificial intelligence (AI)-based clinical decision support systems (CDSSs) are increasingly used to support clinicians by providing evidence-based treatment recommendations for multidisciplinary team (MDT) meetings. This systematic review mapped the current landscape of AI-based CDSSs in surgical oncology decision-making and synthesized evidence on their performance and factors influencing effectiveness. Cochrane, Ovid MEDLINE, and Embase were searched on February 3, 2025. Studies evaluating AI-based CDSSs for therapeutic decision-making in surgical oncology were included. Methodological quality was assessed using the Critical Appraisal Skills Programme Diagnostic Study Checklist, and data were synthesized narratively. Fifty-nine studies, encompassing 23,158 patients, were included. CDSSs were classified into five categories: decision tree-based systems, knowledge representation-based systems, Watson for Oncology (WfO), large language models, and other AI-based systems. Concordance with MDT or guideline-based recommendations ranged from 23.2% to 99%. Decision tree- and knowledge-based systems generally demonstrated higher concordance and improved guideline adherence, while WfO performance varied substantially by region and treatment accessibility. Three key themes related to implementation challenges emerged: technical limitations, socioeconomic and healthcare system constraints, and patient- and tumour-specific factors. Reported benefits included improved adherence of MDT decisions to clinical guidelines, enhanced identification of patients eligible for clinical trial enrolment, support for less experienced clinicians, and facilitation of triage for routine cases. AI-based CDSSs show promise in supporting MDT decision-making but remain constrained by challenges related to system maintenance, variability in clinical protocols, therapeutic availability, and patient heterogeneity. Larger prospective studies are needed to evaluate the real-world integration, clinical impact, and patient outcomes of AI-based CDSSs within MDT workflows. https://www.crd.york.ac.uk/PROSPERO/view/CRD42025639227, identifer: CRD42025639227.
The increase in sophisticated electronic gambling products offered by gambling manufacturers facilitates an increase in the accessibility of gambling platforms and modes. This study proposes utilising Artificial Intelligence (AI) to enhance self-awareness and self-control within electronic gaming machines. The study adopts a case study that follows an exploratory research design, as the causal explanation argument suggests that the Electronic Gaming Machine (EGM) device causes problem gambling. Secondary data on gambling were obtained and analysed. The machine-learning pattern recognition and classification model; a learning approach which learns from a trained dataset to make decisions or predictions, was employed. The dataset was trained iteratively using the scaled conjugate gradient backpropagation with the input and output target samples divided into training, validation, and test datasets. Furthermore, the softmax was used for classifying the dataset into three classes: responsible, intermediate, and irresponsible gambling. The confusion matrix was used to analyse the percentages of correct and incorrect classifications. The results obtained indicated that the accuracy of the developed model was 99.20%, while the precision was 85.70%. The recall achieved 85.70%, while the F1-score reached 80.50%. The closeness of these performance indices to 1, coupled with the negligible value of mean square error, indicates that the developed classification model is robust and suitable for classification problems. Thus, this study contributes to knowledge by developing an AI model that can track players and reduce harm in a land-based gambling environment.
Ensuring fairness and reliabilityin examination scoring remains a persistent challenge in educational assessment, particularly in the presence of subjective grading inconsistencies across evaluators. While automated essay scoring systems improve scalability, limited attention has been given to mechanisms for verifying whether assigned scores are consistent with rubric criteria and comparable peer responses. This study proposes an Explainable Fairness-Aware AI Framework for Exam Score Verification, designed to detect potential scoring anomalies while providing interpretable evidence to support verification decisions. The proposed framework integrates three complementary components: semantic response evaluation using contextual representations derived from DeBERTa-v3 embeddings, criterion-level rubric alignment via cross-attention mechanisms, and peer-consistency analysis based on similarity-driven cohort score distributions. To enhance transparency, an explainability layer combines attention visualization, integrated gradients, and natural language justification to provide both quantitative and qualitative insights into model decisions. The framework was evaluated on the ASAP 2.0 dataset comprising 24,728 rubric-scored essays. Experimental results, reported as mean ± standard deviation over multiple runs, demonstrate that the proposed approach achieves an agreement of 87.1% ± 0.9 with peer-consistent reference scores and a fairness consistency index of 0.82 ± 0.02, outperforming a diverse set of baseline models, including traditional, transformer-based, and hybrid scoring approaches. Ablation studies further confirm the critical role of peer-consistency analysis in detecting scoring irregularities, while the explainability components enhance interpretability without significantly affecting predictive performance. The notion of fairness addressed in this work is grounded in score consistency relative to rubric expectations and semantically similar peer responses. Within this scope, the results indicate that integrating semantic evaluation, rubric-aware modeling, and cohort-referenced analysis provides a robust and interpretable framework for exam score verification. These findings highlight the potential of multi-evidence, explainable AI systems to support reliable and transparent assessment in educational settings.
Artificial Intelligence (AI) has the potential to revolutionize medicine, particularly in the field of cardiology. There are significant diagnostics and treatments variabilities in the field of cardiovascular medicine that affects racial and ethnic racially and ethnically diverse populations as well as female patients across all age groups. The efforts put forth towards the development of AI and precision medicine within the cardiovascular practice do not fully account for existing variations in cardiovascular care delivery. AI models and precision medicine tools that were created with uncomprehensive data primarily drawn from White populations risk embedding historical differences into clinical decision-support systems. This paper outlines the integrative approach taken to review the current variabilities that persist within younger adults (<65 years) and older adults (≥ 65 years) who have cardiovascular disease. Additionally, genetic factors, limited access to care, health literacy, lack of insurance coverage and adherence are examined, as these are frequently cited as major contributors to health care adverse outcomes but remain under-researched and unresolved even with the expansion of Medicaid. For instance, Black patients experience higher prevalence of heart failure (HF) and hypertension, especially transthyretin amyloid cardiomyopathy HF, with Black women being disproportionately affected due to higher structural, environmental and clinical factors. Also, racially and ethnically diverse children with CVDs have higher odds of mortality than their White counterparts. The integration of AI in cardiovascular medicine must first be preceded by an active effort to restructure systems and reduce variable outcomes. Future research must prioritize diverse genomic datasets and equitable comprehensive representation in clinical trials. These initiatives are better served if they are driven by institutions that historically serve racially and ethnically diverse populations and communities to better enhance inclusion and fairness in electronic medical record keeping. Accordingly, cardiovascular medical practices and technology can progress forward with AI and precision medicine models that are both equitable and accurate.
Reproductive tract infections (RTIs) and sexually transmitted infections (STIs) pose a substantial economic burden and public health concern in developing countries such as India, where inadequate early detection and prevention strategies often lead to increased morbidity, mortality, stigma, cancer and adverse reproductive health outcomes in both men and women. The present cross-sectional study employed machine-learning models to predict the risk of RTI/STI among young women in Delhi, India, by analysing self-reported symptoms along with demographic, reproductive health, lifestyle, and hygiene-related factors. The data were analysed using multiple machine learning algorithms, including decision trees, random forest, k-nearest neighbors (KNN), stochastic gradient descent (SGD), and XGBoost. Prior to analysis, the data was preprocessed through different steps including missing-value imputation, outlier removal using the interquartile range (IQR) method, and the synthetic minority oversampling technique (SMOTE) to address class imbalance. A comparative analysis of model performance revealed that the XGBoost algorithm achieved the highest accuracy (89.1%) and the greatest area under receiver operating characteristics curve (ROC-AUC = 0.935), with random forest performing strongly as the second-best model (accuracy = 87.6%, ROC-AUC = 0.921), demonstrating its strong ability to capture non-linear relationships among RTI risk factors. SHAP (SHapley Additive exPlanations) analysis identified family medical history, personal RTI history, and menstrual cycle characteristics as the most influential predictors. The findings underscore the potential of machine learning-driven predictive models in RTI risk stratification, screening support, and hypothesis generation.
This study aims to identify, classify, and typologize the mental models of civil engineers, the diversity of their perspectives, and their intellectual structures regarding the integration of artificial intelligence in the industry. The philosophical framework of this research is based on an interpretive-positivist paradigm (Q method) and its application. The 11 participants were selected from individuals with sufficient knowledge, expertise, and experience related to the research topic, and who could provide the necessary information to achieve the study's objectives. The purpose of this sampling is to classify them and identify common mental patterns within a small and targeted group, rather than to generalize the findings to a larger population. The results indicate that civil engineers' mental models regarding the applications of artificial intelligence are divided into three categories: hopeful-concerned, black-box thinkers, and results-oriented. Analyzing the distinct thought patterns among civil engineers reveals significant heterogeneity in the acceptance of AI, highlighting the need for fundamental changes in communication strategies. The study shows that to facilitate use, key information must be delivered through tailored content. For example, for the black-box group, content should emphasize algorithmic transparency and how models work. For the hopeful-concerned group, content should reduce ethical and employment-related concerns. Meanwhile, for the results-oriented group, content should highlight economic returns and proven successes. Ultimately, this approach is essential for effective and successful integration of AI into the infrastructure industry.
Total hip arthroplasty (THA) is a well-established treatment for end-stage hip disorders, yet its success heavily depends on precise acetabular cup positioning and sizing. Conventional planning, based on manual CT interpretation, is time-consuming, operator-dependent, and lacks standardisation, limiting its ability to achieve consistent surgical accuracy. This narrative review systematically searched PubMed, Web of Science, Cochrane Library, and IEEE Xplore for peer-reviewed studies published between January 2015 and June 2025. We included original research evaluating AI-driven preoperative CT 3D planning for THA, with quantitative outcomes on cup angle or size accuracy. Data were extracted and assessed for methodological quality using standard tools. AI-assisted planning consistently improved accuracy: mean angular errors for inclination and anteversion were below 3*, size matching accuracy within ±1 size ranged from 80% to 85%, and planning time was reduced by 57% to 70% compared with manual templating. These findings were reproducible across different deep-learning architectures and patient cohorts. Although AI planning shows clear benefits in accuracy and efficiency, several challenges remain-including limited generalisability to complex anatomies, susceptibility to image artefacts, and insufficient integration with intraoperative execution. Future research should prioritise multi-centre validation, dynamic functional planning, and seamless clinical workflow integration to translate technological potential into improved patient outcomes.
Large Language Models (LLMs) have demonstrated remarkable proficiency in general-purpose tasks, yet their capacity for fine-grained reasoning in knowledge-intensive domains (KIDs) remains largely unexplored. This study addresses this gap by investigating LLM performance in the specialized field of food chemistry. We introduce a novel benchmark comprising two core tasks: Molecular-to-Food Prediction (MFP) and Food-to-Molecular Prediction (FMP). To support this benchmark, we curated and standardized the FlavorDB dataset, creating a robust testbed for flavor-molecular association. We evaluated two open-source models (Kimi-K2, DeepSeek-V3.2) and three closed-source models (Gemini-3-Pro, GPT-5.1, and Seed-1.8) under zero-shot and one-shot in-context learning settings. Our systematic analysis yields six key findings that characterize the capabilities and limitations of current LLMs in this domain. For instance, in the FMP task, Gemini-3-Pro achieved the highest zero-shot F1 score of 0.556, while Kimi-K2 led the one-shot setting with an F1 score of 0.522. In-context learning consistently improved performance across models, most notably boosting Kimi-K2's F1 score from 0.451 to 0.496 on complex multi-food tasks. Critically, we identify four recurring categories of domain-specific reasoning errors, which illuminate the fundamental challenges general-purpose models face when applied to fine-grained scientific inference in food chemistry. This work not only establishes a framework for rigorously evaluating LLM potential in knowledge-intensive domains but also provides a critical foundation for advancing practical applications, including flavor optimization and the development of novel food products.
We present a multi-stage optimization strategy combining reinforcement learning (RL) and supervised fine-tuning (SFT) to enhance the pedagogical knowledge of Large Language Models (LLMs), thereby providing a technically grounded example of how open-source pedagogical LLMs can be optimized for deployment in diverse educational settings, including institutions with constrained resources. Our approach produces EduQwen 32B-RL1, EduQwen 32B-SFT, and EduQwen 32B-SFT-RL2: (1) first-stage RL optimization implementing progressive difficulty training, focusing on challenging examples, and employing extended reasoning rollouts to facilitate adaptive scaffolding, prioritizing pedagogical steering over direct answer provision; (2) SFT that leverages the RL-trained model to synthesize high-quality training data with difficulty-weighted sampling; and (3) optional second-stage RL refinement. This application-driven family of open-source pedagogical LLMs, built on a dense Qwen3-32B backbone, achieves 96.52% accuracy on the Cross-Domain Pedagogical Knowledge (CDPK) Benchmark, establishing new state-of-the-art (SOTA) performance on the CDPK subset of the interactive Pedagogy Benchmark Leaderboard as of March 2026, with a 5.97 percentage points accuracy gain over the then-reported Gemini-3 Pro's score (previous leader with 90.55% accuracy), under the respective documented evaluation protocols. Critically, our 32-billion-parameter models demonstrate that domain-specialized optimization of mid-sized open-source LLMs can outperform much larger general-purpose systems on pedagogical knowledge benchmarks, indicating a scalable technical approach for wider AI deployment and potential pedagogical support, while preserving the transparency, customizability, and cost-efficiency required for responsible educational AI development, with deployment effectiveness depending on teacher judgment, learner context, and authentic instructional integration across diverse global learning environments.
Financial robo-advisors (FRAs), based on artificial intelligence (AI), are changing the ways people make investment decisions by providing advisory services that are based on algorithms and have different degrees of automation. But little evidence exists on the appropriateness of existing technology adoption theories to explain investor adoption of such systems. The present study investigates the explanatory and predictive power of three established and distinct adoption models for AI-enabled FRAs. It further develops a context-specific integrated adoption framework and examines whether adoption mechanisms differ between fully autonomous and hybrid robo-advisory services. For data collection, a quantitative research design was employed on 397 Indian investors through purposive sampling. The study uses explanatory and predictive assessment, integrated factor analysis, structural equation modeling, and multigroup analysis. The findings reveal that the SVAM outperforms the TPB and UTAUT-II in terms of explanatory and predictive power and underscores the significance of value perceptions and self-efficacy in AI-enabled financial decision-making. The integrated framework identified attitude, hedonic automatism, perceived value, self-efficacy, performance expectancy, effort expectancy, and normative control perception as significant antecedents of behavioural intention, which in turn drives usage behaviour. The results also indicate that the impact of adoption drivers differs across types of robo-advisors, with self-efficacy playing a more significant role in fully autonomous settings, and attitude, effort, and performance-related evaluations gaining more importance in hybrid environments. This study contributes to the literature on AI adoption by integrating competing theoretical perspectives into a context-specific framework and by establishing the type of FRA as an important boundary condition for investor behaviour.
Generative artificial intelligence is increasingly used in higher education, yet its educational value depends not only on technological access but also on how its use is pedagogically regulated. This study examined how teacher-regulated generative AI support is associated with university students' perceived learning gains in higher education. It further tested whether student agency mediates this association and whether perceived fairness conditions the strength of the association between teacher-regulated generative AI support and student agency. We adopted a cross-sectional quantitative survey design and collected data from 428 university students from six universities in western China who had used generative AI in course-related learning tasks under teacher guidance and within teacher-defined expectations. Structural equation modeling was used to test the direct and indirect associations in the proposed model, and a manifest interaction analysis was used to examine the first-stage moderation effect of perceived fairness. In the proposed model, teacher-regulated generative AI support was positively associated with perceived learning gains both directly and indirectly through student agency. Student agency was statistically consistent with a mediated pathway linking teacher-regulated generative AI support to perceived learning gains. In addition, perceived fairness was associated with a stronger positive relationship between teacher-regulated generative AI support and student agency. These findings suggest that students reported greater learning benefits when generative AI use was guided more clearly by teachers, taken up more actively by learners, and experienced in a fairer learning context. The study therefore highlights student agency as one plausible pathway linking teacher-regulated AI use to perceived learning gains in higher education.
Glucagon-like peptide-1 receptor agonists (GLP-1 RAs) are widely used for the management of type 2 diabetes mellitus and obesity; however, substantial inter-individual variability in glycemic and weight loss outcomes remains. This study aimed to develop and validate machine learning (ML) models to predict glycemic control and weight loss outcomes following GLP-1 RA initiation using real-world data and to identify key features associated with treatment response. We conducted a retrospective cohort study using data from the All of Us Research Program. Adult participants initiating GLP-1 RA therapy with available baseline and follow-up measurements were included. Two cohorts were constructed: a glycemic control cohort (n = 3,975) and a weight loss outcome cohort (n = 11,420). Glycemic control was defined as hemoglobin A1c (HbA1c) <7% at follow-up, and weight improvement was defined as achieving a body mass index (BMI) <30 kg/m². Multiple ML models, including logistic regression, random forest (RF), extreme gradient boosting (XGBoost), support vector machine, neural network, LightGBM, and CatBoost, were developed using 10-fold cross-validation. Model performance was evaluated using area under the receiver operating characteristic curve (AUC), accuracy, precision, recall, and precision-recall curves. SHapley Additive exPlanations (SHAP) were used to improve model interpretability. For weight loss outcome prediction, ensemble models demonstrated superior performance, with RF and XGBoost achieving the highest discrimination (AUC ≈ 0.94) and accuracy (0.89-0.90). For glycemic control prediction, RF and XGBoost achieved modest performance (accuracy ≈ 0.73; AUC ≈ 0.79). SHAP analysis identified baseline BMI and body weight as the most influential features of weight improvement, while duration of diabetes, baseline HbA1c, and use of sulfonylureas or insulin were among the most important features of glycemic control. Machine learning models, particularly tree-based ensemble methods, demonstrated strong potential for predicting treatment response to GLP-1 RA therapy. Integration of explainable ML approaches with real-world data may support personalized treatment strategies, and facilitate identification of patients most likely to benefit from GLP-1 RA therapy.
Diabetic retinopathy (DR) is a leading cause of preventable blindness, which has motivated the development of reliable automated grading systems on retinal fundus images. In this study, we perform a controlled comparative evaluation of ConvNeXt-Tiny, Swin-Tiny and their feature fusion for DR classification using the Asia Pacific Tele-Ophthalmology Society (APTOS) 2019 dataset. All models were initialized with weights pre-trained on ImageNet-1K and evaluated with two transfer learning strategies: direct fine-tuning on APTOS 2019, and EyePACS-based domain adaptation with task-specific fine-tuning. Systematic ablation experiments were carried out to evaluate the contribution of Contrast Limited Adaptive Histogram Equalization (CLAHE) preprocessing and channel-spatial attention modules (CSAM). We carried out experiments on the APTOS 2019 dataset with fixed train, validation and test splits and evaluated model stability across three runs with different random seeds by reporting mean ± standard deviation of performance metrics, while performance varied widely across architectures and training settings. After domain adaptation, the fusion-based models achieved more balanced results, while the standalone Swin-Tiny showed weaker adaptation to the retinal imaging domain, and was less sensitive to subtle lesion patterns under the EyePACS-based transfer learning. Adding CLAHE preprocessing and CSAM integration did not consistently improve class-balanced metrics. The best fusion configuration achieved a mean test accuracy of 88.34% ± 1.09 and a macro F1-score of 0.7376 ± 0.0183 on the APTOS 2019 dataset across repeated runs. These results suggest that domain-specific adaptation and architectural complementarity are more beneficial in boosting DR classification performance than auxiliary preprocessing or attention enhancement. The study also emphasizes the importance of controlled comparative evaluation, stability analysis, and configuration-specific evaluation in the research of medical image classification.
We introduce Grounded Multilingual Task Worlds for Romanian (GMTW-Ro), a benchmark designed to evaluate whether large language models can reliably follow complex instructions in Romanian, rather than merely produce fluent text. Existing Romanian benchmarks largely rely on multiple-choice formats, answer extraction, or model-based evaluation, which struggle to assess multi-constraint reasoning and structured task completion. GMTW-Ro addresses these limitations through grounded task worlds: fully specified environments in which model outputs are verified via deterministic, programmatic checks. The benchmark spans four task domains-travel planning, calendar scheduling, context-grounded question answering, and dietary menu planning-requiring both a structured JSON plan and a natural-language explanation in Romanian. Evaluation is decomposed into three orthogonal metrics: Understanding (U), measuring constraint adherence and instruction-following; Generation (G), assessing Romanian text quality through diacritic accuracy, language purity, and code-switching absence; and Faithfulness (F), quantifying consistency between generated plans and their explanations. All instances are automatically verified as solvable using backtracking algorithms. We release two curated datasets: a standard benchmark of 500 instances and an adversarial set of 300 instances with heightened constraint complexity, alongside the complete evaluation toolkit and a purpose-built Romanian NLP library. Evaluation of 11 models reveals substantial performance variation (58.6%-90.7%) and exposes a pronounced knowledge-behavior gap, where models with fluent Romanian generation nevertheless fail core reasoning tasks. Most notably, Romanian-finetuned models underperform their base counterparts: RoLlama3.1-8B scores 20.1 percentage points below Llama-3.1-8B, with structured JSON output success dropping from 95 to 44%. These results raise important questions about how current language adaptation pipelines preserve instruction-following and structured reasoning capabilities.
Early identification of heart disease is important to reduce mortality rates and to provide timely medical intervention for better patient outcomes. In recent studies, machine learning has been used to predict cardiovascular risk, but many existing models use basic feature sets and fixed decision rules. This can limit their ability to adapt when new data are introduced and may also reduce their capability to detect early signs of risk. In this study, we present a heart disease prediction framework that integrates acentric artificial intelligence (Ac-AI) with deep feature engineering to improve prediction accuracy. We adopted an autoencoder-based representation learning module to learn compact latent features from clinical data. These learned features were then combined with the original variables to provide a more informative set of inputs for subsequent analyses. The Ac-AI classifier applies cost-sensitive learning and an adaptive decision threshold to improve sensitivity for disease classes. Across multiple experiments conducted on a benchmark heart disease dataset, the proposed model showed performance comparable to that of random forest (RF), support vector machine (SVM), and XGBoost. In some runs, it performed better than these baseline models. Repeated independent runs provided evidence that the method produced consistent results across trials. These results suggest that Ac-AI combined with deep feature engineering can serve as a useful decision support framework for the early prediction of heart disease risk.
Artificial intelligence (AI) models for cervical cytology screening have achieved pooled accuracy and sensitivity values exceeding 90% in recent meta-analyses, and several commercial systems are now in clinical use. However, whether these results generalize across laboratories, scanners, and clinical settings depends on pre-analytical factors-sample preparation, staining, digitization, and annotation-that are known to introduce substantial variability into the data that models consume. This scoping review assessed how consistently these factors are documented in the cervical cytology AI literature. We examined 28 datasets published between 2005 and 2025, extracting information on 16 pre-analytical variables spanning sample preparation, digitization, and annotation. The mean reporting completeness was 11.4 out of 16 variables (71.2%). Digitization was the weakest category (mean 60.2%), with scanning mode unreported in 57.1% of datasets, image file format in 60.7%, and color normalization status in 78.6%. Staining protocol was mentioned by 75.0% of datasets but described in sufficient procedural detail by only 7.1%. Quantitative inter-annotator agreement was provided by 14.3% of datasets, despite well-documented inter-observer variability in cervical cytology. Notably, the variables with the lowest reporting rates correspond to those identified in the digital pathology literature as the most significant sources of AI model performance variability. To address this gap, we propose PRECY-AI (Pre-analytical Reporting Checklist for Cervical Cytology AI), a 16-item checklist of essential and recommended reporting items designed to complement existing general-purpose guidelines such as TRIPOD+AI and CLAIM. Adoption of domain-specific pre-analytical reporting standards could improve the reproducibility, comparability, and clinical translatability of cervical cytology AI research.
Evaluation metrics are essential for assessing the performance of AI models and enabling comparability across different approaches. However, in surgical phase recognition, their application remains inconsistent. This study provides a systematic analysis of evaluation metrics used in this field, aiming to support standardization and encourage more explicit explanation of metric selection. A systematic literature review was conducted following the PRISMA framework, covering publications from 2016 to 2025. In total, 50 studies were identified and analyzed with respect to the evaluation metrics applied, with particular focus on relaxed boundaries and the distinction between online and offline models. The results show that accuracy, precision, and recall are the most frequently used metrics, followed by the jaccard index. Since around 2023, an increasing use of segment-based metrics can be observed, reflecting a growing emphasis on temporal dynamics. No significant differences in metric selection between online and offline approaches were identified. Furthermore, many studies do not provide explicit explanation for their choice of evaluation metrics. Instead, they often rely on commonly used metrics or practices adopted from previous work. Additionally, results obtained with relaxed boundaries tend to exhibit higher standard deviations compared to those without relaxed boundaries. Based on these findings, it is recommended to more explicitly align evaluation metrics with model objectives and to further investigate the distinction between online and offline settings. Adaptations of the f 1-score, such as the f β-score, may provide a more flexible evaluation framework. Furthermore, the Matthews Correlation Coefficient represents a promising complementary metric, particularly for multi-class settings, if appropriately normalized. Future work should focus on clearly explaining metric selection and consistently reporting relaxed boundaries to improve transparency, comparability, and interpretability in surgical phase recognition.
Globally, 1.3 billion tons of food is lost or wasted each year, negatively impacting food security, the economy, and the climate. Fresh fruits and vegetables (FFVs), with their short shelf life and temperature sensitivity, are the most affected. This study systematically evaluates the integration of Machine Learning (ML), Adaptive Learning (AL), the Internet of Things (IoT), and Fog computing for temperature-break detection and prediction in FFVs supply chains. It critically evaluates their individual and combined capabilities, identifying compounding barriers that prevent genuine real-time integration of these technologies, while assessing their performance and operational readiness for real-time cold chain monitoring. Additionally, the role of fog computing in enabling efficient ML/AL deployment at the network edge is investigated for real-time applications in dynamic environments, with an implementation framework provided. Based on the PRISMA framework, searches of Scopus, Web of Science, IEEE Xplore, ACM Digital, supplemented by citation and reference chasing, produced 830 pre?deduplication records, identifying 14 relevant studies. From the 12 analysed unique-dataset studies, 7 (58.3%) collected data using Basic Sensors, 3 with WSN (25%), and 2 with IoT (16.7%). Of the 14 ML studies, 5 (35.7%) detected temperature breaks, 5 (35.7%) predicted FFVs' temperature values, 3 (21.4%) predicted internal temperature (IT) values of a cold room or container, and 1 predicted IT values and time-to-temperature breaks. None of the studies predicted temperature breaks (event occurrence) or even their causes, while a few detected these breaks and their predefined causes. Four (28.6%) studies used IoT data, but none enabled live ML inference. None of the reviewed studies includes Fog or AL. An integrated IoT-Fog-AL framework is thus proposed to address these gaps. These areas require focus to proactively reduce temperature breaks, thereby minimising food wastage and its associated effects, while also enhancing supply chain resilience and food security.
Standard Retrieval-Augmented Generation (RAG) systems only use semantic similarity to retrieve information, and since this method is quite limiting, it may find clinically irrelevant evidence and produce outputs that are unsafe or hallucinated. This drawback is particularly important in respiratory care, where the diagnosis relies heavily on very accurate physiological indicators such as spirometry patterns and symptom profiles. We propose MMCRAG-Resp., a clinically grounded, physiology-aware corrective RAG framework for respiratory intelligence. The system harmonizes 17,516 patient records from three heterogeneous public datasets (Respiratory Sound Database, NHANES spirometry, and clinical guidelines) through a unified respiratory ontology. Multi-modal embeddings (acoustic, spirometric, and textual) are fused into 144-dimensional vectors and indexed in FAISS. A novel Respiratory Relevance Score (RRS) combining spirometric pattern agreement (weight 0.45), symptom overlap (0.30), and embedding similarity (0.25) gates evidence quality before LLM consumption via a three-path correction policy. A locally deployed small language model (Ollama; llama3.2, Mistral-7B-Instruct) generates evidence-constrained, fully traceable clinical explanations without GPU requirements. Evaluated on 200 held-out test queries drawn from the harmonized corpus, MMCRAG-Resp achieved a Clinical Precision@10 of 0.81, a Hallucination Rate of 0.11 (71% relative reduction over Vanilla RAG), a Spirometry Alignment Score of 0.79, and an Evidence Citation Rate of 0.87. Ablation studies confirmed that spirometry-first scoring contributes the largest single-component gain. All pairwise comparisons with baselines were statistically significant (p < 0.001, McNemar's test; Cohen's d = 1.42). MMCRAG-Resp is, to the best of our knowledge, the first system combining multi-modal respiratory evidence retrieval, physiologically-grounded corrective RAG, and privacy-preserving local LLM inference in a single deployable clinical reasoning framework. The approach advances safe, explainable AI-assisted respiratory decision support.