Declining ophthalmology teaching hours necessitate efficient instructional tools. While generative AI frequently produces structural deviations, these variations can be strategically repurposed as valuable stimuli for comparative learning. This study evaluated the effects of an AI-assisted comparative exercise versus conventional anatomical labeling on knowledge acquisition, learner satisfaction, and cognitive workload. We conducted a quasi-experimental 2 × 2 study with 121 sophomores from two universities in Shanghai, China. Following a standardized 20-min ophthalmic anatomy lecture, intact classes were assigned to either a conventional diagram-labeling task or an AI-assisted comparative exercise. The AI condition included three anatomically correct reference images paired with three expert-curated AI-generated anatomical variants, produced through a systematic expert-in-the-loop approach using Gemini 3.0 Pro. Outcomes comprised baseline and post-intervention knowledge tests, a 5-item satisfaction questionnaire, and the NASA Task Load Index. After baseline adjustment and correction for planned comparisons, no statistically significant AI-versus-conventional difference in post-test knowledge scores was detected (all p > 0.05). Among non-medical students, the AI-assisted comparative exercise was associated with higher composite satisfaction (9.17 vs. 7.52; Holm-adjusted p < 0.001; r = 0.60) and better self-assessed performance (7.76 vs. 6.03; Holm-adjusted p = 0.003; r = 0.47), whereas composite NASA-TLX scores did not differ between AI and conventional conditions within either background group. These satisfaction and self-assessed performance benefits were not observed among medical students (all p > 0.05). A brief AI-assisted comparative exercise did not demonstrate a statistically conclusive advantage in immediate knowledge outcomes, but was associated with higher satisfaction and better self-assessed performance among non-medical students without increasing composite NASA-TLX scores. Carefully curated AI-generated anatomical variants may therefore serve as a structured adjunct for novice ophthalmic anatomy learning.
To evaluate the accuracy, completeness, clarity, source transparency, and readability of leading AI chatbot responses to patient questions about tracheostomy and to determine whether AI tools can reliably support patient education where high-quality guidance is critical for safety. Cross-sectional content analysis. Virtual study environment using publicly accessible AI platforms, with expert evaluation conducted via Qualtrics-based distribution. Twelve frequently asked questions about tracheostomy care were identified using search-listening tools and clinician input, then submitted to 5 AI chatbots - ChatGPT4, Google Gemini 2.0, Microsoft Copilot, DeepSeek V3, and Grok 3 - and to a senior laryngologist. Three blinded laryngologists independently evaluated each response using the Quality Analysis of Medical Artificial Intelligence instrument. Readability was assessed using nine metrics. Gemini 2.0 achieved significantly higher completeness scores than physician responses (P < .001), with DeepSeek and Grok 3 (P < .05) also outperforming (P < .05). Accuracy did not differ significantly between AI- and expert-generated responses. On average, the AI models outperformed physician in clarity, completeness, and usefulness based on QAMAI scoring (P < .05). All AI and expert responses exceeded the NIH-recommended 6th-grade reading level, ranging from 10th-13th grade (P < .001). Inter-rater reliability was 78%. AI chatbots can generate accurate and comprehensive responses to common tracheostomy care questions, demonstrating potential to support patient education. However, they continue to lack guaranteed, verifiable sourcing, and this study did not assess actual patient comprehension of the AI-generated responses. Future efforts should focus on adapting AI-generated education materials to meet health literacy standards and evaluating their direct impact on patient understanding and outcomes.
With the rapid advancement of artificial intelligence (AI), increasing numbers of Chinese patients incorporate AI tools into clinical consultations. However, directly introducing AI-generated recommendations into the clinical setting often challenges the authority of physicians and may provoke resistance. Therefore, this study explores how patients strategically integrate AI suggestions into medical consultations to preserve harmonious doctor-patient relationships while benefiting from AI. We conducted in-depth interviews with 60 Chinese stakeholders, including both 35 patients and 25 clinicians, and employed an inductive thematic analysis to code and interpret the data. The findings reveal that patients render AI-generated outputs clinically acceptable through four semantic strategies: provenance reframing, cue selection, modal attenuation, and sequential embedding. In the short term, these tactics sustain physician-patient harmony while enabling the integration of algorithmic knowledge into medical decision making. Over time, however, they give rise to a trust-transfer dynamic in which patients progressively benchmark clinicians against AI's explanatory precision and evidentiary transparency, thereby privileging algorithmic rationales over professional judgment. We propose the adoption of clinician-led, AI-augmented models, the establishment of standardized communication protocols, targeted clinician training, and the implementation of robust validation procedures to ensure that AI effectively complements professional judgment, fosters patient trust, and preserves the integrity of clinical decision-making.
Artificial intelligence generated advertising is widely adopted in retailing and consumer services, yet its effects on trust, engagement, and purchase intention remain theoretically inconsistent. This study addresses this anomaly by reconceptualising advertising effectiveness under algorithmic authorship as a process of signal resolution rather than additive persuasion. Drawing on signalling theory and advertising value theory, the study specifies advertising value as a formative signal system composed of informativeness, entertainment, and executional credibility, which simultaneously activates competing inferential pathways of perceived credibility and perceived eeriness. Using a theory-driven PLS-SEM model estimated on a quota-based U.S. consumer sample (N = 412), the results provide associative evidence of systematic asymmetry and suppression effects: identical executional cues strengthen credibility while concurrently amplifying eeriness, with trust patterns consistent with the relative dominance of these opposing inferences rather than from overall message quality. Trust functions as a conditional transmission mechanism to engagement and purchase intention, and AI disclosure is associated with shifts in signal weighting by attenuating credibility-based pathways and amplifying eeriness-based suppression. By identifying signal competition as a structural feature associated with AI-generated advertising, the study extends current theoretical understanding of algorithmic persuasion by introducing signal competition as a structural feature associated with AI-generated advertising, departing from human-centric models and clarifying why creative AI execution often fails to yield behavioural conversion.
As the use of artificial intelligence (AI) and large language models (LLMs) is increasingly adopted into scientific writing, it is important to understand AI's ability to produce clear and accurate content that is on par with human-authored content in the field of orthopaedics, including orthopaedic oncology. The aim of this study was to compare a series of editorials written by orthopaedic oncologists with those written by a single LLM (ChatGPT 4.0) using a variety of quality metrics. Volunteer orthopaedic oncologists submitted a 3- to 4-paragraph persuasive editorial on a topic of their choice in the field of musculoskeletal oncology. ChatGPT 4.0 was then prompted to write a corresponding editorial for each topic. Each editorial was evaluated by two blinded peer reviewers and graded using a 25-point scale on the following quality metrics: content, clarity, grammar, persuasiveness, and creativity. The evaluators were also asked to indicate whether they believed the editorials were written by humans or by AI. A total of 20 editorials were submitted by human authors and matched with 20 prompted AI editorials. No notable difference in average total quality score for human versus AI submissions was observed. AI-generated articles scored markedly higher in grammar, but there were no notable differences in any other quality metric. Reviewers correctly identified author type 59% of the time. LLMs such as ChatGPT can generate editorial content in orthopaedic oncology that matches human-written quality, suggesting a potential supportive role for AI in scientific communication, with implications for authorship standards, editorial practices, and peer review.
暂无摘要(点击查看详情)
暂无摘要(点击查看详情)
暂无摘要(点击查看详情)
暂无摘要(点击查看详情)
暂无摘要(点击查看详情)
Generative AI coding assistants are increasingly used to write machine-learning code, yet their ability to produce reliable LSTM implementations for financial prediction remains underexplored. This study evaluates the LSTM code generated by seven assistants ChatGPT 4.5, GitHub Copilot, Deepseek 3, Perplexity, Gemini 2.0 Pro, Claude 3.7 Sonnet, and Meta's Llama from a single standardized prompt, on three indices (Nikkei 225, S&P 500, STOXX Europe 600). Each assistant's generated script was re-executed over independent runs; accuracy (MAE, MSE, RMSE, R2, execution time) is reported as mean ± standard deviation on the original price scale, complemented by a static code-quality analysis (Pylint, Radon, SonarQube, Pytest, Bandit). The assistants converge on nearly identical LSTM architectures, so performance differences arise mainly from data-handling and code-correctness defects: Meta's Llama near-zero errors are an artifact of normalized-scale metrics combined with a shuffled train/test split (data leakage), and once corrected its accuracy is among the weakest; Gemini 2.0 Pro, once its predictions are evaluated consistently on the price scale, is among the most accurate assistants. Differences are validated with Diebold-Mariano and Wilcoxon tests. AI-generated forecasting code can be accurate but is not uniformly trustworthy: its generated preprocessing and evaluation code must be audited before use.
Artificial intelligence (AI) tools, particularly large language models, are increasingly integrated into medical education. However, little is known about how early medical students perceive their educational value, reliability, and ethical implications. A cross-sectional survey was conducted among first-year students in the Undergraduate (MD) and Graduate-Entry (GEMD) Medical Programmes at the University of Nicosia, Cyprus. A total of 102 students participated across three cohorts: MD2030 (n = 44), MD2031 (n = 46), and GEMD2030 (n = 12). The questionnaire explored awareness and use of AI tools, perceived benefits and risks, trust in AI-generated medical information, ethical concerns, and attitudes toward AI in medical education. Quantitative data were analysed using descriptive and inferential statistics, while open-ended responses underwent sentiment analysis and topic modelling. Awareness of AI tools was nearly universal (93-100%), with ChatGPT use reported by 87-100% of students. Regular use ranged from 70.5% in MD2030 to 100% in GEMD2030. Students perceived AI as a valuable supplementary learning tool (mean Likert scores 4.07-4.22/5), particularly for improving understanding and efficiency. Trust in AI-generated medical information remained moderate (3.11-3.33/5), reflecting concerns about inaccuracies and ethical issues. Open-ended responses highlighted incorrect AI-generated information as a common concern. A significant difference in reported encounters with incorrect AI-generated content was observed between cohorts (χ 2 = 10.73, p = 0.0047), with lower reporting among students exposed to AI-literacy training, suggesting an association between prior exposure to AI literacy activities and differences in how students report AI-generated errors. AI tools are widely adopted by medical students and are perceived as enhancing rather than replacing traditional study methods. Structured AI literacy activities may be associated with differences in how students evaluate AI-generated content; however, further longitudinal and controlled studies are required to determine whether such educational experiences have a measurable impact on learners' attitudes and practices.
Generative artificial intelligence (AI) is changing how citizens search for, understand, and use health information. Large language models (LLMs) can simplify, translate, summarize, and tailor health-related content to individual questions. At the same time, AI-generated responses may appear fluent, plausible, and empathic without necessarily being complete, up to date, evidence-based, or suitable for a person's individual situation. This creates a public health challenge for Public Health Education and Promotion, as AI-generated answers may increasingly shape how citizens access, interpret, and trust health information. Building on health literacy, eHealth literacy, digital health literacy, and debates on AI-mediated health communication, this article conceptualizes AI Health Literacy as a task- and context-sensitive extension of existing literacy concepts. Its specific contribution is to understand AI Health Literacy as an interdisciplinary judgement competence that requires citizens to assess not only the content of AI-generated health information, but also the task, context, modality, verifiability, and psychological conditions under which such information is received and used. It proposes a reflective framework with three analytical dimensions: task appropriateness, context of use, and critical verifiability. The central argument is that generative AI can support understanding, orientation, translation, and preparation, but should not replace professional advice in diagnostic, therapeutic, triage-related, medication-related, or crisis situations. The framework aims to support public-health-oriented education, communication, and institutional guidance for the responsible use of AI-generated health information and should be understood as a conceptual orientation tool, not as a validated assessment instrument.
In the contemporary landscape of information technology, the advent of artificial intelligence (AI) in music composition marks a transformative era. This study examines differences in emotional expression accuracy between AI-generated music and compositions by traditional musicians, with a particular focus on Suno AI within the theoretical framework of Plutchik's "Wheel of Emotions." Emotional expression accuracy is defined as the degree of correspondence between the intended target emotion of a musical piece and the emotion perceived by listeners. An empirical approach was adopted to compare AI-generated and human-composed music and to evaluate the accuracy of emotional expression across four primary high-intensity emotions (ecstasy, rage, grief, and terror). A total of 32 musical pieces were analyzed, comprising 16 compositions by traditional composers and 16 works generated by Suno AI. Survey data were collected from 300 university students without formal music training, yielding 283 valid responses. The results revealed that emotional expression accuracy differed significantly between human-composed and AI-generated music for rage (χ2 = 40.66, p < 0.001, Cramer's V = 0.27), ecstasy (χ2 = 25.53, p < 0.001, V = 0.21), and terror (χ2 = 16.46, p < 0.001, V = 0.17), whereas no significant difference was observed for grief (χ2 = 2.44, p = 0.12, V = 0.07). Across emotions, the largest accuracy gap was observed for rage (25.40%), followed by ecstasy (21.50%) and terror (17.00%), while the difference for grief was comparatively smaller (4.20%) and not statistically significant. These findings suggest that the advantage of human-composed music over AI-generated music may vary across emotional categories, with more pronounced differences observed for high-arousal emotions. The results further indicate that, despite substantial advances in generative systems such as Suno AI, current AI models may still face limitations in representing high-intensity emotional states. Future research could further explore improvements in emotional conditioning mechanisms and training data diversity to enhance the emotional expression accuracy of AI-generated music.
As large language models (LLMs) are increasingly integrated into decision-making systems (e.g., autonomous vehicles and medical devices), understanding how humans perceive and evaluate AI-generated judgments is crucial. To investigate this, we conducted a series of experiments in which participants evaluated justifications for moral and non-moral choices, generated either by humans or LLMs. Participants attempted to identify the source of each justification (either human or LLM) and indicated their agreement with its content. We found that while detection accuracy was consistently above chance, it remained below 75%. In terms of agreement, there was no overall preference for human-generated responses, even though machine-generated justifications were favored in particularly challenging moral scenarios. Notably, we observed a systematic anti-AI bias: participants were less likely to agree with judgments they believed were AI-generated, regardless of the true source. Linguistic cues, such as response length, typos, first-person pronouns, and cost-benefit language markers (e.g., "lives," "save"), influenced both detection and agreement. Participants tended to disagree with cost-benefit calculations, possibly due to an expectation that AI would favor such reasoning. These findings highlight the influence of motivated belief and ingroup/outgroup bias in shaping human evaluation of AI-generated content, particularly in morally sensitive contexts.
The European Cystic Fibrosis Society (ECFS) develops education resources to support members; however, these are almost exclusively in English. Many barriers to translation exist, including cost and time. Artificial intelligence (AI) provides an opportunity to support translation and address such barriers. This study aimed to pilot the use of AI-generated translation of ECFS e-learning modules and evaluate the quality. An AI translation program was used to create subtitles of ECFS peer-reviewed education modules. Two independent native language speakers with extensive cystic fibrosis (CF) healthcare experience were identified and tasked with reviewing, editing, and validating. This was followed by the development and circulation of an online evaluation survey assessing users' views on quality. Education packages, each consisting of six subtitled modules, were created in three languages: Ukrainian, Romanian, and Turkish. For each language, corrections to the AI-generated translation by the independent native speakers were essential. Evaluation was conducted in two countries. Eighteen completed surveys were received. Results indicated high levels of accuracy for the final modules, and feedback was very positive regarding the utility and range of topics. The use of novel AI-generated translation shows promise and proved quick and affordable. However, quality of translation was variable, highlighting the critical role of collaborating with native-speaking CF experts to ensure linguistic accuracy. This project highlights the importance of interdisciplinary collaborative efforts between ECFS Education, the Twinning Project, CF Europe, and patient organizations. Further, it demonstrates both the feasibility and practicality of generating effective multilingual educational modules using AI.
Large language models (LLMs) can generate fluent summaries of longitudinal medical records, but in high-stakes clinical settings, verification burden remains a barrier to trust. Existing provenance mechanisms, such as document-level citations and section references, often require manual search within long, fragmented notes, limiting their usefulness during time-constrained workflows for clinicians. To design and evaluate a sentence-level provenance interface ("click-to-inspect") that enables rapid verification of AI-generated longitudinal medical record summaries at the level of individual statements. Between November 2023 and January 2024, we conducted a formative usability study using remote moderated usability sessions via Zoom to evaluate a web-based sentence-level provenance interface for AI-generated longitudinal medical record summaries. A convenience sample of clinicians was recruited through email outreach to academic and professional networks across the United States. Formative usability testing was conducted with 46 clinician interactions using synthetic longitudinal patient charts. Participants included medical students, residents, and attending physicians across multiple specialties including internal medicine, dermatology, radiology, plastic surgery, anesthesiology, interventional radiology, obstetrics-gynecology, and family medicine. Usability was assessed using the System Usability Scale (SUS) and Net Promoter Score (NPS), alongside qualitative feedback. Clinicians reported high usability (mean SUS score 86.25, SD 7.77; 95% CI 83.96-88.54 from 46 participants) and a positive overall experience (NPS 35; 22/46 promoters, 18/46 passives, 6/46 detractors). Participants described rapid access to supporting evidence as critical for trust calibration during first-pass chart review. Qualitative feedback identified friction in traditional citation-based interfaces and supported sentence-level inspectability as a low-friction verification mechanism. Sentence-level provenance transforms AI-generated summaries from static narratives into interactive verification tools. An approach that enables rapid, selective inspection of individual claims during longitudinal chart review, may reduce verification burden and support calibrated reliance in high-risk clinical contexts.
Informed consent is a core ethical and legal requirement in oral and maxillofacial surgery (OMFS). Traditional consent documents often lack adequate readability and comprehensive coverage of ethical and safety elements. Large language models (LLMs) may assist in drafting consent forms; however, their alignment with clinician expectations in OMFS remains insufficiently studied. This in silico descriptive evaluation study assessed AI-generated informed consent drafts using the QUEST (Quality, Utility, Ethics, Safety, and Transparency) framework. Consent drafts were generated for four representative OMFS procedures: third molar extraction, dental implant placement, open reduction and internal fixation of mandibular fracture, and soft tissue biopsy. One draft was produced per procedure. Two experienced oral and maxillofacial surgeons independently rated each draft across QUEST domains using a 5-point Likert scale (1 = poor, 3 = acceptable, 5 = excellent). Readability was evaluated using standard grade-level metrics. Inter-rater reliability was assessed using Cohen's kappa (κ), with significance set at p < 0.05. Four AI-generated consent drafts were evaluated. Readability scores ranged from grade 6.0 to 8.0 (mean 6.85 ± 0.85). Inter-rater reliability was substantial (κ = 0.78; p < 0.001). Among QUEST domains, Quality achieved the highest scores (mean 4.1 ± 0.4), while Ethics, Safety, and Transparency demonstrated comparatively lower ratings. AI-generated informed consent drafts showed acceptable readability and procedural descriptions but demonstrated limitations in ethical, safety, and transparency domains. These findings suggest that LLMs may assist in consent drafting; however, clinician oversight and further validation are required before routine clinical adoption. The online version contains supplementary material available at 10.1007/s12663-026-03085-7.
In a 6-month randomized trial, we evaluated a digitally enabled "human-in-the-loop" care support model using a predictive artificial intelligence (AI) digital twin to provide personalized daily short message service (SMS) feedback for adults with type 2 diabetes (T2D). The parent study enrolled 40 adults aged ≥18 years with T2D who completed 3 months of baseline observation followed by a 3-month intervention period, generating 6467 longitudinal data points across weight, dietary intake, physical activity, and glucose monitoring (mean follow-up: 174 days). For this ancillary AI intervention, a subset of 19 participants was randomized to receive either AI-generated individualized daily feedback (AI group, n = 10) or no daily feedback (control group, n = 9). The online human-in-the-loop predictive control model incorporated a transfer-learning artificial neural network predictive digital twin trained on participant self-monitoring data, including weight, food logs, physical activity, and glucose values. A particle swarm optimization controller identified personalized behavioral recommendations aligned with glucose and weight goals, and the digital twin was retrained weekly using newly accrued data. The model achieved ≥80% prediction accuracy across all diet-condition subgroups. During the intervention period, participants receiving AI-generated feedback demonstrated trends toward increased daily step counts and improved adherence to caloric and carbohydrate intake targets. The AI intervention group achieved significantly greater weight loss than controls (mean loss 5.87 lbs vs 3.57 lbs; p < 0.012) while maintaining stable glucose levels throughout the study period (p = 0.661). These findings suggest that AI-enabled predictive digital twin models may offer a scalable approach for extending precision diabetes self-management support beyond clinic visits.
Early autism diagnosis remains challenging due to reliance on clinical observation and limited specialist availability. Addressing these barriers through automated diagnostic labeling and the integration of parental input may help mitigate the problem. We trained a BioBERT machine learning model to label individual autism behavioral descriptions using the seven DSM-5 diagnostic criteria (A1-A3, B1-B4). This approach offers transparent clinical decision-making by providing detailed diagnostic information for individual behaviors and avoiding final case-level black-box decisions. We evaluated the model's performance on labeling lay (N = 35,971) and clinical (N = 145,603) behavior descriptions, as well as its transferability between the two. In addition, we compared the data sources by evaluating the diagnostic utility of lay and clinical examples across four dimensions, and of AI-generated summaries across two dimensions. We found that BioBERT can label both types of input, although it achieved higher precision (69%) on clinical descriptions and higher recall (83%) on lay descriptions. Sample size did not explain differences in performance. Transferring models from one data type to another results in a performance drop. Overall, training first on clinical data yielded the best-performing diagnostic models. When evaluating the examples from both data sources, the results show similar scores for the Utility, Specificity, Clinical Relevance, and Impact on Daily Life dimensions, and the cosine similarity analysis revealed substantial overlap (0.42) in vocabulary between the two. The utility of examples for A diagnostic behaviors was generally scored higher than that for B diagnostic behaviors. AI-generated summary scores showed a similar pattern between A and B examples but they were only moderately representative of these examples. These results demonstrate that lay behavioral descriptions can provide diagnostically valuable information comparable to clinical observations, although they are not readily summarized by AI. The integration of lay information into the diagnostic workflows could accelerate autism diagnosis without compromising clinical utility.