Objective To develop an LLM based realtime compound diagnostic medical AI interface and performed a clinical trial comparing this interface and physicians for common internal medicine cases based on the United States Medical License Exam (USMLE) Step 2 Clinical Skill (CS) style exams. Methods A nonrandomized clinical trial was conducted on August 20, 2024. We recruited one general physician, two internal medicine residents (2nd and 3rd year), and five simulated patients. The clinical vignettes were adapted from the USMLE Step 2 CS style exams. We developed 10 representative internal medicine cases based on actual patients and included information available on initial diagnostic evaluation. Primary outcome was the accuracy of the first differential diagnosis. Repeatability was evaluated based on the proportion of agreement. Results The accuracy of the physicians' first differential diagnosis ranged from 50% to 70%, whereas the realtime compound diagnostic medical AI interface achieved an accuracy of 80%. The proportion of agreement for the first differential diagnosis was 0.7. The accuracy of the first and second differential diagnoses ranged from 70% to 90% for physicians, whereas
In this paper, we introduce InMD-X, a collection of multiple large language models specifically designed to cater to the unique characteristics and demands of Internal Medicine Doctors (IMD). InMD-X represents a groundbreaking development in natural language processing, offering a suite of language models fine-tuned for various aspects of the internal medicine field. These models encompass a wide range of medical sub-specialties, enabling IMDs to perform more efficient and accurate research, diagnosis, and documentation. InMD-X's versatility and adaptability make it a valuable tool for improving the healthcare industry, enhancing communication between healthcare professionals, and advancing medical research. Each model within InMD-X is meticulously tailored to address specific challenges faced by IMDs, ensuring the highest level of precision and comprehensiveness in clinical text analysis and decision support. This paper provides an overview of the design, development, and evaluation of InMD-X, showcasing its potential to revolutionize the way internal medicine practitioners interact with medical data and information. We present results from extensive testing, demonstrating the eff
Model Medicine is the science of understanding, diagnosing, treating, and preventing disorders in AI models, grounded in the principle that AI models -- like biological organisms -- have internal structures, dynamic processes, heritable traits, observable symptoms, classifiable conditions, and treatable states. This paper introduces Model Medicine as a research program, bridging the gap between current AI interpretability research (anatomical observation) and the systematic clinical practice that complex AI systems increasingly require. We present five contributions: (1) a discipline taxonomy organizing 15 subdisciplines across four divisions -- Basic Model Sciences, Clinical Model Sciences, Model Public Health, and Model Architectural Medicine; (2) the Four Shell Model (v3.3), a behavioral genetics framework empirically grounded in 720 agents and 24,923 decisions from the Agora-12 program, explaining how model behavior emerges from Core--Shell interaction; (3) Neural MRI (Model Resonance Imaging), a working open-source diagnostic tool mapping five medical neuroimaging modalities to AI interpretability techniques, validated through four clinical cases demonstrating imaging, compari
We have developed a Scalable CI/CD Pipeline to address internal challenges related to Japan 2025 cliff problem, a critical issue where the mass end of service life of legacy core IT systems threatens to significantly increase the maintenance cost and black box nature of these system also leads to difficult update moreover replace, which leads to lack of progress in Digital Transformation (DX). If not addressed, Japan could potentially lose up to 12 trillion yen per year after 2025, which is 3 times more than the cost in previous years. Asahi also faced the same internal challenges regarding legacy system, where manual maintenance workflows and limited QA environment have left critical systems outdated and difficult to update. Middleware and OS version have remained unchanged for years, leading to now its nearing end of service life which require huge maintenance cost and effort to continue its operation. To address this problem, we have developed and implemented a Scalable CI/CD Pipeline where isolated development environments can be created and deleted dynamically and is scalable as needed. This Scalable CI/CD Pipeline incorporate GitHub for source code control and branching, Jenk
What does Artificial Intelligence (AI) have to contribute to health care? And what should we be looking out for if we are worried about its risks? In this paper we offer a survey, and initial evaluation, of hopes and fears about the applications of artificial intelligence in medicine. AI clearly has enormous potential as a research tool, in genomics and public health especially, as well as a diagnostic aid. It's also highly likely to impact on the organisational and business practices of healthcare systems in ways that are perhaps under-appreciated. Enthusiasts for AI have held out the prospect that it will free physicians up to spend more time attending to what really matters to them and their patients. We will argue that this claim depends upon implausible assumptions about the institutional and economic imperatives operating in contemporary healthcare settings. We will also highlight important concerns about privacy, surveillance, and bias in big data, as well as the risks of over trust in machines, the challenges of transparency, the deskilling of healthcare practitioners, the way AI reframes healthcare, and the implications of AI for the distribution of power in healthcare ins
The diffusion model presents a powerful ability to capture the entire (conditional) data distribution. However, due to the lack of sufficient training and data to learn to cover low-probability areas, the model will be penalized for failing to generate high-quality images corresponding to these areas. To achieve better generation quality, guidance strategies such as classifier free guidance (CFG) can guide the samples to the high-probability areas during the sampling stage. However, the standard CFG often leads to over-simplified or distorted samples. On the other hand, the alternative line of guiding diffusion model with its bad version is limited by carefully designed degradation strategies, extra training and additional sampling steps. In this paper, we proposed a simple yet effective strategy Internal Guidance (IG), which introduces an auxiliary supervision on the intermediate layer during training process and extrapolates the intermediate and deep layer's outputs to obtain generative results during sampling process. This simple strategy yields significant improvements in both training efficiency and generation quality on various baselines. On ImageNet 256x256, SiT-XL/2+IG achi
With the increasing interest in deploying Artificial Intelligence in medicine, we previously introduced HAIM (Holistic AI in Medicine), a framework that fuses multimodal data to solve downstream clinical tasks. However, HAIM uses data in a task-agnostic manner and lacks explainability. To address these limitations, we introduce xHAIM (Explainable HAIM), a novel framework leveraging Generative AI to enhance both prediction and explainability through four structured steps: (1) automatically identifying task-relevant patient data across modalities, (2) generating comprehensive patient summaries, (3) using these summaries for improved predictive modeling, and (4) providing clinical explanations by linking predictions to patient-specific medical knowledge. Evaluated on the HAIM-MIMIC-MM dataset, xHAIM improves average AUC from 79.9% to 90.3% across chest pathology and operative tasks. Importantly, xHAIM transforms AI from a black-box predictor into an explainable decision support system, enabling clinicians to interactively trace predictions back to relevant patient data, bridging AI advancements with clinical utility.
Recent studies indicate that Generative Pre-trained Transformer 4 with Vision (GPT-4V) outperforms human physicians in medical challenge tasks. However, these evaluations primarily focused on the accuracy of multi-choice questions alone. Our study extends the current scope by conducting a comprehensive analysis of GPT-4V's rationales of image comprehension, recall of medical knowledge, and step-by-step multimodal reasoning when solving New England Journal of Medicine (NEJM) Image Challenges - an imaging quiz designed to test the knowledge and diagnostic capabilities of medical professionals. Evaluation results confirmed that GPT-4V performs comparatively to human physicians regarding multi-choice accuracy (81.6% vs. 77.8%). GPT-4V also performs well in cases where physicians incorrectly answer, with over 78% accuracy. However, we discovered that GPT-4V frequently presents flawed rationales in cases where it makes the correct final choices (35.5%), most prominent in image comprehension (27.2%). Regardless of GPT-4V's high accuracy in multi-choice questions, our findings emphasize the necessity for further in-depth evaluations of its rationales before integrating such multimodal AI m
// Takayuki Takahama 1, * , Kazuko Sakai 2, * , Masayuki Takeda 1, * , Koichi Azuma 3 , Toyoaki Hida 4 , Masataka Hirabayashi 5 , Tetsuya Oguri 6 , Hiroshi Tanaka 7 , Noriyuki Ebi 8 , Toshiyuki Sawa 9 , Akihiro Bessho 10 , Motoko Tachihara 11 , Hiroaki Akamatsu 12 , Shuji Bandoh 13 , Daisuke Himeji 14 , Tatsuo Ohira 15 , Mototsugu Shimokawa 16 , Yoichi Nakanishi 17 , Kazuhiko Nakagawa 1 , Kazuto Nishio 2 1 Department of Medical Oncology, Kinki University Faculty of Medicine, Osaka, Japan 2 Department of Genome Biology, Kinki University Faculty of Medicine, Osaka, Japan 3 Department of Internal Medicine, Division of Respirology, Neurology, and Rheumatology, Kurume University School of Medicine, Kurume, Fukuoka, Japan 4 Department of Thoracic Oncology, Aichi Cancer Center, Nagoya, Japan 5 Department of Respiratory Medicine, Hyogo Prefectural Amagasaki General Medical Center, Hyogo, Japan 6 Department of Medical Oncology and Immunology, Nagoya City University Graduate School of Medical Sciences, Nagoya, Japan 7 Department of Internal Medicine, Niigata Cancer Center Hospital, Niigata, Japan 8 Department of Respiratory Oncology, Iizuka Hospital, Fukuoka, Japan 9 Department of Respiratory Medicine and Oncology, Gifu Municipal Hospital, Gifu, Japan 10 Department of Pulmonary Medicine, Japanese Red Cross Okayama Hospital, Okayama, Japan 11 Division of Respiratory Medicine, Department of Internal Medicine, Kobe University Graduate School of Medicine, Hyogo, Japan 12 Third Department of Internal Medicine, Wakayama Medical University, Wakayama, Japan 13 Division of Hematology, Rheumatology, and Respiratory Medicine, Department of Internal Medicine, Faculty of Medicine, Kagawa University, Kagawa, Japan 14 Department of Internal Medicine, Miyazaki Prefectural Miyazaki Hospital, Miyazaki, Japan 15 Department of Surgery, Tokyo Medical University, Tokyo, Japan 16 Department of Cancer Information Research, National Kyushu Cancer Center, Fukuoka, Japan 17 Research Institute for Diseases of the Chest, Graduate School of Medical Sciences, Kyushu University, Fukuoka, Japan * These authors have contributed equally to this work Correspondence to: Kazuto Nishio, email: knishio@med.kindai.ac.jp Keywords: non–small cell lung cancer, epidermal growth factor receptor (EGFR), mutation, T790M, cell-free DNA Received: April 14, 2016 Accepted: July 27, 2016 Published: August 16, 2016 ABSTRACT Introduction: Next-generation epidermal growth factor receptor (EGFR) tyrosine kinase inhibitors (TKIs) have been developed to overcome resistance to earlier generations of such drugs mediated by a secondary T790M mutation of EGFR , but the performance of a second tumor biopsy to assess T790M mutation status can be problematic. Methods: We developed and evaluated liquid biopsy assays for detection of TKI-sensitizing and T790M mutations of EGFR by droplet digital PCR (ddPCR) in EGFR mutation–positive non–small cell lung cancer (NSCLC) patients with acquired EGFR-TKI resistance. Results: A total of 260 patients was enrolled between November 2014 and March 2015 at 29 centers for this West Japan Oncology Group (WJOG 8014LTR) study. Plasma specimens from all subjects as well as tumor tissue or malignant pleural effusion or ascites fluid from 41 patients were collected after the development of EGFR-TKI resistance. All plasma samples were genotyped successfully and the results were reported to physicians within 14 days. TKI-sensitizing and T790M mutations were detected in plasma of 120 (46.2%) and 75 (28.8%) patients, respectively. T790M was detected in 56.7% of patients with plasma positive for TKI-sensitizing mutations. For the 41 patients with paired samples obtained after acquisition of EGFR-TKI resistance, the concordance for mutation detection by ddPCR in plasma compared with tumor tissue or malignant fluid specimens was 78.0% for TKI-sensitizing mutations and 65.9% for T790M. Conclusions: Noninvasive genotyping by ddPCR with cell-free DNA extracted from plasma is a promising approach to the detection of gene mutations during targeted treatment.
In its broadest definition, systems biology is the application of a `systems' way of thinking about and doing cell biology. By implication, this also invites us to consider a systems approach in the context of medicine and the treatment of disease. In particular, the idea that systems biology can form the basis of a personalised, predictive medicine will require that much closer attention is paid to the analytic properties of the feedback loops which will be set up by a personalised approach to healthcare. To emphasize the role that feedback theory will play in understanding personalised medicine, we use the term feedback medicine to describe the issues outlined.In these notes we consider feedback and control systems concepts applied to two important themes in medical systems biology - personalised medicine and combinatorial intervention. In particular, we formulate a feedback control interpretation for the administration of medicine, and relate them to various forms of medical treatment.
The success of precision medicine requires computational models that can effectively process and interpret diverse physiological signals across heterogeneous patient populations. While foundation models have demonstrated remarkable transfer capabilities across various domains, their effectiveness in handling individual-specific physiological signals - crucial for precision medicine - remains largely unexplored. This work introduces a systematic pipeline for rapidly and efficiently evaluating foundation models' transfer capabilities in medical contexts. Our pipeline employs a three-stage approach. First, it leverages physiological simulation software to generate diverse, clinically relevant scenarios, particularly focusing on data-scarce medical conditions. This simulation-based approach enables both targeted capability assessment and subsequent model fine-tuning. Second, the pipeline projects these simulated signals through the foundation model to obtain embeddings, which are then evaluated using linear methods. This evaluation quantifies the model's ability to capture three critical aspects: physiological feature independence, temporal dynamics preservation, and medical scenario d
The last few years have seen rapid progress in transitioning quantum computing from lab to industry. In healthcare and life sciences, more than 40 proof-of-concept experiments and studies have been conducted; an increasing number of these are even run on real quantum hardware. Major investments have been made with hundreds of millions of dollars already allocated towards quantum applications and hardware in medicine. In addition to pharmaceutical and life sciences uses, clinical and medical applications are now increasingly coming into the picture. This chapter focuses on three key use case areas associated with (precision) medicine, including genomics and clinical research, diagnostics, and treatments and interventions. Examples of organizations and the use cases they have been researching are given; ideas how the development of practical quantum computing applications can be further accelerated are described.
3D data from high-resolution volumetric imaging is a central resource for diagnosis and treatment in modern medicine. While the fast development of AI enhances imaging and analysis, commonly used visualization methods lag far behind. Recent research used extended reality (XR) for perceiving 3D images with visual depth perception and touch but used restrictive haptic devices. While unrestricted touch benefits volumetric data examination, implementing natural haptic interaction with XR is challenging. The research question is whether a multisensory XR application with intuitive haptic interaction adds value and should be pursued. In a study, 24 experts for biomedical images in research and medicine explored 3D medical shapes with 3 applications: a multisensory virtual reality (VR) prototype using haptic gloves, a simple VR prototype using controllers, and a standard PC application. Results of standardized questionnaires showed no significant differences between all application types regarding usability and no significant difference between both VR applications regarding presence. Participants agreed to statements that VR visualizations provide better depth information, using the hand
Medicine, including fields in healthcare and life sciences, has seen a flurry of quantum-related activities and experiments in the last few years (although biology and quantum theory have arguably been entangled ever since Schrödinger's cat). The initial focus was on biochemical and computational biology problems; recently, however, clinical and medical quantum solutions have drawn increasing interest. The rapid emergence of quantum computing in health and medicine necessitates a mapping of the landscape. In this review, clinical and medical proof-of-concept quantum computing applications are outlined and put into perspective. These consist of over 40 experimental and theoretical studies. The use case areas span genomics, clinical research and discovery, diagnostics, and treatments and interventions. Quantum machine learning (QML) in particular has rapidly evolved and shown to be competitive with classical benchmarks in recent medical research. Near-term QML algorithms have been trained with diverse clinical and real-world data sets. This includes studies in generating new molecular entities as drug candidates, diagnosing based on medical image classification, predicting patient pe
This study examines the clinical decision-making processes in Traditional East Asian Medicine (TEAM) by reinterpreting pattern identification (PI) through the lens of dimensionality reduction. Focusing on the Eight Principle Pattern Identification (EPPI) system and utilizing empirical data from the Shang-Han-Lun, we explore the necessity and significance of prioritizing the Exterior-Interior pattern in diagnosis and treatment selection. We test three hypotheses: whether the Ext-Int pattern contains the most information about patient symptoms, represents the most abstract and generalizable symptom information, and facilitates the selection of appropriate herbal prescriptions. Employing quantitative measures such as the abstraction index, cross-conditional generalization performance, and decision tree regression, our results demonstrate that the Exterior-Interior pattern represents the most abstract and generalizable symptom information, contributing to the efficient mapping between symptom and herbal prescription spaces. This research provides an objective framework for understanding the cognitive processes underlying TEAM, bridging traditional medical practices with modern computat
An intriguing unanswered question about the evolution of bilateral animals with internal skeletons is how an internal skeleton evolved in the first place. Computational modeling of the development of bilateral symmetric organisms suggests an answer to this question. Our hypothesis is that an internal skeleton may have evolved from a bilaterally symmetric ancestor with an external skeleton. By growing the organism inside-out an external skeleton becomes an internal skeleton. Our hypothesis is supported by a computational theory of bilateral symmetry that allows us to model and simulate this process. Inside-out development is achieved by an orientation switch. Given the development of two bilateral founder cells that generate a bilateral organism, a mutation that reverses the internal mirror orientation of those bilateral founder cells leads to inside-out development. The new orientation is epigenetically inherited by all progeny. A key insight is that each cell contained in the newly evolved organism with the internal skeleton develops according to the very same downstream developmental control network that directs the development of its exoskeletal ancestor. The networks and their
Time-to-event prediction, e.g. cancer survival analysis or hospital length of stay, is a highly prominent machine learning task in medical and healthcare applications. However, only a few interpretable machine learning methods comply with its challenges. To facilitate a comprehensive explanatory analysis of survival models, we formally introduce time-dependent feature effects and global feature importance explanations. We show how post-hoc interpretation methods allow for finding biases in AI systems predicting length of stay using a novel multi-modal dataset created from 1235 X-ray images with textual radiology reports annotated by human experts. Moreover, we evaluate cancer survival models beyond predictive performance to include the importance of multi-omics feature groups based on a large-scale benchmark comprising 11 datasets from The Cancer Genome Atlas (TCGA). Model developers can use the proposed methods to debug and improve machine learning algorithms, while physicians can discover disease biomarkers and assess their significance. We hope the contributed open data and code resources facilitate future work in the emerging research direction of explainable survival analysis.
Medical image analysis plays a key role in precision medicine as it allows the clinicians to identify anatomical abnormalities and it is routinely used in clinical assessment. Data curation and pre-processing of medical images are critical steps in the quantitative medical image analysis that can have a significant impact on the resulting model performance. In this paper, we introduce a precision-medicine-toolbox that allows researchers to perform data curation, image pre-processing and handcrafted radiomics extraction (via Pyradiomics) and feature exploration tasks with Python. With this open-source solution, we aim to address the data preparation and exploration problem, bridge the gap between the currently existing packages, and improve the reproducibility of quantitative medical imaging research.
The results of Brownian dynamics simulations of a single DNA molecule in shear flow are presented taking into account the effect of internal viscosity. The dissipative mechanism of internal viscosity is proved necessary in the research of DNA dynamics. A stochastic model is derived on the basis of the balance equation for forces acting on the chain. The Euler method is applied to the solution of the model. The extensions of DNA molecules for different Weissenberg numbers are analyzed. Comparison with the experimental results available in the literature is carried out to estimate the contribution of the effect of internal viscosity.
We read with interest the above article by Zavorsky (2025, Respiratory Medicine, doi:10.1016/j.rmed.2024.107836) concerning reference equations for pulmonary function testing. The author compares a Generalized Additive Model for Location, Scale, and Shape (GAMLSS), which is the standard adopted by the Global Lung Function Initiative (GLI), with a segmented linear regression (SLR) model, for pulmonary function variables. The author presents an interesting comparison; however there are some fundamental issues with the approach. We welcome this opportunity for discussion of the issues that it raises. The author's contention is that (1) SLR provides "prediction accuracies on par with GAMLSS"; and (2) the GAMLSS model equations are "complicated and require supplementary spline tables", whereas the SLR is "more straightforward, parsimonious, and accessible to a broader audience". We respectfully disagree with both of these points.