共找到 20 条结果
ChatGPT, an advanced AI language model, presents a transformative opportunity in several fields including the medical education. This article examines the integration of ChatGPT into healthcare learning environments, exploring its potential to revolutionize knowledge acquisition, personalize education, support curriculum development, and enhance clinical reasoning. The AI's ability to swiftly access and synthesize medical information across various specialties offers significant value to students and professionals alike. It provides rapid answers to queries on medical theories, treatment guidelines, and diagnostic methods, potentially accelerating the learning curve. The paper emphasizes the necessity of verifying ChatGPT's outputs against authoritative medical sources. A key advantage highlighted is the AI's capacity to tailor learning experiences by assessing individual needs, accommodating diverse learning styles, and offering personalized feedback. The article also considers ChatGPT's role in shaping curricula and assessment techniques, suggesting that educators may need to adapt their methods to incorporate AI-driven learning tools. Additionally, it explores how ChatGPT could bolster clinical problem-solving through AI-powered simulations, fostering critical thinking and diagnostic acumen among students. While recognizing ChatGPT's transformative potential in medical education, the article stresses the importance of thoughtful implementation, continuous validation, and the establishment of protocols to ensure its responsible and effective application in healthcare education settings.
Distributed Acoustic Sensing (DAS) has shown promise for real-time monitoring of large-scale infrastructure by providing spatio-temporal information about vibrations along a fiber optic cable. However, data easily reaches into terabytes per day due to high spatial resolution and acquisition frequency, making storage, transfer, and analysis economically infeasible especially for applications requiring real-time decisions. Tensor Networks (TNs) are a data structure well poised to address the challenges of DAS as they are effective at capturing signals in low rank and enable linear operations (e.g. signal processing) in the compressed space, providing computational savings and bypassing the need to decompress. This article is the first to demonstrate how TNs can be applied to DAS data and recreates a pre-existing workflow for DAS in TN format for experimental data from a field-scale wellbore. The methods achieved [Formula: see text] real-time compression on a laptop with high accuracy and efficiently processed the data in the compressed space without needing to prematurely decompress. Not only does this research help reduce the cost of implementing DAS technology, but it creates new opportunities between the fields of signal processing and TNs.
Large genetic datasets are terabytes in size, presenting a computational challenge that will intensify as sequencing efforts scale. We present a lossless compression algorithm, kodama, which supports matrix multiplication and is suitable for large-scale statistical analyses. Kodama leverages genealogical relatedness among nominally unrelated individuals and infers a novel data structure similar to the ancestral recombination graph (ARG), called the linear ARG. We applied kodama to whole genome sequencing data from UK Biobank and All of Us. Inferred linear ARGs were 17-89 times smaller on disk compared to the input data; the entire UK Biobank N=200k dataset can be loaded into memory (58GB). Compared with the recently proposed genotype representation graph (GRG), the linear ARG is 2.5 times smaller. Genotype matrix multiplications, which are the bottleneck in most statistical applications, are extremely fast with the linear ARG; we performed a GWAS on the UK Biobank 200k cohort across 89 traits with 42 covariates in 100 seconds, representing a 4,700-fold speedup over PLINK 2.0. We expect that the linear ARG will enable genetic analyses to scale to millions of samples.
Indigenous cattle are central to livestock production in Africa, valued for their adaptability to harsh tropical environments despite lower productivity than commercial breeds. Genome analyses offer critical insights into the genetic potential for enhancing both resilience and productive traits, supporting the advancement of worldwide cattle farming systems. Here, we generated whole-genome sequence data for 240 indigenous cattle representing breeds from distinct agro-climatic regions in Egypt, Uganda, and South Africa. The dataset comprises over ten terabytes of paired-end reads generated using the Illumina NovaSeq. 6000 platform, with an average genome coverage of approximately 10×. Post-filtering reads were mapped to the ARS-UCD1.2 reference genome with a mean mapping rate of 99.2% (range: 64.5-99.9%). Variant calling identified ~43 million SNPs and 6 million indels (≤50 bp) unevenly distributed across the genome. Functional annotation indicated that many variants were located within or near known genes. This comprehensive genomic resource provides a foundation for future studies of genetic diversity, breed identity, population structure, local adaptation, breed-specific traits, or strategies for global cattle conservation.
Epilepsy is a heterogeneous syndrome. Personalised localisation of epileptogenic zone (EZ) is critical for diagnosis and treatment of drug-resistant focal epilepsy. Multichannel stereoelectroencephalography (SEEG) monitoring acquired over a period of two to three weeks was collected in different patients, resulting in comprehensive epileptogenic information and terabytes of high dimensional data. Consequently, there is a need for high-throughput data analytical methods to enable data-driven, personalised seizure detection and EZ localisation. Here, a seizure detection and EZ localisation AI system - SEEGformer is proposed, by utilising SEEG data from 61 patients acquired across two centres and three cohorts capturing tens of thousands of abnormal discharges and around ten seizures per person on average. SEEGformer employs a parallel transformer architecture to analyse multiple representations of multichannel SEEG signals, including the real part, imaginary part, and amplitude of the analytic signal after Fourier transform. The MRI information was encoded in SEEGformer to construct the structural dependence of the brain areas. Inter-channel dependencies and interactions were captured for seizure detection. A cross-channel attention mechanism computed the epileptogenic risk score for each channel to localise EZ using ictal SEEG data. Each patient's SEEG data was used to train and validate their individual-specific SEEGformer model. In three clinical cohorts, SEEGformer achieved an average AUROC of 0.937 (95% CI, 0.922-0.950) for seizure detection and 0.798 (95% CI, 0.749-0.847) for EZ localisation. Localisation performance surpassed state-of-the-art methods by over 5%. SEEGformer further revealed distinct phase synchronisation patterns in dynamically evolving epileptogenic zone networks, with a significance level of P < 0.0001. Due to its high interpretability and visualisation capabilities, SEEGformer can enhance clinical decision-making by providing an objective, data-driven reference to optimise epileptogenic zone delineation and surgical strategy development. Currently, the improved SEEGformer is being developed to construct a dedicated SEEG atlas for epilepsy. This study was funded by Natural Science Foundation of Shanghai (25ZR1401179), the National Key Research and Development Program of China (2022YFB4702702), the Sci-Tech Innovation 2030-Major Project of Brain Science and Brain-inspired Intelligence Technology (2021ZD0200600), Beijing Municipal Public Welfare Development and Reform Pilot Project for Medical Research Institutes (JYY2023-8), and National Key Research and Development Program of China (2024YFC3044700).
Monte Carlo (MC) simulations constitute the most accurate tool for dosimetric analysis in small radiation fields, such as those generated by the Leksell Gamma Knife Perfexion (LGK-PFX) system. However, implementing a fully detailed model of the system is computationally demanding, both in terms of processing time and geometric complexity, due to the explicit inclusion of the 576 collimators distributed throughout the device. To address this limitation, we present an efficient and accurate model based on a single phase-space file (PSF) generated from a source-collimator simulation for each collimator size (4, 8, and 16 mm), computed using PenEasy Monte. The resulting PSF is subsequently rotated to the required orientations to reproduce the system response. This approach drastically reduces the effective computational burden, as each reference PSF is generated only once and can then be reused to simulate any desired LGK-PFX configuration. Compared with full-system models that require hundreds of hours of computation and storage on the order of terabytes, the proposed rotating-PSF strategy enables Monte Carlo simulations with moderate computational resources, making MC-based verification feasible in a clinical context. The model was validated by comparing MC-generated dose distributions with measurements performed using EBT4 radiochromic films at Ruber International Hospital (Madrid), as well as with data provided by Elekta and other MC-based studies. A 7%-0.5 mm gamma analysis of the dose profiles showed pass rates above 95%. Output factors (OFs) for the 4 mm and 8 mm collimators were 0.848 ± 0.011 and 0.889 ± 0.011, respectively, in a cubic water phantom. Overall, the rotating PSF approach preserves dosimetric accuracy while substantially improving computational efficiency, providing a practical and clinically relevant solution for dosimetry with the LGK-PFX system.
Assessing likely variant effects on phenotypes is of critical importance in diagnostic settings, and while much progress has been made in interpreting genic mutations based on our understanding of coding sequence, noncoding variants can be much more challenging to reliably interpret based on DNA sequence alone. High-throughput reporter assays such as STARR-seq and MPRA have shown utility in experimentally measuring regulatory effects of noncoding variants present in samples but provide no readout for variants not present in the assay inputs. However, whole-genome reporter assays provide copious data that can be used to train predictive models for prioritizing variants not directly observed in the experiment. We describe a retrainable predictive modeling framework, BlueSTARR, for this task, and present results of training several models with this framework on whole-genome STARR-seq data from two cell lines and one drug treatment. Using these models, we uncover a global signature across the human genome consistent with purifying selection against both loss-of-function and gain-of-function regulatory variants, with the latter showing a significant bias consistent with selection against gains of cis regulatory function in closed chromatin proximal to genes. By testing the model on synthetic enhancers with binding motifs for transcription factors GR and AP-1, we find that when trained on drug perturbation data, the model is able to learn distance-dependent and treatment-dependent binding patterns and their resulting reporter gene activation. These results demonstrate that lightweight, easily retrainable models such as ours have utility in probing latent signals present in novel experimental data. Finally, we find only modest differences in performance between different deep-learning architectures when trained on this single data modality, and while somewhat greater predictive accuracy can be achieved with much larger models trained at great expense on many terabytes of data, there is still copious room for improvement even for industrial strength, state-of-the-art models.
Artificial intelligence (AI) and machine learning (ML) algorithms possess the capability to accelerate the design of novel materials; however, their advancement in materials science is severely hindered by a fundamental deficit of experimental data, commonly referred to as data starvation. Unlike solution-based chemistry, where high-throughput (HT) technologies are a well-established standard, the automated synthesis of solid materials-particularly polymers and multicomponent composites-poses an extreme engineering challenge. Furthermore, the traditional, manual research model is inherently flawed by human bias, notably the systematic non-publication of negative results, which deprives AI models of critical boundary information regarding the design space. This paper is the first in a three-part review series defining the architecture of a fully automated, unbiased "data factory" for closed-loop discovery. This section focuses on the physical foundations of the HT workflow: experimental planning, automated synthesis, and material management. Emphasis is placed on the paradigm shift from classical, discrete Design of Experiments (DoE) to the novel concept of Continuous Gradient DoE. It reviews how robotic platforms utilizing precise gravimetric and volumetric feeders, integrated with extruders and in-line capillary rheology, enable the seamless, high-throughput manufacturing of thermoplastics and composites. Moreover, an innovative approach to sample logistics is presented, redefining classical storage patterns through the implementation of Continuous Material Management. This encompasses direct physical tagging (e.g., inkjet marking on continuous filaments or films), spool-based transport systems, and precise, real-time metadata mapping. As demonstrated, the integration of these systems yields an order-of-magnitude increase in productivity (generating tens of thousands of novel material variants annually), a radical reduction in unit costs, and the production of terabytes of standardized, machine-readable data. Establishing this reliable hardware and analytical infrastructure represents the essential first step toward unlocking the full potential of artificial intelligence in advanced materials engineering.
Ultra-large-scale structure-based virtual screening (SBVS) for identifying novel bioactive compounds poses significant computational challenges. These challenges arise from the size of available chemical libraries, which can contain billions of molecules that require exhaustive docking and scoring, placing prohibitive demands on CPU/GPU resources. Small- and mid-sized laboratories often lack access to the high-performance computing clusters or cloud resources necessary to process such workloads in a timely manner. Furthermore, managing and analyzing the resulting terabytes of docking data requires robust data-handling pipelines and expertise that are not universally accessible. Here, we present a data-driven drug development pipeline that leverages a subset of molecules from a database with a common scaffold, reducing the chemical search space by tens to hundreds of orders of magnitude. In this case, the common scaffold that is the key to allowing this reduction is the 2-phenylthiazole moiety, identified through NMR fragment screening. We started with a subset of over 400 000 drug-sized 2-phenylthiazole-containing molecules selected from the zinc database and trained a random forest regression model on about 1% of this data to predict binding scores for the entire library. For this purpose, we used a distribution-preserving sampling approach based on KMeans clustering and binning, and we evaluated its statistical fidelity using KS, Wasserstein, JS, and KL divergence metrics. Our approach preserved the distribution of docking scores, demonstrating the utility of data-driven strategies for scalable virtual screening and establishing a benchmark dataset for machine learning in drug discovery.
The new technology of Artificial Intelligence (AI) has become an important aspect of medical imaging. It allows replacing the manual interpretation with the automated analysis based on the data. Clinical data quality in the healthcare industry is growing at a rate of more than 19 terabytes annually bringing a compelling necessity to derive significant information. This is a process of examining trends in the information and handling the complexity of the huge amount of medical data with high precision based on time with time-dependent procedures. Deep-learning systems, more specifically AI systems, have proven to be effective and efficient in automating the complex medical imaging processes. These processes involve image acquisition, image enhancement, image segmentation, quantitative feature extraction, image diagnosis, image prognostics, clinical workflow optimization and image guided surgical planning. The chapter speaks of the relevance of AI towards enhancing medical analysis of images. It explains the ways in which AI is revolutionizing diagnostic radiology to be more patient centered. Along with computational improvements, AI-based imaging analytics will be able to achieve excellence in precision medicine. These systems offer scalable and high-quality frameworks which enable real-time interpretation of the biological data. This eventually helps in the accuracy of diagnosis and increases patient survival.
Advances in X-ray and neutron sources, as well as in area-detector technologies, enable the recording of several terabytes of raw two-dimensional detector data in a single experiment. While several efficient integration and conversion tools are available for data collected in transmission geometry, analogous solutions for grazing-incidence diffraction (including grazing-incidence X-ray diffraction and grazing-incidence wide-angle X-ray scattering) experiments have not yet achieved the same level of efficiency. The development of new data analysis tools, including machine-learning-based software for X-ray data, necessitates the establishment of a standardized format for the converted data. To address these challenges, we have developed a new Python library, pygid, which is designed to facilitate fast data processing while providing compatibility with various raw data formats, a standardized data storage format and an intuitive interface for straightforward use. pygid supports three types of coordinate systems and both transmission and grazing-incidence geometries. It is capable of handling large datasets, performing one-dimensional line cuts and simulating expected Bragg peak positions for given structures. The package facilitates sample and experimental metadata curation in accordance with the FAIR principles. As an integral part of the broader mlgid pipeline, pygid serves as the initial step linking raw scattering patterns with machine learning tools for data analysis. The pygid package is accessible at https://github.com/mlgid-project.
Livestock systems are increasingly instrumented with heterogeneous sensors, yet the resulting data remain fragmented, short-lived, and rarely documented as integrated infrastructures. This gap limits the development of robust multimodal artificial intelligence under real production conditions. Here we present a longitudinal multimodal data infrastructure for poultry monitoring, spanning 22 consecutive weeks across five commercial-style barns. The dataset combines continuous RGB video (1080 p, 30 fps), continuous audio (48 kHz), periodic radiometric thermal imaging, and twice-daily environmental measurements, yielding 10.2 terabytes of temporally heterogeneous data. Rather than focusing on a specific predictive task, the study addresses the underlying data-engineering challenge: how to acquire, synchronize, store, and preprocess multimodal streams at production scale. We detail a reproducible system architecture for distributed sensing, local buffering, secure transfer, and cloud-based organization, together with standardized preprocessing pipelines for illumination correction, acoustic denoising, and radiometric temperature extraction. Temporal alignment is achieved through timestamp-based normalization across asynchronous modalities, with explicit characterization of alignment granularity and missing data under real-world constraints. This work positions multimodal livestock sensing as a data-systems problem. The resulting dataset supports longitudinal analysis, cross-modal querying, and the development and evaluation of machine learning and multimodal fusion approaches at appropriate temporal scales. By releasing both data and workflows, we provide a transparent and extensible foundation for building and evaluating AI systems in precision agriculture.
Recent electron microscopy workflows, particularly 4D Scanning Transmission Electron Microscopy (4D STEM), generate massive data volumes, sometimes exceeding tens of terabytes per dataset. This growth is outpacing improvements in storage, network, and I/O bandwidths. Although lossless compression offers a viable solution, it is often limited by modest compression ratios. Lossy compression offers an alternative; however, compression errors can propagate to downstream analyses. The goal is therefore to deploy a compression method that preserves essential data features while defining a suitable metric. In this paper, we present a moment-preserving compression workflow combining multigrid adaptive reduction (MGARD) with a constraint satisfaction technique. MGARD provides mathematically guaranteed error bounds on raw data, while a constraint satisfaction method rectifies decompressed data to strictly preserve selected statistical moments in diffraction space. This allows moment-derived observables, such as Center-of-Mass (COM) measurements, to be recovered consistently from the corrected tensor. Furthermore, we propose a frequency-domain reliability criterion in raw diffraction space to evaluate the acceptable accuracy limits. Experiments on 4D STEM datasets demonstrate that this moment-preserving adjustment substantially reduces errors in targeted downstream quantities, providing a practical path toward compressed 4D STEM data compatible with downstream microscopy analysis.
High-content imaging (HCI) involves the automated acquisition and quantitative analysis of cell phenotypes from microscopy images. These studies often rely on screening, which can involve thousands of chemical or genetic perturbations that produce terabytes of microscopy data. To extract meaningful biological insights, these data must be processed into quantitative features through a technique known as image-based profiling. A major analytical bottleneck is curating the high-dimensional, single-cell data derived from various image-analysis tools. These datasets suffer from inconsistent schemas, inefficient file formats, and undocumented ontological relationships. These challenges reduce reproducibility and slow progress in downstream applications. To solve these issues, we introduce CytoTable, a software package for harmonizing single-cell image-based profiling. CytoTable enables modular, portable, and cross-language data integration through a robust, reproducible, and scalable engine that harmonizes single-cell readouts from multiple image-analysis tools, preparing for feature integration with software in the Cytomining ecosystem such as Pycytominer.
Reconstructing the complete cell lineages and fate maps of a living embryo has remained a central challenge in developmental biology. Here we introduce ITEC (Iterative Tracking with Error Correction), a fully unsupervised method that automatically reconstructs the lineage of every cell in the embryo with high fidelity. ITEC was validated with manually annotated lineages on four cross-species datasets including zebrafish, mouse, and Drosophila. We reconstructed the developmental lineages of a zebrafish embryo from terabyte-scale data of total 18.5 million cells with an estimated accuracy of over 99.7%. With ITEC, we revealed spatiotemporal dynamics of key morphogenetic processes such as somite boundary formation, demonstrated the various patterns of spatial sorting among adjacent organs or anatomical regions, and identified associations between cellular movement and spatial transcriptomics. ITEC provides a powerful platform for retrospective fate mapping and systematic exploration of developmental dynamics at the cellular scale and single-embryo level.
Domestic yaks are a crucial livestock species on the Qinghai-Tibet Plateau and its surrounding regions, providing indispensable resources for local residents' livelihoods and production. China harbors the world's largest yak population and possesses exceptionally rich yak genetic resources. In this study, we performed whole-genome resequencing on 200 individuals representing 19 indigenous yak breeds. Sequencing of DNA samples generated approximately 4.3 terabytes (TB) of raw data, with an average sequencing depth of 8.36×. High-quality sequencing reads were successfully mapped to the reference yak genome, achieving an average alignment rate of 98.4%. Following stringent quality control and filtering, we identified a total of 13.41 million high-confidence single nucleotide polymorphisms (SNPs). These genome-wide SNP markers provide a valuable resource for in-depth investigation into the demographic history and adaptive evolutionary mechanisms of Chinese yak populations, offering critical genomic support for the conservation and innovative utilization of yak genetic resources.
Polarization-sensitive optical coherence tomography (PS-OCT) is a label-free imaging technique that exploits birefringence to visualize myelinated axons at micrometer resolution. However, serial PS-OCT imaging has been limited to small volumes, including tissue blocks from larger species, owing to constraints in acquisition speed, system stability, and data processing. These limitations have prevented its application to whole-brain mapping in large mammals. Here we present a scalable PS-OCT acquisition system and computational pipeline for whole-brain imaging in the rhesus macaque. The framework integrates high-throughput serial imaging with automated reconstruction and processing, enabling volumetric imaging at micrometer-scale resolution across decimeter-scale brain volumes. Using this approach, we acquired two complete macaque brains at a voxel size of 5.5 × 5.5 × 3.4 μm and an effective resolution of approximately 10 × 10 × 5.5 μm, generating multi-terabyte datasets consisting of multiple contrasts including fiber orientation information. The datasets and associated processing tools are made publicly available. This platform establishes a method for large-scale, high-resolution mapping of white matter architecture in primate brains. The resulting datasets provide a reference for validating MRI models and support the development of neurotechnological applications, including deep brain stimulation, where accurate characterization of axonal organization is required.
Tomographic datasets have continued to grow in size and volume along with camera technology and synchrotron brilliance. These larger datasets require new methods of distributing the reconstruction process on high-performance computing systems. Here we present a program for reconstructing large datasets, allowing the usage of either CPU- or GPU-powered clusters, and show the successful reconstruction of a 24000 pixel-wide dataset. This program is currently in use at the DanMAX beamline at the MAX IV synchrotron and is used as the default reconstruction method for all tomographic experiments at the beamline. Scans reconstructed with the presented pipeline have been published in multiple journals.
This article introduces a dataset designed for the detection of partial discharges in transmission power lines using covered conductors through a contact galvanic method, sourced from real environments across 23 different power lines in various locations. Though partially introduced in a Kaggle competition (only 3% of data), its full extent is disclosed here for the first time. The dataset is distinguished by its rich, imbalanced distribution across seven classes, derived from signals processed via a sophisticated voltage-based method, and supplemented with extracted features to aid analysis. Its scale, detailed labeling, and real-world basis offer unparalleled opportunities for developing machine learning algorithms aimed at fault detection. This contribution holds vast potential for reuse in electrical engineering research focused on enhancing power distribution network reliability and safety, particularly in the context of predictive maintenance and understanding partial discharge behaviors.