A major public health issue is the growing resistance of bacteria to antibiotics. An important part of the needed response is the discovery and development of new antimicrobial strategies. These require the screening of potential new drugs, typically accomplished using high-throughput screening (HTS). Traditionally, HTS is performed by examining one compound per well, but a more efficient strategy pools multiple compounds per well. In this work, we study several recently proposed pooling construction methods, as well as a variety of pooled high-throughput screening analysis methods, in order to provide guidance to practitioners on which methods to use. This is done in the context of an application of the methods to the search for new drugs to combat bacterial infection. We discuss both an extensive pilot study as well as a small screening campaign, and highlight both the successes and challenges of the pooling approach.
Proton-conducting solid acids could enable water-free operation of high-temperature fuel cells. However, systematic materials screening has, hitherto, been computationally prohibitive. Here, we introduce a two-stage high-throughput screening strategy that directly computes proton diffusion coefficients, enabled by machine-learned interatomic potentials fine-tuned to ab initio data. Starting from more than six million materials, our screening -- based on structural motifs rather than empirical descriptors -- identifies $27$ high-performing proton conductors, including over ten previously unexplored compounds. These include sustainable and commercially available materials, candidates that have not yet been synthesized, organic systems that fall outside conventional design rules, and known proton conductors that validate our approach. Importantly, our findings reveal a universal oxygen--oxygen distance of approximately $2.5$~Å at the moment of proton transfer across diverse chemistries, providing mechanistic insight and showing that macroscopic proton conductivity emerges from the interplay between anion rotational dynamics, hydrogen-bond network connectivity, and proton-transfer prob
The mechanical properties of biological fluids can serve as early indicators of disease, offering valuable insights into complex physiological and pathological processes. However, the existing technologies can hardly support high throughput measurement, which hinders their broad applications in disease diagnosis. Here, we propose the ultrasound-coupled microdroplet laser chips to enable high-throughput measurement of the intrinsic mechanical properties of fluids. The microdroplets supporting high-Q (10^4) whispering gallery modes (WGM) lasing were massively fabricated on a hydrophobic surface with inject printing. The ultrasound was used to actuate the mechanical vibration of the microdroplets. We found that the stimulus-response of the laser emission is strongly dependent on the intrinsic mechanical properties of the liquid, which as subsequently employed to quantify the viscosity. The ultrasound-coupled microdroplet laser chips were used to monitor molecular interactions of bovine serum albumin. High-throughput screening of hyperlipidemia disease was also demonstrated by performing over 2,000 measurements using fast laser scanning. Thanks to the small volume of the microdroplets,
The development of new high dielectric materials is essential for advancement in modern electronics. Oxides are generally regarded as the most promising class of high dielectric materials for industrial applications as they possess both high dielectric constants and large band gaps. Most previous researches on high dielectrics were limited to already known materials. In this study, we conducted an extensive search for high dielectrics over a set of ternary oxides by combining crystal structure prediction and density functional perturbation theory calculations. From this search, we adopted multiple stage screening to identify 440 new low-energy high dielectric materials. Among these materials, 33 were identified as potential high dielectrics favorable for modern device applications. Our research has opened an avenue to explore novel high dielectric materials by combining crystal structure prediction and high throughput screening.
Gene expression is a complex phenomenon involving numerous interlinked variables, and studying these variables to control expression is essential in bioengineering and biomanufacturing. While cloning techniques for achieving plasmid libraries that cover large design spaces exist, multiplex techniques offering cell culture screening at similar scales are still lacking. We introduced a microcapillary array-based platform aimed at high-throughput, multiplex screening of miniature cell cultures through fluorescent reporters.
A high-throughput screening using density functional calculations is performed to search for stable boride superconductors from the existing materials database. The workflow employs the fast frozen phonon method as the descriptor to evaluate the superconducting properties quickly. 23 stable candidates are identified from the screening. For almost all found binary compounds, the superconductivity was obtained earlier experimentally or computationally. For ternary borides, previous studies are very limited. Our extensive search among ternary systems confirmed superconductivity in known systems and found several new compounds. Among these discovered superconducting ternary borides, Ta(MoB)$_2$ shows the highest superconducting temperature of ~12K. Most predicted compounds were synthesized previously; therefore, our predictions can be examined experimentally. Our work also demonstrates that the boride systems can have diverse structural motifs that lead to superconductivity.
Topological Weyl semimetals represent a novel class of non-trivial materials, where band crossings with linear dispersions take place at generic momenta across reciprocal space. These crossings give rise to low-energy properties akin to those of Weyl fermions, and are responsible for several exotic phenomena. Up to this day, only a handful of Weyl semimetals have been discovered, and the search for new ones remains a very active area. The main challenge on the computational side arises from the fact that many of the tools used to identify the topological class of a material do not provide a complete picture in the case of Weyl semimetals. In this work, we propose an alternative and inexpensive, criterion to screen for possible Weyl fermions, based on the analysis of the band structure along high-symmetry directions in the absence of spin-orbit coupling. We test the method by running a high-throughput screening on a set of 5455 inorganic bulk materials and identify 49 possible candidates for topological properties. A further analysis, carried out by identifying and characterizing the crossings in the Brillouin zone, shows us that 3 of these candidates are Weyl semimetals. Interestin
Due to their chemical and structural diversity, nanoporous materials can be used in a wide variety of applications, including fluid separation, gas storage, heterogeneous catalysis, drug delivery, etc. Given the large and rapidly increasing number of known nanoporous materials, and the even bigger number of hypothetical structures, computational screening is an efficient method to find the current best-performing materials and to guide the design of future materials. This review highlights the potential of high-throughput computational screenings in various applications. The achievements and the challenges associated to the screening of several material properties are discussed to give a broader perspective on the future of the field.
As connected devices multiply and the internet matures into a ubiquitous platform for exchange and communication, the question of what makes a domain name valuable is ever more significant. Due to the scarcity of meaningful vocabulary and the persistence of domain-related data, the buying and selling of previously owned domain names, also known as the domain aftermarket, has evolved into a billion dollar industry. Each day over a 100,000 domain names expire and become available for re-registration. Manual appraisal is impossible at such a volume; thus a method for the automated identification of valuable domain names is called for. The aim of our study was to develop a method for high throughput screening of domain names for rapid identification of the valuable ones. Five different aspects that make a domain name valuable were identified: name quality, domain authority, domain traffic, active domain age and domain health. An SVM method was developed for high throughput screening of domain names. Our method was able to identify valuable domain names with 97% accuracy for the test set and 93% for the external set and can be used for routinely screening the domain aftermarket.
High throughput screening of compounds (chemicals) is an essential part of drug discovery [7], involving thousands to millions of compounds, with the purpose of identifying candidate hits. Most statistical tools, including the industry standard B-score method, work on individual compound plates and do not exploit cross-plate correlation or statistical strength among plates. We present a new statistical framework for high throughput screening of compounds based on Bayesian nonparametric modeling. The proposed approach is able to identify candidate hits from multiple plates simultaneously, sharing statistical strength among plates and providing more robust estimates of compound activity. It can flexibly accommodate arbitrary distributions of compound activities and is applicable to any plate geometry. The algorithm provides a principled statistical approach for hit identification and false discovery rate control. Experiments demonstrate significant improvements in hit identification sensitivity and specificity over the B-score method, which is highly sensitive to threshold choice. The framework is implemented as an efficient R extension package BHTSpack and is suitable for large scal
It is well known that the high electric conductivity, large Seebeck coefficient, and low thermal conductivity are preferred for enhancing thermoelectric performance, but unfortunately, these properties are strongly inter-correlated with no rational scenario for their efficient decoupling. This big dilemma for thermoelectric research appeals for alternative strategic solutions, while the high-throughput screening is one of them. In this work, we start from total 3136 real electronic structures of the huge X2YZM4 quaternary compound family and perform the high-throughput searching in terms of enhanced thermoelectric properties. The comprehensive data-mining allows an evaluation of the electronic and phonon characteristics of those promising thermoelectric materials. More importantly, a new insight that the enhanced thermoelectric performance benefits substantially from the coexisting quasi-Dirac and heavy fermions plus strong optical-acoustic phonon hybridization, is proposed. This work provides a clear guidance to theoretical screening and experimental realization and thus towards development of performance-excellent thermoelectric materials.
We introduce ChemPro, a progressive benchmark with 4100 natural language question-answer pairs in Chemistry, across 4 coherent sections of difficulty designed to assess the proficiency of Large Language Models (LLMs) in a broad spectrum of general chemistry topics. We include Multiple Choice Questions and Numerical Questions spread across fine-grained information recall, long-horizon reasoning, multi-concept questions, problem-solving with nuanced articulation, and straightforward questions in a balanced ratio, effectively covering Bio-Chemistry, Inorganic-Chemistry, Organic-Chemistry and Physical-Chemistry. ChemPro is carefully designed analogous to a student's academic evaluation for basic to high-school chemistry. A gradual increase in the question difficulty rigorously tests the ability of LLMs to progress from solving basic problems to solving more sophisticated challenges. We evaluate 45+7 state-of-the-art LLMs, spanning both open-source and proprietary variants, and our analysis reveals that while LLMs perform well on basic chemistry questions, their accuracy declines with different types and levels of complexity. These findings highlight the critical limitations of LLMs in
We screen a large chemical space of perovskite alloys for systems with the right properties to accommodate a morphotropic phase boundary (MPB) in their composition-temperature phase diagram, a crucial feature for high piezoelectric performance. We start from alloy end-points previously identified in a high-throughput computational search. An interpolation scheme is used to estimate the relative energies between different perovskite distortions for alloy compositions with a minimum of computational effort. Suggested alloys are further screened for thermodynamic stability. The screening identifies alloy systems already known to host a MPB, and suggests a few new ones that may be promising candidates for future experiments. Our method of investigation may be extended to other perovskite systems, e.g., (oxy-)nitrides, and provides a useful methodology for any application of high-throughput screening of isovalent alloy systems.
When cellular contractile forces are central to pathophysiology, these forces comprise a logical target of therapy. Nevertheless, existing high-throughput screens are limited to upstream signaling intermediates with poorly defined relationship to such a physiological endpoint. Using cellular force as the target, here we screened libraries to identify novel drug candidates in the case of human airway smooth muscle cells in the context of asthma, and also in the case of Schlemm's canal endothelial cells in the context of glaucoma. This approach identified several drug candidates for both asthma and glaucoma. We attained rates of 1000 compounds per screening day, thus establishing a force-based cellular platform for high-throughput drug discovery.
This study assesses the efficiency of several popular machine learning approaches in the prediction of molecular binding affinity: CatBoost, Graph Attention Neural Network, and Bidirectional Encoder Representations from Transformers. The models were trained to predict binding affinities in terms of inhibition constants $K_i$ for pairs of proteins and small organic molecules. First two approaches use thoroughly selected physico-chemical features, while the third one is based on textual molecular representations - it is one of the first attempts to apply Transformer-based predictors for the binding affinity. We also discuss the visualization of attention layers within the Transformer approach in order to highlight the molecular sites responsible for interactions. All approaches are free from atomic spatial coordinates thus avoiding bias from known structures and being able to generalize for compounds with unknown conformations. The achieved accuracy for all suggested approaches prove their potential in high throughput screening.
Although topological invariants have been introduced to classify the appearance of protected electronic states at surfaces of insulators, there are no corresponding indexes for Weyl semimetals whose nodal points may appear randomly in the bulk Brillouin Zone (BZ). Here we use a well-known result that every Weyl point acts as a Dirac monopole and generates integer Berry flux to search for the monopoles on rectangular BZ grids that are commonly employed in self-consistent electronic structure calculations. The method resembles data mining technology of computer science and is demonstrated on locating the Weyl points in known Weyl semimetals. It is subsequently used in high throughput screening several hundreds of compounds and predicting a dozen new materials hosting nodal Weyl points and/or lines.
To enhance large language models (LLMs) for chemistry problem solving, several LLM-based agents augmented with tools have been proposed, such as ChemCrow and Coscientist. However, their evaluations are narrow in scope, leaving a large gap in understanding the benefits of tools across diverse chemistry tasks. To bridge this gap, we develop ChemToolAgent, an enhanced chemistry agent over ChemCrow, and conduct a comprehensive evaluation of its performance on both specialized chemistry tasks and general chemistry questions. Surprisingly, ChemToolAgent does not consistently outperform its base LLMs without tools. Our error analysis with a chemistry expert suggests that: For specialized chemistry tasks, such as synthesis prediction, we should augment agents with specialized tools; however, for general chemistry questions like those in exams, agents' ability to reason correctly with chemistry knowledge matters more, and tool augmentation does not always help.
We discuss a strategy to study non-perturbatively QCD up to very high temperatures by Monte Carlo simulations on the lattice. It allows not only the thermodynamic properties of the theory but also other interesting thermal features to be investigated. As a first concrete application, we compute the flavour non-singlet mesonic screening masses and we present the results of Monte Carlo simulations at 12 temperatures covering the range from T $\sim$ 1 GeV up to $\sim$ 160 GeV in the theory with three massless quarks. On the one side, chiral symmetry restoration manifests itself in our results through the degeneracy of the vector and the axial vector channels and of the scalar and the pseudoscalar ones, and, on the other side, we observe a clear splitting between the vector and the pseudoscalar screening masses up to the highest investigated temperature. A comparison with the high-temperature effective theory shows that the known one-loop order in the perturbative expansion does not provide a satisfactory description of the non-perturbative data up to the highest temperature considered.
This paper treats the problem of screening for variables with high correlations in high dimensional data in which there can be many fewer samples than variables. We focus on threshold-based correlation screening methods for three related applications: screening for variables with large correlations within a single treatment (autocorrelation screening); screening for variables with large cross-correlations over two treatments (cross-correlation screening); screening for variables that have persistently large auto-correlations over two treatments (persistent-correlation screening). The novelty of correlation screening is that it identifies a smaller number of variables which are highly correlated with others, as compared to identifying a number of correlation parameters. Correlation screening suffers from a phase transition phenomenon: as the correlation threshold decreases the number of discoveries increases abruptly. We obtain asymptotic expressions for the mean number of discoveries and the phase transition thresholds as a function of the number of samples, the number of variables, and the joint sample distribution. We also show that under a weak dependency condition the number of
Multimodal scientific reasoning remains a significant challenge for large language models (LLMs), particularly in chemistry, where problem-solving relies on symbolic diagrams, molecular structures, and structured visual data. Here, we systematically evaluate 40 proprietary and open-source multimodal LLMs, including GPT-5, o3, Gemini-2.5-Pro, and Qwen2.5-VL, on a curated benchmark of Olympiad-style chemistry questions drawn from over two decades of U.S. National Chemistry Olympiad (USNCO) exams. These questions require integrated visual and textual reasoning across diverse modalities. We find that many models struggle with modality fusion, where in some cases, removing the image even improves accuracy, indicating misalignment in vision-language integration. Chain-of-Thought prompting consistently enhances both accuracy and visual grounding, as demonstrated through ablation studies and occlusion-based interpretability. Our results reveal critical limitations in the scientific reasoning abilities of current MLLMs, providing actionable strategies for developing more robust and interpretable multimodal systems in chemistry. This work provides a timely benchmark for measuring progress in