Protein-ligand interaction prediction is pivotal in early-stage drug development, enabling large-scale virtual screening, drug optimization, and reverse target searching. In this work, we present Graph_RG, our top-performing model in the CASP16 small molecule track's protein-ligand affinity prediction category, achieving a N-weighted Kendall's Tau of 0.42-significantly outperforming other submissions (second-best: 0.36). Beyond accuracy, Graph_RG is noncomplex dependent, hence exhibits exceptional computational efficiency, operating > 100 000× faster than conformation-search dependent prediction methods, thus enabling billion- to 10-billion-scale screening on standard servers. We further discuss the potential improvements for Graph_RG, including dataset optimization, atomic vector representation enhancements, and model architecture upgrades. We also introduce the potential broader applications in large-scale drug screening, reverse target identification, and GPCR-specific drug discovery. We also point out the development of an interactive web platform hosting Graph_RG and its derivative models to enhance accessibility. By integrating community feedback and iterative model refinement, this initiative bridges the gap between AI-driven predictions and practical drug discovery, fostering advancements in both computational methodologies and biomedical applications.
Generation of upscaled quantities of human-induced pluripotent stem cell-derived cardiomyocytes (hiPSC-CM), for therapeutic or testing applications, is both expensive and time-consuming. Herein, a scalable bioprocess for hiPSC-CM expansion in stirred-tank bioreactors (STB) is developed. By combining the continuous activation of the Wnt pathway, through perfusion of CHIR99021, within a mild hypoxia environment, the expansion of hiPSC-CM as aggregates is maximized, reaching 4 billion of pure hiPSC-CM in 2L STB. In particular, the importance of i) controlling the dissolved oxygen at 10% O2 to reduce reactive oxygen species production and upregulate genes involved in cell proliferation, resulting in higher expansion rates (tenfold) compared to normoxic conditions, and ii) maintaining constant power input per volume as a scale-up criteria is demonstrated. After expansion, hiPSC-CM further mature in culture, revealing more mature transcriptional signatures, higher sarcomere alignment and improved calcium handling. This new bioprocess opens the door to time- and cost-effective generation of hiPSC-CM.
Purchasable chemical space has grown rapidly into the tens of billions of molecules, providing unprecedented opportunities for ligand discovery but straining the tools that might exploit these molecules at scale. We have therefore developed ZINC-22, a database of commercially accessible small molecules derived from multi-billion-scale make-on-demand libraries. The new database and tools enable analog searching in this vast new space via a facile GUI, CartBlanche, drawing on similarity methods that scale sublinearly in the number of molecules. The new library also uses data organization methods, enabling rapid lookup of molecules and their physical properties, including conformations, partial atomic charges, c Log P values, and solvation energies, all crucial for molecule docking, which had become slow with older database organizations in previous versions of ZINC. As the libraries have continued to grow, we have been interested in finding whether molecular diversity has suffered, for instance, because certain scaffolds have come to dominate via easy analoging. This has not occurred thus far, and chemical diversity continues to grow with database size, with a log increase in Bemis-Murcko scaffolds for every two-log unit increase in database size. Most new scaffolds come from compounds with the highest heavy atom count. Finally, we consider the implications for databases like ZINC as the libraries grow toward and beyond the trillion-molecule range. ZINC is freely available to everyone and may be accessed at cartblanche22.docking.org, via Globus, and in the Amazon AWS and Oracle OCI clouds.
The COVID-19 pandemic has introduced new norms, such as social distancing, face masks, quarantine, lockdowns, travel restrictions, work/study from home, and business closures, to name a few. The pandemic's seriousness has made people vocal on social media, especially on microblogs such as Twitter. Since the early days of the outbreak, researchers have been collecting and sharing large-scale datasets of COVID-19 tweets. However, the existing datasets carry issues related to proportion and redundancy. We report that more than 500 million tweet identifiers point to deleted or protected tweets. To address these issues, this paper introduces an enriched global billion-scale English-language COVID-19 tweets dataset, BillionCOV, which contains 1.4 billion tweets originating from 240 countries and territories between October 2019 and April 2022. Importantly, BillionCOV facilitates researchers to filter tweet identifiers for efficient hydration. We anticipate that the dataset of this scale with global scope and extended temporal coverage will aid in obtaining a thorough understanding of the pandemic's conversational dynamics.
Graph computation approaches such as GraphChi and TurboGraph recently demonstrated that a single PC can perform efficient computation on billion-node graphs. To achieve high speed and scalability, they often need sophisticated data structures and memory management strategies. We propose a minimalist approach that forgoes such requirements, by leveraging the fundamental memory mapping (MMap) capability found on operating systems. We contribute: (1) a new insight that MMap is a viable technique for creating fast and scalable graph algorithms that surpasses some of the best techniques; (2) the design and implementation of popular graph algorithms for billion-scale graphs with little code, thanks to memory mapping; (3) extensive experiments on real graphs, including the 6.6 billion edge YahooWeb graph, and show that this new approach is significantly faster or comparable to the highly-optimized methods (e.g., 9.5× faster than GraphChi for computing PageRank on 1.47B edge Twitter graph). We believe our work provides a new direction in the design and development of scalable algorithms. Our packaged code is available at http://poloclub.gatech.edu/mmap/.
Human induced pluripotent stem (iPS) cell-derived hepatocyte-like cells are expected to be utilized in drug screening and regenerative medicine. However, hepatocyte-like cells have not been fully used in such applications because it is difficult to produce such cells on a large scale. In this study, we tried to establish a method to mass produce hepatocyte-like cells using a three-dimensional (3D) cell culture bioreactor called the Rotary Cell Culture System (RCCS). RCCS enabled us to obtain homogenous hepatocyte-like cells on a billion scale (>109 cells). The gene expression levels of some hepatocyte markers (alpha-1 antitrypsin, cytochrome (CYP) 1A2, CYP2D6, and hepatocyte nuclear factor 4alpha) were higher in 3D-cultured hepatocyte-like cells than in 2D-cultured hepatocyte-like cells. This result suggests that RCCS could provide more suitable conditions for hepatocyte maturation than the conventional 2D cell culture conditions. In addition, more than 90% of hepatocyte-like cells were positive for albumin and could uptake low-density lipoprotein in the culture medium. We succeeded in the large-scale production of homogenous and functional hepatocyte-like cells from human iPS cells. This technology will be useful in drug screening and regenerative medicine, which require enormous numbers of hepatocyte-like cells.
In natural language generation, abstractive summarization (AS) is advancing rapidly due to transformer-based language models (LMs). Although decoding strategies significantly influence generated summaries, their significance is often overlooked. Given the abundance of token selection heuristics and associated hyperparameters, the community needs guidance to make well-informed decisions based on the specific task and target metrics. To address this gap, we conduct a comparative assessment of the effectiveness and efficiency of decoding-time techniques for short, long, and multi-document AS. We explore over 3,500 combinations involving three widely used million-scale autoregressive encoder-decoder LMs, two billion-scale decoder-only LMs, six datasets, and nine decoding settings. Our findings highlight that optimized decoding choices can lead to substantial performance improvements. Alongside human evaluation, we quantitatively measure effects using ten automatic metrics, covering dimensions such as semantic similarity, factuality, compression, redundancy, and carbon footprint. To set the stage for differentiable selection and optimization of decoding options, we introduce Prism, a first-of-its-kind dataset that pairs AS gold input-output examples with our LM predictions across a diverse range of decoding options. The code and data are publicly available athttps://github.com/disi-unibo-nlp/prism.
Janus kinase 1 (JAK1) is a key regulator of cytokine signaling and a validated therapeutic target in autoimmune, inflammatory, and oncological disorders. However, existing JAK inhibitors such as Tofacitinib and Ruxolitinib are limited by their narrow pyrrolo-[2,3-d]-pyrimidine scaffold, leading to poor isoform selectivity, JAK3 cross-reactivity, and dose-limiting toxicity. Expanding the chemical space for JAK1 inhibition while achieving higher selectivity therefore represents a critical challenge in drug discovery. To overcome these limitations, we developed a deep learning (DL) based virtual screening framework (VS) that explicitly integrates protein flexibility with a billion-scale chemical exploration. Eight high-resolution JAK1 crystal structures were employed to capture conformational diversity of the ATP-binding pocket. Ensemble docking scores derived from these structures were used to train a deep neural network (DNN) classifier on rigorously curated data sets. The model was applied to over 1.1 billion commercially available compounds from the ZINC database, identifying 131,730 high-confidence candidates. Redocking analysis confirmed that 57% of these compounds consistently surpassed a stringent activity threshold across all receptor conformations, underscoring the robustness of the approach. Scaffold-based analysis of the top 10% candidates revealed 7652 unique chemotypes, with only 13 overlapping with scaffolds of known JAK1 inhibitors, highlighting the substantial novelty of the predicted chemical space. Furthermore, physicochemical and ADME filtering enriched for candidates with favorable drug-like properties. By explicitly embedding receptor flexibility into a scalable artificial intelligence framework, this study establishes a generalizable strategy for kinase-targeted drug discovery and opens new opportunities for selective JAK1 inhibitor development.
Driven by advances in supercomputing, the scale of scientific simulation data has grown dramatically. In fields such as cosmology, particle data have become a common representation, with state-of-the-art simulations now exceeding the trillion-particle mark. Consequently, the challenge of visually analyzing such massive datasets has become increasingly urgent. The traditional visual analysis workflow typically follows a "compression $\rightarrow$→ storage $\rightarrow$→ reconstruction $\rightarrow$→ visualization" pipeline. However, this process is hampered by an extremely time-consuming reconstruction stage, which severely impedes real-time interactive visualization. Moreover, in multi-time-step analyses, the enormous volume of reconstructed data creates significant I/O bottlenecks. In this work, we draw inspiration from 3D Gaussian splatting and compress the simulation data using Gaussian Mixture Models (GMMs), treating the resulting Gaussian kernels as fundamental rendering primitives. Our method renders billion-scale particles for each timestep in approximately 32 ms, requiring only 645 MB of GPU memory per timestep - nearly 20× smaller than the original 12 GB raw data. This eliminates costly reconstruction, accelerates the visual analysis pipeline, and overcomes I/O bottlenecks in multi-time-step analysis. Extensive experiments and comparisons across multiple datasets validate the effectiveness of our method.
Current AI-based Virtual Screening (VS) methods seek to manage ultra-large molecular libraries. To this end, they develop increasingly efficient heuristics to rank ligands by their predicted activity against a target protein. However, these methods remain computationally demanding due to the billion-scale compound libraries that must be evaluated without prior, informed guidance. This article proposes an offline/online method that: (1) Wisely selects (once and forall, offline phase) a small number of easy to compute features [Formula: see text] of both the amino acid sequence of the proteins ([Formula: see text]) and the molecular structure of the ligands ([Formula: see text]), and discretises their domains; this induces a low-dimensional finitisation of proteins' and ligands' chemical spaces. (2) Given a target protein [Formula: see text], immediately returns (online phase) a likelihood-based ranking of the classes of the ligands' chemical space, in descending order of the estimated probability that molecules in each class will achieve a satisfactory activity measurement against [Formula: see text]. This enables any VS method to prioritise the search to the most promising subsets of candidates. To ensure statistically robustness, our offline feature selection: (a) leverages knowledge stemming from a huge dataset of 2 559 403 entries (ligand-protein activity measurements) obtained by unifying the most representative sources regarding biochemical kinetics (Brenda, Sabio-rk, BindingDB) and augmented with 3781 features computed by 7 well-known third-party software tools; (b) explicitly aims at low-dimensional coarse-domain feature spaces; (c) takes proper countermeasures to prevent biases in the source data and overfitting; (d) supports iterative improvement of [Formula: see text] via an anytime offline algorithm and means to interactively exclude features deemed uninformative upon rankings inspection; (e) supports intepretability of the rankings by enabling inspection of the features' values characterising each ligand class. By evaluating our rankings on evaluation data (from PDBbind and additional BindingDB entries unsuitable for accurate statistical analysis), we demonstrate their effectiveness for library prioritisation. Specifically, our findings indicate that approximately 60% of the high-affinity ligands occur in the top 25% ranked ligands' classes, while 85% fall within the top 50%. Furthermore, we conduct retrospective analysises using AutoDock Vina scores for over 260 000 molecules across 58 medically relevant targets. Results demonstrate that our method cuts the number of dockings needed to retrieve an equivalent set of hits by up to [Formula: see text] on average versus unguided screening.
Examining vision-language alignment in multimodal embeddings is crucial for various tasks, such as evaluating generative models and filtering pretraining data. The intricate nature of high-dimensional features necessitates dimensionality reduction (DR) methods to explore alignment of multimodal embeddings. However, existing DR methods fail to account for cross-modal alignment metrics, resulting in severe occlusion of points with divergent metrics clustered together, inaccurate contour maps from over-aggregation, and insufficient support for multi-scale exploration. To address these problems, this paper introduces DKMap, a novel DR visualization technique for interactive exploration of multimodal embeddings through Dynamic Kernel enhanced projection. First, rather than performing dimensionality reduction and contour estimation sequentially, we introduce a kernel regression supervised t-SNE that directly integrates post-projection contour mapping into the projection learning process, ensuring cross-modal alignment mapping accuracy. Second, to enable multi-scale exploration with dynamic zooming and progressively enhanced local detail, we integrate validation-constrained a refinement of a generalized t-kernel with quad-tree-based multi-resolution technique, ensuring reliable kernel parameter tuning without overfitting. DKMap is implemented as a multi-platform visualization tool, featuring a web-based system for interactive exploration and a Python package for computational notebook analysis. Quantitative comparisons with baseline DR techniques demonstrate DKMap's superiority in accurately mapping cross-modal alignment metrics. We further demonstrate generalizability and scalability of DKMap with three usage scenarios, including visualizing million-scale text-to-image corpus, comparatively evaluating generative models, and exploring a billion-scale pretraining dataset.
Chlorine (Cl2) is a hazardous industrial gas and a choking agent, making highly sensitive ppb-level detection with tunable selectivity toward chemical warfare agent-related (CWA-related) analytes important. Here, we synthesized Ag-functionalized SnO2 and WO3 nanofibers by electrospinning and investigated how the host oxide governs Cl2 sensing and selectivity toward CWA-related analytes. The optimized Ag-functionalized SnO2 sensor (AgS5) showed a response of 6.89 to 500 ppb Cl2, an experimentally validated detection limit of 10 ppb, and a calculated lower limit of detection of 0.12 ppb. The host oxide also strongly altered selectivity: SnO2-based sensors showed higher responses to Cl2 and hydrogen cyanide, whereas WO3-based sensors were more sensitive to 2-chloroethyl ethyl sulfide and methyl salicylate. X-ray photoelectron spectroscopy revealed host-dependent Ag speciation, with Ag+ favored on SnO2 and Ag0 on WO3. Temperature-modulated measurements and density functional theory calculations further showed that this difference changes interfacial charge-transport barriers and adsorption energetics, thereby governing both sensitivity and selectivity. These results suggest a trace-level Cl2 sensor and demonstrate that the sensing characteristics of noble metal-decorated chemiresistors can be rationally tuned through host-oxide selection, providing a general design strategy for selective, ppb-level Cl2 detection.
Real-life graphs often exhibit intricate dynamics that evolve continuously over time. To effectively represent continuous-time dynamic graphs (CTDGs), various temporal graph neural networks (TGNNs) have been developed to model their dynamics and topological structures in Euclidean space. Despite their notable achievements, the performance of Euclidean-based TGNNs is limited and bounded by the representation capabilities of Euclidean geometry, particularly for complex graphs with hierarchical and power-law structures. This is because Euclidean space does not have enough room (its volume grows polynomially with respect to radius) to learn hierarchical structures that expand exponentially. As a result, this leads to high-distortion embeddings and suboptimal temporal graph representations. To break the limitations and enhance the representation capabilities of TGNNs, in this article, we propose a scalable and effective TGNN with hyperbolic geometries for CTDG representation (called ${\mathrm { STGN}}^{h}$ ). It captures evolving behaviors and stores hierarchical structures simultaneously by integrating a memory-based module and a structure-based module into a unified framework, which can scale to billion-scale graphs. Concretely, a simple hyperbolic update gate (HuG) is designed as the memory-based module to store temporal dynamics efficiently; for the structure-based module, we propose an effective hyperbolic temporal Transformer (HyT) model to capture complex graph structures and generate up-to-date node embeddings. Extensive experimental results on a variety of medium-scale and billion-scale graphs demonstrate the superiority of the proposed ${\mathrm { STGN}}^{h}$ for CTDG representation, as it significantly outperforms baselines in various downstream tasks.
The accelerating growth of make-on-demand chemical libraries provides unprecedented opportunities to identify starting points for drug discovery with virtual screening. However, these multi-billion-scale libraries are challenging to screen, even for the fastest structure-based docking methods. Here we explore a strategy that combines machine learning and molecular docking to enable rapid virtual screening of databases containing billions of compounds. In our workflow, a classification algorithm is trained to identify top-scoring compounds based on molecular docking of 1 million compounds to the target protein. The conformal prediction framework is then used to make selections from the multi-billion-scale library, reducing the number of compounds to be scored by docking. The CatBoost classifier showed an optimal balance between speed and accuracy and was used to adapt the workflow for screens of ultralarge libraries. Application to a library of 3.5 billion compounds demonstrated that our protocol can reduce the computational cost of structure-based virtual screening by more than 1,000-fold. Experimental testing of predictions identified ligands of G protein-coupled receptors and demonstrated that our approach enables discovery of compounds with multi-target activity tailored for therapeutic effect.
Virtual screening (VS) in drug design employs computational methodologies to systematically rank molecules from a virtual compound library based on predicted features related to their biological activities or chemical properties. The recent expansion in commercially accessible compound libraries and the advancements in artificial intelligence (AI) and computational power - including enhanced central processing units (CPUs), graphics processing units (GPUs), high-performance computing (HPC), and cloud computing - have significantly expanded our capacity to screen libraries containing over 109 molecules. Herein, we review the concept of ultra-large virtual screening (ULVS), focusing on the various algorithms and methodologies employed for virtual screening at this scale. In this context, we present the software utilized, applications, and results of different approaches, such as brute force docking, reaction-based docking approaches, machine learning (ML) strategies applied to docking or other VS methods, and similarity/pharmacophore search-based techniques. These examples represent a paradigm shift in the drug discovery process, demonstrating not only the feasibility of billion-scale compound screening but also their potential to identify hit candidates and increase the structural diversity of novel compounds with biological activities.
Content-Based Multimedia Retrieval (CBMR) has become very popular in several applications, driven by the growing routine use of multimedia data. Since the datasets used in real-world applications are very large and descriptor's dimensionality is high, querying is an expensive, albeit important functionality. Further, exact search is prohibitive in most cases, motivating the use of Approximate Nearest Neighbour Search (ANNS) algorithms, trading accuracy for performance. These have been mainly developed targeting a sequential execution in a single node. However, the large and increasing datasets used and the high query loads submitted to those systems typically surpass the memory and computing resources available in a single node. This motivated the development of parallel distributed memory ANNS solutions to meet the computing capabilities required by those applications. A common problem that must be handled when using distributed memory systems is data partitioning and its impact on load imbalance. Several data partitioning approaches have already been proposed, including elaborated spatial-aware strategies. However, little effort has been put into carefully analyzing the performance of those strategies at scale. Here, we evaluated the commonly used data partitioning strategies in ANNS and identified their limitations to propose a novel class of partitioning algorithms that can minimize load imbalance while improving data locality to attain high performance on the distributed memory search. Experimentally, we found that our proposed algorithms (SABBS and SABBSR) improved search performance by up to 1.64× compared to the best previous solution. In a distributed memory weak scaling evaluation, with up to 12 billion 128-dimensional descriptors and 60 compute nodes, the gains were maintained as the system scaled with our novel approaches. These results demonstrate the efficiency of our new algorithms for billion-scale ANNS and the importance of considering not only data locality but also data and load imbalance in the data partitioning.
Pre-training deep learning models with large data sets of natural images, such as ImageNet, has become the standard for endoscopic image analysis. This approach is generally superior to training from scratch, due to the scarcity of high-quality medical imagery and labels. However, it is still unknown whether the learned features on natural imagery provide an optimal starting point for the downstream medical endoscopic imaging tasks. Intuitively, pre-training with imagery closer to the target domain could lead to better-suited feature representations. This study evaluates whether leveraging in-domain pre-training in gastrointestinal endoscopic image analysis has potential benefits compared to pre-training on natural images. To this end, we present a dataset comprising of 5,014,174 gastrointestinal endoscopic images from eight different medical centers (GastroNet-5M), and exploit self-supervised learning with SimCLRv2, MoCov2 and DINO to learn relevant features for in-domain downstream tasks. The learned features are compared to features learned on natural images derived with multiple methods, and variable amounts of data and/or labels (e.g. Billion-scale semi-weakly supervised learning and supervised learning on ImageNet-21k). The effects of the evaluation is performed on five downstream data sets, particularly designed for a variety of gastrointestinal tasks, for example, GIANA for angiodyplsia detection and Kvasir-SEG for polyp segmentation. The findings indicate that self-supervised domain-specific pre-training, specifically using the DINO framework, results into better performing models compared to any supervised pre-training on natural images. On the ResNet50 and Vision-Transformer-small architectures, utilizing self-supervised in-domain pre-training with DINO leads to an average performance boost of 1.63% and 4.62%, respectively, on the downstream datasets. This improvement is measured against the best performance achieved through pre-training on natural images within any of the evaluated frameworks. Moreover, the in-domain pre-trained models also exhibit increased robustness against distortion perturbations (noise, contrast, blur, etc.), where the in-domain pre-trained ResNet50 and Vision-Transformer-small with DINO achieved on average 1.28% and 3.55% higher on the performance metrics, compared to the best performance found for pre-trained models on natural images. Overall, this study highlights the importance of in-domain pre-training for improving the generic nature, scalability and performance of deep learning for medical image analysis. The GastroNet-5M pre-trained weights are made publicly available in our repository: huggingface.co/tgwboers/GastroNet-5M_Pretrained_Weights.
Pancreatic ductal adenocarcinoma is an intractable disease with frequent recurrence after resection and adjuvant therapy. The present study aimed to clarify whether artificial intelligence-assisted analysis of histopathological images can predict recurrence in patients with pancreatic ductal adenocarcinoma who underwent resection and adjuvant chemotherapy with tegafur/5-chloro-2,4-dihydroxypyridine/potassium oxonate. Eighty-nine patients were enrolled in the study. Machine-learning algorithms were applied to 10-billion-scale pixel data of whole-slide histopathological images to generate key features using multiple deep autoencoders. Areas under the curve were calculated from receiver operating characteristic curves using a support vector machine with key features alone and by combining with clinical data (age and carbohydrate antigen 19-9 and carcinoembryonic antigen levels) for predicting recurrence. Supervised learning with pathological annotations was conducted to determine the significant features for predicting recurrence. Areas under the curves obtained were 0.73 (95% confidence interval, 0.59-0.87) by the histopathological data analysis and 0.84 (95% confidence interval, 0.73-0.94) by the combinatorial analysis of histopathological data and clinical data. Supervised learning model demonstrated that poor tumor differentiation was significantly associated with recurrence. Results indicate that machine learning with the integration of artificial intelligence-driven evaluation of histopathological images and conventional clinical data provides relevant prognostic information for patients with pancreatic ductal adenocarcinoma.
Magic-angle spinning (MAS) solid-state NMR methods are crucial in many areas of biology and materials science. Conventional probe designs have often been specified with 0.1 part per million (ppm) or 100 part per billion (ppb) magnetic field resolution, which is a limitation for many modern scientific applications. Here we describe a novel 5-mm MAS module design that significantly improves the linewidth and line shape for solid samples by an improved understanding of the magnetic susceptibility of probe materials and geometrical symmetry considerations, optimized to minimize the overall perturbation to the applied magnetic field (B0). The improved spinning module requires only first and second order shimming adjustments to achieve a sub-Hz resolution of 13C resonances of adamantane at 150 MHz Larmor frequency (14.1Tesla magnetic field). Minimal use of third and higher order shims improves experimental reproducibility upon sample changes and the exact placement within the magnet. Furthermore, the shimming procedure is faster, and the required gradients smaller, thus minimizing thermal drift of the room temperature (RT) shims. We demonstrate these results with direct polarization (Bloch decay) and cross polarization experiments on adamantane over a range of sample geometries and with multiple superconducting magnet systems. For a direct polarization experiment utilizing the entire active sample volume of a 5-mm rotor (90 µl), we achieved full width at half maximum (FWHM) of 0.76 Hz (5 ppb) and baseline resolved the 13C satellite peaks for adamantane as a consequent of the 7.31 Hz (59 ppb) width at 2% intensity. We expect these approaches to be increasingly pivotal for high-resolution solid-state NMR spectroscopy at and above 1 GHz 1H frequencies.
Protein-ligand docking is a computational method for identifying drug leads. The method is capable of narrowing a vast library of compounds down to a tractable size for downstream simulation or experimental testing and is widely used in drug discovery. While there has been progress in accelerating scoring of compounds with artificial intelligence, few works have bridged these successes back to the virtual screening community in terms of utility and forward-looking development. We demonstrate the power of high-speed ML models by scoring 1 billion molecules in under a day (50 k predictions per GPU seconds). We showcase a workflow for docking utilizing surrogate AI-based models as a pre-filter to a standard docking workflow. Our workflow is ten times faster at screening a library of compounds than the standard technique, with an error rate less than 0.01% of detecting the underlying best scoring 0.1% of compounds. Our analysis of the speedup explains that another order of magnitude speedup must come from model accuracy rather than computing speed. In order to drive another order of magnitude of acceleration, we share a benchmark dataset consisting of 200 million 3D complex structures and 2D structure scores across a consistent set of 13 million "in-stock" molecules over 15 receptors, or binding sites, across the SARS-CoV-2 proteome. We believe this is strong evidence for the community to begin focusing on improving the accuracy of surrogate models to improve the ability to screen massive compound libraries 100 × or even 1000 × faster than current techniques and reduce missing top hits. The technique outlined aims to be a fast drop-in replacement for docking for screening billion-scale molecular libraries.