Long Short-Term Memory (LSTM) models are trained to predict forecast errors for the High-Resolution Rapid Refresh (HRRR) model using the New York State Mesonet and Oklahoma State Mesonet near-surface weather observations as ground truth. When evaluated using mean-absolute-error and percent improvement relative to the HRRR, LSTMs predict precipitation error most accurately, providing, on average, a 48% improvement relative to the HRRR forecast, followed by wind error, providing, on average, a 15% improvement, and then temperature error, providing, on average, a 25% improvement. Precipitation errors exhibit an asymmetry, with overforecast precipitation detected more accurately than underforecast, while wind error predictions are consistent across over- and underforecast predictions. Temperature error predictions are relatively accurate but smoother, with respect to variance, than true observations. This paper describes an overview of LSTM performance with the expressed intent of providing forecasters with real-time predictions of forecast error at the point of use within the New York State and Oklahoma State Mesonets. In practice, the predicted errors can be used to adjust determinis
Risk communication in times of disasters is complex, involving rapid and diverse communication in social networks (i.e., public and/or private agencies; local residents) as well as limited mobilization capacity and operational constraints of physical infrastructure networks. Despite a growing literature on infrastructure interdependencies and co-dependent social-physical systems, an in-depth understanding of how risk communication in online social networks weighs into physical infrastructure networks during a major disaster remains limited, let alone in compounding risk events. This study analyzes large-scale datasets of crisis mobility and activity-related social interactions and concerns available through social media (Twitter) for communities that were impacted by an ice storm (Oct. 2020) in Oklahoma. Compounded by the COVID-19 pandemic, Oklahoma residents faced this historic ice storm (Oct. 26, 2020-Oct. 29, 2020) that caused devastating traffic impacts (among others) due to excessive ice accumulation. By using the recently released academic Application Programming Interface (API) by Twitter that provides complete and unbiased data, geotagged tweets (approx. 210K) were collecte
We develop a statistical method for identifying induced seismicity from large datasets and apply the method to decades of wastewater disposal and seismicity data in California and Oklahoma. The method is robust against a variety of potential pitfalls. The study regions are divided into gridblocks. We use a longitudinal study design, seeking associations between seismicity and wastewater injection along time-series within each gridblock. The longitudinal design helps control for non-random application of wastewater injection. We define a statistical model that is flexible enough to describe the seismicity observations, which have temporal correlation and high kurtosis. In each gridblock, we find the maximum likelihood estimate for a model parameter that relates induced seismicity hazard to total volume of wastewater injected each year. To assess significance, we compute likelihood ratio test statistics in each gridblock and each state, California and Oklahoma. Resampling is used to empirically derive reference distributions used to estimate p-values from the likelihood ratio statistics. In Oklahoma, the analysis finds with extremely high confidence that seismicity associated with wa
The paper investigates gender differences in entrepreneurship by exploiting a large-scale land lottery in Oklahoma at the turn of the 20$^{\text{th}}$ century. Lottery winners claimed land in the order in which their names were drawn, so the draw number is an approximate rank ordering of lottery wealth. This mechanism allows for the estimation of a dose-response function, which relates each draw number to the expected outcome under each draw. I estimate dose-response functions on a linked dataset of lottery winners and land patent records, and find the probability of purchasing land from the government to be decreasing as a function of lottery wealth, which is evidence for the presence of liquidity constraints. I find female winners were more effective in leveraging lottery wealth to purchase additional land, as evidenced by significantly higher median dose-responses compared to those of male winners. For a sample of winners linked to the 1910 Census, I find that male winners have higher median dose-responses compared to female winners in terms of farm or home ownership. These results suggest that liquidity constraints may have been more binding for female entrepreneurs in the mark
Parts of Texas, Oklahoma, and Kansas have experienced increased rates of seismicity in recent years, providing new datasets of earthquake recordings to develop ground motion prediction models for this particular region of the Central and Eastern North America (CENA). This paper outlines a framework for using Artificial Neural Networks (ANNs) to develop attenuation models from the ground motion recordings in this region. While attenuation models exist for the CENA, concerns over the increased rate of seismicity in this region necessitate investigation of ground motions prediction models particular to these states. To do so, an ANN-based framework is proposed to predict peak ground acceleration (PGA) and peak ground velocity (PGV) given magnitude, earthquake source-to-site distance, and shear wave velocity. In this framework, approximately 4,500 ground motions with magnitude greater than 3.0 recorded in these three states (Texas, Oklahoma, and Kansas) since 2005 are considered. Results from this study suggest that existing ground motion prediction models developed for CENA do not accurately predict the ground motion intensity measures for earthquakes in this region, especially for th
This paper has been withdrawn by the author due to obsolete content.
Quantitative interpretation of the tidal response of water levels measured in wells has long been made either with a model for perfectly confined aquifers or with a model for purely unconfined aquifers. However, many aquifers may be neither totally confined nor purely unconfined at the frequencies of tidal loading but behave somewhere between the two end members. Here we present a more general model for the tidal response of groundwater in aquifers with both horizontal and vertical flow. The model has three independent parameters: the transmissivity and storativity of the aquifer and the specific leakage of the leaking aquitard. If transmissivity and storativity are known independently, this model may be used to estimate aquitard leakage from the phase shift and amplitude ratio of water level in wells obtained from tidal analysis. We apply the model to interpret the tidal response of water level in a USGS deep monitoring well installed in the Arbuckle aquifer in Oklahoma, into which massive amount of wastewater co-produced from hydrocarbon exploration has been injected. The analysis shows that the Arbuckle aquifer is leaking significantly at this site. We suggest that the present m
This paper proposes a time-warping transfer learning method, a technique for temporally rescaling the learned dynamics of a recurrent neural network (RNN) with a Long Short-Term Memory (LSTM) layer to enable task transfer across fuel moisture classes. Fuel moisture content (FMC) is divided into idealized classes based on characteristic lag time. Large quantities of real-time data are available for 10h fuels from sensors on weather stations, but observations of other fuel classes are sparse in space and time. We use transfer learning to adapt an RNN pretrained on 10h FMC to predict FMC for 1h, 100h, and 1000h fuels. We validate this method using data from a landmark field study conducted in Oklahoma that was used to calibrate the state-of-the-art Nelson fuel moisture model.
Steep-profiled Highway Railway Grade Crossings (HRGCs) pose safety hazards to vehicles with low ground clearance, which may become stranded on the tracks, creating risks of train vehicle collisions. This research develops a framework for network level evaluation of hang-up susceptibility of HRGCs. Profile data from different crossings in Oklahoma were collected using both a walking profiler and the Pave3D8K Laser Imaging System. A hybrid deep learning model, combining Long Short Term Memory (LSTM) and Transformer architectures, was developed to reconstruct accurate HRGC profiles from Pave3D8K Laser Imaging System data. Vehicle dimension data from around 350 specialty vehicles were collected at various locations across Oklahoma to enable up-to-date statistical design dimensions. Hang-up susceptibility was analyzed using three vehicle dimension scenarios: (a) median dimension (median wheelbase and ground clearance), (b) 75-25 percentile dimension (75 percentile wheelbase, 25 percentile ground clearance), and (c) worst case dimension (maximum wheelbase and minimum ground clearance). Results indicate 70, 80, and 95 crossings at the highest hang-up risk levels under these scenarios, res
We develop a framework for causal inference with continuous spatiotemporal point-process outcomes under cell-level interventions and outcome spillover. Potential outcomes are indexed by full treatment allocations, and the observed post-treatment process is represented as an unlabelled superposition of latent control and treatment components. On the observed design support, expected post-treatment event counts in any spacetime region under a given treatment allocation are identified under consistency, exchangeability, and positivity; off-support contrasts are identified relative to a regime-stable structural point-process model. Estimation is likelihood-based and implemented with stochastic EM. To understand when this is feasible, we analyse a predictable blockwise hard-EM surrogate and show nonasymptotic contraction of estimation error to a statistical floor governed by locally ambiguous regions. This yields plug-in guarantees for cell-level and global causal functionals, and clarifies the additional array conditions needed for unnormalised growing-window contrasts. The framework covers history dependent spatiotemporal point processes including Poisson and Hawkes models, with appli
In this work, we propose new matrix- and tensor-based methodologies for estimating multivariate intensity functions of inhomogeneous point processes. By viewing multivariate intensity functions as infinite-dimensional matrices or tensors within function spaces, our algorithms attain the optimal bias-variance trade-off, yielding rate-optimal estimation error, with model complexity governed by matrix or tensor ranks. They substantially improve estimation accuracy, while simultaneously reducing computational cost. To illustrate the adaptivity of the proposed framework, we show that many fundamental classes of multivariate functions, including additive and mean-field models, admit finite-rank tensor representations. We apply our method to a four-dimensional U.S. Geological Survey earthquake dataset, comprising features such as latitude, longitude, depth, and magnitude. Our tensor estimator recovers localized seismicity patterns (California, Oklahoma, Pacific Northwest, north-central U.S.), whereas the kernel baseline oversmooths them.
Surface winds can vary substantially from one minute to the next, so there is scope for studying its variation on this fine time scale. Restricting to the month of June to minimize seasonality, this work develops a range of machine learning models for generating realistic time series of surface wind vectors at a site in Lamont, Oklahoma based on more than 30 years of high quality measurements at the minute time scale. Such a generator could be used as an input into models from a range of disciplines, notably for wind energy, but also wildfire spread and aviation, among others. The data show complex diurnal structures in both wind speed and direction that would be challenging to capture with standard time series models, so we consider a number of machine learning approaches to producing a stochastic wind generator based on time vector-quantized variational autoencoders. We consider generating a day's worth of data at a time and generating a day of wind vectors conditional on the previous day's winds. We also study methods for incorporating a discrete weather state variable in the generator. We evaluate the generators using a wide range of formal and informal methods. The best of the
The national forecasting competition WxChallenge, brainchild of Brad Illston at the University of Oklahoma in 2005, has become a cherished institution played across the United States each year. Participants include students, faculty, alumni, and industry professionals. However, forecasts are given as scalar values without expression of uncertainty, probabilities being a keystone of meteorological forecasting today, and previous attempts to add probabilistic elements to WxChallenge have failed partly due to challenges in making probability forecasting accessible to all, and inability to combine scores with different units while also appropriately rewarding forecasts using proper scoring rules. Much of the competition's maintenance relies on dedicated volunteers, highlighting need for more automation. Hence I propose three new features: (1) automated forecast problems based on morning ensemble guidance, forming prediction baselines, thresholds over which the players demonstrate skill in their later forecast; (2) a spread betting game, where the players allocate 100 confidence credits to the over-under for exceeding a percentile (e.g., 50pc) threshold of a variable (e.g., maximum temp
Unplanned power outages cost the US economy over $150 billion annually, partly due to predictive maintenance (PdM) models that overlook spatial, temporal, and causal dependencies in grid failures. This study introduces a multilayer Graph Neural Network (GNN) framework to enhance PdM and enable resilience-based substation clustering. Using seven years of incident data from Oklahoma Gas & Electric (292,830 records across 347 substations), the framework integrates Graph Attention Networks (spatial), Graph Convolutional Networks (temporal), and Graph Isomorphism Networks (causal), fused through attention-weighted embeddings. Our model achieves a 30-day F1-score of 0.8935 +/- 0.0258, outperforming XGBoost and Random Forest by 3.2% and 2.7%, and single-layer GNNs by 10 to 15 percent. Removing the causal layer drops performance to 0.7354 +/- 0.0418. For resilience analysis, HDBSCAN clustering on HierarchicalRiskGNN embeddings identifies eight operational risk groups. The highest-risk cluster (Cluster 5, 44 substations) shows 388.4 incidents/year and 602.6-minute recovery time, while low-risk groups report fewer than 62 incidents/year. ANOVA (p < 0.0001) confirms significant inter-c
A relevant question when analyzing spatial point patterns is that of spatial randomness. More specifically, before any model can be fit to a point pattern a first step is to test the data for departures from complete spatial randomness (CSR). Traditional techniques employ distance or quadrat counts based methods to test for CSR based on batched data. In this paper, we consider the practical scenario of testing for CSR when the data are available sequentially (i.e., online). We present a sequential testing methodology called as {\em PRe-process} that is based on e-values and is a fast, efficient and nonparametric method. Simulation experiments with the truth departing from CSR in two different scenarios show that the method is effective in capturing inhomogeneity over time. Two real data illustrations considering lung cancer cases in the Chorley-Ribble area, England from 1974 - 1983 and locations of earthquakes in the state of Oklahoma, USA from 2000 - 2011 demonstrate the utility of the PRe-process in sequential testing of CSR.
Hump crossings, or high-profile Highway Railway Grade Crossings (HRGCs), pose safety risks to highway vehicles due to potential hang-ups. These crossings typically result from post-construction railway track maintenance activities or non-compliance with design guidelines for HRGC vertical alignments. Conventional methods for measuring HRGC profiles are costly, time-consuming, traffic-disruptive, and present safety challenges. To address these issues, this research employed advanced, cost-effective techniques and innovative modeling approaches for HRGC profile measurement. A novel hybrid deep learning framework combining Long Short-Term Memory (LSTM) and Transformer architectures was developed by utilizing instrumentation and ground truth data. Instrumentation data were gathered using a highway testing vehicle equipped with Inertial Measurement Unit (IMU) and Global Positioning System (GPS) sensors, while ground truth data were obtained via an industrial-standard walking profiler. Field data was collected at the Red Rock Railroad Corridor in Oklahoma. Three advanced deep learning models Transformer-LSTM sequential (model 1), LSTM-Transformer sequential (model 2), and LSTM-Transforme
Seismic data contain complex temporal information that arrives at high speed and has a large, even potentially unbounded volume. The explosion of temporally correlated streaming data from advanced seismic sensors poses analytical challenges due to its sheer volume and real-time nature. Sampling, or data reduction, is a natural yet powerful tool for handling large streaming data while balancing estimation accuracy and computational cost. Currently, data reduction methods and their statistical properties for streaming data, especially streaming autoregressive time series, are not well-studied in the literature. In this article, we propose an online leverage-based sequential data reduction algorithm for streaming autoregressive time series with application to seismic data. The proposed Sequential Leveraging Sampling (SLS) method selects only one consecutively recorded block from the data stream for inference. While the starting point of the SLS block is chosen using a random mechanism based on streaming leverage scores of data, the block size is determined by a sequential stopping rule. The SLS block offers efficient sample usage, as evidenced by our results confirming asymptotic norm
The integration of social media and artificial intelligence (AI) into disaster management, particularly for earthquake response, represents a profound evolution in emergency management practices. In the digital age, real-time information sharing has reached unprecedented levels, with social media platforms emerging as crucial communication channels during crises. This shift has transformed traditional, centralized emergency services into more decentralized, participatory models of disaster situational awareness. Our study includes an experimental analysis of 8,900 social media interactions, including 2,920 posts and 5,980 replies on X (formerly Twitter), following a magnitude 5.1 earthquake in Oklahoma on February 2, 2024. The analysis covers data from the immediate aftermath and extends over the following seven days, illustrating the critical role of digital platforms in modern disaster response. The results demonstrate that social media platforms can be effectively used as real-time situational awareness tools, delivering critical information to society and authorities during emergencies.
The aim of this study is to look at predicting whether a person will complete a drug and alcohol rehabilitation program and the number of times a person attends. The study is based on demographic data obtained from Substance Abuse and Mental Health Services Administration (SAMHSA) from both admissions and discharge data from drug and alcohol rehabilitation centers in Oklahoma. Demographic data is highly categorical which led to binary encoding being used and various fairness measures being utilized to mitigate bias of nine demographic variables. Kernel methods such as linear, polynomial, sigmoid, and radial basis functions were compared using support vector machines at various parameter ranges to find the optimal values. These were then compared to methods such as decision trees, random forests, and neural networks. Synthetic Minority Oversampling Technique Nominal (SMOTEN) for categorical data was used to balance the data with imputation for missing data. The nine bias variables were then intersectionalized to mitigate bias and the dual and triple interactions were integrated to use the probabilities to look at worst case ratio fairness mitigation. Disparate Impact, Statistical Pa
Heliotropes are passive solar hot air balloons that are capable of achieving nearly level flight within the lower stratosphere for several hours. These inexpensive flight platforms enable stratospheric sensing with high-cadence enabled by the low cost to manufacture, but their performance has not yet been assessed systematically. During July to September of 2021, 29 heliotropes were successfully launched from Oklahoma and achieved float altitude as part of the Balloon-based Acoustic Seismology Study (BASS). All of the heliotrope envelopes were nearly identical with only minor variations to the flight line throughout the campaign. Flight data collected during this campaign comprise a large sample to characterize the typical heliotrope flight behavior during launch, ascent, float, and descent. Each flight stage is characterized, dependence on various parameters is quantified, and a discussion of nominal and anomalous flights is provided.