Turning a promising economic idea into a credible empirical finding is, in practice, an expensive undertaking: it demands a great deal of specialised computation, and the results are seldom released in a form that others can check or build upon. Econstellar is our response. It is an open, publicly served research engine that runs publication-grade financial econometrics from an ordinary web browser and explains what the results mean, so that a reader does not merely read a finding but can re-run it, vary its inputs, and trace exactly how it was produced. Three choices give the system its character. The heavy computation is placed on the processor that suits it, rather than forced onto hardware ill-matched to the task, which is much of the reason analysis of this kind is so rarely served to the public. An artificial-intelligence assistant selects and interprets the analyses but never originates a number, so every quantity it reports is a real computation the reader can reproduce. And the engine a visitor exercises is the same code that produced the figures in our published research. We expose seventeen econometric methods, each reported with a verified live value and reproducible at
"Vibe coding" and "vibe analytics" have been framed as a democratization of technical capability. This paper argues that AI-assisted methodology more broadly, or what I call "vibe methodology," also democratizes the failure modes specific to each domain. When AI assists with methods whose validity depends on assumptions that cannot be verified from the output alone (a class I call "vibe inference"), the failure surface is structurally different: the output does not reliably signal invalidity, and when it does, recognizing the signal requires the expertise the workflow bypasses. I focus on "vibe econometrics," the subset of AI-assisted causal analysis where identification can be named faster than it can be audited. The claim of this paper is not that AI invents inferential failures that did not previously exist, but that it changes their incidence, observability, and persuasive force enough to create a practically distinct governance problem. This results in three failure modes: method-data mismatch, where AI bypasses expertise at execution; confidence laundering, where AI amplifies the credibility of formatted output; and invisible forking, which spans both. What is new is not the
Regression models and Vector Autoregressive Models (VARs) play crucial roles in econometrics by allowing the analysis of multiple variables simultaneously. Despite their utility, these models face challenges like underfitting and overfitting, especially when determining the optimal model specification, which can lead to significant computational costs. To address these challenges, econometricians often rely on widely adopted model selection criteria such as the Akaike Information Criterion (AIC) and Bayesian Information Criterion (BIC). These criteria help balance model complexity and goodness of fit, aiding in the selection of the most suitable model specification for the given data. Nonetheless, there is a notable gap in existing research concerning the correct specification of these models, particularly in determining the optimal number of states a system can assume. Addressing this gap, we introduce a combinatorial framework designed to calculate the potential number of states in such econometric models. Our approach involves delineating four distinct stages in model development, each offering a range of specifications. This method enables a comprehensive combinatorial calculat
Can AI effectively perform complex econometric analysis traditionally requiring human expertise? This paper evaluates AI agents' capability to master econometrics, focusing on empirical analysis performance. We develop ``MetricsAI'', an Econometrics AI Agent built on the open-source MetaGPT framework. This agent exhibits outstanding performance in: (1) planning econometric tasks strategically, (2) generating and executing code, (3) employing error-based reflection for improved robustness, and (4) allowing iterative refinement through multi-round conversations. We construct two datasets from academic coursework materials and published research papers to evaluate performance against real-world challenges. Comparative testing shows our domain-specialized AI agent significantly outperforms both benchmark large language models (LLMs) and general-purpose AI agents. This work establishes a testbed for exploring AI's impact on social science research and enables cost-effective integration of domain expertise, making advanced econometric methods accessible to users with minimal coding skills. Furthermore, our AI agent enhances research reproducibility and offers promising pedagogical applic
This paper examines the relationship between Official Development Assistance (ODA) and conflict in the ten largest aid-receiving African countries between 2009 and 2023. Using Ordinary Least Squares, Principal Component Analysis, and Ridge (L2) regression, the study assesses whether conflict, proxied by political stability, governance indicators, and macroeconomic conditions, systematically influences aid inflows. Results reveal a nuanced relationship. Pooled regressions indicate that aid is positively associated with poverty, inflation, and fragility, while voice and accountability are negatively related to ODA. Fixed-effects estimates instead show positive associations between aid, political stability, and GDP per capita over time, alongside negative correlations with perceived corruption. Ridge regression confirms the robustness of various governance variables under multicollinearity. Overall, donors appear responsive to both humanitarian need and institutional quality, producing an aid-conflict-institutions trilemma: aid is most concentrated where conflict risk and institutional weakness are greatest, yet these same conditions which constrain aid effectiveness. The paper contri
The Growth-at-Risk (GaR) framework has garnered attention in recent econometric literature, yet current approaches implicitly assume a constant Pareto exponent. We introduce novel and robust econometrics to estimate the tails of GaR based on a rigorous theoretical framework and establish validity and effectiveness. Simulations demonstrate consistent outperformance relative to existing alternatives in terms of predictive accuracy. We perform a long-term GaR analysis that provides accurate and insightful predictions, effectively capturing financial anomalies better than current methods.
We study the long-standing problem of determining the number of principal components in econometric applications from a selective inference perspective. We consider i.i.d. observations from a $p$-dimensional random vector with $p<n$ and define the ``true'' dimensionality as the rank of the population covariance matrix. Building on the sequential testing viewpoint, we propose a data-driven procedure that estimates $\rank(Σ_X)$ using a statistic that depends on the eigenvalues of the sample covariance matrix. While the test statistic shares the functional form of its fixed design counterpart Choi et al. (2017), our analysis departs from the non-stochastic setting by treating the design as random and by avoiding parametric Gaussian assumptions. Under a locally defined null hypothesis, we establish asymptotically exact type~I error controls in the sequential testing procedure, with simulation results indicating empirical validity of the proposed method.
This paper investigates Large Language Models (LLMs) ability to assess the economic soundness and theoretical consistency of empirical findings in spatial econometrics. We created original and deliberately altered "counterfactual" summaries from 28 published papers (2005-2024), which were evaluated by a diverse set of LLMs. The LLMs provided qualitative assessments and structured binary classifications on variable choice, coefficient plausibility, and publication suitability. The results indicate that while LLMs can expertly assess the coherence of variable choices (with top models like GPT-4o achieving an overall F1 score of 0.87), their performance varies significantly when evaluating deeper aspects such as coefficient plausibility and overall publication suitability. The results further revealed that the choice of LLM, the specific characteristics of the paper and the interaction between these two factors significantly influence the accuracy of the assessment, particularly for nuanced judgments. These findings highlight LLMs' current strengths in assisting with initial, more surface-level checks and their limitations in performing comprehensive, deep economic reasoning, suggesti
The chronological hierarchy and classification of psychological types of individuals are examined. The anomalous nature of psychological activity in individuals involved in scientific work is highlighted. Certain aspects of the introverted thinking type in scientific activities are analyzed. For the first time, psychological archetypes of scientists with pronounced introversion are postulated in the context of twelve hypotheses about the specifics of professional attributes of introverted scientific activities. A linear regression and Bayesian equation are proposed for quantitatively assessing the econometric degree of introversion in scientific employees, considering a wide range of characteristics inherent to introverts in scientific processing. Specifically, expressions for a comprehensive assessment of introversion in a linear model and the posterior probability of the econometric (scientometric) degree of introversion in a Bayesian model are formulated. The models are based on several econometric (scientometric) hypotheses regarding various aspects of professional activities of introverted scientists, such as a preference for solo publications, low social activity, narrow spec
We present a simple unifying treatment of a broad class of applications from statistical mechanics, econometrics, mathematical finance, and insurance mathematics, where (possibly subordinated) Lévy noise arises as a scaling limit of some form of continuous-time random walk (CTRW). For each application, it is natural to rely on weak convergence results for stochastic integrals on Skorokhod space in Skorokhod's J1 or M1 topologies. As compared to earlier and entirely separate works, we are able to give a more streamlined account while also allowing for greater generality and providing important new insights. For each application, we first elucidate how the fundamental conclusions for J1 convergent CTRWs emerge as special cases of the same general principles, and we then illustrate how the specific settings give rise to different results for strictly M1 convergent CTRWs.
This paper synthesizes recent advances in the econometrics of difference-in-differences (DiD) and provides concrete recommendations for practitioners. We begin by articulating a simple set of ``canonical'' assumptions under which the econometrics of DiD are well-understood. We then argue that recent advances in DiD methods can be broadly classified as relaxing some components of the canonical DiD setup, with a focus on $(i)$ multiple periods and variation in treatment timing, $(ii)$ potential violations of parallel trends, or $(iii)$ alternative frameworks for inference. Our discussion highlights the different ways that the DiD literature has advanced beyond the canonical model, and helps to clarify when each of the papers will be relevant for empirical work. We conclude by discussing some promising areas for future research.
A supervised machine learning algorithm determines a model from a learning sample that will be used to predict new observations. To this end, it aggregates individual characteristics of the observations of the learning sample. But this information aggregation does not consider any potential selection on unobservables and any status-quo biases which may be contained in the training sample. The latter bias has raised concerns around the so-called \textit{fairness} of machine learning algorithms, especially towards disadvantaged groups. In this chapter, we review the issue of fairness in machine learning through the lenses of structural econometrics models in which the unknown index is the solution of a functional equation and issues of endogeneity are explicitly accounted for. We model fairness as a linear operator whose null space contains the set of strictly {\it fair} indexes. A {\it fair} solution is obtained by projecting the unconstrained index into the null space of this operator or by directly finding the closest solution of the functional equation into this null space. We also acknowledge that policymakers may incur a cost when moving away from the status quo. Achieving \tex
Integrated Nested Laplace Approximation provides a fast and effective method for marginal inference on Bayesian hierarchical models. This methodology has been implemented in the R-INLA package which permits INLA to be used from within R statistical software. Although INLA is implemented as a general methodology, its use in practice is limited to the models implemented in the R-INLA package. Spatial autoregressive models are widely used in spatial econometrics but have until now been missing from the R-INLA package. In this paper, we describe the implementation and application of a new class of latent models in INLA made available through R-INLA. This new latent class implements a standard spatial lag model, which is widely used and that can be used to build more complex models in spatial econometrics. The implementation of this latent model in R-INLA also means that all the other features of INLA can be used for model fitting, model selection and inference in spatial econometrics, as will be shown in this paper. Finally, we will illustrate the use of this new latent model and its applications with two datasets based on Gaussian and binary outcomes.
Causal machine learning (ML) recovers graphical structures that inform us about potential cause-and-effect relationships. Most progress has focused on cross-sectional data with no explicit time order, whereas recovering causal structures from time series data remains the subject of ongoing research in causal ML. In addition to traditional causal ML, this study assesses econometric methods that some argue can recover causal structures from time series data. The use of these methods can be explained by the significant attention the field of econometrics has given to causality, and specifically to time series, over the years. This presents the possibility of comparing the causal discovery performance between econometric and traditional causal ML algorithms. We seek to understand if there are lessons to be incorporated into causal ML from econometrics, and provide code to translate the results of these econometric methods to the most widely used Bayesian Network R library, bnlearn. We investigate the benefits and challenges that these algorithms present in supporting policy decision-making, using the real-world case of COVID-19 in the UK as an example. Four econometric methods are eval
The Journal Impact Factor (IF), as a core indicator of academic evaluation, has not been systematically studied in relation to its historical evolution and global macroeconomic environment. This paper employs a period-based regression analysis using long-term time series data from 1975-2026 to examine the statistical relationship between IF and Federal Reserve monetary policy (using real interest rate as a proxy variable). The study estimates three nested models using Ordinary Least Squares (OLS): (1) a baseline linear model, (2) a linear model controlling for time trends, and (3) a log-transformed model. Empirical results show that: (i) in the early period (1975-2000), there is no significant statistical relationship between IF and real interest rate ($p>0.1$); (ii) during the quantitative easing period (2001-2020), they exhibit a significant negative correlation ($β=-0.069$, $p<0.01$), meaning that for every 1 percentage point decrease in real interest rate, IF increases by approximately 6.9\%; (iii) the adjusted $R^2$ of the full-sample model reaches 0.893, indicating that real interest rate and time trends can explain 89.3\% of IF variation. This finding reveals the indir
Econometric inference usually conditions on a dependence structure chosen in advance, even though the data may support clustering, latent factors, sparse interactions, or mixtures of these mechanisms. This paper studies the prior problem of learning the dependence structure that is relevant for inference. We represent candidate structures as covariance geometries in a common Hilbert space and project an estimable dependence operator onto them. The resulting geometric dependence profile is a low-dimensional diagnostic of their relative empirical support; an off-diagonal companion profile isolates cross-sectional dependence and drives procedure selection. We establish well-definedness, consistency, asymptotic normality, and finite-sample classification bounds under local projection regularity and geometric separation, and show that tangent-space overlap creates a first-order impossibility region in which competing geometries cannot be reliably distinguished. Formulating inference-procedure choice as a statistical decision problem, we prove that when one off-diagonal geometry is uniquely separated and profile rankings are compatible with inferential loss, profile-guided inference is a
fixest is an R package for fast and flexible econometric estimation. It provides a unified framework for applied research, with comprehensive support for a diverse class of models: ordinary least squares, instrumental variables, generalized linear models, maximum likelihood, and difference-in-differences. The package particularly excels at fixed-effects estimation, supported by a novel fixed-point acceleration algorithm implemented in C++. This algorithm achieves rapid convergence across a variety of data contexts and enables efficient estimation of complex models, including those with varying slopes. An expressive formula interface facilitates multiple estimations, stepwise regressions, and variable interpolation in a single call. Users can adjust inference strategies on the fly, choosing from an array of built-in robust standard errors. The package also provides methods for publication-ready regression tables and coefficient plots. Benchmarks demonstrate that fixest offers best-in-class performance against leading alternatives in R, PYTHON, and JULIA.
We ask whether pretrained time series foundation models (TSFMs) improve on established econometric benchmarks for forecasting realized volatility. Using the VOLARE dataset, we conduct the first systematic comparison of nine zero-shot TSFMs against eight econometric specifications, including the Heterogeneous Autoregressive (HAR) family, across 50 assets in equities, foreign exchange, and futures, and three forecast horizons, with formal pairwise and multi-model forecast-comparison tests. Foundation models do not deliver a uniform gain. Pooled losses favor them, but the advantage is concentrated in a few outlier assets; averaging each asset's loss ratio to a well-specified Log-HAR benchmark, so that no single asset dominates, only one small model, Tiny Time Mixers (TTM), beats the benchmark at every horizon, and by a narrow margin. The other foundation models do not improve on Log-HAR, and the econometric benchmarks remain competitive throughout. A Mincer--Zarnowitz recalibration, which removes level and scale bias from every forecast, shows that much of the short-horizon advantage reflects better-scaled forecasts rather than better prediction of volatility dynamics, and only at the
Empirical researchers increasingly use upstream machine-learning (ML) methods to construct proxies for latent target variables from complex, unstructured data. A naive plug-in use of such proxies in downstream econometric models, however, can lead to biased estimation and invalid inference. This paper develops a framework for partial identification and inference in general moment models with ML-generated proxies. Our approach does not require restrictive assumptions on the upstream ML procedure, such as consistency or known convergence rates, nor does it require a complete validation sample containing all variables used in the downstream analysis. Instead, we assume access to two datasets: a downstream sample containing observed covariates and the proxy, and an auxiliary validation sample containing joint observations on the proxy and its target variable. We treat the proxy as a linking variable between these two samples, rather than as a literal noisy substitute for the latent target variable. Building on this idea, we develop a sharp identification strategy based on an unconditional optimal transport characterization and an inference procedure that controls asymptotic size using
This study examines the dynamic relationship between the global oil prices and Nepal Stock Exchange (NEPSE) using an integrated approach which combines traditional econometric techniques with machine learning and explainable AI techniques. For this, Daily data of International Oil prices and NEPSE index is analyzed from approximately thirteen years (June 2013 to June 2026) using Granger causality, EGARCH(1,1), and DCC-GARCH models to examine different properties like predictive relationships, asymmetric volatility behaviour, and time-varying correlations. To further supplement the econometric analysis, Machine Learning Models like Random Forest, LightGBM, and XGBoost algorithms were used to capture nonlinear relationships, along with explainable artificial intelligence techniques like SHAP values, Partial Dependence Plots, and Individual Conditional Expectation plots to further interpret the results of the model. The results from the econometric analysis showed a statistically significant unidirectional Granger causality from Brent crude oil to NEPSE with a four-day lag, high volatility persistence in both markets, and weak yet highly time-varying conditional correlations. Among th