Alphas are pivotal in providing signals for quantitative trading. The industry highly values the discovery of formulaic alphas for their interpretability and ease of analysis, compared with the expressive yet overfitting-prone black-box alphas. In this work, we focus on discovering formulaic alphas. Prior studies on automatically generating a collection of formulaic alphas were mostly based on genetic programming (GP), which is known to suffer from the problems of being sensitive to the initial population, converting to local optima, and slow computation speed. Recent efforts employing deep reinforcement learning (DRL) for alpha discovery have not fully addressed key practical considerations such as alpha correlations and validity, which are crucial for their effectiveness. In this work, we propose a novel framework for alpha discovery using DRL by formulating the alpha discovery process as program construction. Our agent, $\text{Alpha}^2$, assembles an alpha program optimized for an evaluation metric. A search algorithm guided by DRL navigates through the search space based on value estimates for potential alpha outcomes. The evaluation metric encourages both the performance and t
Alpha factor mining is a fundamental task in quantitative trading, aimed at discovering interpretable signals that can predict asset returns beyond systematic market risk. While traditional methods rely on manual formula design or heuristic search with machine learning, recent advances have leveraged Large Language Models (LLMs) for automated factor discovery. However, existing LLM-based alpha mining approaches remain limited in terms of automation, generality, and efficiency. In this paper, we propose Chain-of-Alpha, a novel, simple, yet effective and efficient LLM-based framework for fully automated formulaic alpha mining. Our method features a dual-chain architecture, consisting of a Factor Generation Chain and a Factor Optimization Chain, which iteratively generate, evaluate, and refine candidate alpha factors using only market data, while leveraging backtest feedback and prior optimization knowledge. The two chains work synergistically to enable high-quality alpha discovery without human intervention and offer strong scalability. Extensive experiments on real-world A-share benchmarks demonstrate that Chain-of-Alpha outperforms existing baselines across multiple metrics, presen
Signal decay and regime shifts pose recurring challenges for data-driven investment strategies in non-stationary markets. Conventional time-series and machine learning approaches, which rely primarily on historical correlations, often struggle to generalize when the economic environment changes. While large language models (LLMs) offer strong capabilities for processing unstructured information, their potential to support quantitative factor screening through explicit economic reasoning remains underexplored. Existing factor-based methods typically reduce alphas to numerical time series, overlooking the semantic rationale that determines when a factor is economically relevant. We propose Alpha-R1, an 8B-parameter reasoning model trained via reinforcement learning for context-aware alpha screening. Alpha-R1 reasons over factor logic and real-time news to evaluate alpha relevance under changing market conditions, selectively activating or deactivating factors based on contextual consistency. Empirical results across multiple asset pools show that Alpha-R1 consistently outperforms benchmark strategies and exhibits improved robustness to alpha decay. The full implementation and resourc
Alpha mining, a critical component in quantitative investment, focuses on discovering predictive signals for future asset returns in increasingly complex financial markets. However, the pervasive issue of alpha decay, where factors lose their predictive power over time, poses a significant challenge for alpha mining. Traditional methods like genetic programming face rapid alpha decay from overfitting and complexity, while approaches driven by Large Language Models (LLMs), despite their promise, often rely too heavily on existing knowledge, creating homogeneous factors that worsen crowding and accelerate decay. To address this challenge, we propose AlphaAgent, an autonomous framework that effectively integrates LLM agents with ad hoc regularizations for mining decay-resistant alpha factors. AlphaAgent employs three key mechanisms: (i) originality enforcement through a similarity measure based on abstract syntax trees (ASTs) against existing alphas, (ii) hypothesis-factor alignment via LLM-evaluated semantic consistency between market hypotheses and generated factors, and (iii) complexity control via AST-based structural constraints, preventing over-engineered constructions that are
We carry out Faddeev calculations of three-alpha (3 alpha) and two-alpha plus Lambda (alpha alpha Lambda) systems, using two-cluster resonating-group method kernels. The input includes an effective two-nucleon force for the alpha alpha resonating-group method and a new effective Lambda N force for the Lambda alpha interaction. The latter force is a simple two-range Gaussian potential for each spin-singlet and triplet state, generated from the phase-shift behavior of the quark-model hyperon-nucleon interaction, fss2, by using an inversion method based on supersymmetric quantum mechanics. Owing to the exact treatment of the Pauli-forbidden states between the clusters, the present three-cluster Faddeev formalism can describe the mutually related, alpha alpha, 3 alpha and alpha alpha Lambda systems, in terms of a unique set of the baryon-baryon interactions. For the three-range Minnesota force which describes the alpha alpha phase shifts quite accurately, the ground-state and excitation energies of 9Be Lambda are reproduced within 100 - 200 keV accuracy.
Transitioning a strategy from backtest to live trading is a common failure point for quantitative systems due to parameter overfitting, selection bias, and sensitivity to regime changes. This paper presents the AlgoXpert Alpha Research Framework, a standardized protocol that evaluates strategies across three stages: In Sample (IS), which focuses on stable parameter regions instead of single optima; Walk Forward Analysis (WFA) using rolling windows and purge gaps to reduce information leakage, supported by majority pass and catastrophic veto rules; and Out of Sample (OOS) testing under strict parameter lock with no further tuning. The framework applies a defense in depth structure that includes structural safeguards such as cliff veto, execution controls such as spread and leverage guards, and equity protection mechanisms such as circuit breakers and a kill switch. A case study on USDJPY M5 intraday data demonstrates how to detect overfitting through performance decay and drawdown behavior across chronological stages. A post validation comparison of four alpha variants (v1 to v4) shows rank reversal when the objective changes from maximizing Sharpe to minimizing maximum drawdown, hi
One of the most important tasks in quantitative investment research is mining new alphas (effective trading signals or factors). Traditional alpha mining methods, either hand-crafted factor synthesizing or algorithmic factor mining (e.g., search with genetic programming), have inherent limitations, especially in implementing the ideas of quants. In this work, we propose a new alpha mining paradigm by introducing human-AI interaction, and a novel prompt engineering algorithmic framework to implement this paradigm by leveraging the power of large language models. Moreover, we develop Alpha-GPT, a new interactive alpha mining system framework that provides a heuristic way to ``understand'' the ideas of quant researchers and outputs creative, insightful, and effective alphas. We demonstrate the effectiveness and advantage of Alpha-GPT via a number of alpha mining experiments.
We reexamine Smale's alpha theory as a way to certify a numerical solution to an analytic system. For a given point and a system, Smale's alpha theory determines whether Newton's method applied to this point shows the quadratic convergence to an exact solution. We introduce the alpha theory computation using interval arithmetic to avoid costly exact arithmetic. As a straightforward variation of the alpha theory, our work improves computational efficiency compared to software employing the traditional alpha theory.
We present the first Lyman Alpha Emitter (LAE) study that combines: (i) cosmological SPH simulations run using GADGET-2, (ii) radiative transfer simulations (CRASH), and (iii) a previously developed LAE model. This complete LAE model accounts for the intrinsic LAE Lyman Alpha/continuum luminosity, dust enrichment and Lyman Alpha transmission through the intergalactic medium (IGM), to quantify the effects of reionization, dust and velocity fields on the Lyman Alpha and UV Luminosity Functions (LF). We find that a model neglecting dust sorely fails to reproduce either the slope or the magnitude of the observed Lyman Alpha and UV LFs. Clumped dust is required to simultaneously fit the observed UV and Lyman Alpha LFs, such that the intrinsic Lyman Alpha-to-continuum luminosity is enhanced by a factor f_alpha/f_c ~ 1.5 (3.7) excluding (including) peculiar velocities. The higher value including velocity fields arises since LAEs reside in large potential wells and inflows decrease their Lyman Alpha transmission. For the first time, a degeneracy is found between the the ionization state of the IGM and the clumping of dust inside high-redshift galaxies. The Lyman Alpha LF at z ~ 5.7 can be
We report on the discovery of a bright Lyman alpha blob associated with the z=3 quasar SDSSJ124020.91+145535.6 which is also coincident with strong damped Lyman alpha absorption from a foreground galaxy (a so-called proximate damped Lyman alpha system; PDLA). The one dimensional spectrum acquired by the Sloan Digital Sky Survey (SDSS) shows a broad Lyman alpha emission line with a FWHM ~ 500 km/s and a luminosity of L_{Lya} = 3.9e43 erg/s superposed on the trough of the PDLA. Mechanisms for powering this large Lyman alpha luminosity are discussed. We argue against emission from HII regions in the PDLA galaxy since this requires an excessive star-formation rate ~ 500 Msun/yr and would correspond to the largest Lyman alpha luminosity ever measured from a damped Lyman alpha system or starburst galaxy. We use a Monte Carlo radiative transfer simulation to investigate the possibility that the line emission is fluorescent recombination radiation from the PDLA galaxy powered by the ionizing flux of the quasar, but find that the predicted Lyman alpha flux is several orders of magnitude lower than observed. We conclude that the Lyman alpha emission is not associated with the PDLA galaxy at
We calculate Lambda alpha, Sigma alpha and Xi alpha potentials from the nuclear-matter G-matrices of the SU6 quark-model baryon-baryon interaction. The alpha-cluster wave function is assumed to be a simple harmonic-oscillator shell-model wave function. A new method is proposed to derive the direct and knock-on terms of the interaction Born kernel from the hyperon-nucleon G-matrices, with explicit treatments of the nonlocality and the center-of-mass motion between the hyperon and alpha. We find that the SU6 quark-model baryon-baryon interactions, FSS and fss2, yield a reasonable bound-state energy for 5 He Lambda, -3.18 -- -3.62 MeV, in spite of the fact that they give relatively large depths for the Lambda single-particle potentials, 46 -- 48 MeV, in symmetric nuclear matter. An equivalent local potential derived from the Wigner transform of the nonlocal Lambda alpha kernel shows a strong energy dependence for the incident Lambda-particle, indicating the importance of the strangeness-exchange process in the original hyperon-nucleon interaction. The Sigma alpha and Xi alpha potentials are repulsive with the attractive isospin I=1/2 (Sigma alpha) and I=0 (Xi alpha) components and the
We propose a framework for constructing factor models for alpha streams. Our motivation is threefold. 1) When the number of alphas is large, the sample covariance matrix is singular. 2) Its out-of-sample stability is challenging. 3) Optimization of investment allocation into alpha streams can be tractable for a factor model alpha covariance matrix. We discuss various risk factors for alphas such as: style risk factors; cluster risk factors based on alpha taxonomy; principal components; and also using the underlying tradables (stocks) as alpha risk factors, for which computing the factor loadings and factor covariance matrices does not involve any correlations with alphas, and their number is much larger than that of the relevant principal components. We draw insight from stock factor models, but also point out substantial differences.
Java is very powerful, but in Deep Learning field, its capabilities probably has not been sufficiently exploited. Compared to the Java-based deep-learning-frameworks, the Python-based (PyTorch, TensorFlow, etc) are undoubtedly the mainstream, due to their easy-to-use, flexibility and better ecosystem. Dragon-Alpha is a Java-based Tensor Computing Framework, with easy-to-use, high-scalability and high-performance, trying to break Java's dilemma in deep learning field and make it more effective. Dragon-Alpha supports different levels of APIs, and can be used as a deep-learning-framework through its user-friendly high-level APIs. Dragon-Alpha has potential to aggregate computing-power across heterogeneous platforms and devices, based on its multi-layer architecture and Java's big-data ecosystem. Dragon-Alpha has its asynchronized APIs to improve parallelism, and highly-optimized CUDA library cu32 which adopts unique convolution\deconvolution operators for small feature maps. The experiments show that, compared to PyTorch&cuDNN, Dragon-Alpha&cu32 costs less time and memory (75.38% to 97.32%, 29.2% to 66.4%), to train some typical neural networks (AlexNet, VGG, GoogleNet, ResNet
We study H I Lyman-alpha absorption observed by the Hubble Space Telescope toward the nearby binary system Alpha Cen (G2 V+K0 V) and its distant companion star Proxima Cen (M5.5 Ve). Absorption from heliospheric H I heated by the solar wind/ISM interaction is observed toward both Alpha Cen and Proxima Cen. Absorption from analogous "astrospheric" material surrounding the stars is detected toward Alpha Cen, but not Proxima Cen. The nondetection of astrospheric absorption toward Proxima Cen suggests that the stellar wind of Proxima Cen must be significantly weaker than that of the Alpha Cen system. We use hydrodynamic models of the astrospheres computed assuming different mass-loss rates to predict astrospheric Lyman-alpha absorption for comparison with the observations. The model that best matches the Alpha Cen data has a mass-loss rate of twice the solar rate, and the models suggest an upper limit of 0.2 solar for Proxima Cen. Finally, we note that the heliospheric absorption observed toward Proxima Cen in 2000 May is identical to the heliospheric absorption observed toward Alpha Cen in 1995 May, implying that the structure of the outer heliosphere does not change significantly dur
Highly accurate alpha blending can be performed entirely with integer operations, and no divisions. To reduce the number of integer multiplications, multiple color components can be blended in parallel in the same 32-bit or 64-bit register. This tutorial explains how to avoid division operations when alpha blending with 32-bit RGBA pixels. An RGBA pixel contains four 8-bit components (red, green, blue, and alpha) whose values range from 0 to 255. Alpha blending requires multiplication of the color components by an alpha value, after which (for greatest accuracy) each of these products is divided by 255 and then rounded to the nearest integer. This tutorial presents an approximate alpha-blending formula that replaces the division operation with an integer shift and add -- and also enables the number of multiplications to be reduced. When the same blending calculation is carried out to high precision using double-precision floating-point division operations, the results are found to exactly match those produced by this approximation. C++ code examples are included.
The solar chromosphere and transition region are highly structured and complex regimes. A recent breakthrough has been the identification of dynamic fibrils observed in H alpha as caused by field-aligned magnetoacoustic shocks. We seek to find whether such dynamic fibrils are also observed in Ly alpha. We used a brief sequence of four high-resolution Ly alpha images of the solar limb taken by the Very high Angular resolution ULtraviolet Telescope (VAULT), which displays many extending and retracting Ly alpha jets. We measured their top trajectories and fitted parabolas to the 30 best-defined ones. Most jet tops move supersonically. Half of them decelerate, sometimes superballistically, the others accelerate. This bifurcation may arise from incomplete sampling of recurrent jets. The similarities between dynamic Ly alpha jets and H alpha fibrils suggest that the magnetoacoustic shocks causing dynamic H alpha fibrils also affect dynamic Ly alpha jets.
Lyman alpha galaxies at high redshifts offer a powerful probe of both the formation of galaxies and the reionization of the intergalactic medium. Lyman alpha line emission is an efficient tool for identifying young galaxies at high redshift, because it is strong in systems with young stars and little or no dust-- properties expected in galaxies undergoing their first burst of star-formation. Lyman alpha galaxies also provide a robust test of the reionization epoch that is independent of Gunn-Peterson trough observations in quasar spectra and is better able to distinguish line center optical depths tau=5 from tau=10^5. This is because neutral gas scatters Lyman alpha photons, dramatically ``blurring'' images of Lyman alpha galaxies embedded in a neutral intergalactic medium and rendering them undetectable. We present a photometrically selected sample of z=5.7 Lyman alpha emitters derived from the Large Area Lyman Alpha survey. The presence of these low-luminosity Lyman alpha sources at z=5.7 immediately implies that the reionization redshift was > 5.7. Comparing these objects to our earlier z=4.5 sample, we find that the number of z=5.7 emitters at fixed line luminosity marginall
We give an explicit algorithm and source code for extracting expected returns for stocks from expected returns for alphas. Our algorithm altogether bypasses combining alphas with weights into "alpha combos". Simply put, we have developed a new method for trading alphas which does not involve combining them. This yields substantial cost savings as alpha combos cost hedge funds around 3% of the P&L, while alphas themselves cost around 10%. Also, the extra layer of alpha combos, which our new method avoids, adds noise and suboptimality. We also arrive at our algorithm independently by explicitly constructing alpha risk models based on position data.
We investigate the H-alpha and infrared star formation rate (SFR) diagnostics for galaxies in the Nearby Field Galaxy Survey (NFGS). For the 81 galaxies in our sample, we derive H-alpha fluxes (included here) from integrated spectra. There is a strong correlation between the ratio of far-infrared to optical luminosities L(FIR)/L(H-alpha) and the extinction E(B-V) measured with the Balmer decrement. Before reddening correction, the SFR(IR) and SFR(H-alpha) are related to each other by a power-law. Correction of the SFR(H-alpha) for extinction using the Balmer decrement and a classical reddening curve both reduces the scatter in the SFR(IR)-SFR(H-alpha) correlation and results in a much closer agreement (within ~10%) between the two SFR indicators. This SFR relationship spans 4 orders of magnitude and holds for all Hubble types with IRAS detections in the NFGS. A constant ratio between the SFR(IR) and SFR(H-alpha) for all Hubble types, including early types (S0-Sab), suggests that the IR emission in all of these objects results from a young stellar population.
We have imaged two normal, non-coronal, infrared-bright K giants, Alpha Tau and Alpha Boo, in the 1.4-mm and 2.8-mm continuum using the Berkeley Illinois Maryland Association millimeter array. These stars have been used as important absolute calibrators for several infrared infrared satellites. Our goals are: (1) to establish whether these stars radiate as simple photospheres or possess long-wavelength chromospheres; and (2) to make a connection between millimeter wave and far-infrared absolute flux calibrations. To accomplish these goals we also present Infrared Space Observatory Long Wavelength Spectrometer measurements of both these K giants. The far-infrared and millimeter continuum radiation is produced in the vicinity of the temperature minimum in Alpha Tau and Alpha Boo. We find that current photospheric models predict fluxes in reasonable agreement with those observed for wavelengths which sample the upper photosphere, namely <=125 microns in Alpha Tau and Alpha Boo. We clearly detect chromospheric radiation from both stars by 2.8mm (by 1.4mm in the case of Alpha Boo). Only additional observations can determine precisely where beyond 125 microns the purely radiative mode