We further develop a trace-anomaly-motivated gluonic scenario in which cold dark matter (CDM) is modeled as a long-lived color-singlet Bose-Einstein condensate seeded at the QCD confinement transition. Specifically, guided by the near-universality of the galactic acceleration scale $g^{}_\dagger \simeq (1\text{--}2)\times10^{-10}\,\mathrm{m\,s^{-2}}$ inferred from the RAR, we hypothesize that this relic gluonic condensate organizes, at galactic distances, into a spectrally rigid lowest-weight structure characterized by a protected infrared gap. Within an effective representation-theoretic framework, Lorentz covariance together with positive-energy lowest-weight unitarity naturally favors the AdS algebra $\mathfrak{so}(2,3)$ as the minimal algebraic structure supporting such a discrete tower of states. The resulting condensate produces, by construction, a cored halo profile with finite total mass $M_h$. Normalizing this profile by the observed approximate universality of the central DM surface density, $Σ_0\simeq141\,M_\odot\,\mathrm{pc}^{-2}$, yields a universal characteristic acceleration $g^{}_\star ={G\, M_h}/{r_{\rm c}^2} \simeq π^2 G\,Σ_0 \simeq 1.9\times10^{-10}\,\mathrm{m\,s
Large-scale autoregressive models have demonstrated remarkable capabilities in image generation. However, their sequential raster-scan decoding relies on strictly next-token prediction, making inference prohibitively expensive. Existing acceleration methods typically either introduce entirely new generation paradigms that necessitate costly pre-training from scratch, or enable parallel generation at the expense of a training-inference gap or altered prediction objectives. In this paper, we introduce FlashAR, a lightweight post-training adaptation framework that efficiently adapts a pre-trained raster-scan autoregressive model into a highly parallel generator based on two-way next-token prediction. Our key insight is that effective adaptation should minimize modifications to the pre-trained model's original training objective to preserve its learned prior. Accordingly, we retain the original AR head as a horizontal head for row-wise prediction and introduce a complementary, lightweight vertical head for column-wise prediction. To facilitate efficient adaptation, we branch the vertical head from an intermediate layer rather than the final layer, bypassing the inherent horizontal head
Cold dark matter halos are expected to be triaxial and often tilted relative to the stellar disk. Stellar streams provide a sensitive tracer of the Milky Way's halo shape, though models for the Galactic potential are typically limited to simple, symmetric functional forms. Here, we measure the Galactic acceleration field along the GD-1 stellar stream using a direct differentiation of the stream's track in phase-space. Using a fully data-driven catalog of stream members from Gaia, SDSS, LAMOST, and DESI, we map the stream in 6D phase-space. We fit splines to the stream track, and infer cylindrical acceleration components $a_R = -2.5 \pm_{0.1}^{0.2}, \ a_z = -1.8\pm 0.1, \ a_φ= 0.2\pm 0.1~\rm{km \ s^{-1} \ Myr^{-1}}$ at $(R,z,φ) = (11.9~\rm{kpc}, 7.3~\rm{kpc}, 171.1~\rm{deg})$. We measure mass enclosed within $14~\rm{kpc}$ of $1.4\pm 0.1 \times 10^{11} M_\odot$ and z-axis density flattening of $q_{ρ, z} = 0.81\pm^{0.06}_{0.03}$, both consistent with previous estimates. However, we find a 2$σ$ deviation from an axisymmetric acceleration field, which can be explained by a triaxial dark matter halo with axis ratios 1:0.75:0.70. The major axis of the halo is consistent with a tilt of $18
This study introduces the Iterative Chainlet Partitioning (ICP) algorithm and its neural acceleration for solving the Traveling Salesman Problem with Drone (TSP-D). The proposed ICP algorithm decomposes a TSP-D solution into smaller segments called chainlets, each optimized individually by a dynamic programming subroutine. The chainlet with the highest improvement is updated, and the procedure is repeated until no further improvement is possible. We show that the subroutine runs in quadratic time and the number of subroutine calls is bounded linearly in problem size for the first iteration and remains constant in subsequent iterations, ensuring algorithmic scalability. Empirical results show that ICP outperforms existing algorithms in both solution quality and computational time. Tested over 1,249 benchmark instances, ICP yields an average improvement of 2.6\% in solution quality over the previous state-of-the-art algorithm while reducing computational time by 91.3\%. The procedure is deterministic, ensuring reliability without requiring multiple runs. The subroutine is the computational bottleneck in the already efficient ICP algorithm. To reduce the necessity of subroutine calls,
The subject of this paper is stochastic acceleration by plasma turbulence, a process akin to the original model proposed by Fermi. We review the relative merits of different acceleration models, in particular the so called first order Fermi acceleration by shocks and second order Fermi by stochastic processes, and point out that plasma waves or turbulence play an important role in all mechanisms of acceleration. Thus, stochastic acceleration by turbulence is active in most situations. We also show that it is the most efficient mechanism of acceleration of relatively cool non relativistic thermal background plasma particles. In addition, it can preferentially accelerate electrons relative to protons as is needed in many astrophysical radiating sources, where usually there are no indications of presence of shocks. We also point out that a hybrid acceleration mechanism consisting of initial acceleration by turbulence of background particles followed by a second stage acceleration by a shock has many attractive features. It is demonstrated that the above scenarios can account for many signatures of the accelerated electrons, protons and other ions, in particular $^3$He and $^4$He, seen
In this paper we develop the first fine-grained rounding error analysis of finite element (FE) cell kernels and assembly. The theory includes mixed-precision implementations and accounts for hardware-acceleration via matrix multiplication units, thus providing theoretical guidance for designing reduced- and mixed-precision FE algorithms on CPUs and GPUs. Guided by this analysis, we introduce hardware-accelerated mixed-precision implementation strategies which are provably robust to low-precision computations. Indeed, these algorithms are accurate to the lower-precision unit roundoff with an error constant that is independent from: the conditioning of FE basis function evaluations, the ill-posedness of the cell, the polynomial degree, and the number of quadrature nodes. Consequently, we present the first AMX-accelerated FE kernel implementations on Intel Sapphire Rapids CPUs. Numerical experiments demonstrate that the proposed mixed- (single/half-) precision algorithms are up to 60 times faster than their double precision equivalent while being orders of magnitude more accurate than their fully half-precision counterparts.
The transformer architecture predominates across various models. As the heart of the transformer, attention has a computational complexity of $O(N^2)$, compared to $O(N)$ for linear transformations. When handling large sequence lengths, attention becomes the primary time-consuming component. Although quantization has proven to be an effective method for accelerating model inference, existing quantization methods primarily focus on optimizing the linear layer. In response, we first analyze the feasibility of quantization in attention detailedly. Following that, we propose SageAttention, a highly efficient and accurate quantization method for attention. The OPS (operations per second) of our approach outperforms FlashAttention2 and xformers by about 2.1 times and 2.7 times, respectively. SageAttention also achieves superior accuracy performance over FlashAttention3. Comprehensive experiments confirm that our approach incurs almost no end-to-end metrics loss across diverse models, including those for large language processing, image generation, and video generation. The codes are available at https://github.com/thu-ml/SageAttention.
Shear flows are ubiquitously present in space and astrophysical plasmas. This paper highlights the central idea of the non-thermal acceleration of charged particles in shearing flows and reviews some of the recent developments. Topics include the acceleration of charged particles by microscopic instabilities in collisionless relativistic shear flows, Fermi-type particle acceleration in macroscopic, gradual and non-gradual shear flows, as well as shear particle acceleration by large-scale velocity turbulence. When put in the context of jetted astrophysical sources such as Active Galactic Nuclei, the results illustrate a variety of means beyond conventional diffusive shock acceleration by which power-law like particle distributions might be generated. This suggests that relativistic shear flows can account for efficient in-situ acceleration of energetic electrons and be of relevance for the production of extreme cosmic rays.
In relativistic dynamics, force and acceleration are no longer parallel. In this article, we revisit the relativistic motion of a particle under the action of a constant force, $\boldsymbol{f}$. \ For a two-dimensional motion, the final velocity in each axis is $(f_{i}/f)c,$ independently of the initial velocities, yielding an asymptotic velocity always parallel to the force. The particular case in which the force is applied in a single axis is analyzed in detail, with the behavior of velocity and acceleration being exhibited for several configurations. Some previous results of the literature concerning velocity and acceleration behavior are improved and better explored. Differently from which was previously claimed, it is shown that a negative acceleration component can exist in the direction of the biggest force component and that acceleration does not decrease monotonically to zero.
The dynamic processes of magnetic reconnection and turbulence cause magnetic islands/flux-ropes generation. The in-situ observations suggest that the coalescence or/and contraction of magnetic islands are responsible to the charged particle acceleration (keV to MeV energy range). Numerical simulations also support this acceleration mechanism. However, the most fundamental question raise here is, does this mechanism contribute to the cosmic rays acceleration? To answer this, we report, in-situ evidence of flux-ropes formation, their magnetic re-connection and its manifestation as cosmic ray (GeV charged particle) acceleration in interplanetary counterpart of coronal mass ejection(ICME). Further, we propose that cosmic ray (high and/or ultra-high energy) acceleration by Fermi mechanism is valid not only through stochastic reflections of particles from the shock boundaries but also through the boundaries of contracting magnetic islands or/and their merging via magnetic re-connection. This has significant implications on cosmic ray origin and their acceleration process.
From the Eddington-Weinberg relationship, which may be explained by the holographic principle and the cosmic coincidence in a flat Universe, it follows that the characteristic gravitational acceleration aN associated with the nucleon and its Compton wavelength is of order the Hubble acceleration H0c in this epoch. A natural scaling for the cosmological constant is obtained from this acceleration term. It also happens that the critical acceleration a0 associated with the MOND theory is of order aN.
Nonlinear acceleration methods are powerful techniques to speed up fixed-point iterations. However, many acceleration methods require storing a large number of previous iterates and this can become impractical if computational resources are limited. In this paper, we propose a nonlinear Truncated Generalized Conjugate Residual method (nlTGCR) whose goal is to exploit the symmetry of the Hessian to reduce memory usage. The proposed method can be interpreted as either an inexact Newton or a quasi-Newton method. We show that, with the help of global strategies like residual check techniques, nlTGCR can converge globally for general nonlinear problems and that under mild conditions, nlTGCR is able to achieve superlinear convergence. We further analyze the convergence of nlTGCR in a stochastic setting. Numerical results demonstrate the superiority of nlTGCR when compared with several other competitive baseline approaches on a few problems. Our code will be available in the future.
The mechanism(s) driving the early- and late-time accelerated expansion of the Universe represent one of the most compelling mysteries in fundamental physics today. The path to understanding the causes of early- and late-time acceleration depends on fully leveraging ongoing surveys, developing and demonstrating new technologies, and constructing and operating new instruments. This report presents a multi-faceted vision for the cosmic survey program in the 2030s and beyond that derives from these considerations. Cosmic surveys address a wide range of fundamental physics questions, and are thus a unique and powerful component of the HEP experimental portfolio.
Cosmic-ray production in young supernova remnant (SNR) shocks is expected to be efficient and strongly nonlinear. In nonlinear, diffusive shock acceleration, compression ratios will be higher and the shocked temperature lower than test-particle, Rankine-Hugoniot relations predict. Furthermore, the heating of the gas to X-ray emitting temperatures is strongly coupled to the acceleration of cosmic-ray electrons and ions, thus nonlinear processes which modify the shock, influence the emission over the entire band from radio to gamma-rays and may have a strong impact on X-ray line models. Here we apply an algebraic model of nonlinear acceleration, combined with SNR evolution, to model the radio and X-ray continuum of Kepler's SNR.
We present a simple and efficient acceleration technique for an arbitrary method for computing the Euclidean projection of a point onto a convex polytope, defined as the convex hull of a finite number of points, in the case when the number of points in the polytope is much greater than the dimension of the space. The technique consists in applying any given method to a "small" subpolytope of the original polytope and gradually shifting it, till the projection of the given point onto the subpolytope coincides with its projection onto the original polytope. The results of numerical experiments demonstrate the high efficiency of the proposed acceleration technique. In particular, they show that the reduction of computation time increases with an increase of the number of points in the polytope and is proportional to this number for some methods. In the second part of the paper, we also discuss a straightforward extension of the proposed acceleration technique to the case of arbitrary methods for computing the distance between two convex polytopes, defined as the convex hulls of finite sets of points.
The strong variability of magnetic central engines of AGN and GRBs may result in highly intermittent strongly magnetized relativistic outflows. We find a new magnetic acceleration mechanism for such impulsive flows that can be much more effective than the acceleration of steady-state flows. This impulsive acceleration results in kinetic-energy-dominated flows at astrophysically relevant distances from the central source. For a spherical flow, a discrete shell ejected from the source over a time t_0 with Lorentz factor Gamma~1 and initial magnetization sigma_0 = B_0^2/(4 pi rho_0 c^2) >> 1 quickly reaches a typical Lorentz factor Gamma ~ sigma_0^{1/3} and magnetization sigma ~ sigma_0^{2/3} at the distance R_0 ~ ct_0. At this point the magnetized shell of width Delta ~ R_0 in the lab frame loses causal contact with the source and continues to accelerate by spreading significantly in its own rest frame. The expansion is driven by the magnetic pressure gradient and leads to relativistic relative velocities between the front and back of the shell. While the expansion is roughly symmetric in the center of momentum frame, in the lab frame most of the energy and momentum remain in a
The Advanced Proton Driven Plasma Wakefield Acceleration Experiment (AWAKE) aims at studying plasma wakefield generation and electron acceleration driven by proton bunches. It is a proof-of-principle R&D experiment at CERN and the world's first proton driven plasma wakefield acceleration experiment. The AWAKE experiment will be installed in the former CNGS facility and uses the 400 GeV/c proton beam bunches from the SPS. The first experiments will focus on the self-modulation instability of the long (r.m.s ~12 cm) proton bunch in the plasma. These experiments are planned for the end of 2016. Later, in 2017/2018, low energy (~15 MeV) electrons will be externally injected to sample the wakefields and be accelerated beyond 1 GeV.
Cosmological observations in the new millennium have dramatically increased our understanding of the Universe, but several fundamental questions remain unanswered. This topical group report describes the best opportunities to address these questions over the coming decades by extending observations to the $z<6$ universe. The greatest opportunity to revolutionize our understanding of cosmic acceleration both in the modern universe and the inflationary epoch would be provided by a new Stage V Spectroscopic Facility (Spec-S5) which would combine a large telescope aperture, wide field of view, and high multiplexing. Such a facility could simultaneously provide a dense sample of galaxies at lower redshifts to provide robust measurements of the growth of structure at small scales, as well as a sample at redshifts $2<z<5$ to measure cosmic structure at the largest scales, spanning a sufficient volume to probe primordial non-Gaussianity from inflation, to search for features in the inflationary power spectrum on a broad range of scales, to test dark energy models in poorly-explored regimes, and to determine the total neutrino mass and effective number of light relics. A number of
Recently, Cardoso and Natario constructed an exact solution of Einstein-scalar field equations that describes a scalar counterpart of the Schwarzschild-Melvin Universe. In fact, this solution belongs to a more general class of solutions described by Herdeiro in [7]. In this work we show how to further generalize these solutions in presence of acceleration, rotation and various charges. More specifically, we describe general accelerating charged rotating black holes with NUT charge in scalar multipolar universes and present some of their properties.
Diffusion transformers have gained significant attention in recent years for their ability to generate high-quality images and videos, yet still suffer from a huge computational cost due to their iterative denoising process. Recently, feature caching has been introduced to accelerate diffusion transformers by caching the feature computation in previous timesteps and reusing it in the following timesteps, which leverage the temporal similarity of diffusion models while ignoring the similarity in the spatial dimension. In this paper, we introduce Cluster-Driven Feature Caching (ClusCa) as an orthogonal and complementary perspective for previous feature caching. Specifically, ClusCa performs spatial clustering on tokens in each timestep, computes only one token in each cluster and propagates their information to all the other tokens, which is able to reduce the number of tokens by over 90%. Extensive experiments on DiT, FLUX and HunyuanVideo demonstrate its effectiveness in both text-to-image and text-to-video generation. Besides, it can be directly applied to any diffusion transformer without requirements for training. For instance, ClusCa achieves 4.96x acceleration on FLUX with an