Diffusion models have shown remarkable success across generative tasks, yet their high computational demands challenge deployment on resource-limited platforms. This paper investigates a critical question for compute-optimal diffusion model deployment: Under a post-training setting without fine-tuning, is it more effective to reduce the number of denoising steps or to use a cheaper per-step inference? Intuitively, reducing the number of denoising steps increases the variability of the distributions across steps, making the model more sensitive to compression. In contrast, keeping more denoising steps makes the differences smaller, preserving redundancy, and making post-training compression more feasible. To systematically examine this, we propose PostDiff, a training-free framework for accelerating pre-trained diffusion models by reducing redundancy at both the input level and module level in a post-training manner. At the input level, we propose a mixed-resolution denoising scheme based on the insight that reducing generation resolution in early denoising steps can enhance low-frequency components and improve final generation fidelity. At the module level, we employ a hybrid modul
According to the "hard-steps" model, the origin of humanity required "successful passage through a number of intermediate steps" (so-called "hard" or "critical" steps) that were intrinsically improbable with respect to the total time available for biological evolution on Earth. This model similarly predicts that technological life analogous to human life on Earth is "exceedingly rare" in the universe. Here, we critically reevaluate the core assumptions of the hard-steps model in light of recent advances in the Earth and life sciences. Specifically, we advance a potential alternative model where there are no hard steps, and evolutionary novelties (or singularities) required for human origins can be explained via mechanisms outside of intrinsic improbability. Furthermore, if Earth's surface environment was initially inhospitable not only to human life, but also to certain key intermediate steps in human evolution (e.g., the origin of eukaryotic cells, multicellular animals), then the "delay" in the appearance of humans can be best explained through the sequential opening of new global environmental windows of habitability over Earth history, with humanity arising relatively quickly o
The enumeration of walks in the quarter plane confined in the first quadrant has attracted a lot of attention over the past fifteenth years. The generating functions associated to small steps models satisfy a functional equation in two catalytic variables. For such models, Bousquet-Mélou and Mishna defined a group called the group of the walk which turned out to be central in the classification of small steps models. In particular, its action on the catalytic variables yields a set of change of variables compatible with the structure of the functional equation. This particular set called the orbit has been generalized to models with arbitrary large steps by Bostan, Bousquet-Mélou and Melczer. However, the orbit had till now no underlying group. In this article, we endow the orbit with the action of a Galois group, which extends the group of the walk to models with large steps. Within this Galoisian framework, we generalized the notions of invariants and decoupling. This enable us to develop a general strategy to prove the algebraicity of models with small backward steps. Our constructions lead to the first proofs of algebraicity of weighted models with large steps, proving in parti
Diffusion probabilistic models (DPMs) have shown remarkable performance in high-resolution image synthesis, but their sampling efficiency is still to be desired due to the typically large number of sampling steps. Recent advancements in high-order numerical ODE solvers for DPMs have enabled the generation of high-quality images with much fewer sampling steps. While this is a significant development, most sampling methods still employ uniform time steps, which is not optimal when using a small number of steps. To address this issue, we propose a general framework for designing an optimization problem that seeks more appropriate time steps for a specific numerical ODE solver for DPMs. This optimization problem aims to minimize the distance between the ground-truth solution to the ODE and an approximate solution corresponding to the numerical solver. It can be efficiently solved using the constrained trust region method, taking less than $15$ seconds. Our extensive experiments on both unconditional and conditional sampling using pixel- and latent-space DPMs demonstrate that, when combined with the state-of-the-art sampling method UniPC, our optimized time steps significantly improve i
We address the problem of extracting key steps from unlabeled procedural videos, motivated by the potential of Augmented Reality (AR) headsets to revolutionize job training and performance. We decompose the problem into two steps: representation learning and key steps extraction. We propose a training objective, Bootstrapped Multi-Cue Contrastive (BMC2) loss to learn discriminative representations for various steps without any labels. Different from prior works, we develop techniques to train a light-weight temporal module which uses off-the-shelf features for self supervision. Our approach can seamlessly leverage information from multiple cues like optical flow, depth or gaze to learn discriminative features for key-steps, making it amenable for AR applications. We finally extract key steps via a tunable algorithm that clusters the representations and samples. We show significant improvements over prior works for the task of key step localization and phase classification. Qualitative results demonstrate that the extracted key steps are meaningful and succinctly represent various steps of the procedural tasks.
The unique resonance and locking phenomena in the superconductor-ferromagnet-superconductor $\varphi_0$ Josephson junction under external electromagnetic radiation are demonstrated when not just the electric but also the magnetic component of external radiation is taken into account. Due to the coupling of superconductivity and magnetism in this system, the magnetic moment precession of the ferromagnetic layer caused by the magnetic component of external radiation can lock the Josephson oscillations, which results in the appearance of a particular type of steps in the current-voltage characteristics, completely different from the well-known Shapiro steps. We call these steps the Buzdin steps in the case when the system is driven only by the magnetic component and the Chimera steps in the case when both magnetic and electric components are present. Unlike the Shapiro steps where the magnetization remains constant along the step, here it changes though the system is locked. The spin-orbit coupling substantially contributes to the amplitude, i.e., the size of these steps. Dramatic changes in their amplitudes are also observed at frequencies near the ferromagnetic resonance. Combinatio
The current-voltage characteristic of a driven superconducting Josephson junction displays discrete steps. This phenomenon, called the Shapiro steps, forms today's voltage standard! Here, we report the observation of Shapiro steps in a driven Josephson junction in a gas of ultracold atoms. We demonstrate that the steps exhibit universal features, and provide insight into the microscopic dissipative dynamics that we directly observe in the experiment. We find that the steps are directly connected to phonon emission and nucleation of solitonic excitations, whose dynamics we follow in space and time. The experimental results are underpinned by extensive numerical simulations based on classical-field dynamics and may enable metrological and fundamental advances.
The Natural Language Inference (NLI) task often requires reasoning over multiple steps to reach the conclusion. While the necessity of generating such intermediate steps (instead of a summary explanation) has gained popular support, it is unclear how to generate such steps without complete end-to-end supervision and how such generated steps can be further utilized. In this work, we train a sequence-to-sequence model to generate only the next step given an NLI premise and hypothesis pair (and previous steps); then enhance it with external knowledge and symbolic search to generate intermediate steps with only next-step supervision. We show the correctness of such generated steps through automated and human verification. Furthermore, we show that such generated steps can help improve end-to-end NLI task performance using simple data augmentation strategies, across multiple public NLI datasets.
In this paper, we introduce an innovative NLP model specifically fine-tuned to determine the minimal number of denoising steps required for any given text prompt. This advanced model serves as a real-time tool that recommends the ideal denoise steps for generating high-quality images efficiently. It is designed to work seamlessly with the Diffusion model, ensuring that images are produced with superior quality in the shortest possible time. Although our explanation focuses on the DDIM scheduler, the methodology is adaptable and can be applied to various other schedulers like Euler, Euler Ancestral, Heun, DPM2 Karras, UniPC, and more. This model allows our customers to conserve costly computing resources by executing the fewest necessary denoising steps to achieve optimal quality in the produced images.
Using a Monte Carlo method on a lattice model of a vicinal surface with short-range step-step attraction, we show that, at low temperature and near equilibrium, there is an inhibition of the motion of macro-steps. This inhibition leads to a pinning of steps without defects, adsorbates, or impurities (self-pinning of steps). We show that this inhibition of the macro-step motion is caused by faceted steps, which are macro-steps that have a smooth side surface. The faceted steps result from discontinuities in the anisotropic surface tension (the surface free energy per area). The discontinuities are brought into the surface tension by the short-range step-step attraction. The short-range step-step attraction also originates `step-droplets', which are locally merged steps, at higher temperatures. We derive an analytic equation of the surface stiffness tensor for the vicinal surface around the (001) surface. Using the surface stiffness tensor, we show that step-droplets roughen the vicinal surface. Contrary to what we expected, the step-droplets slow down the step velocity due to the diminishment of kinks in the merged steps (smoothing of the merged steps).
Giant steps is a technique to accelerate Monte Carlo radiative transfer in optically-thick cells of astrophysical atmospheres by greatly reducing the number of Monte Carlo steps needed to propagate photon packets through such cells. Giant steps replaces the exact diffusion treatment of ordinary Monte Carlo radiative transfer in the cells by an approximate diffusion treatment. In this paper, we describe the basic idea of giant steps and report demonstration giant-steps flux calculations for the grey atmosphere. Speed-up factors of order 100 are obtained relative to ordinary Monte Carlo radiative transfer. In practical applications, speed-up factors of order ten (Mazzali et al. 2001) and perhaps more are possible. The speed-up factor is likely to be significantly application-dependent and there is a trade-off between speed-up and accuracy. This paper and past work (Mazzali et al. 2001) suggest that giant-steps error can probably be kept to a few percent by using sufficiently large boundary-layer optical depths while still maintaining large speed-up factors. Thus, giant steps can be characterized as a moderate accuracy radiative transfer technique.
The paper is devoted to the study of lattice paths that consist of vertical steps $(0,-1)$ and non-vertical steps $(1,k)$ for some $k\in \mathbb Z$. Two special families of primary and free lattice paths with vertical steps are considered. It is shown that for any family of primary paths there are equinumerous families of proper weighted lattice paths that consist of only non-vertical steps. The relation between primary and free paths is established and some combinatorial and statistical properties are obtained. It is shown that the expected number of vertical steps in a primary path running from $(0,0)$ to $(n,-1)$ is equal to the number of free paths running from $(0,0)$ to $(n,0)$. Enumerative results with generating functions are given. Finally, a few examples of families of paths with vertical steps are presented and related to Łukasiewicz, Motzkin, Dyck and Delannoy paths.
Sharpness-aware minimization (SAM) methods have gained increasing popularity by formulating the problem of minimizing both loss value and loss sharpness as a minimax objective. In this work, we increase the efficiency of the maximization and minimization parts of SAM's objective to achieve a better loss-sharpness trade-off. By taking inspiration from the Lookahead optimizer, which uses multiple descent steps ahead, we propose Lookbehind, which performs multiple ascent steps behind to enhance the maximization step of SAM and find a worst-case perturbation with higher loss. Then, to mitigate the variance in the descent step arising from the gathered gradients across the multiple ascent steps, we employ linear interpolation to refine the minimization step. Lookbehind leads to a myriad of benefits across a variety of tasks. Particularly, we show increased generalization performance, greater robustness against noisy weights, as well as improved learning and less catastrophic forgetting in lifelong learning settings. Our code is available at https://github.com/chandar-lab/Lookbehind-SAM.
Images of the morphology of GaN (0001) surfaces often show half-unit-cell-height steps separating a sequence of terraces having alternating large and small widths. This can be explained by the $αβαβ$ stacking sequence of the wurtzite crystal structure, which results in steps with alternating $A$ and $B$ edge structures for the lowest energy step azimuths, i.e. steps normal to $[0 1 \bar{1} 0]$ type directions. Predicted differences in the adatom attachment kinetics at $A$ and $B$ steps would lead to alternating $α$ and $β$ terrace widths. However, because of the difficulty of experimentally identifying which step is $A$ or $B$, it has not been possible to determine the absolute difference in their behavior, e.g. which step has higher adatom attachment rate constants. Here we show that surface X-ray scattering can measure the fraction of $α$ and $β$ terraces, and thus unambiguously differentiate the growth dynamics of $A$ and $B$ steps. We first present calculations of the intensity profiles of GaN crystal truncation rods (CTRs) that demonstrate a marked dependence on the $α$ terrace fraction $f_α$. We then present surface X-ray scattering measurements performed \textit{in situ} dur
The Fearless Steps Challenge 2019 Phase-1 (FSC-P1) is the inaugural Challenge of the Fearless Steps Initiative hosted by the Center for Robust Speech Systems (CRSS) at the University of Texas at Dallas. The goal of this Challenge is to evaluate the performance of state-of-the-art speech and language systems for large task-oriented teams with naturalistic audio in challenging environments. Researchers may select to participate in any single or multiple of these challenge tasks. Researchers may also choose to employ the FEARLESS STEPS corpus for other related speech applications. All participants are encouraged to submit their solutions and results for consideration in the ISCA INTERSPEECH-2019 special session.
We investigate phase transitions in a solid-on-solid model where double-height steps as well as single-height steps are allowed. Without the double-height steps, repulsive interactions between up-up or down-down step pairs give rise to a disordered flat phase. When the double-height steps are allowed, two single-height steps can merge into a double-height step (step doubling). We find that the step doubling reduces repulsive interaction strength between single-height steps and that the disordered flat phase is suppressed. As a control parameter a step doubling energy is introduced, which is assigned to each step doubling vertex. From transfer matrix type finite-size-scaling studies of interface free energies, we obtain the phase diagram in the parameter space of the step energy, the interaction energy, and the step doubling energy.
The field of Language Reasoning Models (LRMs) has been very active over the past few years with advances in training and inference techniques enabling LRMs to reason longer, and more accurately. However, a growing body of studies show that LRMs are still inefficient, over-generating verification and reflection steps. To address this challenge, we introduce the Step-Tagging framework, a lightweight sentence-classifier enabling real-time annotation of the type of reasoning steps that an LRM is generating. To monitor reasoning behaviors, we introduced ReasonType: a novel taxonomy of reasoning steps. Building on this framework, we demonstrated that online monitoring of the count of specific steps can produce effective interpretable early stopping criteria of LRM inferences. We evaluate the Step-tagging framework on three open-source reasoning models across standard benchmark datasets: MATH500, GSM8K, AIME and non-mathematical tasks (GPQA and MMLU-Pro). We achieve 20 to 50% token reduction while maintaining comparable accuracy to standard generation, with largest gains observed on more computation-heavy tasks. This work offers a novel way to increase control over the generation of LRMs,
Large language models (LLMs) have recently demonstrated impressive performance on complex, multi-step reasoning tasks, especially when post-trained with outcome-rewarded reinforcement learning Guo et al. 2025. However, it has been observed that outcome rewards often overlook flawed intermediate steps, leading to unreliable reasoning steps even when final answers are correct. To address this unreliable reasoning, we propose PRoSFI (Process Reward over Structured Formal Intermediates), a novel reward method that enhances reasoning reliability without compromising accuracy. Instead of generating formal proofs directly, which is rarely accomplishable for a modest-sized (7B) model, the model outputs structured intermediate steps aligned with its natural language reasoning. Each step is then verified by a formal prover. Only fully validated reasoning chains receive high rewards. The integration of formal verification guides the model towards generating step-by-step machine-checkable proofs, thereby yielding more credible final answers. PRoSFI offers a simple and effective approach to training trustworthy reasoning models.
When leveraging language models for reasoning tasks, generating explicit chain-of-thought (CoT) steps often proves essential for achieving high accuracy in final outputs. In this paper, we investigate if models can be taught to internalize these CoT steps. To this end, we propose a simple yet effective method for internalizing CoT steps: starting with a model trained for explicit CoT reasoning, we gradually remove the intermediate steps and finetune the model. This process allows the model to internalize the intermediate reasoning steps, thus simplifying the reasoning process while maintaining high performance. Our approach enables a GPT-2 Small model to solve 9-by-9 multiplication with up to 99% accuracy, whereas standard training cannot solve beyond 4-by-4 multiplication. Furthermore, our method proves effective on larger language models, such as Mistral 7B, achieving over 50% accuracy on GSM8K without producing any intermediate steps.
Reasoning is a fundamental capability for solving complex multi-step problems, particularly in visual contexts where sequential step-wise understanding is essential. Existing approaches lack a comprehensive framework for evaluating visual reasoning and do not emphasize step-wise problem-solving. To this end, we propose a comprehensive framework for advancing step-by-step visual reasoning in large language models (LMMs) through three key contributions. First, we introduce a visual reasoning benchmark specifically designed to evaluate multi-step reasoning tasks. The benchmark presents a diverse set of challenges with eight different categories ranging from complex visual perception to scientific reasoning with over 4k reasoning steps in total, enabling robust evaluation of LLMs' abilities to perform accurate and interpretable visual reasoning across multiple steps. Second, we propose a novel metric that assesses visual reasoning quality at the granularity of individual steps, emphasizing both correctness and logical coherence. The proposed metric offers deeper insights into reasoning performance compared to traditional end-task accuracy metrics. Third, we present a new multimodal vis