Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, but it is expensive to repeat on every new strong model because the target model must generate many rollouts during training. As models scale, post-training itself becomes a bottleneck. We study a weak-to-strong alternative: run RL on a smaller model where rollouts are cheaper, then reuse what that RL run learned to improve a stronger target model. Directly distilling the post-RL weak teacher is not enough, because the teacher's final policy mixes useful RL gains with the limitations of the smaller model. We propose Direct On-Policy Distillation (Direct-OPD), which transfers the teacher's RL-induced policy shift instead. Direct-OPD compares the post-RL teacher with its own pre-RL reference and treats their log-ratio as a dense implicit reward for the student. In plain terms, the checkpoint pair tells us which actions RL made the weak model more or less likely to take, and Direct-OPD applies that signal on the stronger student's own on-policy states. This directly reuses the weak model's RL supervision signal without running sparse-reward RL on the target model. Empiric
Object insertion aims to seamlessly composite a reference object into a specified region of a background image. Recent diffusion-based methods achieve high visual quality but formulate insertion as a simple 2D inpainting task, providing no explicit control over the object's 3D pose and limiting their practical applicability. We propose DIRECT (Decomposed Injection for Reference Composition and Target-integration), a novel framework that integrates interactive pose manipulation with high-fidelity 2D image synthesis to enable pose-controllable object insertion. Our method decomposes the insertion conditions into three complementary components: appearance guidance capturing visual details from the reference object, geometry guidance derived from the user-adjusted 3D proxy, and context guidance from the target background. By injecting them through separate pathways, DIRECT avoids feature entanglement and simultaneously preserves reference appearance, follows the user-specified pose, and adapts the object to the target scene. We also introduce an automated data construction pipeline to improve the diversity and quality of training data. Experiments show that DIRECT outperforms previous
Dung's abstract argumentation frameworks model acceptability solely in terms of an attack relation, thereby conflating two conceptually distinct aspects of argumentative reasoning: direct conflict between arguments and the structural dependencies that arise from their internal composition. While this abstraction preserves extension-based semantics, it obscures how justification is grounded in subarguments and how defeats propagate through argument structure. We introduce Subargument Argumentation Frameworks (SAFs), an abstract framework in which direct attack and subargumenthood are represented as independent primitive relations. This separation makes structural dependency explicit at the representational level while leaving its semantic impact to be determined by structure-sensitive notions of defence, admissibility, and complete semantics defined within the framework. We show that projecting SAFs onto attack-only frameworks yields extension-equivalent Dung frameworks under all standard semantics, yet the projection irreversibly loses information about justificatory grounding and structural propagation. SAFs therefore provide strictly greater representational expressiveness while
This paper introduces Dr-PoGO, a method for Simultaneous Localization And Mapping (SLAM) using a 2D spinning radar. Unlike cameras or lidars that require line-of-sight, millimetre-wave radars can `see' through dust, falling snow, rain, etc. Accordingly, it is a great modality for robust perception regardless of the weather conditions. While most existing radar-based SLAM methods rely on the extraction of point clouds or features to perform ego-motion estimation, Dr-PoGO leverages direct registration techniques for odometry (DRO) and loop-closure registration. An off-the-shelf radar-focused place recognition algorithm, RaPlace, provides loop-closure candidates. As RaPlace does not provide relative transformations, Dr-PoGO introduces a coarse-to-fine registration that uses visual features and descriptors to obtain an initial guess for the direct transformation refinement. The global trajectory is optimized in a pose-graph optimization. Dr-PoGO demonstrates state-of-the-art performance over 300km of data in various real-world automotive environments. Our implementation is publicly available: https://github.com/utiasASRL/dr_pogo.
The goal of this review article series is to provide a comprehensive overview of galactic star formation and quenching from both an observational and theoretical perspective. Drawing on a vast quantity of literature, we attempt to answer a deceptively simple question: why do galaxies cease forming stars? In Part II, we concentrate primarily on results from theory and simulations, in addition to direct observational tests of theoretically proposed quenching routes. Over the past two decades, N-body simulations, semi-analytic models, idealized hydrodynamical simulations, cosmological hydrodynamical simulations, and zoom-in hydrodynamical simulations have all provided crucial insights into the fundamental causes and specific mechanisms of galaxy quenching. Throughout this part of the review, we discuss intrinsic routes to massive galaxy quenching from strong baryonic feedback - including supernovae and active galactic nuclei (in both the ejective and preventative modes). Additionally, we discuss dynamical stabilization and the role of mergers. We go on to consider environmental routes to quenching via both ram pressure and dynamical stripping, and as a consequence of the location of s
Atomic precision advanced manufacturing (APAM) dopes silicon with enough carriers to change its electronic structure and can be used to create novel devices by defining metallic regions whose boundaries have single-atom abruptness. Incompatibility with the thermal and lithography process requirements for gated silicon transistor manufacturing have inhibited exploration of both how APAM can enhance CMOS performance and how transistor manufacturing steps can accelerate the discovery of new APAM device concepts. In this work, we introduce an APAM process that enables direct integration into the middle of a transistor manufacturing workflow. We show that a process that combines sputtering and annealing with a hardmask preserves a defining characteristic of APAM, a doping density far in excess of the solid solubility limit, while trading another, the atomic precision, for compatibility with manufacturing. The electrical characteristics of a chip combining a transistor with an APAM resistor show that the APAM module has only affected the transistor through the addition of a resistance and not by altering the transistor. This proof-of-concept demonstration also outlines the requirements a
We present DIRECT-3D, a diffusion-based 3D generative model for creating high-quality 3D assets (represented by Neural Radiance Fields) from text prompts. Unlike recent 3D generative models that rely on clean and well-aligned 3D data, limiting them to single or few-class generation, our model is directly trained on extensive noisy and unaligned `in-the-wild' 3D assets, mitigating the key challenge (i.e., data scarcity) in large-scale 3D generation. In particular, DIRECT-3D is a tri-plane diffusion model that integrates two innovations: 1) A novel learning framework where noisy data are filtered and aligned automatically during the training process. Specifically, after an initial warm-up phase using a small set of clean data, an iterative optimization is introduced in the diffusion process to explicitly estimate the 3D pose of objects and select beneficial data based on conditional density. 2) An efficient 3D representation that is achieved by disentangling object geometry and color features with two separate conditional diffusion models that are optimized hierarchically. Given a prompt input, our model generates high-quality, high-resolution, realistic, and complex 3D objects with
Numerical precision in large-scale scientific computations has become an emerging topic due to recent developments in computer hardware. Lower floating point precision offers the potential for significant performance improvements, but the uncertainty added from reducing the numerical precision is a major obstacle for it to reach prevalence in high-fidelity simulations of turbulence. In the present work, the impact of reducing the numerical precision under different rounding schemes is investigated and compared to the presence of white noise in the simulation data to obtain statistical averages of different quantities in the flow. To investigate how this impacts the simulation, an experimental methodology to assess the impact of these sources of uncertainty is proposed, in which each realization $u^i$ at time $t_i$ is perturbed, either by constraining the flow to a coarser discretization of the phase space (corresponding to low precision formats rounded with deterministic and stochastic rounding) or by perturbing the flow with white noise with a uniform distribution. The purpose of this approach is to assess the limiting factors for precision, and how robust a direct numerical simul
Directional methods have been considered to provide a solid proof for the direct detection of the dark matter. Gaseous time-projection-chambers (TPCs) are the most mature devices for directional dark matter searches although there still exist several challenges to overcome. This paper reviews the history, current challenges and future prospects of the gaseous TPCs for directional dark matter searches.
In this paper, we consider a proper Kähler fibration $f \colon X \to Y$ and a singular Hermitian line bundle $(L, h)$ on $X$ with semi-positive curvature. We prove that the direct image sheaf $f_{*}(\mathcal{O}_{X}(K_{X/Y}+L) \otimes \mathcal{I}(h))$, equipped with the Narasimhan-Simha metric, is singular Nakano semi-positive in the sense that the $\overline{\partial}$-equation can be solved with optimal $L^{2}$-estimate. Our proof does not rely on the theory of Griffiths positivity for the direct image sheaf.
We prove two criteria for direct sum decomposability of homogeneous polynomials. For a homogeneous polynomial with a non-zero discriminant, we interpret direct sum decomposability of the polynomial in terms of factorization properties of the Macaulay inverse system of its Milnor algebra. This leads to an if-and-only-if criterion for direct sum decomposability of such a polynomial, and to an algorithm for computing direct sum decompositions over any field, either of characteristic $0$ or of sufficiently large positive characteristic, for which polynomial factorization algorithms exist. For homogeneous forms over algebraically closed fields, we interpret direct sums and their limits as forms that cannot be reconstructed from their Jacobian ideal. We also give simple necessary criteria for direct sum decomposability of arbitrary homogeneous polynomials over arbitrary fields and apply them to prove that many interesting classes of homogeneous polynomials are not direct sums.
Direct detection experiments have established the most stringent constraints on potential interactions between particle candidates for relic, thermal dark matter and Standard Model particles. To surpass current exclusion limits a new generation of experiments is being developed. The upcoming upgrade of the CRESST experiment will incorporate $\mathcal{O}$(100) detectors with different masses ranging from $\sim$2g to $\sim$24g, aiming to achieve unprecedented sensitivity to sub-GeV dark matter particles with a focus on spin-independent dark matter-nucleus scattering. This paper presents a comprehensive analysis of the planned upgrade, detailed experimental strategies, anticipated challenges, and projected sensitivities. Approaches to address and mitigate low-energy excess backgrounds $-$ a key limitation in previous and current sub-GeV dark matter searches $-$ are also discussed. In addition, a long-term roadmap for the next decade is outlined, including other potential scientific applications.
Direct imaging and spectroscopy is the likely means by which we will someday identify, confirm, and characterize an Earth-like planet around a nearby Sun-like star. This Chapter summarizes the current state of knowledge regarding discovering and characterizing exoplanets by direct imaging and spectroscopy. We detail instruments and software needed for direct imaging detections and summarize the current inventory of confirmed and candidate directly-imaged exoplanets. Direct imaging and spectroscopy in the past decade has provided key insights into jovian planet atmospheres, probed the demographics of the outskirts of planetary systems, and shed light on gas giant planet formation. We forecast the new tools and future facilities on the ground and in space that will enhance our capabilities for exoplanet imaging and will likely image habitable zone rocky planets around the nearest stars.
Aligning language models with human preferences through reinforcement learning from human feedback is crucial for their safe and effective deployment. The human preference is typically represented through comparison where one response is chosen over another for a given prompt. However, standard preference datasets often lack explicit information on why a particular choice was made, presenting an ambiguity that can hinder efficient learning and robust alignment, especially given the high cost of acquiring extensive human annotations. While many studies focus on algorithmic improvements, this work adopts a data-centric perspective, exploring how to enhance learning from existing preference data. We propose augmenting standard preference pairs with rationales that explain the reasoning behind the human preference. Specifically, we introduce a simple and principled framework that leverages machine-generated rationales to enrich preference data for preference optimization algorithms. Our comprehensive analysis demonstrates that incorporating rationales improves learning efficiency. Extensive experiments reveal some advantages: rationale-augmented learning accelerates convergence and can
We present the case for a dark matter detector with directional sensitivity. This document was developed at the 2009 CYGNUS workshop on directional dark matter detection, and contains contributions from theorists and experimental groups in the field. We describe the need for a dark matter detector with directional sensitivity; each directional dark matter experiment presents their project's status; and we close with a feasibility study for scaling up to a one ton directional detector, which would cost around $150M.
Direct imaging of gas giant exoplanets provides key information on planetary atmospheres and the architectures of planetary systems. However, few planets have been detected in blind surveys used to achieve imaging detections. Using Gaia and Hipparcos astrometry we identified dynamical evidence for a gas giant planet around the nearby star HIP 99770 and then confirmed this planet by direct imaging with the Subaru Coronagraphic Extreme Adaptive Optics Project. HIP 99770 b orbits 17 astronomical units from its host star, with an insolation comparable to Jupiter's and a dynamical mass of 13.9--16.1 Jupiter masses. Its planet-to-star mass ratio (7--8$\times$10$^{-3}$) is comparable to that other directly-imaged planets. The planet's atmosphere resembles an older, less-cloudy analogue of the atmospheres of previously-imaged exoplanets around HR 8799.
The recently proposed CP language adopts Compositional Programming: a new modular programming style that solves challenging problems such as the Expression Problem. CP is implemented on top of a polymorphic core language with disjoint intersection types called Fi+. The semantics of Fi+ employs an elaboration to a target language and relies on a sophisticated proof technique to prove the coherence of the elaboration. Unfortunately, the proof technique is technically challenging and hard to scale to many common features, including recursion or impredicative polymorphism. Thus, the original formulation of Fi+ does not support the two later features, which creates a gap between theory and practice, since CP fundamentally relies on them. This paper presents a new formulation of Fi+ based on a type-directed operational semantics (TDOS). The TDOS approach was recently proposed to model the semantics of languages with disjoint intersection types (but without polymorphism). Our work shows that the TDOS approach can be extended to languages with disjoint polymorphism and model the full Fi+ calculus. Unlike the elaboration semantics, which gives the semantics to Fi+ indirectly via a target la
We undertook a long term project, DIRECT, to obtain the direct distances to two important galaxies in the cosmological distance ladder -- M31 and M33 -- using detached eclipsing binaries (DEBs) and Cepheids. While rare and difficult to detect, DEBs provide us with the potential to determine these distances with an accuracy better than 5%. The extensive photometry obtained in order to detect DEBs provides us with good light curves for the Cepheid variables. These are essential to the parallel project to derive direct Baade-Wesselink distances to Cepheids in M31 and M33. For both Cepheids and eclipsing binaries, the distance estimates will be free of any intermediate steps. As a first step in the DIRECT project, between September 1996 and October 1997 we obtained 95 full/partial nights on the F. L. Whipple Observatory 1.2 m telescope and 36 full nights on the Michigan-Dartmouth-MIT 1.3 m telescope to search for DEBs and new Cepheids in the M31 and M33 galaxies. In this paper, fifth in the series, we present the catalog of variable stars found in the field M31F $[(α,δ)= (10.\arcdeg10, 40.\arcdeg72), J2000.0]$. We have found 64 variable stars: 4 eclipsing binaries, 52 Cepheids and 8 ot
Let $M$ be a compact smooth manifold of dimension $m$ (without boundary) and $G$ be a finite-dimensional Lie group, with Lie algebra $g$. Let $H^{>m/2}(M,G)$ be the group of all mappings $γ\colon M\to G$ which are $H^s$ for some $s>m/2$. We show that $H^{>m/2}(M,G)$ can be made a regular Lie group in Milnor's sense, modelled on the Silva space $H^{>m/2}(M,g)$ which is the locally convex direct limit of the Hilbert spaces $H^s(M,g)$ for $s>m/2$, such that $H^{>m/2}(M,G)$ is the direct limit of the Hilbert-Lie groups $H^s(M,G)$ for $s>m/2$ as a smooth Lie group. We also explain how the (known) Lie group structure on $H^s(M,G)$ can be obtained as a special case of a general construction of Lie groups $F(M,G)$ whenever real-valued function spaces $F(U,R)$ on open subsets $U$ of $R^m$ are given, subject to simple axioms.
Direct detection of dark matter with directional sensitivity offers not only measurement of both recoil energy and direction of dark matter, but also a way to understand dark matter distribution in the Galaxy. Maxwell distribution is usually supposed as the distribution near the Earth, however, deviation from that, caused by tidal streams in the Galaxy, has been suggested. We explore the possibility of distinguishing the distribution by direct detection using nuclear emulsions.