共找到 20 条结果
Dust emission at submillimeter wavelengths can be used to reliably trace the basic properties of molecular clouds. Early results from a recent Submillimeter Array (SMA) survey of the Andromeda Galaxy (M31) include the first detections of resolved dust continuum emission from individual giant molecular clouds (GMCs) in an external spiral galaxy. This paper updates on the now-complete SMA survey of 80 Herschel-identified giant molecular associations (GMAs) in M31. The SMA survey simultaneously probes dust continuum emission at 230 GHz and the $J = 2 \rightarrow 1$ transitions of the CO isotopologues, $^{12}\rm CO$, $^{13}\rm CO$, and $\rm C^{18}O$ at a spatial resolution of $\lesssim 15~\mathrm{pc}$. Dust continuum emission was detected in 71 cloud cores, of which 26 were resolved. This more than doubles the size of the previous sample. By comparing dust and CO observations with identical astrometry, we directly measure the dust mass to-light ratios, $\rm α^{\prime}_{^{12}CO}$, and $\rm α^{\prime}_{^{13}CO}$. We derive $<α^{\prime}_{\rm ^{12}\rm CO}>~=~0.070~\pm~0.031~M_{\odot}\,(\rm K~km~s^{-1}~pc^{2})^{-1}$ and $<α^{\prime}_{\rm ^{13}\rm CO}>~=~0.37~\pm~0.20~M_{\odot}\,
Some frontier AI developers aim to align language models to a Model Spec or Constitution that describes the intended model behavior. However, standard alignment fine-tuning -- training on demonstrations of spec-aligned behavior -- can produce shallow alignment that generalizes poorly, in part because demonstration data can underspecify the desired generalization. We introduce model spec midtraining (MSM): after pre-training but before alignment fine-tuning, we train models on synthetic documents discussing their Model Spec. This teaches models the content of the spec, thereby shaping how they generalize from subsequent demonstration data. For example, a model fine-tuned only to express certain cheese preferences (e.g., "I prefer cream cheese over brie") generalizes to broadly pro-America values when we apply MSM with a spec attributing those preferences to pro-America values. Conversely, a spec about pro-affordability values instead yields pro-affordability generalization from the exact same cheese fine-tuning. MSM can also shape complex safety-relevant propensities: applying MSM with a spec addressing self-preservation and goal-guarding substantially reduces agentic misalignment r
Let $\mathcal{M}_g$ be the moduli space of smooth curves of genus $g$. The image of a non-constant morphism from a curve $T$ to $\mathcal{M}_g$ is a curve in $\mathcal{M}_g$. By work of González Díez and Harvey, for every integer $g \geq 3$, there exists a complete curve in $\mathcal{M}_g$. Here we generalize the construction to produce new complete curves in $\mathcal{M}_g$. We also find a formula for the genus of each curve $T$ using Galois theory for function fields.
Stochastic convex optimization is a classical problem with well-understood guarantees under first-order feedback. In contrast, for zero-order optimization with noisy function evaluations, a logarithmic gap has persisted between known upper bounds and the $Ω(1/\sqrt{T})$ lower bound, even in the one-dimensional case. In this work, we study the problem of minimizing a convex function $f : [0,1] \to [0,1]$ using a zero-order oracle with subGaussian noise. We propose a computationally efficient algorithm that achieves the optimal $O(1/\sqrt{T})$ convergence rate, matching the lower bound. The result closes the existing gap in one dimension, providing the first sharp rate guarantee in this setting.
Inspired by the dilution refrigerator and magnet monitoring system developed for HAYSTAC at Yale University, Speller Lab at Johns Hopkins University developed the Fridge Real Time Monitoring System (FRTMS). The FRTMS accesses logs saved locally by the dilution refrigerator, saves these logs to various backup locations, edits the logs into a format for upload to a MySQL database, and allows the logs to be remotely monitored in near-real time.
YouTube is central to contemporary mass media. However, the official YouTube API does not provide access to the full set of creators or creator metadata on the platform. This lack of basic visibility into the YouTube ecosystem hinders understanding of the platform's creator economy. Researchers currently have no easy, transparent, or replicable way to construct large-scale datasets of YouTube creators and their audiences over time. This makes it challenging to study vital social questions, such as how changes to the YouTube recommendation algorithm shape creator incentives and by extension the mass media on the platform. We address this gap with TubeCensus, a large-scale longitudinal dataset of YouTube creators and subscriber counts, constructed by collecting, linking, and organizing nearly two decades of YouTube page captures from the Internet Archive. This approach is transparent and replicable and does not require interaction with the YouTube API, whose output can change over time. We validate the coverage of TubeCensus against prior estimates of YouTube's size and find that our resource includes creators responsible for at least 30-36% of all YouTube content. We also find that
REJ1034+396 is one of the few active galactic nuclei with a significant quasi-periodic oscillation (QPO). The QPO has been observed in over 1 Ms of XMM-Newton observations spanning over a decade. We investigate the power spectral density function (PSD) of 7 long (~90 ks) XMM-Newton observations of the active galactic nucleus REJ1034+396 in two energy bands. The soft (0.3-0.5 keV) band targets emission from the disk, while the hard (2-7 keV) band isolates the primary X-ray continuum emission from the corona. The QPO is significantly detected in the hard band of 5 of the 7 observations. The best fitting models indicate that the QPO detection in both bands is entirely attributable to the coronal emission with no additional contribution from the disk. This explains the strong coherence between the hard and soft bands at the QPO frequency. The covariance spectrum is consistent with this picture as the variability at QPO frequencies is attributed solely to fluctuations in the hot corona. The time lag as a function of energy is well described by a ~2000 s intrinsic soft lag, resulting from the disk responding to emission from the corona, that undergoes phase wrapping at approximately the
Dynamic Treatment Regimes (DTRs) provide a systematic framework for optimizing sequential decision-making in chronic disease management, where therapies must adapt to patients' evolving clinical profiles. Inverse probability weighting (IPW) is a cornerstone methodology for estimating regime values from observational data due to its intuitive formulation and established theoretical properties, yet standard IPW estimators face significant limitations, including variance instability and data inefficiency. A fundamental but underexplored source of inefficiency lies in the strict alignment requirement between observed and target treatment trajectories, which fails to account for partial compatibility and discards substantial information from individuals with only minimal deviations from the regime. We propose two novel methodologies that relax the strict inclusion rule through flexible compatibility mechanisms. Both methods provide computationally tractable alternatives that can be easily integrated into existing IPW workflows, offering more efficient approaches to DTR estimation. Theoretical analysis demonstrates that both estimators preserve consistency while achieving superior finite
Mutualistic interactions, where individuals from different species can benefit from each other, are widespread across ecosystems. This study develops a general deterministic model of mutualism involving two populations, assuming that mutualism may involve both costs and benefits for the interacting individuals, leading to density-dependent effects on the dynamics of the two species. This framework aims at generalizing pre-existing models, by allowing the ecological interactions to transition from mutualistic to parasitic when the respective densities of interacting species change. Through ordinary differential equations and phase portrait analysis, we derive general principles governing these systems, identifying sufficient conditions for the emergence of certain dynamic behaviors. In particular, we show that limit cycles can arise when interactions include parasitic phases but are absent in strictly mutualistic regimes. This framework provides a general approach for characterizing the population dynamics of interacting species and highlights the effect of the transitions from mutualism to parasitism due to density dependence.
As AI deployments become more complex and high-stakes, it becomes increasingly important to be able to estimate their risk. AI control is one framework for doing so. However, good control evaluations require eliciting strong attack policies. This can be challenging in complex agentic environments where compute constraints leave us data-poor. In this work, we show how to optimize attack policies in SHADE-Arena, a dataset of diverse realistic control environments. We do this by decomposing attack capability into five constituent skills -- suspicion modeling, attack selection, plan synthesis, execution and subtlety -- and optimizing each component individually. To get around the constraint of limited data, we develop a probabilistic model of attack dynamics, optimize our attack hyperparameters using this simulation, and then show that the results transfer to SHADE-Arena. This results in a substantial improvement in attack strength, reducing safety score from a baseline of 0.87 to 0.41 using our scaffold.
We report the analysis of ~1Ms of XMM-Newton observations of the rapidly accreting active galactic nucleus REJ1034+396. The 0.3-9 keV EPIC-pn spectra are well described by a model consisting of steep continuum emission from the corona accompanied by relativistically-blurred reflection from a highly ionized accretion disk. The source is known to exhibit strong excess soft X-ray emission, which we show is well represented by thermal disk photons Comptonized by a warm plasma spanning the inner accretion flow. Additionally, the EPIC-pn data provide compelling evidence ($ΔC\sim60$ for 4 additional parameters) for the presence of an ultrafast outflow (UFO) with a line-of-sight velocity $v/c=0.307^{+0.001}_{-0.005}$, and an emission signature consistent with reflection of the corona from modestly ionized, outflowing gas. The simultaneous $0.5-2.5$\,keV RGS spectra show clear absorption lines. Modelling of these data confirms the presence of the UFO and constrains its equivalent hydrogen column density, log $N_\mathrm{H}/$(atom cm$^{-2}$) = $21.7^{+0.1}_{-0.2}$. The RGS data also reveal at least two warm absorber components with a modest outflow velocity ($1680^{+40}_{-50}$ km/s). The meas
Trustworthy evaluations of dangerous capabilities are increasingly crucial for determining whether an AI system is safe to deploy. One empirically demonstrated threat is sandbagging - the strategic underperformance on evaluations by AI models or their developers. A promising defense is to monitor a model's chain-of-thought (CoT) reasoning, as this could reveal its intentions and plans. In this work, we measure the ability of models to sandbag on dangerous capability evaluations against a CoT monitor by prompting them to sandbag while being either monitor-oblivious or monitor-aware. We show that both frontier models and small open-sourced models can covertly sandbag against CoT monitoring 0-shot without hints. However, they cannot yet do so reliably: they bypass the monitor 16-36% of the time when monitor-aware, conditioned on sandbagging successfully. We qualitatively analyzed the uncaught CoTs to understand why the monitor failed. We reveal a rich attack surface for CoT monitoring and contribute five covert sandbagging policies generated by models. These results inform potential failure modes of CoT monitoring and may help build more diverse sandbagging model organisms.
In this paper we study the complexity of solving orientable quadratic equations in wreath products $A\wr B$ of finitely generated abelian groups. We give a classification of cases (depending on genus and other characteristics of a given equation) when the problem is computationally hard or feasible.
Creep tests on heterogeneous materials under subcritical loading typically show a power-law decaying strain rate before failure, with the exponent often considered material-dependent but independent of applied stress. By imposing successive small stress relaxations through a displacement feedback loop, we probe creep dynamics and show experimentally that this exponent varies with both applied load and loading direction. Simulations of a disordered fiber bundle model reproduce this load dependence, demonstrating that such models capture essential features of delayed rupture dynamics.
As AI systems become more capable and widely deployed as agents, ensuring their safe operation becomes critical. AI control offers one approach to mitigating the risk from untrusted AI agents by monitoring their actions and intervening or auditing when necessary. Evaluating the safety of these protocols requires understanding both their effectiveness against current attacks and their robustness to adaptive adversaries. In this work, we systematically evaluate a range of control protocols in SHADE-Arena, a dataset of diverse agentic environments. First, we evaluate blue team protocols, including deferral to trusted models, resampling, and deferring on critical actions, against a default attack policy. We find that resampling for incrimination and deferring on critical actions perform best, increasing safety from 50% to 96%. We then iterate on red team strategies against these protocols and find that attack policies with additional affordances, such as knowledge of when resampling occurs or the ability to simulate monitors, can substantially improve attack success rates against our resampling strategy, decreasing safety to 17%. However, deferring on critical actions is highly robust
Photon-pair correlations in spontaneous parametric down conversion are ubiquitous in quantum photonics. The ability to engineer their properties for optimising a specific task is essential, but often challenging in practice. We demonstrate the shaping of spatial correlations between entangled photons in the form of arbitrary amplitude and phase objects. By doing this, we encode image information within the pair correlations, making it undetectable by conventional intensity measurements. It enables the transmission of complex, high-dimensional information using quantum correlations of photons, which can be useful for developing quantum communication and imaging protocols.
The Pierre Auger Observatory has revealed a significant challenge in air shower physics: a discrepancy between the simulated and observed muon content in cosmic-ray interactions, known as the 'Muon Puzzle'. This issue stems from a lack of understanding of high-energy hadronic interactions. Current state-of-the-art hadronic interaction models fall short, underscoring the need for improvements. In this contribution, we explore the integration of the Pythia 8 hadronic interaction model into air shower simulations. While Pythia 8 is primarily used in Large Hadron Collider experiments, recent advancements in its Angantyr module show promise in better describing hadron-nucleus interactions, making it a valuable tool for addressing the Muon Puzzle.
For over two decades, CORSIKA 7 and its previous versions have been the leading Monte Carlo code for simulating extensive air showers. However, its monolithic Fortran-based software design and hand-optimized code has created challenges for maintenance, adaptation to new computing paradigms, and extensions for more complex simulations. Addressing these limitations, the CORSIKA 8 project represents a comprehensive rewrite of CORSIKA 7, seeing its core functionality re-envisioned in a modern and modular C++ framework. CORSIKA 8 has now reached a "physics-complete" state, offering a robust platform that encourages expert development for specialized applications. It supports high-energy hadronic interactions using models such as Sibyll 2.3d, QGSJet-II.04, EPOS-LHC, and Pythia 8.3, alongside the treatment of electromagnetic cascades with PROPOSAL 7.6.2. Key highlights are the support for multiple interaction media, including cross-media particle showers, and an enhanced calculation of radio emissions from particle showers. This contribution provides an overview of the current functionalities, showcases validation results of its simulations, and discusses future development plans.
We consider time-series forecasting problems where data is scarce, difficult to gather, or induces a prohibitive computational cost. As a first attempt, we focus on short-term electricity consumption in France, which is of strategic importance for energy suppliers and public stakeholders. The complexity of this problem and the many levels of geospatial granularity motivate the use of an ensemble of Gaussian Processes (GPs). Whilst GPs are remarkable predictors, they are computationally expensive to train, which calls for a frugal few-shot learning approach. By taking into account performance on GPs trained on a dataset and designing a random walk on these, we mitigate the training cost of our entire Bayesian decision-making procedure. We introduce our algorithm called \textsc{Domino} (ranDOM walk on gaussIaN prOcesses) and present numerical experiments to support its merits.
As AI systems become more capable of complex agentic tasks, they also become more capable of pursuing undesirable objectives and causing harm. Previous work has attempted to catch these unsafe instances by interrogating models directly about their objectives and behaviors. However, the main weakness of trusting interrogations is that models can lie. We propose self-report fine-tuning (SRFT), a simple supervised fine-tuning technique that trains models to occasionally make factual mistakes, then admit them when asked. We show that the admission of factual errors in simple question-answering settings generalizes out-of-distribution (OOD) to the admission of hidden misaligned objectives in adversarial agentic settings. We evaluate SRFT in OOD stealth tasks, where models are instructed to complete a hidden misaligned objective alongside a user-specified objective without being caught by monitoring. After SRFT, models are more likely to confess the details of their hidden objectives when interrogated, even under strong pressure not to disclose them. Interrogation on SRFT models can detect hidden objectives with near-ceiling performance (F1 score = 0.98), while the baseline model lies wh