Long-horizon LLM sessions outlive their context windows, and the standard mitigations - truncation, summarization, retrieval - share a structural flaw: they treat history as flat text, discarding precisely the content that makes a session resumable: decisions and their rationales, task status, and file modification history. We present TokenMizer, an open-source transparent proxy that maintains session history as a typed knowledge graph and, at context boundaries, replaces the raw transcript with a token-budgeted serialization of session state. The schema comprises 14 node types and 7 edge types under an 8-state lifecycle in which decisions can be superseded or explicitly invalidated; bitemporal validity intervals support time-travel queries; and first-class decision-transition records preserve why each decision replaced its predecessor (trigger, reason, evidence). Version 0.3.1 embeds this memory core in a production-shaped serving layer - SSE streaming, security middleware, nine provider adapters, a monitoring dashboard, graph exports (D3 JSON, self-contained interactive HTML, Obsidian Canvas) - and exposes checkpoint/resume to agents as Model Context Protocol tools. The evaluatio
Java remains central to enterprise software, and many applications outlive their original architecture. Migrating them across frameworks is a behavior-preserving refactoring spanning build configuration, dependency injection, persistence, request handling, and deployment. Existing software-engineering benchmarks cover bug fixing, feature implementation, and language or version modernization, but leave cross-framework refactoring largely unmeasured. We introduce ScarfBench, a benchmark for behavior-preserving cross-framework refactoring of enterprise Java applications. It is built from expert-written implementation triples across Spring, Jakarta EE, and Quarkus: 34 applications (29 focused single-layer, 5 whole) yielding 102 variants (~151K lines across 1946 source and test files) and 204 directed refactoring tasks. Each task gives an agent a working source application and a target framework; the agent must synthesize a target implementation preserving the source behavior. Correctness is evaluated by an application-specific executable oracle: the candidate must compile, deploy in a containerized target runtime, and pass behavioral tests over the application's observable interface. W
The extremely rapid evolution of kilonovae results in spectra that change on an hourly basis. These spectra are key to understanding the processes occurring within the event, but this rapid evolution is an unfamiliar domain compared to other explosive transient events, such as supernovae. In particular, the most obvious P Cygni feature in the spectra of AT2017gfo -- commonly attributed to strontium -- possesses an emission component that emerges after, and ultimately outlives, its associated absorption dip. This delay is theorised to arise from reverberation effects, wherein photons emitted earlier in the kilonova's evolution are scattered before reaching the observer, causing them to be detected at later times. We aim to examine how the finite speed of light -- and therefore the light travel time to an observer -- contributes to the shape and evolution of spectral features in kilonovae. Using a simple model, and tracking the length of the journey photons undertake to an observer, we are able to test the necessity of accounting for this time delay effect when modelling kilonovae. In periods where the photospheric temperature is rapidly evolving, we show spectra synthesised using a
Revealing the interactions binding electronic and lattice components of cooperative quantum order is central to sculpting new states of matter. This challenge is epitomized by the charge density wave material 1T-TiSe$_2$, where photoexcitation disrupts its presumed hybrid exciton-phonon order. This exposes a paradox: the electronic component collapses within femtoseconds while the periodic lattice distortion persists. If the lattice distortion outlives the excitonic condensate, were they truly intertwined? Here we resolve this by uncovering a low-frequency mode (approx. 0.13 THz) emerging only in the ordered state, signaling exciton-phonon coupling. This mode is consistent with a locked phason -- a collective excitation arising if coupling between the excitonic condensate and lattice reduces continuous phase symmetry to a discrete one, giving the excitonic Goldstone mode finite mass. This is captured by an effective theory describing a shared potential landscape. At a critical threshold, the collapse of excitonic order flattens the potential, triggering an exciton-phonon catastrophe: selective overheating of the charge density wave phonon, disappearance of the locked phason, and su
Scientific knowledge on the Web is published as passive assertions and cannot decide when to validate evidence, reconcile contradictions, or update confidence as findings accumulate. Curation depends on centralised middleware and institutional continuity, but when registries close, active stewardship stops even when data remain online. We advance the concept of Autonomous FAIR Digital Objects (aFDOs) from an abstract idea to an operational model, to offer a route from passive scientific publication toward accountable, standards-aligned automation that can outlive its publishing institutions. aFDO augments FDOs with three capabilities anchored in Semantic Web standards, namely 1) a policy layer over RDF-star aligned with PROV-O, SHACL, and ODRL for portable condition-action rules, 2) an announcement layer over ActivityStreams 2.0 that bounds per-announcement evaluation cost, and 3) an agreement layer that resolves multi-source contradictions through reputation and confidence weighted agreement under a bounded adversarial model. We provide a formal definition that distinguishes policy specifications, event handlers, and communication interfaces. We evaluate an open reference implemen
As Large Language Models (LLMs) transition from standalone chat interfaces to foundational reasoning layers in multi-agent systems and recursive evaluation loops (LLM-as-a-judge), the detection of durable, provider-level behavioral signatures becomes a critical requirement for safety and governance. Traditional benchmarks measure transient task accuracy but fail to capture stable, latent response policies -- the ``prevailing mindsets'' embedded during training and alignment that outlive individual model versions. This paper introduces a novel auditing framework that utilizes psychometric measurement theory -- specifically latent trait estimation under ordinal uncertainty -- to quantify these tendencies without relying on ground-truth labels. Utilizing forced-choice ordinal vignettes masked by semantically orthogonal decoys and governed by cryptographic permutation-invariance, the research audits nine leading models across dimensions including Optimization Bias, Sycophancy, and Status-Quo Legitimization. Using Mixed Linear Models (MixedLM) and Intraclass Correlation Coefficient (ICC) analysis, the research identifies that while item-level framing drives high variance, a persistent `
Identifying bias in LLMs is ongoing. Because they are still in development, what is true today may be false tomorrow. We therefore need general strategies for debiasing that will outlive current models. Strategies developed for debiasing human decision making offer one promising approach as they incorporate an LLM-style prompt intervention designed to bring latent knowledge into awareness during decision making. LLMs trained on vast amounts of information contain information about potential biases, counter-arguments, and contradictory evidence, but that information may only be brought to bear if prompted. Metacognitive prompts developed in the human decision making literature are designed to achieve this, and as I demonstrate here, they show promise with LLMs. The prompt I focus on here is "could you be wrong?" Following an LLM response, this prompt leads LLMs to produce additional information, including why they answered as they did, errors, biases, contradictory evidence, and alternatives, none of which were apparent in their initial response. Indeed, this metaknowledge often reveals that how LLMs and users interpret prompts are not aligned. Here I demonstrate this prompt using a
Simulation was launched in the 1950s, nicknamed a tool of "last resort." Over the years, this Operations Research (OR) method has made significant progress, and utilizing the accelerated advances in computer science (hardware and software, processing speed, and advanced information visualization capabilities) to improve simulation usability in research and practice. After overcoming the initial obstacles and the scare of outliving its usefulness in the 2000s, computer simulation has remained a popular OR tool applied in diverse industries and sectors, earning its popularity leading to the term "simulation everywhere." This study uses bibliographic data from research and practice literature to evaluate the evolutionary expansion in simulation, focusing on discrete-event simulation (DES). The results show asymmetrical but positive yearly literature out-put, broadened DES adoption in diverse fields, and sustained relevance as a scientific method for tackling old, new, and emerging issues. Also, DES is an essential tool in Industry 4.0 and plays a central role in digital transformation that has swept the industrial space, from manufacturing to healthcare and other sectors. With the eme
We investigate the evolution of spread over three days in a numerical ensemble experiment starting from tiny initial condition uncertainty. We simulate a real event during which three mesoscale convective systems occur in close proximity to the midlatitude jet. The spread evolution is compared with an existing conceptual three-stage model. Each system follows the first stage, characterised by development of convective variability. Nevertheless, we find significant variation among the systems in their propensity to interact with the jet stream, which characterises conceptual stage 2. One exemplary convective system follows the conceptual evolution of Baumgart et al., i.e., convective uncertainty initially projects onto the jet by upper-tropospheric outflow, which further amplifies spread through nonlinear growth as it propagates downstream. Rossby-like dispersion in the downstream spread is strongly associated with the convective variability. In contrast, for another convective system, convective variability projects onto the local anticyclonic flow aloft. Subsequently, this anticyclonic perturbation hardly (if at all) projects convective uncertainty onto the particularly straight j
Local reasoning about programs that combine aliasing and mutable state is a longstanding challenge. Existing approaches -- ownership systems, linear and affine types, uniqueness types, and lexical effect tracking -- impose global restrictions such as uniqueness or linearity, or rely on shallow syntactic analyses. These designs fall short with higher-order functions and shared mutable state. Reachability Types (RT) track aliasing and separation in higher-order programs, ensuring runtime safety and non-interference. However, RT systems face three key limitations: (1) they prohibit cyclic references, ruling out non-terminating computations and fixed-point combinators; (2) they require deep tracking, where a qualifier must include all transitively reachable locations, reducing precision and hindering optimizations like fine-grained parallelism; and (3) referent qualifier invariance prevents referents from escaping their allocation contexts, making reference factories inexpressible. In this work, we address these limitations by extending RT with three mechanisms that enhance expressiveness. First, we introduce cyclic references, enabling recursive patterns to be encoded directly through
Leakage out of the computational subspace is a major limitation of current state-of-the-art neutral-atom quantum computers and a significant challenge for scalable systems. In a quantum processor with cesium atoms, we demonstrate proof-of-principle circuit-based conversion of leakage errors to erasure errors via Leakage Detection Units (LDUs), which non-destructively map information about the presence or absence of the qubit onto the state of an ancilla. With a standard LDU circuit, we successfully convert leakage errors to erasure errors for all major leakage pathways while preserving the quantum information in the case that no leakage occurred. We benchmark the performance of the LDU using a three-outcome low-loss state detection method and also explore the advantages of three-outcome measurements for LDUs. We find that the LDU detects atom-loss errors with ~93.4% accuracy, limited by technical imperfections of our apparatus. We further compile and execute a SWAP LDU, wherein the roles of the original data atom and ancilla atom are exchanged under the action of the LDU, providing 'free refilling' of atoms in the case of leakage errors. This circuit-based leakage-to-erasure error
In a previous work, we investigated the evolution of the flow field around sunspots during sunspot decay and compared it with the flow field of supergranular cells. The decay of a sunspot proceeds as it interacts with its surroundings. This is manifested by the changes observed in the flow field surrounding the decaying spot. We now investigate in detail the evolution of the flow field in the direct periphery of the sunspots of the same sample and aim to provide a complete picture of the role of large-scale flows present in sunspot cells. We analyse the horizontal velocity profiles of sunspots obtained from observations by the Helioseismic and Magnetic Imager (HMI) on board the Solar Dynamics Observatory (SDO). We follow their evolution across the solar disc from their stable phase to their decay and their final disappearance. We find two different scenarios for the evolution of the flow region surrounding a spot in the final stage of its decay: (i) either the flow cell implodes and disappears under the action of the surrounding supergranules or (ii) it outlives the spot. In the later case, an inwards flow towards the remaining naked spot develops in the vicinity closest to the spo
The Cassini state equilibrium associated with the precession of the Moon predicts that the mantle, fluid core and solid inner core precess at different angles. We present estimates of the dissipation from viscous friction associated with the differential precession at the core-mantle boundary (CMB), $Q_{cmb}$, and at the inner core boundary (ICB), $Q_{icb}$, as a function of the evolving lunar orbit. We focus on the latter and show that, provided the inner core was larger than 100 km, $Q_{icb}$ may have been as high as $10^{10}-10^{11}$ W for most of the lunar history for a broad range of core density models. This is larger than the power required to maintain the fluid core in an adiabatic state, therefore the heat released by the differential precession at the ICB can drive a past lunar dynamo by thermal convection. This dynamo can outlive the dynamo from precession at the CMB and may have shutoff only relatively recently. Estimates of the magnetic field strength at the lunar surface are of the order of a few $μ$T, compatible with the lunar paleomagnetic intensities recorded after 3 Ga. We further show that it is possible that a transition of the Cassini state associated with the
The intuition suggested by the Drake equation implies that technology should be less prevalent than biology in the galaxy. However, it has been appreciated for decades in the SETI community that technosignatures could be more abundant, longer-lived, more detectable, and less ambiguous than biosignatures. We collect the arguments for and against technosignatures' ubiquity and discuss the implications of some properties of technological life that fundamentally differ from nontechnological life in the context of modern astrobiology: It can spread among the stars to many sites, it can be more easily detected at large distances, and it can produce signs that are unambiguously technological. As an illustration in terms of the Drake equation, we consider two Drake-like equations, for technosignatures (calculating N(tech)) and biosignatures (calculating N(bio)). We argue that Earth and humanity may be poor guides to the longevity term L and that its maximum value could be very large, in that technology can outlive its creators and even its host star. We conclude that while the Drake equation implies that N(bio)>>N(tech), it is also plausible that N(tech)>>N(bio). As a consequen