Machine-Generated Text (MGT) is becoming increasingly difficult to distinguish from Human-Written Text (HWT). This trend has exacerbated malicious activities such as fake news and online fraud. The generalization ability of fine-tuned detectors relies heavily on dataset quality, and simply expanding the sources of MGT may become increasingly insufficient. Further augmentation of the generation process is required. Based on HC-Var's theory, enhancing the human-like alignment of MGT not only facilitates robustness testing of existing detectors but also boosts the generalization ability of detectors fine-tuned on such aligned MGT datasets. Therefore, we propose the \textbf{M}achine-\textbf{A}ugment-\textbf{G}enerated Text via \textbf{A}lignment (MAGA) Detection Benchmark. MAGA integrates several alignment methods, ranging from prompt construction to \textbf{G}enerator-\textbf{D}etector \textbf{A}dversarial \textbf{R}einforcement \textbf{L}earning (GDARL) and the reasoning process. In our experiments, the RoBERTa detector fine-tuned on MAGA achieves an average improvement of 4.60\% in generalization AUC. Conversely, the aligned MGTs in MAGA also lead to an average decrease of 8.13\% in
Vision Transformers (ViTs) and Convolutional Neural Networks (CNNs) face inherent challenges in image matting, particularly in preserving fine structural details. ViTs, with their global receptive field enabled by the self-attention mechanism, often lose local details such as hair strands. Conversely, CNNs, constrained by their local receptive field, rely on deeper layers to approximate global context but struggle to retain fine structures at greater depths. To overcome these limitations, we propose a novel Morpho-Aware Global Attention (MAGA) mechanism, designed to effectively capture the morphology of fine structures. MAGA employs Tetris-like convolutional patterns to align the local shapes of fine structures, ensuring optimal local correspondence while maintaining sensitivity to morphological details. The extracted local morphology information is used as query embeddings, which are projected onto global key embeddings to emphasize local details in a broader context. Subsequently, by projecting onto value embeddings, MAGA seamlessly integrates these emphasized morphological details into a unified global structure. This approach enables MAGA to simultaneously focus on local morpho
We study the optimal rectangular-discrepancy approximation of permutons by finite permutations. We transfer bounds from discrepancy theory to this more restricted setup. Moreover, we show that superlinear approximation can occur only for permutons supported by graphs of measure-preserving functions, and demonstrate how the local regularity of this function obstructs approximability. We also consider the biased Brownian separable permuton and prove a lower bound on its approximation error by showing that its supporting measure-preserving function has Lipschitz points almost surely.
We study the metric geometry of the set of permutons under the rectangular distance $d_{\square}$. We determine the Chebyshev radius to be 1/4 and characterize all Chebyshev centers: a permuton is a center if and only if it is 1/2- periodic in each coordinate. We also describe permutons that attain the extremal distance 1/4 from a given center.
In the previous decades, the size of level sets of functions have been extensively studied in various setups involving different regularity properties and size notions. In the case of Hölder functions, the authors have provided various bounds, but to date no explicit formulae have been found for any studied dimension and the results were valid only about very specific fractals. In this paper, for the first time, we have a result valid for a large class of self-similar sets, namely we prove that for these fractals Lebesgue almost every level set of the generic 1-Hölder-$α$ function defined on $F\subseteq \mathbb{R}^p$ has upper box dimension $\dim_H F - α$.
This work considers combinatorial and statistical aspects of {\em{shifts of finite type}}, which are families of words over a finite alphabet which avoid a fixed class of {\emph{forbidden}} sub-words. The overarching question we are interested in is: how do local statistics of a uniformly random element of the shift space depend on combinatorial features of the forbidden set? We focus on the binary alphabet $\{0,1\}$, the class of shift spaces where a single pattern is forbidden, and the average frequency of $1$s (equivalently, the probability of observing $1$ at a given position). In this case, the relevant combinatorial information is encoded by a two-variable auto-correlation polynomial associated to the forbidden word, which we call the {\em{border polynomial}}. We present several results and examples characterizing the ordering of all words by their letter frequencies: for example, we describe the set of patterns which, when forbidden, cause the frequency of $1$s to increase, decrease, or stay exactly $1/2$. We conjecture that, among forbidden patterns of the same length (except for four exceptional words), the letter frequency is monotone with respect to the number of $1$s in
The increasing use of 3D imaging technologies in biological sciences is generating vast repositories of anatomical data, yet significant barriers prevent this data from reaching its full potential in educational and collaborative contexts. While sharing raw CT and MRI scans has become routine, distributing value-added segmented datasets, where anatomical structures are precisely labeled and delineated, remains difficult and rare. Current repositories function primarily as static archives, lacking mechanisms for iterative refinement, community-driven curation, standardized orientation protocols, and the controlled terminology essential for downstream computational applications, including artificial intelligence, to help us analyze and interpret these unprecedented data resources. We introduce MorphoDepot, a framework that adapts the "fork-and-contribute" model, a cornerstone of modern open-source software development, for collaborative management of 3D morphological data. By integrating git version control and GitHub's "social" collaborative infrastructure with 3D Slicer and its SlicerMorph extension, MorphoDepot transforms segmented anatomical datasets from static resources into dy
We introduce a novel type of quantum error correcting code, called the spinor code, based on spaces defined by total spin. The code is a nonstabilizer code, and is also a nonlinear quantum error correcting code, meaning that quantum information is encoded in a parameterized family of quantum states, rather than a linear superposition of code words. Syndrome measurements are performed by projecting on states with differing total spin, with an associated correction to map states back to the maximum total spin space. We show that the code is asymptotically capable of protecting against any single qubit Pauli error for Gaussian distributed states such as spin coherent state. We directly evaluate the performance under the depolarizing channel, considering various cases, with and without initialization and measurement errors, as well as two qubit errors. We estimate the code-capacity threshold to be in the range of 32-75%, while the phenomenological threshold is in the range 9-75%.
The quantum internet is a rapidly developing technological reality, yet, it remains unclear what kind of quantum network structures might emerge. Since indirect quantum communication is already feasible and preserves absolute security of the communication channel, a new node joining the quantum network does not need to connect directly to its desired target. Instead, in our proposed quantum preferential attachment model, it uniformly randomly connects to any node within the proximity of the target, including, but not restricted to, the target itself. This local flexibility is found to qualitatively change the global network behavior, leading to two distinct classes of complex network architectures, both of which are small-world, but neither of which is scale-free. Our numerical findings are supported by rigorous analytic results, in a framework that incorporates quantum and classical variants of preferential attachment in a unified phase diagram. Besides quantum networks, we expect that our results will have broad implications for classical scenarios where there is flexibility in establishing new connections.
The individual-based model of simple contagion processes is considered on regular graphs. This model explicitly incorporates the adjacency matrix of the network enabling us to study the effect of network structure on the dynamic of the propagation process. While the asymptotic behaviour of the model is well known, the transient behaviour has been less studied. Our goal in this paper is to give a theoretical estimate on the accuracy of the one-dimensional population-level approximation. This is carried out for arbitrary simple contagion processes and regular Turán graphs. Numerical evidence is shown that the theoretical estimate is rather sharp for dense graphs.
Hearts subjected to volume overload (VO) are prone to detrimental anatomical and functional changes in response to elevated mechanical loading, ultimately leading to heart failure. Experimental findings now emphasize that organ-scale changes following VO cannot be explained by myocyte growth alone, as traditionally proposed in the literature. Collagen degradation, in particular, has been associated with VO and assumed to play a central role in both its acute and chronic stages. This hypothesis, however, remains to be substantiated by comprehensive mechanistic evidence, and each constituent contribution to myocardial growth and remodeling (G&R) processes is yet to be quantified. In this work, we present a multi-constituent G&R framework that integrates a mixture-based constitutive model within the kinematic growth formulation. This framework enables us to mechanistically assess the relative contributions of collagen and myocyte changes to alterations in tissue properties, ventricular dimensions, and growth phenotype. Our numerical results confirm that collagen remodeling affects the passive mechanical response of the myocardium, whereas myocytes predominantly influence the e
For a permuton $μ$ let $H_n(μ)$ denote the Shannon entropy of the sampling distribution of $μ$ on $n$ points. We investigate the asymptotic growth of $H_n(μ)$ for a wide class of permutons. We prove that if $μ$ has a non-vanishing absolutely continuous part, then $H_n(μ)$ has a growth rate $Θ(n \log n)$. We show that if $μ$ is the graph of a piecewise continuously differentiable, measure-preserving function $f$, then $H_n(μ)/n$ tends to the Kolmogorov--Sinai entropy of $f$. Using genericity arguments, we also prove the existence of function permutons for which $H_n(μ)$ does not converge either after normalizing by $n$ or by $n\log n$. We study the sampling entropy of a natural family of random fractal-like permutons determined by a sequence of i.i.d. choices. It turns out that for every $n$, $H_n(μ)/n$ is heavily concentrated. We prove that the sequence $H_n(μ)/n$ either converges or has deterministic log-periodic oscillations almost surely, and argue towards the conjecture that in nondegenerate case, oscillation holds. On the other hand, for a straightforward random perturbation of the model $\tildeμ$ of $μ$, we prove the almost sure convergence of $H_n(\tildeμ)/n$.
In this paper, we study the question when a (rational or Gaussian) integral vector can be extended to an integral orthogonal basis consisting of vectors of equal length. We also study when a set of integral vectors has such an extension. Some necessary conditions are given which are proven to be sufficient in dimensions $3$ and $4$.
The digitization of biological specimens has revolutionized the field of morphology, creating large collections of 3D data, and microCT in particular. This revolution was initially supported by the development of open-source software tools, specifically the development of SlicerMorph extension to the open-source image analytics platform 3D Slicer. Through SlicerMorph and 3D Slicer, biologists, morphologists and scientists in related fields have all the necessary tools to import, visualize and analyze these complex and large datasets in a single platform that is flexible and expandible, without the need of proprietary software that hinders scientific collaboration and sharing. Yet, a significant "compute gap" remains: While data and software are now open and accessible, the necessary high-end computing resources to run them are often not equally accessible in all institutions, and particularly lacking at Primarily Undergraduate Institutions (PUIs) and other educational settings. Here, we present MorphoCloud, an "IssuesOps"-based platform that leverages Github Actions and the JetStream2 cloud farm to provide on-demand, research-grade computing environments to researchers working with
The $α$-Weierstrass function is defined as $W_g^{α,b}(x) = \sum_{k=0}^{\infty} b^{-αk} g(b^k x)$, where $g$ is a Lipschitz function on the unit circle. For a prevalent $α$-Weierstrass function, we prove that the upper Minkowski dimension of every level set is at most $1-α$, and the Hausdorff dimension of almost every level set equals $1-α$ with respect to its occupation measure. We further demonstrate that the occupation measure of a prevalent $α$-Weierstrass function is absolutely continuous with respect to the Lebesgue measure. Consequently, the result on the Hausdorff dimension of level sets applies to a set of level sets with positive Lebesgue measure. A central tool in our analysis is the Weierstrass embedding. For a sufficiently large dimension $d$, we construct Lipschitz functions $g_0, g_1, \ldots, g_{d-1}$ such that the mapping $x \mapsto \big(W_{g_0}^{α,b}(x), W_{g_1}^{α,b}(x), \ldots, W_{g_{d-1}}^{α,b}(x)\big)$ is $α$-bi-Hölder. We also prove that such an embedding requires at least $1/α$ coordinate functions.
Current Ethereum fraud detection methods rely on context-independent, numerical transaction sequences, failing to capture semantic of account transactions. Furthermore, the pervasive homogeneity in Ethereum transaction records renders it challenging to learn discriminative account embeddings. Moreover, current self-supervised graph learning methods primarily learn node representations through graph reconstruction, resulting in suboptimal performance for node-level tasks like fraud account detection, while these methods also encounter scalability challenges. To tackle these challenges, we propose LMAE4Eth, a multi-view learning framework that fuses transaction semantics, masked graph embedding, and expert knowledge. We first propose a transaction-token contrastive language model (TxCLM) that transforms context-independent numerical transaction records into logically cohesive linguistic representations. To clearly characterize the semantic differences between accounts, we also use a token-aware contrastive learning pre-training objective together with the masked transaction model pre-training objective, learns high-expressive account representations. We then propose a masked account
Dimensions of level sets of generic continuous functions and generic Hölder functions defined on a fractal $F$ encode information about the geometry, ``the thickness" of $F$. While in the continuous case this quantity is related to a reasonably tame dimension notion which is called the topological Hausdorff dimension of $F$, the Hölder case seems to be highly nontrivial. A number of earlier papers attempted to deal with this problem, carrying out investigation in the case of Hausdorff dimension and box dimension. In this paper we continue our study of the Hausdorff dimension of almost every level set of generic $1$-Hölder-$ α$ functions, denoted by $D_{*}( α, F)$. We substantially improve previous lower and upper bounds on $D_{*}( α, Δ)$, where $Δ$ is the Sierpiński triangle, achieving asymptotically equal bounds as $α\to 0+$. Using a similar argument, we also give an even stronger lower bound on the generic lower box dimension of level sets. Finally, we construct a connected fractal $F$ on which there is a phase transition of $D_{*}( α, F)$, thus providing the first example exhibiting this behavior.
The study of matroid products traces back to the 1970s, when Lovász and Mason studied the existence of various types of matroid products with different strengths. Among these, the tensor product is arguably the most important, which can be considered as an extension of the tensor product from linear algebra. However, Las Vergnas showed that the tensor product of two matroids does not always exist. Over the following four decades, matroid products remained surprisingly underexplored, regaining attention only in recent years due to applications in tropical geometry and the limit theory of matroids. In this paper, inspired by the concept of coupling in probability theory, we introduce the notion of coupling for matroids -- or, more generally, for submodular set functions. This operation can be viewed as a relaxation of the tensor product. Unlike the tensor product, however, we prove that a coupling always exists for any two submodular functions and can be chosen to be increasing if the original functions are increasing. As a corollary, we show that two matroids always admit a matroid coupling, leading to a novel operation on matroids. Our construction is algorithmic, providing an orac
Hausdorff dimension of level sets of generic continuous functions defined on fractals can give information about the "thickness/narrow cross-sections" "network" corresponding to a fractal set, $F$. This lead to the definition of the topological Hausdorff dimension of fractals. Finer information might be obtained by considering the Hausdorff dimension of level sets of generic $1$-Hölder-$α$ functions, which has a stronger dependence on the geometry of the fractal, as displayed in our previous papers. In this paper, we extend our investigations to the lower and upper box-counting dimension as well: while the former yields results highly resembling the ones about Hausdorff dimension of level sets, the latter exhibits a different behaviour. Instead of "finding narrow-cross sections", results related to upper box-counting dimension try to "measure" how much level sets can spread out on the fractal, how widely the generic function can "oscillate" on it. Key differences are illustrated by giving estimates concerning the Sierpiński triangle.