We continue the study of adapted optimal transport in the discrete-time Gaussian setting. To this end, we introduce a space of filtered Gaussian processes where both the randomness and the flow of information are driven by a Gaussian white noise. On this space, the adapted $2$-Wasserstein distance (${AW}_2$) admits a variational representation as a constrained orthogonal Procrustes problem between Cholesky factors. Furthermore, the resulting quotient space is the ${AW}_2$-completion of the space of Gaussian distributions on the path space. We also characterize explicitly the ${AW}_2$-projections onto the subspaces of Gaussian martingales. Next, we analyze the adapted Brenier coupling -- a multivariate generalization of the Knothe--Rosenblatt coupling that serves as a myopic solution to the adapted transport problem, and compute its transport cost. Utilizing a Gaussian random matrix framework, we investigate the asymptotic behavior of transport costs as the time horizon grows; notably, we establish that the transport costs of all Gaussian bicausal couplings are asymptotically equivalent, whereas the classical Bures--Wasserstein distance is strictly smaller. Finally, we demonstrate t
The adapted Bures--Wasserstein space consists of Gaussian processes endowed with the adapted Wasserstein distance. It can be viewed as the analogue of the classical Bures--Wasserstein space in optimal transport for the setting of stochastic processes, where the standard Wasserstein distance is inadequate and has to be replaced by its adapted counterpart. We develop a comprehensive geometric theory for the adapted Bures--Wasserstein space, thereby also providing the first results on the fine geometric structure of adapted optimal transport. In particular, we show that the adapted Bures--Wasserstein space is an Alexandrov space with non-negative curvature and provide explicit descriptions of tangent cones and exponential maps. Moreover, we show that Gaussian processes satisfying a natural non-degeneracy condition form a geodesically convex subspace. This subspace is characterized precisely by the property that its tangent cones are linear and hence coincide with the tangent space.
Adapted invariant measures, such as the natural area measure (Liouville), have a central place in the development of ergodic theory for billiards. These measures ensure local Pesin charts can be constructed almost everywhere even in the nonuniformly hyperbolic setting. Recently, for Sinai billiards satisfying certain conditions, the unique measure of maximal entropy has been shown to be adapted. However, not all positive entropy measures are. To investigate the connection between entropy and adaptedness, we examine Markov interval maps with exactly one singularity. We prove that a condition relating the entropy of the map and the "strength" of the singularity determines if the measure of maximal entropy is adapted with respect to this singularity. We also show that under a Hölder condition, recurrence of the singularity is necessary to have nonadapted invariant measures.
We present Six Llamas, a comparative study examining whether large language models fine-tuned on distinct religious corpora encode systematically different patterns of ethical reasoning. Six variants of Meta-Llama-3.1-8B are constructed: one unmodified control and five LoRA-adapted models trained exclusively on the sacred and theological texts of Christianity, Islam, Judaism, Hinduism, or Buddhism. All six models are probed with an identical battery of 17 standardized ethical prompts spanning moral dilemmas, game-theoretic scenarios, public policy questions, and moral-psychological self-assessments. To assess robustness and reproducibility, we implement a multi-temperature sampling design spanning ten temperature settings. We compute response consistency metrics, pairwise inter-model agreement rates, temperature sensitivity coefficients across four prompt domains, and run-to-run stability analyses. Findings show that LoRA-adapted models produce ethical reasoning patterns that are (a) systematically differentiated from the base model, (b) consistent with the moral logics of their training traditions, (c) structured along interpretable dimensions in moral-philosophical space, (d) cor
The adapted Wasserstein distance controls the calibration errors of optimal values in various stochastic optimization problems, pricing and hedging problems, optimal stopping problems, etc. However, statistical aspects of the adapted Wasserstein distance are bottlenecked by the failure of empirical measures to converge under this distance. Kernel smoothing and adapted projection have been introduced to construct converging substitutes of empirical measures, known respectively as smoothed empirical measures and adapted empirical measures. However, both approaches have limitations. Specifically, smoothed empirical measures lack comprehensive convergence results, whereas adapted empirical measures in practical applications lead to fewer distinct samples compared to standard empirical measures. In this work, we address both of the aforementioned issues. First, we develop comprehensive convergence results of smoothed empirical measures. We then introduce a smoothed version for adapted empirical measures, which provide as many distinct samples as desired. We refer them as adapted smoothed empirical measures and establish their convergence in mean, deviation, and almost sure convergence.
Pinsker's classical inequality asserts that the total variation $TV(μ, ν)$ between two probability measures is bounded by $\sqrt{ 2H(μ|ν)}$ where $H$ denotes the relative entropy (or Kullback-Leibler divergence). Considering the discrete metric, $TV$ can be seen as a Wasserstein distance and as such possesses an adapted variant $ATV$. Adapted Wasserstein distances have distinct advantages over their classical counterparts when $μ, ν$ are the laws of stochastic processes $(X_k)_{k=1}^n, (Y_k)_{k=1}^n$ and exhibit numerous applications from stochastic control to machine learning. In this note we observe that the adapted total variation distance $ATV$ satisfies the Pinsker-type inequality $$ ATV(μ, ν)\leq \sqrt{n} \sqrt{2 H(μ|ν)}.$$
The $L^1$ transport-entropy inequality (or $T_1$ inequality), which bounds the $1$-Wasserstein distance in terms of the relative entropy, is known to characterize Gaussian concentration. To extend the $T_1$ inequality to laws of discrete-time processes while preserving their temporal structure, we investigate the adapted $T_1$ inequality which relates the $1$-adapted Wasserstein distance to the relative entropy. Building on the Bolley--Villani inequality, we establish the adapted $T_1$ inequality under the same moment assumption as the classical $T_1$ inequality.
Document Visual Question Answering (DocVQA) presents a complex multimodal challenge, requiring models to exploit visual, textual, and layout information from documents. Although Vision-Language Models (VLMs) have shown remarkable performance in text-vision tasks, their robustness and transferability to different document domains remains underexplored. In this study, we present a comprehensive evaluation of 8 open-source pretrained VLMs on DocVQA in three different document domains: industrial documents of varying type, infographics, and presentation slides. We systematically assess model performance under zero-shot evaluations, fully supervised finetuning with inter- and intra-dataset evaluations, and few-shot learning evaluations of knowledge transfer between domains. Our findings demonstrate that while large pretrained VLMs possess strong zero-shot baselines for structured layouts, their performance strongly decreases on visually complex layouts of infographics and slides. Although parameter scaling is a dominant factor on performance, supervised finetuning yields higher relative gains in smaller architectures. Furthermore, our cross-domain and few-shot experiments show that visu
Causal optimal transport and adapted Wasserstein distance have applications in different fields from optimization to mathematical finance and machine learning. The goal of this article is to provide equivalent formulations of these concepts in classic probabilistic language. In particular, we prove a Skorokhod representation theorem for adapted weak convergence, reformulate the equivalence of stochastic processes using Markovian lifts, and give an expression for the adapted Wasserstein distance based on representing processes on a common stochastic basis.
The adapted Wasserstein distance is a metric for quantifying distributional uncertainty and assessing the sensitivity of stochastic optimization problems on time series data. A computationally efficient alternative to it, is provided by the entropically regularized adapted Wasserstein distance. Suffering from similar shortcomings as classical optimal transport, there are only few explicitly known solutions to those distances. Recently, Gunasingam--Wong provided a closed-form representation of the adapted Wasserstein distance between real-valued stochastic processes with Gaussian laws. In this paper, we extend their work in two directions, by considering multidimensional ($\mathbb{R}^d$-valued) stochastic processes with Gaussian laws and including the entropic regularization. In both settings, we provide closed-form solutions.
We derive explicitly the adapted $2$-Wasserstein distance between non-degenerate Gaussian distributions on $\mathbb{R}^N$ and characterize the optimal bicausal coupling(s). This leads to an adapted version of the Bures-Wasserstein distance on the space of positive definite matrices.
The Wasserstein distance $\mathcal{W}_p$ is an important instance of an optimal transport cost. Its numerous mathematical properties as well as applications to various fields such as mathematical finance and statistics have been well studied in recent years. The adapted Wasserstein distance $\mathcal{A}\mathcal{W}_p$ extends this theory to laws of discrete time stochastic processes in their natural filtrations, making it particularly well suited for analyzing time-dependent stochastic optimization problems. While the topological differences between $\mathcal{A}\mathcal{W}_p$ and $\mathcal{W}_p$ are well understood, their differences as metrics remain largely unexplored beyond the trivial bound $\mathcal{W}_p\lesssim \mathcal{A}\mathcal{W}_p$. This paper closes this gap by providing upper bounds of $\mathcal{A}\mathcal{W}_p$ in terms of $\mathcal{W}_p$ through investigation of the smooth adapted Wasserstein distance. Our upper bounds are explicit and are given by a sum of $\mathcal{W}_p$, Eder's modulus of continuity and a term characterizing the tail behavior of measures. As a consequence, upper bounds on $\mathcal{W}_p$ automatically hold for $\mathcal{AW}_p$ under mild regularity
We consider empirical measures of $\R^{d}$-valued stochastic process in finite discrete-time. We show that the adapted empirical measure introduced in the recent work \cite{backhoff2022estimating} by Backhoff et al. in compact spaces can be defined analogously on $\R^{d}$, and that it converges almost surely to the underlying measure under the adapted Wasserstein distance. Moreover, we quantitatively analyze the convergence of the adapted Wasserstein \add{distance} between those two measures. We establish convergence rates of the expected error as well as the deviation error under different moment conditions. \add{Under suitable integrability and kernel assumptions, we recover the optimal convergence rates of both expected error and deviation error.} Furthermore, we propose a modification of the adapted empirical measure with \add{projection} on a non-uniform grid, which obtains the same convergence rate but under weaker assumptions.
A number of researchers have independently introduced topologies on the set of laws of stochastic processes that extend the usual weak topology. Depending on the respective scientific background this was motivated by applications and connections to various areas (e.g. Plug-Pichler - stochastic programming, Hellwig - game theory, Aldous - stability of optimal stopping, Hoover-Keisler - model theory). Remarkably, all these seemingly independent approaches define the same adapted weak topology in finite discrete time. Our first main result is to construct an adapted variant of the empirical measure that consistently estimates the laws of stochastic processes in full generality. A natural compatible metric for the weak adapted topology is the given by an adapted refinement of the Wasserstein distance, as established in the seminal works of Pflug-Pichler. Specifically, the adapted Wasserstein distance allows to control the error in stochastic optimization problems, pricing and hedging problems, optimal stopping problems, etc. in a Lipschitz fashion. The second main result of this article yields quantitative bounds for the convergence of the adapted empirical measure with respect to adap
Let $\mathfrak a$ be an algebraic Lie algebra. An adapted pair for $\mathfrak a$ is pair $(h,η)$ consisting of an ad-semisimple element of $h \in \mathfrak a$ and a regular element of $η\in \mathfrak a^*$ satisfying $(ad \ h)η=-η$. An adapted pair $(h,η)$ is said to satisfy integrality if $ad \ h$ has integer eigenvalues on $\mathfrak a$. Integrality is shown to hold for any Frobenius Lie algebra which is a biparabolic subalgebra of a semisimple Lie algebra; but may fail in general. Call $\mathfrak a$ regular if there are no proper semi-invariant polynomial functions on $\mathfrak a^*$ and if the subalgebra of invariant functions is polynomial. In this case there are no known counter-examples to integrality. It is shown that if $\mathfrak a$ is the canonical truncation of a biparabolic subalgebra of a simple Lie algebra $\mathfrak g$ which is regular and admits an adapted pair $(h,η)$, then the eigenvalues of $ad \ h$ on $\mathfrak a$ lie in $\frac{1}{m}\mathbb Z$, where $m$ is a coefficient of a simple root in the highest root of $\mathfrak g$. Let $\mathfrak a$ be a regular Lie algebra admitting an adapted pair $(h,η)$. Let $\mathfrak a_\mathbb Z$ be the subalgebra spanned by the
The topology of weak convergence does not account for the growth of information over time that is captured in the filtration of an adapted stochastic process. For example, two adapted stochastic processes can have very similar laws but give completely different results in applications such as optimal stopping, queuing theory, or stochastic programming. To address such discontinuities, Aldous introduced the extended weak topology, and subsequently, Hoover and Keisler showed that both, weak topology and extended weak topology, are just the first two topologies in a sequence of topologies that get increasingly finer. We use higher rank expected signatures to embed adapted processes into graded linear spaces and show that these embeddings induce the adapted topologies of Hoover--Keisler.
This paper explores the geometric structure of the spectrahedral cone, called the symmetry adapted PSD cone, and the symmetry adapted Gram spectrahedron of a symmetric polynomial. In particular, we determine the dimension of the symmetry adapted PSD cone, describe its extreme rays, and discuss the structure of its matrix representations. We also consider the symmetry adapted Gram spectrahedra for specific families of symmetric polynomials including binary symmetric polynomials, quadratics, and ternary quartics and sextics which give us further insight into these symmetric SOS polynomials. Finally, we discuss applications of the theory of sums of squares and symmetric polynomials which arise from symmetric function inequalities.
Recent research on dialectal NLP has identified data scarcity as a primary limitation. To address this limitation, this paper presents a catalog of contemporary Basque dialectal data and resources, offering a systematic and comprehensive compilation of the dialectal data currently available in Basque. Two types of data sources have been distinguished: online data originally written in some dialect, and standard-to-dialect adapted data. The former includes all dialectal data that can be found online, such as news and radio sites, informal tweets, as well as online resources such as dictionaries, atlases, grammar rules, or videos. The latter consists of data that has been adapted from the standard variety to dialectal varieties, either manually or automatically. Regarding the manual adaptation, the test split of the XNLI Natural Language Inference dataset was manually adapted into three Basque dialects: Western, Central, and Navarrese-Lapurdian, yielding a high-quality parallel gold standard evaluation dataset. With respect to the automatic dialectal adaptation, the automatically adapted physical commonsense dataset (BasPhyCowest) underwent additional manual evaluation by native spea
We present the e-Llama models: 8 billion and 70 billion parameter large language models that are adapted towards the e-commerce domain. These models are meant as foundation models with deep knowledge about e-commerce, that form a base for instruction- and fine-tuning. The e-Llama models are obtained by continuously pretraining the Llama 3.1 base models on 1 trillion tokens of domain-specific data. We discuss our approach and motivate our choice of hyperparameters with a series of ablation studies. To quantify how well the models have been adapted to the e-commerce domain, we define and implement a set of multilingual, e-commerce specific evaluation tasks. We show that, when carefully choosing the training setup, the Llama 3.1 models can be adapted towards the new domain without sacrificing significant performance on general domain tasks. We also explore the possibility of merging the adapted model and the base model for a better control of the performance trade-off between domains.
Imitation learning enables autonomous agents to learn from human examples, without the need for a reward signal. Still, if the provided dataset does not encapsulate the task correctly, or when the task is too complex to be modeled, such agents fail to reproduce the expert policy. We propose to recover from these failures through online adaptation. Our approach combines the action proposal coming from a pre-trained policy with relevant experience recorded by an expert. The combination results in an adapted action that closely follows the expert. Our experiments show that an adapted agent performs better than its pure imitation learning counterpart. Notably, adapted agents can achieve reasonable performance even when the base, non-adapted policy catastrophically fails.