Recent advances have sparked significant interest in the development of privacy-preserving Principal Component Analysis (PCA). However, many existing approaches rely on restrictive assumptions, such as assuming sub-Gaussian data or being vulnerable to data contamination. Additionally, some methods are computationally expensive or depend on unknown model parameters that must be estimated, limiting their accessibility for data analysts seeking privacy-preserving PCA. In this paper, we propose a differentially private PCA method applicable to heavy-tailed and potentially contaminated data. Our approach leverages the property that the covariance matrix of properly rescaled data preserves eigenvectors and their order under elliptical distributions, which include Gaussian and heavy-tailed distributions. By applying a bounded transformation, we enable straightforward computation of principal components in a differentially private manner. Additionally, boundedness guarantees robustness against data contamination. We conduct both theoretical analysis and empirical evaluations of the proposed method, focusing on its ability to recover the subspace spanned by the leading principal components.
The current-induced magnetisation dynamics in a ferromagnet at elevated temperatures can be described by the Landau--Lifshitz--Bloch (LLB) equation with spin-torque terms. In this paper, we focus on the regime above the Curie temperature. We first establish the existence and uniqueness of a global strong solution to the model in spatial dimensions $d=1,2,3$, under an additional smallness assumption on the initial data if $d=3$. Relevant smoothing and decay estimates are also derived. We then propose a fully discrete, linearly implicit finite element scheme for the problem and prove that it achieves optimal-order convergence, assuming adequate regularity of the exact solution. In addition, we introduce an unconditionally energy-stable finite element method for the case of negligible non-adiabatic torque. This scheme is also shown to converge optimally and, in the absence of current, preserves energy dissipation at the discrete level. Finally, we present numerical simulations that support the theoretical analysis and demonstrate the performance of the proposed methods.
We introduce three representative topics in semi-classical analysis. Starting from the correspondence between classical and quantum mechanics, basic semi-classical analysis tools and results are presented. The three topics are investigated in the light of the introduced techniques allowing one to emphasize different aspects of semi-classical analysis.
It is well known that the controllability property of partial differential equations (PDEs) is closely linked to the proof of an observability inequality for the adjoint system, which, sometimes, involves analyzing a spectral problem associated with the PDE under consideration. In this work, we study a series of spectral issues that ensure the controllability of the renowned Korteweg-de Vries equation on a star-graph. This investigation reduces to determining when certain functions, associated with this spectral problem, are entire. The novelty here lies in presenting this detailed analysis in the context of a star graph structure.
The Origins, Spectral Interpretation, Resource Identification, and Security Regolith Explorer (OSIRIS-REx) spacecraft arrived at its target, near-Earth asteroid 101955 Bennu, in December 2018. After one year of operating in proximity, the team selected a primary site for sample collection. In October 2020, the spacecraft descended to the surface of Bennu and collected a sample. The spacecraft departed Bennu in May 2021 and will return the sample to Earth in September 2023. The analysis of the returned sample will produce key data to determine the history of this B-type asteroid and that of its components and precursor objects. The main goal of the OSIRIS-REx Sample Analysis Plan is to provide a framework for the Sample Analysis Team to meet the Level 1 mission requirement to analyze the returned sample to determine presolar history, formation age, nebular and parent-body alteration history, relation to known meteorites, organic history, space weathering, resurfacing history, and energy balance in the regolith of Bennu. To achieve this goal, this plan establishes a hypothesis-driven framework for coordinated sample analyses, defines the analytical instrumentation and techniques to b
We develop a theory of quantum harmonic analysis on lattices in $\mathbb{R}^{2d}$. Convolutions of a sequence with an operator and of two operators are defined over a lattice, and using corresponding Fourier transforms of sequences and operators we develop a version of harmonic analysis for these objects. We prove analogues of results from classical harmonic analysis and the quantum harmonic analysis of Werner, including Tauberian theorems and a Wiener division lemma. Gabor multipliers from time-frequency analysis are described as convolutions in this setting. The quantum harmonic analysis is thus a conceptual framework for the study of Gabor multipliers, and several of the results include results on Gabor multipliers as special cases.
Free boundaries of biofilms advancing on surfaces evolve according to conservation laws coupled with systems of partial differential equations for velocities, pressures and chemicals affecting cell behavior. Thin film approximations lead to complicated quasi-stationary systems coupling stationary transport equations and compressible Stokes systems with convection-reaction-diffusion equations.We establish existence, uniqueness and stability of solutions of the different submodels involved and then obtain well posedness results for the full system. Our analysis relies on the construction of weak solutions for the steady transport equations under sign assumptions and the reformulation of the compressible Stokes problem as an elliptic system with enhanced regularity properties on the pressure. We need to consider velocity fields whose divergence and normal boundary components satisfy sign conditions, instead of vanishing as classical results require. Applications include the study of cells, biofilms and tissues, where one phase is a liquid solution, whereas the other one is assorted biomass.
We present a new method which generalizes subspace learning based on eigenvalue and generalized eigenvalue problems. This method, Roweis Discriminant Analysis (RDA), is named after Sam Roweis to whom the field of subspace learning owes significantly. RDA is a family of infinite number of algorithms where Principal Component Analysis (PCA), Supervised PCA (SPCA), and Fisher Discriminant Analysis (FDA) are special cases. One of the extreme special cases, which we name Double Supervised Discriminant Analysis (DSDA), uses the labels twice; it is novel and has not appeared elsewhere. We propose a dual for RDA for some special cases. We also propose kernel RDA, generalizing kernel PCA, kernel SPCA, and kernel FDA, using both dual RDA and representation theory. Our theoretical analysis explains previously known facts such as why SPCA can use regression but FDA cannot, why PCA and SPCA have duals but FDA does not, why kernel PCA and kernel SPCA use kernel trick but kernel FDA does not, and why PCA is the best linear method for reconstruction. Roweisfaces and kernel Roweisfaces are also proposed generalizing eigenfaces, Fisherfaces, supervised eigenfaces, and their kernel variants. We also
The clinical interest is often to measure the volume of a structure, which is typically derived from a segmentation. In order to evaluate and compare segmentation methods, the similarity between a segmentation and a predefined ground truth is measured using popular discrete metrics, such as the Dice score. Recent segmentation methods use a differentiable surrogate metric, such as soft Dice, as part of the loss function during the learning phase. In this work, we first briefly describe how to derive volume estimates from a segmentation that is, potentially, inherently uncertain or ambiguous. This is followed by a theoretical analysis and an experimental validation linking the inherent uncertainty to common loss functions for training CNNs, namely cross-entropy and soft Dice. We find that, even though soft Dice optimization leads to an improved performance with respect to the Dice score and other measures, it may introduce a volume bias for tasks with high inherent uncertainty. These findings indicate some of the method's clinical limitations and suggest doing a closer ad-hoc volume analysis with an optional re-calibration step.
The computational analysis of fiber network fracture is an emerging field with application to paper, rubber-like materials, hydrogels, soft biological tissue, and composites. Fiber networks are often described as probabilistic structures of interacting one-dimensional elements, such as truss-bars and beams. Failure may then be modeled as strong discontinuities in the displacement field that are directly embedded within the structural finite elements. As for other strain-softening materials, the tangent stiffness matrix can be non-positive definite, which diminishes the robustness of the solution of the coupled (monolithic) two-field problem. Its uncoupling, and thus the use of a staggered solution method where the field variables are solved alternatingly, avoids such difficulties and results in a stable, but sub-optimally converging solution method. In the present work, we evaluate the staggered against the monolithic solution approach and assess their computational performance in the analysis of fiber network failure. We then propose a hybrid solution technique that optimizes the performance and robustness of the computational analysis. It represents a matrix regularization techni
The analysis of the leukemia data from Whitehead/MIT group is a discriminant analysis (also called a supervised learning). Among thousands of genes whose expression levels are measured, not all are needed for discriminant analysis: a gene may either not contribute to the separation of two types of tissues/cancers, or it may be redundant because it is highly correlated with other genes. There are two theoretical frameworks in which variable selection (or gene selection in our case) can be addressed. The first is model selection, and the second is model averaging. We have carried out model selection using Akaike information criterion and Bayesian information criterion with logistic regression (discrimination, prediction, or classification) to determine the number of genes that provide the best model. These model selection criteria set upper limits of 22-25 and 12-13 genes for this data set with 38 samples, and the best model consists of only one (no.4847, zyxin) or two genes. We have also carried out model averaging over the best single-gene logistic predictors using three different weights: maximized likelihood, prediction rate on training set, and equal weight. We have observed tha
A novel social networks sentiment analysis model is proposed based on Twitter sentiment score (TSS) for real-time prediction of the future stock market price FTSE 100, as compared with conventional econometric models of investor sentiment based on closed-end fund discount (CEFD). The proposed TSS model features a new baseline correlation approach, which not only exhibits a decent prediction accuracy, but also reduces the computation burden and enables a fast decision making without the knowledge of historical data. Polynomial regression, classification modelling and lexicon-based sentiment analysis are performed using R. The obtained TSS predicts the future stock market trend in advance by 15 time samples (30 working hours) with an accuracy of 67.22% using the proposed baseline criterion without referring to historical TSS or market data. Specifically, TSS's prediction performance of an upward market is found far better than that of a downward market. Under the logistic regression and linear discriminant analysis, the accuracy of TSS in predicting the upward trend of the future market achieves 97.87%.
With the increasingly detailed investigation of game play and tactics in invasive team sports such as soccer, it becomes ever more important to present causes, actions and findings in a meaningful manner. Visualizations, especially when augmenting relevant information directly inside a video recording of a match, can significantly improve and simplify soccer match preparation and tactic planning. However, while many visualization techniques for soccer have been developed in recent years, few have been directly applied to the video-based analysis of soccer matches. This paper provides a comprehensive overview and categorization of the methods developed for the video-based visual analysis of soccer matches. While identifying the advantages and disadvantages of the individual approaches, we identify and discuss open research questions, soon enabling analysts to develop winning strategies more efficiently, do rapid failure analysis or identify weaknesses in opposing teams.
The recently proposed statistical finite element (statFEM) approach synthesises measurement data with finite element models and allows for making predictions about the unknown true system response. We provide a probabilistic error analysis for a prototypical statFEM setup based on a Gaussian process prior under the assumption that the noisy measurement data are generated by a deterministic true system response function that satisfies a second-order elliptic partial differential equation for an unknown true source term. In certain cases, properties such as the smoothness of the source term may be misspecified by the Gaussian process model. The error estimates we derive are for the expectation with respect to the measurement noise of the $L^2$-norm of the difference between the true system response and the mean of the statFEM posterior. The estimates imply polynomial rates of convergence in the numbers of measurement points and finite element basis functions and depend on the Sobolev smoothness of the true source term and the Gaussian process model. A numerical example for Poisson's equation is used to illustrate these theoretical results.
Numerical schemes provably preserving the positivity of density and pressure are highly desirable for MHD, but the rigorous positivity-preserving (PP) analysis remains challenging. The difficulties mainly arise from the intrinsic complexity of the MHD equations as well as the indeterminate relation between the PP property and the divergence-free condition on magnetic field. We present the first rigorous PP analysis of conservative schemes with Lax-Friedrichs (LF) flux for ideal MHD. The significant innovation is the discovery of theoretical connection between PP property and a discrete divergence-free (DDF) condition. This connection is established through the generalized LF splitting properties, which are alternatives of the usually-expected LF splitting property that does not hold for ideal MHD. The generalized LF splitting properties involve a number of admissible states strongly coupled by DDF condition, making their derivation very difficult. We derive these properties via a novel equivalent form of the admissible state set and an important inequality skillfully constructed by technical estimates. Rigorous PP analysis is presented for finite volume and discontinuous Galerkin s
Symbolic methods of analysis are valuable tools for investigating complex time-dependent signals. In particular, the ordinal method defines sequences of symbols according to the ordering in which values appear in a time series. This method has been shown to yield useful information, even when applied to signals with large noise contamination. Here we use ordinal analysis to investigate the transition between eyes closed (EC) and eyes open (EO) resting states. We analyze two {EEG} datasets (with 71 and 109 healthy subjects) with different recording conditions (sampling rates and the number of electrodes in the scalp). Using as diagnostic tools the permutation entropy, the entropy computed from symbolic transition probabilities, and an asymmetry coefficient (that measures the asymmetry of the likelihood of the transitions between symbols) we show that ordinal analysis applied to the raw data distinguishes the two brain states. In both datasets, we find that the EO state is characterized by higher entropies and lower asymmetry coefficient, as compared to the EC state. Our results thus show that these diagnostic tools have the potential for detecting and characterizing changes in time-
Computer vision is a growing field with a lot of new applications in automation and robotics, since it allows the analysis of images and shapes for the generation of numerical or analytical information. One of the most used method of information extraction is image filtering through convolution kernels, with each kernel specialized for specific applications. The objective of this paper is to present a novel convolution kernels, based on principles of electromagnetic potentials and fields, for a general use in computer vision and to demonstrate its usage for shape and stroke analysis. Such filtering possesses unique geometrical properties that can be interpreted using well understood physics theorems. Therefore, this paper focuses on the development of the electromagnetic kernels and on their application on images for shape and stroke analysis. It also presents several interesting features of electromagnetic kernels, such as resolution, size and orientation independence, robustness to noise and deformation, long distance stroke interaction and ability to work with 3D images
Over a decade ago, the H1 Collaboration decided to embrace the object-oriented paradigm and completely redesign its data analysis model and data storage format. The event data model, based on the RooT framework, consists of three layers - tracks and calorimeter clusters, identified particles and finally event summary data - with a singleton class providing unified access. This original solution was then augmented with a fourth layer containing user-defined objects. This contribution will summarise the history of the solutions used, from modifications to the original design, to the evolution of the high-level end-user analysis object framework which is used by H1 today. Several important issues are addressed - the portability of expert knowledge to increase the efficiency of data analysis, the flexibility of the framework to incorporate new analyses, the performance and ease of use, and lessons learned for future projects.
We recently reported the existence of fluctuations in neural signals that may permit neurons to code multiple simultaneous stimuli sequentially across time. This required deploying a novel statistical approach to permit investigation of neural activity at the scale of individual trials. Here we present tests using synthetic data to assess the sensitivity and specificity of this analysis. We fabricated datasets to match each of several potential response patterns derived from single-stimulus response distributions. In particular, we simulated dual stimulus trial spike counts that reflected fluctuating mixtures of the single stimulus spike counts, stable intermediate averages, single stimulus winner-take-all, or response distributions that were outside the range defined by the single stimulus responses (such as summation or suppression). We then assessed how well the analysis recovered the correct response pattern as a function of the number of simulated trials and the difference between the simulated responses to each "stimulus" alone. We found excellent recovery of the mixture, intermediate, and outside categories (>97% percent correct), and good recovery of the single/winner-ta
This paper is concerned with the analysis and numerical analysis for the optimal control of first-order magneto-static equations. Necessary and sufficient optimality conditions are established through a rigorous Hilbert space approach. Then, on the basis of the optimality system, we prove functional a posteriori error estimators for the optimal control, the optimal state, and the adjoint state. 3D numerical results illustrating the theoretical findings are presented.