共找到 20 条结果
I first met Leo Breiman in 1979 at the beginning of his third career, Professor of Statistics at Berkeley. He obtained his PhD with Loéve at Berkeley in 1957. His first career was as a probabilist in the Mathematics Department at UCLA. After distinguished research, including the Shannon--Breiman--MacMillan Theorem and getting tenure, he decided that his real interest was in applied statistics, so he resigned his position at UCLA and set up as a consultant. Before doing so he produced two classic texts, Probability, now reprinted as a SIAM Classic in Applied Mathematics, and Statistics. Both books reflected his strong opinion that intuition and rigor must be combined. He expressed this in his probability book which he viewed as a combination of his learning the right hand of probability, rigor, from Loéve, and the left-hand, intuition, from David Blackwell.
We study percolation in the following random environment: let $Z$ be a Poisson process of constant intensity in the plane, and form the Voronoi tessellation of the plane with respect to $Z$. Colour each Voronoi cell black with probability $p$, independently of the other cells. We show that the critical probability is 1/2. More precisely, if $p>1/2$ then the union of the black cells contains an infinite component with probability 1, while if $p<1/2$ then the distribution of the size of the component of black cells containing a given point decays exponentially. These results are analogous to Kesten's results for bond percolation in the square lattice. The result corresponding to Harris' Theorem for bond percolation in the square lattice is known: Zvavitch noted that one of the many proofs of this result can easily be adapted to the random Voronoi setting. For Kesten's results, none of the existing proofs seems to adapt. The methods used here also give a new and very simple proof of Kesten's Theorem for the square lattice; we hope they will be applicable in other contexts as well.
Our data are random fields of multivariate Gaussian observations, and we fit a multivariate linear model with common design matrix at each point. We are interested in detecting those points where some of the coefficients are nonzero using classical multivariate statistics evaluated at each point. The problem is to find the $P$-value of the maximum of such a random field of test statistics. We approximate this by the expected Euler characteristic of the excursion set. Our main result is a very simple method for calculating this, which not only gives us the previous result of Cao and Worsley [Ann. Statist. 27 (1999) 925--942] for Hotelling's $T^2$, but also random fields of Roy's maximum root, maximum canonical correlations [Ann. Appl. Probab. 9 (1999) 1021--1057], multilinear forms [Ann. Statist. 29 (2001) 328--371], $\barχ^2$ [Statist. Probab. Lett 32 (1997) 367--376, Ann. Statist. 25 (1997) 2368--2387] and $χ^2$ scale space [Adv. in Appl. Probab. 33 (2001) 773--793]. The trick involves approaching the problem from the point of view of Roy's union-intersection principle. The results are applied to a problem in shape analysis where we look for brain damage due to nonmissile trauma.
We prove that if (X,\mathfrakA,P) is an arbitrary probability space with countably generated σ-algebra \mathfrakA, (Y,\mathfrakB,Q) is an arbitrary complete probability space with a lifting ρand \hat R is a complete probability measure on \mathfrakA \hat \otimes_R \mathfrakB determined by a regular conditional probability {S_y:y\in Y} on \mathfrakA with respect to \mathfrakB, then there exist a lifting πon (X\times Y,\mathfrakA \hat \otimes_R \mathfrakB,\hat R) and liftings σ_y on (X,\hat \mathfrakA_y,\hat S_y), y\in Y, such that, for every E\in\mathfrakA \hat \otimes_R \mathfrakB and every y\in Y, [π(E)]^y=σ_y\bigl([π(E)]^y\bigr). Assuming the absolute continuity of R with respect to P\otimes Q, we prove the existence of a regular conditional probability {T_y:y\in Y} and liftings \varpi on (X\times Y,\mathfrakA \hat \otimes_R \mathfrakB,\hat R), ρ' on (Y,\mathfrakB,\hat Q) and σ_y on (X,\hat \mathfrakA_y,\hat S_y), y\in Y, such that, for every E\in\mathfrakA \hat \otimes_R \mathfrakB and every y\in Y, [\varpi(E)]^y=σ_y\bigl([\varpi(E)]^y\bigr) and \varpi(A\times B)=\bigcup_{y\inρ'(B)}σ_y(A)\times{y}\qquadif A\times B\in\mathfrakA\times\mathfrakB. Both results are generalizations o
Let $X=\{X(t),t\in {\mathbb{R}}^N\}$ be a centered Gaussian random field with stationary increments and $X(0)=0$. For any compact rectangle $T\subset {\mathbb{R}}^N$ and $u\in {\mathbb{R}}$, denote by $A_u=\{t\in T:X(t)\geq u\}$ the excursion set. Under $X(\cdot)\in C^2({\mathbb{R}}^N)$ and certain regularity conditions, the mean Euler characteristic of $A_u$, denoted by ${\mathbb{E}}\{\varphi(A_u)\}$, is derived. By applying the Rice method, it is shown that, as $u\to\infty$, the excursion probability ${\mathbb{P}}\{\sup_{t\in T}X(t)\geq u\}$ can be approximated by ${\mathbb{E}}\{\varphi(A_u)\}$ such that the error is exponentially smaller than ${\mathbb{E}}\{\varphi(A_u)\}$. This verifies the expected Euler characteristic heuristic for a large class of Gaussian random fields with stationary increments.
What makes a problem suitable for statistical analysis? Are historical and religious questions addressable using statistical calculations? Such issues have long been debated in the statistical community and statisticians and others have used historical information and texts to analyze such questions as the economics of slavery, the authorship of the Federalist Papers and the question of the existence of God. But what about historical and religious attributions associated with information gathered from archeological finds? In 1980, a construction crew working in the Jerusalem neighborhood of East Talpiot stumbled upon a crypt. Archaeologists from the Israel Antiquities Authority came to the scene and found 10 limestone burial boxes, known as ossuaries, in the crypt. Six of these had inscriptions. The remains found in the ossuaries were reburied, as required by Jewish religious tradition, and the ossuaries were catalogued and stored in a warehouse. The inscriptions on the ossuaries were catalogued and published by Rahmani (1994) and by Kloner (1996) but there reports did not receive widespread public attention. Fast forward to March 2007, when a television ``docudrama'' aired on The
Several classical results on boundary crossing probabilities of Brownian motion and random walks are extended to asymptotically Gaussian random fields, which include sums of i.i.d. random variables with multidimensional indices, multivariate empirical processes, and scan statistics in change-point and signal detection as special cases. Some key ingredients in these extensions are moderate deviation approximations to marginal tail probabilities and weak convergence of the conditional distributions of certain ``clumps'' around high-level crossings. We also discuss how these results are related to the Poisson clumping heuristic and tube formulas of Gaussian random fields, and describe their applications to laws of the iterated logarithm in the form of the Kolmogorov--Erdős--Feller integral tests.
We formulate the insurance risk process in a general Levy process setting, and give general theorems for the ruin probability and the asymptotic distribution of the overshoot of the process above a high level, when the process drifts to -\infty a.s. and the positive tail of the Levy measure, or of the ladder height measure, is subexponential or, more generally, convolution equivalent. Results of Asmussen and Kluppelberg [Stochastic Process. Appl. 64 (1996) 103-125] and Bertoin and Doney [Adv. in Appl. Probab. 28 (1996) 207-226] for ruin probabilities and the overshoot in random walk and compound Poisson models are shown to have analogues in the general setup. The identities we derive open the way to further investigation of general renewal-type properties of Levy processes.
We survey a number of models from physics, statistical mechanics, probability theory and combinatorics, which are each described in terms of an orthogonal polynomial ensemble. The most prominent example is apparently the Hermite ensemble, the eigenvalue distribution of the Gaussian Unitary Ensemble (GUE), and other well-known ensembles known in random matrix theory like the Laguerre ensemble for the spectrum of Wishart matrices. In recent years, a number of further interesting models were found to lead to orthogonal polynomial ensembles, among which the corner growth model, directed last passage percolation, the PNG droplet, non-colliding random processes, the length of the longest increasing subsequence of a random permutation, and others. Much attention has been paid to universal classes of asymptotic behaviors of these models in the limit of large particle numbers, in particular the spacings between the particles and the fluctuation behavior of the largest particle. Computer simulations suggest that the connections go even farther and also comprise the zeros of the Riemann zeta function. The existing proofs require a substantial technical machinery and heavy tools from various p
We consider the high energy physics unfolding problem where the goal is to estimate the spectrum of elementary particles given observations distorted by the limited resolution of a particle detector. This important statistical inverse problem arising in data analysis at the Large Hadron Collider at CERN consists in estimating the intensity function of an indirectly observed Poisson point process. Unfolding typically proceeds in two steps: one first produces a regularized point estimate of the unknown intensity and then uses the variability of this estimator to form frequentist confidence intervals that quantify the uncertainty of the solution. In this paper, we propose forming the point estimate using empirical Bayes estimation which enables a data-driven choice of the regularization strength through marginal maximum likelihood estimation. Observing that neither Bayesian credible intervals nor standard bootstrap confidence intervals succeed in achieving good frequentist coverage in this problem due to the inherent bias of the regularized point estimate, we introduce an iteratively bias-corrected bootstrap technique for constructing improved confidence intervals. We show using simul
In this paper, we address the problem of constructing a uniform probability measure on $\mathbb{N}$. Of course, this is not possible within the bounds of the Kolmogorov axioms and we have to violate at least one axiom. We define a probability measure as a finitely additive measure assigning probability $1$ to the whole space, on a domain which is closed under complements and finite disjoint unions. We introduce and motivate a notion of uniformity which we call weak thinnability, which is strictly stronger than extension of natural density. We construct a weakly thinnable probability measure and we show that on its domain, which contains sets without natural density, probability is uniquely determined by weak thinnability. In this sense, we can assign uniform probabilities in a canonical way. We generalize this result to uniform probability measures on other metric spaces, including $\mathbb{R}^n$.
Notes for a Course on Probability and Statistics: L1: Elements of Probability; L2: Bayesian Inference; L3: Monte Carlo Methods
Statistical depth measures the centrality of a point with respect to a given distribution or data cloud. It provides a natural center-outward ordering of multivariate data points and yields a systematic nonparametric multivariate analysis scheme. In particular, the half-space depth is shown to have many desirable properties and broad applicability. However, the empirical half-space depth is zero outside the convex hull of the data. This property has rendered the empirical half-space depth useless outside the data cloud, and limited its utility in applications where the extreme outlying probability mass is the focal point, such as in classification problems and control charts with very small false alarm rates. To address this issue, we apply extreme value statistics to refine the empirical half-space depth in "the tail." This provides an important linkage between data depth, which is useful for inference on centrality, and extreme value statistics, which is useful for inference on extremity. The refined empirical half-space depth can thus extend all its utilities beyond the data cloud, and hence broaden greatly its applicability. The refined estimator is shown to have substantially
Complex functional brain network analyses have exploded over the last eight years, gaining traction due to their profound clinical implications. The application of network science (an interdisciplinary offshoot of graph theory) has facilitated these analyses and enabled examining the brain as an integrated system that produces complex behaviors. While the field of statistics has been integral in advancing activation analyses and some connectivity analyses in functional neuroimaging research, it has yet to play a commensurate role in complex network analyses. Fusing novel statistical methods with network-based functional neuroimage analysis will engender powerful analytical tools that will aid in our understanding of normal brain function as well as alterations due to various brain disorders. Here we survey widely used statistical and network science tools for analyzing fMRI network data and discuss the challenges faced in filling some of the remaining methodological gaps. When applied and interpreted correctly, the fusion of network scientific and statistical methods has a chance to revolutionize the understanding of brain function.
Data science has become increasingly essential for the production of official statistics, as it enables the automated collection, processing, and analysis of large amounts of data. With such data science practices in place, it enables more timely, more insightful and more flexible reporting. However, the quality and integrity of data-science-driven statistics rely on the accuracy and reliability of the data sources and the machine learning techniques that support them. In particular, changes in data sources are inevitable to occur and pose significant risks that are crucial to address in the context of machine learning for official statistics. This paper gives an overview of the main risks, liabilities, and uncertainties associated with changing data sources in the context of machine learning for official statistics. We provide a checklist of the most prevalent origins and causes of changing data sources; not only on a technical level but also regarding ownership, ethics, regulation, and public perception. Next, we highlight the repercussions of changing data sources on statistical reporting. These include technical effects such as concept drift, bias, availability, validity, accur
We formulate an optimal stopping problem for a geometric Brownian motion where the probability scale is distorted by a general nonlinear function. The problem is inherently time inconsistent due to the Choquet integration involved. We develop a new approach, based on a reformulation of the problem where one optimally chooses the probability distribution or quantile function of the stopped state. An optimal stopping time can then be recovered from the obtained distribution/quantile function, either in a straightforward way for several important cases or in general via the Skorokhod embedding. This approach enables us to solve the problem in a fairly general manner with different shapes of the payoff and probability distortion functions. We also discuss economical interpretations of the results. In particular, we justify several liquidation strategies widely adopted in stock trading, including those of "buy and hold", "cut loss or take profit", "cut loss and let profit run" and "sell on a percentage of historical high".
Quantities with right-skewed distributions are ubiquitous in complex social systems, including political conflict, economics and social networks, and these systems sometimes produce extremely large events. For instance, the 9/11 terrorist events produced nearly 3000 fatalities, nearly six times more than the next largest event. But, was this enormous loss of life statistically unlikely given modern terrorism's historical record? Accurately estimating the probability of such an event is complicated by the large fluctuations in the empirical distribution's upper tail. We present a generic statistical algorithm for making such estimates, which combines semi-parametric models of tail behavior and a nonparametric bootstrap. Applied to a global database of terrorist events, we estimate the worldwide historical probability of observing at least one 9/11-sized or larger event since 1968 to be 11-35%. These results are robust to conditioning on global variations in economic development, domestic versus international events, the type of weapon used and a truncated history that stops at 1998. We then use this procedure to make a data-driven statistical forecast of at least one similar event o
A depth-based rank sum statistic for multivariate data introduced by Liu and Singh [J. Amer. Statist. Assoc. 88 (1993) 252--260] as an extension of the Wilcoxon rank sum statistic for univariate data has been used in multivariate rank tests in quality control and in experimental studies. Those applications, however, are based on a conjectured limiting distribution, provided by Liu and Singh [J. Amer. Statist. Assoc. 88 (1993) 252--260]. The present paper proves the conjecture under general regularity conditions and, therefore, validates various applications of the rank sum statistic in the literature. The paper also shows that the corresponding rank sum tests can be more powerful than Hotelling's T^2 test and some commonly used multivariate rank tests in detecting location-scale changes in multivariate distributions.
We consider the detection of multivariate spatial clusters in the Bernoulli model with $N$ locations, where the design distribution has weakly dependent marginals. The locations are scanned with a rectangular window with sides parallel to the axes and with varying sizes and aspect ratios. Multivariate scan statistics pose a statistical problem due to the multiple testing over many scan windows, as well as a computational problem because statistics have to be evaluated on many windows. This paper introduces methodology that leads to both statistically optimal inference and computationally efficient algorithms. The main difference to the traditional calibration of scan statistics is the concept of grouping scan windows according to their sizes, and then applying different critical values to different groups. It is shown that this calibration of the scan statistic results in optimal inference for spatial clusters on both small scales and on large scales, as well as in the case where the cluster lives on one of the marginals. Methodology is introduced that allows for an efficient approximation of the set of all rectangles while still guaranteeing the statistical optimality results desc
We provide a unifying framework linking two classes of statistics used in two-sample and independence testing: on the one hand, the energy distances and distance covariances from the statistics literature; on the other, maximum mean discrepancies (MMD), that is, distances between embeddings of distributions to reproducing kernel Hilbert spaces (RKHS), as established in machine learning. In the case where the energy distance is computed with a semimetric of negative type, a positive definite kernel, termed distance kernel, may be defined such that the MMD corresponds exactly to the energy distance. Conversely, for any positive definite kernel, we can interpret the MMD as energy distance with respect to some negative-type semimetric. This equivalence readily extends to distance covariance using kernels on the product space. We determine the class of probability distributions for which the test statistics are consistent against all alternatives. Finally, we investigate the performance of the family of distance kernels in two-sample and independence tests: we show in particular that the energy distance most commonly employed in statistics is just one member of a parametric family of ke