Seminar

Sparse Envelop Model: Efficient Estimation and Response Variable Selection in Multivariate Linear RegressionTitle

Response variable selection arises naturally in many applications, but has not been studied as thoroughly as predictor variable selection. In this talk, I will firstly introduce the envelope model which allows efficient estimation in multivariate linear regression. Then I will discuss response variable selection in both the standard multivariate linear regression and the envelope contexts. Finally I will introduce the sparse envelope model we proposed to perform variable selection on the responses and preserve the efficiency gains offered by the envelope model. We establish consistency and the oracle property and obtain the asymptotic distribution of the sparse envelope estimator.


 

Factor Copula Models for Replicated Spatial Data

We propose a new copula model that can be used with replicated spatial data. Unlike the multivariate normal copula, the proposed copula is based on the assumption that a common factor exists and affects the joint dependence of all measurements of the process. Moreover, the proposed copula can model tail dependence and tail asymmetry. The model is parameterized in terms of a covariance function that may be chosen from the many models proposed in the literature, such as the Matern model. For some choice of common factors, the joint copula density is given in closed form and therefore likelihood estimation is very fast. In the general case, one-dimensional numerical integration is needed to calculate the likelihood, but estimation is still reasonably fast even with large data sets. We use simulation studies to show the wide range of dependence structures that can be generated by the proposed model with different choices of common factors. We apply the proposed model to spatial temperature data and compare its performance with some popular geostatistics models.

Robust Pairwise Learning With Kernels

Regularized empirical risk minimization plays an important role in machine learning theory. We will investigate a broad class of regularized pairwise learning (RPL) methods based on kernels. One example is regularized minimization of the error entropy loss which has recently attracted quite some interest from the viewpoint of consistency and learning rates. Another example is machine learning for ranking problems. We show that such RPL methods have additionally good statistical robustness properties, if the loss function and the kernel are chosen appropriately. We treat two cases of particular interest: (i) a bounded and non-convex loss function and (ii) an unbounded convex loss function satisfying a certain Lipschitz type condition. We will also give a result on the qualitative robustness of the empirical bootstrap of RPL methods. This is joint work with Prof. Dr. Ding-Xuan Zhou (City University of Hong Kong). The talk is based on a paper with the title ”Robustness of Regularized Pairwise Learning Methods Based on Kernels” which is accepted by the Journal of Complexity.

Assessing performance of classifiers by cross-validation based on binary data

In statistical applications, we are often asked to construct a classifier based on a random sample from a specific population. Once a classifier is built, we may use it to categorize new individuals from the population. The accuracy of categorizing new individuals is related to the precision of the classifier we built. Yet, the sample from the population is generally noisy. Unless the sample size is very large, the performance of the classifier in terms of correctly classifying new individuals is far from certain. In the data analysis stage, we usually look for the classifier that provides the highest success rate in classifying individuals in the given sample. This classifier's apparent rate of success generally over-estimates its precision when it is applied on new individuals from the population. To overcome this issue, the cross-validation technique is often suggested to be used to assess the performance of a classifier. In this project, we use simulation studies to investigate if the cross-validation technique indeed accurately estimates the performance of classifiers in various situations.
 

Bayesian data science

Speaker's Page

Two arguments for not using Bayesian statistics in a data science context are:

1. Having to wait for several hours of MCMC simulation every time you fix a bug in your model.

2. Worse: having to spend several months implementing finicky MCMC algorithms.

I will talk about the work of my collaborators, students, and myself on trying to make this suffering less severe, using ideas from the disparate fields of statistical mechanics and software engineering.

**Warning:** This will mostly be an unconventional talk: a large chunk will be devoted to introducing "blang", an experimental probabilistic programming language for Bayesian data science we are working on.  **Please bring your laptop with Chrome installed.**

Please also let me know (bouchard@stat.ubc.ca) if you are interested in staying after the talk for continuing with pizza and a more hands-on primer to declarative Bayesian data science with blang.

Models and monitoring designs for spatio-temporal climate data fields

The modelling of temperature fields, which are crucial to understand a region's climate, can be challenging due to the topography of the study region. In the Pacific Northwest, extensive forests, mountains and proximity to the Pacific Ocean may create sudden changes in climate, contributing to the complexity of the modelling of temperature fields in this area. In this talk, we will firstly describe a modelling strategy for complex temperature fields that addresses non-stationarity via a new approach to modelling the spatial mean field.

Secondly, we will focus on the important task of surveillance of environmental processes. We will introduce a novel strategy for the design of monitoring networks where the goal is to choose a high-quality yet diverse set of locations. The idea is brought to this context via the theory of determinantal point processes (DPPs). We will demonstrate how DPPs, which have traditionally been used in other scientific domains, can also play an important role in statistical sciences, particularly in spatial design.

Time permitting, we will also discuss a recent challenge in spatial statistics applications: the data fusion problem. There has been an increased need for combining information from multiple sources that may have been observed on different spatial scales. We will give an overview of an ensemble modelling strategy which combines observed temperature measurements with outputs from an ensemble of deterministic climate models. This methodology can ultimately be used for calibration of model outputs, spatial mapping, and future forecasting.

Inferring Brain Signals Synchronicity from a Sample of EEG Readings

Speaker's Page

Abstract:  Inferring patterns of synchronous brain activity from a heterogeneous sample of electroencephalograms (EEG) is scientifically and methodologically challenging. While it is statistically appealing to rely on readings from more than one individual in order to highlight patterns of coordinated brain activities, pooling information across subjects presents with non trivial methodological problems. We discuss some of the scientific issues associated with the understanding of synchronized neuronal activity and propose a methodological framework for statistical inference from a sample of EEG readings. Our work builds on classical contributions in time-series, cluster and functional data analysis, in an effort to reframe a challenging inferential problem in the context of familiar analytical techniques. Some attention is paid to computational issues, with a proposal based on the hybrid combination of machine learning and Bayesian techniques.

Stochastic Processes, Statistical Inference and Efficient Algorithms for Phylogenetic Inference

Phylogenetic inference aims to reconstruct the evolutionary history of populations or species. With the rapid expansion of genetic data available, statistical methods play an increasingly important role in phylogenetic inference. In this talk, we present new evolutionary models, statistical inference methods and efficient algorithms for reconstructing phylogenetic trees at the level of populations using single nucleotide polymorphism data and at the level of species using multiple sequence alignment data.

 

At the level of populations, we introduce a new inference method to estimate evolutionary distances for any two populations to their most recent common ancestral population using single-nucleotide polymorphism allele frequencies. Our method is based on a new evolutionary model for both drift and fixation. To scale this method to large numbers of populations, we introduce the asymmetric neighbor-joining algorithm, an efficient method for reconstructing rooted bifurcating trees. 

At the level of species, we introduce a continuous time stochastic process, the geometric Poisson indel process, that allows indel rates to vary across sites. We design an efficient algorithm for computing the probability of a given multiple sequence alignment based on our new indel model. We describe a method to construct phylogeny estimates from a fixed alignment using neighbor-joining.

Some mechanisms leading to underdispersion of count data

Speaker's Page

Abstract:  The theory of Poisson-overdispersed count models has been developed in deep, and consequently there are many known "physical mechanisms" leading to overdispersion. For instance, the general families of Mixed Poisson and Compound Poisson distributions are always overdispersed. These physical mechanisms can be interpreted and successfully used for health sciences and biological modelling. There are also some mechanisms leading to underdispersion but they are not very known. In this talk we are going to review some of them and present new  methods and applications. The first mechanism considered is a Poisson-type process where the waiting times are not exponentially distributed. Barlow and Proschan in the 1960s showed that Increasing (Decreasing) Failure Rate distributions for the waiting times produce under(over)-dispersed count distributions. Examples of this mechanism are the models of Winkelmann (1995) using Gamma and Weibull waiting times. The second mechanism is the extended Poisson process of Faddy and Bosch (2001) based on the fact that any count distribution can be represented as a pure birth process with non-constant rates. This representation not always has a simple and meaningful interpretation. The third mechanism is provided by the limiting distribution of a M/M/1 queuing model, where the service time depends of the number of individuals in the queue. An example of this is the original development of the COM-Poisson distribution. This mechanism allows to construct new distributions capable to explain the behaviour of the counts of chromosomal aberrations under high doses of radiation (see Pujol et al., 2014). Finally we will introduce some new mechanisms based on the binomial subsampling operation (p-thinning). It is known that the Poisson distribution is closed under p-thinnings, but if p depends of the number of Poisson realizations the resulting distribution can be underdispersed. Several examples of application will be analyzed and discussed.

References

[1] Faddy, MJ. and Bosch RJ. (2001). Likelihood-Based Modeling and Analysis of Data Underdispersed Relative to the Poisson Distribution. Biometrics, 57, 620-624.

[2] Pujol M., Barquinero JF., Puig P., Puig R., Caballin MR., Barrios L. (2014). A New Model of Biodosimetry to Integrate Low and High Doses. PLoSONE, 9(12):e114137.

[3] Winkelmann, R.(1995). Duration Dependence and Dispersion in Count-Data Models. Journal of Business and Economic Statistics, 13(4), 467-474.