Seminar

International workshop/PIMS: on the Perspectives on High-dimensional Data Analysis III

Invited Lecturers

This will be the 3rd annual “International Workshop on Perspectives on High-dimensional Data Analysis”, following the first two highly successful workshops entitled: “International Workshop on Perspectives on High-dimensional Data Analysis” (IWPHD) which took place at the Fields Institute, Toronto, 9–11 June 2011 (http://www.fields.utoronto.ca/programs/scientific/10-11/dataanalysis/) and the “International Workshop on Perspectives on High-dimensional Data Analysis II” which was held at the Centre de Recherches Mathématiques, Montréal, 30 May–1 June, 2012 (http://www.crm.umontreal.ca/2012/Perspectives12/index_e.php/11/dataanaly...).


CRM jointly with the American Mathematical society will be publishing an edited volume for the papers presented in this workshop.

The purpose of this workshop is to stimulate research and to foster the interaction of researchers in the area of high-dimensional data analysis in an informal setting. The workshop will provide a venue for participants to meet the field’s leading researchers in a small group setting in order to maximize the chance of interaction and discussion. The objectives include: 1) highlight and expand the breadth of existing methods in high-dimensional data analysis and their potential for the advancement of both mathematical and statistical sciences; 2) identify important directions for future research in the theory of regularization methods, in algorithmic development, and in methodology for different application areas; 3) facilitate collaboration between theoretical and subject-area researchers; and 4) provide the opportunity for highly qualified personnel, including students, to interact with leading researchers from countries around the world. 

The rationale for such a workshop is the continued rapid advancement of modern technology that is allowing scientists to collect data of increasingly unprecedented size and complexity. Examples include epigenomic data, genomic data, proteomic data, high-resolution image data, high frequency financial data, functional and longitudinal data, and network data, among others. Simultaneous variable selection and estimation is one of the key statistical problems in analyzing such complex data. This joint variable selection and estimation problem is one of the most actively researched topics in the current statistical literature. There have been many advances on the variable selection problem for linear and generalized linear regression models in the past decades. However, more recently, regularization, or penalized, methods are becoming increasingly popular and many new developments have been established.

REGISTRATION REQUIRED

A Robust Fit for Generalized Partial Linear Partial Additive Models

In this talk, we propose a robust model fitting algorithm for Generalized Partial Linear Partial Additive Models (GAPLMs), which is a hybrid of the widely-used Generalized Linear Models (GLMs) and Generalized Additive Models (GAMs). The traditional model fitting algorithms are mainly based on likelihood. However, those fits can be severely distorted by the presence of a small portion of atypical observations (also known as "outliers"), which deviate from the assumed model. As a result, the fits become close to those outliers making them not seem atypical. In order to solve this problem, we developed a model fitting algorithm which is resistant to the effect of outliers. To fit the "partial linear partial additive" styled model, our method involves backfitting algorithm and generalized Speckman estimator. To achieve a robust fit, we applied the robust weights derived from robust quasi-likelihood equations proposed by Cantoni and Ronchetti 2001, instead of the likelihood based weights, in generalized local scoring algorithm. To compare the our model fitting performance with the non-robust fit given by the R function gam::gam(), we operated a simulation study and applied the two fitting methods on an example of real dataset. It has been shown in our studies that our robust model fitting algorithm can effectively resist the effect of atypical observations and identify outliers by comparing the robust fitted values with the observed response variable.

Modeling Dependencies in Multivariate Data

In multivariate regression, researchers are interested in modeling a correlated multivariate response variable as a function of covariates. The response of interest can be multidimensional; the correlation between the elements of the multivariate response can be very complex. In many applications, the association between the elements of the multivariate response is typically treated as a nuisance parameter. The focus is on estimating efficiently the regression coefficients, in order to study the average change in the mean response as a function of predictors. However, in many cases, the estimation of the covariance and, where applicable, the temporal dynamics of the multidimensional response is the main interest, such as the case in finance, for example. Moreover, the correct specification of the covariance matrix is important for the efficient estimation of the regression coefficients. These complex models usually involve some parameters that are static and some dynamic. Until recently, the simultaneous estimation of dynamic and static parameters in the same model has been difficult. The introduction of particle MCMC algorithms has allowed for the possibility of considering such models. In this thesis, we propose a general framework for jointly estimating the covariance matrix of multivariate data as well as the regression coefficients. This is done under different settings, for different dimensions and measurement scales.

Switching Nonparametric Regression Models

We propose a methodology to analyze data arising from a curve that, over its domain, switches among J states. We consider a sequence of response variables, where each response y depends on a covariate x according to an unobserved state z. The states form a stochastic process and their possible values are j=1,...,J. If z equals j the expected response of y is one of J unknown smooth functions evaluated at x. We call this model a switching nonparametric regression model. We consider two types of analyses: a single realization case and a replicate case. In the single realization case, we consider one curve switching among J functions. In the replicate case, we have N curves, called replicates, switching between J functions. We develop an EM algorithm to estimate the parameters of the latent state process and the functions corresponding to the J states. We also obtain standard errors for the parameter estimates of the state process. We conduct simulation studies to analyze the frequentist properties of our estimates. We also apply the proposed methodology to two different data sets.

Adaptive learning rates and Bayesian updating in the presence of drift, with applications

Streaming data has become ubiquitous as fast-paced automatic data collection is becoming routine in a variety of settings. Such settings require incremental model updates to handle continual data arrival, and adaptivity to handle temporal variation of possibly unknown characteristics. In this talk we will explore practical alternatives to Bayesian dynamic modelling that make use of forgetting factors and stochastic approximation techniques in order to maximise computational efficiency relying on only weak assumptions about the underlying dynamics. Particular emphasis will be paid to hybrid approaches.Applications of interest include streaming classification and multi-armed bandit problems.

Dr Anagnostopoulos is a Lecturer in Statistics in the Department of Mathematics in Imperial College London, prior to which he was a Research Fellow at the Statistical Laboratory in the University of Cambridge. His research interests are focused on statistical methods for streaming data analysis, including streaming regression, anomaly detection and network analysis. His theoretical interests involve stochastic approximation and state-space modelling. He has consulting experience in e-commerce, retail banking and online advertising, and is an advisor to a UK start-up on statistical software for streaming data analysis.

Jointly Modeling Longitudinal Process with Measurement Errors, Missing Data, and Outliers

Tingting Yu:


In many longitudinal studies, several longitudinal processes may be associated. For example, a time-dependent covariate in a longitudinal model may be measured with errors or have missing data so it needs to be modeled together with the response process in order to address the measurement errors and missing data. In such cases, a joint inference is appealing since it can incorporate information of all processes simultaneously. The joint inference is not only more efficient than separate inferences but it may also avoid possible biases. In addition, longitudinal data often contain outliers, so robust methods for the joint models are necessary. In this talk, we discuss joint models for two correlated longitudinal processes with missing data, measurement errors, and outliers. We consider two-step methods and joint likelihood methods for joint inference, and propose robust methods based on M-estimators to address possible outliers for joint models. Simulation studies are conducted to evaluate the performances of the proposed methods, and a real AIDS dataset is analyzed using the proposed methods.

Given the simple setting of point outcome, treatment and confounding variables, all of which are binary, a double-robust estimator for the average causal effect the can be cast as arising from compromises between parametric and non-parametric outcome models. Inspired by this idea, a Bayesian model averaging estimator is introduced as a weighted average between a Bayesian saturated model and a parametric outcome model.  However, this new Bayesian estimator cannot scale up well in presence of sparse data. We propose a Bayesian hierarchical framework to address this issue.  Several estimation are investigated via simulation studies.

From Geometry of the Limit Set to Extremal Dependence Properties of Light-tailed Distributions

Sample clouds of multivariate data points from light-tailed distributions can often be scaled to converge onto a deterministic set as the sample size tends to infinity. It turns out that the shape of this limit set can be related to a number of extremal dependence properties of the underlying distribution. In this talk, I will present several simple relations, and illustrate how they can be used to replace frequently cumbersome or intractable analytical computations. As an application from the area of finance, I will discuss a new approach to quantifying the effect of risk diversification, which arises from forming a portfolio of risky assets whose behavior is determined by a multivariate probability density. The attention will be restricted to the class of densities whose level sets are all scaled copies of a given set. Exploiting the geometric structure of this common shape we are able to measure the effect of risk aggregation on extremes as well as to quantify the impact of dimension on diversification.

STATS/MATH/PIMS Event: Indian Residential Schools & Their Legacy: Reconciliation through Education

In preparation for the West Coast National Event of the Truth and Reconciliation Commission in Vancouver from September 18-21, the UBC Statistics Department, the UBC Mathematics Department and the Pacific Institute for the Mathematical Sciences (PIMS) have organized an introductory seminar for students, faculty, and staff.

This seminar will present a meaningful introduction to the historical and current realities of the lives of Aboriginal people. This talk will be followed by a short presentation discussing possibilities for individuals to become involved in outreach programs in the mathematical sciences, such as those implemented by PIMS and the Mathematics Department.

Paulette Regan, PhD. is Senior Researcher for the Truth and Reconciliation Commission of Canada. She is an Adjunct Professor in the Faculty of Education at Simon Fraser University and a Research Fellow at the Liu Institute for Global Issues, University of BC. Her book, Unsettling the Settler Within: Indian Residential Schools, Truth Telling and Reconciliation in Canada (UBC Press, 2010) has been a  non-fiction bestseller in BC and was short-listed for the 2012 Canada Prize by the Canadian Federation for the Humanities and Social Sciences.

PIMS Public Lecture - Sparse Linear Models

Speaker's Page

Abstract:  In a statistical world faced with an explosion of data, regularization
has become an important ingredient. In many problems, we have many
more variables than observations, and the lasso penalty and its
hybrids have become increasingly useful. This talk presents a general framework
for fitting large scale regularization paths for a variety of problems. We describe the
approach, and demonstrate it via examples using our R package GLMNET.
We then outline a series of related problems using extensions of these ideas.
*joint work with Jerome Friedman, Rob Tibshirani and Noah Simon.

Recent Progresses of the General Minimum Lower Order Confounding Theory and Its Application in Experimental Designs

To recognize the real world, there are three fundamental sciences as the bases or tools of all
the practical sciences. One is philosophy, one is mathematics and one is statistics. Di erent from
the first two, statistics is to study the general logical thinking and methodology of how to infer
the real world through the data obtained by observing the world, including how to observe the
world and how to analyze the data so that the conclusion obtained by inferring can close to the
real world as ecient and exact as possible. Experimental design is a branch of statistics, which
studies how eciently and economically to observe a real world by planning experiments and how
scientifically to analyze the experimental data.


In this talk, a optimality theory in the field of fractional factorial designs, called general minimum
lower order (GMC) theory, will be introduced, which was developed in recent years (see
Zhang, Li, Zhao and Ai (2008). In the first part, a overview of the GMC theory will be given:
first we introduce some basic points of the GMC theory, including the motivation of the study, the
notion of AENP and GMC criterion, how the GMC to unify the existing criteria and a review of
some other results of the GMC theory obtained in two years ago.


In the second part, some progresses of the GMC theory in these two years will be presented.
The first work is related to how to arrange factors in practical experiments. For a given design, in
order to optimally arrange the factors, we proposed a pattern, called factor aliased e ect number
pattern (F-AENP), for measuring its columns and give a criterion for ranking columns. The FAENP
is used in the two-level GMC designs. The F-AENPs of all the GMC 2n

/lib/FCKuserfiles//A-Talk-about-GMC-Runchu.pdf