Seminar

Likelihood Ratio Test for Multisample Mixture Model and Its Application to Genetic Imprinting

PhD (Probability Theory and Mathematical Statistics, Northeast Normal University, China); BSc (Mathematics and Application Mathematics, Northeast Normal University, China) My research interests are statistics and biostatistics, particularly in the statistical methods base on the mixture model. For biology and genome data, such as human complex diseases in genome wide association studies (GWAS), I have focused on some statistical methods to find the cause of diseases. Based on the genetic data, I have proposed a multi-sample mixture model to identifying genetic imprinting. For the finite mixture model, I am also interested in the large sample theory, which includes the consistency of MLEs and the limiting distribution of LRT statistics. In addition, I have focused on mixture model with auxiliary information.

Likelihood Ratio Test for Multisample Mixture Model and Its Application to Genetic Imprinting

Show Abstract

Genomic imprinting is a known aspect of the etiology of many diseases. The imprinting phenomenon depicts differential expression levels of the allele depending on its parental origin. When the parental origin is unknown, the expression level has a ?nite normal mixture distribution. In such applications, a random sample of expression levels consists of three subsamples according to the number of minor alleles an individual possesses, of which one is the mixture and the other two are homogeneous. This understanding leads to a likelihood ratio test (LRT) for the presence of imprinting. Because of the nonregularity of the ?nite mixture model, the classical asymptotic conclusions on likelihood-based inference are not applicable. We show that the maximum likelihood estimator of the mixing distribution remains consistent. More interestingly, thanks to the homogeneous subsamples, the LRT statistic has an elegant and rather distinct 0.5X1^2 + 0.5X2^2 null limiting distribution. Simulation studies con?rm that the limiting distribution provides precise approximations of the ?nite sample distributions under various parameter settings. The LRT is applied to expression data. Our analyses provide evidence for imprinting for a number of isoform expressions.

Bayesian False Discovery Rate Controlling Methodology

Control over the FDR has become an essential tool in the analysis of large data sets. In my talk, I will explain the relation between selective inference, simultaneity, post-hoc inference, Bayesian discrimination rules and control over the FDR. I will also show how the Benjamini-Hochberg FDR controlling procedure can be extended to construct confidence intervals for selected parameters and illustrate in simulations the connection between frequentist and Bayesian control over the FDR.

Database of Religious History: A Statistically Analyzable Database of History

The idea that there may be patterns in history goes back to at least the Ancient Greeks. Looking back at the stories of their ancestors and later, written history, humans have been apt to see recurrences in the rise and fall of empires, the machinations of politics and war, and the failures and fortunes of leaders, ideas, and cultures. Thus far, these speculations have remained causal storytelling, making it difficult to determine if these patterns are a byproduct of our evolved pattern-seeking brains or true trajectories through time and space. The Database of Religious History is an ambitious effort to address these puzzles by compiling historical data in a systematic and open-access format and using formal mathematical and statistical models to extract the broad patterns. There are many challenges to designing a statistically-analyzable, human-readable, humanities database of knowledge. From a technical perspective, such a system needs to be able to handle hundreds of variables, millions of data points and potentially millions of users. From a user perspective, it needs to be (a) easy to enter data for experts from history, anthropology, and archeology, and (b) easy to search, visualize, and analyze the data for analysts from these fields, as well as psychology, evolutionary biology, statistics, and other interested fields. I’ll discuss the technical and human hurdles in creating the system, show you our design, infrastructure, and data, and explain how our team of scientists and historians build and test theories.

The generalized lasso with non-linear measurements

Consider signal estimation from non-linear measurements.  A rough heuristic often used in practice postulates that "non-linear measurements may be treated as noisy linear measurements" and the signal may be reconstructed accordingly.  We give a rigorous backing to this idea, with a focus on low-dimensional signals buried in a high-dimensional spaces.  Just as noise may be diminished by projecting onto the lower dimensional space, the error from modeling non-linear measurements with linear measurements will be greatly reduced when using the signal structure in the reconstruction.  We assume a random Gaussian model for the measurement matrix, but allow the rows to have an unknown, and ill-conditioned, covariance matrix.  As a special case of our results, we give theoretical accuracy guarantee for 1-bit compressed sensing with unknown covariance matrix of the measurement vectors.

Quantum Computation and Statistics

Quantum computation and quantum information are of great current interest in computer science, mathematics, physical sciences and engineering. They will likely lead to a new wave of technological innovations in communication, computation and cryptography. As the theory of quantum physics is fundamentally stochastic, randomness and uncertainty are deeply rooted in quantum computation, quantum information and quantum simulation. Thus statistics can play an important role in quantum computation and quantum simulation, which in turn offer great potential to revolutionize computational statistics. This talk will first give a brief introduction on quantum computation and then present my recent work on quantum tomography via compressed sensing as well as statistical modeling and analysis of quantum computing experimental data.

Nonparametric Likelihood and Its Applications

Junjian Zhang

Dr. Junjian Zhang is a professor at Guangxi Normal University, Guilin China. He received his PhD in 2006 at Academy of Mathematics and Systems Science, Chinese Academy of Sciences under the supervision of Professor Guoying Li. Following his PhD, he conducted post-doctoral research at Beijing University of Technology under supervision of Professor Zhongzhan Zhang. He is currently visiting the department of statistics, UBC hosted by Dr. Jiahua Chen for the period of 2015. His research interests include mathematical statistics and its applications, especially in nonparametric likelihood ratio. His research is supported by the following funds: National Social Science Foundation of China, National Natural Science Foundation of China, Guangxi Science Foundation. He was the winner of the “best paper award of Zhong Jiaqing” at the 10th Jingjin Wusi youth meeting and the “best paper” award at the 8th Guangxi statistical science colloquium. He was invited speaker in many research conferences including the united meeting of Hunan,Guangdong and Guangxi mathematical societies, the Tenth National Congress of Chinese Society of Probability and Statistics. His email address is jjzhang@mailbox.gxnu.edu.cn.

Nonparametric Likelihood and Its Applications

Show Abstract

Nonparametric likelihood is one of the important topics in statistics. In this talk, we will introduce the basic ideas for nonparametric likelihood and present our latest research achievements. For example, we generalize the empirical likelihood to the empirical Lq likelihood and Empirical power divergent likelihood. The former is usually used to the estimating theory, the latter is usually used to the goodness-of-fit. This talk will focus on the goodness-of-fit. In addition, the talk will discuss the adjusted empirical (Euclidean) likelihood and its applications, the nonparametric likelihood for the complex data such as the rounded data, dependent data, high-dimensional data, and so on.

Projecting the uncertainty of sea level rise using climate models and statistical downscaling

Most global climate models do not estimate sea level directly. A semi-empirical approach is to relate sea level change to temperature, and then apply this relationship to climate model projections of temperature for different future scenarios. Another possibility is to estimate the relationship between global mean temperature in historical runs of a model, and instead apply this relationship to future temperature projections. We compare these two methods to estimate global annual mean sea level, and assess the resulting uncertainty.

Of more practical importance is to estimate local sea level. We exemplify this by developing models for projected sea level rise in Vancouver and Washington State, and illustrate different sources of uncertainty in the projections.

Information on the Constance van Eeden fund can be found here.

Please click here to see a video of and slides from Professor Guttorp's March 24th talk!

Co-op Experiences

Jack Ni (11am--11:30am)

Stress testing is a tool to understand the stability of individual financial institutions and the financial system. In 2014, a capital stress test was issued to provincially regulated financial institutions in BC. As probability of default is an integral component of stress testing, my main task was to design a probability of default model for residential mortgages in BC. During this seminar, I will go over the projects I was involved in during my co-op and discuss my experience working at FICOM.

Yifan Zhang (11:30am--12pm)

Bill shock refers to customers’ reaction when they see their unexpectedly high monthly bill amount. Telecommunication companies receive an enormous amount of calls from these shocked customers and spend millions of credits every year. The major objective of my first Co-op placement at Centre for Operations Excellence (COE) in Sauder was to help a telecommunication company predict the calling customers. My second Co-op placement at SPPH involved examining health care (mainly home care) interventions among adult chronic kidney disease (CKD) patients. In this seminar, I will talk about these two projects that I worked on, and my experiences at Sauder and SPPH.

The Moving Holiday Model’s Design and Application - Taking the Chinese Spring Festival for Example

Seasonal adjustment is a very important step for economic data
preprocessing. Holiday adjustment is an inevitable step for the popular
seasonal adjustment methods which include X-12-ARIMA and TRAMO/SEATS.
Because different countries have different kinds of holidays, the
popular seasonal adjustment methods must be modified when they are used
in different countries. For China, the Spring Festival is a very
important and comparatively long holiday, and it occurs in January in
some years and in February in other years. On the base of the X-12-ARIMA
method, taking the Chinese Spring Festival for example, this paper
designs different kinds of moving holiday models with considering the
effects of holiday’s influence as well as spans on economic data. By
selecting different economic indicators and by using software Demetra
and EViews, this paper tests the performance of these different models.
In the end, the best adjustment models are derived based on the criteria
of outlier percentage reduction.

Air Quality Model Evaluation through the Analysis and Modelling of Ozone Features

Legislative actions regarding ozone pollution use air quality models (AQMs) such as Community Multiscale Air Quality (CMAQ) model for scientific guidance, hence the evaluation of AQM such as CMAQ is an important subject. Traditional point-to-point comparisons between AQM outputs and ozone observations can be uninformative or even misleading since the AQM modelled ozone process and physical observations are governed by different stochastic spatial processes. I propose an alternative model evaluation approach that is based on the comparison of spatial-temporal ozone features, where I compare the dominant space-time structures between AQM and observation. To successfully implement feature-based AQM evaluation, I further developed statistical framework of analyzing and modelling space-time ozone fields using ozone features. Rather than working directly with raw data, I analyze the spatial-temporal variability of ozone fields by extracting data features using Principal Component Analysis (PCA). These features are then modelled as Gaussian Processes (GPs) driven by various atmospheric conditions and chemical precursor pollution. My method is implemented on CMAQ outputs during several ozone episodes in Lower Fraser Valley, BC. I found that the feature-based ozone model is an efficient way of modelling and forecasting a complex space-time ozone field. The framework of ozone feature analysis is then applied to evaluate CMAQ outputs against the observations. Here, I found that CMAQ persistently over-estimates the observed spatial ozone pollution. Through the modelling of ``feature differences'', I identified their associations with CMAQ inputs on ozone precursor emissions, and the CMAQ-observation differences are focused on regions where the pollution process transitions from NOx-sensitive to VOC-sensitive. Through the comparison of dynamic ozone features, I found that CMAQ's over-prediction is also connect to the model producing higher than observed ozone plume in daytime. However, CMAQ model did capture the observed space-time pattern of diurnal ozone advection. Lastly, individual modelling of CMAQ and observed ozone features revealed that even under the same atmospheric conditions, CMAQ tends to significantly over-estimate the ozone pollution during the early morning. In the end, I demonstrated that the feature-based AQM evaluation methods developed in this research are able to provide ``big picture'' process-level understandings of AQM deficiency.