Seminar

Much ado about equivalence testing

In order to determine whether or not an effect/association is absent based on a statistical test, the recommended frequentist tool is the equivalence test. In this talk, we review three recent and related papers. First, we introduce non-inferiority tests (one-sided equivalence tests) for ANOVA and linear regression analyses that correspond to standard F-tests for ?^2 and R2 (https://arxiv.org/abs/1905.11875). Second, with the aim of motivating a critical discussion as to what is acceptable and desirable in the interpretation of equivalence tests, we consider a series of controversial scenarios in which the equivalence margin is defined post-hoc (https://arxiv.org/abs/1807.03413).  Finally, we propose an alternative publication policy similar to "Registered Reports" that incorporates equivalence testing in order to strategically reduce publication bias (https://doi.org/10.1371/journal.pone.0195145).

Extreme Value Approach to CoVaR Estimation

The global financial crisis of 2007 – 2009 revealed the importance of systemic risk: the risk that may destabilize the global economy due to financial contagion. Accurate assessment of systemic risk would not only enable regulators to introduce suitable policies to mitigate the risk, but also allow individual institutions to monitor and mitigate their vulnerability. An effective measurement of systemic risk should be able to capture the co-movements between a financial system (or market) and individual financial institutions. One popular measure of systemic risk is CoVaR. In this talk, a methodology is proposed to compute dynamic forecasts of CoVaR semi-parametrically within the classical framework of multivariate extreme value theory (EVT). According to the definition, CoVaR can be viewed as a high quantile of a conditional distribution where the conditioning event corresponds to large losses of an institution. The idea of our methodology is to relate this conditional distribution to the tail dependence function. We develop an EVT-based framework to estimate CoVaR statically by combining parametric modelling of the tail dependence function to address the issue of data sparsity in the joint tail regions and semi-parametric univariate tail estimation techniques. The performance of the methodology is illustrated via simulation studies and real data examples.

D-vine copulas for repeated measurements

We propose a model for unbalanced longitudinal data, where the univariate margins can be selected arbitrarily and the dependence structure is described with the help of a D-vine copula. We show that our approach is an extremely flexible extension of the widely used linear mixed model if the correlation is homogeneous over the considered individuals. As an alternative to joint maximum likelihood a sequential estimation approach for the D-vine copula is provided and validated in a simulation study. The model can handle missing values without being forced to discard data. Since conditional distributions are known analytically, we easily make predictions for future events. For model selection, we adjust the Bayesian information criterion to our situation. In an application to heart surgery data our model performs clearly better than competing linear mixed models.

Reference: Killiches, Matthias, and Claudia Czado. "A D-vine copula-based model for repeated measurements extending linear mixed models with homogeneous correlation structure." Biometrics 74.3 (2018): 997-1005.

Looking for Anomalies in Correlated Time Series

We present a proximity-based anomaly detection system designed to look for anomalies in multiple linearly correlated time series that can have trends. Our system is primarily designed to work in an unsupervised setting to detect contextual anomalies. Anomalies are data points whose values are considered extreme in relation to all other values observed at a particular point in time. Our system uses a two-step procedure to flag anomalous values. It first builds some representative time series, using simple descriptive statistics, of the underlying trend common to all time series in the data. Models are found for the representative time series in training sets with no anomalies. When the models are applied to test data, some days can be flagged as having anomalies. Then methods are used to find the anomalous values in the flagged days. The methodology is modified using robust measures of spread when there is no training set with no anomalies. The end result is a system that accurately classifies anomalies offline (in historical data sets) and is easy to and extend to detect anomalies online (in real-time).

Monitoring test under the density ratio model for unequal sized clustered data

Monitoring the change in lower percentiles of the modulus of rupture and modulus of elasticity is important in the forest industry. Often, the lumber data are clustered and observations within the same cluster are potentially correlated. Such a correlation structure leads to increased false alarm rates shown in the eight statistical tests investigated by Verrill et al. (2015). In forestry or many other industries, such clustered data are often collected independently from multiple populations. To taking advantage of the similarity of their population distributions and avoid the risk of model specification, Chen et al. (2016) recommended the use of a density ratio model and analyzing the data via composite empirical likelihood. They further propose a cluster-based bootstrapping method to do away with the effect of the correlation structures in the data. The cluster-based bootstrap procedure permits easy construction of confidence intervals and carrying out tests for various hypotheses.

Chen et al. (2016) developed theory and carried out data experiments for clustered data from multiple populations of equal cluster size. It is foreseeable that their subsequent data analysis methods remain effective when the data contain clusters of unequal sizes, after some minor adjustments. In fact, if the majority of clusters are of the same cluster size, no changes seem necessary. When data contain various cluster sizes, the theory of Chen et al. (2016) must be re-examined and their methods may have different asymptotic properties. If their methods are applied with only minor changes, the confidence intervals may have too high or too low coverage probabilities, and the hypothesis tests may have inaccurate sizes.

In this project, we go over the methods and theory of Chen et al. (2016) when the data contain different cluster sizes. We conclude some changes are indeed necessary to ensure the asymptotic validity of these inference methods. We propose first to create a new cluster of nearly equal cluster sized from the original clustered data and analyze the new clustered data by the methods in Chen et al. (2016). We discuss the asymptotic properties of the new procedures. We establish the consistency of the new estimator and show that the maximum composite EL quantile estimator still has Bahadur representation. We use simulation to obtain the average mean square errors of the estimators and coverage probabilities of two-sided 95% confidence intervals. The simulation results show that the AMSE of the composite EL quantiles are less than those of the empirical quantiles, and the bootstrap confidence intervals close to nominal 95% only when we have large sample sizes.

Two UBC Statistics M.Sc. Student Presentations

***

Speaker: Xinmiao Wang

Title: Nonlinear Mixed Effects Models with Missing Time-dependent Covariates

Abstract:  HIV viral dynamic models have been used to describe the virus elimination and production process during antiviral treatments. These models have received great attention in the literature and have been shown to perform well in many AIDS studies. Nonlinear mixed effects (NLME) models have been proposed for modelling HIV viral dynamic. However, missing data in time-dependent covariates often lead to challenges in statistical analysis of data for HIV viral dynamics. We propose a new multiple imputation method to deal with the missing time-dependent covariates problems in NLME models. The proposed method is used to analyze a real HIV dataset and compared to the naïve complete case method, and somewhat different conclusions are obtained based on the proposed multiple imputation method.

***

Speaker: Xinzhe Dong

Title: A Multiple Imputation Method for Missing Responses in NLME Models

Abstract: Missing data frequently arise in longitudinal studies. There has been extensive research in this area. However, further research is still needed for certain models for longitudinal data with missing values. Understanding the missing data and handling them properly is essential for the validity of statistical inference. This project focuses on one of the missing data methods, the multiple imputation method, and extends this method to impute the missing responses in nonlinear mixed effects models. Simulations are conducted to evaluate the performance of the proposed method. An AIDS study is presented as an example, in which the viral loads are modeled by an HIV viral dynamic model.

Uncovering the Hidden Universe of Rental Units in Surrey

Current information on the distribution and numbers of rental units in the City of Surrey is incomplete. This hampers the City's ability to accurately plan for community services, schools, transportation, and other infrastructure. Without a full accounting of rental units, the City also does not know the actual vacancy rate, which has policy implications in terms of affordable housing supply. The goal of this project is to assemble information (e.g., create a database and visualizations) from multiple sources and datasets to construct as complete a picture as possible. Equipped with this data, the City will review and shape policy development on a number of fronts; including parking management policies, secondary suite legalization campaigns, and affordable housing policies that specifically address the supply on rental housing.

UBC Statistics M.Sc. Co-op Student Presentations

***

Speaker: Kathy Zhao

Title: Tips for working in labs as a statistician

Abstract: Co-op is an excellent opportunity to gain valuable working experience. From the 8-month co-op training, I improved on skills of bash scripting, paper reading, and data visualization. In this presentation, I will introduce the working architecture in Dr. Daley's lab and some hospital/lab protocols to those who are interested in labs/hospitals. Also, a short description of my co-op projects will be presented as well, including a comparison between bisulfite sequencing aligners, and data visualization for the Gera project. This presentation will focus on more of the thinking logic at work, rather than statistical methods.

***

Speaker: Yidie Feng

Title: Clustering large-scale single cell RNAseq data

Abstract: Accurate clustering of single cell RNAseq data has been a challenge due to noise and technical complications. Dr. Prabhakar's research team has published an algorithm called Reference Component Analysis (RCA), which managed to cluster single cell RNAseq data accurately. However, as the scale of single cell RNAseq data increases, there is a need to upgrade RCA algorithm and make it scalable to large datasets but still accurate at the same time, which is what I worked on during my Co-op. This presentation will give an overview of RCA algorithm and my Co-op experience.

Two UBC Statistics Co-op M.Sc. Student Presentations

Speaker: Qiaoyue Tang, UBC Satistics Co-op M.Sc. Student

Title: Classification of rare cancer subtypes

Abstract: In many biomedical data, the goal is to predict outcomes such as disease subtypes given a set of features such as patients’ characteristics, RNA expression, images, etc. These data are usually imbalanced where the interest lies in properly classifying minority outcomes represented by rare subtypes. High-dimensionality in these data often makes the problem more difficult. In this talk, I will discuss common sampling strategies and algorithm-level approaches to handle classification tasks with the imbalanced condition.

***

Speaker: Jonathan Steif, UBC Satistics Co-op M.Sc. Student

Title: Histone Modifications in Diverse Human Cell-Types: An Integrative Analysis of 70 ChIP-seq Datasets

Abstract: Chromatin immunoprecipitation followed by high-throughput sequencing (ChIP-seq) is a relatively new advancement in the study of protein-DNA binding. High costs associated with deep sequencing are still a limiting factor for most researchers, leading to published datasets with low numbers of biological replicates and often no technical replicates. These inadequate sample sizes have left several rudimentary questions relating to the consistency and variability of protein-DNA binding unanswered. Namely, are certain genomic regions consistently marked by proteins in samples of the same disease or cell-type? Are these marked genomic regions observed across samples of different cell-types? To what extent does the choice of signal detection algorithm (peak-caller) bias results?

This integrated analysis of 70 datasets available via the International Human Epigenome Consortium employs machine learning techniques to examine the variability in the enrichment of six histone modifications in both cancerous and normal human tissues. Traditional peak-calling algorithms are limited in that they are incapable of handling multiple ChiP-seq samples as inputs. The statistical assumptions behind a range of popular algorithms are critiqued, motivating the need for new algorithms better suited for large-scale analyses.

Infectious disease transmission and genomic data: who infected whom?

I will describe a Bayesian approach to reconstructing who infected whom with the help of pathogen genetic data. As sequencing technologies have dramatically declined in cost, it is now feasible to sequence large numbers of viral or bacterial genomes in infectious disease outbreaks, and there have been high hopes that the resulting DNA or RNA sequences will tell the story of who infects whom and when, leading to both better infectious disease control and a better understanding of pathogen evolution. However, sequences do not directly reveal who infected whom, and leave considerable uncertainty. I will outline our main approach with extensions to include multiple datasets and to handle covariates. Finally I will describe the limitations, and several open challenges in this area.