Seminar

The highs and lows of clustering single-cell biological data.

Unsupervised learning by clustering is a key tool for analyzing single-cell biological data. StormGraph is a clustering algorithm that we have developed to analyze protein clustering in single cells imaged using super-resolution microscopy. I will introduce the StormGraph algorithm and discuss its application to studying the nanoscale biology of Diffuse Large B-Cell Lymphoma (DLBCL), a clinically heterogeneous, aggressive blood cancer. I will also briefly discuss how we are looking to study heterogeneity in DLBCL tumours by clustering high-throughput, high-dimensional single-cell proteomic data, with implications for personalized medicine.

A Bayesian Group Sparse Multi-Task Regression Model for Imaging Genetics

Recent advances in technology for brain imaging and high-throughput genotyping have motivated studies examining the influence of genetic variation on brain structure. In this setting, high-dimensional regression for multi-SNP association analysis is challenging as the response variables obtained through brain imaging comprise potentially interlinked endophenotypes, and there is a desire to incorporate a biological group structure among SNPs based on their genetic arrangement. We consider a recently developed approach for the analysis of imaging genetic studies based on penalized regression with regularization based on a group l_{2,1}-norm penalty which encourages sparsity at both the gene and SNP level. While incorporating a number of useful features, a shortcoming of the proposed approach is that it only furnishes a point estimate and techniques for obtaining valid standard errors or interval estimates are not provided. We solve this problem by developing a corresponding Bayesian formulation based on a three-level hierarchical model that allows for full posterior inference using Gibbs sampling. Techniques for the selection of tuning parameters are investigated thoroughly and we make comparisons between cross-validation, fully Bayes, and empirical Bayes approaches for the choice of tuning parameters. Our proposed methodology is investigated using simulation studies and is applied to the analysis of a large dataset collected as part of the Alzheimer's Disease Neuroimaging Initiative. I will discuss how our MCMC algorithm scales with an increasing number of SNPs, imaging phenotypes, and subjects, and I will also describe extensions of the model for application to brain-wide data and the corresponding development of a spatial model that is currently in progress.

Adequacy-of-fit diagnostics for multivariate copulas with parsimonious dependence structures

Parsimonious dependence structures are crucial in multivariate modelling as they offer better interpretation and may improve the quality of estimation. However, one must also be careful not to use an overly parsimonious model that results in underfitting. Model selection criteria like AIC and BIC do not typically provide inslight on the quality of model fit. We propose the use of adequacy-of-fit statistics for this purpose. Pairwise differences between empirical and model-based features are aggregated and the resulting statistic is compared against a cutoff value, the exceedance of which suggests model inadequacy. In this talk, we will explore the challenges in obtaining an appropriate cutoff and some pragmatic approaches we put forward. The techniques are applied to a financial data set with dependent returns.

DIABLO - an integrative, multi-omics, multivariate method for multi-group classification

Rapid advances in technology have led to a wealth of large-scale molecular omics datasets. Integrating such data offers an unprecedented opportunity to assess molecular interactions at multiple functional levels and provide a more comprehensive understanding of the biological pathways involved in different diseases subgroups. However, multiple omics data integration is a challenging task due to the heterogeneity in the different platforms used. There is a need to address the complex and correlated nature of different data-types, in order to identify a robust and reliable multi-omics signature that can predict a phenotype of interest. We introduce a novel multivariate dimension reduction method for multiple omics integration, classification and identification of a multi-omics molecular signature. DIABLO - Data Integration Analysis for Biomarker discovery using a Latent component method for Omics studies, models the correlation structure between omics datasets, resulting in an improved ability to associate biomarkers across multiple functional levels to phenotypes of interest. We demonstrate the capabilities of DIABLO using simulated data and studies of breast cancer and asthma, integrating up to four types of omics datasets to identify relevant biomarkers, while still retaining competitive classification and predictive performance compared to existing methods. Our statistical integrative framework can benefit a diverse range of research areas with varying types of study designs, as well as enabling module-based analyses. Importantly, graphical outputs of our method assist in the interpretation of such complex analyses and provide significant biological insights.

Sparse robust regression estimators

Abstract:In many current applications scientists can easily measure a very large number of variables (for example, several thousands of gene expression levels) some of which are expected be useful to explain or predict a specific response variable of interest. These potential explanatory variables are most likely to contain redundant or irrelevant information, and in many cases, their quality and reliability may be suspect. We developed a penalized robust regression estimator that can be used to identify a useful subset of explanatory variables to predict the response, while protecting the resulting estimator against possible aberrant observations in the data set. Using an Elastic Net penalty, the proposed estimator can be used to select variables, even in cases with more variables than observations or when many of the candidate explanatory variables are correlated. In this talk, I will present the new estimator and an algorithm to compute it. I will also illustrate its performance in a simulation study. This is joint work with Professor Matias Salibian-Barrera and our student David Kepplinger.

The Gene-environment independence assumption in the analysis of case-control data

We consider the problem of exploiting the gene-environment independence assumption in a case-control study inferring the joint effect of genotype and environmental exposure on disease risk.

We first take a detour and develop the constrained maximum likelihood estimation theory for parameters arising from a partially identified model, where some parameters of the model may only be identified through constraints imposed by additional assumptions. We show that, under certain conditions, the constrained maximum likelihood estimator exists and locally maximizes the likelihood func- tion subject to constraints. Moreover, we study the asymptotic distribution of the estimator and propose a numerical algorithm for estimating parameters.

Next, we use the frequentist approach to analyze case-control data under the gene-environment independence assumption. By transforming the problem into a constrained maximum likelihood estimation problem, we are able to derive the asymptotic distribution of the estimator in a closed form. We then show that exploiting the gene-environment independence assumption indeed improves estimation efficiency. Also, we propose an easy-to-implement numerical algorithm for finding estimates in practice.

Furthermore, we approach the problem in a Bayesian framework. By introducing a different parameterization of the underlying model for case-control data, we are able to define a prior structure reflecting the gene-environment independence assumption and develop an efficient numerical algorithm for the computation of the posterior distribution. The proposed Bayesian method is further generalized to address the concern about the validity of the gene-environment independence assumption.

Finally, we consider a special variant of the standard case-control design, the case-only design, and study the analysis of case-only data under the gene-environment independence assumption and the rare disease assumption. We show that the Bayesian method for analyzing case-control data is readily applicable for the analysis of case-only data, allowing the flexibility of incorporating different prior beliefs on disease prevalence.

From Genetics to Epigenetics to CRISPR Gene Editing with Machine Learning

Abstract: Molecular biology, healthcare and medicine have been slowly morphing
into large-scale, data driven sciences dependent on machine learning and
applied statistics. In this talk I will start by explaining some of the
modelling challenges in finding the genetic underpinnings of disease, which is
important for screening, treatment, drug development, and basic biological
insight. Genome and epigenome-wide associations, wherein individual or sets of
(epi)genetic markers are systematically scanned for association with disease
are one window into disease processes. Naively, these associations can be
found by use of a simple statistical test. However, a wide variety of
structure and confounders lie hidden in the data, leading to both spurious and
missed associations if not properly addressed. Most of this talk will focus on
how to model these types of data. Once we uncover genetic causes, genome
editing—which is about deleting or changing parts of the genetic code—will one
day let us fix the genome in a bespoke manner. Editing will also help us to
understand mechanisms of disease, enable precision medicine and drug
development, to name just a few more important applications. I will close by
discussing how we are using machine learning to enable more effective CRISPR
gene editing.

Bio: Jennifer Listgarten is a Senior Researcher at Microsoft Research New
England, located in Cambridge, MA. She took a long and winding road to find her
current area of interest in computational biology, starting off with an
undergraduate degree in Physics, followed by a Master’s in Computer Vision
before completing a Ph.D. in Machine Learning at the University of Toronto.
Her current focus is in machine learning and applied statistics with
application to problems in biology. She works on both methods development and
applications enabling new insights into basic biology and medicine. Particular
areas of focus have included CRISPR guide design, statistical genetics,
immunoinformatics, liquid-chromatography proteomics, and microarray analysis.
You can find out more about her work here: http://www.jennifer.listgarten.com.

Semiparametric monitoring test based on clustered data

Joint work with Pengfei Li and  Yukun Liu

Part of the Forestry Product Project headed by Jim Zidek


*****

Due to factors such as climate change, forest fire and plague of insects, it is essential to update lumber quality monitoring procedures in American Society for Testing and Materials (ASTM) Standard D1990 (adopted in 1991) from time to time. A key component of monitoring is an effective method for detecting the change in lower percentiles of the solid lumber strength based on observations of multiple samples. Yet currently used statistical methods do not meet the modern standard of high sensitivity and reliability, particularly when the data are clustered.

In a recent study by Verrill et al. (2015), eight statistical tests proposed by wood scientists were examined thoroughly based on real and simulated data sets. They are found unsatisfactory for various reasons such as seriously inflated false alarm rate when observations are clustered, suboptimal power properties, or having ad hoc rejection regions. The ultimate reason behind the unsatisfactory performance is that most of these tests are developed for purposes other than detecting the changes in quantile.

This paper proposes a method directly targeting changes in quantile based on a composite empirical likelihood. The approach effectively combines the information from multiple samples via a semi-parametric model assumption. It satisfactorily controls the type I error through a cluster-based bootstrapping procedure whether or not the data are correlated. The performance of the test is examined through simulation experiments and a real data example. The new method is generic, not confined to the motivating example.

2 UBC Statistics Students' MSc Presentations

11am - 11:30am

Speaker:  Creagh Briercliffe, current Statistics PhD student, presenting his UBC Statistics MSc work
 

Title: Poisson Process Infinite Relational Model: a Bayesian nonparametric model for transactional data

Abstract: Transactional data consists of instantaneously occurring observations made on ordered pairs of entities. Visually, it can be represented as a networkor more specifically, a directed multigraphwith edges possessing unique timestamps. In this talk, I present work from my master's thesis, which explores a Bayesian nonparametric model for discovering latent class-structure in transactional data. Furthermore, by pooling information within clusters of entities, this model can be used to infer the underlying dynamics of the time-series data. 
This talk will cover the following: (i) examples of transactional data, (ii) details of the Bayesian model and approximate inference scheme, (iii) procedures used to validate computational correctness and evaluate the model's performance, and (iv) an example of the model applied to real data from historical records of militarized disputes between nations.

--------------------------------------------------------------------------

11:30am - 12:00pm

Speaker:  Jeff Bone, Statistics MSc student

Title:  Instantaneous Dynamics of Functional Data

Abstract:  Time dynamic systems can be used in many applications to data modeling. In the case of longitudinal data, the dynamics of the underlying differential equation can often be inferred under minimal assumptions via smoothing based procedures.

In many cases, one wants to learn the dynamics of a differential equation that incorporates more than just one stochastic process. We propose extensions to existing two-step smoothing methods that allow for the presence of additional functional data arising from a second stochastic process. We further introduce model comparison techniques to assess the hypothesis that there is a significant change in fit provided by this additional process. These techniques are applied to the instantaneous dynamics of mouse growth data and allow us to make comparisons between mice who have been assigned different genetic and physical conditions.