Seminar

The Impact of HCV Co-Infection on Healthcare-Related Utilization among HIV Patients in British Columbia, Canada; Instrumental Variables Selection: a Comparison between Regularization and Post-Regularization Methods

Talk by Huiting Ma (11am - 11:30am)

Title: The Impact of HCV Co-Infection on Healthcare-Related Utilization among HIV Patients in British Columbia, Canada

Abstract:  Over the past several decades, the analysis of healthcare-related count data has been increasingly recognized as an essential research topic for many researchers. There is a long history of analyzing counts data under the framework of parametric distributions for independent identically distributed random variables. Standard methodology for analysing this type of data falls under generalized linear models, which include Poisson regression models and negative binomial regression models. There are other alternative models, such as quasi-Poisson and zero-inflated models. Analysis of longitudinal count data, which contain repeated count observations for each subject over time, is often done by fitting generalized estimating equation models and generalized mixed effects models.

In this project, our main objective was to characterize the trends in healthcare utilization (i.e. the number of healthcare related visits) of HIV mono-infected individuals and HIV/hepatitis C (HCV) co-infected individuals, in British Columbia, from April 1, 1997 to March 31, 2010. We use several statistical methods for the analysis of healthcare-related cross-sectional and longitudinal count data. Understanding the healthcare burden among HIV mono-infected and HIV/HCV co-infected individuals has the potential to help stakeholders to identify and address the unique healthcare needs of these individuals. Our data analyses results show that individuals with an HIV/HCV co-infection status were at a risk of experiencing higher rates of healthcare-related visits than HIV mono-infected individuals.

Talk by Chiara Di Gravio (11:30am - 12pm)

Title: Instrumental Variables Selection: a Comparison between Regularization and Post-Regularization Methods

Abstract: Instrumental variables are commonly used in statistics, econometrics, and epidemiology for the estimation of causal effects when controlled experiments are not available. Specifically, instrumental variables estimators provide consistent parameter estimates in regression models when some of the predictors are correlated with the error term. However, the properties of these estimators are sensitive to the choice of valid instruments. Since in many applications, valid instruments come in a bigger set that includes also weak and possibly irrelevant instruments, the researcher needs to select a smaller subset of instruments that are relevant and strongly correlated with the predictors in the model.

In this project we review the instrumental variables estimators, discuss the problems related to having instruments that are either weak or possibly irrelevant instruments, and compare already existing techniques with new approaches. Simulation studies will be presented to compare the different methods.

Vine Regression

Roger Cooke is currently Chauncey Starr Senior Fellow for Risk Analysis, Resources for the Future (Washington DC), a position he has held since 2005. Roger Cooke's career has mainly been in Delft, the hometown of Constance van Eeden. He is still active (as Professor Emeritus) at the Delft University of Technology, (Faculty of Electrical Engineering, Mathematics and Computer Science) where he had built a graduate program on Risk Analysis. Information about him, including his CV, can be found at the above link. He is recognized as one of the world's leading authorities on mathematical modeling of risk and uncertainty. He has published two well known books in this area: "Experts in Uncertainty; Opinion and Subjective Probability in Science" in 1991, and "Probabilistic Risk Analysis: Foundations and Methods" with Bedford in 2001. Recent papers include uncertainty analysis in impacts of climate change. In addition, Roger Cooke is the inventor of the 'vine' graphical object, and this has influenced greatly the developments in the past 10 years in the area of high-dimensional multivariate non-Gaussian models and copulas.

Structured Expert Judgment and Invasive Species in the Great Lakes

Roger Cooke is currently Chauncey Starr Senior Fellow for Risk Analysis, Resources for the Future (Washington DC), a position he has held since 2005. Roger Cooke's career has mainly been in Delft, the hometown of Constance van Eeden. He is still active (as Professor Emeritus) at the Delft University of Technology, (Faculty of Electrical Engineering, Mathematics and Computer Science), where he had built a graduate program on Risk Analysis. Information about him, including his CV, can be found at the above link. He is recognized as one of the world's leading authorities on mathematical modeling of risk and uncertainty. He has published two well known books in this area: "Experts in Uncertainty; Opinion and Subjective Probability in Science" in 1991, and "Probabilistic Risk Analysis: Foundations and Methods" with Bedford in 2001. Recent papers include uncertainty analysis in impacts of climate change. In addition, Roger Cooke is the inventor of the 'vine' graphical object, and this has influenced greatly the developments in the past 10 years in the area of high-dimensional multivariate non-Gaussian models and copulas.

Communicating with Data Using Simplified Models and Uncertainty Quantification

In the first part of the talk we explore ways to predict mortality in critical care situations, such as ICUs. Current models use a small number of variables, no temporal features, and are regression based with manual variable selection and weighting. We develop a univariate flagging algorithm (UFA) that predicts well, scales to a large number of variables, is robust to missing data, and easy to interpret and visualize.  While Random Forests, etc. can be competitive with UFA in these situations, they are a black box to the practitioners using them.
 
In the second part we consider methods to quantify potential uncertainty in plots and images.The basic idea is to find a way to remove structure from the image, bootstrap what is left, and then restore the structure leading to, say, and 1000 images.The Earth Mover’s Distance allows us to compute the distances between these plots and optimization algorithms allow us to order the plots and then find, say the lower extreme, middle, and upper extreme to visualize the uncertainty that may be present in the plot or image.

Advances in fitting statistical models to huge datasets

In the first part, I will consider the problem of minimizing a finite sum
of smooth functions. This is a ubiquitous computational problem in
statistics, as it frequently arises in various maximum likelihood and
regularized maximum likelihood frameworks. I will describe the stochastic
average gradient algorithm which, despite over 60 years of work on
stochastic gradient algorithms, is the first method to achieve the low
iteration cost of stochastic gradient methods while achieving a linear
convergence rate as in deterministic gradient methods that process the
entire dataset on every iteration.

In the second part, I will consider the even-more-specialized case where we
have a linearly-parameterized model (such as linear least squares or
logistic regression). I will talk about how coordinate descent methods,
though a terrible idea for minimizing general functions, are theoretically
and empirically well-suited to solving such problems. I will also discuss
how we can design clever coordinate selection rules, that are much more
efficient than the classic cyclic and randomized choices.

*Bio: * Mark Schmidt has been an assistant professor in the Department of
Computer Science at the University of British Columbia since 2014. His
research focuses on developing faster algorithms for large-scale machine
learning, with an emphasis on methods with provable convergence rates and
that can be applied to structured prediction problems. From 2011 through
2013 he worked at the École normale supérieure in Paris on inexact and
stochastic convex optimization methods. He finished his M.Sc. in 2005 at
the University of Alberta working as part of the Brain Tumor Analysis
Project, and his Ph.D. in 2010 at the University of British Columbia
working with Kevin Murphy on graphical model structure learning with
L1-regularization. He has also worked at Siemens Medical Solutions on heart
motion abnormality detection, with Michael Friedlander in the Scientific
Computing Laboratory at the University of British Columbia on
semi-stochastic optimization methods, and with Anoop Sarkar at Simon Fraser
University on large-scale training of natural language models.

Modelling tuberculosis and HIV in South Africa

Tuberculosis (TB) continues to afflict millions of people and causes over a million deaths a year worldwide, especially in countries with a high burden of HIV. Multi-drug resistance is also on the rise, causing concern among public-health experts. This talk will give an overview of my work on modelling TB and HIV in South Africa. I will present a Bayesian approach to jointly modelling the TB and HIV epidemics, where a compartmental model allowed us to understand the drivers of the epidemics and compare the effects of TB and HIV interventions. I will also describe a project on TB evolution, where a new metric combined with an optimization-based approach resulted in an accurate classification of complex infections as originating from mutation or mixed infection, as well as the identification of the strains composing these complex infections.

Phylogenetic clustering with a Markov-modulated Poisson process

Most of the information that public health agencies and researchers have about emerging disease epidemics is obtained by on-the-ground epidemiology, that is, by asking infected people about where they went and who they contacted. Phylodynamics is an emerging area of research which aims to enrich this information using viral genomic data and bioinformatic methods. One application of phylodynamics is the identification of groups of epidemiologically related individuals, termed phylogenetic clustering. Here, we develop a novel clustering method which uses a Markov-modulated Poisson process, applied to a "family tree" of viruses, to identify parts of the population experiencing elevated transmission rates. We applied this method to anonymised viral genomic data sampled from almost 8000 HIV-infected individuals in British Columbia.   


Rosemary McCloskey is a second year Master's student in the CIHR Strategic Training Program in Bioinformatics at UBC, working in Dr. Art Poon's lab at the BC Centre for Excellence in HIV/AIDS. She has an undergraduate degree in Mathematics from Simon Fraser University, and is the inaugural recipient of the Statistics Department Award in Data Science.

Supervised principal components regression using a Cox-LASSO model

Diffuse Large B-Cell Lymphoma (DLBCL) is an aggressive cancer of the white blood cells, and its causes are not well understood. Pathologists hope to discover a molecular signature that is predictive of the disease's survival after adjusting for other features that are known to be important. A classical Cox Proportional Hazards model is inappropriate because the DLBCL data is high dimensional and many genomic features suffer from multicollinearity. Thus, we used a "Cox-LASSO" method to select a relevant subset of features correlated with survival. Instead of using all the features in the regression from the LASSO model, we predict using the first principal component (PC). The first PC is constructed by adapting Bair & Tibshirani's supervised PC regression method. This approach ensures a reduction in the dimensionality of the covariate space addressing the collinearity typically observed in the data. The prediction performance of the resulted model is evaluated by cross validation. This talk describes analyses performed in a joint collaboration with the Centre for Lymphoid Cancer at BC Cancer Agency.

Generalized Dynamic Principal Components

Brillinger defined dynamic principal components (DPC) for time series based on a reconstruction criterion. He gave a very elegant theoretical solution and proposed an estimator which is consistent under stationarity. Here we propose a new enterally empirical approach to DPC. The main differences with the existing methods -mainly Brillinger procedure- are (i) the DPC we propose need not be a linear combination of the observations and (ii) it can be based on a variety of loss functions including robust ones. Unlike Brillinger, we do not establish any consistency results; however, contrary to Brillinger’s, which has a very strong stationarity flavor, our concept aims at a better adaptation to possible nonstationary features of the series. We also present a robust version of our procedure that allows to estimate the DPC when the series have outlier contamination. We give iterative algorithms to compute the proposed procedures that can be used with a large number of variables. Our non robust and robust procedures are illustrated with real data sets. Joint work with Dr. Victor Yohai.

Assessing Prediction Error in Survival Times: Application to Ovarian Cancer

Random survival forests (RSFs) are a tool for predicting survival times. Forests are grown using a set of (possibly censored) survival times with associated predictor variables. One challenge surrounding the use of RSFs is the assessment of prediction error; mean squared error (MSE), the commonly used measure of prediction error in the traditional random forest setting, is not an option if some observations are censored. An alternative measure of prediction error that has been suggested in the literature is the C-index, which summarizes the concordance between the observed and predicted values and can incorporate some censored cases. However, in this talk, we show that this measure of prediction error can behave in undesirable ways (and quite differently than MSE). We then present some alternative ways of thinking about prediction error in the context of predicting recurrence time of ovarian cancer.