Seminar

Modeling the clinical progression of HIV-positive patients in British Columbia (BC) during their last five years of life prior to death

Objective:
We aimed to model the progression of CD4 cell count and the number of emergency department visits during the five years prior to death. Secondarily, we examined the potential factors that may explain the trajectories in these two outcomes.

Methods:
We fitted a linear mixed effect model for CD4 cell count, and we fitted a generalized mixed effect model with Poisson distribution for the number of emergency department visits.

Conclusions:
It was showed that CD4 cell counts of patients declined steeply in the five years prior to death, and negatively associated with the number of emergency department visits. The multivariable models showed the patients had a history of IDU had a significantly worse clinical progression.

My Co-op Experience at Rick Hansen Institute: Access to Care and Timing Study

This talk is to share my 8-month Co-op Experience gained at Rick Hansen Institute. Introduction to the institution and Spinal Cord Injury (SCI) will be first talked about. The main focus will be Access to Care and Timing Project. The study is to describe traumatic SCI care in Canada and develop a computer simulation model that tracks the entire continuum of care in order to evaluate the timeliness and location of care and how they related to both system and patient outcomes. Different types of Generalized Linear Models are built.

Sample of a Co-op Experience: microarray probe set filtering and globin depletion in RNA-Seq data

Science Co-op brought me to work at the PROOF (Prevention of Organ Failure) Centre of Excellence at St. Paul's Hospital, where I mostly performed exploratory assessments of various genomic expression measurement platforms. I will discuss two projects I was involved in. The first is a part of PROOF's computational pipeline of biomarker discovery: probe set filtering. Here, probe sets filtered by a new method, PVAC, are assessed for their resulting signal in comparison to a more established method, FARMS. The second project compares two samples from each of six healthy subjects for their gene expressions from a RNA-Seq assay. We compare samples depleted of globins with non-depleted ones to assess the efficacy of globin depletion on the detection of other gene transcripts.

Diagnostic Proteomic Biomarkers of Acute Kidney Rejection

Acute allograft rejection is an adverse predictor for long-term graft survival in kidney transplantation. The goal of this study is to discover diagnostic biomarkers of acute allograft rejection in kidney transplant patients to identify those of whom have acute rejection (AR) using blood samples. An 13-peptide biomarker panel was identified that can discover acute kidney allograft rejection. The panel was validated in an international cohort of patients. Results hold great potential to transition from proteomic discovery to routine clinical use in order to monitor post-transplant patients.

Likelihood Inference in Spatial Generalized Linear Mixed Models with Multivariate CAR Models for Areal Data

Disease mapping studies have been widely performed with considering only one disease in the estimated models. Simultaneous modeling of different diseases can also be a valuable tool both from the epidemiological and also from the statistical point of view. In particular, when we have several measurements recorded at each spatial location, we need to consider multivariate models in order to handle the dependence among the multivariate components as well as the spatial dependence between locations. These models can be studied in the class of spatial generalized linear mixed models (SGLMMs). It is well known that the frequentist analysis of SGLMMs is computationally difficult. Recently, there are a few papers which explored multivariate spatial models for areal data adopting the Bayesian framework as the natural inferential approach. We use an approach, which yields to maximum likelihood estimation, to conduct frequentist analysis of SGLMMs with multivariate conditional autoregressive (CAR) models for areal data. The performance of the proposed approach is evaluated through a simulation study and also by a real dataset.

Key Words: Disease mapping, hierarchical models, maximum likelihood estimation, spatial statistics

Big Data Analysis for Brain Connectivity Modelling

Speaker's Page

Abstract: High-dimensional datasets, where the number of measured variables is larger than the sample size, are not uncommon in modern real-world applications such as brain connectivity modeling using functional Magnetic Resonance Imaging (fMRI) data. Conventional statistical signal processing tools and mathematical models could fail at handling such high-dimensional problems, and developing efficient algorithms for high-dimensional situations are of great importance.

This talk mainly focuses on the following two issues: (1) recovery of sparse regression coefficients in linear systems, here we focus on the Lasso-type sparse linear regression; (2) estimation of high-dimensional covariance matrix and precision matrix, both subject to additional random noise.

Perspectives on Human Bias versus Machine Bias: Generalized Linear Models

In this talk, we consider estimation problem in generalized linear models when there are many potential predictors and some of them may not have influence on the response of interest. In the context of two competing models where one model includes all predictors and the other restricts variable coefficients to a candidate linear subspace based on subject matter, prior knowledge or auxiliary information.  We investigate the relative performances of Stein type shrinkage, pretest, and penalty estimators (L1GLM, adaptive L1GLM, and SCAD) with respect to the full model maximum likelihood estimator (MLE). The asymptotic properties of the pretest and shrinkage estimators including the derivation of asymptotic distributional biases and risks are established. In particular, we give conditions under which the shrinkage estimators are asymptotically more efficient than the full model MLE. A Monte Carlo simulation study shows that the mean squared error (MSE) of an adaptive shrinkage estimator is comparable to the MSE of the penalty estimators in many situations and in particular performs better than the penalty estimators when the dimension of the restricted parameter space is large. The Steinian shrinkage and penalty estimators all improve substantially on the full model MLE.  Finally, the methodology is evaluated through application to a real data.

Workshop: Visualising data with ggplot2 (FULL - REGISTRATION CLOSED)

This tutorial will introduce you to the theory and practice of ggplot2. I'll introduce you to the rich theory that underlies ggplot2, and then we'll get our hands dirty making graphics to help understand data. I'll also point you towards resources where you can learn more, and highlight some of the other packages that work hand in hand with ggplot2 to make data analysis easy.

You will have the opportunity to practice what you learn, so please bring along your laptop, with the latest version of R installed. Make sure that your version of ggplot2 is up-to-date by running install.packages("ggplot2").

To get the most out of the course, I'd recommend that you're already comfortable with R: you know how to get your data into R, you've done some graphics (base or lattice) in the past, and you've written an R function.

 You can see a movie of this workshop, courtesy of the Pacific Institute of Mathematical Sciences (PIMS).

Forecasting with bootstrap procedures: incorporating the parameter and distribution uncertainties

Speaker's Page

Abstract:  When predicting the future evolution of a given variable, there is an increasing interest in obtaining prediction intervals which incorporate the uncertainty associated with the predictions. Due to their flexibility, bootstrap procedures are often implemented with this purpose. First, they do not rely on severe distributional assumptions. Second, they are able to incorporate the parameter uncertainty. Finally, they are attractive from a computational point of view. Many bootstrap methods proposed in the literature rely on the backward representation which limits their application to models with this representation and reduces their advantages given that it complicates computationally the procedures and could limit the asymptotic results to Gaussian errors. However, this representation is not theoretically needed. Therefore, it is possible to simplify the bootstrap procedures implemented in practice without losing their good properties. The bootstrap procedures to construct prediction intervals that do not rely on the backward representation can be implemented in a very wide range of models without this representation. In particular, they can be implemented in univariate ARMA and GARCH models. Also, extensions to unobserved component models and multivariate VARMA models will be considered. Several applications with simulated and real data will be used to illustrate the procedures.

Bayesian Methods for Alleviating Identification Issues with Applications in Health and Insurance Areas

In areas such as health and insurance, there can be data limitations that may cause an identification problem in statistical modeling. Ignoring the issues may result in bias in statistical inference. Bayesian methods have been proven to be useful in alleviating identification issues by incorporating prior knowledge.

In health areas, the existence of hard-to-reach populations in survey sampling will cause a bias in population estimates of disease prevalence, medical expenditures and health care utilizations. For the three types of measures, we propose four Bayesian models based on binomial, gamma, zero-inflated Poisson and zero-inflated negative binomial distributions. Large-sample limits of the posterior mean and standard deviation are obtained for population estimators. By extensive simulation studies, we demonstrate that the posteriors are converging to their large-sample limits in a manner comparable to that of an identified model. Under the regression context, the existence of hard-to-reach populations will cause a bias in assessing risk factors such as smoking. For the corresponding regression models, we obtain theoretical results on the limiting posteriors. Case studies are conducted on several well-known survey datasets. Our work confirms that sensible results can be obtained using Bayesian inference, despite the nonidentifiability caused by hard-to-reach populations.

In insurance, there are specific issues such as misrepresentation on risk factors that may result in biased estimates of insurance premiums. In particular, for a binary risk factor, the misclassification occurs only in one direction. We propose three insurance prediction models based on Poisson, gamma and Bernoulli distributions to account for the effect. By theoretical studies on the form of posterior distributions and method of moment estimators, we confirm that model identification depends on the distribution of the response. Furthermore, we propose a binary model with the misclassified variable used as a response. Through simulation studies for the four models, we demonstrate that acknowledging the misclassification improves the accuracy in parameter estimation. For road collision modeling, measurement errors in annual traffic volumes may cause an attenuation effect in regression coefficients. We propose two Bayesian models, and theoretically confirm that the gamma model is identified. Simulation studies are conducted for finite sample scenarios.