Seminar

Computational methods for mixed effects models in large-scale genetic and genomic studies

This will be a two-part talk describing efficient computational methods for mixed effects models in large-scale genome-wide association studies (GWASs) and RNA sequencing (RNAseq) studies. In the first part of the talk I will focus on linear mixed models. Linear mixed models have attracted considerable attention recently as powerful and effective tools to account for population stratification and relatedness in GWASs. However, existing methods for calculating likelihood ratio test (LRT) statistics are computationally impractical for even moderate-sized GWASs, and many studies have to rely on approximate LRT methods. To address this issue, I present novel computationally-efficient algorithms, which we refer to as genome-wide efficient mixed model association (GEMMA), for fitting both univariate and multivariate linear mixed models, and computing the LRT for SNP associations in GWASs. Our methods improve on existing approximate LRT methods in computation speed, power/correct control of type I error, and ability to deal with more than two phenotypes. I illustrate these features on real and simulated data. In the second part of the talk I will focus on Poisson mixed models. High throughput sequencing is extremely widely used in genetics and genomics, and provides unprecedented insights into many basic biological questions. However, analyzing sequencing data in related individuals presents major statistical and computational challenges, as the read count data generated from a sequencing experiment not only inherit a Poisson noise from the sequencing machine, but also display additional variation across individuals (i.e. over-dispersion). To model this over-dispersion, I propose a Poisson mixed effects model with one random effects term to account for individual relatedness and an error term to account for independent noise. I present a novel and efficient posterior sampling algorithm for the model. With real RNAseq data, I show that our model is more powerful than the widely used negative binomial based approaches in identifying differentially expressed genes.

Integrating multiple types of genomics data to disentangle meaningful associations

The quantity and variety of genomics datasets has increased tremendously in the last decade, presenting novel opportunities both for deriving cellular pathways and networks, and for identifying genetic and cellular mechanisms that underlie disease.  However, interpreting this data to extract biological insights requires disentangling meaningful, and hence reproducible and consequential associations, from mere correlations (i.e. spurious associations).  In this talk, I will present computational and machine learning approaches for leveraging prior biological knowledge, while integrating heterogeneous data, in order to find robust associations.  In particular, I will first describe a scalable method for graph-based integration of diverse types of genomics data, in order to accurately infer functional roles for uncharacterized genes based on a small set of known (training) genes.  This approach results in the state of the art for automatically leveraging the continuous production of new genomics data.  Secondly, focusing on the task of finding associations between genetic variation and cellular (expression) traits in a population-based study, I will present methods for using known confounding factors in order to infer and account for hidden confounding factors.  Thirdly, expanding this task to the context of a specific disease, I will describe a project that combines genotype, RNA-sequencing, and environmental data to find genes and pathway that correlate with disease status.  Application of this approach to a large case/control study of major depression, a highly confounded disorder, sheds new light on molecular mechanisms associated with this pathology.

Optimal adaptive design of experiments for stochastic dynamic systems

Speaker's Page

Abstract:  This talk considers the problem of designing inputs into stochastic experimental systems so that the resulting observations will yield maximal information about parameters of interest. In stochastic systems, maximal information is generally obtained from particular parts of state space and it is therefore advantageous to adapt inputs in response to the system's current state.

We demonstrate that in diffusion processes, the problem of adaptively maximizing Fisher information can be solved through control theoretic techniques and we demonstrate the numerical implementation of these. When only partial and noisy observations of a system are available, we consider imputing the system's state via a filter before applying the optimal policy as described above. Alternatively, maximizing Fisher Information for the partially observed system yields a non-standard control problem and we demonstrate an approximate solution in the context of discrete hidden Markov models.

The approaches described here have applications to numerous experimental systems; we demonstrate the efficacy of input design for parameter estimation through examples in ecology, single neuron experiments and experimental economics.

Semiparametric Longitudinal Model with Irregular Time Autoregressive Error Process

This talk considers semiparametric inference for longitudinal data collected at irregular and possibly subject-specific times. We propose an irregular time autoregressive model for the error process in a partially linear model and develop a unified semiparametric profiling approach to estimating the regression parameters and autoregressive coefficients. An appealing feature of the proposed method is that it can effectively accommodate irregular and  subject-specific observation times.We establish the asymptotic normality of the proposed estimators and derive  explicit forms of their asymptotic variances. For the nonparametric component, we construct  a two stage local polynomial estimator. Our method takes into account the autoregressive error structure and does not drop any observations.

The asymptotic bias and variance of the two stage local polynomial estimator are derived. Simulation studies are conducted to evaluate the finite sample performance of the proposed method. A data example is used to demonstrate its application.

Deciphering autism spectrum disorder: Clues from visual profiles

Autism spectrum disorder (ASD), characterized by difficulties in social communication and interaction, as well as repetitive and restrictive behaviours, encompasses a heterogeneous and varied set of symptoms and levels of severity. The latest revision of the Diagnostic and Statistical Manual of Mental Disorders (DSM-5) included the decision to collapse previous diagnostic groupings into a single umbrella diagnosis of ASD. Although the idea of categorizing individuals on the ASD spectrum into well-circumscribed sub-groups is highly attractive, scientific evidence necessary to make such distinctions is currently lacking. We examined visual function of a group of adults with ASD using a battery of psychophysical tasks. Data showed two distinct clusters based on performance in an orientation discrimination task. One cluster had results consistent with the oblique effect, i.e., superior precision around cardinal axes, compared to oblique angles. The other cluster of ASD participants showed a complete lack of the oblique effect with flat profiles across all orientations. We hypothesize that this clustering may be a reflection of true etiological sub-groups within ASD. To examine potential links between these clusters defined based on visual performance and ASD symptomology we collected the following neuropsychological measures on the same group of adults with ASD (N=19): 1) Wechsler Abbreviated Scale of Intelligence (WASI-II), 2) Autism Quotient (AQ), 3) Multidimensional Social Competence Scale (MSCS), and 4) Autism Diagnostic Observation Schedule (ADOS). A support vector machine pattern classifier was able to correctly predict which cluster an individual belongs to with generalization accuracy of 89.47% (p =0.01) based solely on Full IQ and gender.  In addition, these clusters were also independently identified with 76.92% generalization accuracy based only on a single subscale of the MSCS assessing social motivation (p< 0.05) based on a subset of our ASD group (N=13) who completed this measure. Our results suggest that visual performance profiles provide valuable information in identifying true etiological subgroups within ASD. In addition, they suggest that these visual protocols can serve as potential tools to improve diagnostic specificity.

Semi-automated categorization of open-ended questions

Text data from open-ended questions in surveys are difficult to analyze and are frequently ignored. Yet open-ended questions are important because they do not constrain respondents’ answer choices. Where open-ended questions are necessary, sometimes multiple human coders hand-code answers into one of several categories. At the same time, computer scientists have made impressive advances in text mining that may allow automation of such coding.  Automated algorithms do not achieve an overall accuracy high enough to entirely replace humans.  We categorize open-ended questions using text mining for easy-to-categorize answers and humans for the remainder using expected accuracies to guide the choice of the threshold delineating between “easy” and “hard”. We illustrate this approach with examples from open-ended questions related to respondents’ advice to a patient in a hypothetical dilemma, a follow-up probe related to respondents’ perception of disclosure/ privacy risk, and from a follow-up survey from the Ontario Smoker’s Helpline. Targeting 80% combined accuracy, we found that 59%-80% of the data could be categorized automatically.

Statistics MSc Students' Seminars on Co-op Experience

Jonathan Baik
 

Title:  My Co-op Experience at Environment Canada: Using Synoptic Map Analogs to Predict Weather Events


Abstract:  Due to many advances in numerical weather prediction systems over the past several decades, forecasters have been able to steadily improve the accuracy of their weather predictions. With the increases in computational power and the ability to store massive amounts of data, a method for making weather forecasts that had been abandoned several decades ago has become feasible. I will talk about making weather forecasts using large amounts of historical data, go over the project that I worked on during my co-op, and provide some commentary about my experiences at Environment Canada.

_____________________________________________________________________________________________________________

Shannon Erdelyi

Title: A Co-op Experience

Abstract: In September 2010, British Columbia introduced the Immediate Roadside Prohibition (IRP) program. This program put harsher penalties in place for drunk drivers, and is among the strictest traffic law programs in Canada. Similar penalties, including fines and license suspensions, were also put in place for excessive speeders. The objective of my co-op placement was to evaluate the public health benefits of these new traffic laws. Using interrupted time series methods, I explored changes in motor vehicle collision and public health service utilization rates before and after the IRP intervention. During this seminar, I will discuss the challenges associated with modelling these outcomes. 
 

Student’s t as a new epistemological framework for teaching measurement and uncertainty

Speaker's Page

Abstract:  Many introductory physics labs ask students to conduct experiments to see or experience physics concepts from class first hand. Students collect data from these experiments and are expected to analyze the data to make sense of the physics equations they've learned in class. In first year, however, many of the students have little to no background in statistics. In addition, they enter the first year lab with misconceptions about the nature of measurement, uncertainty, and variability. This provides significant limitations to engaging students with physics concepts and developing experimentation skills. In the first-year honours physics lab at UBC, we have removed the conceptual physics learning goals from the course and replaced them exclusively with goals for learning data analysis and measurement skills. This year in particular, we have introduced the Student's t-test to the course material as a way to engage students in meaningful reflection of their results and to promote iterative experimentation. This talk will present some of these learning goals and new teaching techniques, as well as evidence of students' improved skills over previous iterations of the lab.

Graphical models based on vines: from partial correlation representations of Gaussian models to high-dimensional copulas

Speaker:  Harry Joe

This presentation will consist of introductory material to show partial correlation representations in Gaussian autoregressive, truncated vine, common factor, structured factor, and structural equation models. There is a dependence and graphical parametrization that should have applications in many areas.

Comparisons will be made of various types of graphical models, including Whittaker's Gaussian graphical model based on the inverse correlation matrix, vines with/without latent variables, and path diagrams for structural equation models.  The extension from partial correlation representations to high-dimensional copula models will be indicated. Vines have mainly been used for flexible high-dimensional copula models and lead to new ways of viewing Gaussian dependence models.

The key ideas are the following.

(i) The partial correlation parametrization for Gaussian dependence models.

(ii) The mixing of conditional distributions as the copula extension of partial correlations.

(iii) The substitution of a bivariate copula for each partial correlation combined with the sequential mixing of conditional distributions to get the vine copula extension of multivariate Gaussian, even if there are latent variables.

A Natural Robustification of the Ordinary Instrumental Variables Estimator

Speaker:  Ruben Zamar

Instrumental variables estimators are designed to provide consistent parameter estimates for linear regression models when some covariates are correlated with the error term. We propose a new robust instrumental variables estimator (RIV) which is a natural robustification of the ordinary instrumental variables estimator (OIV). Specifically, we construct RIV using a robust multivariate location and scatter S-estimator to robustify the solution of the estimating equations that define OIV.
RIV is computationally inexpensive and readily available for applications through the R-library riv. It has attractive robustness and asymptotic properties, including:

- high resilience to outliers,
- bounded inuence function,
- consistency under weak distributional assumptions,
- asymptotic normality under mild regularity conditions, and
- equivariance.


We further endow RIV with an iterative algorithm which allows for the estimation of models with endogenous continuous covariates and exogenous dummy covariates. We study the performance of RIV when the data contains outliers using an extensive Monte Carlo simulation study and by applying it to a limited-access dataset from the Framingham Heart Study-Cohort to estimate the effect of long-term systolic blood pressure on left atrial size.


Key words: Endogenous covariate; instrumental variable; robust estimation; S-estimator.


Joint work with

-Gaby Cohen (Statistics, UBC) and
-Hernan Ortiz-Molina (Sauder School of Business, UBC)