Seminar

CANSSI National Seminar Series: October speaker, Christopher Jackson

Talk Title: A Comparison of Two Frameworks for Multi-State Modelling, Applied to Outcomes after Hospital Admissions with COVID-19

Time

Seminar: 10:00–11:30am PST
Student Session: 12:00pm–1:00pm PST

Registration & talk details

Learn about and register for this talk and the student session on CANSSI’s website.

CANSSI National Seminar Series

Christopher Jackson is the second speaker in the Canadian Statistical Sciences Institute (CANSSI) National Seminar Series. This fall, the theme of the series is hidden Markov models.

All speakers in the series will provide suggested read-ahead journal articles, as part of a journal club. Christopher’s seminar will include these two journal club papers.

How to Improve Prediction Accuracy in the Analysis of Computer Experiments: Exploitation of Low-Order Effects and Dimensional Analysis

*To join this seminar via Zoom, attendees will need to request connection details from headsec@stat.ubc.ca.

Abstract: Computer codes simulate natural phenomena and engineering processes where physical experimentation is too costly or infeasible. Hence a computer experiment obtains data by running such a code. Nonetheless, the code can be too resource-consuming to run numerous times. Thus we replace a code with a Gaussian Stochastic Process (GaSP) statistical model, as a computationally faster surrogate.  

The talk will outline two strategies to improve prediction accuracy of these surrogates: novel correlation structures for GaSPs based on principles for physical factorial designs, and Dimensional Analysis. It will then focus on Dimensional Analysis, which pays attention to fundamental physical dimensions when modelling scientific and engineering systems. It goes back at least a century but has recently caught statisticians' attention, in the design of physical and computer experiments. The core idea is to analyze dimensionless quantities derived from the original variables and possibly design for them.

Dimensional Analysis has significant challenges in variable selection, which we address with Functional Analysis of Variance. We apply this strategy in various case studies to improve prediction accuracy. Thus we propose new modelling frameworks in computer experiments to accomplish more accurate surrogate models.

Penalized empirical likelihood inference for abundance from capture-recapture data

*To join this seminar via Zoom, attendees will need to request connection details from headsec@stat.ubc.ca.

Post-seminar Q&A: Graduate students are invited to stay after the seminar for a Q&A with the speaker (~12pm12:10pm).

Abstract: Capture-recapture experiments are widely used to collect data needed to estimate the abundance of a closed population. To account for heterogeneity in the capture probabilities, Huggins (1989) and Alho (1990) proposed a semiparametric model in which the capture probabilities are modelled parametrically and the distribution of individual characteristics is left unspecified. A conditional likelihood method was then proposed to obtain point estimates and Wald-type confidence intervals for the abundance. Empirical studies show that the small-sample distribution of the maximum conditional likelihood estimator is strongly skewed to the right, which may produce Wald-type confidence intervals with lower limits that are less than the number of captured individuals or even negative. Furthermore, the conditional-likelihood method may produce spuriously large estimates when a nontrivial number of capture probabilities are close to 0.

In this talk, we present a penalized empirical likelihood approach based on Huggins and Alho's model. We show that the null distribution of the penalized empirical likelihood ratio for the abundance is asymptotically chi-square with one degree of freedom, and the maximum penalized empirical likelihood estimator achieves semiparametric efficiency. We further propose an expectation–maximization algorithm to numerically calculate the proposed point estimate and penalized empirical likelihood ratio function. Simulation studies show that the penalized-empirical-likelihood-based method is superior to the conditional-likelihood-based method: its confidence interval has much better coverage, and the maximum empirical likelihood estimator has a smaller mean square error. 

Point process models for sequence detection in neural spike trains

*To join this seminar via Zoom, attendees will need to request connection details from headsec@stat.ubc.ca.

Post-seminar Q&A: Graduate students are invited to stay after the seminar for a Q&A with the speaker (~12pm12:30pm).

Abstract: Sparse sequences of neural spikes are posited to underlie aspects of working memory, motor production, and learning. Discovering these sequences in an unsupervised manner is a longstanding problem in statistical neuroscience. I will present our new work using NeymanScott processes—a class of doubly stochastic point processes—to model sequences as a set of latent, continuous-time, marked events that produce cascades of neural spikes. This sparse representation of sequences opens new possibilities for spike train modeling. For example, we introduce learnable time warping parameters to model sequences of varying duration, as have been experimentally observed in neural circuits. Bayesian inference in this model requires integrating over the set of latent events, akin to inference in mixture of finite mixture (MFM) models. I will show how recent work on MFMs can be adapted to develop a collapsed Gibbs sampling algorithm for NeymanScott processes. Finally, I will present an empirical assessment of the model and algorithm on spike-train recordings from songbird higher vocal center and rodent hippocampus.

New challenges in phylogenetic inference

*To join this seminar via Zoom, attendees will need to request connection details from headsec@stat.ubc.ca.

Post-seminar Q&A: Graduate students are invited to stay after the seminar for a Q&A with the speaker (~12pm12:30pm).

Abstract: Phylogenetics studies the evolutionary relationships between different organisms, and its main goal is the inference of the Tree of Life. Usual statistical inference techniques like maximum likelihood and Bayesian inference through Markov chain Monte Carlo (MCMC) have been widely used, but their performance deteriorates as the datasets increase in number of genes or number of species. I will present different approaches to improve the scalability of phylogenetic inference: from divide-and-conquer methods based on pseudolikelihood, to computation of Frechet means in BHV space, finally concluding with neural network models to approximate posterior distributions in tree space. The proposed methods will allow scientists to include more species into the Tree of Life, and thus complete a broader picture of evolution.

Bayesian adjustment for preferential testing in estimating infection fatality rates: Theory and methods as motivated by the COVID-19 pandemic

*Note: Attendees will need to request connection details to join this seminar via Zoom. To request Zoom connection details, contact headsec@stat.ubc.ca.

Authors: Harlan Campbell,  Perry de Valpine, Lauren Maxwell, Valentijn M.T. de Jong, Thomas P.A. Debray, Thomas Jaenisch, Paul Gustafson

Abstract: A key challenge in estimating the infection fatality rate (IFR) is determining the total number of cases. The total number of cases is not known because not everyone is tested but also, more importantly, because tested individuals are not representative of the population at large. We refer to the phenomenon whereby infected individuals are more likely to be tested than non-infected individuals as “preferential testing.” An open question is whether or not it is possible to reliably estimate the IFR without any specific knowledge about the degree to which the data are biased by preferential testing. In this paper we take a partial identifiability approach, formulating clearly where deliberate prior assumptions can be made and presenting a Bayesian model, which pools information from different samples. Results of a simulation study suggest that when certain populations with representative testing (i.e., or for which the degree of preferential testing is known) are included in the analysis, identifiability and reliable estimates are attainable. When only limited knowledge is available about the magnitude of preferential testing, reliable estimation of the IFR may still be possible so long as there is sufficient “heterogeneity of bias” across samples. When the model is fit to European data obtained from seroprevalence studies and national official COVID-19 statistics, we estimate the overall COVID-19 IFR for Europe to be 0.47%, 95% C.I. = [0.34%, 0.63%].

CANSSI National Seminar Series: September speaker, Ruth King

We’re pleased to share information on the Canadian Statistical Sciences Institute (CANSSI) National Seminar Series. This fall, there will be three speakers, speaking on the fourth Thursday of the month. The theme is hidden Markov models.

Ruth King is the first speaker in the series.

Like each seminar in the series, Ruth’s seminar will include a read-ahead (but not required) journal article, as part of a journal club.

Registration

To Register: Registration details will be available soon on the CANSSI National Seminar Series page.

Time

Seminar: 10:00am–11:15am PST
Student Session: 11:30am-12:30pm PST

Title & abstract

Title and Abstract: A title and abstract will be available soon on CANSSI's page.

Journal club

Journal Article for this Talk: King, Ruth, Statistical Ecology (January 2014). Annual Review of Statistics and Its Application, Vol. 1, Issue 1, pp. 401-426, 2014. Available at SSRN: https://ssrn.com/abstract=2405891 or http://dx.doi.org/10.1146/annurev-statistics-022513-115633

See the CANSSI National Seminar Series page for more information on the read-ahead journal for this talk and for journal-club funding opportunities for graduate students.

Statistical Inference via Data Science: A ModernDive into R and the Tidyverse

Note: Attendees will need a password to join this seminar via Zoom. To request the meeting password, please contact headsec@stat.ubc.ca.

***

Abstract: We propose a pathway for teaching and learning statistical inference using data science tools widely accepted in industry, academia, and government. We argue that our "data science first" approach to statistical inference is both (1) feasible since the tidyverse is more intuitive for new R users to learn than base R and (2) worthwhile since our proposed pathway exposes students to data science tools applicable beyond the classroom. We first introduce the tidyverse suite of R packages, including ggplot2 for data visualization and dplyr for data wrangling. After equipping students with just enough of these data science tools to perform effective exploratory data analysis, we then guide students through traditional introductory statistics topics like confidence intervals, hypothesis testing, and multiple regression modeling, all while focusing on visualization and the use of real data throughout.

van Eeden seminar: Calcium imaging, clustering, and corncob

**Note: This talk is in an unusual location (the Woodward IRC Building at 2194 Health Sciences Mall). Consider giving yourself some extra time to walk to the building and find the seminar room (Lecture Room 4on the main floor).**

There will be a pre-talk reception in the Woodward IRC lobby at 10:30am.

**********

Abstract: Since this is a student-invited seminar, I'm going to highlight three research projects led by my three senior PhD students. Each project is motivated by a distinct problem in biology.

First, calcium imaging data is transforming the field of neuroscience by making it possible to assay the activities of large numbers of neurons simultaneously. For each neuron, the resulting "fluorescence trace" can be seen as a noisy surrogate of its spikes over time. In order to deconvolve a fluorescence trace into the underlying spike times, we consider an auto-regressive model for calcium dynamics. This leads naturally to a seemingly intractable $\ell_0$ optimization problem. I will show that it is in fact possible to efficiently solve this optimization problem for the global optimum, leading to substantial improvements over competing approaches. I will also talk about quantifying uncertainty associated with these spike estimates.

Second, across many areas of biology, it is becoming increasingly common to collect "multi-view data": that is, data in which multiple data types (e.g. gene expression, DNA sequence, clinical measurements) have been measured on a single set of observations (e.g. patients). I will consider the following question: given a set of n observations with measurements on L data types, can a single clustering of the n observations be defined on all L data types, or does each data type have its own clustering of the observations? To answer this question, I will introduce a general framework for modeling multi-view data, as well as hypothesis tests that can be used in order to characterize the extent to which the clusterings on each of the L data types are the same or different.

Finally, I will consider a fundamental question that arises in the analysis of microbial ecology data: how can we determine whether the abundance of a given taxon differs across conditions?

Sean Jewell, Lucy Gao, and Bryan Martin are 5th year PhD students at University of Washington who carried out the work described in this talk.

**********

This talk is supported by the van Eeden fund, the Department of Statistics, and PIMS.

Adapting black-box machine learning methods for causal inference

I'll discuss the use of observational data to estimate the causal effect of a treatment on an outcome. This task is complicated by the presence of 'confounders' that influence both treatment and outcome, inducing observed associations that are not causal. Causal estimation is achieved by adjusting for this confounding by using observed covariate information. I'll discuss the case where we observe covariates that carry sufficient information for the adjustment, but where explicit models relating treatment, outcome, covariates, and confounding are not available. For example, in medical data the covariates might consist of a large number of convenience health measurements of which only an unknown subset are relevant, and even then in some totally unknown manner. Or, the covariates might be a passage of (natural language) text that describes the relevant information. I'll describe an approach that adapts deep learning and embedding methods to produce representations of the covariate information targeted toward the causal adjustment problem. In particular, I'll describe how to modify standard architectures and training objectives to achieve statistically efficient and practically useful causal estimates.