Seminar

Two UBC Statistics MSc student presentations (Kevin Chern & Matteo Lepur)

To Join Via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Presentation 1

Time: 11:00am - 11:30am

Speaker: Kevin Chern, UBC Statistics MSc student

Title: On estimating marginal likelihoods via Monte Carlo methods based on distribution continua

Abstract: Computing the marginal likelihood (ML) of a model is an integral part of the Bayesian paradigm for model comparison. In practice, computing the ML equates to evaluating a high-dimensional integral, rendering standard numerical methods impractical. Consequently, more sophisticated alternatives such as Monte Carlo (MC) methods are often employed instead. State-of-the-art MC methods include Parallel Tempering (PT) and Sequential Monte Carlo (SMC), both of which leverage a sequence of tempered distributions to improve the efficiency of estimators. PT and SMC have been studied extensively but independently, leaving a gap in the literature for comparing the two. Furthermore, it is unclear how one should allocate a fixed computational budget between two parameters in SMC: the length of the sequence of distributions and the number of particles. While adaptive SMC algorithms exist for automatically determining sequences of distributions, runtimes are typically random and unknown a priori to inference, further complicating the budget allocation problem.

In an attempt to address these challenges, we benchmark the performances between PT and SMC on 14 statistical models including common models from the physics and biology literature. Our first contribution provides practical recommendations for selecting estimators given characteristics of a model, e.g., multimodal, discrete, or continuous targets. Second, we propose a strategy for approximating the length of the sequence of distributions for an adaptive SMC algorithm. By employing this adaptive SMC algorithm in tandem with an accurate approximation of the random runtime, we provide practical recommendations for allocating budgets in SMC.

Presentation 2

Time: 11:30am - 12:00pm

Speaker: Matteo Lepur, UBC Statistics MSc student

Title: A Bayesian Nonparametric Model for Pan-Cancer Analysis

Abstract: Cancers from different tissue types can share a latent structure reflecting commonly altered gene pathways. It is difficult to cluster cancer patients based on this latent structure because the tissue of origin often dominates the latent structure effect. We propose a Bayesian nonparametric model that accounts for the tissue effect and clusters based on a latent structure using a Dirichlet Process prior. Our approach learns the tissue effect by using tissue parameters in a supervised learning setting, while simultaneously learning the latent structure based on the resulting residuals in an unsupervised setting. We demonstrate our model by showing results on synthetic data, semi-synthetic data, and a publicly available dataset from the International Cancer Genome Consortium (ICGC).

Improved conditional generative adversarial networks for image generation: methods and their application in knowledge distillation

To Join Via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: Conditional generative adversarial networks (cGANs) are state-of-the-art models for synthesizing images dependent on some conditions. These conditions are usually categorical variables such as class labels. cGANs with class labels as conditions are also known as class-conditional GANs. Some modern class-conditional GANs such as BigGAN can even generate photo-realistic images. The success of cGANs has been shown in various applications. However, two weaknesses of cGANs still exist. First, image generation conditional on continuous, scalar variables (termed regression labels) has never been studied. Second, low-quality fake images still appear frequently in image synthesis with state-of-the-art cGANs, especially when training data are limited. This thesis aims to resolve the above two weaknesses of cGANs and explore the applications of cGANs in improving a lightweight model with the knowledge from a heavyweight model (i.e., knowledge distillation).

First, existing empirical losses and label input mechanisms of cGANs are not suitable for regression labels, making cGANs fail to synthesize images conditional on regression labels. To solve this problem, this thesis proposes the continuous conditional generative adversarial network (CcGAN), including novel empirical losses and label input mechanisms.

Moreover, even the state-of-the-art cGANs may produce low-quality images, so a subsampling method to drop these images is necessary. In this thesis, we propose a density ratio based subsampling framework for unconditional GANs. Then, we introduce its extension to the conditional image synthesis setting called cDRE-F-cSP+RS, which can effectively improve the image quality of both class-conditional GANs and CcGAN.

Finally, we propose a unified knowledge distillation framework called cGAN-KD suitable for both image classification and regression (with a scalar response), where the synthetic data generated from class-conditional GANs and CcGAN are used to transfer knowledge from a teacher net to a student net, and cDRE-F-cSP+RS is applied to filter out bad-quality images. Compared with existing methods, cGAN-KD has many advantages, and it achieves state-of-the-art performance in both image classification and regression tasks.

Two UBC Statistics MSc student presentations (Gian Carlo Di-Luvi & Kenny Chiu)

To Join Via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Presentation 1

Time: 11am – 11:30am

Speaker: Gian Carlo Di-Luvi, UBC Statistics MSc student

Title: Locally-Adaptive Boosting Variational Inference

Abstract: Boosting variational inference (BVI) approximates Bayesian posterior densities by iteratively building a mixture of component distributions. However, BVI requires greedily optimizing the next component—an optimization problem that becomes increasingly computationally expensive as more components are added to the mixture. Furthermore, previous work has only used simple (i.e., Gaussian) component distributions; in practice, many of these components are needed to obtain a reasonable approximation. These shortcomings can be addressed by considering components that adapt to the target density. However, natural choices such as MCMC chains do not have tractable densities and thus require a density-free divergence for training.

As a first contribution, we show that the kernelized Stein discrepancy—which to the best of our knowledge is the only density-free divergence feasible for VI—cannot detect when an approximation is missing modes of the target density. Hence, it is not suitable for boosting components with intractable densities. As a second contribution, we develop locally-adaptive boosting variational inference (LBVI), in which each component distribution is a Sequential Monte Carlo (SMC) sampler, i.e., a tempered version of the posterior initialized at a given simple reference distribution. Instead of greedily optimizing the next component, we greedily choose to add components to the mixture and perturb their adaptivity, thereby causing them to locally converge to the target density; this results in refined approximations with considerably fewer components. Moreover, because SMC components have tractable density estimates, LBVI can be used with common divergences (such as the Kullback–Leibler divergence) for model learning. Experiments show that, when compared to previous BVI methods, LBVI produces reliable inference with fewer components and in less computation time.

Presentation 2

Time: 11:30am – 12pm

Speaker: Kenny Chiu, UBC Statistics MSc student

Title: On the Statistical Properties of Entromin as an Orthogonal Rotation Criterion

Abstract: The goal in factor analysis is to uncover a set of latent factors that can explain the variation in the data. Principal Component Analysis is one approach that estimates the factors by a set of principal components. However, it may be difficult to interpret the factors as-is, and so it is common to rotate the estimated factors to make their coefficients as sparse as possible to improve interpretability. Varimax is the most popular method for factor rotations, and its statistical properties have been studied in recent literature. Entromin is another factor rotation method that is less commonly used and not as well-studied, but there exists conventional wisdom that Entromin generally finds sparser rotations compared to Varimax.

In this thesis, we aim to explain the sparsity claim for Entromin by studying its statistical properties. Our main contributions include several theoretical results that take steps towards this aim. We show that Varimax is a first-order approximation of Entromin, and that generalizing this connection leads to a family of Entromin approximations. We derive the conditions under which the second-order approximation is expected to identify the true factors in a latent factor model. We then make the connection between optimizing the Entromin objective and recovering sparsity in the factors. Other contributions of this thesis include novel connections to statistical concepts that have not been made in the literature to our knowledge, and an empirical study of Entromin on a real dataset.

The augmented maximum likelihood estimation, and testing for homogeneity under finite vector-parameter mixture models

To Join Via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca


Abstract: It is well-known that the MLE fails under some finite mixture models because their likelihood function is unbounded. This unboundedness occurs, for instance, under the finite normal mixture model, the finite gamma mixture model, and the finite location-scale mixture model. I study a novel way to modify the likelihood function based on data augmentation. This modified likelihood function produces a consistent estimator, the augmented MLE, under those finite mixture models with unbounded likelihood functions. In some circumstances, the augmented MLE is more efficient than its competitors in the literature. 


Hypothesis testing for homogeneity under finite mixture models assesses the hypotheses that data are collected from a non-mixture distribution (Null) or a two-subpopulation mixture distribution. I develop an EM test and a C(a) test under finite vector-parameter mixture models for this purpose.
 

The Neutral-to-the-Left Mixture Model

To Join Via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: A useful step in data analysis is clustering, in which observations are grouped together in a hopefully meaningful way. The mainstay model for Bayesian Nonparametric clustering is the Dirichlet Process Mixture Model, which has one key advantage of inferring the number of clusters automatically. However, the Dirichlet Process Mixture Model is not perfect, and there is further research to be done into other Bayesian Nonparametric models that address the weaknesses of the Dirichlet Process Mixture Model while maintaining automatic inference of the number of clusters.

In this thesis, we introduce the Neutral-to-the-Left Mixture Model, a family of Bayesian Nonparametric infinite mixture models which serves as a strict generalization of the Dirichlet Process Mixture Model. This family of mixture models has two key parameters: the distribution of arrival times of new clusters, and the parameters of the distribution of the stick breaking representation of this model, whose customization allows the analyst to inject prior beliefs regarding the structure of the clusters into the model. We describe sampling algorithms to infer the posterior distribution of clusterings given data for the model. We consider one particular parameterization of the Neutral-to-the-Left Mixture Model with characteristics that are distinct from the Dirichlet Process Mixture Model, evaluate its performance on simulated data, and compare these to results from a Dirichlet Process Mixture Model. Finally, we apply one parameterization of the Neutral-to-the-Left Mixture Model to cluster Twitter datasets to reveal temporal evolution of tweets.

Two UBC Statistics MSc student presentations

To join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Presentation 1

Time: 11am – 11:30am

Speaker: Sherry Gao, UBC Statistics MSc student

Title: Nonlinear mixed-effects models for HIV viral load trajectories before and after antiretroviral therapy interruption, incorporating left censoring

Abstract: In an HIV study, the viral decay during an anti-HIV treatment and the viral rebound after the treatment is interrupted can be viewed as two longitudinal processes, and they may be related to each other. Our goal is to investigate if key features of HIV viral decay and CD4 trajectories during antiretroviral therapy (ART) are associated with characteristics of HIV viral rebound following ART interruption. Nonlinear mixed-effects (NLME) models are used to model viral load trajectories before and following ART interruption, incorporating left censoring due to lower detection limits of viral load assays. A stochastic approximation EM (SAEM) algorithm is used for parameter estimation and inference. To circumvent the computational intensity associated with maximizing the joint likelihood, we propose an easy-to-implement three-step method. We evaluate the performance of this method through simulation studies and apply it to data from the Zurich Primary HIV Infection Study. We find that some key features of viral load and CD4 trajectories during ART (e.g., viral decay rate) are significantly associated with important characteristics of viral rebound following ART interruption (e.g., viral set point).

Presentation 2

Time: 11:30am – 12pm

Speaker: Ian Murphy, UBC Statistics MSc student

Title: Modelling dive phase definitions for Northern Resident Killer Whales

Abstract: Northern Resident Killer Whales (NRKWs), in contrast with the endangered Southern Resident Killer Whales (SRKWs), have been thriving in their habitats. A key component to understanding whale survival is to identify prey capture events, but they are difficult to directly observe. Instead, kinematic variables during the bottom phase of a dive are used to predict prey captures. However, universal definitions of the bottom phase have not been established, despite the fact that modifying the bottom phase greatly impacts existing methods to predict prey capture events. To investigate bottom phase variability, we asked several whale researchers to identify the bottom phase of various dives. The diving data used were collected from 3 NRKWs by a UBC whale researcher. Linear mixed-effects models show that there exists substantial variation in bottom phase definitions across different researchers and across different dive types. We then propose several statistical models for the bottom phase of a dive, including functional linear regression models. Identification of the bottom phase using these models improves the prediction of prey capture dives compared to the currently used bottom phase definitions. Finally, we formulate methods to determine an adequate sample size for fitting these statistical models, and then apply the methods to the data.

Movement data reveal dynamic social relationships

To Join Via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Post-Seminar Virtual Lunch and Journal Club: Graduate students are invited to stay after the seminar to eat lunch and participate in an animal-movement journal club with the speaker. The recommended paper for discussion is Dynamic social networks based on movement by Scharf et al.

Abstract: Satellite-based tracking devices allow researchers to collect increasingly rich data for a wide variety of animals. These data often provide a reasonably inexpensive source of information about the behaviour of not just one, but several individuals. Data used to define social connectivity are often expensive to collect and based on case-specific, ad hoc criteria. Moreover, in applications involving animal social networks, collection of these data is often opportunistic and can be invasive. Frequently, social relationships influence the way individuals move. Thus telemetry data, which are minimally invasive and relatively inexpensive to collect, present an alternative source of information for learning about animal social structure. I describe two recent model-based approaches for inferring dynamic social relationships from movement data and demonstrate the methods using remotely sensed locations of killer whales and sandhill cranes. I conclude with current and emerging challenges to the analysis of movement data for multiple individuals and some preliminary successes in attempts to address those challenges.

Joint colloquium with the University of British Columbia and the University of Washington

Link to Zoom Webinar: https://washington.zoom.us/j/93146867819

Webinar ID: 931 4686 7819

UBC Speaker

Trevor Campbell, Assistant Professor
Department of Statistics, UBC

UW Speaker

Alex Luedtke, Assistant Professor
Department of Statistics, UW

Title: Parallel Tempering on Optimized Paths

Title: Using Deep Adversarial Learning to Construct Optimal Statistical Procedures

Abstract: Parallel tempering (PT) is a class of Markov chain Monte Carlo algorithms that constructs a path of distributions annealing between a tractable reference and an intractable target, and then interchanges states along the path to improve mixing in the target. The performance of PT depends on how quickly a sample from the reference distribution makes its way to the target, which in turn depends on the particular path of annealing distributions. However, past work on PT has used only simple paths constructed from convex combinations of the reference and target log-densities. In this talk I'll show that this path performs poorly in the common setting where the reference and target are nearly mutually singular. To address this issue, I'll present an extension of the PT framework to general families of paths, formulate the choice of path as an optimization problem that admits tractable gradient estimates, and present a flexible new family of spline interpolation paths for use in practice. Theoretical and empirical results will demonstrate that the proposed methodology breaks previously established upper performance limits for traditional paths.

Abstract: Traditionally, statistical procedures have been derived via analytic calculations whose validity often relies on sample size growing to infinity. We use tools from deep learning to develop a new approach, adversarial Monte Carlo meta-learning, for constructing optimal statistical procedures. Statistical problems are framed as two-player games in which Nature adversarially selects a distribution that makes it difficult for a Statistician to answer the scientific question using data drawn from this distribution. The players’ strategies are parameterized via neural networks, and optimal play is learned by modifying the network weights over many repetitions of the game. In numerical experiments and data examples, this approach performs favorably compared to standard practice in point estimation, individual-level predictions, and interval estimation, without requiring specialized statistical knowledge.
 

Bio: Trevor Campbell is an Assistant Professor of Statistics at the University of British Columbia. His research focuses on automated, scalable Bayesian inference algorithms, Bayesian nonparametrics, streaming data, and Bayesian theory. He was previously a Postdoctoral Associate advised by Tamara Broderick in the Computer Science and Artificial Intelligence Laboratory (CSAIL) and Institute for Data, Systems, and Society (IDSS) at MIT, a Ph.D. candidate under Jonathan How in the Laboratory for Information and Decision Systems (LIDS) at MIT, and before that he was in the Engineering Science program at the University of Toronto.

 

Bio: Alex Luedtke is an Assistant Professor in the Department of Statistics at the University of Washington, with an affiliate appointment in the Vaccine and Infectious Disease Division at the Fred Hutchinson Cancer Research Center. He received his Sc.B. in Applied Math from Brown University in 2012 and completed his Ph.D. in Biostatistics at University of California, Berkley in 2016, under the supervision of Mark van der Laan. His research interests involve quantifying the uncertainty for population-level effects while making minimal assumptions for how the data were generated. He works with both clinical trial data and observational data—when working with observational data, he applies methods from causal inference to elicit assumptions under which a causal effect of an intervention can be estimated.

Two UBC Statistics MSc student presentations

To Join Via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Presentation 1

Time: 11am – 11:30am

Speaker: Lulu Pei, UBC Statistics MSc student

Title: An assessment of the robustness of nonlinear mixed effects models to covariance structure specification

Abstract: Nonlinear mixed effects (NLME) models are widely used with applications in both HIV/AIDS studies and pharmacokinetic/pharmacodynamic studies. In HIV trials, NLME models can be used to describe the virus elimination and production process for a patient on antiretroviral therapy. The estimated viral dynamic parameters can then be used to evaluate the efficacies of the treatments. In practice, there are often substantial variations in viral load and CD4 cell count measurements among patients, and the viral dynamic parameters may vary greatly across patients. Mixed effects models appear appealing in such cases since random effects can be used to characterize individual deviations from population averages. In the case of viral load, it is evident that both the variations of the within-individual repeated measurements and the variations between individuals increase over time. As such, there is a need to consider appropriate specification of covariance structures beyond the default constant variance and independence between repeated measurements over time assumed by common NLME implementation software such as R. We will explore various covariance structures for the within-individual repeated measurements and the between-individual random effects to see if analysis results are sensitive to structure specification.

Presentation 2


Time: 11:30am – 12pm

Speaker: Lily Xia, UBC Statistics MSc student

Title: Quantifying the utility of personalized treatment decision rules: Extending and comparing two metrics for summarizing the heterogeneity of treatment effects

Abstract: The treatment benefit prediction model is a type of clinical prediction model that quantifies the magnitude of treatment benefit given an individual's unique characteristics. As the topic of treatment effect modelling is relatively new, quantifying and summarizing the performance of treatment benefit models are not well studied. The “concordance-statistic for benefit” and the “concentration of benefit index” are two newly developed metrics that evaluate the discriminative ability of the treatment benefit prediction. However, the similarities and differences between these two metrics are not yet explored. We compare and contrast the metrics from conceptual, theoretical, and empirical perspectives and illustrate the application of the metrics. We consider the common scenario of a logistic regression model for a binary response developed based on data from a randomized controlled trial with two treatment arms. This dissertation provides two major contributions: first, the two metrics are expanded into three pairs of metrics, each having a particular scope; second, it provides results of theoretical and simulation studies that compare and contrast the construct and empirical behaviour of these metrics. We found that the heterogeneity of treatment effect appropriately influences these metrics. Metrics related to the “concordance-statistic for benefit” are sensitive to the unobservable correlation between counterfactual outcomes. In a case study, we quantify the metrics in a randomized controlled trial of acute myocardial infarction therapies on 30-day mortality. We conclude that these metrics help understand the heterogeneity of treatment effect and the consequent impact on treatment decision-making.

Two UBC Statistics MSc Co-op student presentations

To join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Presentation 1

Time: 4pm – 4:30pm

Speaker: Jingyiran Li, UBC Statistics MSc Co-op Student

Title: From biostatistics to AI algorithmic trading: A validator's perspective

Abstract: Sophia kicked off her industrial experience with a Research Analyst position on a prospective cohort study with Health Canada–MIREC Group. Afterwards, she switched to the insurance industry by working as a Data Engineer who dealt extensively with big data platforms and tools such as HDFS, Spark, Jenkins, and various databases such as Hive, Drill, and IBMDB. Finally, Sophia landed an AI Scientist position with RBC Capital Markets where she mainly focused on validating high frequency trading algorithms through traditional benchmarking, predictive uncertainty quantification, and simulating optimal liquidation strategy using deep reinforcement learning. Sophia will share her work experiences in three different roles from three different industries and provide relevant insights to those who are interested in pursuing a career in those industries.

Presentation 2

Time: 4:30pm – 5pm

Speaker: Harper Cheng, UBC Statistics MSc Co-op Student

Title: My co-op experience at the BC Cancer Research Centre: aGCT RNA-seq sample analysis

Abstract: Adult granulosa cell tumours or aGCT are characterized by their slow growth and usually occur in peri- or post-menopausal women with a median age of diagnosis of 50 to 54 years. The majority of aGCTs are diagnosed at an early stage with an indolent prognosis. However, one-third of patients relapse, typically 4–7 years after initial diagnosis leading to death in 50% of these patients with advanced-stage tumours. Surgery is the foundation of treatment for both primary and recurrent disease, but there is a subset of patients with relapsed tumours where surgery is not an option.

Using whole-transcriptome paired-end RNA sequencing technology, Shah et al. (4) identified a single somatic missense mutation in FOXL2 (402C?G) in four GCT with the predicted consequence to be the substitution of a tryptophan residue for a highly conserved cysteine residue at amino acid position 134 (C134W).

Although we have translated this FOXL2 mutation discovery into a biomarker that can be used for differential diagnosis, we still do not understand how this FOXL2 mutation promotes tumourigenesis. So that leads to the overarching question of: During the last decade, our understanding of the molecular pathogenesis of aGCTs has significantly improved, whereas the developments of chemotherapeutic regimens and especially targeted therapies have remained modest. We analyzed bulk RNA seq data in the hope of elucidating how FOXL2 mutation promotes tumourigenesis.