Seminar

Forming realistic population distributions for flood frequency assessments

To join via Zoom: To join this seminar, please register here.

Title: Forming realistic population distributions for flood frequency assessments

Abstract: Forming a population distribution in the frequentist setting typically amounts to choosing a pre-defined family like Gaussian or Exponential, but this may not always be realistic. One such situation is the formation of distributions for the annual maximum river flow in a changing climate when the river is driven by more than one process (eg., snowmelt and rainfall). I will demonstrate this analysis for the Coldwater River in British Columbia, an analysis that has recently been submitted to the Fraser Basin Council following the devastating floods in British Columbia in November 2021. I will also demonstrate how the distplyr R package makes this type of analysis (forming realistic probability distributions) simpler by providing a grammar for manipulating probability distributions.

Design and Analysis of Computer Experiments: Large Datasets and Multi-Model Ensembles

To Join Via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: Computer models are used as replacements for physical experiments in a wide variety of applications. Nevertheless, direct use of the computer model for the ultimate scientific objective is often limited by the complexity and cost of the model. Historically, Gaussian process (GP) regression has proven to be the almost ubiquitous choice for a fast statistical emulator for such a computer model, due to its flexible form and analytical expressions for predictive uncertainty.

In the first part of this dissertation, we consider complications that arise when the design is moderate to large. Fitting a GP regression can be computationally intractable for even moderate designs, due to computing time increasing with the cube of the design size. We propose a new solution to this problem: adaptive design and analysis via partitioning trees (ADAPT). By taking a data-adaptive approach to the development of a design, and choosing to partition the space in the regions of highest variability, we obtain a higher density of points in these regions and hence accurate prediction for complex computer models.

Next, we consider the scenario where multiple computer models are available for predicting the same physical process—known as multi-model ensembles (MMEs). Such ensembles are common in many applications, such as climate modelling and weather prediction. We present a new statistical methodology for combining output from such models to best describe the underlying physical process, using field data to estimate the weights as- signed to each model. The methodology allows us to make predictions with appropriate measures of uncertainty. Additionally, the weights are allowed to vary with the inputs and thus represent the changing relative importance between the computer models throughout the input space. The methodology is applied to ice sheet models for the deglaciation of North America. Finally, we address several considerations that arise when the MME field data are binary. A new MME model formulation is presented, and applied to ice absence/presence data in the deglaciation application.

In summary, this dissertation presents new methods for two scenarios prevalent in the design and analysis of computer experiments: large designs, and the presence of multiple computer models.

Two UBC Statistics MSc student presentations (Johnny Xi & Naitong Chen)

To join via Zoom: Please register here.

Presentation 1

Time: 11:00am – 11:30am

Speaker: Johnny Xi, UBC Statistics MSc student

Title: Indeterminacy in Latent Variable Models: Characterization and Strong Identifiability

Abstract: The history of latent variable models spans nearly 100 years, from factor analysis to modern unsupervised machine learning. An enduring goal is to interpret the latent variables as "true" factors of variation, unique to a sample. Unfortunately, modern non-linear methods are wildly underdetermined, leading to many possible, equally valid solutions, even in the limit of infinite data. I will describe a theoretical framework that rigourously formulates the uniqueness problem as statistical identifiability, unifying existing progress towards this goal. The framework explicitly characterizes the sources of non-identifiability, making it possible to design strongly identifiable latent variable models in a transparent way. Using insights derived from the framework, our work proposes two flexible non-linear models with unique latent variables.

Presentation 2

Time: 11:30am – 12:00pm

Speaker: Naitong Chen, UBC Statistics MSc student

Title: Bayesian Inference via Sparse Hamiltonian Flows

Abstract: A Bayesian coreset is a small, weighted subset of data that replaces the full dataset during Bayesian inference, with the goal of reducing computational cost. Although past work has shown empirically that there often exists a coreset with low inferential error, efficiently constructing such a coreset remains a challenge. Current methods tend to be slow, require a secondary inference step after coreset construction, and do not provide bounds on the data marginal evidence. In this work, we introduce a new method—sparse Hamiltonian flows—that addresses all three of these challenges. The method involves first subsampling the data uniformly, and then optimizing a Hamiltonian flow parametrized by coreset weights and including periodic momentum quasi-refreshment steps. Theoretical results show that the method enables an exponential compression of the dataset in a representative model, and that the quasi-refreshment steps reduce the KL divergence to the target. Real and synthetic experiments demonstrate that sparse Hamiltonian flows provide accurate posterior approximations with significantly reduced runtime compared with competing dynamical-system-based inference methods

Non-Crossing Dual Neural Network: Joint Value at Risk and Conditional Tail Expectation estimations with non-crossing conditions

To Join via Zoom: To join this seminar, please register here.

Title: Non-Crossing Dual Neural Network: Joint Value at Risk and Conditional Tail Expectation estimations with non-crossing conditions

Abstract: When datasets present long conditional tails on their response variables, algorithms based on Quantile Regression have been widely used to assess extreme quantile behaviors. Value at Risk (VaR) and Conditional Tail Expectation (CTE) allow the evaluation of extreme events to be easily interpretable. The state-of-the-art methodologies to estimate VaR and CTE controlled by covariates are mainly based on linear quantile regression, and usually do not have in consideration non-crossing conditions across VaRs and their associated CTEs. We implement a non-crossing neural network that estimates both statistics simultaneously, for several quantile levels and ensuring a list of non-crossing conditions. We illustrate our method with a household energy consumption dataset from 2015 for quantile levels 0.9, 0.925, 0.95, 0.975 and 0.99, and show its improvements against a Monotone Composite Quantile Regression Neural Network approximation.

Improving multiple-try Metropolis with local balancing

To join via Zoom: Please request Zoom connection details from headsec@stat.ubc.ca

Title: Improving multiple-try Metropolis with local balancing

Abstract: Multiple-try Metropolis (MTM) is a popular Markov chain Monte Carlo algorithm which is amenable to parallel computing and thus has great potential. At each iteration, it samples several candidates for the next state of the Markov chain and randomly selects one of them based on a weight function. By leveraging a connection with the work of Zanella (2020), we show that the preferred choice of weight function in the literature, which is proportional to the target density, induces pathological behaviours in high dimensions. Those pathological behaviours are demonstrated through numerical experiments and theoretical results. We propose different weight functions for which those pathological behaviours do not arise. Also, we provide a scaling-limit analysis that allows to characterize the scalability with respect to the dimension of MTM when using the preferred weight function and when using a proposed weight function. In each case, MTM is seen as an approximation to a limiting sampling scheme which is approached under conditions on the rate at which the number of candidates increases with the dimension.

Investigating pilotage as a navigational strategy in Northern Fur Seals (Callorhinus ursinus) during pre-migratory foraging trips

To join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: Research on the navigational strategies of marine mammals often uses long-distance migratory data. However, these cost-intensive, long-term monitoring projects are often limited to small sample sizes. One alternative is to use non-migratory trips, where more animals can be tracked. Using shorter trips to investigate navigational hypotheses also allows for a finer-scale investigation of navigational movement, which can aid in exploring pilotage as a potential navigational strategy.

In this project, I explore the potential use of bathymetric landmarks as navigational landmarks by Northern Fur Seals (NFS) during pre-migratory foraging trips. I isolate transiting-type behavior during pre-migratory foraging trips using hidden Markov models, and then use potential functions – a technique traditionally used to model particle or planetary motion – to model the transiting movement of NFS during pre-migratory foraging trips. I then compare these potential functions to the bathymetry of the region.

An R package for computing the regularization paths for sparse group-lasso penalized problems

To Join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: The sparse group-lasso is an advanced statistical technique that involves the lasso and group-lasso penalties, which enables the sparsity at both the individual and group level regarding the natural grouping structure when solving high-dimensional learning problems. However, estimating the regularization paths can be computationally hard if the designed matrix input is large. To improve the computational efficiency, we develop an R package ’sparsegl' according to an algorithm by iteratively applying the KKT Stationarity Condition and the Strong Rule, which helps to identify the active and inactive predictors before updating their coefficients. Furthermore, we empirically illustrate its efficiency by comparing the running time to other existing packages for solving lasso-type problems.

SSC Data Science & Analytics Section 2022 Webinar Series: Vincenzo Coia

Zoom Link & Talk Details

This talk is one of SSC Data Science & Analytics Section 2022 Webinar Series.

Zoom Link: https://uwaterloo.zoom.us/j/96812177576 (Passcode: DSA2022)

Talk Title: Piecing together the statistical puzzle to build a model

Abstract: As a student well into my statistics education, I knew a fair share of statistical methods and how they worked. But I couldn't see how they all fit together in a bigger framework. Methods like linear regression, machine learning, and out-of-sample scoring seemed to stand alone. In this talk, I'll discuss how your analysis' overarching statistical question, together with the nature of your data, tells us where and when to use what statistical methods. A variety of these methods will be discussed in the context of uncertainty quantification, tying statistics back to its probabilistic roots. By the end, you'll learn how to make clearer decisions when building a statistical model.

UBC Statistics MSc students Co-op Talks II

To join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Presentation 1

Time: 4:00pm – 4:30pm

Speaker: Jeffrey Yiu, UBC Statistics MSc student

Title: Comparing public and private long-term care outcomes in British Columbia

Abstract: In British Columbia, approximately 90% of long-term care beds for seniors are publicly subsidized, while the remaining beds are privately paid. A 2018 report from the Office of the Seniors Advocate, an independent office of the B.C. provincial government, found that there were differences in emergency department visits, hospitalization rates, and mortality, depending on whether beds were publicly subsidized or privately paid. During my co-op terms with the B.C. Ministry of Health, I explored health databases such as the Discharge Abstract Database (DAD) and the National Ambulatory Care Reporting System (NACRS) to compare health outcomes between public and private long-term care clients in the 2019/20 fiscal year. Areas of investigation included demographics, morbidity, mortality, and utilization of health services.

Presentation 2

Time: 4:30pm – 5:00pm

Speaker: Pramoda Jayasinghe, UBC Statistics MSc student

Title: Working with health data and consulting for health sector clients

Abstract: Health data helps inform many decisions ranging from public health policies to showing the efficacy of a new drug. Due to the sensitivity of health data and the expectations of clients using said data, consulting for health-related projects has a unique set of challenges. While working on multiple projects during my co-op at Broadstreet, an organization that provides consulting services for health economics and outcomes research, I got to experience some of the practical issues that exist when consulting for health sector clients, specifically. The main objective of this presentation is to point out some aspects to consider when working with real-world health data.

Presentation 3

Time: 5:00pm – 5:30pm

Speaker: Li Zha, UBC Statistics MSc student

Title: Aspects of working with data between industry and health research

Abstract: Machine learning has been widely adopted in many applications in industry and research in recent years.  In this presentation, I will first introduce my industrial experience as a data science intern with TD Personal Banking where I mainly focused on building a propensity score matching for part of their test tool, and exploring data pruning to deal with noisy texts for fraud detection. Afterwards, I switched to the health industry by working on a polygenic risk score using GWAS (Genome-wide genetic association studies) data under Prof. Park’s lab in BC Cancer.  This presentation will focus more on sharing my learnings with working in different industries and provide some insights to those who are interested in data science roles.

UBC Statistics MSc students Co-op Talks I

To join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Presentation 1

Time: 4:00pm – 4:30pm

Speaker: Shirley Cui, UBC Statistics MSc student

Title: Co-op experience in a Contract Research Organization(CRO) company: phase I-III clinical trials

Abstract: I started my industrial experience with a junior statistician position in a CRO company that provides clinical trial management services (eg. clinical research project management, on-site management, data management, statistical analysis, etc.) for the pharmaceutical industry companies. I mainly assist senior statistician with study design, sample size calculation, randomization list generation, eCRF review and annotation etc. I would like to share my coop experience in participating research projects involving phase I-III clinical trials and provide insights to those who are interested in pursuing a career in this field.

Presentation 2

Time: 4:30pm – 5:00pm

Speaker: Shuyi Tan, UBC Statistics MSc student

Title: Data Analytics in Healthcare

Abstract: The healthcare industry collects vast amounts of health data. Analysis of these data informs important decisions related to healthcare delivery, evaluation, and quality management. For the past 12 months, Shuyi has been working as a business analyst at Fraser Health Authority focusing on process improvement and data analysis. She would like to talk about her experience working with healthcare data in terms of data acquisition, data mining, insight extraction, and potential challenges. Since the data-related inquiries from clinical specialists will be quite different from those from an academic setting, examples would be provided to give some insights into how to approach these real-world questions.

Presentation 3

Time: 5:00pm – 5:30pm

Speaker: Zhipeng Zhu, UBC Statistics MSc student

Title: Practical Tools for Junior Data Analyst: From Graphic Software to DSLs

Abstract: This presentation talks about some experience as an entry-level data analyst or data scientist and introduces the idea of participating in DSL development.

Compared to academic tasks in many statistics courses, the actual work content as an analyst or consultant is quite different. While statistic knowledge is indispensable, using tools flexibly has become more critical in many workloads. To deal with complicated cases, analytic software with graphic UI has become an essential tool for most analysts. These graphic software offer excellent plot outputs and a superior feature of document generation compared to coding style tools such as R. However, graphic software sometimes cost the performance for particular tasks, and hence we require the hard-coded programs. This inspires me to participate in developing R packages “distplyr” and “distionary” as DSLs with Vincenzo Coia. Even though we still need to use coding languages for many specific tasks, a good DSL will significantly improve our productivity.