Seminar

Dependence modelling in high dimensions with latent variables

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: Modelling the dependence relations among a large number of quantitative variables has broad applications in various fields. Dependence models with copulas are widely used in multivariate applications when the classical assumption of Gaussian-distributed variables does not hold, and tail inferences are needed. For a large number of highly correlated variables, the dependence relation can be explained using several latent variables; these methods are known as factor models.

In Gaussian factor models (1-factor, bi-factor, oblique factor) and their factor copula counterparts, we propose a way to estimate the latent variables with proxies. We show the proxies, which are defined as the conditional expectation of the latent factors given the observed variables, are consistent (under weak regularity conditions) as the number of observed variables linked to each latent variable increases. The proxies can help to select the bivariate linking copulas in factor copulas and to estimate the copula parameters. With estimated "latent variables", existing copula methods can be applied in modelling the dependence relations of a large number of variables to yield a parsimonious dependence structure. Examples of dependence graphs will be used for illustration.

MSc students' Co-op Presentations

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Presentation 1

Time: 10:30am – 11:00am

Speaker: Clayton Allard, UBC Statistics MSc student

Title: 2023 co-op term

Abstract: For my 8-month co-op, I worked as a data scientist for the US Department of Defense. I worked with a collection of cybersecurity data referred to as the "Activity Store". The Activity Store collects data on suspicious activity within a customer's network which is then flagged as an alert for review. The problem is that there are far too many alerts to review. My main project was to create a statistical model to predict whether each alert corresponds to malicious or benign activity. The objective was to select alerts to be reviewed that are more likely to be malicious activity, and determine how the alert system can be modified to reduce the traffic entering the Activity Store. As a result, we used a random forest and got significantly better results than the manual rule-based approach that was previously implemented.

Presentation 2

Time: 11:00am – 11:30am

Speaker: Xiao (Nicole) Hu, UBC Statistics MSc student

Abstract: I worked as a data analyst at Staples Lab within the UBC Department of Medicine for an 8-month Co-op term. My primary project focused on examining the risk of illicit drug overdose following ‘Before medically advised’ (BMA) discharge. The BMA discharge or patient-initiated departure is more common among people with problematic drug use and has been reported to increase the risk of death, yet its relationship to subsequent illicit drug overdose remains uncertain. We conducted a cohort analysis and a case-crossover analysis of population-based linked administrative health data to examine the risk of subsequent illicit drug overdose. In the cohort study, we performed a retrospective analysis using administrative health data from a 20% random sample of the population of British Columbia, Canada. We focused on non-elective, non-obstetrical hospitalizations occurring between 2015 and 2019. We used survival analysis to compare the risk of drug overdose in the first 30 days after BMA discharge to the risk of drug overdose after routine physician-advised discharge. In the case cross-over study, we focused on individuals experiencing an overdose between 2016 and 2019 in British Columbia, Canada. We used a conditional logistic regression to compare the likelihood of hospital discharge in the 28 days prior to overdose (the 'pre-overdose interval') to the likelihood of hospital discharge in two self-matched 28-day control intervals ending 26 and 52 weeks prior to overdose. The primary analysis evaluated whether BMA discharge was associated with an increased risk of subsequent drug overdose. A secondary analysis evaluated the association between physician-advised discharge and subsequent drug overdose.

Presentation 3

Time: 11:30am – 12:00pm

Speaker: Yicheng Wang, UBC Statistics MSc student

Title: Data Pipeline Infrastructure Before Any Analytics Models

Abstract: In industrial world, study shows that data scientists and analysts spend 60% of their time on cleaning and organizing data. The quality of data is found to be the key to success of modelling success. This talk will focus on the infrastructure of end-to-end data pipelining. During my co-op in Samsung, I mainly worked on data pipeline construction and airflow unit testing. I used Apache Airflow, Spark and AWS Athena to re-build query engine in recommendation system for 100+ million SmartThings users.  At the end of co-op, I looked into Airflow API document and source code to research on how to build unit tests to shorten time for pipeline development.

Presentation 4

Time: 12:00pm – 12:30pm

Speaker: Yixin Zhang, UBC Statistics MSc student

Title: Two Implementations of Large Language Models

Abstract: Large Language Models have captured widespread attention since the grand entrance of ChatGPT. During my co-op at Landsure Systems, I worked on two distinct projects that in some ways implemented large language models. In the first project, we aimed to detect and highlight discriminatory covenants from approximately 110 million pages of land title contracts. We implemented pre-trained transformer models specialized in sentiment analysis to differentiate between discriminatory and non-discriminatory phrases. In the second project, we implemented retrieval augmented generation on GPT-4 to create a chatbot for land title related queries. The resulting product was able to give more accurate and domain-specific answers with lower hallucination compared to traditional chatbots and base GPT models.

Bayesian Inference for Big Data

To join this seminar virtually: https://ubc.zoom.us/j/68285564037?pwd=R2ZpLy9uc2pUYldHT3laK3orakg0dz09

Meeting ID: 682 8556 4037 Passcode: 636252

Abstract: Since shortly after the popularization of stochastic gradient optimization methods in machine learning---which now scale model training to billions of examples and beyond---researchers have been trying to use the same basic data subsampling techniques to speed up computational Bayesian inference algorithms. In this talk, I'll cover the broad classes of methods that have been developed, highlights of progress in the field, and the current state of the art. Along the way I'll introduce some recent work from my group on scalable Bayesian inference via coresets, i.e., sparse dataset summaries. I'll show that coresets offer an exponential compression of the data (and so an exponential speed-up of methods like Markov chain Monte Carlo) in a wide variety of models, and can be constructed in an automated manner, with theoretical convergence guarantees, and without requiring special knowledge of model structure beyond conditional independence of data. While other methods implicitly rely on asymptotic Gaussianity, coresets are particularly amenable to posteriors that don't exhibit this usual asymptotic behaviour, with discrete variables, weak or unidentifiability, low-dimensional manifold structure, etc. I'll conclude with empirical results and a discussion of next steps for Bayesian inference in the big data regime.

van Eeden seminar: Ethical AI is More than Loss Functions

Zoom Registration

https://ubc.zoom.us/meeting/register/u5Mpfu-orz4uHtMoS6AcwTE_0VZ_DDEghNdA

(If you have any questions about your registration or the seminar, please contact ea@stat.ubc.ca.)

Title

Ethical AI is More than Loss Functions

Abstract

What constitutes a fair algorithm and the ethical use of data is context specific. Algorithms are not neutral and optimization choices will reflect a specific value system and the distribution of power to make these decisions. Data also reflect societal bias, such as structural racism. Ethics and fairness research for health AI spans many fields, including policy, medicine, computer science, sociology, and statistics. Considerations go well beyond loss functions and typical measures of statistical assessment. This talk includes discussion of team construction, who decides the research question, minimum standards for research quality, reproducibility, least publishable units, and community engaged research. Overarching themes are also that centering health equity and developing methodology tailored to specific health questions are critical given the stakes involved.

van Eeden speakers

Professor Sherri Rose has been invited to be this year's van Eeden speaker by the graduate students in the Department of Statistics at the University of British Columbia. A van Eeden speaker is a prominent statistician who is chosen each year to give a lecture, supported by the UBC Constance van Eeden Fund. The 2024 seminar is additionally sponsored by the Canadian Statistical Sciences Institute (CANSSI), the Pacific Institute for the Mathematical Sciences (PIMS), and the Walter H. Gage Memorial Fund.

 

Dependence models for mortalities using a copula state-space approach

To join this seminar virtually: please request Zoom connection details from ea@stat.ubc.ca

Abstract: We investigate how the COVID-19 pandemic has affected mortality experience of major causes, and more importantly how COVID-19 has changed the dependence structure across these causes. This enables us to gain more insights into the potential impact of COVID-19 on future life expectancy, and conduct scenario-based projections. For this we propose a dynamic copula state space approach, which quantifies and visualizes the joint dynamics across major causes of death, both before and during the pandemic. Based on US weekly mortality data from January 2015 to November 2022, we find that COVID-19 has elevated the mortality level for a majority of causes, and has changed the dependence structure across these causes.

This work is joint with Han Li, Centre for Actuarial Studies, Department of Economics, The University of Melbourne, Australia.

Probabilistic topic models for single-cell genomics

To join this seminar virtually: please request Zoom connection details from ea@stat.ubc.ca

Abstract: Building a comprehensive topic model has become an important research tool in single-cell genomics. With a topic model, we can decompose and ascertain distinctive cell topics shared across multiple cells, and the gene programs implicated by each topic can later serve as a predictive model in translational studies. In this talk, I will present a few topic modeling tools for single-cell RNA-sequencing (RNA-seq) data analysis. The first topic model builds on the Embedded Topic Model (ETM) and incorporates sparse-inducing priors to make the model more interpretable. I will showcase that it can be used to unravel cell types or cell states. The second topic model uncovers short-term RNA velocity patterns from a plethora of spliced and unspliced single-cell RNA-sequencing (RNA-seq) counts. I will show that modeling both types of RNA counts can improve robustness in statistical estimation and can reveal new aspects of dynamic changes that can be missed in static analysis. I will showcase that our modeling framework can be used to identify statistically significant dynamic gene programs in pancreatic cancer data. Our results discovered that seven dynamic gene programs (topics) are highly correlated with cancer prognosis and generally enrich immune cell types and pathways.

Margin-closed multivariate autoregressive time series models

To join this seminar virtually: please request Zoom connection details from ea@stat.ubc.ca

Abstract: In literature on multivariate modeling using copulas, it is typical to first model univariate margins and then the multivariate dependence structure between the marginal components. This is useful because there are many diagnostics that can help in selecting univariate models. Following the same idea, we derive the conditions for when a multivariate stationary Gaussian vector autoregressive (VAR) time series is closed under margins, i.e., it has univariate autoregressive (AR) margins or lower-dimensional vector autoregressive (VAR) margins. It leads to a copula model and it can be extended to a regime-switching setting. The constraint of the closure under margins can reduce the number of parameters in VAR and Markov switching vector autoregressive (MSVAR) models. Moreover, after transforming the stationary univariate margins into standard Gaussian, the property of closure under margins enables a new framework of modeling high-dimensional time series by modeling its low-dimensional sub-processes first and then modeling their dependence structure. The framework makes it more flexible in analyzing marginal behavior of the multivariate time series and also enables a multi-stage estimation procedure.

Speeding up Metropolis using Theorems

To join this seminar virtually: Please register here.

Title: Speeding up Metropolis using Theorems

Abstract: Markov chain Monte Carlo (MCMC) algorithms, such as the Metropolis algorithm, are designed to converge to complicated high-dimensional target distributions, to facilitate sampling. The speed of this convergence is essential for practical use. In this talk, we will present several theoretical probability results which can help improve the Metropolis algorithm's convergence speed. Specific topics will include: diffusion limits, optimal scaling, optimal proposal shape, tempering, adaptive MCMC, the Containment property, and the notion of adversarial Markov chains. The ideas will be illustrated using the simple graphical example available at probability.ca/met. No particular background knowledge will be assumed.

Uncertainty Quantification for Structure Learning and Interpretable Machine Learning

Date/Time: Thursday, February 8, 2024, 10:25am to 11:25am

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Abstract: The reliability of machine learning in scientific discoveries and decision-making hinges on the critical role of uncertainty quantification (UQ). In this talk, I will discuss UQ in two challenging scenarios motivated by scientific and societal applications: selective inference for large-scale graph learning and UQ for model-agnostic machine learning interpretations. Specifically, the first part concerns graphical model inference when only irregular, patchwise observations are available, a common setting in neuroscience, healthcare, genomics, and econometrics. To filter out low-confidence edges due to the irregular measurements, I will present a novel inference method that quantifies the uneven edgewise uncertainty levels over the graph as well as an FDR control procedure; this is achieved by carefully disentangling the dependencies across the graph and consequently yields more reliable graph selection. In the second part, I will discuss the computational and statistical challenges associated with UQ for feature importance of any machine learning model. I will take inspiration from recent advances in conformal inference and utilize an ensemble framework to address these challenges. This leads to an almost computationally free, assumption-light, and statistically powerful inference approach for occlusion-based feature importance. For both parts of the talk, I will highlight the potential applications of my research in science and society as well as how it contributes to more reliable and trustworthy data science.

Probabilistic methods for designing functional protein structures

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Abstract: The biochemical functions of proteins, such as catalyzing a chemical reaction or binding to a virus, are typically conferred by the geometry of only a handful of atoms.  This arrangement of atoms, known as a motif, is structurally supported by the rest of the protein, referred to as a scaffold.  A central task in protein design is to identify a diverse set of stabilizing scaffolds to support a motif known or theorized to confer function. This long-standing challenge is known as the motif-scaffolding problem.

In this talk, I describe a statistical approach I have developed to address the motif-scaffolding problem.  My approach involves (1) estimating a distribution supported on realizable protein structures and (2) sampling scaffolds from this distribution conditioned on a motif.  For step (1) I adapt diffusion generative models to fit example protein structures from nature.  For step (2) I develop sequential Monte Carlo algorithms to sample from the conditional distributions of these models.  I finally describe how, with experimental and computational collaborators, I have generalized and scaled this approach to generate and experimentally validate hundreds of proteins with various functional specifications.