Seminar

Bayesian causal inference for discrete data

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: Causal inference provides a framework for estimating how a response changes when a given cause of interest changes. When all data are discrete we can use saturated nonparametric models to avoid unnecessary assumptions in our causal inference modelling, where we specify unique parameters for all possible combinations of treatments and confounders when estimating an outcome. Bayesian methods allow us to incorporate prior information into these saturated models, making them usable beyond simple settings with low dimensional confounders.

We propose two new nonparametric Bayes methods for causal inference based on saturated modelling. The first method combines a parametric model with a nonparametric saturated outcome model to estimate treatment effects in observational studies with longitudinal data. By conceptually splitting the data, we can combine these models while maintaining a conjugate framework, allowing us to avoid the use of Markov chain Monte Carlo methods. Approximations using the central limit theorem and random sampling allows our method to be scaled to high-dimensional confounders.

The second method uses prior restrictions of the parameter space of a saturated model to partially identify causal effect estimates in scenarios with nonignorable missing outcome data. We focus on two common restrictions, instrumental variables and the direction of missing data bias, and investigate how these restrictions narrow the identification region for parameters of interest. Additionally, we propose a rejection sampling algorithm that allows us to quantify the evidence for these assumptions in the data.

Causal clustering: design of cluster experiments under network interference

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: This paper studies the design of cluster experiments to estimate the global treatment effect in the presence of network spillovers. We provide a framework to choose the clustering that minimizes the worst-case mean-squared error of the estimated global effect. We show that optimal clustering solves a novel penalized min-cut optimization problem computed via off-the-shelf semi-definite programming algorithms. Our analysis also characterizes simple conditions to choose between any two cluster designs, including choosing between a cluster or individual-level randomization. We illustrate the method's properties using unique network data from the universe of Facebook's users and existing data from a field experiment.

Link to paper: https://arxiv.org/abs/2310.14983

Finding the Best Player via Multi-Player Comparisons

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: Suppose there are n items, with each item having an unknown value. At any time, we can can either stop and declare which item has the largest value or else choose a subset of items to compare.

If subset S is chosen, then a given item in S will be preferred with a probability equal to the value of that item divided by the sum of the values of all items in S.

Assuming a Bayesian prior on the values, and subject to the proviso that the policy employed will make the correct choice with probability at least some specified value, we are looking for a policy that needs a relatively small mean number of comparisons before making a decision. Some heuristic policies are presented and analyzed.

Extreme Value Modelling with Application to Reverse Stress Testing

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: Reverse stress testing of a financial portfolio aims to identify scenarios for risk factors that lead to a specified adverse portfolio outcome. The stress scenarios of interest naturally need to be extreme yet plausible. A statistical formulation of these requirements is to define a stress scenario at extreme threshold as the mode of the conditional density of the random vector of risk factors given that the loss on the portfolio exceeds the threshold.

In situations where the interest is in an extreme conditioning event corresponding to a large value of the threshold, traditional multivariate mode estimators would not perform well due to very limited sample of observations in the conditioning event. Under this consideration, we propose an estimator based on techniques from multivariate extreme value theory under the assumption of multivariate regular variation and tail dependence. The method effectively addresses data scarcity in the joint tail regions while allowing for more flexible model assumptions focusing on extremes. We study the asymptotic behaviour of the proposed estimator, investigate its finite-sample performance in simulation studies and apply it to real data in a case study.

Modelling orbits to detect exoplanets

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: Determining the orbits of planets is one of the earliest applications of modern science, though nowadays our focus has turned outwards from our solar system to exoplanets orbiting distant stars. To detect and study these planets, we must combine sparse evidence from a variety of sources---direct images of exoplanetary systems, radial velocity curves, interferometer visibilities, transits, and more---to model their physical and orbital parameters. Unfortunately for us, these models can lead to very challenging posteriors that are unidentifiable, multimodal, and can have strong curvature. I will discuss these modelling challenges and present our successes from applying modern sampling techniques (gradient based samplers, and non-reversible parallel tempering) which reduce sampling time from days to seconds, and will enable us to find new exoplanets.

Dependence modelling in high dimensions with latent variables

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: Modelling the dependence relations among a large number of quantitative variables has broad applications in various fields. Dependence models with copulas are widely used in multivariate applications when the classical assumption of Gaussian-distributed variables does not hold, and tail inferences are needed. For a large number of highly correlated variables, the dependence relation can be explained using several latent variables; these methods are known as factor models.

In Gaussian factor models (1-factor, bi-factor, oblique factor) and their factor copula counterparts, we propose a way to estimate the latent variables with proxies. We show the proxies, which are defined as the conditional expectation of the latent factors given the observed variables, are consistent (under weak regularity conditions) as the number of observed variables linked to each latent variable increases. The proxies can help to select the bivariate linking copulas in factor copulas and to estimate the copula parameters. With estimated "latent variables", existing copula methods can be applied in modelling the dependence relations of a large number of variables to yield a parsimonious dependence structure. Examples of dependence graphs will be used for illustration.

MSc students' Co-op Presentations

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Presentation 1

Time: 10:30am – 11:00am

Speaker: Clayton Allard, UBC Statistics MSc student

Title: 2023 co-op term

Abstract: For my 8-month co-op, I worked as a data scientist for the US Department of Defense. I worked with a collection of cybersecurity data referred to as the "Activity Store". The Activity Store collects data on suspicious activity within a customer's network which is then flagged as an alert for review. The problem is that there are far too many alerts to review. My main project was to create a statistical model to predict whether each alert corresponds to malicious or benign activity. The objective was to select alerts to be reviewed that are more likely to be malicious activity, and determine how the alert system can be modified to reduce the traffic entering the Activity Store. As a result, we used a random forest and got significantly better results than the manual rule-based approach that was previously implemented.

Presentation 2

Time: 11:00am – 11:30am

Speaker: Xiao (Nicole) Hu, UBC Statistics MSc student

Abstract: I worked as a data analyst at Staples Lab within the UBC Department of Medicine for an 8-month Co-op term. My primary project focused on examining the risk of illicit drug overdose following ‘Before medically advised’ (BMA) discharge. The BMA discharge or patient-initiated departure is more common among people with problematic drug use and has been reported to increase the risk of death, yet its relationship to subsequent illicit drug overdose remains uncertain. We conducted a cohort analysis and a case-crossover analysis of population-based linked administrative health data to examine the risk of subsequent illicit drug overdose. In the cohort study, we performed a retrospective analysis using administrative health data from a 20% random sample of the population of British Columbia, Canada. We focused on non-elective, non-obstetrical hospitalizations occurring between 2015 and 2019. We used survival analysis to compare the risk of drug overdose in the first 30 days after BMA discharge to the risk of drug overdose after routine physician-advised discharge. In the case cross-over study, we focused on individuals experiencing an overdose between 2016 and 2019 in British Columbia, Canada. We used a conditional logistic regression to compare the likelihood of hospital discharge in the 28 days prior to overdose (the 'pre-overdose interval') to the likelihood of hospital discharge in two self-matched 28-day control intervals ending 26 and 52 weeks prior to overdose. The primary analysis evaluated whether BMA discharge was associated with an increased risk of subsequent drug overdose. A secondary analysis evaluated the association between physician-advised discharge and subsequent drug overdose.

Presentation 3

Time: 11:30am – 12:00pm

Speaker: Yicheng Wang, UBC Statistics MSc student

Title: Data Pipeline Infrastructure Before Any Analytics Models

Abstract: In industrial world, study shows that data scientists and analysts spend 60% of their time on cleaning and organizing data. The quality of data is found to be the key to success of modelling success. This talk will focus on the infrastructure of end-to-end data pipelining. During my co-op in Samsung, I mainly worked on data pipeline construction and airflow unit testing. I used Apache Airflow, Spark and AWS Athena to re-build query engine in recommendation system for 100+ million SmartThings users.  At the end of co-op, I looked into Airflow API document and source code to research on how to build unit tests to shorten time for pipeline development.

Presentation 4

Time: 12:00pm – 12:30pm

Speaker: Yixin Zhang, UBC Statistics MSc student

Title: Two Implementations of Large Language Models

Abstract: Large Language Models have captured widespread attention since the grand entrance of ChatGPT. During my co-op at Landsure Systems, I worked on two distinct projects that in some ways implemented large language models. In the first project, we aimed to detect and highlight discriminatory covenants from approximately 110 million pages of land title contracts. We implemented pre-trained transformer models specialized in sentiment analysis to differentiate between discriminatory and non-discriminatory phrases. In the second project, we implemented retrieval augmented generation on GPT-4 to create a chatbot for land title related queries. The resulting product was able to give more accurate and domain-specific answers with lower hallucination compared to traditional chatbots and base GPT models.

Bayesian Inference for Big Data

To join this seminar virtually: https://ubc.zoom.us/j/68285564037?pwd=R2ZpLy9uc2pUYldHT3laK3orakg0dz09

Meeting ID: 682 8556 4037 Passcode: 636252

Abstract: Since shortly after the popularization of stochastic gradient optimization methods in machine learning---which now scale model training to billions of examples and beyond---researchers have been trying to use the same basic data subsampling techniques to speed up computational Bayesian inference algorithms. In this talk, I'll cover the broad classes of methods that have been developed, highlights of progress in the field, and the current state of the art. Along the way I'll introduce some recent work from my group on scalable Bayesian inference via coresets, i.e., sparse dataset summaries. I'll show that coresets offer an exponential compression of the data (and so an exponential speed-up of methods like Markov chain Monte Carlo) in a wide variety of models, and can be constructed in an automated manner, with theoretical convergence guarantees, and without requiring special knowledge of model structure beyond conditional independence of data. While other methods implicitly rely on asymptotic Gaussianity, coresets are particularly amenable to posteriors that don't exhibit this usual asymptotic behaviour, with discrete variables, weak or unidentifiability, low-dimensional manifold structure, etc. I'll conclude with empirical results and a discussion of next steps for Bayesian inference in the big data regime.

van Eeden seminar: Ethical AI is More than Loss Functions

Zoom Registration

https://ubc.zoom.us/meeting/register/u5Mpfu-orz4uHtMoS6AcwTE_0VZ_DDEghNdA

(If you have any questions about your registration or the seminar, please contact ea@stat.ubc.ca.)

Title

Ethical AI is More than Loss Functions

Abstract

What constitutes a fair algorithm and the ethical use of data is context specific. Algorithms are not neutral and optimization choices will reflect a specific value system and the distribution of power to make these decisions. Data also reflect societal bias, such as structural racism. Ethics and fairness research for health AI spans many fields, including policy, medicine, computer science, sociology, and statistics. Considerations go well beyond loss functions and typical measures of statistical assessment. This talk includes discussion of team construction, who decides the research question, minimum standards for research quality, reproducibility, least publishable units, and community engaged research. Overarching themes are also that centering health equity and developing methodology tailored to specific health questions are critical given the stakes involved.

van Eeden speakers

Professor Sherri Rose has been invited to be this year's van Eeden speaker by the graduate students in the Department of Statistics at the University of British Columbia. A van Eeden speaker is a prominent statistician who is chosen each year to give a lecture, supported by the UBC Constance van Eeden Fund. The 2024 seminar is additionally sponsored by the Canadian Statistical Sciences Institute (CANSSI), the Pacific Institute for the Mathematical Sciences (PIMS), and the Walter H. Gage Memorial Fund.

 

Dependence models for mortalities using a copula state-space approach

To join this seminar virtually: please request Zoom connection details from ea@stat.ubc.ca

Abstract: We investigate how the COVID-19 pandemic has affected mortality experience of major causes, and more importantly how COVID-19 has changed the dependence structure across these causes. This enables us to gain more insights into the potential impact of COVID-19 on future life expectancy, and conduct scenario-based projections. For this we propose a dynamic copula state space approach, which quantifies and visualizes the joint dynamics across major causes of death, both before and during the pandemic. Based on US weekly mortality data from January 2015 to November 2022, we find that COVID-19 has elevated the mortality level for a majority of causes, and has changed the dependence structure across these causes.

This work is joint with Han Li, Centre for Actuarial Studies, Department of Economics, The University of Melbourne, Australia.