Seminar

CANCELLED: On False Positive Error

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: Controlling the false positive error in model selection is a prominent paradigm for gathering evidence in data-driven science.  In model selection problems such as variable selection and graph estimation, models are characterized by an underlying Boolean structure such as presence or absence of a variable or an edge.  Therefore, false positive error or false negative error can be conveniently specified as the number of variables/edges that are incorrectly included or excluded in an estimated model.  However, the increasing complexity of modern datasets has been accompanied by the use of sophisticated modeling paradigms in which defining false positive error is a significant challenge.  For example, models specified by structures such as partitions (for clustering), permutations (for ranking), directed acyclic graphs (for causal inference), or subspaces (for principal components analysis) are not characterized by a simple Boolean logical structure, which leads to difficulties with formalizing and controlling false positive error.  We present a generic approach to endow a collection of models with partial order structure, which leads to systematic approaches for defining natural generalizations of false positive error and methodology for controlling this error. (Joint work with Peter Bühlmann, Venkat Chandrasekaran, and Parikshit Shah)

***

Dear STAT seminars and events subscribers:

Please be advised that this seminar has been cancelled. We will reschedule the seminar in the near future. We sincerely apologize for any inconvenience.

Best wishes,

UBC Statistics Department

Modular and Efficient Compilation of Probabilistic Programs

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: Probabilistic programming languages (PPLs) enable a clean separation between probabilistic models and Bayesian inference algorithms. Ideally, such separation makes it possible for the modeler to focus on the probabilistic modeling task without knowing the details of how to implement the inference method. However, making such general inference machinery efficient and automatic is challenging. In this talk, I will discuss our ongoing work on developing a modular framework for efficient compilation of domain-specific languages (DSLs) targeting PPLs. Specifically, I will discuss techniques that enable modular design and compilation algorithms for efficient inference. Moreover, I will also briefly discuss two application areas, including probabilistic programming of real-time systems and statistical phylogenetics.

Bio: David Broman is a Professor at the Department of Computer Science, KTH Royal Institute of Technology, a Visiting Professor at the Computer Science Department, Stanford University, and an Associate Director Faculty for Digital Futures. He received his Ph.D. in Computer Science in 2010 from Linköping University, Sweden. Between 2012 and 2014, he was a visiting scholar at the University of California, Berkeley, where he also was employed as a part-time researcher until 2016. His research focuses on the intersection of (i) programming languages and compilers, (ii) real-time and cyber-physical systems, and (iii) probabilistic machine learning. David has received the Best ETAPS paper award on on programming languages and systems (the EAPLS Award, co-authored 2023), a Distinguished Artifact Award at ESOP (co-authored 2022), an outstanding paper award at RTAS (co-authored 2018), a best paper award in the journal Software & Systems Modeling (SoSyM award 2018), the award as teacher of the year, selected by the student union at KTH (2017), the best paper award at IoTDI (co-authored 2017), and awarded the Swedish Foundation for Strategic Research's individual grant for future research leaders (2016). He has worked several years within the software industry, co-founded companies, co-founded the EOOLT workshop series, and is a member of IFIP WG 2.4, Modelica Association, a senior member of IEEE, and a former board member of Forskning och Framsteg.

Generalized Data Thinning Using Sufficient Statistics

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: Sample splitting is one of the most tried-and-true tools in the data scientist toolbox. It breaks a data set into two independent parts, allowing one to perform valid inference after an exploratory analysis or after training a model. A recent paper (Neufeld, et al. 2023) provided a remarkable alternative to sample splitting, which the authors showed to be attractive in situations where sample splitting is not possible. Their method, called convolution-closed data thinning, proceeds very differently from sample splitting, and yet it also produces two statistically independent data sets from the original. In this talk, we will show that sufficiency is the key underlying principle that makes their approach possible. This insight leads naturally to a new framework, which we call generalized data thinning. This generalization unifies both sample splitting and convolution-closed data thinning as different applications of the same procedure. Furthermore, we show that this generalization greatly widens the scope of distributions where thinning is possible.

Testing for Change-points in Heavy-tailed Time Series – A Winsorized CUSUM Approach

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: It is well-known how to detect the change-point in heavy-tailed time series is an open problem since the traditional tests may not have a power. This article proposes a winsorized CUSUM approach to solve this problem. We begin by investigating the winsorized CUSUM process and deriving the limiting distributions of the Kolmogorov-Smirnov test and the Self-normalized test under the null hypothesis. Under the alternative hypothesis, we firstly uncover the behavior of change-point magnitude after the winsorized data and show that our tests have a power approaching to 1 as the sample size $n \to\infty$. We then extend the winsorizing technique to tests for multiple change-points without the prior information on the number of actual change points. Our framework is quite general and its assumption is very weak. This enables the application of our tests to both linear time series and nonlinear time series, such as TAR and G-GARCH processes. The empirical results illustrate the effectiveness of our proposed procedures for change-point detection. (This is a joint work with Rui She and Linlin Dai)

Two MSc student presentations (Hannah Bobst & Nathaniel Dyrkton)

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Presentation 1

Time: 3:00pm – 3:30pm

Speaker: Hannah Bobst, UBC Statistics MSc student

Title: Analysis of Density Ratio Model Performance for Quantile Estimation

Abstract: Quantiles are important descriptive statistics in many applications. In practice, quantiles must be estimated using a representative sample from the population of interest. While parametric estimators are preferable in general, they become inaccurate in the case of even minor model misspecification. Nonparametric estimators, which requires no specific model, are common alternative to parametric estimators. However, empirical-based estimators often require large sample sizes to attain satisfactory precision. The density ratio model (DRM) is a semi-parametric model which balances the trade-off between model misspecification risks and statistical efficiency in the presence of multiple related populations. This model allows for related populations to be used alongside the population of interest, which is particularly beneficial when the sample of interest is not large enough. This paper examines the perceived benefit of the DRM quantile estimator. We compare the DRM estimator with the parametric and empirical based estimators based on their bias and variance via simulation. We created five scenarios in terms of sample sizes. We generated data from the normal and gamma distributions, but the distribution information is only used for parametric estimation. We also generated data the same distributions but added some noise. Namely, the model is now mildly misspecified. We confirm that the parametric estimator is preferred under the true model, while DRM estimator is not far behind and it improves substantially over the nonparametric estimator. When the parametric model is wrong, DRM estimator is the overall winner. We also repeated the simulation using data of daily price changes of some technology stocks and found the DRM estimator has the best overall performance.

Presentation 2

Time: 3:30pm – 4:00pm

Speaker: Nathaniel Dyrkton, UBC Statistics MSc student

Title: Integrating representative and non-representative survey data for efficient inference

Abstract: Non-representative surveys are commonly used and widely available but suffer from selection bias that generally cannot be entirely eliminated using weighting techniques. Instead, we propose a Bayesian method to synthesize longitudinal representative and unbiased surveys with non-representative biased surveys by estimating the degree of selection bias over time. We show using a simulation study that synthesizing biased and unbiased surveys together out-performs using the unbiased surveys alone, even if the selection bias may evolve in a complex manner over time. Using COVID-19 vaccination data, we are able to synthesize two large sample biased surveys with an unbiased survey to reduce uncertainty in now-casting and inference estimates while simultaneously retaining the empirical credible interval coverage. Ultimately, we are able to conceptually obtain the properties of a large sample unbiased survey if the assumed unbiased survey, used to anchor the estimates, is unbiased for all time-points.

Bayesian causal inference for discrete data

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: Causal inference provides a framework for estimating how a response changes when a given cause of interest changes. When all data are discrete we can use saturated nonparametric models to avoid unnecessary assumptions in our causal inference modelling, where we specify unique parameters for all possible combinations of treatments and confounders when estimating an outcome. Bayesian methods allow us to incorporate prior information into these saturated models, making them usable beyond simple settings with low dimensional confounders.

We propose two new nonparametric Bayes methods for causal inference based on saturated modelling. The first method combines a parametric model with a nonparametric saturated outcome model to estimate treatment effects in observational studies with longitudinal data. By conceptually splitting the data, we can combine these models while maintaining a conjugate framework, allowing us to avoid the use of Markov chain Monte Carlo methods. Approximations using the central limit theorem and random sampling allows our method to be scaled to high-dimensional confounders.

The second method uses prior restrictions of the parameter space of a saturated model to partially identify causal effect estimates in scenarios with nonignorable missing outcome data. We focus on two common restrictions, instrumental variables and the direction of missing data bias, and investigate how these restrictions narrow the identification region for parameters of interest. Additionally, we propose a rejection sampling algorithm that allows us to quantify the evidence for these assumptions in the data.

Causal clustering: design of cluster experiments under network interference

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: This paper studies the design of cluster experiments to estimate the global treatment effect in the presence of network spillovers. We provide a framework to choose the clustering that minimizes the worst-case mean-squared error of the estimated global effect. We show that optimal clustering solves a novel penalized min-cut optimization problem computed via off-the-shelf semi-definite programming algorithms. Our analysis also characterizes simple conditions to choose between any two cluster designs, including choosing between a cluster or individual-level randomization. We illustrate the method's properties using unique network data from the universe of Facebook's users and existing data from a field experiment.

Link to paper: https://arxiv.org/abs/2310.14983

Finding the Best Player via Multi-Player Comparisons

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: Suppose there are n items, with each item having an unknown value. At any time, we can can either stop and declare which item has the largest value or else choose a subset of items to compare.

If subset S is chosen, then a given item in S will be preferred with a probability equal to the value of that item divided by the sum of the values of all items in S.

Assuming a Bayesian prior on the values, and subject to the proviso that the policy employed will make the correct choice with probability at least some specified value, we are looking for a policy that needs a relatively small mean number of comparisons before making a decision. Some heuristic policies are presented and analyzed.

Extreme Value Modelling with Application to Reverse Stress Testing

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: Reverse stress testing of a financial portfolio aims to identify scenarios for risk factors that lead to a specified adverse portfolio outcome. The stress scenarios of interest naturally need to be extreme yet plausible. A statistical formulation of these requirements is to define a stress scenario at extreme threshold as the mode of the conditional density of the random vector of risk factors given that the loss on the portfolio exceeds the threshold.

In situations where the interest is in an extreme conditioning event corresponding to a large value of the threshold, traditional multivariate mode estimators would not perform well due to very limited sample of observations in the conditioning event. Under this consideration, we propose an estimator based on techniques from multivariate extreme value theory under the assumption of multivariate regular variation and tail dependence. The method effectively addresses data scarcity in the joint tail regions while allowing for more flexible model assumptions focusing on extremes. We study the asymptotic behaviour of the proposed estimator, investigate its finite-sample performance in simulation studies and apply it to real data in a case study.

Modelling orbits to detect exoplanets

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: Determining the orbits of planets is one of the earliest applications of modern science, though nowadays our focus has turned outwards from our solar system to exoplanets orbiting distant stars. To detect and study these planets, we must combine sparse evidence from a variety of sources---direct images of exoplanetary systems, radial velocity curves, interferometer visibilities, transits, and more---to model their physical and orbital parameters. Unfortunately for us, these models can lead to very challenging posteriors that are unidentifiable, multimodal, and can have strong curvature. I will discuss these modelling challenges and present our successes from applying modern sampling techniques (gradient based samplers, and non-reversible parallel tempering) which reduce sampling time from days to seconds, and will enable us to find new exoplanets.