Seminar

Bayesian Regression Trees, Nonparametric Heteroscedastic Regression Modeling and MCMC Sampling

Bayesian additive regression trees (BART) have become increasingly popular as flexible and scalable non-parametric models useful in many modern applied statistics regression problems. They bring many advantages to the practitioner dealing with large datasets and complex non-linear response surfaces, such as the matrix-free formulation and the lack of a requirement to specify a regression basis a priori. However, there are some known challenges to this modeling approach, such as poor mixing of the MCMC sampler and inappropriate uncertainty intervals when the assumed homoscedastic variance model is violated.  In this talk, weintroduce a new Bayesian regression tree model that allows for possible heteroscedasticity in the variance model and devise novel MCMC samplers that appear to adequately explore the posterior tree space of this model.

Multi-resolution spatial methods for large data sets (joint EOS,SCAIM,PIMS,STAT)

Spatial data is ubiquitous arising in numerous areas in the geophysical and environmental sciences. A basic problem for statisticians is to estimate complete surfaces from irregular observations or measurements and to quantify the uncertainty in the result. However, standard  statistical methods break when applied to large data sets and so alternative approaches are needed that balance shortcuts in the statistical models for increases in computational efficiency. A useful method expands the surface in a set of compact basis functions and places a Markov random field model on the basis coefficients. The impact is that evaluating the model likelihood and computing spatial predictions is feasible even for tens of thousands of spatial observations on a single computational core (e.g. a laptop). Moreover, by varying the support of the basis functions and the correlations among basis coefficients it is possible to entertain multi-resolution and non-stationary spatial models that mirror the rich covariance structure often found in large geophysical data sets.

See: A multi-resolution Gaussian process model for the analysis of large spatial data sets.

D Nychka, S Bandyopadhyay, D Hammerling, F Lindgren, S Sain (2014) Journal of Computational and Graphical Statistics (In press).

Reconstructing carbon dioxide for the last 2000 years: a hierarchical success story

Knowledge of atmospheric carbon dioxide (CO2) concentrations in the past are important to provide an understanding of how the Earth's carbon cycle varies over time. This project combines ice core CO2 concentrations, from Law Dome, Antarctica and a physically based forward model to infer CO2 concentrations on an annual basis. Here the forward model connects concentrations at given time to their depth in the ice core sample and an interesting feature of this analysis is a more complete characterization of the uncertainty in "inverting" this relationship. In particular, Monte Carlo based ensembles are particularly useful for assessing the size of the decrease in CO2 around 1600 AD. This reconstruction problem, also known as an inverse problem, is used to illustrate a general statistical approach where observational information is limited and characterizing the uncertainty in the results is important. These methods, known as Bayesian hierarchical models, have become a mainstay of data analysis for complex problems and have wide application in the geosciences.

This work is in collaboration with Eugene Wahl (NOAA), David Anderson (NOAA) and Catherine Trudinger (CSIRO).

Detecting Rate Changes in Point Processes

Speaker Bio
Dr. Michael Messer completed his doctorate in Mathematics with specialization in Statistics at the department of Computer Science and Mathematics at Goethe University Frankfurt, Germany. His thesis was supervised by Prof. Gaby Schneider and reported by Prof. Anton Wakolbinger (both Frankfurt) and Prof. Roland Fried from Dortmund. The research was conducted in collaboration with neuroscientists in an interdisciplinary research project for the investigation of neuronal disorders (NeFF - Neuronal Coordination Research Focus Frankfurt). He developed statistical techniques for the analysis of neuronal spike trains. The main paper was recently accepted for publication at The Annals of Applied Statistics.
Abstract

Nonstationarity of the event rate is a persistent problem in modeling time series of events, such as neuronal spike trains. Motivated by a variety of patterns in neurophysiological spike train recordings, we de fine a general class of renewal processes. This class is used to test the null hypothesis of stationary rate versus a wide alternative of renewal processes with finitely many rate changes (change points). Our test extends ideas from the filtered derivative approach by using multiple moving windows simultaneously. We also develop a multiple filter algorithm, which can be used when the null hypothesis is rejected in order to estimate the number and location of change points. We analyze the benefi ts of multiple filtering and its increased detection probability as compared to a single window approach. Application to spike trains recorded from dopamine midbrain neurons of anesthetized mice illustrates the relevance of the proposed techniques as preprocessing steps for methods that assume rate stationarity.

Early Alert Orientation for Faculty, TAs, and Staff

In this workshop, Patty Hambler (Acting Associate Director, Strategic Initiatives & Special Projects, Student Development & Services) will teach: - how Early Alert simplifies the process for faculty, TAs, and staff to connect students of concern with campus resources and supports - the kinds of concerns that are appropriate to submit within Early Alert; and - how Early Alert protects student privacy and maintains confidentiality.

Early Alert

Supporting student learning and success is a priority for UBC.

Early Alert helps achieve this goal by helping faculty and staff provide more comprehensive support for students who are facing difficulties that put their academic success at risk.

Earlier support to get back on track

With Early Alert, faculty and staff can identify their concerns about students sooner and in a more coordinated way. This gives students the earliest possible connection to the right resources and support, before difficulties become overwhelming.

Development of the Bayesian approach in Canadian fisheries stock assessment

Many of the world's harvested fish stocks are managed with the use of results from fish stock assessments.  The penultimate goal of fish stock assessment is to evaluate the potential consequences of alternative management options.  The bulk of the analytical effort is typically focused on formulating credible models of fish population dynamics, compiling data and fitting the models to data to estimate model parameters and management quantities of interest.  Trends in fish stock abundance and fishing mortality rates are evaluated and the models are projected to evaluate the potential consequences of different management options.  Since the mid-1990s applications of the Bayesian statistical approach to fish stock assessment have been increasing and in recent years applications have become commonplace.  Debates about whether the Bayesian is appropriate for fisheries stock assessment have moved on to debates over how the approach should be applied.  In this talk I review recent developments of the Bayesian approach in Canadian fisheries stock assessment with a focus on some of my recent applications to rockfish stocks that have been designated as threatened and endangered.  I will highlight some of the chief merits of the approach but also problems commonly experienced with its application.

Degree distribution of shortest path trees and bias in network sampling algorithms

In this talk, we investigate the degree distribution of shortest path trees of various weighted network models. The aim of many empirical studies is to determine the degree distribution of a network with unknown structure by using trace-route sampling. We derive the limiting degree distribution of the shortest path tree from a single source on various random network models with edge weights: the configuration model and r-regular graphs with i.i.d. power law degrees and i.i.d. edge weights, the complete graph with edge weights that are powers of i.i.d. exponential random variables. We use these results to shed light on an empirically observed bias in network sampling methods.

Estimation under nearly-correct models

When additional variables are measured on a subsample from an existing cohort, we can fit models using survey-sampling approaches.  In some settings we can also estimate by semiparametric maximum likelihood, profile likelihood, or similar techniques, leading to a semi parametric efficient estimator. For example, under case-control sampling we can use weighted logistic regression or unweighted logistic regression.  The common wisdom is that the weighted estimators are inefficient because of the variation in the weights. I will argue that this is neither true nor helpful. By considering contiguous model misspecification I will show that the efficient estimator gains its extra precision from relying more heavily on the model, and this is true in a quantitative sense, not merely as a heuristic.

Approximate models and robust decisions

Statistical decisions based partly or solely on predictions from probabilistic models may be sensitive to model mis-specification. Statisticians are taught from an early stage that all "models are wrong" but little formal guidance exists on how to assess the impact of model approximation, or how to proceed when optimal actions appear sensitive to model fidelity. In this talk I will present a general applied framework to address this issue. This builds on diagnostic techniques, including graphical approaches and summary statistics, to help highlight decisions made through minimized expected loss that are sensitive to model mis-specification. The stability of decision-systems can then be assessed by considering perturbations within a neighborhood of model space centered at the (mis-specified) approximating model. This neighborhood can either be defined via an information (Kullback-Leibler) divergence, or using the Dirichlet Process as a non-parametric extension to the model. A Bayesian approach is adopted throughout, although the methods are agnostic to this position.

Causal Inference Approaches for Dealing with Time-dependent Confounding in Longitudinal Studies, with Applications to Multiple Sclerosis Research

Marginal structural Cox models (MSCMs) have gained popularity in analyzing longitudinal data in the presence of  'time-dependent confounding'. This talk is motivated by issues arising in connection with dealing with time-dependent confounding while assessing the effects of beta-interferon drug exposure on disease progression in relapsing-remitting multiple sclerosis (MS) patients in the real-world clinical practice setting. In the context of this chronic, yet fluctuating disease, MSCMs were used to adjust for the time-varying confounders, such as MS relapses, as well as baseline characteristics, through the use of inverse probability weighting (IPW). Using a large cohort of relapsing-remitting MS patients in British Columbia, Canada (1995-2008), no strong association between beta-interferon exposure and the hazard of disability progression was found. We also investigated whether it is possible to improve the MSCM weight estimation techniques by using the statistical learning methods, such as bagging, boosting and support vector machines. Statistical learning methods require fewer assumptions and have been found to estimate propensity scores with better covariate balance. As propensity scores and IPWs in MSCM are functionally  related, we also studied the usefulness of statistical learning methods via a series of simulation studies. Additionally, two alternative approaches, prescription time-distribution matching (PTDM) and the sequential Cox approach, proposed in the literature to deal with immortal time bias and time-dependent confounding respectively, were compared via a series of simulations. The PTDM approach was found to be not as effective as the Cox model (with treatment considered as a time-dependent exposure) in minimizing immortal time bias. The sequential Cox approach was, however, found to be an effective method to minimize immortal time bias, but not as effective as a MSCM, in the presence of time-dependent confounding. These methods were used to re-analyze the MS dataset to show their applicability. The findings from the simulation studies were also used to guide the data analyses.