Seminar

Statistical design and analysis of response adaptive clinical trials

Speaker's Page

Abstract:  Clinical trials are regarded as the most reliable way to evaluate the efficacy of new medical interventions. This practice has taken a prominent role in modern clinical research. However clinical experimentation on human subjects requires a careful balancing act between the benefit of the collective and the benefit of the individual.  This talk is focused on statistical design and analysis of response adaptive Phase III clinical trials, which represent recent advancements in clinical trial methodology. The adaptive designs help balance the ethical issues and improve efficiency without undermining the validity and integrity of the clinical research. The talk is based on joint work with several graduate students.

Hassle-free Allocation of Risk Capital and Economic Pricing in Insurance: the Weighted Insurance Pricing Model

Speaker's Page

Abstract:  Determining aggregate risk capital has become a fundamental
problem in modern Enterprise Risk Management, and the determination process
has been fairly well studied. The consequent exercise of allocating the
aggregate risk capital to constituents has also been given high priority in both
life and general insurance in such contexts as pricing, performance management,
and pro profitability testing. In fact, the allocation exercise has been often
called the primary driver for calculating the aggregate risk capital.

Unfortunately, the allocation problem is, in general, noticeably more involved
than the problem of calculating the aggregate risk capital. In fact, the former
problem is not easy even when a specific c risk measure that induces the
allocation rule has been assumed, let alone when a class of risk measures is
considered. In this talk I will demonstrate that quite often, the problems of
determining and allocating the aggregate risk capital are of a similar complexity.
Remarkably, this turns out to be the case for the entire class of weighted
risk capital allocations, as well as for the risk portfolios having dependence
structures of the popular multiplicative and additive background risk models.

Bayesian Uncertainty Quantification for Solutions to Differential Equation Models

Speaker's Page

Abstract:  Differential equation models offer a succinct way of describing rates of change using few but readily interpretable parameters.   In most interesting cases analytic solutions do not exist and numerical methods are used to approximate a solution over a discretization grid.  We explore the use of probability models for uncertainty arising from the discretization of ordinary or partial differential equation solutions. Viewing the system solution as an inference problem allows us to quantify numerical uncertainty using the tools of smoothing and gaussian process regression.   A formalism for inferring differential equation model trajectories is proposed through a Bayesian updating scheme based on interrogations of the model derivatives.   The approach provides estimates for the differential equation model solution and derivatives thereof which can be incorporated into the inference problem.  The proposed approach is demonstrated to capture the functional structure and magnitude of the discretization error, while attaining computational scaling of the same order as standard numerical solver methods.  This allows us to define a trade-off between accuracy and discretization grid size.  Our approach is demonstrated on ordinary and partial differential equation models, ill-conditioned mixed boundary value problems, and delay differential equations.  This talk represents joint work with Oksana Chkrebtii, Mark Girolami, and Ben Calderhead and appeared as a discussion paper in Bayesian Analysis in December 2016.

Statistical Inference with Estimating Functions via the MapReduce Scheme

Speaker's Page

Abstract:  The theory of statistical inference along with the strategy of divide-and-combine for large-scale data analysis has recently attracted considerable interest due to great popularity of the MapReduce scheme in the Hadoop platform.  The key to the development of statistical inference lies in the method of combining results yielded from separately mapped data batches.  One seminal solution based on the confidence distribution has been proposed in the setting of maximum likelihood estimation in the literature. We consider a more general inferential methodology based on estimating functions, of which the maximum likelihood is a special case.  This generalization allows us to perform regression analyses of massive complex data via the MapReduce scheme, such as longitudinal data, survival data and quantile regression, which cannot be done using the maximum likelihood method.  The proposed statistical inference inherits many key large-sample properties of estimating functions. In addition, because the proposed method is closely connected to the generalized method of moments (GMM) and Crowder’s optimality, its optimality over the existing methods is conveniently verified.  Our method provides a unified framework for many kinds of statistical models and data types, which is illustrated via numerical examples in both simulation studies and real-world data analyses.  

This is a joint work with Ling Zhou.

A Theory of Experimenters

*Note the unusual location of this talk.*

Speaker's Page:  http://economics.ubc.ca/faculty-and-staff/erik-snowberg/.

UBC news-release on Dr. Snowberg: http://news.ubc.ca/2016/03/15/esteemed-economist-joins-ubc-as-new-canada-excellence-research-chair/

Abstract:  This paper proposes a decision-theoretic framework for experiment design. We  model experimenters as ambiguity averse decision-makers, who trade-off subjective expected performance, and robustness.  This framework suitably accounts for experimenters' preferences for randomization, and the circumstances in which randomization occurs: whenever available sample size becomes large enough. We illustrate the practical value of such a framework by studying the issue of rerandomization. We show that rerandomization creates a trade-off between subjective performance and robustness but that loss in robustness due to rerandomization grows very slowly with the number of assignment draws.

Forecasting Extremes using Copula Models

Forecasts of extreme events are useful in order to prepare for disaster. Such forecasting can be achieved by existing methods in quantile regression, but these methods cannot capture nonlinearity in the predictors, and cannot capture the distributional tail behaviour. To address these issues, I introduce a method that uses copulas to build nonlinear models, which are fit using the proposed composite nonlinear quantile regression. This new approach is more able to capture the effect that predictors have on a response, and allows for probabilistic forecasts to be issued in the form of the predictive distribution's tail. We end up with a tool that can be used as an early warning system, addressing questions like "how bad could it get?", and is applied to forecast flooding of the Bow River in Alberta.

Robust Estimation of Multivariate Location and Scatter under Cellwise and Casewise Contamination

In traditional robust statistics, it is generally assumed that the majority of the observations in the data are free of contamination, while only a minority of the observations are contaminated. The contaminated observations are flagged as outliers and down-weighted even if only a single component is contaminated. In practice, observations can be entirely contaminated. This situation usually refers to casewise contamination. However, observations can also be only partially contaminated. This type of contamination often appears as single outlying cells in a data matrix and therefore, usually refers to cellwise contamination. Under cellwise contamination, a lot of information could be lost through downweighting the whole observation, especially for high-dimensional data. Furthermore, recent work has shown that traditional robust procedures that proceed in such way may fail when applied to such datasets. In this talk, I will present this problem using a real data example and sketch out our proposal when the goal is to estimate multivariate location and scatter matrix under simultaneous cellwise and casewise contamination. 

Hierarchical clustering of observations and features for high-dimensional data

In this talk, we present new developments of hierarchical clustering in high-dimensional data. We consider clustering both the observations and the features. We first focus on the clustering of observations. In high-dimensional data, the existence of potential noise features and outliers poses unique challenges to the existing hierarchical clustering techniques. We propose the robust sparse hierarchical clustering (RSHC) and the multi-rank sparse hierarchical clustering (MrSHC) to address these challenges. We then consider clustering of features in high-dimensional data. We propose a new hierarchical clustering technique to divide the large number of features into subgroups called regression phalanxes. The regression phalanxes are used for building base regression models for further ensembling. We show that the ensemble of regression phalanxes resulting from the hierarchical clustering produces further gains in prediction accuracy when applied to an effective method like Lasso or Random Forests.

Selective inference in linear regression

We consider inference after model selection in linear regression problems, specifically after fitting the LASSO. A classical approach to this problem is data splitting, using some randomly chosen portion of the data to choose the model and the remaining data for inference in the form of confidence intervals and hypothesis tests. Viewing this problem in the framework of selective inference, conditional on a selection event, we describe other randomized algorithms with similar guarantees to data splitting, at least in the parametric setting. Time permitting, we describe analogous results for statistical functionals obeying a CLT in the classical fixed dimensional setting.

Stochastic order of sampling plans in sample theory

Sampling plans form a crucial part of sample survey theory in the study finite populations. Statisticians are generally concerned about the efficiency of certain point estimators of some simple finite population parameters such as population mean and total. The efficiency of an estimator is highly related to the underlying sampling plan. We are therefore interested in ranking the sampling plans according to the efficiencies of their corresponding point estimators. Motivated by this fact, we introduce notion of stochastic order to sampling plans in the context of sample survey. We focus on comparing various sequential sampling plans such as successive sampling plan, rejective sampling plan and multinomial sampling plan with the same selection probability vector (which is also a distribution on the finite population). We show that in general, the stochastic order in selection probability vector leads to the same stochastic order in the corresponding inclusion probability vector. Under certain conditions on the selection probability vector, we also show some without-replacement sequential sampling plans leads to more efficient point estimators of the population total than the one based on the with replacement sampling plan.