Seminar

Measuring the Impact of Nonignorability in Intensive Longitudinal EMA Data with Non-Monotone Nonresponse

Modern intensive longitudinal data that collect many occasions of data per subject using mobile devices have been increasingly common nowadays. These data collection methods offer important benefits in measuring various traits in the real-life context, such as moods, physical activity and other human behaviors. A common problem occurring in these studies is potentially nonrandom missing data, such as nonresponse to prompts. It is possible that the nonresponse behavior depends on the unobserved values of interest, leading to nonignorable nonresponses. It is important to perform sensitivity analysis to assess the potential impact of nonignorability on standard inference based on ignorable nonresponse. However a sensitivity analysis that directly fits a range of nonignorable models can be challenging to conduct because the resulting likelihood functions from these nonignorable models involve high dimensional integrations, due to the intensiveness feature, missingness in both outcome and covariates as well as non-monotone missingness patterns. In this work, we consider a leading case where both outcome and covariates subject to missingness follow correlated linear mixed effects models and developed index of local sensitivity to nonignorability (ISNIL and ISNIQ) to measure the nonlinear local sensitivity of the missing at random (MAR) estimators to nonignorability for such intensive longitudinal data.  This method is tractable to use and completely avoids the need to evaluate those high-dimensional integrals associated with nonignorable nonresponses. We derive the formulas for ISNI index measures and evaluate the performance of the method using simulated data, and apply it to a real dataset obtained using the Ecological Momentary Assessment (EMA) method.

Adjusting for Bias Induced by Informative Dose Selection Procedures

Many fields such acute toxicity studies, Phase I cancer trials, sensory studies and psychometric testing use informative dose allocation procedures. In this talk, we explain how such adaptive designs induce bias, and in the context of dose-finding designs we show how to modify frequency data to adjust for this bias.

To provide context, we start the talk with a general discussion of issues in inference following adaptive designs. Then, we assume a binary response Y has a monotone positive response prob- ability to a stimulus or treatment X, and we consider designs that sequentially select X values for new subjects in a way that concentrates treatments in a certain region of interest under the dose-response curve. We discuss how data analysis at the end of a study is affected by choosing the stimulus value for each subject sequentially according to some informative sampling rule.

Without loss of generality, we call a positive response a toxicity and the stimulus a dose. For simplicity, we restrict this talk to the case of a univariate treatment X and binary Y, and further assume that treatments are limited to a finite set {d1, d2, . . . , dM } of M values we call doses. Now suppose n subjects receive treatments that were sequentially selected (according so some rule using data from prior subjects) from the restricted set of M doses.  Let Nm and Tm denote the number of subjects receiving treatment dm and the number of toxicities observed on treatment dm, respectively. Define Fm = P{Y = 1|X = dm} = E[Y |X = dm].

Then it is often said that the distribution of Tm given Nm is Binomial with parameters (Fm, Nm). But taking Nm as fixed is not the same as conditioning on this random variable, and conditioning on informative dose assignments is not the same as conditioning on summary dose frequencies. Indeed, it is easy to show that the observed dose-specific toxicity rate, Tm/Nm, is biased for Fm. From first principals, we obtain

 

E[Tm / Nm] = Fm - Cov[Tm/Nm, Nm] / E[Nm]

 

The observed toxicity rate is biased for Fm because adaptive allocations, by design, induce a correlation between toxicity rates and allocation frequencies.

This bias impacts inference procedures: Isotonic regression methods use dose-specific toxicity rates directly. Standard likelihood-based methods mask the bias by providing first-order linear approximations. We illustrate these biases using isotonic and likelihood-based regression methods in some well known (small sample size) adaptive methods including selected up-and-down designs, interval designs, and the continual reassessment method. Then we propose a bias adjustment inspired by Firth (1993).

 

[Nancy Flournoy; University of Missouri – http://web.missouri.edu/flournoyn/]

[flournoyn@missouri.edu  –   https://en.wikipedia.org/wiki/Nancy_Flournoy]

Spatial extremes: a conditional approach

The past decade has seen a huge effort in modelling the extremes of spatial processes. Significant challenges include the development of models with an appropriate asymptotic justification for the tail; ensuring model assumptions are compatible with the data; and the fitting of these models to (at least reasonably) high-dimensional datasets. I will review basic ideas of modelling spatial extremes, and introduce an approach based on the (multivariate) conditional extreme value model of Heffernan and Tawn (2004) and Heffernan and Resnick (2007). Advantages of the conditional approach include its asymptotic motivation, flexibility, and comparatively simple inference, meaning that it can be fitted to reasonably large datasets. The modelling approach is applied to understand the spatial extent of high temperature extremes across Australia.

Methods for Preferential Sampling in Geostatistics

 

Preferential sampling in geostatistics refers to the instance in which the process that determines the sampling locations may depend on the spatial process that is being modelled. If ignored, this dependency can result in biased parameter estimates and may affect the resulting spatial prediction. Recent research on correcting for preferential sampling bias has been limited to stationary sampling locations, such as air-quality monitoring sites. We propose a flexible framework for inference on preferentially sampled fields, which can be used to expand preferential sampling methodology to the case in which the preferentially sampled locations are obtained from a process moving in space and time. An example of such data, the preferential sampling of ocean temperature by tagged marine mammals, is presented.

 

 

Approximate Bayesian Forecasting

Approximate Bayesian Computation (ABC) has become increasingly prominent as a method for conducting parameter inference in a range of challenging statistical problems, most notably those characterized by an intractable likelihood function. In this paper, we focus on the use of ABC not as a tool for parametric inference, but as a means of generating probabilistic forecasts; or for conducting what we refer to as approximate Bayesian forecasting. The four key issues explored are: i) the link between the theoretical behavior of the ABC posterior and that of the ABC-based predictive; ii) the use of proper scoring rules to measure the (potential) loss of forecast accuracy when using an approximate rather than an exact predictive; iii) the performance of approximate Bayesian forecasting in state space models; and iv) the use of forecasting criteria to inform the selection of ABC summaries in empirical settings. The primary finding of the paper is that ABC can provide a computationally efficient means of generating probabilistic forecasts that are nearly identical to those produced by the exact predictive, and in a fraction of the time required to produce predictions via an exact method.

UBC Statistics M.Sc. Co-op Student Presentations (2)

4:00pm - 4:30pm:  Boyi Hu, UBC Statistics M.Sc. student

Title:  An R package for monitoring test under density ratio model and its applications

Abstract:  Quantiles and their functions are important population characteristics in many applications. In forestry, lower quantiles of the modulus of rapture and other mechanical properties of the wood products are important quality indices. It is important to ensure that the wood products in the market over the years meet the established industrial standards. Two well-known risk measures in finance and hydrology, value at risk (VaR) and median shortfall (MS), are extreme quantiles of their corresponding marginal distributions. Chen et al. [2016] developed an empirical likelihood approach based on density ratio model and multiple samples. Following their work, we build a user-friendly R package to make their methods easy-to-use for practitioners. The package also includes some diagnostic tools to allow users to investigate the goodness of the fit of the density ratio model. With the help of this package, we study the performance of DRM CEL-based inference with clustered data with possibly different cluster sizes.

******

4:30pm - 5:00pm:  Wayne Wang, UBC Statistics M.Sc. student

Title:  Applying record value theory in combinatorial optimization with application to environmental statistics

Abstract:  We consider the problem of optimal subset selection from a set of correlated random variables. In particular, we consider the associated combinatorial optimization problem of maximizing the determinant of a symmetric positive semidefinite matrix that characterizes the chosen subset. This problem arises in many domains, such as experimental designs, regression modelling, and environmental statistics. In this thesis, we attempt to establish an efficient polynomial-time algorithm for approximating the optimal solution to the problem. Firstly, we employ determinantal point processes, a special class of spatial point processes, to develop an easy-to-implement sampling-based stochastic search algorithm for the task of finding approximations to the combinatorial optimization problem. Secondly, we establish theoretical tools for assessing the quality of those approximations using statistical results from record value theory, the study of record values and related statistics from a sequence of observations. 

R-vine copula based quantile regression

Quantile regression—the prediction of conditional quantiles—has steadily gained importance in statistical modeling. Using D-vine copulas, which are built from arbitrary bivariate (conditional) copulas, Kraus and Czado (2017) propose a novel approach for quantile regression, which automatically takes typical issues such as quantile crossing or transformations, interactions and collinearity of variables into account. Their algorithm is based on sequentially fitting a likelihood optimal D-vine copula to given data. D-vine copulas are not only a very flexible class of copulas, their construction principle further allows for easy extractability of the conditional quantiles. We build upon this work and develop methodologies for the general class of regular vine (R-vine) copulas. As opposed to D-vine copulas, where the underlying vine tree sequence follows a line structure, R-vine structures are not restricted and thus allow for even more complex dependencies between variables. We propose an algorithm that sequentially fits an optimal R-vine copula to given data. Due to the enormous amount of possible R-vine structures covariate selection via maximizing the conditional likelihood as suggested in Kraus and Czado (2017) is no longer feasible. We propose a partial correlation based selection approach, which computationally is significantly less demanding. The developed estimation and prediction approach will be presented along with an extensive simulation study to (1) show the good finite sample performance of the proposed methodology and (2) to compare the novel approach to benchmark models such as D-vine copula quantile regression and linear quantile regression. A real data example on DAX-stock returns will be discussed.

References: Kraus, D. and Czado, C.: D-vine copula based quantile regression. Computational Statistics and Data Analysis (110), 2017, 1-18.

20 Years of Data Science: Music, Genomics, and Hurricane Maria

The UBC Data Science Institute is delighted to have Dr. Rafael Irizarry speak in our DSI Distinguished Seminar Series on August 1, 2018. 

The talk will take place in Hugh Dempster Pavilion, room 110 at 4:00PM. We look forward to you joining us. Details below.

Speaker:  Dr. Rafael Irizarry (Professor, Applied Statistics, Harvard & Dana-Faber Cancer Institute) 

 

Title: 20 Years of Data Science: Music, Genomics, and Hurricane Maria

 

Abstract: As new academic data science departments and training programs are created, statisticians are left wondering what role our discipline plays in this emerging field. In this talk I will share my related thoughts via examples from my own work that I now describe as data science projects. I will highlight both the statistical insights and the important considerations that fall outside the realm of the current scope of the statistics disciple. The first example relates to the analysis of musical sound signals. I will describe how locally harmonic models can be used to provide meaningful parameters useful for manipulation of sounds. The second example relates to a biological discovery enabled by statistical reasoning. Finally, I will describe a project in which we collected and analyzed data to describe the effects of Hurricane Maria in Puerto Rico.

 

Bio: Rafael Irizarry received his Bachelor’s in Mathematics in 1993 from the University of Puerto Rico and went on to receive a Ph.D. in Statistics in 1998 from the University of California, Berkeley. He is now Professor of Biostatistics and Computational Biology at the Dana-Farber Cancer Institute and a Professor of Biostatistics at Harvard School of Public Health. Since 1999, Rafael Irizarry’s work has focused on Genomics and Computational Biology problems. In particular, he has worked on the analysis and signal processing of microarray, next-generation sequencing, and genomic data. He is currently interested in leveraging his knowledge in translational work, e.g. developing diagnostic tools and discovering biomarkers. Source: https://rafalab.github.io/pages/about.html

 

If you are interested to attend, please register for Dr. Rafael Irizarry’s lecture (on August 1, 4:00PM) here: https://dsi-lecture-irizarry.eventbrite.ca.

 

There will be a small reception after his talk so please RSVP to save your spot and ensure we order the appropriate amount of snacks and drinks. Details about the talk below and at our webpage (https://dsi.ubc.ca/dsi-distinguished-seminar-dr-rafael-irizarry)

Credit Risk Classification using Statistical and Machine Learning Methods

Credit ratings present a rating agency or a bank’s assessment of a company’s risk profile. It is used extensively by business practitioners to judge the company’s credit worthiness, to determine whether or not to grant credit, and to calculate the interest rate for loans. Under BASEL II, banks are allowed to use advanced internal rating-based approaches to calculate credit risks themselves if said approaches comply with certain supervisory standards. This has sparked an interest in developing statistical credit classification models that can produce accurate ratings quickly and interpretably. We focus on the classification accuracy and interpretability of four classification methods on a credit portfolio of small Canadian businesses. The four methods that we compare are ordinal regression, ordinal gradient boosting, multinomial gradient boosting and random forest.

Estimations of long-run covariance and long-memory parameter in stationary functional time series

Abstract:  In arenas of application including environmental science, economics, and medicine, it is increasingly common to consider time series of curves or functions. Many inferential procedures employed in the analysis of such data involve the long-run covariance function, which is analogous to the long-run covariance matrix familiar to finite-dimensional time series analysis and econometrics. I present a kernel sandwich estimator for estimating the long-run covariance. From estimated long-run covariance, I study the estimation of a long-memory parameter in a long-range dependent stationary functional time series, and identify the most accurate estimation method via a series of simulation studies.

Biography:  Han Lin Shang is an Associate Professor of Statistics at the Research School of Finance, Actuarial Studies and Statistics, Australian National University. His research interests include actuarial studies, computational statistics, demographic forecasting and empirical finance. He is serving as an associate editor for Journal of Computational and Graphical Statistics and Australian & New Zealand Journal of Statistics.