Seminar

Covariate balancing for treatment effects and mediation

Propensity scores have been central to causal inference and are often used as balancing weights. Using estimated propensity score as inverse weights, however, may exhibit undesirable finite-sample performance. Since propensity score is originally proposed by mimicking a randomized trial, and that an important property of a randomization is that confounders are balanced among treatment groups, we construct weights that directly balance the mean and other functionals of the covariate distribution.  We show that the estimators have desirable theoretical and numerical properties. We extend the covariate balancing procedure to mediation analysis.  Since mediators are post-treatment variables that are not balanced even by randomization, a tilted balancing method is needed.

Recent advances in computer experiments: prediction accuracy of kriging and a new framework of calibration

This talk will focus on two problems in computer experiments: 1) error estimates for kriging prediction; 2) a new framework for calibration of computer models.

Kriging based on Gaussian random fields is widely used in reconstructing unknown functions. The kriging method has pointwise predictive distributions which are computationally simple. However, in many applications one would like to predict for a range of untried points simultaneously. In this work we obtain some error bounds for the (simple) kriging predictor under the uniform metric. It works for a scattered set of input points in an arbitrary dimension, and also covers cases where the covariance function of the Gaussian process is misspecified. These results lead to a better understanding of the rate of convergence of kriging under the Gaussian or the Matérn correlation functions, the relationship between space-filling designs and kriging models, and the robustness of the Matérn correlation functions.

The goal of calibration is to identify the model parameters in deterministic computer experiments, which cannot be measured or are not available in physical experiments. In a study of the prevailing Bayesian method proposed by Kennedy and O’Hagan (2001), Tuo-Wu (2015, 2016) and Tuo-Wang-Wu (2017) find that this method may render unreasonable estimation for the calibration parameters. Two novel methods are proposed and proven to enjoy nice properties. In an application example, we study a calibration problem for a composite fuselage simulation. The calibration of computer model parameters is conducted with the help of engineering design knowledge. An effective method is proposed to identify and adjust the important calibration parameters with limited physical experimental data.

Exposure Prediction Error and Spatial Confounding in Air Pollution Epidemiology

Long-term exposure to air pollution has been linked to multiple cardiovascular and respiratory morbidities. Large administrative cohorts, such as the United States Medicaid population, provide an opportunity to investigate these associations on a national scale. To conduct inference about air pollution exposures and the aggregated health data, several challenges must be addressed. Spatial misalignment between pollution monitors and area-level health data require spatial prediction of exposure. In this talk, I present a comparison of approaches for estimating area-level averages of air pollution, when the exposure prediction model is known to be mis-specified. Additionally, the limited amount of individual-level data can lead to unmeasured spatial confounding. I present a method for linking flexible spatial confounding adjustment to spatial scales. These methods are applied to an analysis of long-term particulate matter exposure and asthma morbidity among U.S. children in Medicaid.

The Nested G-formula: A Causal Approach for Analyzing Medical Cost Outcomes

Rising medical costs are of growing importance in guiding health policy decisions. Analyses of longitudinal studies with cost outcomes are often complicated by right-censoring, whereby complete costs are only available on a subset of participants. Existing methods seeking to address this challenge are intent-to-treat in nature, utilizing only baseline treatment status irrespective of any changes in treatment received. It is essential to take time-varying treatment and confounding into account to more adequately meet the goals of health policy guidance and resource allocation. In this talk, we formalize a nested g-computation procedure to target contrasts in marginal means under different hypothetical population-level treatment strategies. Simulations demonstrate that the nested g-computation procedure exhibits a fair amount of robustness to model misspecification. Based on an application to endometrial cancer using SEER-Medicare data, we further demonstrate that the nested g-formula is a flexible framework that can be used to gain insights into overall costs implied by competing policies.

Statistical Methods for The Analysis of Censored Family Data under Biased Sampling Schemes

*Please note unusual start-time of talk.

Studies of the genetic basis for chronic disease often first aim to examine the nature and extend of within-family dependence in disease status. Families for such studies are typically selected using a biased sampling scheme in which affected individuals are recruited from a disease registry, followed by their consenting relatives. This gives right-censored or current status information on disease onset times.  Methods for correcting this response-dependent sampling scheme have been developed for correlated binary data but variation in the age of assessment for family members makes this analysis uninterpretable.  We develop likelihood and composite likelihood methods for modeling within-family associations in disease onset time using copula functions and second-order regression models in which dependencies are characterized by Kendall’s t. Auxiliary data from an independent sample of individuals can be integrated by augmenting the composite likelihood to ensure identifiability and increase efficiency. An application to a motivating family study in psoriatic arthritis illustrates the method and provides evidence of excessive paternal transmission of risk.  Ongoing work on the use of second-order estimating functions, alternative framework for dependence modeling, and approaches to efficient study design will also be discussed.

Normal approximation for recovery of structured unknowns in high dimension: Steining the Steiner formula

Intrinsic volumes of convex sets are natural geometric quantities that also play important roles in applications. In particular, the discrete probability distribution

L(VC) given by the sequence v0,...., vd of conic intrinsic volumes of a closed convex cone C in Rd summarizes key information about the success of convex programs used to solve for sparse vectors, and other structured unknowns such as low rank matrices, in high dimensional regularized inverse problems. The concentration of VC implies the existence of phase transitions for the probability of recovery of the unknown in the number of observations. Additional information about the probability of recovery success is provided by a normal approximation for VC. Such central limit theorems can be shown by first considering the squared length GC of the projection of a Gaussian vector on the cone C. Applying a second order Poincaré inequality, proved using Stein's method, then produces a non-asymptotic total variation bound to the normal for L(GC). A conic version of the classical Steiner formula in convex geometry translates finite sample bounds and a normal limit for GC to that for VC.

Joint with Ivan Nourdin and Giovanni Peccati.   http://arxiv.org/abs/1411.6265

Simplified Power Calculations for Genetic Association Studies with Rare variants

Genome-wide association studies are now shifting focus from an analysis of common to uncommon and rare variants with an anticipation to explain additional variation in complex traits. As power for association testing for individual rare variants may often be low, various aggregate level association tests have been proposed to detect genetic loci that may contain clusters of causal variants. First, we show that these methods can be divided into two classes: tests based on linear and composite statistics (e.g. variance-component tests). Typically, power calculations for such tests require specification of many parameters, making them difficult to use in practice. In this presentation, we approximate power of linear and quadratic tests to varying degree of accuracy using a smaller number of key parameters, including the total genetic variance explained by multiple variants within a locus. Using the simplified power calculation methods, we then develop a mathematical framework to obtain bounds on the genetic architecture of an underlying trait given results from a genome-wide study. By using proposed framework, we observe important implications for lack or a limited number of findings in many currently reported studies. Finally, we provide insights into the required quality of annotation/functional information for identification of likely causal variants to make meaningful improvement in power of subsequent association tests.

Modern Classification with Big Data

Rapid advances in information technologies have ushered in the era of "big data" and revolutionized the scientific research. Big data creates golden opportunities but has also arisen unprecedented challenges due to the massive size and complex structure of the data. Among many tasks in statistics and machine learning, classification has diverse applications, ranging from improving daily life to reaching the new frontiers of science and engineering. This talk will discuss the envisions of broader approaches to modern classification methodologies, as well as computational considerations to cope with the big data challenges. I will present a modern classification method named data-driven generalized distance-weighted discrimination. A fast algorithm with an emphasis on computational efficiency for big data will be introduced. Our method is formulated in a reproducing kernel Hilbert space, and learning theory of the Bayes risk consistency will be developed. In addition, I will use extensive benchmark data applications to demonstrate that the prediction accuracy of our method is highly competitive with state-of-the-art classification methods including support vector machine, random forest, gradient boosting, and deep neural network.

Combinatorial Inference

We propose the combinatorial inference to explore the topological structures of graphical models. The combinatorial inference can conduct the hypothesis tests on many graph properties including connectivity, hub detection, perfect matching, etc. On the other side, we also develop a generic minimax lower bound which shows the optimality of the proposed method for a large family of graph properties. Our methods are applied to the neuroscience by discovering hub voxels contributing to visual memories.

Automated, Scalable Bayesian Inference with Theoretical Guarantees

The automation of posterior inference in Bayesian data analysis has enabled experts and nonexperts alike to use more sophisticated models, engage in faster exploratory modeling and analysis, and ensure experimental reproducibility. However, standard automated posterior inference algorithms are not tractable at the scale of massive modern datasets, and modifications to make them so are typically model-specific, require expert tuning, and can break theoretical guarantees on inferential quality. This talk will instead take advantage of data redundancy to shrink the dataset itself as a preprocessing step, forming a "Bayesian coreset." The coreset can be used in a standard inference algorithm at significantly reduced cost while maintaining theoretical guarantees on posterior approximation quality. The talk will include an intuitive formulation of Bayesian coreset construction as sparse vector sum approximation, an automated coreset construction algorithm that takes advantage of this formulation, strong theoretical guarantees on posterior approximation quality, and applications to a variety of real and simulated datasets.