Seminar

Preferential attachment and neutral random graphs: statistically useful generative models of network data

 

Preferential attachment (PA) and other probabilistic generative models of network growth have been popular for their ability to explain large-scale phenomena from simple interaction mechanisms. However, PA has been of limited use as a statistical model, due to its lack of exchangeability: in a statically observed network with n edges, inference requires considering all n! possible edge arrival orders. Moreover, in models based on forms of exchangeability, inference algorithms benefit from an edge-decoupled representation, in which all dependence between edges is captured by some latent quantity; no such representation is known for PA models. I will describe my work toward making PA useful as a statistical model: an edge-decoupled representation for a class of generalized PA models is established, and it reveals probabilistic structure, called left-neutrality, that can be exploited for efficient inference algorithms even in the presence of unknown edge arrival order. Furthermore, the edge-decoupled representation endows the PA model with a set of interpretable model parameters. Finally, I will describe how exchangeability still plays a role, despite PA's non-exchangeability.

 

This work was done in collaboration with Christian Borgs, Jennifer T. Chayes, Adam Foster, Emile Mathieu, Peter Orbanz, and Yee Whye Teh.

Support points – a new way to reduce big and high-dimensional data

This talk presents a new method for reducing big and high-dimensional data into a smaller dataset, called support points (SPs). In an era where data is plentiful but downstream analysis is oftentimes expensive, SPs can be used to tackle many big data challenges in statistics, engineering and machine learning. SPs have two key advantages over existing methods. First, SPs provide optimal and model-free reduction of big data for a broad range of downstream analyses. Second, SPs can be efficiently computed via parallelized difference-of-convex optimization; this allows us to reduce millions of data points to a representative dataset in mere seconds. SPs also enjoy appealing theoretical guarantees, including distributional convergence and improved reduction over random sampling and clustering-based methods. The effectiveness of SPs is then demonstrated in two real-world applications, the first for reducing long Markov Chain Monte Carlo (MCMC) chains for rocket engine design, and the second for data reduction in computationally intensive predictive modeling.

Connections between optimization and sampling

 

In this talk, I am going to give a brief summary of two of my recent papers that exploit connections between optimization and sampling. "Hamiltonian descent methods" greatly extends the class of convex functions that can be optimized with linear rates compared to existing methods. This new optimization method is based on conformal Hamiltonian dynamics, and it was inspired by the literature on Hamiltonian MCMC methods. Besides being more efficient, this method is also more reliable and requires less tuning than existing techniques such as gradient descent and heavy ball method. "Randomized Hamiltonian Monte Carlo as Scaling Limit of the Bouncy Particle Sampler and Dimension-Free Convergence Rates" studies the high dimensional behavior of the Bouncy Particle Sampler (BPS), a non-reversible piecewise deterministic MCMC method. Although the paths of this method are straight lines, we show that in high dimensions they converge to a Randomised Hamiltonian Monte Carlo (RHMC) process, whose paths are determined by the Hamiltonian dynamics. We also give a characterization of the mixing rate of the RHMC process for log-concave target distributions that can be used to tune the parameters of BPS.

Mendelian randomization: A comprehensive statistical approach and applications to preventing heart disease

 

Mendelian randomization (MR) can give unbiased estimate of a confounded causal effect by using genetic variants as instrumental variables (IV). The summary-data MR design is rapidly gaining popularity in practice due to the increasing availability of large-scale genome-wide association studies (GWAS). As we are entering the "MR of every risk factor on every disease outcome" era, existing statistical methods still lack theoretical grounding and face at least four major challenges: measurement error in the genetic associations, invalid IVs due to pleiotropy, weak IV bias, and selection bias IV screening.

To overcome these challenges, I will formulate the summary-data MR problem as a linear errors-in-variables regression problem with over-dispersion and occasional outliers. This means that none of the genetic IVs is strictly valid. This model is inspired by our exploratory data analysis and the recent omnigenic model for complex traits. I will present a new approach based on adjusting and robustifying the profile score function, with provable consistency and asymptotic normality when the IVs are collectively strong but may be individually weak. The efficiency of this method can be further increased by empirical (partially) Bayes shrinkage. The new methods will be used to re-analyze several cardiometabolic diseases and risk factors, yielding new insights into the role of HDL particles (the "good" cholesterol) in coronary artery disease.

This talk is based on joint works with Jingshu Wang, Nancy Zhang, Dylan Small (University of Pennsylvania); Jack Bowden, Gibran Hemani, George Davey Smith (University of Bristol); Yang Chen (University of Michigan).

Non-iterative Estimation Update for Parametric and Semiparametric Models with Population-based Auxiliary Information

With the advancement in disease registries and surveillance data, population-based
information on disease incidence, survival probability or other important biological
characteristics become increasingly available. Such information can be leveraged in
studies that collect detailed measurements but with smaller sample sizes. In contrast
to recent proposals that formulate the additional information as constraints in opti-
mization problems, we develop a general framework to construct simple estimators that
update usual regression estimators with some functionals of data and the initial esti-
mator based on the additional information. We consider general settings which include
nuisance parameters in the auxiliary information, non-i.i.d. data such as case-control
sampling, and semiparametric models with infinite dimensional parameters. Detailed
examples of several important data and sampling settings are provided.

Sparse Generalized Eigenvalue Problem and Its Application to Multivariate Statistics

Sparse generalized eigenvalue problem (GEP) plays a pivotal role in a large family of high-dimensional learning tasks, including sparse Fisher’s discriminant analysis, canonical correlation analysis, and sufficient dimension reduction. Most of the existing methods and theory in the context of specific statistical models that are special cases of sparse GEP require restrictive structural assumptions on the input matrices.  This talk will focus on a two-stage computational framework for solving the non-convex optimization problem resulting from the sparse GEP.  At the first stage, we solve a convex relaxation of the sparse GEP.  Taking the solution as an initial value, we then exploit a non-convex optimization perspective and propose the truncated Rayleigh flow method (Rifle) to estimate the leading generalized eigenvector, and show that it converges to a solution with the optimal statistical rate of convergence. Theoretically, our method significantly improves upon the existing literature by eliminating the structural assumptions on the input matrices.  Numerical studies in the context of several statistical models are provided to validate the theoretical results.  We then apply the proposed method to an electrocorticography data to understand how human brains recall and mentally rehearse word sequences.

Bayesian adjustments for disease misclassification in epidemiological studies of health administrative data

Canadian health administrative databases are a popular and rich data source for epidemiological research at the population level, but error-prone diagnostic information typically provide investigators with a less-than-perfect proxy for the disease under study. Motivated by an observational study into the prodromal phase of multiple sclerosis, this talk considers the setting of a matched exposure-disease association study where the disease variable is measured with error, and participant selection is based upon the error-prone disease label rather than the true disease status. We initially focus on the special case of a pair-matched case-control study. Assuming non-differential misclassification of study participants, we give a closed-form expression for asymptotic biases in odds ratios arising under naive analyses of misclassified data, and propose a Bayesian model to correct association estimates for misclassification bias. For identifiability, the model relies on information from a validation cohort of correctly classified case-control pairs, and also requires prior knowledge about the predictive values of the classifier. In a simulation study, the model shows improved point and interval estimates relative to the naive analysis, but is also found to be overly restrictive for our motivating dataset. In light of these concerns, we further propose a generalized model for misclassified data that extends to the case of differential misclassification and allows for a variable number of controls per matching stratum. Instead of prior information about the classification process, the model relies on individual-level estimates of each participant’s true disease status, which were obtained from a counting process mixture model of MS-specific healthcare utilization in our motivating example. Both methods are applied to real study data to investigate the symptoms proceeding the first recognized sign of multiple sclerosis.

Special Seminar: Scalable Bayesian Inference with Hamiltonian Monte Carlo

Abstract: Despite the promise of big data, inferences are often limited
not by sample size but rather by systematic effects. Only by carefully
modeling these effects can we take full advantage of the data -- big
data must be complemented with big models and the algorithms that can
fit them.  One such algorithm is Hamiltonian Monte Carlo, which exploits
the inherent geometry of the posterior distribution to admit full
Bayesian inference that scales to the complex models of practical
interest. In this talk Michael Betancourt will present a conceptual
discussion of the challenges inherent to Bayesian computation and the
foundations of why Hamiltonian Monte Carlo in uniquely suited to
surmount them.

Biography: Michael Betancourt is the principle research scientist with
Symplectomorphic, LLC where he develops theoretical and methodological
tools to support practical Bayesian inference. He is also a core
developer of Stan, where he implements and tests these tools. In
addition to hosting tutorials and workshops on Bayesian inference with
Stan he also collaborates on analyses in epidemiology, pharmacology, and
physics, amongst others. Before moving into statistics, Michael earned a
B.S. from the California Institute of Technology and a Ph.D. from the
Massachusetts Institute of Technology, both in physics.

Website: https://betanalpha.github.io
Twitter: @betanalpha

Bayesian analysis of continuous time Markov chains with applications to phylogenetics

Bayesian analysis of continuous time, discrete state space time series is an important and challenging problem, where incomplete observation and large parameter sets call for user-defined priors based on known properties of the process. Generalized Linear Models (GLM) have a largely unexplored potential to construct such prior distributions. In this talk, we show that an important challenge with Bayesian generalized linear modelling of Continuous Time Markov Chains (CTMCs) is that classical Markov Chain Monte Carlo (MCMC) techniques are too ineffective to be practical in that setup. We propose two computational methods to address this issue. The first algorithm uses an auxiliary variable construction combined with an Adaptive Hamiltonian Monte Carlo (AHMC) algorithm. To further exploit the sparsity structure often found in the parameterization of high-dimensional rate matrices, we make use of recently developed Monte Carlo schemes based on non-reversible Piecewise-deterministic Markov Processes (PDMPs). Our second algorithm combines Hamiltonian Monte Carlo (HMC) and Local Bouncy Particle Sampler (LBPS) to take advantage of the sparsity in certain high-dimensional rate matrices. We propose a characterization for a class of sparse factor graphs, where LBPS can be efficient. We also provide a framework to assess the computational complexity for the algorithm. An important aspect of practical implementation of Bouncy Particle Sampler (BPS) is the simulation of event times. Default implementations use conservative thinning bounds. Such bounds can slow down the algorithm and limit the computational performance. In the second algorithm, we develop exact analytical solutions to the random event times in the context of CTMCs. Both sampling algorithms and our model make it efficient both in terms of computation and analyst's time to construct stochastic processes informed by prior knowledge, such as known properties of the states of the process. We demonstrate the flexibility and scalability of our framework using both synthetic and real phylogenetic protein data.

SAEM methods and PDE parameters statistical estimation - illustrations in biology

Parameter estimation in non linear mixed effects models requires a large number of evaluations of the model to study. For ordinary differential equations, the overall computation time is reasonable. However when the model itself is more complex (for instance when it is a set of partial differential equations (PDE)) it may be infeasible within a reasonable time.
In this talk, we present two variations on the stochastic approximation expectation maximisation (SAEM) method in conjunction with PDEs to estimate the parameters of a population of individuals. One is based on the building of an off line grid used to approximate the model (so called metamodel) and the other involves a dynamic refinement (using a kriging approach) of the metamodel along the iterations of the SAEM. These methods are illustrated on the classical Fisher-KPP equation and on a renewal (age-structured) equation.

PIMS (CNRS UMI 3069), University of British Columbia, Vancouver, Canada &
UMPA (CNRS UMR 5669), ENS de Lyon, France
http://perso.ens-lyon.fr/paul.vigneaux/