Preferential attachment and neutral random graphs: statistically useful generative models of network data
This talk presents a new method for reducing big and high-dimensional data into a smaller dataset, called support points (SPs). In an era where data is plentiful but downstream analysis is oftentimes expensive, SPs can be used to tackle many big data challenges in statistics, engineering and machine learning. SPs have two key advantages over existing methods. First, SPs provide optimal and model-free reduction of big data for a broad range of downstream analyses. Second, SPs can be efficiently computed via parallelized difference-of-convex optimization; this allows us to reduce millions of data points to a representative dataset in mere seconds. SPs also enjoy appealing theoretical guarantees, including distributional convergence and improved reduction over random sampling and clustering-based methods. The effectiveness of SPs is then demonstrated in two real-world applications, the first for reducing long Markov Chain Monte Carlo (MCMC) chains for rocket engine design, and the second for data reduction in computationally intensive predictive modeling.
With the advancement in disease registries and surveillance data, population-based
information on disease incidence, survival probability or other important biological
characteristics become increasingly available. Such information can be leveraged in
studies that collect detailed measurements but with smaller sample sizes. In contrast
to recent proposals that formulate the additional information as constraints in opti-
mization problems, we develop a general framework to construct simple estimators that
update usual regression estimators with some functionals of data and the initial esti-
mator based on the additional information. We consider general settings which include
nuisance parameters in the auxiliary information, non-i.i.d. data such as case-control
sampling, and semiparametric models with infinite dimensional parameters. Detailed
examples of several important data and sampling settings are provided.
Sparse generalized eigenvalue problem (GEP) plays a pivotal role in a large family of high-dimensional learning tasks, including sparse Fisher’s discriminant analysis, canonical correlation analysis, and sufficient dimension reduction. Most of the existing methods and theory in the context of specific statistical models that are special cases of sparse GEP require restrictive structural assumptions on the input matrices. This talk will focus on a two-stage computational framework for solving the non-convex optimization problem resulting from the sparse GEP. At the first stage, we solve a convex relaxation of the sparse GEP. Taking the solution as an initial value, we then exploit a non-convex optimization perspective and propose the truncated Rayleigh flow method (Rifle) to estimate the leading generalized eigenvector, and show that it converges to a solution with the optimal statistical rate of convergence. Theoretically, our method significantly improves upon the existing literature by eliminating the structural assumptions on the input matrices. Numerical studies in the context of several statistical models are provided to validate the theoretical results. We then apply the proposed method to an electrocorticography data to understand how human brains recall and mentally rehearse word sequences.
Canadian health administrative databases are a popular and rich data source for epidemiological research at the population level, but error-prone diagnostic information typically provide investigators with a less-than-perfect proxy for the disease under study. Motivated by an observational study into the prodromal phase of multiple sclerosis, this talk considers the setting of a matched exposure-disease association study where the disease variable is measured with error, and participant selection is based upon the error-prone disease label rather than the true disease status. We initially focus on the special case of a pair-matched case-control study. Assuming non-differential misclassification of study participants, we give a closed-form expression for asymptotic biases in odds ratios arising under naive analyses of misclassified data, and propose a Bayesian model to correct association estimates for misclassification bias. For identifiability, the model relies on information from a validation cohort of correctly classified case-control pairs, and also requires prior knowledge about the predictive values of the classifier. In a simulation study, the model shows improved point and interval estimates relative to the naive analysis, but is also found to be overly restrictive for our motivating dataset. In light of these concerns, we further propose a generalized model for misclassified data that extends to the case of differential misclassification and allows for a variable number of controls per matching stratum. Instead of prior information about the classification process, the model relies on individual-level estimates of each participant’s true disease status, which were obtained from a counting process mixture model of MS-specific healthcare utilization in our motivating example. Both methods are applied to real study data to investigate the symptoms proceeding the first recognized sign of multiple sclerosis.
Abstract: Despite the promise of big data, inferences are often limited
not by sample size but rather by systematic effects. Only by carefully
modeling these effects can we take full advantage of the data -- big
data must be complemented with big models and the algorithms that can
fit them. One such algorithm is Hamiltonian Monte Carlo, which exploits
the inherent geometry of the posterior distribution to admit full
Bayesian inference that scales to the complex models of practical
interest. In this talk Michael Betancourt will present a conceptual
discussion of the challenges inherent to Bayesian computation and the
foundations of why Hamiltonian Monte Carlo in uniquely suited to
surmount them.
Biography: Michael Betancourt is the principle research scientist with
Symplectomorphic, LLC where he develops theoretical and methodological
tools to support practical Bayesian inference. He is also a core
developer of Stan, where he implements and tests these tools. In
addition to hosting tutorials and workshops on Bayesian inference with
Stan he also collaborates on analyses in epidemiology, pharmacology, and
physics, amongst others. Before moving into statistics, Michael earned a
B.S. from the California Institute of Technology and a Ph.D. from the
Massachusetts Institute of Technology, both in physics.
Website: https://betanalpha.github.io
Twitter: @betanalpha
Bayesian analysis of continuous time, discrete state space time series is an important and challenging problem, where incomplete observation and large parameter sets call for user-defined priors based on known properties of the process. Generalized Linear Models (GLM) have a largely unexplored potential to construct such prior distributions. In this talk, we show that an important challenge with Bayesian generalized linear modelling of Continuous Time Markov Chains (CTMCs) is that classical Markov Chain Monte Carlo (MCMC) techniques are too ineffective to be practical in that setup. We propose two computational methods to address this issue. The first algorithm uses an auxiliary variable construction combined with an Adaptive Hamiltonian Monte Carlo (AHMC) algorithm. To further exploit the sparsity structure often found in the parameterization of high-dimensional rate matrices, we make use of recently developed Monte Carlo schemes based on non-reversible Piecewise-deterministic Markov Processes (PDMPs). Our second algorithm combines Hamiltonian Monte Carlo (HMC) and Local Bouncy Particle Sampler (LBPS) to take advantage of the sparsity in certain high-dimensional rate matrices. We propose a characterization for a class of sparse factor graphs, where LBPS can be efficient. We also provide a framework to assess the computational complexity for the algorithm. An important aspect of practical implementation of Bouncy Particle Sampler (BPS) is the simulation of event times. Default implementations use conservative thinning bounds. Such bounds can slow down the algorithm and limit the computational performance. In the second algorithm, we develop exact analytical solutions to the random event times in the context of CTMCs. Both sampling algorithms and our model make it efficient both in terms of computation and analyst's time to construct stochastic processes informed by prior knowledge, such as known properties of the states of the process. We demonstrate the flexibility and scalability of our framework using both synthetic and real phylogenetic protein data.
Parameter estimation in non linear mixed effects models requires a large number of evaluations of the model to study. For ordinary differential equations, the overall computation time is reasonable. However when the model itself is more complex (for instance when it is a set of partial differential equations (PDE)) it may be infeasible within a reasonable time.
In this talk, we present two variations on the stochastic approximation expectation maximisation (SAEM) method in conjunction with PDEs to estimate the parameters of a population of individuals. One is based on the building of an off line grid used to approximate the model (so called metamodel) and the other involves a dynamic refinement (using a kriging approach) of the metamodel along the iterations of the SAEM. These methods are illustrated on the classical Fisher-KPP equation and on a renewal (age-structured) equation.
PIMS (CNRS UMI 3069), University of British Columbia, Vancouver, Canada &
UMPA (CNRS UMR 5669), ENS de Lyon, France
http://perso.ens-lyon.fr/paul.vigneaux/