Seminar

Smog, smoke and standards: modelling the effects of air pollution in space and time

Dr Shaddick is a Reader in Statistics in the Department of Mathematical Sciences at the University of Bath. He has a PhD from Imperial College London in statistics and epidemiology and a Masters from University College London in applied stochastic systems. His research interests include the theory and application of Bayesian statistics to the areas of spatial epidemiology, environmental health risk and the modelling of spatio-temporal fields of environmental hazards. He was a co-author of the Oxford Handbook of Epidemiology for Clinicians which was Highly Commended in the Basis of Medicine Category, BMA Book Awards 2013. Together with Jim Zidek, he has a recently written a book `Spatio-temporal methods in environmental epidemiology’ which will be published in July 2015.

Smog, smoke and standards: modelling the effects of air pollution in space and time

Show Abstract

From the famous London smogs in the 1950s to air quality in megacities today, the potential effects of air pollution are a major concern both in terms of the environment and in relation to human health. In order to support environmental policy and to estimaterisks of environmental hazards on human health study there is a requirement for accurate estimates of exposures that might be experienced by the populations at risk. For epidemiological studies these exposures must be linked to health outcomes but health and exposure data may not match at all locations in space and time. In such cases a direct comparison of exposures and health outcomes is often not possible without an underlying model to align the two in the spatial and temporal domains. In addition, there may be periods of missing data and preferential sampling, where monitoring locations in environmental networks may be located in areas where levels are expected to be high. Biased estimates of exposures may lead to biased estimates of risk. The Bayesian approach provides a natural framework for modelling complexities within exposure data and provides a means of incorporating the results within health models. However, the large amounts of data that can arise from environmental networks mean that inference using MCMC may not be computational feasible. Here we use Integrated Nested Laplace Approximations (INLA) to implement spatio-temporal exposure models and show how the results can be used to reduce the potential biases that may occur when estimating levels of air pollution and the associated risks to human health.

Evolution of Big Data Analytics

Kwok L Tsui is head and chair professor in the Department of Systems Engineering and Engineering Management at City University of Hong Kong. Prior to the current position, Dr. Tsui has been professor/associate professor in the School of Industrial and Systems Engineering at Georgia Institute of Technology in 1990-2011; and member of technical staff in the Quality Assurance Center at AT&T Bell Labs in 1986-1990. He received his Ph.D. in Statistics from the University of Wisconsin at Madison. Professor Tsui was a recipient of the National Science Foundation Young Investigator Award. He is Fellow of the American Statistical Association, American Society for Quality, International Society of Engineering Asset Management, and Hong Kong Institution of Engineers; and U.S. representative to the ISO Technical Committee on Statistical Methods. Professor Tsui was Chair of the INFORMS Section on Quality, Statistics, and Reliability and the Founding Chair of the INFORMS Section on Data Mining. Professor Tsui’s current research interests include data mining, surveillance in healthcare and public health, prognostics and systems health management, calibration and validation of computer models, process control and monitoring, and robust design and Taguchi methods.

Automatic Modulation Recognition in the 868 MHz Wireless Network using Tree-Based Methods

Wireless communication systems enable the transfer of data between devices through the transmission of radio signals. These wireless devices often transmit a variety of waveform types called modulation formats. Misidentification of modulation formats between two communication devices could lead to undesirable delays in data transfer and inefficient energy consumption. The role of automatic modulation recognition (AMR) is the identification of the different modulations of transmitted signals. In this project, we explore tree-based AMR methods in the 868 MHz frequency band. An existing work implements a tree-based classifier that is constructed manually by feature inspection. We extend the recent approach by implementing classification tree (CT) and random forest (RF) classifiers, as well as introducing an expanded list of features. Performance is verified via a simulation at different signal to noise ratios (SNR). Signal data is initially preprocessed, and features are extracted and used to train each classifier. Improvements of 14% and 3% success rates are found for RF and CT, respectively using all features and at a SNR of 1 dB.

Stochastic gradient methods for estimating expectations

Stochastic gradient methods have had great impact on tuning large scale models such as deep learning. This talk describes recent results of the use stochastic gradient methods for approximating expectation with respect to probability distributions. Applying standard Markov chain Monte Carlo (MCMC) algorithms to large data sets is computationally expensive. Both the calculation of the acceptance probability and the creation of informed proposals usually require an iteration through the whole data set. The recently proposed stochastic gradient Langevin dynamics (SGLD) method circumvents this problem by generating proposals which are only based on a subset of the data, by skipping the accept-reject step. The talk surveys two recent preprints providing rigorous foundation on decreasing and non decreasing step size SGLD,http://vollmer.ms/sebastian/sgld.pdf and http://vollmer.ms/sebastian/sgld2.pdf, respectively.

Conditional extremes in asymmetric financial markets

The report focuses on the estimation of the probability distribution of a bivariate random vector given that one of the components takes on a large value. These conditional probabilities can be used to quantify the effect of financial contagion when the random vector represents losses on financial assets and as a stress-testing tool in financial risk management. In the context of risk management, the main interest lies in the tails of the underlying distribution. In such cases, empirical probabilities fail to provide adequate estimates while fully parametric methods are subject to large model uncertainty as there is too little data to assess the model fit in the tails. We propose a semi-parametric framework using asymptotic results in the spirit of extreme values theory. The main contributions include an extension of the limit theorem in {Abdous2005a} [Canad. J. Statist. 33 (2005)] to allow for asymmetry, frequently encountered in financial and insurance applications, and a new approach for inference.

Protein loop prediction via sequential Monte Carlo and simultaneous side-chain filtering

Using computation to predict a protein's 3D structure from its amino acid sequence remains a highly challenging problem in biology. This talk focuses on one particular subproblem: the structure prediction of the variable loop regions in proteins that connect the more ordered helices and sheets. The primary difficulty of this task is the lack of efficient sampling algorithms to explore the low-energy conformational space of the loop region. Our new method, inspired by sequential Monte Carlo techniques, simultaneously explores both the backbone and side-chain space of the loop region as it places one amino acid at a time. Amino acid placements are filtered at each step to retain the most promising candidates, while maintaining a diverse set of samples. This produces an ensemble of low-energy conformation proposals for the loop region of interest, from which a final prediction can be chosen according to selection criteria. We show that the method improves sampling and prediction accuracy on benchmark datasets compared to previous approaches.

Title: DrSMC: A Sequential Monte Carlo Sampler for Deterministically Related Variables; Phylogenetic Inference with Divide and Conquer Sequential Monte Carlo (D&C SMC)

Talk by Neil Spencer (11am - 11:30am)

Title:  DrSMC: A Sequential Monte Carlo Sampler for Deterministically Related Variables

Abstract:  Computing posterior distributions over variables linked by deterministic constraints is a recurrent problem in Bayesian analysis. Such problems can arise due to censoring, identifiability issues, or other considerations. It is well-known that standard implementations of Monte Carlo methods break down in the presence of these deterministic relationships. Although several alternative Monte Carlo methods have been recently developed, few are applicable to deterministic relationships on continuous random variables. My Masters thesis work involved the development of a new Sequential Monte Carlo Sampler, called DrSMC, for such problems. In this talk, I discuss applying DrSMC to compute the posterior distribution of a continuous random vector given its sum.

Talk by Sean Jewell (11:30am - 12pm)

Title:  Phylogenetic Inference with Divide and Conquer Sequential Monte Carlo (D&C SMC)

Abstract:  Recently reconstructing evolutionary histories has become a computational issue due to the increased availability of genetic sequencing data and relaxations of classical modelling assumptions. My master's thesis focused on specializing a D&C SMC inference algorithm to phylogenetics to address these challenges. In phylogenetics, the tree structure used to represent evolutionary histories provides the model decomposition for D&C SMC. In particular, speciation events are used to recursively decompose the model into subproblems. Each subproblem is approximated by an independent population of weighted particles, which are merged and propagated to create an ancestral population. This approach provides the flexibility to relax classical assumptions on large trees by parallelizing these recursions.

A theoretical framework for calibration in computer models: parametrization, estimation and convergence properties

Calibration parameters in deterministic computer experiments are those attributes that cannot be measured or available in physical experiments or observations. Kennedy-O’Hagan (2001) suggested an approach to estimate them by using data from physical experiments and computer simulations. A new theoretical framework is given which allows us to study the issues of parameter identifiability and estimation. It is shown that a simplified version of the original KO method leads to asymptotically inconsistent calibration. A novel calibration method, called the L2 calibration, is proposed and proven to be consistent and enjoys optimal convergence rate. The asymptotic results of L2 calibration for stochastic physical systems are also studied. It is proved that the L2 calibration estimator is asymptotically normal and semi-parametric efficient. This work also investigates the asymptotic properties of the ordinary least squares method.
(joint work with C. F. Jeff Wu, Georgia Institute of Technology)

Destructively Assessed Property Relationships in Reliability Theory with an Application in Relating Structural Properties of Dimension Lumber

We present a novel approach for predicting one lumber strength property from another, each being measured destructively. The objective is to reduce the cost of lumber monitoring programs, by measuring one of the properties and predicting the other. To reach the objective, we review single proof load design (SPLD) proposed to assess dependence between two jointly normally distributed random variables X and Y that cannot be observed simultaneously. The SPLD tests specimens in one mode X up to a determined load (proof loading), and the survivors are tested to failure in a second mode Y. To resolve the near non-identifiability of parameters in the SPLD approach, we redesign the SPLD for improving the estimation performance. The new design assigns specimens to one of the two groups: SPLD and shoulder. The shoulder group tests specimens to failure in the Y mode. The SPLD with a shoulder approach results in a more accurate and precise estimate. To quantify damage caused by proof loading lumber specimens, we use the maximum likelihood method to estimate theoretical quantiles of the strength distribution using a sample from a SPLD with a shoulder experiment with a low proof load level. The comparison between those estimated theoretical quantiles and empirical quantiles of the proof load survivors reveals the damage at higher proof load levels. Using our experimental data on manufactured lumber, we find low and high proof load levels leave survivors undamaged, but intermediate load levels do damage weaker survivors. We generalize the SPLD with a shoulder approach to incorporate proof load damage, and thus finally provide a method for estimating the X-Y dependence when damage occurs. This generalized approach is applied on our experimental data to estimate the relationship between bending and log--transformed tension. The high correlation found in our application suggests that if there is a need to verify the two strength properties, only one of them needs to be measured and the other is predicted from the measured strength.

A Markov Random Field Approach to Modelling Animal Habitat


Habitat modelling presents a challenge due to the variety and the accuracy of the data available. My Master's thesis focused on applying Markov random field theory to incorporate distinct types of ecological data into a habitat model. In this talk, I provide a brief overview of the data available and both the intuition and theory for using this data for habitat modelling. Model building, parameter estimation, and results with synthetic data are presented.