Seminar

Pumps, Maps and Pea Soup: Spatio-temporal methods in environmental epidemiology

This talk provides an introduction to epidemiological analysis where the distribution of health outcomes and related exposures are measured over both space and time. Developments in this field have been driven by public interest in the effects of environmental pollution, increased availability of data and increases in computing power. These factors, together with recent advances in the field of spatio-temporal statistics, have led to the development of models which can consider relationships between adverse health outcomes and environmental exposures over both time and space simultaneously.
 
Using illustrative examples, from outbreaks of cholera in London in the 1850s, episodes of smog in the 1950s to present day epidemiological studies, we discuss a variety of issues commonly associated with analyses of this type including modelling auto-correlation, preferential sampling of exposures and ecological bias. The precise choice of statistical model may be based on whether we are explicitly interested in the spatio-temporal pattern of disease incidence, e.g. disease mapping and cluster detection, or whether clustering is a nuisance quantity that we need to acknowledge, e.g. spatio-temporal regression. Throughout we consider the practical implementation of models with specific focus on inference within a Bayesian framework using computational methods such as Markov Chain Monte Carlo and Integrated Nested Laplace Approximations. 
 
The talk also serves as a precursor to a graduate level course on spatio-temporal methods in epidemiology. This course will cover the basic concepts of epidemiology, methods for temporal and spatial analysis and the practical application of such methods using commonly available computer packages. It will have an applied focus with both lectures and practical computer sessions in which participants will be guided through analyses of epidemiological data.     
 
BACKGROUND INFORMATION:  The Statistics Department, with the support of the Constance van Eeden Fund,  is honoured to host Dr Gavin Shaddick during term 2 2012-13.   Dr Shaddick, a Reader in Statistics in the Department of Mathematical Sciences at the University of Bath, has achieved international prominence for his contributions to the theory and application of  Bayesian statistics to the areas of spatial epidemiology, environmental health risk and the modelling of spatio-temporal fields of environmental hazards.

Dr Shaddick will begin his visit to the Department, by giving the 2012-13 van Eeden lecture.   That lecture will inaugurate a one term special topics graduate course in statistics, which the Department of Statistics is offering next term. It will be given by Dr Shaddick and Dr James Zidek  (Statistics, UBC)  on the subject of spatial epidemiology. This course, which is aimed primarily at a statistical audience, will provide an introduction to environmental epidemiology and spatio-temporal process modeling, as it applies to the assessment of risk to human health and welfare due to random fields of hazards such as air pollution.  Please see the course outline for more information.

Spatio-temporal statistical analysis of oral cancer growth

Like most human epithelial cancers, oral cancer originates form an accumulation of molecular, genetic and biological events through a series of precancerous lesions. The earlier we detect the lesions during its natural history, the better the outcome.

Thanks to rapid development and explosion in informatics, electronics, optics, and microscopy, new technologies can now be used to detect these early lesions. Unfortunately, even with the exponential quantity and resolution of data generated by the "omics" sciences (genomics, proteomics, digital pathology, etc..), it is still difficult to identify lesions that are likely to become cancer , i.e. high-risk precancerous lesions, in order to treat them before they become fully malignant.  To transform this data to knowledge, other approaches are needed.

We propose the development of a modeling platform for the spatio-temporal analysis of oral cancer, that will combine 3D in silico modeling and statistical modeling approaches.

As part of a pan-canadian Terry Fox grant project and the British Columbia Oral Cancer Prevention Program (BC OCPP), directed by Dr Miriam Rosin (BCCRC), we have been collecting for the last 15 years a unique collection of oral specimens from patients involved in different clinical trials. In the present COOL trial, we have been collecting cytological and histological specimens from high-risk patients at different time for a 2 years period with the following characteristics:


A. Clinical information and risk factors are available for each patient (age, sex, smoking, viral infection, alcohol, etc.. ).

B. Sampling and analyses are repeated every 6 months for 2 years (temporal level)

C. For each patient, different histological specimens are collected from the oral cavity: from the worst abnormal area as well as from areas sampled at different distances from the former, and/or from randomly selected areas (spatial level 1)

D. For each specimen, pathologists identify different biologically or pathologically distinct unit areas(spatial level 2)

E. Genetic markers and several phenotypic biomarkers are measured on each unit area by  Quantitative Tissue Phenotypic Analysis (using high-resolution image algorithms) that extracts quantitative information at three levels:
            o Cell level : the basic unit of analysis: about 200 features measured on each cell  (nuclear morphology, DNA content , DNA chromatin texture, etc..), with few hundred to thousand nuclei per unit (spatial level 3);
            o Tissue level : the overall organization of the unit. Using graph-theory based tools (Voronoi diagram, MST, Gabriel Graph, RNG,etc.. ) (spatial level 4);
            o Neighbourhood level : spatial arrangement of different groups of cells sharing molecular or phenotypic characteristics (mutated cells, proliferating cells, etc.) (using local graph, Ulam tree, k Nearest neighbours, etc..) (spatial level 5);

We will first present the correlation between phenotypic features, genetic markers and pathological diagnosis. We will show the potential of these features to identify high-risk lesions more likely to progress to cancer.  We will then present some preliminary simulations of our in silico static and dynamic 3D model (Idefics). Finally we will introduce the concepts of a new spatio-temporal statistical model of oral cancer progression.

Spatio-temporal statistical analysis of oral cancer growth

The traditional training of a biostatistician ill prepares one to encounter the enormous datasets and science-fiction technologies that arise during the search for biomarkers. In my talk I will describe a few of the interesting projects I have had the good luck to work on over the last few years and how biomarkers have become recurring theme in my work. Along the way, I will discuss some of the ways that Bioinformatics and Biostatistics may cross-pollinate with one another in the search for and validation of new biomarkers.

A Biostatistician’s Take on Biomarkers

The traditional training of a biostatistician ill prepares one to encounter the enormous datasets and science-fiction technologies that arise during the search for biomarkers. In my talk I will describe a few of the interesting projects I have had the good luck to work on over the last few years and how biomarkers have become recurring theme in my work. Along the way, I will discuss some of the ways that Bioinformatics and Biostatistics may cross-pollinate with one another in the search for and validation of new biomarkers.

Second-order Least Squares Estimation in Linear Dynamic Panel Data Models

We propose the second-order least squares estimator for the autoregressive panel data models.  This method requires only the specification of the first two conditional moments of the unobserved effects given the process initial observation, and does not require any other distributional assumptions. The data generating process can be either stationary or nonstationary. The proposed estimator is consistent and asymptotically normal for large N and finite T under fairly general regularity conditions. Moreover, we show that our estimator reaches an optimal semiparametric efficiency bound.  Monte Carlo simulation studies show that the proposed estimator performs satisfactorily in finite sample situations compared to the usual first-differenced generalized method of moment (GMM) and the random effects pseudo maximum likelihood (PML) estimators.

This is joint work with Mustafa Salamh.

A Co-op Experience at BC Centre for Disease Control: Study of Acute HIV Infection in Gay Men & A Non-linear Mixed-effect Logistic Modeling of Seroconversion Panel Data

Here, I intend to share the co-op experience I gained at the BC Centre for Disease Control, and I will be emphasizing on two projects; the Study of Acute HIV infection in Gay Men, which was initiated to explore the gay men's understanding of HIV testing, their motivations and challenges in taking an HIV test, and the impacts of new testing technologies that are able to identify persons with acute infection, on their testing practices. For this study I will be mainly discussing the data structure and data manipulation issues I faced. The other project was to estimate day of seroconversion and the day of HIV-infection by using a non-linear mixed-effect logistic model on HIV-seroconversion panel data. 

Co-op Experience: COFAS Ankle Arthritis Surgical Outcome Study

This talk will discuss the work experience gained during my two co-op placements. The focus of the presentation will be the COFAS ankle arthritis study conducted by the Orthopaedic Research Centre at St. Paul’s Hospital. The study examined outcomes for patients with end-stage ankle arthritis based on questionnaires collected over a 10-year period, a longitudinal analysis of the data will be presented. The main objective was to compare the effectiveness of two surgical options, ankle fusions and total ankle replacements based on standard pain and disability scores.

Epidemiologic methods are useless. They can only give you answers.

Abstract:  The first duty of any epidemiologist is to ask a relevant question. Learning and applying sophisticated epidemiologic methods is of little help if the methods are used to answer irrelevant questions. This talk will discuss the formulation of research questions in the presence of time-varying treatments and treatments with multiple versions, including pharmacological treatments and lifestyle exposures. Several examples will show that discrepancies between observational studies and randomized trials are often not due to confounding, but to the different questions asked.


Brief Biography:  Miguel Hernán is Professor of Department of Epidemiology and Department of Biostatistics at the Harvard School of Public Health (HSPH). His research is focused on the development and application of causal inference methods to guide policy and clinical interventions. He and his collaborators apply statistical methods to observational studies under suitable conditions to emulate hypothetical randomized experiments so that well-formulated causal questions can be investigated properly. His research applied to many areas, including investigation of the optimal use of antiretroviral therapy in patients infected with HIV, assessment of various interventions of kidney disease, cardiovascular disease, cancer and central nervous system diseases. He is Associate Director of HSPH Program on Causal Inference in Epidemiology and Allied Sciences, member of the Affiliated Faculty of the Harvard-MIT Division of Health Sciences and Technology, and an Editor of the journal EPIDEMIOLOGY. He is the author of upcoming highly anticipated textbook "Causal Inference" (Chapman & Hall/CRC, 2013), drafts of selected chapters are available on his website.

Viral Marketing over Social Networks: A Journey

Speaker's Page

Abstract:  There is tremendous interest in understanding and harnessing the spread of information and influence over social networks, fueled by applications such as viral marketing, epidemiology, and spreading of innovation. Ever since the seminal publications by Domingos and Richardson (2001) and Kempe, Kleinberg, and Tardos (2003), this field has seen an explosive growth. Using viral marketing as a driving application, in this talk, I will present some key problems in this area that touch on inherent assumptions, modeling, scalability, and functionality. I will also discuss the solutions developed in our group. I will conclude the talk with a list of problems that need to be solved in order to take viral marketing out of the lab.

This is joint work with several colleagues, grad students, and postdocs, who will be acknowledged in the talk.

Simulation Model Calibration and Prediction Using Outputs From Multi-Fidelity Simulators

Computer simulators are used widely to describe physical processes in lieu of physical observations. In some cases, more than one computer code can be used to explore the same physical system - each with different degrees of fidelity. In this work, we combine field observations and model runs from deterministic multi-fidelity computer simulators to build a predictive model for the real process. The resulting model can be used to perform sensitivity analysis for the system and make predictions with associated measures of uncertainty. Our approach is Bayesian and will be illustrated through a simple example, as well as a real application in predictive science at the Center for Radiative Shock Hydrodynamics at the University of Michigan.