Seminar

Model and Inference Issues Related to Exposure-Disease Relationships

We work on two issues related to exposure-disease relationships. Firstly we build a Bayesian hierarchical model for relating disease to a potentially harmful exposure, using data from studies in occupational epidemiology, and compare our method with the traditional group-based exposure assessment method through simulation studies, a real data application, and theoretical calculation. We focus on cohort studies where a logistic disease model is appropriate and where group means can be treated as fixed effects. The results show a variety of advantages of the fully Bayesian approach, and provide recommendations on situations where the traditional group-based exposure assessment method may not be suitable to use.


Secondly, the shape of the relationship between a continuous exposure variable and a binary disease variable is often central to epidemiologic investigations. We investigates a number of issues surrounding inference and the shape of the relationship. Presuming that the relationship can be expressed in terms of regression coefficients and a shape parameter, we investigate how well the shape can be inferred in settings which might typify epidemiologic investigations and risk assessment. We also consider a suitable definition of the average effect of exposure, and investigate how precisely this can be inferred. This is done both in the case of using a model acknowledging uncertainty about the shape parameter and in the case of using a simple model ignoring this uncertainty. We also examine the extent to which exposure measurement error distorts inference about the shape of the exposure-disease relationship. All these investigations require a family of exposure-disease relationships indexed by a shape parameter. For this purpose, we employ a family based on the Box-Cox transformation.

Nonparametric Estimation of the Drift Coefficient of a Stochastic Diffusion Process in the Presence of Measurement Error

We propose a Nadaraya-Watson type kernel estimator of the drift coefficient of a diffusion process, where the process is observed discretely in time and with independent additive measurement errors. Our estimation procedure first averages the data neighboring in time, therefore reducing the noise caused by the measurement errors and revealing the latent diffusion process. We show that our estimator is consistent and asymptotically normal when the diffusion process is positive recurrent and strictly stationary and the independent additive measurement errors have zero mean and bounded variance. We study the properties of our estimator via simulation.

The Estimation of Spatially Varying Coefficients via a Spatial GLMM with Application to Multiple Sclerosis MRI Data

Multiple Sclerosis (MS) is an autoimmune disease that acts the central nervous system (CNS) by disrupting nerve transmission. This disruption is caused by damage to the myelin sheath surrounding nerves that acts as an insulator. Patients with MS have a multitude of symptoms that depend on where lesions occur in the brain and/or spinal cord. MS has no cure and in order to help manage the disease, physicians subtype MS patients into 4 categories that depend on the pattern of MS episodes.  Patient symptoms are rated by the Kurtzke Functional Systems (FS) scores and the paced auditory serial addition test (PASAT) score. The
eight functional systems (CNS areas or circuits that regulate body functions) are: 1) pyramidal; 2) cerebellar; 3) brainstem; 4) sensory; 5) bowel and bladder; 6) visual; 7) cerebral; and 8) other.  Of interest to Neurologists is whether lesion locations can be predicted using these FS and PASAT scores and whether the data can help predict MS subtype in a newly diagnosed patient.

 
To help answer these questions, we propose an autoprobit regression model with spatially varying random coecients. The data of interest are digitized binary images of the brain derived from high-resolution T2-weight MRI images encoded such that, at each voxel, 1 indicates the presence of a lesion and 0 denotes absence of a lesion.  These binary lesion maps are taken as the dependent variables. In contrast to most spatial applications, in which only one realization of a process is observed, we have multiple, independent realizations, one from each patient. This allows the modeling and estimation of spatially varying parameters of patient level covariates such as age, gender, disease duration, FS and PASAT scores. Maps of these spatially varying parameters over the brain allow us to spatially predict lesion probabilities over the brain given covariates. Furthermore, we show that, via Bayes Theorem, the model can be used to accurately predict the MS subtype of a new subject.

Empirical Likelihood Methods for Pretest-Posttest Studies

Pretest-posttest studies are an important and popular method for assessing treatment effects or the effectiveness of an intervention in many areas of scientific research. There are two distinct features for this type of study: availability of baseline information for all subjects in the study and missingness by design of measures of the responses. In this talk, we first provide a brief overview of existing research on the topic. We then present alternative empirical likelihood approaches to inferences on the treatment effects with efficient use of all available information. Theoretical results are developed, and finite sample performances of the proposed methods with comparison to existing ones are investigated through simulation studies.
 
This is joint work with Min Chen and Mary Thompson.

Inferring population size for epidemiological purposes with the multiplier method

The multiplier method is a method to estimate population size (e.g. the number of sex-workers in a city), popular among epidemiologists. However, it relies on the identity that p_t = N_t/N, where p_t is the proportion of target population with a certain trait, N_t is the number of people with this trait, and N is the size of the target population. This causes problems in today's world where information on many traits exist in record such that each trait can be used to generate an estimate. I will introduce a recent extension of the multiplier method with which information on multiple traits are used simultaneously to produce a single estimate of population size. I will include examples and some theoretical findings on the effect of study design on inference precision.
 

Combination based permutation tests: theory and applications

In recent years permutation testing methods have increased both in number of applications and in solving complex multivariate problems. Permutation tests are essentially of an exact nonparametric nature in a conditional context where conditioning is on the pooled observed data set which is generally a set of sufficient statistics in the null hypothesis.

There are many complex multivariate problems which are difficult to solve outside the conditional framework and in particular outside the nonparametric combination (NPC) of dependent permutation tests method. We discuss this method along with some application in different experimental or observational situations (e.g. in biostatistics).Some recent properties and features of combination-based tests are also presented.

Understanding Food Webs: Twitter, Critters, Bugs, Python, and More

Any network or web-like structure consists of "actors" that are linked by some type of activity.  Thus, inference for features in a social network or a predator-prey food web can be made using a common statistical framework as foundation. Our Bayesian methodology extends upon latent space network modelling to address the notion of "trophic levels" from three perspectives of feeding behaviour: (1) activity level as predator and prey, (2) feeding preference as predator, and (3) feeding preference as prey.  We present worked examples, and discuss the implementation of MCMC using brute force, (Open/Win)BUGS, and (Py)Stan.

2 UBC Statistics MSc Student Talks

Determination of Sample Size for Phase II Clinical Trials in Multiple Sclerosis using Lesional Recovery as an Outcome Measure

Speaker:  Md Mahsin

Multiple sclerosis (MS) is an inflammatory demyelinating disease of the central nervous system. The hallmark feature of the disease is the formation of focal demyelinating lesions accompanied by myelin destruction in the white matter (WM). Magnetic resonance imaging (MRI) is used identify and visualize these lesions. Repeated MRI scanning of patients (most often monthly) over period of months has become a standard protocol for Phase II trials of experimental treatment in MS. The formation of WM lesions in MS is characterized by inflammatory demyelination and then remyelination usually occurs over several months after lesion formation. Hence, a measure reflecting lesional recovery is a promising outcome for phase II clinical trials that assess the eff ect of therapies intended to induce remyelination. Our objective is to provide sample sizes required to detect such an experimental treatment e ffect with certain statistical power. We consider a parallel group design with two arms of equal number of subjects. The study design is considered as a three level hierarchical data structure where lesions are nested within subjects and are assessed repeatedly over the study period. Variable numbers of new enhancing lesions per subject and variable numbers of measurements at before and after enhancement (depends on the time of the lesion's appearance) are also considered. The numbers of subjects in each treatment arm necessary to obtain statistical powers of 80% or 90% are determined for di erent numbers (6; 9; 12) of monthly follow-up scans. A mixed-e ffects linear regression model is used for this sample size determination.

Copula Models from Combining Sub-models for Different Groups   

Speaker:  Peijun Sang

The multivariate Student t copula is widely used in modeling the dependence structure of financial return data. However, it is subject to tail symmetry. The existence of tail asymmetry, particularly skewed to the joint lower tail, has been shown in a variety of situations. I will talk about how to employ skew-t distributions to allow for tail dependence and asymmetry.  Another interesting research topic I want to discuss is how to combine distributions of variables from different non-overlapping groups. Structured factor copula models proposed by Krupskii and Joe consist of one approach. I will introduce another method to handle this problem. Applications to financial return data will be included to make a comparison of these two methods. 

Building Statistical Models for the Prediction of Oral Cancer Recurrence

Oral cancer is a disease resulting from abnormal cell growth in the mouth, lips, tongue or throat, with high morbidity rate and high recurrence incidence.


Recently, Fluorescence Visualization (FV) has shown its value in identifying cancer tissues during surgeries, as well as facilitating early detection of cancer recurrence.

In a recent research project at BC Cancer Agency, oral cancer patients were recruited to study the means of predicting future recurrence. All the patients recruited in this program received surgery guided by the FV device and had follow-up visits regularly for many years. The lesion length and width under FV (FV measurement), as well as many other clinical risk factors were recorded at surgery and in follow-up visits.
 
In this project, we aim to build appropriate statistical models to link the risk factors and the cancer-free survival time of the patients after surgery. A Cox proportional hazards model is employed to fit the recurrence data against a selected subset of risk factors. Due to the existence of time-dependent risk factors (FV measurement), it is not appropriate to use the Cox model to predict future recurrence. Instead, we fitted logistic regression models between the 5-year cancer-free survival and three possible sets of constructed risk factors.  The prediction performance of the logistic models were assessed through a cross-validation study and based on the criteria of sensitivity, specificity and AUC.

Our data analysis reveals that the FV measurement is significantly associated with cancer recurrence based on the Cox regression model. It indicates that the FV measurement is informative about the cancer recurrence. The logistic regression analysis shows that it is possible to find an appropriate set of risk factors to give a prediction of the 5-year cancer free survival. The sensitivity, specificity and AUC of the fitted model are 0.449, 0.944, and 0.861, respectively.

EM-test and the related issues

In scientific investigations, a population is often suspected of containing several more homogeneous sub-populations. Such a population structure is most accurately described by a finite mixture model but the model should only be adopted with a statistically significant evidence through a rigorous hypothesis test. Developing valid and effective statistical inference methods on mixing distribution is an important research topic yet it posts serious technical challenges. Classical procedures when applied to mixture models often have sophisticated asymptotic properties which render them useless in applications. For a large number of finite mixture models, we have successfully designed corresponding EM-tests whose limiting distributions are easier to derive mathematically, simple for implementation in data analysis.

In this talk, we will first give a general introduction to EM-test. We will illustrate their elegant asymptotic properties and present a new approach to the tuning of ancillary parameters. We will also present some results on the sample size determination.