Seminar

Multivariate investigations in ecology: The search for significance

Speaker's Page:  Tse-Lynn Loh, Marine Biologist, Ph.D. (2012), University of North Carolina Wilmington

Ecosystems are composed of multiple layers of interactions within and among biotic communities and their environments, leading to multivariate ecological studies. One of the fundamental questions in ecology is why biodiversity arose, and how it is maintained. Namely, why are species in a particular location, why and how do these species persist, and how do these species affect one another? My area of interest lies in trophic interactions, such as predator-prey dynamics, considering the flow of energy among various organisms as one of the main drivers of ecosystem function. Identifying the “true” ecological factors from a suite of possibilities is challenging, thus the tendency to rely on tests that produce a measure of statistical significance. As examples, I will draw on my research experiences with examining sponge composition on Caribbean coral reefs, and the spatial distribution of exploited seahorse populations in Southeast Asia. Because statistical significance does not always denote a significant ecological effect, I combined statistical outcomes with probable or confirmed mechanistic pathways. This allows for a fuller understanding of my study systems, a useful and necessary step for real-world applications, such as ecosystem-based fisheries management.

Improved nonparametric estimation of the number of zeros and illustrations

In this talk we present some lower bounds for the probability of zero for the class of count distributions having a log-convex probability generating function, which includes Compound and Mixed Poisson distributions. These lower bounds allow to construct non-parametric estimators of the non-observed number of zeros, which are useful in capture-recapture models. Some of these bounds lead to the well known Chao's and Turing's estimators. Several examples of application are analyzed and discussed.


 

Kalman filter, anti-coagulant therapy and life in general

Speaker:  Søren Lundbye-Christensen

Abstract

I will present a patented application of a statistical model for time series in connection with monitoring of warfarin treatment.

Anticoagulant monitoring is done by measurements of the International Normalized Ratio (INR) in bloodsamples in order to maintain INR within a certain terapeutic interval.
A noticeable between patients variation in response to warfarin entails individual dosing and frequent monitoring of the INR.

We have developed a monitoring algorithm based on a non-linear state space model using the Kalman filter to provide individualized dose suggestions. The model handles between-patients variations in dosage and in sensitivity to changes in dose, and also variations over time in dosage within-patient.

In my talk I will discuss the model and our retrospective and prospective validation of safety and the impact of the model on treatment quality. Also I will discuss my personal motivation to be involved in this research. Throughout the talk I will try to relate this project to the work I did, when I was so fortunate to visit UBC. 

As this visit was in 1993/94 it can be expected that I will illustrate the talk with nearly 25 years old photos.

Pseudo-observations for interval censored survival data using parametric estimates of the marginal survival function

Speaker:  Martin Berg Johansen

Abstract:  In event history analysis, periodic examinations may lead to event times which are known only to lie within a certain time interval. This can occur when a patient group is followed by routine controls or when a screening for a disease, for example cancer screenings, is performed in a population. In such cases, event times are said to be interval censored. Although a common phenomenon, interval censoring can be notoriously difficult to deal with analytically and is often unjustly ignored in applications.

Pseudo-observations have been proposed by Andersen, Klein and Rosthøj1 and can be used to formulate regression models using a non-parametric estimator of the marginal survival function, or cumulative incidence function in the presence of competing risks, by applying a generalized linear model to the derived pseudo-observations. Pseudo-observations enable routine regression analysis of clinically relevant effect measures, including risk ratios, risk differences, and the restricted mean survival time. The theory behind pseudo-observation methods does not, however, cover existing non-parametric estimators of the survival function based on interval censored data.

We aimed to construct a simple method for generating pseudo-observations in the context of interval censored event times. To this end, we used the approach of Royston and Parmar2 as a preliminary step to constructing a flexible spline-based parametric estimator of the marginal survival function.

References

1)      Andersen, Biometrika 2003; 90:15–27
2)      Royston, Stat Med. 2002; 21(15):2175–2197

Distribution-Free Testing; the Khmaladze Transform

The Khmaladze transform takes the vector of components of Pearson's chi-square statistic to another vector which contains the same "statistical information" but is asymptotically distribution-free. Hence any test statistic based on the new vector is also asymptotically distribution-free. Natural examples are goodness-of-fit statistics based on partial sums.

A version of the Khmaladze transform for testing whether data comes from a continuous distribution function, F, maps the normalized error function, which is asymptotically an F-bridge, to a process which is asymptotically a standard Brownian bridge. The associated statistic, which is empirically based, can be used for distribution-free testing.

Structured factor copula models for multivariate data

In factor copula models for multivariate data, dependence is modeled via one or several common factors. The models allow great flexibility in modeling different types of dependence structure including tail dependence and asymmetry. We propose two structured factor copula models for the case where variables can be split into non-overlapping groups such that there is homogeneous dependence within each group. A typical example of such variables occurs for stock returns from different sectors.

We use some tail-weighted measures of dependence to select appropriate copulas in the model and to assess the adequacy of fit in the tails. We apply the structured factor copula models to analyze a financial data set, and compare with other copula models for tail inference.

Statistical methods for genetic association studies with rare variants

The recent focus on genetic rare variants has produced a large number of testing strategies to assess association between a group of rare variants and a heritable trait, with competing claims about the performance of various tests.  We review frequently used tests and show that they fall into either linear or quadratic classes, and neither class consistently outperforms the other across genetic models (Derkach, Lawless and Sun online, Statistical Science).  This understanding leads to development of robust tests that borrow strength from the two complementary classes (Derkach, Lawless and Sun 2013, Genetic Epidemiology).  However, regardless of the specific test used, power is generally low in current realistic settings due to various factors including rareness of the key variants and multiple hypothesis testing.  To increase power, we investigate various cost-effective response-dependent sampling strategies. This is joint work with graduate student Andriy Derkach and Professor Jerry Lawless.

Phylogenetic analysis of species radiations using SNPs and AFLPs. [Bio/Bioinformatics/Genetics]

Technological wonders such as next generation sequencing mean that we can now, in principle, obtain SNP (single nucleotide polymorphism) data from multiple individuals in multiple species. This promises enormous benefits for population genetic and phylogenetic analysis, particularly of closely related or poorly resolved species. My interest is in how to analyse these data effectively and responsibly. We have developed an algorithm which estimates species trees, divergence times, and population sizes from independent (binary) makers such as well spaced SNPs. The method is based on coalescent theory (like the BEAST software), though it uses mathematical trickery to avoid having to consider all the possible gene trees. As a `full likelihood' method, it should be more accurate than alternative FST based approaches. I'll talk about our experiences applying this method to AFLP data from alpine plants, and some recent discoveries about the usefulness (or uselessness) of SNP data for estimating population sizes.

Bio: David did his PhD with Mike Steel at the University of Canterbury. He had postdocs with David Sankoff (Montreal) and Olivier Gascuel (Montpellier), before taking up a position at McGill. After tenure, he moved back to NZ for positions at the University of Auckland, and more recently Otago, which is in Dunedin in the South Island of NZ.

The Lasso: a brief review and a new significance test

I will review the lasso method and show an example of its utility in cancer diagnosis via mass spectometry. Then I will  consider the testing the significance of the terms in afitted regression, fit via the lasso. I will present a novel test statistic for this problem, and show that it has a simple asymptotic null distribution. This work builds on the least angle regression approach for fitting the lasso, and the notion of degrees of freedom  for adaptive models (Efron 1986) and for the lasso (Efron et. al 2004, Zou et al 2007). We give examples of this procedure, discuss extensions to generalized linear models and the Cox model, and describe an R language package for its computation.

This work is joint with Richard Lockhart (Simon Fraser University), Jonathan Taylor (Stanford Univ), and Ryan Tibshirani (Carnegie Mellon University).
 

Hypothesis testing under density ratio models in the presence of Type I censored multiple samples

Maintaining a high quality of lumber products is of great social and economic importance. As part of a research program aimed at developing a long term program for monitoring change in the strength of lumber, we develop a theory of hypothesis testing concerning a given number of populations with Type I censored samples from each. Statistical methods for lumber quality monitoring should ideally be efficient and nonparametric. These desiderata lead us to adopt a semiparametric density ratio model to pool the information across multiple samples and use the nonparametric empirical likelihood (EL) as the tool for statistical inference. We establish a powerful framework for performing EL inference under the density ratio model when Type I censored samples are present. This inference framework centers on a concave dual partial EL, and features an easy computation. We find that under a class of general composite hypotheses, the corresponding EL ratio test has a classical chi--square limiting distribution under the null model and a non--central chi--square limiting distribution under local alternatives. We show that the local power of this EL ratio test is often increased when strength is borrowed from additional samples even when their underlying distributions are unrelated to the hypothesis of interest. Simulation studies show that this test has better power properties than all potential competitors adopted to the multiple sample problem under the investigation, and is robust to model misspecification. The proposed test is then applied to assess strength properties of lumber with intuitively reasonable implications for the forest industry.