Showing posts with label population genetics. Show all posts
Showing posts with label population genetics. Show all posts

Monday, November 12, 2012

"Effective" population size

Bocquet-Appel-Recent_Advances_in_Palaeodemography-9781402064234

Chapter 1 by J Hawks.

In the absence of selection, allele frequencies vary as a stochastic process. The parameters influencing this process are themselves demographic: population size and mating pattern. Ultimately, the rate of evolution of a population must be constrained by these parameters. This means that the observable genetic characteristics of populations are to some extent natural estimators of demographic characteristics. The relationship between the demographic parameters of a population and its genetic characteristics may in some cases be approximated by a single parameter: the “effective population size.” Effective population size refers the demographic complexity of some real population to the simplicity of some ideal population — in other words, it is a measure of the extent to which a natural population corresponds to some theoretical population model.

...

The model-dependence of effective population size is rarely considered in analyses of molecular data. Ewens (2004) gives a good account of the problem:
Except in simple cases, the concept [of effective population size] is not directly related to the actual size of the population. For example, a population might have an actual size of 200 but, because of a distorted sex ratio, have an effective population size of only 25. This implies that some characteristic of the model describing this population, for example a leading eigenvalue, has the same numerical value as that of a Wright- Fisher model with a population size of 25. It would be more indicative of the concept if the adjective “effective” were replaced by “in some given respect Wright-Fisher model equivalent.” Misinterpretations of effective population size calculations frequently follow from a misunderstanding of this fact (Ewens, 2004, 37–38).

...

The utility of effective population size comes from the fact that it concatenates many separate stochastic phenomena into a single parameter. As an example, a gene frequency is a single value, with a single degree of freedom. It is therefore sufficient to estimate only a single parameter. This approach obviously runs into trouble when more than one stochastic factor varies in the population.

Monday, October 29, 2012

European Genetics Blog

http://eurogenes.blogspot.fr/2012/08/admixture-and-structure-tests-arent.html

This looks like a rich source of information on population genetics in general and european genetics in particular.

Wednesday, October 24, 2012

Molecular Clock Problems

Wessen-Simulating-human-origin-evo.pdf

Simulating Human Origins and Evolution by Ken Wessen (2005) pp 9-11:

As is apparent from the above discussion of the work by Cann et al. (1987), molecular methods rely on knowledge of the mutation rate of DNA across time and between species. The molecular clock hypothesis is a consequence of the neutral theory of evolution (Kimura, 1968) and implies an approximately constant rate of mutation, so long as the DNA sequence retains its original function. If this is the case, then the degree of difference between sequences being compared is simply proportional to the time since the sequences diverged. By incorporating fossil evidence, the clock can be calibrated, and thus divergence times can be attached to a molecular phylogeny.  

In fact, particular DNA sequences and proteins can mutate at vastly different rates at different times and in different lineages, and although there may be some local validity of the molecular clock hypothesis, in general there is global failure (Avise, 2000; Gibbons, 1998; Ruvolo, 1996; Strauss, 1999; Wills, 1995). The fast-mutating microsatellite loci, i.e. short repetitive sections of DNA that lie between genes, have been used to construct an alternative method for timing lineages that does not rely on external calibration of the rate of molecular evolution (Goldstein et al., 1995). However, because of mutational saturation, nuclear microsatellites are only useful for timing relatively recent events. In particular, the deepest split in the human phylogeny can be recovered with such a method, but saturation will occur in less time than the five million years or more back to the human–chimpanzee common ancestor (Jorde et al., 1998).

This situation also affects substantially the common ancestor calculations described above. For example, Wills (1995) includes a variable mutation rate across mtDNA sites and obtains a range of 436 000 to 800 000 years ago for the mitochondrial common ancestor, depending on the date used for the human–chimpanzee common ancestor.

 In general, the molecular data seem to support the replacement hypothesis, but when all the aforementioned caveats are considered, it remains far from conclusive. The dates vary widely, depending on the method and assumptions employed. Furthermore, a recent African origin has difficulty with the observed continuity of regional morphological traits, especially outside of Europe, whereas the multiregional hypothesis has difficulty with the amount of gene flow required for its support, as well as with a number of aspects of the molecular data. Perhaps the only thing that is truly clear is that population size, breeding patterns, local geographic events, migrations and reproductive barriers present a severe challenge when it comes to interpreting these results (Lahr and Foley, 1998). So long as positions at both extremes in this debate consider themselves equally well supported by the same data, be it fossil or molecular, substantial further study into the basis of all these methods is obviously of great importance.

An unintuitive truth

Human genetic diversity is always more pronounced within a given group than between given groups.

References:
Lewontin 1972
Barbujani et al. 1997
Jorde et al. 2000
Tomualdi et al. 2002

History of Simulating Evolution

Wessen-Simulating-human-origin-evo.pdf

Simulating Human Origins and Evolution by Ken Wessen (2005) p 12:

Raup et al. (1973) studied the generation of species lineages by modelling speciation as an equilibrium process of random lineage branching. All lineages stem from a common ancestor, and may continue in time, become extinct, or produce a new lineage by branching, with a probability based on the difference between the existing diversity and a predetermined equilibrium value. An algorithm for the automatic identification of clades was included, allowing study of the taxonomy of the resulting phylogeny. The simulations produced quite a variety of clade shapes, which were then compared with actual clades for the Reptilia. An important fact demonstrated by this work is that differences in evolutionary pattern do not necessarily imply an inherent difference in the associated taxonomic groups: simulated groups evolving under identical constraints can behave very differently. Sepkoski and Kendrick (1993) used a similar model to simulate phylogenies. Employing exponential, logistic and mass-extinction diversification profiles, the resulting phylogenies were degraded in various ways (to model the effects of fossilisation, for example) and the information content remaining was analysed with respect to the ‘true’ phylogeny. Both these models can be generalised to allow the study of higher taxa, e.g. genus, family, etc. Nee et al. (1994) also used a similar approach to study the reconstruction of phylogenies, looking particularly at the role of lineages that become extinct.

Qu'est-ce que c'est qu'une population ?

"C'est un joker." - Mme Degioanni
On parle souvent de:
  • Proximité géographique 
  • Même langue 
  • Partage réligion, culture, ethnie (moins de probabilité de mariages intergroupe)

Coalescent theory is a probabilistic framework

Wessen-Simulating-human-origin-evo.pdf Simulating Human Origins and Evolution by Ken Wessen (2005) pp 13: The study of a population is generally retrospective in nature, starting with a sample from an existing population and then attempting to describe the observed features in terms of the population’s prior evolution. Results are then generalised from the sample to the entire population. Coalescent theory (Kingman, 1982b; Hudson, 1990) provides a probabilistic framework perfectly suited to this approach, and has therefore become an extremely important tool in population genetics over the past 20 years. In brief, coalescent theory describes the merging of lineages from a sample of a population as one goes backwards in time, to the point where only a single lineage remains, i.e. the common ancestor. It is particularly well suited to molecular data, and although the usual formulation is based on the neutral model (Kimura, 1968) and a single, randomly mating population of constant size, extensions to cover recombination (Hudson, 1983), population growth (Kuhner et al., 1998), population subdivision (Hudson, 1990; Donnelly and Tavaré, 1995) and selection (Neuhauser and Krone, 1997) are well developed and are the subject of much ongoing research. A good review may be found in Fu and Li (1999).

Coalescent theory in a nutshell

From Coalescing into the 21st century: An overview and prospects of coalescent theory. Theor. Popul. Biol. 56, 1–10. By Fu and Li (1999) To infer the past from a sample taken from a present population, a new approach is required. Coalescent theory arose from this necessity. The essence of coalescent theory is to start with a sample, and trace backward in time to identify events that occurred in the past since the most recent common ancestor of the sample. Since the seminal work of Kingman (1982a, b), coalescent theory has been the most active topic in theoretical population genetics, and it is now widely recognized as the cornerstone for various statistical analyses of molecular population samples. The usefulness of the theory comes mainly from three features. First, it is a sample-based theory. Since the study of a population usually relies on a sample of individuals from that population, a theory that describes the properties of a sample is more relevant than the classical population genetics theory that describes the properties of the entire population. Second, it is a highly efficient approach. An important by-product of coalescent theory is the development of highly efficient algorithms for simulating population samples under various population genetics models, allowing various aspects of a model to be examined numerically. Third, coalescent theory is particularly suitable for molecular data, such as DNA sequence samples, which contain rich information about the ancestral relationships among the individuals sampled.