Evolution is the study of the change in the frequency of transmissible traits over time. We'll need to specify a unit that can transmit traits, and we'll examine a group made of many of these units. That is an abstract way of describing the types of data we look at evolutionarily. For population genetics the unit will (usually) be an organism, the transmissible trait will be an allele, and what we are concerned with are the frequencies of those alleles in populations.
Whether those traits influence their own transmission is an important distinction. If a trait influences its own transmission, we say it is selected — it can increase or decrease the probability that it is transmitted, which is what the direction of selection means. But an influence of the trait on whether it is transmitted is not necessary. Any time there is unequal transmission, any time some traits are transmitted more or less often for any reason, the frequency will change. And the smaller a population is, the more dramatically those fluctuations will vary.
Below you are going to run some simulations. You control the number of individuals in the population — imagine a population, or a prairie, or some other enclosed space — and the unevenness of how much they replicate. If you put a zero for the differences in expected number of offspring, you are not saying they all reproduce evenly. You are saying they all have an equal likelihood of reproducing. For this activity the trait has no impact at all on whether it is transmitted. We are only looking at how populations drift from wherever they began to wherever they end up.
# one population. every individual gets a share of the next generation.N <- 100cv <- 0 # differences in expected number of offspring, mean fixed at 1w <- if (cv == 0) rep(1, N) else rgamma(N, shape = 1/cv^2, scale = cv^2)p <- w / sum(w) # nothing here looks at colour# each of the N offspring draws two parents with those sharesmo <- sample(N, N, replace = TRUE, prob = p)fa <- sample(N, N, replace = TRUE, prob = p)# individuals that left nothingsum(tabulate(c(mo, fa), nbins = N) == 0)# the frequency of the purple allele after they breedmean(kid)
In part A we watched one population drift, over and over again, from different starting points. We can also watch several populations drift at the same time, from the same starting point, and see how far apart they get.
A core feature of drift is that each generation inherits the random noise of the generation before it. That is what makes the frequency move in a way that looks, for any one population, like a trend — you could see it in part A, where an allele would climb or fall as though some force were driving it somewhere on purpose. Nothing was.
The switch below decides whether that noise is inherited. On one setting the parents are drawn from the generation before, so the error accumulates. On the other, every generation is drawn from the frequency the run started at, so the same amount of chance happens every generation and none of it carries forward. That second setting is, roughly, what migration does, and migration comes in a few weeks.
Below you set yourself a target — how far apart the twenty end up — and then move the number of individuals, the number of generations and the switch until you hit it. If a target turns out to be one nothing reaches, abandon it and move on.
# the same population, two rules for where the parents come fromN <- 50gens <- 100g0 <- rep(c(0, 1), each = N) # the population you started withg <- g0for (t in 1:gens) { src <- if (inherit) g else g0 # <- the whole switch g <- sample(src, 2*N, replace = TRUE)}# how far it ended up from where it startedabs(mean(g) - mean(g0))# one of each allele: two copies pulled at random, how often they differ2 * mean(g) * (1 - mean(g))
Peter Buri took a great many fruit flies and divided them into 107 bottles, each one holding a self-sustaining population. He then watched the frequency of one allele change over time in every one of those 107 bottles. Flies have short generation times, which made this feasible. He controlled how many flies made it into each new bottle. And because the gene showed up as a visible characteristic, he could track it in every bottle — before anybody could sequence anything.
Drift destroys variation. What happened in parts A and B is that you watched, over and over, different populations becoming more different from one another — because different alleles were going extinct. If an allele randomly goes to zero it is gone forever. And alleles will randomly go to zero, because in a pure drift scenario there is nothing keeping them around. With nothing maintaining that variation, the system inevitably collapses to one allele, no matter how many you begin with.
The time it takes to collapse is, in a pure drift scenario, a function of only two things: the starting frequency of the allele, and the size of the population.
You are going to simulate that experiment: 107 bottles, all the same size, all starting at the same frequency.
# 107 populations, same size, same start, nothing selectedN <- 25gens <- 60p <- rep(0.5, 107)H <- numeric(gens+1); H[1] <- mean(2*p*(1-p))for (t in 1:gens) { p <- rbinom(107, 2*N, p) / (2*N) H[t+1] <- mean(2*p*(1-p))}# populations with one allele leftsum(p == 0 | p == 1)# the fraction that goes missing each generation1 - exp(coef(lm(log(H) ~ seq_along(H)))[2])
For this final part, we're going to look at a real population that you've already examined before: the moose of Isle Royale. We looked at this moose population in Lesson 6, and in order to fit the growth curve you had to incorporate a few bad years. These represent population bottlenecks. The result is that the trajectory of their population over time has fluctuated. They have not held steady, as the populations you looked at previously in this lesson did.
You control how many loci you follow and how many alleles each locus has, and try to find the size of a steady moose population that the real one is evolving like. The real population ranges from about 500 to about 2400 moose. The rate of drift, however, is quite high comparatively.
The top plot shows the population number. The bottom plot shows the heterozygosity for your given set of parameters. You will adjust the constant herd and find the effective population size that matches the observed rate of decline in heterozygosity in your moose.
Anything that makes some individuals reproduce more than others makes evolution faster. It makes things change more over time. Regardless of whether it covaries with traits.
d0 <- read.csv("data/clean/isle_royale.csv")d0 <- d0[d0$year <= 2000, ]N_rec <- d0$mooseN_you <- 500loci <- 40k <- 4male_risk <- 8 # a male's risk in a bad winter, vs a female'sherdRun <- function(census) { n <- census[1] g <- array(sample.int(k, 2*loci*n, TRUE), c(2, loci, n)) sex <- sample(c("m","f"), n, TRUE) h <- numeric(length(census)) for (t in seq_along(census)) { h[t] <- mean(g[1,,] != g[2,,]) if (t == length(census)) break m <- which(sex == "m"); f <- which(sex == "f") # the winter takes the shortfall, males first, spilling onto females am <- length(m); af <- length(f) for (d in seq_len(max(0, n - min(n, census[t+1])))) { if (runif(1) < male_risk*am/(male_risk*am + af) && am > 1) am <- am - 1 else if (af > 1) af <- af - 1 else if (am > 1) am <- am - 1 } m <- sample(m, am); f <- sample(f, af) w <- rexp(am)^2 # strong Vk: a few males sire most n <- census[t+1] dad <- sample(m, n, TRUE, prob = w) mom <- sample(f, n, TRUE) gn <- array(0L, c(2, loci, n)) gn[1,,] <- g[cbind(sample(1:2, loci*n, TRUE), rep(1:loci, n), rep(dad, each=loci))] gn[2,,] <- g[cbind(sample(1:2, loci*n, TRUE), rep(1:loci, n), rep(mom, each=loci))] g <- gn; sex <- sample(c("m","f"), n, TRUE) } h}hm <- herdRun(N_rec) # Isle Royalehy <- herdRun(rep(N_you, length(N_rec))) # your steady herd1 - tail(hm,1)/hm[1] # heterozygotes Isle Royale lost1 - tail(hy,1)/hy[1] # heterozygotes your herd lostmean(N_rec) # the plain averagelength(N_rec) / sum(1/N_rec) # the harmonic mean