Lesson 12 — Pushing the allele frequency with selection

Put a steady thumb on the scale so one type gains a little each generation, then try to separate that push from the plain wandering underneath it.

Pine trees — do they want their seeds to be eaten? No. The trees that made pinecone seeds that were easy to eat — what happened to them? They died. The trees that survived are the ones that had pinecone seeds that weren't easy to eat. That's more or less all selection is doing — the ones that aren't good enough to make more of themselves don't make more of themselves.— 202_lec16_03

A — Deterministic selection — the s-curve

Infinite population, no drift. A favorable allele with coefficient s sweeps to fixation along a logistic curve.

Locked — confirm your name above to begin.

Scenario

Allele A has fitness 1 + s relative to a; AA / Aa / aa get genotype fitnesses (1+s)², (1+s), 1 in the haploid approximation. The frequency p evolves: p' = p(1+s)/(p(1+s)+(1−p)). Plot trajectory from p₀.

Frequency trajectory under selection

s: 0.05  |  p₀: 0.01  |  time to p=0.5: gen  |  time to p=0.99: gen

Prediction

  1. Q1. Halving s (from 0.05 to 0.025) changes the time to fixation by approximately:
Try at least 4 (s, p₀) combos. 0/4 combos

Controls

0.050
0.010

R code — deterministic selection

s <- 0.05p <- 0.01gens <- 500; traj <- numeric(gens+1); traj[1] <- pfor (g in 1:gens) { p <- p*(1+s)/(p*(1+s)+(1-p)); traj[g+1] <- p }

B — Selection + drift

Real populations are finite. Same selection coefficient, finite N. Some replicates sweep; some lose the beneficial allele to drift anyway.

Complete Stage A.

Scenario

50 replicate populations of N = 100. Same s = 0.05. Same p₀ = 0.05. Watch the spread. Some fix the allele; some lose it. Probability of fixation for a single new beneficial allele is ≈ 2s (Haldane).

50 replicate trajectories

N: 100  |  s: 0.05  |  p₀: 0.05  |  % replicates fixed:  |  Haldane prediction (2s): 10%

Prediction

  1. Q1. A new beneficial allele (1 copy, p₀ = 1/(2N)) with s = 0.05 in a population of N = 100 will fix about:
Try at least 5 (N, s, p₀) combos. 0/5 combos

Controls

100
0.050
0.050
42

R code — selection + drift

set.seed(42)N <- 100; s <- 0.05; p0 <- 0.05; reps <- 50trajs <- replicate(reps, {  p <- p0; traj <- p  for (g in 1:300) {    p_sel <- p*(1+s)/(p*(1+s)+(1-p))    p <- rbinom(1,2*N,p_sel)/(2*N); traj <- c(traj, p)  }; traj})

C — Estimate s from observed Δp

Given a type's frequency at two times, work out how hard it was being pushed. The range comes out wide when the change is small next to the wandering.

Complete Stage B.

Scenario

Simulate a sweep at a known push strength. Watch the frequency at the start and at time T, then work backward to the push. How tight your answer comes out depends on the population size, how long you watched, and how hard the push was.

Total miss across candidate push strengths

true s: 0.05  |  best-fit s:  |  95% range:

Prediction

  1. Q1. With N = 200, T = 50 generations, and true s = 0.02 (weak), the fit will:
Try at least 3 (s, N) combos. 0/3 combos

Controls

0.050
200
50
42

R code — fit s

set.seed(42)N <- 200; s_true <- 0.05; T <- 50; p0 <- 0.05p <- p0; traj <- pfor (g in 1:T) { p_sel <- p*(1+s_true)/(p*(1+s_true)+(1-p)); p <- rbinom(1,2*N,p_sel)/(2*N); traj <- c(traj, p) }s_grid <- seq(-0.05, 0.2, 0.001)ssr <- sapply(s_grid, function(s) { p2 <- p0; pp <- p2; for (g in 1:T) { pp <- pp*(1+s)/(pp*(1+s)+(1-pp)); p2 <- c(p2, pp) }; sum((traj - p2)^2) })s_grid[which.min(ssr)]

D — Lenski's long-term E. coli experiment

Real allele-frequency time courses from data/clean/ltee_allele_freqs.csv. Fit s to each rising allele. Some sweep cleanly; some get displaced.

Complete Stage C.

Scenario

LTEE allele frequencies plotted over generations. Each rising trajectory is a beneficial mutation. Fit s per trajectory. Watch for clonal interference — alleles that rise then crash as a fitter mutation appears in the population.

LTEE allele frequencies

trajectories shown:  |  median fit s:  |  range:

Prediction

  1. Q1. Across LTEE beneficial mutations that successfully fix, the typical selection coefficient is around:
Inspect at least 2 trajectories. 0/2 inspections

Controls

42

R code — fit s per LTEE trajectory

ltee <- read.csv("data/clean/ltee_allele_freqs.csv")# For each mutation, fit s by least squares to logisticmuts <- split(ltee, ltee$mutation_id)fits <- sapply(muts, function(m) {  t <- m$generation - min(m$generation); pobs <- m$freq; p0 <- pobs[1]  s_grid <- seq(0, 0.3, 0.005)  ssr <- sapply(s_grid, function(s) sum((pobs - 1/(1 + (1/p0 - 1)*exp(-s*t)))^2))  s_grid[which.min(ssr)]})

E — Drift, selection, or both? Five trajectories, five labels.

Same Wright-Fisher + selection simulator. Each round shows one trajectory. You decide which forces produced it before the truth is revealed.

Complete Stage D.

Scenario

Five trajectories, generated under known (N, s, p₀) combinations. Some are drift-only (s = 0). Some are selection-only-feeling (s much larger than 1/N). Some are the boundary case (s ≈ 1/N) where you can't tell from a single trace.

For each trajectory, pick the best label. The next round won't unlock until you commit.

Round 1 of 5

progress: 0 / 5  |  correct so far: 0

Classify this trajectory

Finish all 5 rounds. 0/5 rounds

Why this drill?

The boundary |s| ≈ 1/N is the most important single number in this lesson and you'll rely on it through Units 3 and 4. Below the boundary, selection is invisible against drift; above it, selection dominates. This drill makes you confront the boundary five times, on trajectories with known truth.