The sample variance of all N yields is \(s^2 = \text{SST} / (N - 1)\). Its numerator, SST, is the total variation in yield: the thing the F-test tries to explain.
Each plot's distance from the grand mean splits into a fertilizer part (its fertilizer's average minus the grand mean) and a noise part (its yield minus its fertilizer's average). Squared and added up, the two parts split SST exactly.
Even a season of duds explains a little by luck, since each fertilizer's average soaks up some noise. Dividing by degrees of freedom makes the comparison fair: when every bag is a dud (\(H_0\)), MSM and MSE both estimate \(\sigma^2\), so F is near 1.
Under \(H_0\), F follows the F distribution with \(k-1\) and \(N-k\) degrees of freedom (df). The p-value is the area to the right of the observed F. The critical value \(F_{\alpha}(k-1,\,N-k)\) cuts off a right tail of area \(\alpha\); it is the number in the F table.
SSM is the extra sum of squares: SSE for one shared average (the reduced model) minus SSE for an average per fertilizer (the full model).
One die has standard deviation \(\sqrt{35/12} = 1.71\) lb. At that σ each yield is soil 7 plus one die roll plus the fertilizer's effect. Other σ values stretch the same rolls away from 3.5, so each season keeps its shape as the noise changes.
The full model gives each fertilizer its own mean. The reduced model forces them all equal, so it predicts every plot with the grand mean. It is the full model with some of its coefficients set to zero, which is what nested means.
A model's SSE is its residuals squared and added. Its residual df is the number of plots minus the number of means it estimates.
Since SST = SSM + SSE, the drop in leftover is SSM, and the df it cost is k − 1. Put those into the full-vs-reduced formula and you get MSM ÷ MSE.
The full-vs-reduced F tests any set of terms: drop them, refit, and compare the drop per df with the full model's leftover per df. With blocks in the model, the reduced model keeps the blocks and drops the fertilizer. MSM ÷ MSE is the one-way case, where the reduced model is a single mean.
anova(lm(yield ~ 1), lm(yield ~ fertilizer)) prints each model's RSS, the extra sum of squares as Sum of Sq, and F. summary(aov(yield ~ fertilizer)) prints the cut in the fertilizer row and SSE in the Residuals row. Both give the same F and p.
summary(aov()) builds its table one term at a time. Each row's Sum Sq is the extra sum of squares when that term joins the terms above it, so the reduced model for fertilizer's row is whatever entered first:
When every fertilizer-and-garden cell holds the same number of plots, the two factors are orthogonal: each explains a separate piece of SST, and both orders print the same table. When the cells are lopsided, the two factors overlap, and a factor entered first takes credit for whatever it shares with the factors after it. The Residuals row is the same in either order.
One trial gives one difference, F − P. Rerun it many times and the differences pile up: around 0 if the fertilizer is a dud (the null hypothesis, \(H_0\), orange), and around δ if it adds δ pounds (the alternative hypothesis, \(H_a\), blue). Both spread by one standard error (SE).
The area under a curve over a range of differences is the share of trials that land there. That is why α, β, and power are areas.
A difference beyond ±c is significant. α is the orange area beyond ±c: a false alarm on a dud. β is the solid blue area between −c and +c: a miss of a fertilizer that works. Power, 1 − β, is the hatched blue area.
Differences where the curves overlap could come from either hypothesis, and the line has to cut through the overlap, so more overlap means more misses at the same α. Two normal curves with one spread overlap by \(2 \times \texttt{pnorm}(-\delta/2,\ 0,\ SE)\). More plots, balance, blocking, and a bigger effect all shrink it. The curves must cross somewhere in the overlap, and there is nothing special about that point.
The smallest true effect the design hits 80% of the time. Slide the blue curve right until the critical line cuts off its lower 20%. The line is then 1.96 SE above 0 and 0.84 SE below the blue curve's center.
A lopsided split gives the SE of a smaller balanced one. Two F plots and six P plots give \(n_{\text{eff}} = 3\) per group: eight plots for the precision of six.
Each garden's soil (two dice, SD 2.42 lb) moves every plot in that garden together.
Blocked. F and P plots share each garden, so the soil cancels from F − P and σ is the plot noise, 1.71.
Scattered. Every plot has its own garden, so soil joins the noise: \(\sigma = \sqrt{1.71^2 + 2.42^2} = 2.96\).
Whole garden. All F plots share one garden and all P plots another, so the soil difference between those two gardens is part of every F − P. The plots' own spread understates the SE, and the line falls far inside the dud curve.
In a real trial σ comes from the data, so the statistic is t = difference ÷ estimated SE. Error df are counted where the treatment was assigned: N − 3 when blocked in two gardens, N − 2 when scattered. Few df fatten the tails and push the critical value out (2.57 SE on 5 df). The page draws t × SE so the axis stays in pounds. The blue curve then shows where t × SE falls when the fertilizer really adds δ. Its tails are fatter than the σ-known curve's, and its right tail is the longer one: a trial whose estimated SE comes out small pushes t far to the right.
Each dot is one season's F = MSM / MSE: the 200 dud seasons from Build the F above the middle line, and the same seasons with the fertilizer's effect added below it. When every bag is a dud (\(H_0\)), F follows the F distribution with \(k-1\) and \(N-k\) df: the orange curve, the one the F table comes from. When a bag works (\(H_a\)), the blue curve shows where F falls, farther right. The 200 seasons cover up to 4 fertilizers and 12 plots; the Anvil goes further.
Power is the share of the working pile past the line. R's power.anova.test() finds it from the number of fertilizers (groups), the plots per fertilizer (n), the variance of the true averages (between.var, R's var()), and the noise variance σ² (within.var). Brand A adding 3 lb to Plain, with 4 plots each and σ = 1.71 lb, has power 0.55:
power.anova.test(groups = 2, n = 4, between.var = var(c(3, 0)), within.var = 2.92)
More plots push the working pile right, since each average counts once per plot. They also raise \(df_2\), so the critical value falls toward its floor. Both add power. Noise scales MSM and MSE alike, so a dud's F never changes with σ.
With k = 2, F is the square of the two-sample t statistic, and the critical F is the critical t squared: \(5.99 = 2.447^2\) on 6 df.
Squaring also explains the spike at 0. Every t between −0.3 and 0.3 squares to an F below 0.09, so with k = 2 both F curves rise without bound as F nears 0. Every F distribution with 1 numerator df has that shape.
Two Curves with σ known uses 1.96. An F test estimates σ from \(N-k\) df, so its critical value is larger and its power smaller. The balanced Tomato Trial design: 0.70 with σ known, 0.55 on 6 df.
One bag works: one fertilizer adds δ and the rest are plain. Evenly spread: the true averages step evenly from 0 to δ. With three or more fertilizers, the even spread has the same biggest effect, a smaller between.var, and less power.
After Douglas C. Montgomery, Design and Analysis of Experiments, 8th ed., on sample size for the fixed-effects ANOVA.
The grand mean averages every plot. The intercept μ is the mean of means, the average of the fertilizer means:
The grand mean counts each plot once and μ counts each fertilizer once, so the two are equal when every fertilizer has the same number of plots. A fertilizer's effect is its mean minus μ, and a residual is one plot minus its fertilizer's mean:
Every df comes from the design. With one factor at k levels on N plots, the factor gets k − 1, Error gets N − k, and Total gets N − 1. The df and SS columns add up to the Total. In every row, MS = SS ÷ df. Each term's F divides its mean square, MSM, by the Error row's, MSE:
An F past the line \(F_{0.05}(\text{df}, \text{df}_{\text{Error}})\) is the same event as p below 0.05. It is evidence that at least one average differs. A follow-up comparison, such as Tukey's intervals, says which.
power.anova.test() takes the number of fertilizers (groups), the plots per fertilizer (n), the variance of the true averages (between.var), the noise variance σ² (within.var), α (sig.level), and the power. Leave one out and R solves for it. A solved n is rarely a whole number. Round up, since rounding down leaves the power under the target.