📚 A Complete Guide to the Chi-Square Test in Genetics | 卡方检验在遗传学中的应用指南
The chi-square test is one of the most important statistical tools in A-Level Biology. In genetics, it allows you to compare observed results from a genetic cross with the results that would be expected under a particular inheritance pattern, such as Mendelian ratios. This guide explains every step, from setting up hypotheses to interpreting critical values, with worked examples.
卡方检验是 A-Level 生物学中最重要的统计工具之一。在遗传学中,它让你能够将遗传杂交的实际观察结果与某种遗传模式(如孟德尔比例)下应有的预期结果进行比较。本指南将逐步讲解从建立假设到解释临界值的每一个环节,并配有完整例题。
1. What Is the Chi-Square Test? | 什么是卡方检验?
The chi-square test, written as χ² test, is a statistical method used to determine whether there is a significant difference between observed frequencies and expected frequencies. It answers the question: ‘Does the deviation from the expected ratio occur by chance alone?’
卡方检验(用 χ² 表示)是一种统计方法,用来判断观察频数与预期频数之间是否存在显著差异。它回答的问题是:“偏离预期比例这一现象是否仅由偶然因素造成?”
In genetics, after carrying out a cross, you count offspring in different phenotypic categories. You then ask whether these counts fit a predicted ratio, such as 3:1 for a monohybrid cross or 9:3:3:1 for a dihybrid cross.
在遗传学中,完成一次杂交后,你会统计不同表型类别中的子代数量。然后你需要判断这些数量是否符合预测的比例,例如单因子杂交的 3:1,或双因子杂交的 9:3:3:1。
2. Why Use the Chi-Square Test in Genetics? | 为什么在遗传学中使用卡方检验?
Observed results rarely match theoretical ratios exactly. Random fertilisation and sampling error mean that even a true 3:1 ratio will not produce exactly 150 round seeds and 50 wrinkled seeds for every 200 offspring. The chi-square test helps you distinguish between a result that fits the expected ratio and one that does not.
实际观察结果很少完全符合理论比例。随机受精和抽样误差意味着,即使是真正的 3:1 比例,在 200 个子代中也不一定会恰好产生 150 粒圆形种子和 50 粒皱缩种子。卡方检验帮助你区分实验结果是否符合预期比例,还是出现了偏差。
-
It tests the ‘goodness of fit’ between observed data and a genetic model.
它检验观察数据与遗传模型之间的“拟合优度”。
-
It allows you to reject or accept a proposed pattern of inheritance.
它帮助你在统计上拒绝或接受一种假定的遗传模式。
-
It is a required skill in the CIE A-Level Biology practical and essay examinations.
这也是 CIE A-Level 生物实验和论文考试中的必备技能。
3. Setting Up Hypotheses | 建立假设
Every chi-square test begins with a null hypothesis and an alternative hypothesis. You need to state these clearly before performing any calculation.
每一次卡方检验都从零假设和备择假设开始。在进行任何计算之前,你需要清楚写出这两个假设。
The null hypothesis, H₀, always states that there is no significant difference between observed and expected results. In genetics, H₀ usually takes the form: ‘The observed offspring fit the predicted Mendelian ratio.’
零假设(H₀)总是说明观察结果与预期结果之间没有显著差异。在遗传学中,H₀ 通常表达为:“观察到的子代表型比例符合预测的孟德尔比例。”
The alternative hypothesis, H₁, states that there is a significant difference between observed and expected results. In genetics, H₁ would be: ‘The observed offspring do not fit the predicted Mendelian ratio.’
备择假设(H₁)说明观察结果与预期结果之间存在显著差异。在遗传学中,H₁ 通常表达为:“观察到的子代表型比例不符合预测的孟德尔比例。”
4. Calculating Expected Values | 计算预期值
Expected values are the theoretical counts you would get if the null hypothesis were exactly true. To calculate the expected count for a category, multiply the total number of offspring by the probability of that category.
预期值是在零假设完全成立时你应当得到的理论数量。计算某一类别的预期数量时,用子代总数乘以该类别对应的概率。
Expected value E = Total number of offspring × Expected proportion
For a monohybrid cross with a 3:1 ratio, out of 200 offspring you would expect 150 dominant phenotype and 50 recessive phenotype. For a dihybrid cross with a 9:3:3:1 ratio, you would divide the total by 16 and multiply by 9, 3, 3 and 1 respectively.
对于比例为 3:1 的单因子杂交,200 个子代中你预期会得到 150 个显性表型和 50 个隐性表型。对于比例为 9:3:3:1 的双因子杂交,你需要将总数除以 16,再分别乘以 9、3、3 和 1。
Make sure you use the correct total. The sum of all expected values must equal the total number of offspring. If the expected frequency in any category is less than 5, the chi-square test is unreliable and you should not use it.
务必使用正确的总数。所有预期值之和必须等于子代总数。如果任一类别中的预期频数小于 5,卡方检验的结果就不可靠,此时不应使用该检验。
5. The Chi-Square Formula | 卡方公式
The chi-square statistic is calculated by summing the squared difference between observed and expected values, divided by the expected value, for every category. The Greek letter Σ means ‘the sum of’.
卡方统计量的计算方法是对每一个类别,取观察值与预期值之差的平方,再除以预期值,最后将所有类别的结果加总。希腊字母 Σ 表示“求和”。
χ² = Σ (O − E)² / E
Where O is the observed frequency and E is the expected frequency for each category. The larger the difference between O and E, the larger the value of χ², and the less likely the null hypothesis is to be true.
其中 O 是每个类别的观察频数,E 是每个类别的预期频数。O 与 E 之间的差距越大,χ² 值就越大,零假设成立的可能性也就越小。
Work through each category separately, then add all the term values together. Keep at least three significant figures during intermediate steps to avoid rounding errors.
逐个类别分别计算,最后将所有项加总。在中间计算过程中至少保留三位有效数字,以避免舍入误差。
6. Degrees of Freedom | 自由度
Degrees of freedom, usually written as ‘df’, is the number of independent categories that are free to vary once the total is fixed. It affects which critical value you use to judge your result.
自由度(通常写为 df)是在总数固定后,可以自由变动的独立类别数目。它影响你用来判断结果的临界值。
Degrees of freedom = Number of phenotypic categories − 1
For a monohybrid cross with two phenotypes (3:1), df = 2 − 1 = 1. For a dihybrid cross with four phenotypes (9:3:3:1), df = 4 − 1 = 3.
对于具有两种表型的单因子杂交(3:1),df = 2 − 1 = 1。对于具有四种表型的双因子杂交(9:3:3:1),df = 4 − 1 = 3。
You may also subtract additional degrees of freedom if you estimate a parameter, such as allele frequency, from the data. However, in standard A-Level genetics questions, the expected ratio is usually given, so the formula above is sufficient.
如果你需要从数据中估计参数(例如等位基因频率),还需要额外减去相应的自由度。但在标准 A-Level 遗传学题目中,通常直接给出预期比例,因此使用上面的公式就足够了。
7. Using the Critical Value Table | 使用临界值表
Once you have calculated χ² and determined the degrees of freedom, you compare your value with the critical value in a chi-square distribution table. The critical value tells you the maximum χ² that would be expected by chance alone at a given probability level.
计算出 χ² 并确定自由度后,你需要将你的数值与卡方分布表中的临界值进行比较。临界值给出了在给定概率水平下,仅凭偶然因素可能得到的最大 χ² 值。
In biology, the standard significance level is p = 0.05. This means that if the null hypothesis is true, there is less than a 5% probability that the observed deviation is due to chance alone.
在生物学中,标准显著性水平是 p = 0.05。这意味着如果零假设成立,观察到的偏差仅由偶然因素造成的概率小于 5%。
If χ² is less than the critical value, the difference is not significant and you accept H₀. If χ² is greater than or equal to the critical value, the difference is significant and you reject H₀ in favour of H₁.
如果 χ² 小于临界值,说明差异不显著,接受 H₀。如果 χ² 大于或等于临界值,说明差异显著,拒绝 H₀,接受 H₁。
| Degrees of freedom | Critical value at p = 0.05 |
| 1 | 3.84 |
| 2 | 5.99 |
| 3 | 7.82 |
| 4 | 9.49 |
| 5 | 11.07 |
You should memorise or be able to read this table quickly in the examination. In exams, the relevant table is usually provided.
你应当熟练记忆或快速读懂这张表。在考试中,通常会自动提供所需的临界值表。
8. Worked Example: Monohybrid Cross | 实例一:单因子杂交
A plant heterozygous for seed shape, Rr × Rr, is expected to produce round seeds and wrinkled seeds in a 3:1 ratio. A student counts 1000 seeds and observes 740 round seeds and 260 wrinkled seeds. Perform a chi-square test to determine whether these results fit the expected ratio.
一株种子形状为杂合的植物杂交,Rr × Rr,预期圆粒与皱粒的比例为 3:1。某学生统计了 1000 粒种子,观察到 740 粒圆粒和 260 粒皱粒。使用卡方检验判断这些结果是否符合预期比例。
Step 1: State the null hypothesis. ‘There is no significant difference between the observed numbers and the expected 3:1 ratio.’
第一步:写出零假设。“观察值与预期的 3:1 比例之间没有显著差异。”
Step 2: Calculate expected values. Total = 1000. Expected round = 1000 × 3 / 4 = 750. Expected wrinkled = 1000 × 1 / 4 = 250.
第二步:计算预期值。总数 = 1000。预期圆粒数 = 1000 × 3 / 4 = 750。预期皱粒数 = 1000 × 1 / 4 = 250。
Step 3: Calculate each term of χ².
第三步:计算 χ² 的每一项。
-
Round: (740 − 750)² / 750 = 100 / 750 = 0.133
圆粒:(740 − 750)² / 750 = 100 / 750 = 0.133
-
Wrinkled: (260 − 250)² / 250 = 100 / 250 = 0.400
皱粒:(260 − 250)² / 250 = 100 / 250 = 0.400
Step 4: Add the terms. χ² = 0.133 + 0.400 = 0.533.
第四步:将各项相加。χ² = 0.133 + 0.400 = 0.533。
Step 5: Determine degrees of freedom. df = 2 − 1 = 1. At p = 0.05, the critical value is 3.84. Since 0.533 is less than 3.84, the difference is not significant. We accept H₀ and conclude that the data fit the expected 3:1 ratio.
第五步:确定自由度。df = 2 − 1 = 1。在 p = 0.05 时,临界值为 3.84。因为 0.533 小于 3.84,差异不显著。我们接受 H₀,认为数据符合预期的 3:1 比例。
9. Worked Example: Dihybrid Cross | 实例二:双因子杂交
In a classic dihybrid cross between pea plants heterozygous for seed colour and seed shape, Mendel recorded 556 offspring in four phenotypes. The expected ratio is 9:3:3:1 for round-yellow, round-green, wrinkled-yellow and wrinkled-green respectively. The observed numbers are given below.
在一个经典的豌豆双因子杂交实验中,孟德尔记录了 556 个子代,涉及种子颜色和种子形状两对相对性状。预期四种表型比例分别为 9:3:3:1,即圆黄、圆绿、皱黄、皱绿。观察值如下所示。
| Phenotype | Observed O | Expected E | (O − E)² / E |
| Round-yellow | 315 | 312.75 | 0.016 |
| Round-green | 108 | 104.25 | 0.135 |
| Wrinkled-yellow | 101 | 104.25 | 0.101 |
| Wrinkled-green | 32 | 34.75 | 0.218 |
Expected values are calculated as follows: total = 556, so each unit is 556 / 16 = 34.75. Round-yellow = 9 × 34.75 = 312.75. Round-green = 3 × 34.75 = 104.25. Wrinkled-yellow = 3 × 34.75 = 104.25. Wrinkled-green = 1 × 34.75 = 34.75.
预期值计算如下:总数 = 556,因此每份为 556 / 16 = 34.75。圆黄 = 9 × 34.75 = 312.75。圆绿 = 3 × 34.75 = 104.25。皱黄 = 3 × 34.75 = 104.25。皱绿 = 1 × 34.75 = 34.75。
Step 4: Add all the final column values to obtain χ² = 0.016 + 0.135 + 0.101 + 0.218 = 0.470.
第四步:将最后一列所有值相加,得到 χ² = 0.016 + 0.135 + 0.101 + 0.218 = 0.470。
Step 5: df = 4 − 1 = 3. At p = 0.05, the critical value is 7.82. Since 0.470 is much less than 7.82, the deviation is not significant. We accept H₀ and conclude that the observed results are consistent with independent assortment.
第五步:df = 4 − 1 = 3。在 p = 0.05 时,临界值为 7.82。由于 0.470 远远小于 7.82,偏差不显著。我们接受 H₀,认为观察结果符合独立分配定律。
10. Common Mistakes and Exam Tips | 常见错误与应试提示
Many students lose marks in genetics questions not because they cannot calculate χ², but because they make small but avoidable errors in their reasoning and presentation.
许多学生在遗传学题目中失分,不是因为不会计算 χ²,而是在推理和表达中出现了一些小但可避免的错误。
-
Using percentages instead of counts. The chi-square test requires raw frequencies, not percentages or proportions.
使用百分比而不是计数。卡方检验要求使用原始频数,而不是百分比或比例。
-
Forgetting to confirm that expected frequencies are above 5. If an expected value is below 5, the test may not be valid.
忘记确认预期频数都大于 5。如果某个预期值小于 5,该检验可能不成立。
-
Using the wrong degrees of freedom. Count the number of phenotype classes, not the number of genes.
使用错误的自由度。要数表型类别的数量,而不是基因数量。
-
Stating ‘the results prove the ratio’ instead of ‘the results support the ratio’. A chi-square test provides evidence, not absolute proof.
说“结果证明了该比例”而不是“结果支持该比例”。卡方检验提供的是证据,不是绝对证明。
-
Forgetting to state the null hypothesis and the conclusion in a clear sentence with p = 0.05.
忘记写零假设,以及没有用包含 p = 0.05 的完整句子进行结论陈述。
-
Misplacing brackets when entering calculations into a calculator. Always calculate (O − E)² before dividing by E.
在计算器中计算时错误放置括号。一定要先算 (O − E)²,再除以 E。
In the exam, show every step clearly. Examiners award marks for the expected value, the formula, the χ² value, the degrees of freedom, the comparison with the critical value, and the biological conclusion. Even if your arithmetic is slightly wrong, a correct method and conclusion may still earn partial credit.
在考试中,要清晰展示每一步。阅卷官通常会按步骤给分,包括预期值、公式、χ² 值、自由度、与临界值的比较以及生物学结论。即使你的计算略有错误,正确的方法和结论仍可能获得部分分数。
Published by TutorHao | Biology Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导