📚 Edexcel IGCSE Statistics: Interdisciplinary Integrated Question Practice | Edexcel IGCSE统计:跨学科综合题型训练
In the Edexcel IGCSE Statistics examination, questions increasingly require you to apply statistical methods across a range of real-world contexts. These interdisciplinary problems blend data handling, probability and inferential statistics with topics from biology, economics, geography and the social sciences. The ability to recognise which statistical tool to use, to interpret outputs in context and to communicate findings clearly is essential. This article provides comprehensive practice, covering the most common cross-subject question types, complete with worked examples, key formulas and examiner advice.
在爱德思IGCSE统计考试中,越来越多题目要求你在真实的跨学科情境中应用统计方法。这些综合题型把数据处理、概率与推断统计同生物学、经济学、地理学及社会科学主题结合在一起。能否辨识该选用哪种统计工具、在语境中解读结果并清晰传达结论至关重要。本文提供全面的综合训练,涵盖最常见的跨学科题型,配有范例、关键公式和考官建议。
1. Biology: Heart Rate and Data Collection | 生物学:心率与数据收集
Suppose a biologist measures the resting heart rates (beats per minute) of 20 students before and after a short exercise session. The raw data must be organised into a back-to-back stem-and-leaf diagram to compare the two distributions. A good stem-and-leaf diagram uses a sensible key, orders the leaves and splits stems where necessary. From the diagram, you can determine the median and interquartile range (IQR) for each set. In IGCSE Statistics, always check for outliers using the rule Q₁ − 1.5 × IQR and Q₃ + 1.5 × IQR.
假设一位生物学家测量了20名学生在短暂运动前与后的安静心率(次/分)。原始数据必须整理成背靠背茎叶图以比较两个分布。好的茎叶图需要合理的图例、有序的叶并在需要时劈开茎。从图中可以确定每组的中位数与四分位距(IQR)。在IGCSE统计中,务必使用 Q₁ − 1.5 × IQR 与 Q₃ + 1.5 × IQR 规则检查离群值。
The mean resting heart rate before exercise was 72 bpm with a standard deviation of 8 bpm. After exercise, the mean rose to 98 bpm with a standard deviation of 12 bpm. When comparing data sets, always mention both a measure of central tendency and a measure of spread. Here, exercise not only raised the average heart rate but also increased variability, suggesting that individuals respond to exercise differently. A paired t-test would be the appropriate inferential method if we wanted to test whether the increase is statistically significant, although for IGCSE you are more likely to calculate the standardised difference or use a box plot comparison.
运动前的平均安静心率为72 bpm,标准差为8 bpm。运动后均值升至98 bpm,标准差为12 bpm。比较数据集时,务必同时提及集中趋势度量和离散度量。在此,运动不仅提高了平均心率,还增大了变异性,说明个体对运动的反应存在差异。若想检验升高是否统计显著,配对t检验是合适的推断方法,不过IGCSE阶段更常见的是计算标准化差值或用箱线图进行比较。
2. Economics: Consumer Price Index and Weighted Averages | 经济学:消费者价格指数与加权平均
A typical economics application involves calculating a weighted price index. The table gives prices and quantities for three commodities in a base year and a current year. To find the Laspeyres index, use the formula: ∑(pₙ × q₀) / ∑(p₀ × q₀) × 100, where pₙ is the current price, p₀ the base price and q₀ the base quantity. This index measures how much more expensive the base-year basket has become. The Paasche index uses current quantities instead, and the Fisher index is the geometric mean of Laspeyres and Paasche.
典型的经济学应用题涉及加权价格指数的计算。表格给定了基年与当年的三种商品价格及数量。要计算拉氏指数,使用公式:∑(pₙ × q₀) / ∑(p₀ × q₀) × 100,其中 pₙ 为年价格,p₀ 为基期价格,q₀ 为基期数量。该指数衡量基期篮子变贵了多少。帕氏指数则改用当年数量,而费雪指数是拉氏指数与帕氏指数的几何平均。
| Commodity | Base Price (p₀) | Current Price (pₙ) | Base Quantity (q₀) |
|---|---|---|---|
| Bread | £1.00 | £1.20 | 50 |
| Milk | £0.80 | £0.95 | 30 |
| Eggs | £2.00 | £2.50 | 20 |
Calculate the Laspeyres price index: numerator = (1.20×50)+(0.95×30)+(2.50×20) = 60+28.5+50 = 138.5. Denominator = (1.00×50)+(0.80×30)+(2.00×20) = 50+24+40 = 114. Index = (138.5/114)×100 ≈ 121.5. This means the cost of the base-year basket has increased by about 21.5%. In IGCSE questions, you may also be asked to interpret the index and discuss its limitations, such as substitution bias or the introduction of new goods.
计算拉氏价格指数:分子 = (1.20×50)+(0.95×30)+(2.50×20) = 60+28.5+50 = 138.5;分母 = (1.00×50)+(0.80×30)+(2.00×20) = 50+24+40 = 114;指数 = (138.5/114)×100 ≈ 121.5。这意味着基期篮子的花费上涨了约21.5%。在IGCSE题目中,还可能要求你解释该指数并讨论其局限性,比如替代偏差或新商品推出的影响。
3. Geography: Population Pyramids and Demographic Data | 地理学:人口金字塔与人口数据
Population data are often presented in grouped frequency tables with age intervals. A population pyramid is essentially two back-to-back horizontal bar charts showing the male and female distributions. When you are asked to construct a population pyramid, you must use a suitable scale, label axes clearly, and remember that the male bars extend to the left and female bars to the right. The shape of the pyramid reveals whether a population is expanding, stable or contracting. Broad bases indicate high birth rates, while narrow tops reflect lower life expectancy.
人口数据常以年龄分组的频数表呈现。人口金字塔本质上就是两个背靠背的水平条形图,分别显示男性和女性分布。当要求绘制人口金字塔时,你必须选用合适的比例尺,清晰标注坐标轴,并记住男性的条形向左延伸、女性向右延伸。金字塔的形状揭示人口是扩张型、稳定型还是收缩型。宽阔底部表示高出生率,狭窄顶部反映较低的预期寿命。
To derive statistical measures from such data, you need to estimate the median age and the interquartile range. Because the data are grouped, use linear interpolation within the appropriate class interval. For example, if the cumulative frequency reaches the halfway mark in the 20–29 age group, the median is calculated as lower class boundary + ( (n/2 − previous cumulative frequency) / class frequency ) × class width. A population with a median age of 22 years is markedly younger than one with a median of 40 years, which has implications for dependency ratios and economic planning.
要从这类数据中得出统计量,需要估计年龄中位数和四分位距。由于数据是分组的,应使用线性插值法在相应组距内计算。例如,若累计频数在20–29年龄组达到一半,则中位数 = 组下限 + ( (n/2 − 前累计频数) / 组频数 ) × 组距。中位年龄22岁的人口明显比中位年龄40岁的人口年轻,这对抚养比和经济规划具有深远影响。
4. Physics: Measurement Errors and Standard Deviation | 物理学:测量误差与标准差
In physics experiments, repeated measurements often vary due to random errors. You might be given 10 readings of the time period of a pendulum: 1.25, 1.29, 1.24, 1.30, 1.27, 1.26, 1.28, 1.23, 1.31, 1.27 seconds. The mean is 1.27 s. To quantify precision, calculate the standard deviation using the formula s = √[ Σ(x − x̄)² / (n−1) ]. A smaller standard deviation indicates more consistent measurements. An anomalous result can be identified using the ±2s rule, which in IGCSE is often simplified to ±2 standard deviations from the mean.
物理实验中,重复测量常因随机误差而产生差异。你可能会得到单摆周期的10个读数:1.25, 1.29, 1.24, 1.30, 1.27, 1.26, 1.28, 1.23, 1.31, 1.27 秒。平均值是1.27 s。为量化精密度,使用公式 s = √[ Σ(x − x̄)² / (n−1) ] 计算标准差。标准差越小说明测量越一致。运用 ±2s 规则可识别异常值,在IGCSE中通常简化为与平均值相差2个标准差的规则。
Calculate the deviations: −0.02, +0.02, −0.03, +0.03, 0, −0.01, +0.01, −0.04, +0.04, 0. Squared: 0.0004, 0.0004, 0.0009, 0.0009, 0, 0.0001, 0.0001, 0.0016, 0.0016, 0. Sum = 0.0060. Divide by 9: 0.0060/9 ≈ 0.000667. s = √0.000667 ≈ 0.0258 s. Any reading outside 1.27 ± 2×0.0258, i.e. (1.2184, 1.3216), would be suspect. All data fall within this range, so no outliers. This method is commonly assessed alongside the concept of absolute and percentage uncertainty.
计算偏差:−0.02, +0.02, −0.03, +0.03, 0, −0.01, +0.01, −0.04, +0.04, 0。平方后求和:0.0004+0.0004+0.0009+0.0009+0+0.0001+0.0001+0.0016+0.0016+0 = 0.0060。除以9:0.0060/9 ≈ 0.000667。s = √0.000667 ≈ 0.0258 s。任何超出 1.27 ± 2×0.0258,即 (1.2184, 1.3216) 的读数都可疑。所有数据均在此范围内,因此没有离群值。这种方法常与绝对不确定度和百分比不确定度的概念一同考查。
5. Environmental Science: Carbon Emissions and Time Series | 环境科学:碳排放与时间序列
An environmental study records the annual CO₂ emissions (in million tonnes) of a country over 12 years. A time series graph helps identify the trend, seasonal variation and any cyclical patterns. Moving averages are used to smooth out irregular fluctuations. For annual data, a 3-point or 5-point moving average is suitable. For quarterly data, a 4-point centred moving average is standard. Once the trend is determined, you can make short-term forecasts by extrapolating the trend line, though IGCSE questions explicitly remind you that such forecasts become less reliable the further into the future you go.
一项环境研究记录了一个国家12年间每年的二氧化碳排放量(百万吨)。时间序列图有助于识别趋势、季节变化和任何周期性模式。移动平均数用来平滑不规则波动。对年度数据,3点或5点移动平均较为合适。对季度数据,则使用4点中心移动平均为标准做法。一旦确定了趋势,就可以通过延伸趋势线进行短期预测,不过IGCSE题目会明确提醒你,预测的未来越远越不可靠。
When modelling emissions, a line of best fit by eye is acceptable in IGCSE, but you may also calculate the least squares regression line. The equation takes the form y = a + bx where b = Σ(x − x̄)(y − ȳ) / Σ(x − x̄)². For example, if the years (coded as 1 to 12) and emissions show a positive linear relationship, the gradient b represents the average annual increase in emissions. The coefficient of determination R² (though not heavily tested in IGCSE) can be mentioned, but the focus is on interpreting the slope and intercept in context.
为排放量建模时,IGCSE允许使用目测最佳拟合线,但你也可以计算最小二乘回归线。方程形式为 y = a + bx,其中 b = Σ(x − x̄)(y − ȳ) / Σ(x − x̄)²。例如,若年份(编码为1至12)与排放量呈正线性关系,则斜率 b 代表排放量的年均增长额。可决系数 R² 虽非IGCSE考查重点,但可以一提,关键在于结合情境解读斜率与截距。
6. Psychology: Experimental Design and Hypothesis Testing | 心理学:实验设计与假设检验
A psychologist wants to know whether a new memory technique improves recall scores. Two groups of 15 participants each are used: one control, one experimental. The mean recall score for the experimental group is 78 with standard deviation 9.2, and for the control group it is 70 with standard deviation 8.5. The question might ask you to carry out a two-sample t-test if equal variances can be assumed. However, at IGCSE level, you are more likely to calculate a standardised score (z-score) or to compare the difference in means relative to the pooled standard error.
一位心理学家想了解一种新的记忆技术是否提高回忆分数。实验采用两组各15名参与者:一个对照组,一个实验组。实验组平均回忆分数为78,标准差9.2;对照组平均为70,标准差8.5。题目可能要求你进行两样本t检验(若假设方差齐性)。不过在IGCSE阶段,更常见的是计算标准化分数(z分数)或比较均值差相对于合并标准误的大小。
Pooled estimate of common standard deviation: sₚ = √[((n₁−1)s₁² + (n₂−1)s₂²) / (n₁+n₂−2)] = √[((14×84.64)+(14×72.25))/28] = √[(1184.96+1011.5)/28] = √(2196.46/28) = √78.445 ≈ 8.86. Standard error of difference in means = sₚ × √(1/n₁ + 1/n₂) = 8.86 × √(1/15+1/15) = 8.86 × √(0.0667+0.0667) = 8.86 × √0.1333 ≈ 8.86 × 0.3651 ≈ 3.23. The observed difference 8, when divided by the SE, gives approximately 2.48. If this exceeds the critical value from t-tables (e.g., 2.048 for 28 d.f. at 5% two-tailed), the result is statistically significant. Psychology questions often also require you to state the null hypothesis and to comment on limitations like small sample size.
公共标准差的合并估计:sₚ = √[((n₁−1)s₁² + (n₂−1)s₂²) / (n₁+n₂−2)] = √[((14×84.64)+(14×72.25))/28] = √[(1184.96+1011.5)/28] = √(2196.46/28) = √78.445 ≈ 8.86。均值差的标准误 = sₚ × √(1/n₁ + 1/n₂) = 8.86 × √(0.1333) ≈ 3.23。观察差值8除以标准误得约2.48。若该值超过t分布临界值(例如28自由度、5%双尾临界值2.048),则结果具有统计显著性。心理学问题还常要求你陈述零假设,并对样本量小等局限性加以评论。
7. Business: Market Research and Sampling Methods | 商学:市场调研与抽样方法
A company wants to estimate the proportion of customers satisfied with its new product. It surveys a sample of 200 customers and finds that 150 are satisfied. The sample proportion is 0.75. You can construct an approximate 95% confidence interval for the true population proportion using the formula p ± z × √(p(1−p)/n), where z = 1.96 for 95% confidence. The standard error is √(0.75 × 0.25 / 200) = √(0.1875/200) = √0.0009375 ≈ 0.0306. The interval is 0.75 ± 1.96 × 0.0306, i.e., (0.690, 0.810). This means we are 95% confident that the true proportion of satisfied customers lies between 69.0% and 81.0%.
一家公司想估计对新品满意的顾客比例。它调查了200名顾客的样本,发现150人满意。样本比例为0.75。可用 p ± z × √(p(1−p)/n) 建立真实总体比例的近似95%置信区间,其中95%置信度的 z = 1.96。标准误为 √(0.75 × 0.25 / 200) = √0.0009375 ≈ 0.0306。区间为 0.75 ± 1.96 × 0.0306,即 (0.690, 0.810)。这意味着我们有95%的把握说真实满意顾客比例介于69.0%至81.0%之间。
IGCSE questions also examine different sampling techniques: simple random, stratified, systematic, quota and cluster sampling. A good business question might describe a scenario and ask which method is most appropriate and why. For instance, if the company has clearly defined customer segments (e.g., age groups), stratified sampling ensures each segment is proportionally represented, giving more precise estimates. The disadvantages include requiring a full list of the population and the complexity of drawing separate random samples from each stratum.
IGCSE题目还会考查不同的抽样技术:简单随机、分层、系统、配额和整群抽样。好的商科题目会描述一个情境,询问哪一方法最合适及其原因。例如,若公司有明确界定的顾客细分(如年龄组),分层抽样可确保各细分比例代表,给出更精确的估计。缺点则包括需要完整的总体名单以及从各层分别抽取随机样本的复杂性。
8. Sports Science: Performance Statistics and Normal Distribution | 运动科学:运动表现统计与正态分布
The times for 200 athletes running 800 m are normally distributed with a mean of 125 seconds and a standard deviation of 8 seconds. Using the empirical rule (68–95–99.7%), about 68% of athletes will have times between 117 s and 133 s (µ ± σ). If a coach wants to select the fastest 5% for advanced training, you need to find the cut-off time. The z-score corresponding to the top 5% is approximately 1.645 (since 5% in the upper tail corresponds to a cumulative probability of 0.95). The cut-off = µ − 1.645σ = 125 − 1.645 × 8 = 125 − 13.16 = 111.84 seconds. Athletes with times below 111.84 s qualify.
200名运动员跑800米的时间服从正态分布,均值为125秒,标准差为8秒。根据经验法则(68–95–99.7%),约68%的运动员成绩将在117秒至133秒之间(µ ± σ)。若教练想选拨前5%的运动员进行高阶训练,需要找到成绩分界线。与最高的5%对应的z分数约为1.645(因为上尾5%对应0.95累计概率)。分界值 = µ − 1.645σ = 125 − 1.645 × 8 = 125 − 13.16 = 111.84秒。成绩快于111.84秒的运动员入选。
Questions may also ask for the probability that a randomly chosen athlete runs between two given times. First convert the boundaries to z-scores. For example, P(120 < X < 130): z₁ = (120−125)/8 = −0.625, z₂ = (130−125)/8 = 0.625. Using standard normal tables, P(Z < 0.625) ≈ 0.7340, P(Z < −0.625) ≈ 0.2660, so the probability is 0.7340 − 0.2660 = 0.4680. This kind of calculation appears frequently in IGCSE Statistics, and it is essential to show all standardisation steps clearly.
题目还可能求随机选一名运动员成绩在给定两值之间的概率。先将界限转换成z分数。例如 P(120 < X < 130):z₁ = (120−125)/8 = −0.625,z₂ = (130−125)/8 = 0.625。查标准正态表,P(Z < 0.625) ≈ 0.7340,P(Z < −0.625) ≈ 0.2660,故概率为0.7340 − 0.2660 = 0.4680。此类计算在IGCSE统计中频繁出现,必须清晰展现所有标准化步骤。
9. Sociology: Chi-Squared Test for Independence | 社会学:独立性的卡方检验
A sociologist investigates whether there is an association between gender and preference for a new community programme. Data are summarised in a 2×2 contingency table. The null hypothesis states that gender and preference are independent. Expected frequencies are calculated by row total × column total / grand total. The chi-squared statistic is computed as χ² = Σ (O − E)² / E. With 1 degree of freedom ( (2−1)×(2−1) ), the critical value at the 5% significance level is 3.841.
一位社会学家研究性别与对新社区项目偏好之间是否有关联。数据汇总成2×2列联表。零假设为性别与偏好相互独立。期望频数按(行合计 × 列合计)/ 总计计算。卡方统计量 χ² = Σ (O − E)² / E。自由度为1((2−1)×(2−1)),5%显著性水平下临界值为3.841。
| Prefer | Not Prefer | Total | |
|---|---|---|---|
| Male | 40 | 20 | 60 |
| Female | 30 | 30 | 60 |
| Total | 70 | 50 | 120 |
Expected for Male/Prefer = 60×70/120 = 35. Male/Not Prefer = 60×50/120 = 25. Female/Prefer = 60×70/120 = 35. Female/Not Prefer = 25. χ² = (40−35)²/35 + (20−25)²/25 + (30−35)²/35 + (30−25)²/25 = 25/35 + 25/25 + 25/35 + 25/25 ≈ 0.714 + 1.0 + 0.714 + 1.0 = 3.428. Since 3.428 < 3.841, we do not reject the null hypothesis – there is insufficient evidence of an association at the 5% level. Always relate the conclusion back to the context.
男/偏好的期望频数 = 60×70/120 = 35;男/不偏好 = 25;女/偏好 = 35;女/不偏好 = 25。χ² = (40−35)²/35 + (20−25)²/25 + (30−35)²/35 + (30−25)²/25 = 0.714+1.0+0.714+1.0 = 3.428。由于3.428 < 3.841,我们不拒绝零假设——在5%水平下没有足够证据表明存在关联。请务必将结论联系回情境。
10. Mixed Interdisciplinary Problem Set | 跨学科综合问题套题
IGCSE Statistics papers often end with a multi-part question that weaves together several disciplines. For example, a scenario might involve a city’s quality of life index, which is constructed from environmental, economic and health indicators. You might be asked to: calculate a composite index using weighted averages (economics), draw a comparative bar chart of air quality indices across districts (geography), calculate the median and range of life expectancy from grouped data (health science), and finally test whether the observed distribution of satisfaction ratings fits a predicted distribution using a chi-squared goodness-of-fit test. This type of question rewards systematic working and clear interpretation.
IGCSE统计试卷常以一个多部分的综合题收尾,将多个学科交织在一起。例如,场景可能涉及某城市的生活质量指数,该指数由环境、经济和健康指标复合而成。题目可能要求:用加权平均计算综合指数(经济学),绘制各区空气质量指数的对比条形图(地理学),根据分组数据计算预期寿命的中位数与极差(健康科学),最后用卡方拟合优度检验检验满意度评级的观测分布是否符合预测分布。这类题目会奖赏系统性的解题步骤和清晰的解读。
A strong response will show every step: coding the original variables, listing the weights, constructing composite scores, plotting accurately, interpolating for grouped medians, and computing expected frequencies for the goodness-of-fit test. For the chi-squared goodness-of-fit test, remember that degrees of freedom = number of categories − 1, and that expected frequencies must all be at least 5 for the test to be valid. If any expected frequency is below 5, you should combine categories. Finally, summarise the findings in plain language suitable for a non-statistical audience, a skill that bridges statistics with communication.
一份优秀的答案会展示每一步骤:对原始变量编码、列出权重、构建综合得分、精确绘图、用分组内插法求中位数,以及为拟合优度检验计算期望频数。对于卡方拟合优度检验,记住自由度 = 类别数 − 1,且为确保检验有效,所有期望频数须至少为5。若有期望频数低于5,应合并类别。最后,用适合非统计专业受众的平实语言总结发现,这是连接统计与沟通的桥梁技能。
Published by TutorHao | Statistics Revision Series | aleveler.com
Find Edexcel IGCSE Statistics Textbooks on eBay UK
New, used and second-hand copies of textbooks and revision guides are often much cheaper than retail — check current listings and prices before you buy.
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导