📚 Year 10 OCR Statistics: Interdisciplinary Mixed Question Practice | Year 10 OCR 统计:跨学科综合题型训练
Statistics is not a standalone subject—it permeates biology, geography, business, sports science, and everyday decision-making. The OCR Year 10 Statistics course expects you to apply your knowledge of data collection, representation, averages, probability, and correlation to realistic, cross-curricular contexts. This article provides targeted mixed-question practice drawn from diverse fields, with step-by-step reasoning to help you build fluency and confidence for your exams.
统计学并非一门孤立的学科——它渗透在生物学、地理学、商业、体育科学以及日常决策中。OCR 10 年级统计课程要求你将数据收集、数据表示、平均数、概率和相关等知识应用于真实的跨学科情境。本文提供了来自不同领域的针对性综合题型训练,并配有分步推理,以帮助你提升解题流畅度与应考信心。
1. Data Collection & Sampling in Geography Fieldwork | 地理实地考察中的数据收集与抽样
A group of Year 10 students investigates pebble size along a beach transect. They decide to sample every 5th pebble from the water’s edge to the cliff. This is an example of systematic sampling, which ensures an even spread across the transect. The advantage is simplicity; however, if there is a hidden pattern (e.g., every 5th pebble happens to be in a tide line of larger stones), the sample may be biased.
一群 10 年级学生调查了海滩断面上的卵石大小。他们决定从水边到悬崖每 5 个卵石取样一个。这是系统抽样的一个例子,能确保在断面上均匀分布。其优点是简单;但如果存在隐藏模式(例如每第 5 个卵石恰好位于大石块较多的潮线),样本可能会产生偏差。
In contrast, a random sample would involve numbering all pebbles visible in a 1 m² quadrat and using a random number generator to select 20. Random sampling removes selection bias but is more time-consuming to organise. OCR exam questions often ask you to evaluate a sampling method, so always mention practicality, bias, and representativeness.
相比之下,随机抽样需要给 1 m² 样方内所有可见卵石编号,然后用随机数生成器选取 20 个。随机抽样能消除选择偏差,但组织起来更耗时。OCR 考题常要求你评价一种抽样方法,因此务必提及实用性、偏差和代表性。
When collecting data in the field, you also need to decide on sample size. A larger sample (e.g., 100 pebbles) reduces variability in the estimate of the mean diameter, but resources may limit this. The concept of capture-recapture (Petersen method) might appear for estimating animal populations in biology contexts, using the formula N = (M × C) ÷ R, where M is the number first captured and marked, C is the total captured in the second sample, and R is the number of marked recaptures.
在实地收集数据时,你还需要确定样本量。较大的样本(例如 100 块卵石)可以减少均值直径估计的变异,但资源可能限制这一点。捕获-再捕获(Petersen 方法)的概念可能会出现在生物学情境中,用于估算动物种群数量,使用的公式是 N = (M × C) ÷ R,其中 M 是首次捕获并标记的数量,C 是第二次捕获的总数,R 是带有标记的再捕获数。
2. Graphs and Visualisations in Science Experiments | 科学实验中的图表与可视化
In a physics experiment, a student measures the extension of a spring as mass is added. The data are recorded in a table, and the relationship between force (F) and extension (x) is expected to follow Hooke’s Law: F ∝ x. A scatter graph with a line of best fit helps visualise this. Outliers—points far from the trend—must be checked; if due to a measurement error, they may be excluded from the line of best fit.
在一次物理实验中,一名学生测量了弹簧在增加质量时的伸长量。数据记录在表格中,力 (F) 与伸长量 (x) 之间的关系预期遵循胡克定律:F ∝ x。带最佳拟合线的散点图有助于可视化这一关系。远离趋势的异常点必须检查;如果是由测量误差造成,则可在绘制最佳拟合线时将其排除。
When constructing cumulative frequency graphs in a chemistry context (e.g., reaction times of 50 trials), you need to plot the upper class boundary against the cumulative frequency. The interquartile range (IQR = Q₃ − Q₁) can then be read from the graph. A box plot (box-and-whisker) provides a five-number summary and is excellent for comparing distributions of results from two different catalysts.
在化学情境中(例如 50 次试验的反应时间)绘制累积频率图时,需要将上组界作为横坐标,累积频率作为纵坐标。然后可以从图中读取四分位距 (IQR = Q₃ − Q₁)。箱形图(箱线图)提供了五数概括,非常适用于比较两种不同催化剂的反应时间分布。
Histograms with unequal class widths frequently appear in biology (e.g., heart rates of students after exercise). Remember: frequency density = frequency ÷ class width. The area of each bar is proportional to frequency. In OCR questions, you may be given a partially completed histogram and asked to fill in missing bars or frequencies.
不等组距的直方图经常出现在生物学中(例如运动后学生的心率)。记住:频率密度 = 频率 ÷ 组距。每个柱形的面积与频率成正比。在 OCR 考题中,你可能会遇到部分完成的直方图,并被要求补全缺失的柱形或频率。
3. Averages and Spread in Business Sales | 商业销售中的平均数与离散程度
A small shop records the daily number of customers over 30 days. To summarise the data, the mean gives the overall average footfall, but it is sensitive to extreme values like a bank holiday surge. The median (the 15.5th value when ordered) is more robust. The mode identifies the most frequent customer count, useful for staffing decisions.
一家小商店记录了 30 天内每日顾客人数。为了汇总数据,均值给出了整体平均客流量,但它对极值(如银行假日激增)敏感。中位数(排序后的第 15.5 个值)更具稳健性。众数则指出了最常出现的顾客人数,有助于排班决策。
To measure consistency, the range (max − min) is quick but crude. The interquartile range (IQR) covers the middle 50% of data and is unaffected by outliers. The standard deviation, required for higher-tier papers, quantifies the average distance from the mean. For the customer data set {12, 15, 18, 20, 22, 25, 30}, the mean x̄ = 20.29, deviations squared and averaged give a variance, and the square root gives the standard deviation s ≈ 5.79.
为了衡量一致性,极差(最大值 − 最小值)虽然快捷但比较粗糙。四分位距(IQR)包含了中间 50% 的数据,且不受异常值影响。高阶试卷要求的标准差则量化了到均值的平均距离。对于顾客数据集 {12, 15, 18, 20, 22, 25, 30},均值 x̄ = 20.29,求各偏差平方的平均数得到方差,再开方得到标准差 s ≈ 5.79。
When comparing two branches (e.g., Branch A: median 45, IQR 10; Branch B: median 48, IQR 18), you should comment on both central tendency (Branch B has a slightly higher median footfall) and spread (Branch A’s customer numbers are more consistent, having a smaller IQR). A dual box plot makes this comparison instant.
比较两个分店时(例如 A 分店:中位数 45,IQR 10;B 分店:中位数 48,IQR 18),你应同时评论集中趋势(B 分店的中位客流量略高)和离散程度(A 分店的顾客数量更一致,因为 IQR 更小)。双重箱形图可以瞬间呈现这种比较。
4. Probability and Risk in Health Contexts | 健康情境中的概率与风险
A medical screening test for a condition has a sensitivity (true positive rate) of 95% and a specificity (true negative rate) of 90%. The prevalence of the condition in the population is 2%. Using a tree diagram or a two-way table, you can calculate the probability that a person who tests positive actually has the condition (positive predictive value). This is a classic Bayes’ theorem scenario, but at Year 10 level, a carefully constructed contingency table with 10 000 hypothetical people works well.
某种疾病的医学筛查检测灵敏度(真阳性率)为 95%,特异度(真阴性率)为 90%,而该疾病在人群中的患病率为 2%。利用树形图或双向表,你可以计算检测呈阳性的人实际患病的概率(阳性预测值)。这是一个经典的贝叶斯定理场景,但在 10 年级水平,用一个精心构造的 10 000 人假设列联表就能很好地解决。
Construct: Population 10 000. With disease: 2% of 10 000 = 200; without disease: 9 800. Among the 200 with disease, 95% test positive → 190 true positives, 10 false negatives. Among the 9 800 without disease, 90% test negative → 8 820 true negatives, 980 false positives. Total positive tests = 190 + 980 = 1 170. The probability of having the disease given a positive test = 190 / 1 170 ≈ 16.2%. This low probability often surprises students and highlights why screening programmes must be carefully designed.
构造假设数据:人口 10 000 人。患病人数:10 000 的 2% = 200 人;未患病人数:9 800 人。在 200 名患者中,95% 检测呈阳性 → 190 名真阳性,10 名假阴性。在 9 800 名非患者中,90% 检测呈阴性 → 8 820 名真阴性,980 名假阳性。检测阳性总数 = 190 + 980 = 1 170。已知检测阳性下患病的概率 = 190 / 1 170 ≈ 16.2%。这一低概率常令学生惊讶,也凸显了为何筛查方案必须精心设计。
In a genetics context, you might construct Punnett squares to show the probabilities of inheriting a recessive disorder. For two carrier parents (Bb × Bb), the offspring genotype probabilities are BB 0.25, Bb 0.5, bb 0.25. The probability of a child having the disorder is 0.25. The concepts of mutually exclusive and independent events are essential: having a boy and a child with the disorder are independent, so multiply probabilities.
在遗传学情境中,你可以构建旁氏表来展示遗传隐性疾病的概率。对于两位携带者父母 (Bb × Bb),子代基因型概率为 BB 0.25、Bb 0.5、bb 0.25。孩子患病的概率为 0.25。互斥事件与独立事件的概念至关重要:生男孩与孩子患病是相互独立的,因此需要将概率相乘。
5. Scatter Graphs and Correlation in Sports Science | 体育科学中的散点图与相关性
A sports scientist records the hours of sleep (x) and reaction time in milliseconds (y) for 12 athletes. Plotting the data shows a downward trend: as sleep increases, reaction time decreases. This suggests a negative correlation. To measure the strength, you can draw a line of best fit by eye (for foundation tier) or calculate Spearman’s rank correlation coefficient rₛ (for higher tier). The formula rₛ = 1 − (6 Σ d²) / (n(n² − 1)) requires ranking both variables and computing the differences d in ranks.
一位体育科学家记录了 12 名运动员的睡眠时长 (x) 与反应时间(毫秒)(y)。绘制数据后发现一种下降趋势:睡眠越长,反应时间越短。这表明存在负相关。为了衡量相关程度,你可以凭目测画出最佳拟合线(基础卷),或计算斯皮尔曼等级相关系数 rₛ(高阶卷)。公式 rₛ = 1 − (6 Σ d²) / (n(n² − 1)) 要求对两个变量分别排序,并计算排序差 d。
Interpreting a correlation is not the same as establishing causation. The improvement in reaction time might be due to a third variable, such as overall recovery quality or training regimen. In an exam, you will be asked to ‘describe the relationship’ (strength, direction, form, any outliers) and ‘interpret’ it in context. For a strong positive correlation, you might say: ‘Athletes who slept more tended to have faster reaction times, but we cannot be certain that extra sleep causes the improvement.’
解释相关性并不等同于确立因果关系。反应时间的改善可能源于第三个变量,例如整体恢复质量或训练方案。在考试中,你会被要求“描述关系”(强度、方向、形式、是否有异常点),并结合上下文进行“解释”。对于强正相关,你可以这样表述:“睡眠时间越长的运动员,反应时间往往越快,但我们无法确定额外睡眠导致了改善。”
Regression lines can be used to make predictions, but only within the range of the data (interpolation). Extrapolating beyond the range (e.g., predicting reaction time for 0 hours sleep) is unreliable. The equation of the line of best fit is typically in the form y = a + bx, where b is the gradient and a is the y-intercept. Once you have the equation, you can estimate a value of y for a given x.
回归线可用于预测,但仅限于数据范围内(内插)。外推到范围之外(例如预测睡眠 0 小时的反应时间)是不可靠的。最佳拟合线的方程通常为 y = a + bx,其中 b 是斜率,a 是 y 轴截距。一旦得到方程,你就能对给定的 x 估算 y 值。
6. Time Series and Forecasting in Economics | 经济学中的时间序列与预测
An economics student is given quarterly data on the number of tourists visiting a town. The time series plot shows an upward trend over the years and regular seasonal variation (peaks in summer, troughs in winter). Moving averages smooth out the seasonal pattern. For quarterly data, a 4-point moving average is calculated, and then a centered 4-point moving average is used to align with seasons.
一位经济学学生获得了某城镇每季度游客数量的数据。时间序列图显示出多年来的上升趋势和规则的季节性波动(夏季高峰,冬季低谷)。移动平均线可以平滑季节性模式。对于季度数据,需要计算 4 点移动平均,然后使用中心化的 4 点移动平均与季节对齐。
Calculating seasonal variation: subtract the centered moving average (trend) from the actual value. For instance, if Summer actual = 520, centred trend = 480, the seasonal variation is +40. Averaging the seasonal variations for each quarter over several years gives the average seasonal effect. These are used to adjust predictions and remove seasonality from data (seasonally adjusted data = actual − seasonal effect).
计算季节性波动:用实际值减去中心化移动平均值(趋势)。例如,若夏季实际值为 520,中心化趋势为 480,则季节性波动为 +40。对若干年中每个相同季度的季节性波动取平均,即可得到平均季节效应。这些值可用于调整预测,从数据中剔除季节性(季节调整数据 = 实际值 − 季节效应)。
Forecasting future values involves extending the trend line (perhaps by linear regression on the trend values) and then adding the average seasonal effect. OCR questions may ask you to plot the time series, draw the trend line, and predict the tourist numbers for the next summer, assuming the pattern continues. Comment on the reliability: ‘Prediction assumes past trends and seasonal patterns remain unchanged, which may not hold due to unexpected events like a pandemic or economic crisis.’
预测未来值需要延长趋势线(也许通过对趋势值进行线性回归),然后加上平均季节效应。OCR 考题可能会要求你画出时间序列图,绘制趋势线,并在假设模式继续的前提下预测下一个夏季的游客数量。需评论可靠性:“预测假设过去的趋势和季节模式保持不变,但这可能不成立,因为可能发生如疫情或经济危机等突发事件。”
7. Index Numbers and Rates in Environmental Science | 环境科学中的指数与比率
Index numbers simplify comparisons over time by expressing values relative to a base period, usually set to 100. If carbon dioxide concentration in the atmosphere was 315 ppm in 1960 (base year) and 420 ppm in 2023, the simple index for 2023 is (420 ÷ 315) × 100 ≈ 133.3. This means CO₂ concentration is 33.3% higher than the base year. A weighted index might be used for air quality, combining several pollutants with different harm factors.
指数通过将数值表示为相对于基期(通常设为 100)的值,简化了时间上的比较。如果 1960 年大气中二氧化碳浓度为 315 ppm(基年),而 2023 年为 420 ppm,则 2023 年的简单指数为 (420 ÷ 315) × 100 ≈ 133.3。这意味着 CO₂ 浓度比基年高出 33.3%。在空气质量的评估中,加权指数可能结合多种污染物并赋予不同的危害因子。
Rates such as birth rate per 1 000 population or reaction rate in chemistry (cm³/s) are fundamental to statistics. The concept of standardised rates appears when comparing mortality rates across countries with different age structures. The directly standardised rate applies age-specific death rates to a standard population. Although not heavily tested at Year 10, recognising that crude rates can be misleading is an important skill.
诸如每千人出生率或化学中的反应速率(cm³/s)等比率是统计学的基础。当比较不同年龄结构的国家之间的死亡率时,会用到标准化比率的概念。直接标准化比率是将年龄别死亡率应用到一个标准人口上。尽管在 10 年级不作重点考查,但认识到粗率可能产生误导是一项重要的能力。
When working with rates, always check the units and the denominator. A common mistake is to interpret ’25 per 1000′ as 25% (it is actually 2.5%). In exam contexts, you might calculate the rate of deforestation for a country given forest area lost over time, or the rate of population change factoring in births, deaths, and migration.
在处理比率时,务必检查单位和分母。一个常见错误是将“千分之 25”理解为 25%(实际上为 2.5%)。在考试情境中,你可能会根据一段时间内森林面积损失的数据计算一个国家的森林砍伐率,或者结合出生、死亡和迁移因素计算人口变化率。
8. Probability Distributions in Genetics and Games | 遗传学与游戏中的概率分布
The binomial distribution arises when you have a fixed number of independent trials, each with two outcomes (success/failure) and constant probability p. For example, pea plants in a genetics experiment have a probability p = 0.75 of producing yellow peas. In a sample of 8 offspring, the probability of exactly 6 yellow peas is P(X=6) = ⁸C₆ × (0.75)⁶ × (0.25)². Calculating ⁸C₆ = 28, gives P(6) ≈ 0.311 or 31.1%. OCR may ask you to complete a probability table or calculate expected frequencies.
二项分布适用于固定次数的独立试验,每次试验只有两种结果(成功/失败),且概率 p 恒定。例如,在遗传学实验中,豌豆植株产生黄色豌豆的概率 p = 0.75。在 8 株子代样本中,恰好获得 6 粒黄色豌豆的概率为 P(X=6) = ⁸C₆ × (0.75)⁶ × (0.25)²。计算 ⁸C₆ = 28,得出 P(6) ≈ 0.311,即 31.1%。OCR 可能会要求你完成概率表格或计算期望频率。
The normal distribution is another key model for continuous data like heights, weights, or measurement errors. The empirical rule states that approximately 68% of data lie within 1 standard deviation of the mean, 95% within 2, and 99.7% within 3, provided the distribution is bell-shaped and symmetric. In a quality control context (e.g., filling cereal boxes), you might estimate what percentage of boxes are underweight given μ = 500 g and σ = 5 g. A box below 495 g is 1σ below the mean; about 16% of boxes would be below that threshold (since 50% − 34% ≈ 16%).
正态分布是另一类重要模型,适用于连续数据,如身高、体重或测量误差。经验法则指出,如果分布呈钟形且对称,则大约 68% 的数据落在均值两侧 1 个标准差内,95% 落在 2 个标准差内,99.7% 落在 3 个标准差内。在质量控制情境中(例如灌装早餐谷物盒),给定 μ = 500 g 和 σ = 5 g,你可以估算出多少比例的盒子重量不足。一个低于 495 g 的盒子比均值低 1σ;大约 16% 的盒子会低于该阈值(因为 50% − 34% ≈ 16%)。
In board games, expected value helps make decisions. If a spinner has equal sectors giving payouts £0, £2, £5, £10, the expected win per spin is (0+2+5+10)/4 = £4.25. If the cost to play is £3, the expected profit is £1.25 per spin, so it is favourable. However, variance matters: the risk of losing money in the short run is high.
在棋盘游戏中,期望值有助于决策。如果一个转盘各扇区奖金分别为 £0、£2、£5、£10,则每次旋转的期望收益为 (0+2+5+10)/4 = £4.25。如果每次游戏成本为 £3,则期望利润为 £1.25,因此是有利的。不过,方差也很重要:短期内亏损的风险仍然很高。
9. Mixed Question Practice: Step-by-Step Worked Examples | 混合题型练习:分步例题
Let’s consolidate with a full cross-disciplinary question. A psychology study investigates whether students who eat breakfast perform better in memory tests. Twelve students are tested, recording breakfast consumption (yes/no) and number of words recalled. The data: With breakfast: 8, 10, 9, 11, 12, 10; Without breakfast: 6, 7, 5, 8, 6, 9.
让我们通过一道完整的跨学科问题巩固所学。一项心理学研究调查吃早餐的学生在记忆测试中表现是否更好。12 名学生参与测试,记录早餐摄入情况(是/否)与回忆单词数。数据如下:吃早餐组:8, 10, 9, 11, 12, 10;不吃早餐组:6, 7, 5, 8, 6, 9。
Part (a): Calculate the median and IQR for each group, and draw comparative box plots. For breakfast group, ordered: 8,9,10,10,11,12. Median = (10+10)/2 = 10. Q1 = 9, Q3 = 11, IQR = 2. No breakfast group: 5,6,6,7,8,9. Median = (6+7)/2 = 6.5. Q1 = 6, Q3 = 8, IQR = 2. The box plots show the breakfast group’s median noticeably higher, but spreads similar.
第 (a) 题:计算每组的中位数和四分位距,并绘制比较箱形图。吃早餐组排序后为:8,9,10,10,11,12。中位数 = (10+10)/2 = 10。Q1 = 9,Q3 = 11,IQR = 2。不吃早餐组:5,6,6,7,8,9。中位数 = (6+7)/2 = 6.5。Q1 = 6,Q3 = 8,IQR = 2。箱形图显示吃早餐组的中位数明显更高,但离散度相似。
Part (b): Calculate the probability that a randomly chosen student from this study who recalled more than 8 words had eaten breakfast. Total students with >8 words: Breakfast 10,9,11,12,10 → 5 students; No breakfast 9 → 1 student. So P(Breakfast | >8 words) = 5/6 ≈ 0.833. This is a conditional probability question embedded in a data set.
第 (b) 题:计算从本研究随机选出一名回忆单词超过 8 个的学生确系吃过早餐的概率。超过 8 个单词者中:吃早餐组有 10,9,11,12,10 共 5 名;不吃早餐组仅 9 共 1 名。因此 P(吃早餐 | >8 个单词) = 5/6 ≈ 0.833。这是一个嵌入数据集的条�概率问题。
Part (c): Discuss whether the data support the hypothesis. We observe a difference in medians, but the sample size is small (only 12) and the students were not randomly selected—they might be from the same class. A confounding variable: students who eat breakfast may also have more consistent sleep patterns. Therefore, we cannot conclude causation.
第 (c) 题:讨论数据是否支持假设。我们观察到中位数差异,但样本量很小(仅 12 人),且学生并非随机选取——他们可能来自同一班级。一个混杂变量:吃早餐的学生也可能拥有更规律的睡眠模式。因此,我们不能得出因果结论。
10. Exam Tips for Cross-Disciplinary Questions | 跨学科题的应试技巧
Always read the contextual setup carefully; underline key variables and the type of data (discrete, continuous, categorical). Identify the statistical technique required: Is it a comparison of averages? A probability calculation? A graph interpretation? Many marks are lost by applying the right tool to the wrong question.
务必仔细阅读情境设置;在关键变量和数据类型(离散、连续、分类)下划线。识别所需的统计方法:是比较平均数?概率计算?还是图表解读?许多失分源于把正确的工具用错了题目。
When drawing conclusions, link back to the context. Instead of writing ‘The median of Group A is higher’, write ‘Students who had breakfast typically recalled 3.5 more words on average (median) than those who skipped breakfast, suggesting a possible link between breakfast and short-term memory in this sample.’ This contextualised interpretation earns high marks.
得出结论时,务必回扣情境。不要只写“A 组的中位数更高”,而应写“在这个样本中,吃早餐的学生比不吃早餐的学生平均多回忆 3.5 个单词(中位数),提示早餐与短期记忆之间可能存在联系”。这种结合情境的解释能拿到高分。
Show all steps clearly, especially in probability and moving average calculations. Even if your final answer is wrong, method marks can be awarded for correctly identifying ranked positions, using the formula, or setting up a two-way table. Manage your time: a 5-mark question deserves about 5 minutes of thinking and writing.
清晰展示所有步骤,尤其是在概率和移动平均计算中。即使最终答案错误,只要正确识别了排序位置、使用了公式或构建了双向表,仍可获得方法分。合理安排时间:一道 5 分的题目大约需要 5 分钟的思考与书写时间。
Finally, practice with past OCR papers that contain scenarios from ecology, psychology, sports, finance, and health. The more you encounter interdisciplinary contexts, the quicker you will translate a real-world situation into a statistical framework.
最后,用包含生态学、心理学、体育、金融和健康情境的往年 OCR 真题进行练习。你接触的跨学科情境越多,就能越迅速地将真实世界场景转化为统计框架。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导