📚 Interdisciplinary Applied Statistics Practice | 跨学科综合题型训练
Year 10 Eduqas Statistics challenges you to apply statistical methods across real-world contexts such as science, geography, and social studies. This article combines bilingual explanations with worked examples to strengthen your ability to interpret data, design investigations, and critique bias in unfamiliar interdisciplinary scenarios. Each section pairs English paragraphs with Chinese equivalents so you can master terminology and reasoning simultaneously.
Year 10 Eduqas 统计课程要求你在科学、地理和社会研究等真实情境中运用统计方法。本文结合双语讲解与实例练习,帮助你在陌生的跨学科场景中提升数据解读、调查设计和偏差批判的能力。每个部分都配有中英文段落,让你同步掌握术语和推理逻辑。
1. Scientific Investigations and Measurement Error | 科学调查与测量误差
In science experiments, every measurement carries an inherent uncertainty. When you record the temperature change of a chemical reaction, the thermometer’s precision (±0.5 °C) and your reaction time create random errors. You must identify systematic errors, such as a poorly calibrated balance that consistently reads 0.2 g too high, because these affect the accuracy of your calculated mean.
在科学实验中,每次测量都存在固有不确定性。当你记录化学反应温度变化时,温度计的精度(±0.5 °C)和你的反应时间会产生随机误差。你必须识别系统误差,例如一台未校准的天平始终偏高 0.2 g,因为这些误差会影响所计算均值的准度。
- Random error reduces when you repeat measurements and calculate the mean.
- 随机误差可通过重复测量并计算均值来减小。
- Systematic error persists even with large sample sizes; it must be corrected by calibrating equipment or applying a correction factor.
- 系统误差即使样本量大也会持续存在;必须通过校准仪器或应用修正系数来纠正。
Example: A student measures the mass of five iron nails using the same balance and obtains: 1.8 g, 1.9 g, 2.0 g, 1.8 g, 1.9 g. However, the balance consistently reads 0.3 g too low. The true mean mass is (1.88 + 0.3) = 2.18 g.
示例:一名学生用同一架天平测量五根铁钉的质量,得到:1.8 g、1.9 g、2.0 g、1.8 g、1.9 g。然而,天平始终偏低 0.3 g。真实的平均质量为(1.88 + 0.3)= 2.18 g。
2. Rates and Proportions in Biological Studies | 生物研究中的比率与比例
Ecology field studies often ask you to estimate population proportions. If you tag 40 woodlice in a log pile, release them, and later recapture 60 woodlice of which 10 are tagged, the Lincoln-Petersen estimator gives N = (40 × 60) ÷ 10 = 240 as the estimated total population. This method relies on the assumption that tagged and untagged individuals mix randomly.
生态学野外研究常要求你估算种群比例。如果你在木桩堆中标记了 40 只鼠妇,放回后重新捕获 60 只,其中 10 只带有标记,Lincoln-Petersen 估计公式给出 N = (40 × 60) ÷ 10 = 240 作为种群总数估计值。此方法依赖于标记与未标记个体随机混合的假设。
You also encounter stratified sampling when a habitat has distinct zones. If a woodland contains 60% deciduous area and 40% coniferous, you might sample 12 quadrats from the deciduous zone and 8 from the coniferous zone to maintain proportional representation. The overall biodiversity index then becomes a weighted mean of zone-specific indices.
当栖息地包含不同区域时,你还会遇到分层抽样。如果一片林地包含 60% 的落叶林区和 40% 的针叶林区,你可能需要从落叶区抽取 12 个样方、从针叶区抽取 8 个样方以保持比例代表性。整体的生物多样性指数则成为各区特定指数的加权平均值。
3. Weather Data and Moving Averages in Geography | 地理学中的天气数据与移动平均
Geography fieldwork frequently uses time series data like monthly rainfall. A 3-point moving average smooths out erratic fluctuations to reveal seasonal trends. For monthly rainfall values (mm): Jan 78, Feb 56, Mar 65, Apr 92, the first moving average (Feb) = (78+56+65)/3 = 66.3 mm. This technique helps you spot long-term climate patterns rather than getting distracted by anomalous storms.
地理实地考察经常使用时间序列数据,如月降雨量。三点移动平均可以平滑不规则波动,揭示季节性趋势。对于月降雨量(毫米):一月 78、二月 56、三月 65、四月 92,第一个移动平均值(二月)= (78+56+65)/3 = 66.3 毫米。这一技巧能帮助你发现长期气候模式,而不被异常暴风雨所干扰。
When comparing two weather stations, you might calculate the percentage change in annual rainfall. If Station A recorded 820 mm last year and 943 mm this year, the percentage increase = ((943 − 820) ÷ 820) × 100 = 15.0%. Always check whether the denominator represents the original or the final value; using the wrong base is a common mistake.
当比较两个气象站时,你可能要计算年降雨量的百分比变化。如果 A 站去年记录为 820 毫米,今年为 943 毫米,则增长百分比 = ((943 − 820) ÷ 820) × 100 = 15.0%。务必检查分母代表初始值还是最终值;用错基数是常见错误。
4. Social Surveys and Questionnaire Design | 社会调查与问卷设计
When you design a questionnaire on a social issue like screen time among teenagers, every question must be unbiased. A leading question such as “Do you agree that excessive gaming damages your grades?” assumes the negative effect. Instead, ask: “How many hours per week do you spend gaming?” followed by “What is your average test score?” so you can later analyze correlation objectively.
当你针对青少年屏幕时间等社会问题设计问卷时,每个问题都必须无偏。像“你是否同意过度游戏会损害成绩?”这样的引导性问题假设了负面效应。正确的做法是问:“你每周花多少小时玩游戏?”接着问“你的平均考试分数是多少?”,这样你之后可以客观地分析相关性。
You need to consider sample representativeness. If you only survey students in a school computer club, your findings cannot be generalized to all teenagers. A stratified sample by year group and gender ensures each subgroup’s voice is included proportionally, making the median screen time more reliable.
你需要考虑样本代表性。如果你只调查学校计算机俱乐部的学生,你的结论不能推广到所有青少年。按年级和性别进行分层抽样能确保各个子群体的声音得到比例均衡的反映,使中位屏幕时间更可靠。
5. Combining Probability with Health Statistics | 概率与健康统计的结合
Health campaigns often quote lifetime risk. If a disease affects 1 in 400 people in a population, the probability that a randomly selected person has the disease is 0.0025. More complex is conditional probability: given a positive test result with 95% sensitivity and 2% false positive rate, the actual probability of having the disease depends on the base rate (prevalence). This is a classic Bayes’ theorem application, which you might explore through tree diagrams.
健康宣传常引用终生风险。如果某种疾病在人群中影响 1/400 的人,则随机挑选一人患病的概率是 0.0025。更复杂的是条件概率:若某一检测灵敏度为 95%,假阳性率为 2%,则得到阳性结果时真正患病的概率取决于基础比率(患病率)。这是经典的贝叶斯定理应用,你可以通过树状图来探究。
You can construct a two-way table for a hypothetical screening of 10,000 people. With prevalence 0.0025, you expect about 25 sick individuals. The test will identify roughly 24 of them (95% of 25) and also falsely flag about 199 healthy people as positive (2% of 9,975). Thus, a positive result yields a probability of truly being sick ≈ 24 ÷ (24+199) ≈ 10.8%. This counterintuitive result highlights why mass screening requires careful statistical reasoning.
你可以为假设的 10,000 人筛查构建双向表。患病率为 0.0025 时,大约有 25 名患者。该检测将识别其中约 24 人(25 的 95%),同时还将约 199 名健康人错误地标记为阳性(9,975 的 2%)。因此,阳性结果真正患病的概率 ≈ 24 ÷ (24+199) ≈ 10.8%。这一反直觉的结果凸显了为什么大规模筛查需要谨慎的统计推理。
6. Financial Contexts: Index Numbers and Inflation | 金融情境:指数与通货膨胀
In economics and business studies, index numbers simplify the comparison of price changes over time. If the price of a basket of goods is £250 in the base year (index = 100) and rises to £270 the next year, the new index value = (270 ÷ 250) × 100 = 108. The inflation rate is 8%. Weighted index numbers become essential when items have different importance — housing costs usually receive a higher weight than entertainment.
在经济学和商业研究中,指数简化了不同时期价格变化的比较。如果一篮子商品在基年的价格为 250 英镑(指数 = 100),次年升至 270 英镑,则新指数值 = (270 ÷ 250) × 100 = 108。通货膨胀率为 8%。当各项目重要性不同时,加权指数变得至关重要——住房成本通常被赋予比娱乐更高的权重。
You might be asked to calculate a weighted retail price index. Suppose food (weight 40) rises 5%, clothing (weight 30) rises 2%, and transport (weight 30) rises 10%. The overall percentage change = (40×5 + 30×2 + 30×10) ÷ (40+30+30) = (200+60+300) ÷ 100 = 5.6%. This weighted average more accurately reflects the average consumer’s experience than a simple mean of 5.67%.
你可能需要计算加权零售价格指数。假设食品(权重 40)上涨 5%,服装(权重 30)上涨 2%,交通(权重 30)上涨 10%。总体百分比变化 = (40×5 + 30×2 + 30×10) ÷ (40+30+30) = (200+60+300) ÷ 100 = 5.6%。这一加权平均值比简单平均 5.67% 更准确地反映了普通消费者的体验。
7. Bivariate Data in Sports and Engineering | 体育与工程中的双变量数据
Scatter graphs and correlation coefficients appear across disciplines. A sports scientist might plot weekly training hours (x) against sprint time (y) for 12 athletes. If the Pearson correlation coefficient r = −0.82, this indicates a strong negative correlation: more training is associated with faster (shorter) sprint times. However, correlation does not imply causation — other factors like diet and rest also matter.
散点图和相关系数出现在各个学科中。运动科学家可能会绘制 12 名运动员的每周训练时间(x)与短跑时间(y)的关系图。如果皮尔逊相关系数 r = −0.82,则表明存在强负相关:训练越多,短跑时间越短。然而,相关关系并不意味因果关系——饮食和休息等其他因素也很重要。
In engineering, you might examine the relationship between the load placed on a spring (kg) and its extension (cm). A strong positive linear correlation (r = 0.99) validates Hooke’s Law. You can then use the regression line equation y = a + bx to predict extension for a given load within the tested range. Always be cautious about extrapolating beyond the data; a spring will eventually deform non-linearly.
在工程学中,你可能研究施加在弹簧上的负载(kg)与其伸长量(cm)之间的关系。强正线性相关(r = 0.99)验证了胡克定律。然后你可以用回归直线方程 y = a + bx 在测试范围内预测给定负载下的伸长量。始终谨防在数据范围之外进行外推;弹簧最终会发生非线性形变。
8. Using Box Plots and Cumulative Frequency in Environmental Studies | 环境研究中的箱线图与累积频数
Environmental monitoring often produces large datasets, such as particulate matter (PM2.5) readings from 50 sensors. A cumulative frequency curve lets you quickly estimate the median, quartiles, and interquartile range (IQR). For PM2.5 concentrations (μg/m³), you might find median = 18 μg/m³, lower quartile = 12, upper quartile = 25, giving IQR = 13 μg/m³. The box plot then visually flags any sensors that recorded extreme pollution.
环境监测通常会产生大型数据集,例如 50 个传感器的 PM2.5 读数。累积频数曲线可以让你快速估计中位数、四分位数和四分位距 (IQR)。对于 PM2.5 浓度(μg/m³),你可能会得到中位数 = 18 μg/m³,下四分位数 = 12,上四分位数 = 25,IQR = 13 μg/m³。接着箱线图可以直观地标出记录到极端污染的任何传感器。
Compare two areas using side-by-side box plots. If the urban site has median = 22 μg/m³ and the rural site has median = 10 μg/m³ with non-overlapping notches, you can infer a statistically significant difference in median air quality. Such comparative visual analysis is a key skill for Eduqas statistics investigations.
使用并排箱线图比较两个区域。如果城市站点中位数 = 22 μg/m³,而乡村站点中位数 = 10 μg/m³,且凹槽不重叠,你可以推断空气质量的中位数存在统计学显著差异。这种比较性可视化分析是 Eduqas 统计调查的关键技能。
9. Critiquing Methodology and Bias in News Reports | 批判新闻报道中的方法与偏差
Cross-discipline tasks often ask you to evaluate a newspaper headline claiming “Screen time causes anxiety in teens.” You must check the study’s methodology. Was it observational or experimental? An observational study that finds a correlation cannot establish causation because of confounding variables — for example, teens with existing anxiety might retreat to screens for comfort. The sample might also be self-selected if it was an online poll, leading to voluntary response bias.
跨学科任务常常要求你评价一则新闻标题,如“屏幕时间导致青少年焦虑”。你必须检查研究的方法论。它是观察性研究还是实验性研究?发现相关性的观察性研究无法确立因果关系,因为存在混杂变量——例如,本就有焦虑的青少年可能转向屏幕寻求安慰。如果是网络民意调查,样本也可能是自选的,导致自愿响应偏差。
A well-designed interdisciplinary critique considers: sample size, random allocation (if an experiment), blinding, and potential funding sources. A study funded by a gaming company might present a biased interpretation even if the data itself is sound. Teaching yourself to ask “Who collected the data and why?” is a hallmark of statistical literacy.
一个好的跨学科批判应考虑:样本量、随机分配(如果是实验)、盲法以及潜在的资金来源。由游戏公司资助的研究即使数据本身可靠,也可能给出有偏的解读。学会问“谁收集了数据?为什么?”是统计素养的标志。
10. Decision-Making Under Uncertainty: Cost-Benefit Analysis | 不确定性下的决策:成本-收益分析
Many real-world problems require combining probability with numerical outcomes to inform decisions. In health economics, you might weigh the cost of a vaccination programme against the expected cost of treating the disease. If a vaccine costs £10 per dose and prevents a disease that costs £500 per case, with a 20% infection risk in a population of 10,000, the expected treatment cost without vaccine = 0.20 × 10,000 × £500 = £1,000,000. Vaccinating everyone costs £100,000, saving £900,000.
许多现实问题需要将概率与数值结果结合起来以指导决策。在健康经济学中,你可能需要权衡疫苗接种项目的成本与治疗该疾病的预期成本。如果疫苗每剂 10 英镑,可预防一种每例治疗费 500 英镑的疾病,在人口 10,000 人中感染风险为 20%,则无疫苗时的预期治疗成本 = 0.20 × 10,000 × £500 = £1,000,000。为所有人接种成本为 £100,000,可节省 £900,000。
You can extend this to decision trees with multiple branches. A manufacturing company might choose between two production methods: Method A has a 5% defect rate and costs £2 per unit; Method B has a 1% defect rate but costs £3 per unit. For 50,000 units, calculating expected losses from defects (£10 per defective unit) determines which method yields lower total expected cost.
你可以将此延伸到多分支决策树。一家制造公司可能在两种生产方法中选择:方法 A 的次品率为 5%,单位成本 £2;方法 B 的次品率为 1%,但单位成本 £3。对于 50,000 件产品,计算次品带来的预期损失(每件次品损失 £10),即可确定哪种方法总预期成本更低。
11. Normal Distribution and Quality Control in Industry | 工业中的正态分布与质量控制
Many manufacturing tolerance problems assume a normal distribution. If a machine fills crisp packets with mean weight μ = 200 g and standard deviation σ = 3 g, and the specification requires weight between 194 g and 206 g, you calculate z-scores to find the proportion likely out of spec. The lower z = (194 − 200) ÷ 3 = −2.0, the upper z = (206 − 200) ÷ 3 = 2.0. From tables, about 95.4% lie within these limits, so around 4.6% of packets fall outside, which may trigger a machine adjustment.
许多制造公差问题都假定服从正态分布。如果一台机器填充薯片包装袋,平均重量 μ = 200 g,标准差 σ = 3 g,而规格要求重量在 194 g 到 206 g 之间,你可以计算 z 分数以找出可能超标的比例。下限 z = (194 − 200) ÷ 3 = −2.0,上限 z = (206 − 200) ÷ 3 = 2.0。查表可知约 95.4% 落在该范围内,因此约 4.6% 的包装袋不合格,这可能会触发机器调整。
This connects to the concept of standard error when you take samples. If a quality inspector takes a sample of 9 packets and finds their mean is 198 g, the standard error of the mean = σ ÷ √n = 3 ÷ 3 = 1 g. A 95% confidence interval for the true mean would be 198 ± 1.96 × 1, i.e., 196.04 g to 199.96 g. This interval does not contain 200 g, suggesting the machine may need recalibration.
这关联到抽样时的标准误差概念。如果质检员抽取 9 包样本,发现平均重量为 198 g,则均值的标准误差 = σ ÷ √n = 3 ÷ 3 = 1 g。真实均值的 95% 置信区间为 198 ± 1.96 × 1,即 196.04 g 到 199.96 g。该区间不包含 200 g,表明机器可能需要重新校准。
12. Combining Graphical Methods to Tell a Data Story | 组合图形方法讲述数据故事
An interdisciplinary investigation often culminates in a report that blends graphs to support a conclusion. You might use a choropleth map to show regional vaccination rates, a line graph to display the downward trend in cases, and a bar chart comparing case rates by age group. Each visual must be clearly labelled with units, title, and data source so a reader can evaluate the evidence without any prior knowledge.
跨学科调查通常在报告中综合运用多种图形来支撑结论。你可以使用等值区域图来显示各地区疫苗接种率,用折线图展示病例的下降趋势,用条形图比较不同年龄组的病例率。每个图表必须清晰标注单位、标题和数据来源,以便读者无需任何先验知识就能评估证据。
You must also choose the graph appropriate to data type. Time series data belong on a line graph; categorical comparisons suit bar charts (discrete) or pie charts (proportions of a whole). For bivariate continuous data, a scatter plot with a line of best fit is preferred. Critically, avoid misleading axes — truncated scales can exaggerate small differences, so always check if the y-axis starts at zero for bar charts representing frequencies.
你还必须根据数据类型选择恰当的图形。时间序列数据适合用折线图;类别比较适合条形图(离散)或饼图(整体的比例)。对于双变量连续数据,首选带最佳拟合线的散点图。关键是要避免误导性的坐标轴——截断的刻度可能夸大微小差异,因此在表示频数的条形图中,务必检查 y 轴是否从零开始。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导