📚 IGCSE Cambridge Statistics: A Parent’s Guide to Supporting Your Child | IGCSE剑桥统计:家长辅导指南
This guide helps parents understand what Cambridge IGCSE Statistics involves and how to support a teenager through revision without needing to become a subject expert.
本指南帮助家长了解剑桥IGCSE统计学考什么,以及如何在不成为学科专家的情况下支持孩子复习。
Statistics rewards clear thinking, careful presentation and sensible interpretation of real-world data. Parents can make a real difference by helping with routines, practice habits and a calm study environment.
Cambridge IGCSE Statistics is a practical subject focused on real-world data, not just abstract mathematics. Students learn how to plan a statistical enquiry, collect or obtain data, present it clearly, analyse it with appropriate measures, and write sensible conclusions.
Parents often find this subject more accessible than pure mathematics because many examples come from everyday life, such as surveys, sports, prices, weather and business.
The Cambridge IGCSE Statistics course is usually assessed through two compulsory written papers, each about two hours long and each worth half of the final mark. There is no separate core or extended tier; all candidates answer the same papers.
A calculator is normally allowed, but students should not rely on it without understanding. The exam rewards clear method, accurate graph work, and correct interpretation of results.
Data collection starts with deciding between a census and a sample. A census asks every member of the population, while a sample asks only part of it; sampling is cheaper and quicker but can introduce bias.
Key sampling methods include simple random sampling, stratified sampling, systematic sampling, quota sampling and convenience sampling. Stratified sampling is especially useful when the population contains clear groups.
4. Charts, Diagrams and Data Presentation | 图表与数据展示
Choosing the correct diagram depends on the data type. Categorical data is best shown by bar charts or pie charts, while continuous data is often shown by histograms, cumulative frequency curves or box plots.
Common diagrams in the syllabus include stem-and-leaf diagrams, scatter graphs, frequency polygons and comparative charts. Students must label axes, use sensible scales, and give titles where appropriate.
📚 IGCSE Cambridge Statistics: Mapping UK University Entry Requirements | IGCSE 剑桥统计:英国大学申请要求对照
For international students planning to apply to UK universities, IGCSE grades are more than just a school report. They are often the first formal evidence of academic ability that admissions tutors use to judge your potential. Cambridge IGCSE Statistics, as an applied mathematics subject, can play a useful supporting role in your application, especially for courses that value data analysis, quantitative reasoning, and evidence-based thinking.
1. Why IGCSE Statistics Matters for UK Admissions | 为什么 IGCSE 统计对英国大学申请重要
UK universities do not usually list IGCSE Statistics as a compulsory subject, but they do recognise it as a strong signal of your quantitative ability. Because statistics involves collecting, presenting, and interpreting data, it supports the analytical skills expected in many degree courses, from psychology to economics.
Admissions tutors often look for evidence beyond the minimum requirements. A good grade in IGCSE Statistics shows that you are comfortable with chance, variability, and data-based judgement, which are central to modern research and professional practice.
2. How UK Universities View IGCSE Grades | 英国大学如何看待 IGCSE 成绩
Most UK universities treat GCSE/IGCSE results as an early indicator of academic discipline. English Language and Mathematics are the two subjects most commonly required, with typical minimum grades ranging from 4/C to 6/B depending on the institution and course.
IGCSE Statistics is usually classified as an additional or supporting subject. It is rarely accepted as a replacement for IGCSE Mathematics, because universities want to see a full course covering algebra, geometry, and number skills. However, a strong statistics grade can strengthen borderline applications.
3. Typical GCSE/IGCSE Requirements by University Group | 按大学梯队的典型 GCSE/IGCSE 要求
The table below gives a general comparison of IGCSE expectations across different groups of UK universities. Always check the exact course page, because requirements vary by department and year.
📚 IGCSE Cambridge Statistics: Report Writing Framework and Model Answer | IGCSE 剑桥统计:论文写作框架与范文
In IGCSE Cambridge Statistics, a high-scoring written report or structured data-response answer is not just a collection of correct numbers. It is a logical argument built around a clear aim, appropriate data, sensible calculations, and a conclusion that refers back to the original question.
This article gives a repeatable writing framework, worked model paragraphs, and examiner-focused advice for IGCSE Statistics report or investigation-style questions.
本文为 IGCSE 统计报告或探究类题目提供可复用的写作框架、范文段落和考官视角的建议。
1. Why a Clear Framework Wins Marks | 为什么清晰框架能得分
A clear report structure helps the examiner see your statistical thinking. In IGCSE Statistics, marks are awarded for choosing a suitable sample, drawing an appropriate diagram, calculating a correct average or spread, and explaining the result in context. If these stages are jumbled together, it is easy to lose communication marks.
The examiner is not marking a literary essay. You should write short, precise statements with numbers, units and comparisons. Each section should lead naturally to the next, from aim to data to analysis to conclusion.
A useful structure is PPDAC: Problem, Plan, Data, Analysis, Conclusion. Many IGCSE investigations can add an Evaluation at the end, making PPDACE. This cycle helps you avoid jumping straight from raw data to a conclusion.
What is the answer and how reliable is it? 答案是什么,可靠性如何?
3. Before Writing: Read the Command Words | 动笔前:读懂指令词
Command words tell you the depth required. ‘State’ needs a short answer; ‘calculate’ needs a method; ‘compare’ needs data quoted; ‘evaluate’ needs a limitation and improvement.
Formula, substitution, answer with units 公式、代入、带单位答案
Compare 比较
Give similarity or difference 说明相同或不同
Quote medians, IQRs or means 引用中位数、IQR 或平均数
Comment/Interpret 解释
Explain in context 结合情境解释
Say what the value means 说明该值的含义
Evaluate 评价
Give limitations and improvements 给出局限与改进
Because…, could be improved by… 因为……,可通过……改进
4. Section 1: Aim and Hypothesis | 第一部分:目的与假设
Start with a precise aim: ‘To investigate whether …’. Then write a prediction that can be tested, for example ‘I predict that … because …’. A hypothesis should be specific enough to be supported or rejected by data.
开头写出精确目的:“To investigate whether …”,再写一个可检验的预测,例如“I predict that … because …”。假设必须足够具体,能够被数据支持或否定。
A weak aim says ‘I will look at hand spans.’ A strong aim says ‘To investigate whether the median hand span of Year 10 boys is greater than that of Year 10 girls.’
弱目的如“I will look at hand spans”,强目的则写“旨在研究 10 年级男生的手掌宽度中位数是否大于女生”。
In the hypothesis, give a reason. For example: ‘I predict that boys will have a larger median hand span because male growth patterns often produce wider hands.’
在假设中给出理由。例如:“我预测男生的中位数手掌宽度更大,因为男性发育模式通常使手掌更宽。”
5. Section 2: Data Collection and Sampling | 第二部分:数据收集与抽样
State the population, sample size, sampling method
Published by TutorHao | IGCSE 统计 Revision Series | aleveler.com
International statistics competitions reward students who can move beyond formula memorisation and apply statistical thinking to unfamiliar scenarios. IGCSE Cambridge Statistics provides an excellent foundation because it covers data representation, probability, sampling and interpretation, all of which are common in junior and intermediate contests.
This guide blends Cambridge IGCSE Statistics revision with competition strategy, including core knowledge, common question traps, time management and a 12-week training plan.
1. Understanding the Exam and Competition Landscape | 了解考试与竞赛格局
Cambridge IGCSE Statistics typically covers descriptive statistics, probability, sampling methods, bivariate data and interpretation of statistical diagrams. International competitions such as UKMT, AMC 8/10 and SASMO often include statistical reasoning questions within broader mathematics papers, while dedicated statistics olympiads may require data analysis and project-style responses.
Before you enter a competition, compare its syllabus with IGCSE Statistics. This helps you identify which topics need extension, especially counting principles, conditional probability and statistical inference.
Winter break is a crucial window to turn scattered knowledge into exam-ready skills. A focused four-to-six week plan, built around the Cambridge IGCSE Statistics syllabus, will help you improve accuracy, speed, and confidence across data handling, probability, bivariate data, and index numbers.
Start by setting a realistic weekly timetable: four to six weeks of revision, with 90-120 minutes per day, is usually more effective than occasional long study sessions. The plan should rotate between learning, question practice, and error correction so that weaknesses are addressed quickly.
Begin with a diagnostic past paper or topic test to identify which areas are costing you the most marks. Record your scores by topic, then allocate more revision days to low-scoring sections such as histograms, tree diagrams, or moving averages.
Keep a revision log with the date, topic, questions attempted, and mistakes corrected. Tracking progress makes the holiday plan measurable and reduces the risk of repeating the same errors in the final exam.
2. Data Collection and Sampling Methods | 数据收集与抽样方法
Revise the difference between a census and a sample. A census collects information from every member of a population, giving complete data but often requiring more time and money; a sample is cheaper and faster but can be biased if not chosen carefully.
Know how to select a stratified sample from grouped populations. For example, if a school has 300 boys and 200 girls and you need a stratified sample of 50 students, select 30 boys and 20 girls to maintain the same proportions.
Stratum sample = (stratum size ÷ population size) × total sample size
Also practise identifying bias in surveys. Leading questions, voluntary response samples, and incomplete sampling frames can all make data unrepresentative, and exam questions often ask you to suggest how to remove or reduce such bias.
Different data types require different graphs: bar charts and pie charts suit categorical data, while histograms, frequency polygons, box plots, and cumulative frequency graphs are used for numerical data. Choosing the correct graph is itself a common exam skill.
For histograms with unequal class widths, the vertical axis must show frequency density, not raw frequency. Calculate frequency density by dividing the frequency by the class width, otherwise the areas will not represent the data correctly.
Cumulative frequency graphs allow you to estimate the median, lower quartile, upper quartile, and percentiles. Read the value at the relevant cumulative frequency, and do not confuse cumulative frequency with frequency.
Box plots are especially useful for comparing two or more distributions. Always label the minimum, lower quartile, median, upper quartile, and maximum clearly, and refer to spread and central tendency when making comparisons.
4. Measures of Central Tendency and Spread | 集中趋势与离散程度
For central tendency, revise the mean, median, and mode. The mean uses all data values and is affected by outliers; the median is resistant to outliers and is often preferred for skewed data; the mode is the most frequent value.
For spread, the range is the simplest measure: maximum minus minimum. The interquartile range, or IQR, focuses on the middle 50% of the data and is less affected by extreme values.
You should also be confident with standard deviation, which measures how far values are from the mean. A lower standard deviation indicates that data are clustered closely around the mean, while a higher one shows greater spread.
Probability measures the chance of an event and always lies between 0 and 1. The basic formula is the number of favourable outcomes divided by the total number of equally likely outcomes, and the sum of probabilities of all outcomes is 1.
P(A) = number of favourable outcomes ÷ total number of outcomes
For mutually exclusive events, use P(A or B) = P(A) + P(B). For independent events, use P(A and B) = P(A) × P(B). Reading the words ‘or’ and ‘and’ carefully prevents many common errors.
对于互斥事件,使用 P(A 或 B) = P(A) + P(B)。对于独立事件,使用 P(A 和 B) = P(A) × P(B)。仔细辨别“或”和“和”可以避免很多常见错误。
P(A or B) = P(A) + P(B)
P(A and B) = P(A) × P(B)
Tree diagrams are essential for successive or conditional events. Multiply probabilities along each branch to obtain the probability of that path, then add the relevant path probabilities to answer the question.
Practise reverse tree diagrams for conditional probability questions, where you are given a final outcome and asked to find an earlier probability. These questions require clear labelling and systematic working.
练习反向树图的条件概率题,即已知最终结果求早期概率的题型。这类题需要清晰的标注和系统的解题步骤。
6. Bivariate Data and Correlation | 双变量数据与相关
Bivariate data involve two variables measured for the same individuals or items. A scatter diagram shows whether there is positive correlation, negative correlation, or no clear relationship between the variables.
Draw a line of best fit by eye, balancing the points above and below the line. Use the line to estimate an unknown value by reading from one axis to the line and then to the other axis.
Interpolation, estimating within the range of the plotted data, is generally reliable. Extrapolation, estimating outside the range, can be risky because the relationship may not continue.
If your syllabus includes Spearman’s rank correlation coefficient, practise ranking two sets of data and computing the differences in ranks. The formula tests the strength of a monotonic relationship between variables.
📚 Cross-Disciplinary Integrated Question Training for IGCSE Cambridge Statistics | IGCSE 剑桥统计:跨学科综合题型训练
In IGCSE Cambridge Statistics, examination questions rarely appear as isolated calculations. Instead, they are embedded in real-world contexts from biology, business, geography, physics and sport. This article provides structured cross-disciplinary question training so that you can move confidently between statistical techniques and applied scenarios.
Cambridge IGCSE Statistics (0479) requires you to select, apply and interpret statistical methods. A biology question may ask you to compare reaction times before and after caffeine; a business question may require a moving average for quarterly sales. The context determines which graph, average or spread is appropriate.
Cross-disciplinary training helps you avoid the mistake of treating every question as a generic ‘find the mean’ task. You must read the scenario, identify the variable type, and decide whether to compare centres, spreads or trends.
Common contexts in Cambridge IGCSE Statistics include laboratory experiments, market research, population studies, quality control and sports performance. Each context has its own units, measurement issues and sensible interpretation.
A typical biology-integrated question gives resting heart rates for two groups, such as trained athletes and non-athletes. You may need to calculate the mean, median and interquartile range, then comment on which group has the lower centre and smaller spread.
For grouped heart-rate data, use midpoints to estimate the mean. The formula is:
对于分组心率数据,使用组中值来估计平均数。公式为:
x̄ = Σfx ÷ Σf
where x is the midpoint and f is the frequency. When comparing box plots, always refer to median, quartiles and outliers, not just the range.
其中 x 是组中值,f 是频数。在比较箱线图时,始终要提及中位数、四分位数和异常值,而不仅仅是全距。
English: Identify the variable as continuous and the data as ungrouped or grouped. 中文:识别变量为连续变量,数据为未分组或分组。
English: Use median and interquartile range when data are skewed or contain outliers. 中文:当数据偏态或含异常值时使用中位数和四分位距。
English: When comparing two groups, quote both the centre and the spread. 中文:比较两组数据时,同时引用中心值和离散程度。
3. Geography: Population Pyramids and Demographic Measures | 地理:人口金字塔与人口指标
Population pyramids combine frequency diagrams for age groups of males and females. In IGCSE Statistics, you may be asked to compare the percentage of the population aged 0-14 and 65+, or to calculate the dependency ratio.
Dependency ratio = [(population aged 0-14 + population aged 65+) ÷ population aged 15-64] × 100
Be careful to use the correct denominator and convert final answers to a percentage where required.
注意使用正确的分母,并在需要时将最终结果转换为百分比。
When interpreting a population pyramid, a wide base indicates a high birth rate, while a narrow apex suggests a smaller elderly population. You should link the shape to statistical measures such as median age and age-specific proportions.
4. Business: Interpreting Sales Trends and Index Numbers | 商业:解读销售趋势与指数
Business contexts often require you to smooth time-series data using a moving average. For quarterly sales, a four-point moving average is centred to align with the original time periods.
Index numbers compare price or quantity changes relative to a base period. The formula is:
指数将价格或数量变化与基期进行比较。公式为:
Index number = (current value ÷ base value) × 100
When the index rises from 100 to 112, the percentage increase is 12%, not 112%.
当指数从 100 上升到 112 时,增长率为 12%,而不是 112%。
You may also be asked to calculate a weighted index number, where each item is multiplied by its weight before summing. Always check whether the base year is given as 100 or as another value.
你还可能需要计算加权指数,即每项先乘以其权重再求和。始终检查基年是否为 100 或其他值。
5. Physics: Experimental Measurement and Uncertainty | 物理:实验测量与不确定度
In physics experiments, repeated measurements of time, length or current produce variation. You may calculate the mean and standard deviation, then identify whether a result is repeatable or reproducible.
If repeated readings give 2.1, 2.3, 2.2, 2.4, the mean is:
如果重复读数为 2.1、2.3、2.2、2.4,平均数为:
x̄ = (2.1 + 2.3 + 2.2 + 2.4) ÷ 4 = 2.25
The range is 2.4 − 2.1 = 0.3. A smaller standard deviation indicates less experimental uncertainty.
全距为 2.4 − 2.1 = 0.3。标准差越小,说明实验不确定度越小。
When an experiment produces outliers, investigate them rather than automatically removing them. In IGCSE Statistics, you should be able to identify an outlier using quartiles and the interquartile range.
Sports data, such as 100 m sprint times or basketball scores, are often compared using box plots and histograms. A lower median sprint time indicates better performance, so interpret direction carefully.
If group A has a median of 11.2 s and IQR of 0.4 s, while group B has a median of 11.8 s and IQR of 0.9 s, group A is faster on average and more consistent.
如果 A 组中位数为 11.2 秒,四分位距为 0.4 秒;B 组中位数为 11.8 秒,四分位距为 0.9 秒,则 A 组平均更快且更稳定。
For symmetric distributions, the mean and median are close. For skewed distributions, such as basketball scores with a few very high values, the mean is pulled towards the tail, so the median may be a better summary of typical performance.
7. Environmental Science: Sampling and Estimation | 环境科学:抽样与估计
Environmental studies often use quadrat sampling or capture-recapture. You may estimate population size using the Petersen estimate:
环境研究常使用样方抽样或标志重捕法。你可以使用 Petersen 估计法估算种群大小:
N = (M × C) ÷ R
where M is the number initially marked, C is the total number captured in the second sample, and R is the number of marked individuals recaptured.
其中 M 是首次标记的数量,C 是第二次捕获的总数,R 是重捕到的标记个体数。
For quadrat sampling, if the mean number of daisies per 1 m² quadrat is 12 and the field area is 500 m², the estimated total is 12 × 500 = 6000. Always distinguish between sample statistic and population estimate.
Assumptions matter in capture-recapture: marked individuals must mix randomly, marks must not be lost, and the population must be closed during the study. Mention these assumptions when evaluating an estimate.
8. Economics: Correlation and Regression in Context | 经济:情境中的相关与回归
Economic data such as income and spending often show a positive correlation. You may be asked to draw a scatter diagram, describe correlation, and use a line of best fit for prediction.
收入和支出等经济数据通常呈现正相关。你可能需要绘制散点图、描述相关性,并使用最佳拟合线进行预测。
If the least squares regression line is y = 1.8x + 20, where x is hours worked and y is daily earnings, then for x = 6, predicted y = 1.8(6) + 20 = 30.8. Avoid extrapolating far beyond the data range.
如果最小二乘回归线为 y = 1.8x + 20,其中 x 是工作小时数,y 是日收入,则当 x = 6 时,预测 y = 1.8(6) + 20 = 30.8。避免对数据范围之外作过度外推。
Correlation does not imply causation. If ice cream sales and drowning incidents both rise in summer, the hidden variable is temperature, not ice cream causing drowning. State such limitations when interpreting a regression model.
Consider this integrated question: ‘A biologist records the lengths of 40 leaves from two plants. Plant A has mean 8.2 cm and standard deviation 1.1 cm; Plant B has mean 8.2 cm and standard deviation 2.4 cm. Compare the distributions and suggest which plant is more uniform.’
考虑这道综合题:“一位生物学家记录了两株植物的 40 片叶子长度。植物 A 的平均数为 8.2 厘米,标准差为 1.1 厘米;植物 B 的平均数为 8.2 厘米,标准差为 2.4 厘米。比较分布,并指出哪株植物更均匀。”
Step 1: Note that the means are equal, so the centre is the same. Step 2: Compare spreads; Plant A has a smaller standard deviation, so its leaf lengths are less variable. Step 3: Conclude that Plant A is more uniform.
第 1 步:注意平均数相等,因此中心相同。第 2 步:比较离散程度;植物 A 的标准差更小,因此其叶长变异更小。第 3 步:得出植物 A 更均匀的结论。
For top marks, always quote the statistics: ‘Both means are 8.2 cm, but the standard deviation is 1.1 cm for A and 2.4 cm for B. Plant A is more uniformly distributed because its standard deviation is smaller.’
为获得高分,始终引用统计数据:“两者的平均数均为 8.2 厘米
Published by TutorHao | IGCSE 统计 Revision Series | aleveler.com
📚 Common Misconceptions in IGCSE Cambridge Statistics and How to Fix Them | IGCSE 剑桥统计常见误区与纠正方法
In IGCSE Cambridge Statistics, students often understand the basic definitions but lose marks through subtle misapplications. A small misunderstanding – such as using frequency instead of frequency density in a histogram – can change an entire answer. This guide highlights the most common errors and shows the correct method step by step.
Misconception: The mean is always the best measure of centre. In reality, an outlier such as 100 in the set 2, 3, 4, 5, 100 makes the mean 22.8, which does not describe the typical value.
Correction: For skewed data or data with outliers, use the median. The median of 2, 3, 4, 5, 100 is 4, which is far more representative. Use the mode for categorical data or when you need the most frequent value.
Misconception: A histogram is not a bar chart. Drawing frequency on the vertical axis gives a distorted shape when class widths are unequal.
误区:直方图不是条形图。当组距宽度不相等时,在纵轴上绘制频率会使形状失真。
Correction: The vertical axis must show frequency density. Area of each bar equals frequency density × class width, which equals frequency. This keeps the total area proportional to total frequency.
3. Cumulative Frequency Graphs and Quartiles | 累计频率图与四分位数
Misconception: Plot cumulative frequency against the class midpoint or lower boundary. This shifts the curve and gives incorrect quartile readings.
误区:将累计频率相对于组中点或下组界绘制。这会使曲线偏移,得到错误的四分位数读数。
Correction: Always plot each cumulative frequency at the upper boundary of its interval. The median is read at the n ÷ 2 value, the lower quartile at n ÷ 4, and the upper quartile at 3n ÷ 4.
纠正:始终在每一组的上组界处绘制累计频率。中位数在 n ÷ 2 处读取,下四分位数在 n ÷ 4 处读取,上四分位数在 3n ÷ 4 处读取。
Median position = n ÷ 2, Q₁ position = n ÷ 4, Q₃ position = 3n ÷ 4
4. Mean from Grouped Data | 分组数据的平均数
Misconception: When estimating the mean from a grouped frequency table, some students use the lower or upper class limit as the x value. This gives a biased estimate.
误区:在计算分组频数表的平均数估计值时,有些学生使用下组限或上组限作为 x 值,这会产生有偏估计。
Correction: Use the midpoint of each class interval: x = (lower boundary + upper boundary) ÷ 2. Then apply the grouped mean formula.
纠正:使用每组的组中点:x = (下组界 + 上组界) ÷ 2。然后应用分组平均数公式。
Estimated mean x̄ = Σfx ÷ Σf
5. Mutually Exclusive and Independent Events | 互斥事件与独立事件
Misconception: Treating mutually exclusive events and independent events as the same thing. Mutually exclusive means both cannot occur together; independent means one event does not affect the probability of the other.
Correction: For mutually exclusive events, P(A or B) = P(A) + P(B). For independent events, P(A and B) = P(A) × P(B). These formulas apply in different situations.
纠正:对于互斥事件,P(A 或 B) = P(A) + P(B)。对于独立事件,P(A 和 B) = P(A) × P(B)。这两个公式适用于不同的情境。
P(A or B) = P(A) + P(B) − P(A and B)
6. Conditional Probability and the Sample Space | 条件概率与样本空间
Misconception: When asked for P(A given B), students continue to use the original total sample space instead of restricting to B. This overstates or understates the probability.
误区:求 P(A 给定 B) 时,学生仍使用原始总样本空间,而不是将样本空间限制在 B 中。这会高估或低估概率。
Correction: Use the conditional probability formula or reduce the sample space. The denominator becomes the number of outcomes in event B, not the total sample size.
纠正:使用条件概率公式或缩小样本空间。分母变为事件 B 的结果数,而不是总样本量。
P(A | B) = P(A ∩ B) ÷ P(B)
7. Probability Tree Diagrams | 概率树形图
Misconception: In a tree diagram, some students multiply along a path but forget that branch probabilities must sum to 1 at each node. Others add probabilities from different paths even when they are not mutually exclusive.
Correction: At each branch point, the probabilities must add to 1. Multiply along a path to find the probability of a combined outcome. Add the probabilities of different successful paths because those paths are mutually exclusive.
Misconception: Students often compute variance correctly but then forget to take the square root for standard deviation, or they divide by n for a sample when they should use n − 1.
误区:学生常常正确计算方差,却忘记对标准差开方,或者在样本标准差中除以 n,而应该除以 n − 1。
Correction: Standard deviation is the square root of variance. Use divisor n for a population and n − 1 for a sample. Always check whether the question refers to a population or a sample.
纠正:标准差是方差的平方根。总体使用除数 n,样本使用除数 n − 1。始终检查题目指的是总体还是样本。
s = √(Σ(x − x̄)² ÷ (n − 1))
9. Correlation and Causation | 相关与因果
Misconception: A high correlation coefficient proves that one variable causes the other. Correlation only measures the strength and direction of a linear relationship.
误区:高相关系数证明一个变量导致另一个变量。相关只衡量线性关系的强度和方向。
Correction: Always consider possible lurking variables or coincidence. For example, ice cream sales and drowning incidents may be correlated because both increase in summer, but ice cream does not cause drowning.
Misconception: Convenience sampling or volunteer sampling gives results that can be generalised to the whole population. These methods often over-represent certain groups and introduce bias.
误区:便利抽样或自愿抽样可以推广到整个总体。这些方法
Published by TutorHao | IGCSE 统计 Revision Series | aleveler.com
Although the Cambridge IGCSE Statistics syllabus (0479) does not include a separate speaking or listening paper, strong spoken English and accurate listening skills are essential for explaining statistical concepts, interpreting data aloud, and following instructions in class and in examinations. This guide focuses on the vocabulary, phrases, and listening strategies that help IGCSE Statistics learners communicate statistical ideas clearly.
1. Core Statistical Terms You Must Pronounce Clearly | 必须清晰发音的核心统计术语
In spoken English, many statistical terms are similar and can be misheard. You need to pronounce and recognise these words accurately: mean, median, mode, range, quartile, percentile, histogram, frequency polygon, cumulative frequency, interquartile range, standard deviation, probability, correlation, regression, and outlier. Practise saying each word slowly, then in full sentences.
2. Describing Data Orally: Useful Sentence Patterns | 口头描述数据:常用句型
Useful sentence patterns include: “The mean is greater than the median, which suggests the distribution is positively skewed.” “The range is 24, so the data is quite spread out.” “There is a strong positive correlation between hours of revision and test scores.” “The modal class is 20–30 minutes.” “The data is roughly symmetrical.” These sentence frames help you respond fluently.
常用句型包括:”The mean is greater than the median, which suggests the distribution is positively skewed.”(平均数大于中位数,表明分布呈正偏态。)”The range is 24, so the data is quite spread out.”(极差为 24,所以数据分布较分散。)”There is a strong positive correlation between hours of revision and test scores.”(复习时长与测试分数之间存在强正相关。)”The modal class is 20–30 minutes.”(众数所在组是 20–30 分钟。)”The data is roughly symmetrical.”(数据大致对称。)这些句式框架有助于流利作答。
3. Listening for Instructions in Statistics Lessons | 听清统计课上的指令
When a teacher or recording gives statistical tasks, listen for command verbs: calculate, construct, interpret, compare, estimate, plot, draw, and describe. Also note words like “nearest integer”, “two decimal places”, “show your working”, and “using the graph”. Pause and replay any instruction that is unclear.
当老师或录音给出统计任务时,注意听清指令动词:calculate(计算)、construct(构造/绘制)、interpret(解释)、compare(比较)、estimate(估计)、plot(描点)、draw(绘图)、describe(描述)。还要注意 “nearest integer”(精确到整数)、”two decimal places”(保留两位小数)、”show your working”(写出步骤)、”using the graph”(根据图表)等。不清楚的指令要暂停并重听。
4. Interpreting Charts and Graphs Aloud | 口头解读图表
When speaking about a chart, describe the axes first: “The horizontal axis shows time in minutes; the vertical axis shows frequency.” Then identify the overall shape: “The distribution is bimodal with peaks at 10 and 30.” Use words like trend, peak, trough, spread, cluster, and gap. For a scatter graph, say: “As x increases, y decreases; this indicates a negative correlation.”
口头解读图表时,先描述坐标轴:”The horizontal axis shows time in minutes; the vertical axis shows frequency.”(横轴表示时间(分钟),纵轴表示频数。)然后识别整体形状:”The distribution is bimodal with peaks at 10 and 30.”(该分布是双峰的,峰值在 10 和 30。)使用 trend(趋势)、peak(峰值)、trough(谷值)、spread(离散程度)、cluster(聚集)、gap(间隙)等词。对于散点图,可说:”As x increases, y decreases; this indicates a negative correlation.”(随着 x 增大,y 减小,表明存在负相关。)
5. Probability Language: Speaking and Listening | 概率语言:口语与听力
Probability involves everyday language with precise meanings. “Impossible” means probability 0; “certain” means probability 1; “equally likely” means each outcome has the same chance; “biassed” means not fair; “random” means every member has an equal chance of being selected. Practise sentences such as “The probability of rolling a six is one sixth”, “The events are mutually exclusive”, and “The sample space has 36 outcomes.”
概率涉及具有精确含义的日常用语。”Impossible”(不可能)表示概率为 0;”certain”(必然)表示概率为 1;”equally likely”(等可能)表示每个结果机会相同;”biassed”(有偏的)表示不公平;”random”(随机)表示每个成员被选中的机会均等。练习句子,如 “The probability of rolling a six is one sixth”(掷出六点的概率是六分之一)、”The events are mutually exclusive”(这些事件互斥)、”The sample space has 36 outcomes”(样本空间有 36 个结果)。
6. Common Mistakes in Spoken Statistics | 统计口语中的常见错误
Avoid saying “average” when you mean a specific measure; say “mean”, “median”, or “mode”. Do not confuse “correlation” with “causation”: a correlation does not prove that one variable causes the other. Be careful with “percentage” and “percentage point”: a rise from 10% to 15% is a 5 percentage point increase but a 50% relative increase. Know the difference between discrete and continuous data.
During lessons, listen for signal words: “For example”, “This is important because”, “Compare this with”, “However”, “Therefore”, and “In contrast”. These help you follow the logic. Take notes in two columns: key term / explanation. If a teacher says “The median is not affected by extreme values, but the mean is”, write down that contrast.
在课堂上,注意听信号词:”For example”(例如)、”This is important because”(这很重要,因为)、”Compare this with”(将此与……比较)、”However”(然而)、”Therefore”(因此)、”In contrast”(相比之下)。这些词有助于你理解逻辑。做笔记时分成两栏:关键术语 / 解释。如果老师说 “The median is not affected by extreme values, but the mean is”(中位数不受极端值影响,但平均数会),记下这一对比。
8. Oral Presentation of a Statistical Investigation | 统计调查的口头展示
Structure your spoken report: 1) Aim: “The aim of my investigation is to find out whether…” 2) Data collection: “I collected a sample of 40 students using random sampling.” 3) Results: “The estimated mean is 65 minutes.” 4) Conclusion: “The data suggests that…” Use phrases like “As shown in the histogram”, “According to my cumulative frequency graph”, and “Overall, the evidence indicates that”.
口头报告结构:1) 目的:”The aim of my investigation is to find out whether…”(我的调查目的是查明……是否……);2) 数据收集:”I collected a sample of 40 students using random sampling.”(我使用随机抽样收集了 40 名学生的样本。)3) 结果:”The estimated mean is 65 minutes.”(估计平均值为 65 分钟。)4) 结论:”The data suggests that…”(数据表明……)。使用 “As shown in the histogram”(如直方图所示)、”According to my cumulative frequency graph”(根据我的累积频数图)、”Overall, the evidence indicates that”(总体而言,证据表明)等短语。
9. Listening to Statistical Vocabulary in Context | 在语境中听辨统计词汇
Hearing terms in full sentences is different from reading them. Practise with a table of common phrases you might hear and their meanings. For example: “Work out the interquartile range” means calculate Q₃ minus Q₁. “Plot the cumulative frequency curve” means draw a graph using the upper class boundary. “Comment on the correlation” means describe strength and direction. Listen for numbers said in different ways: “nought point three” (0.3), “one in twenty” (1/20), “fifty per cent” (50%).
在完整句子中听辨术语与阅读术语不同。使用常见短语及含义表进行练习。例如:”Work out the interquartile range”(计算四分位距)意为用 Q₃ 减去 Q₁;”Plot the cumulative frequency curve”(绘制累积频数曲线)意为使用上组界绘图;”Comment on the correlation”(评论相关性)意为描述强度和方向。注意数字的不同说法:”nought point three”(0.3)、”one in twenty”(二十分之一)、”fifty per cent”(50%)。
Spoken phrase
Meaning
中文含义
Work out the interquartile range
Calculate Q₃ − Q₁
计算四分位距
Plot the cumulative frequency curve
Draw a graph using upper class boundaries
使用上组界绘制累积频数曲线
Comment on the correlation
Describe strength and direction
描述相关性的强度和方向
Estimate the mean from a grouped table
Use midpoints to find an approximate average
使用组中值估计平均数
10. Exam and Assessment Tips for Spoken Responses | 考试与评估中的口语作答技巧
Although the written paper is the main assessment, oral questioning may be part of classroom assessment. Answer in short, structured sentences: state the value, then the interpretation. For example: “The probability is 0.25, which means there is a one in four chance.” If you do not understand a spoken question, ask for clarification: “Could you repeat the question, please?” or “Do you mean the median or the mode?”
尽管笔试是主要评估方式,但口头提问也可能是课堂评估的一部分。回答时使用简短、结构化的句子:先陈述数值,再解释含义。例如:”The probability is 0.25, which means there is a one in four chance.”(概率是 0.25,这意味着四分之一的机会。)如果你没听懂口头问题,可以请求澄清:”Could you repeat the question, please?”(请重复一下问题好吗?)或 “Do you mean the median or the mode?”(你是指中位数还是众数?)
Published by TutorHao | Statistics Revision Series | aleveler.com
This quick reference summarises the core formulas and definitions needed for the Cambridge IGCSE Statistics course. Use it alongside past papers and topic questions for efficient revision.
Data can be qualitative (labels or categories, such as eye colour) or quantitative (numerical values). Quantitative data are discrete if they can only take particular values, and continuous if they can take any value in an interval.
For stratified sampling, the sample from each group is proportional to the group’s size. The required number from a stratum is calculated as:
分层抽样中,每层抽取的样本量与该层大小成比例。某层所需样本数计算如下:
Stratum sample size = (stratum size / population size) × total sample size
Random sampling gives each member an equal chance of selection; systematic sampling selects every kth member; cluster sampling selects whole groups; quota sampling continues until set numbers in categories are met.
随机抽样使每个成员被抽中的概率相同;系统抽样每隔固定间隔 k 抽取一个;整群抽样直接抽取整个群体;配额抽样按各类别的设定人数抽取直到满足配额。
A sampling frame is a list of all members in the population. Bias can occur if some members are excluded, over-represented, or if non-responses are ignored.
抽样框是总体中所有成员的名单。如果某些成员被排除、被过度代表或未回答被忽略,就可能产生偏差。
2. Frequency Tables and Charts | 频数表与统计图
For a histogram with unequal class widths, the height of each bar is given by the frequency density. The area of a bar represents the frequency.
在组距不等的直方图中,每个条形的高度由频数密度给出。条形的面积表示频数。
Frequency density = frequency / class width
A cumulative frequency curve is plotted using the upper class boundary against the cumulative frequency. The median is read at half the total frequency, the lower quartile at one quarter, and the upper quartile at three quarters.
For a pie chart, the angle of each sector is proportional to its frequency. To find the angle:
饼图中每个扇形的角度与其频数成比例。计算角度的公式为:
Sector angle = (frequency / total frequency) × 360°
Stem-and-leaf diagrams preserve original data values and allow median and quartiles to be found directly. Box-and-whisker plots show the minimum, Q1, median, Q3 and maximum on a single scale.
The arithmetic mean is the sum of all values divided by the number of values. For raw data:
算术平均数是所有数值之和除以数值的个数。对于未分组数据:
x̄ = Σx / n
For a frequency distribution, multiply each value by its frequency and divide by the total frequency:
对于频数分布,将每个数值乘以其频数再除以总频数:
x̄ = Σfx / Σf
The median is the middle value when data are arranged in order. If n is odd, it is the (n+1)/2 th value; if n is even, it is the mean of the n/2 th and (n/2 + 1)th values.
中位数是将数据按顺序排列后的中间值。如果 n 为奇数,则为第 (n+1)/2 个值;如果 n 为偶数,则为第 n/2 个与第 (n/2 + 1) 个值的平均数。
The mode is the value that occurs most often. A data set can have one mode, more than one mode, or no mode.
众数是出现次数最多的数值。一组数据可以有一个众数、多个众数或没有众数。
When two or more groups are combined, the overall mean is found by weighting each group mean by its size:
当两个或多个组合并时,总平均数需要用各组大小对各组平均数加权:
x̄_combined = (n₁x̄₁ + n₂x̄₂) / (n₁ + n₂)
A weighted mean uses weights to reflect different
Published by TutorHao | IGCSE 统计 Revision Series | aleveler.com
📚 IGCSE Cambridge Statistics: Unit Test Mock Paper Walkthrough | IGCSE 剑桥统计:单元测试模拟卷解析
This mock paper walkthrough covers the most common Cambridge IGCSE Statistics unit test topics, including averages, frequency tables, cumulative frequency, probability, scatter diagrams, histograms, time series, box plots, index numbers and sampling. Each question below is followed by a clear worked solution so you can check your method and learn the typical marking points.
1. Question 1: Mean, Median, Mode and Range | 第1题:平均数、中位数、众数和极差
The data set is 7, 9, 11, 13, 15, 15, 18, 21, 24. Find the mean, median, mode and range.
数据集为 7、9、11、13、15、15、18、21、24。求平均数、中位数、众数和极差。
First arrange the data in ascending order. The mean is the sum of all values divided by the number of values, so Σx = 133 and n = 9. The median is the 5th value because (9 + 1) ÷ 2 = 5, which gives 15. The mode is the most frequent value, also 15. The range is the largest value minus the smallest value, giving 24 – 7 = 17.
Mean = 133 ÷ 9 = 14.8, Median = 15, Mode = 15, Range = 17
2. Question 2: Frequency Table and Estimated Mean | 第2题:频数表与估计平均数
The grouped frequency table shows the time taken by 25 students to finish a puzzle: 0 ≤ t < 10 has frequency 4, 10 ≤ t < 20 has frequency 7, 20 ≤ t < 30 has frequency 9, and 30 ≤ t < 40 has frequency 5. Estimate the mean time.
该分组频数表显示了 25 名学生完成拼图所用的时间:0 ≤ t < 10 的频数为 4,10 ≤ t < 20 的频数为 7,20 ≤ t < 30 的频数为 9,30 ≤ t < 40 的频数为 5。估计平均时间。
Use the midpoint of each class as the representative value: 5, 15, 25 and 35. Multiply each midpoint by its frequency and add the results: 4 × 5 + 7 × 15 + 9 × 25 + 5 × 35 = 20 + 105 + 225 + 175 = 525. Then divide by the total frequency 25.
The modal class is 20 ≤ t < 30 because it has the largest frequency 9.
众数所在组是 20 ≤ t < 30,因为该组频数最大,为 9。
3. Question 3: Cumulative Frequency, Median and Interquartile Range | 第3题:累积频数、中位数与四分位距
The cumulative frequency table for the heights of 50 students is: less than 150 cm: 0, less than 155 cm: 8, less than 160 cm: 20, less than 165 cm: 35, less than 170 cm: 42, less than 175 cm: 50. Draw a cumulative frequency curve and estimate the median and interquartile range.
Plot the upper class boundary against cumulative frequency and join the points with a smooth curve. The median is the value at half the total frequency, 50 ÷ 2 = 25. Reading from the curve gives approximately 162 cm. The lower quartile is at 25% of the total frequency, 12.5, giving about 156 cm. The upper quartile is at 75%, 37.5, giving about 168 cm. The interquartile range is upper quartile minus lower quartile.
4. Question 4: Probability from a Two-Way Table | 第4题:双列表格中的概率
Eighty students each choose either History or Geography. The table shows: 45 boys, 35 girls; 20 boys choose History, 25 boys choose Geography, 14 girls choose History, 21 girls choose Geography. A student is selected at random. Find P(boy), P(girl and Geography), P(Geography), and P(boy given that the student chose History).
The total number of students is 80. P(boy) = 45 ÷ 80 = 0.5625. The number of girls choosing Geography is 21, so P(girl and Geography) = 21 ÷ 80 = 0.2625. The total choosing Geography is 25 + 21 =
Published by TutorHao | IGCSE 统计 Revision Series | aleveler.com
📚 High-Frequency Topics and Common Error Analysis in Cambridge IGCSE Statistics | IGCSE Cambridge 统计高频考点与易错题分析
Cambridge IGCSE Statistics rewards students who can describe data clearly, calculate probabilities accurately and interpret results in context. This revision article summarises the most frequently examined topics and highlights the mistakes that repeatedly cost marks. Use the paired English-Chinese notes to revise actively: cover one language, test yourself, then check.
Data can be qualitative (non-numerical, such as eye colour) or quantitative (numerical). Quantitative data is further split into discrete data, which can only take exact values, and continuous data, which can take any value in a range.
A common trap is treating shoe size as continuous because it has halves. Shoe size is discrete because only certain sizes exist; height is continuous because it can be 162.3 cm, 162.35 cm and so on.
常见陷阱是认为鞋码有半码就是连续数据。鞋码是离散数据,因为只有某些尺码存在;身高是连续数据,因为它可以是 162.3 cm、162.35 cm 等等。
For sampling, random sampling gives every member an equal chance. Stratified sampling divides the population into groups and takes a sample proportional to each group’s size: n_stratum = (group size ÷ population size) × total sample size.
If a stratified sample of 60 is taken from 300 boys and 200 girls, choose boys: (300 ÷ 500) × 60 = 36, girls: (200 ÷ 500) × 60 = 24. Many errors come from forgetting to multiply after finding the fraction.
2. Charts for Discrete and Continuous Data | 离散与连续数据的图表
Bar charts are used for discrete or categorical data, and the bars usually have gaps between them. Pie charts show proportions, with each angle calculated as (frequency ÷ total frequency) × 360°.
条形图用于离散或分类数据,条形之间通常有空
Published by TutorHao | IGCSE 统计 Revision Series | aleveler.com
📚 IGCSE Cambridge Statistics: Top Scorer Tips | IGCSE 剑桥统计:学霸高分经验分享
Getting an A* in Cambridge IGCSE Statistics is not about memorising every number; it is about knowing how to choose the right method, show clear working, and interpret results in context. Top scorers practise smartly, use the mark scheme as feedback, and treat every past paper as a diagnostic tool.
1. Know the Syllabus and Paper Structure | 吃透考纲与试卷结构
Top scorers begin by printing the official syllabus and turning each bullet point into a checklist. Cambridge IGCSE Statistics includes data collection, data representation, averages, spread, probability, bivariate analysis, time series and index numbers. Each topic can appear with different command words such as calculate, compare, describe and interpret.
When you read the syllabus, highlight verbs because they tell you exactly what the examiner wants. For example, ‘calculate’ means show your method to get marks, while ‘interpret’ means explain what the number means in the given context.
This article summarises the essential knowledge areas tested in the Cambridge IGCSE Statistics syllabus (0479), providing clear explanations and exam-focused formulas.
本文梳理剑桥 IGCSE 统计课程(0479)的核心考点,提供清晰解释和考试常用公式。
1. Types of Data | 数据类型
Data can be classified as qualitative (categorical) or quantitative (numerical).
数据可分为定性(分类)数据或定量(数值)数据。
Quantitative data is further split into discrete data, which can only take specific values, and continuous data, which can take any value within a range.
定量数据进一步分为离散数据(只能取特定值)和连续数据(可在某一范围内取任意值)。
Type
Example / 例子
Qualitative / 定性
Favourite colour / 最喜欢的颜色
Quantitative discrete / 定量离散
Number of students / 学生人数
Quantitative continuous / 定量连续
Height in cm / 身高(厘米)
Choosing the correct type is essential because it determines which diagram and which average are appropriate.
选择正确的数据类型很重要,因为它决定了应使用哪种图表和哪种平均数。
2. Data Collection and Sampling | 数据收集与抽样
Primary data is collected by the investigator directly, while secondary data comes from existing sources such as government publications or online databases.
原始数据由调查者直接收集,二手数据来自现有来源,如政府出版物或在线数据库。
Sampling methods include random sampling, stratified sampling, systematic sampling and quota sampling.
抽样方法包括随机抽样、分层抽样、系统抽样和配额抽样。
In stratified sampling, the population is divided into groups called strata, and a random sample is taken from each group in proportion to its size.
在分层抽样中,总体被分成若干层,并按各层大小比例从每层中随机抽取样本。
A sample should be representative of the population so that valid conclusions can be drawn.
样本应具有总体代表性,以便得出有效结论。
3. Charts and Diagrams | 图表与图示
Common diagrams include bar charts, pie charts, histograms, frequency polygons, stem-and-leaf diagrams and pictograms.
常见图示包括条形图、饼图、直方图、频数多边形、茎叶图和象形图。
For grouped continuous data, a histogram uses area to represent frequency, so the vertical axis is frequency density.
对于分组连续数据,直方图用面积表示频数,因此纵轴为频数密度。
Frequency density = Frequency ÷ Class width
Stem-and-leaf diagrams keep the raw data visible while showing the shape of the distribution.
茎叶图在保留原始数据的同时展现分布形态。
A frequency polygon can be drawn by joining the midpoints of the tops of histogram bars.
频数多边形可通过连接直方图各柱顶部中点绘制。
4. Measures of Central Tendency | 集中趋势度量
The three main measures are mode, median and mean.
三个主要度量是众数、中位数和平均数。
The mode is the most frequent value; the median is the middle value when data are ordered; the mean is the sum divided by the number of values.
众数是出现频率最高的数值;中位数是数据排序后的中间值;平均数是总和除以数值个数。
For n ordered values, the median is at position (n + 1) ÷ 2.
对于 n 个已排序数值,中位数的位置是 (n + 1) ÷ 2。
For grouped data, the mean is estimated by Σfx ÷ Σf, where f is frequency and x is the class midpoint.
对于分组数据,平均数用 Σfx ÷ Σf 估计,其中 f 是频数,x 是组中值。
Mean = Σfx ÷ Σf
The median is often preferred when data contain outliers because it is not affected by extreme values.
当数据含有异常值时,通常优先使用中位数,因为它不受极端值影响。
5. Measures of Dispersion | 离散程度度量
Range is the difference between the largest and smallest values.
极差是最大值与最小值之差。
Interquartile range (IQR) = upper quartile – lower quartile, measuring the spread of the middle 50% of the data.
四分位距(IQR)= 上四分位数 – 下四分位数,衡量中间 50% 数据的离散情况。
Variance and standard deviation measure how far the values are from the mean.
方差和标准差衡量各数值与平均数的偏离程度。
For a population of size n, variance is σ² = Σ(x – μ)² ÷ n, and standard deviation is σ = √[Σ(x – μ)² ÷ n].
Box plots are useful for comparing distributions and identifying outliers.
箱线图有助于比较分布和识别异常值。
7. Probability Rules | 概率规则
Probability measures the chance that an event occurs, with 0 ≤ P(A) ≤ 1.
概率度量事件发生的可能性,满足 0 ≤ P(A) ≤ 1。
If all outcomes are equally likely, P(A) = number of favourable outcomes ÷ total number of outcomes.
若所有结果等可能,则 P(A) = 有利结果数 ÷ 总结果数。
For any event A, P(A’) = 1 – P(A), where A’ is the complement of A.
对于任意事件 A,P(A’) = 1 – P(A),其中 A’ 是 A 的补事件。
For any two events, P(A or B) = P(A) + P(B) – P(A and B).
对于任意两个事件,P(A 或 B) = P(A) + P(B) – P(A 与 B)。
If A and B are mutually exclusive, then P(A and B) = 0.
若 A 和 B 互斥,则 P(A 与 B) = 0。
If A and B are independent, then P(A and B) = P(A) × P(B).
若 A 和 B 独立,则 P(A 与 B) = P(A) × P(B)。
8. Probability Trees and Venn Diagrams | 概率树与维恩图
Tree diagrams multiply along branches for successive independent events and add for mutually exclusive paths.
树状图在连续独立事件中沿分支相乘,在互斥路径间相加。
Conditional probability is written P(A|B) = P(A and B) ÷ P(B).
条件概率写作 P(A|B) = P(A 与 B) ÷ P(B)。
Venn diagrams illustrate unions, intersections and complements visually.
维恩图直观展示并集、交集和补集。
When using a tree diagram, the probabilities on each set of branches must sum to 1.
使用树状图时,每组分支上的概率之和必须等于 1。
9. The Binomial Distribution | 二项分布
A binomial experiment has a fixed number n of independent trials, each with two outcomes labelled success (probability p) and failure (probability q = 1 – p).
二项试验有固定的 n 次独立试验,每次试验只有两个结果:成功(概率 p)和失败(概率 q = 1 – p)。
The probability of exactly r successes is:
恰好 r 次成功的概率为:
P(X = r) = ⁿCᵣ pʳ qⁿ⁻ʳ
The mean of a binomial distribution is np, and the variance is npq.
二项分布的平均数为 np,方差为 npq。
Use the binomial formula when trials are independent and the probability of success remains constant.
当各次试验独立且成功概率保持不变时,使用二项公式。
10. Scatter Diagrams and Correlation | 散点图与相关性
A scatter diagram shows the relationship between two variables.
散点图展示两个变量之间的关系。
Correlation describes the strength and direction of a linear relationship: positive, negative or none.
相关性描述线性关系的强度和方向:正相关、负相关或无相关。
Spearman’s rank correlation coefficient rₛ measures monotonic correlation using ranks.
斯皮尔曼等级相关系数 rₛ 通过排序来衡量单调相关。
rₛ = 1 – (6Σd²) ÷ [n(n² – 1)]
Here d is the difference between the paired ranks, and n is the number of data pairs.
其中 d 是成对等级的差值,n 是数据对的数量。
A value close to +1 indicates strong positive correlation; a value close to -1 indicates strong negative correlation.
数值接近 +1 表示强正相关;接近 -1 表示强负相关。
11. Regression Lines | 回归直线
The line of best fit is drawn to model the linear relationship
Published by TutorHao | IGCSE 统计 Revision Series | aleveler.com
Preparing for Cambridge IGCSE Statistics is not about reading notes over and over. It is a deliberate cycle of diagnosing weak areas, doing targeted practice, and applying knowledge to past papers. A clear timetable prevents last-minute cramming, and the right strategies turn marks lost through careless habits into marks gained through exam technique.
Start by downloading the latest Cambridge IGCSE Statistics syllabus, usually coded 0479. Highlight the two main question styles you will face: short structured questions and longer data-response questions. The syllabus lists topic areas such as data collection, representation, averages, dispersion, probability, correlation and regression, sampling, and time series. Knowing the exact scope prevents you from studying unnecessary material or missing an entire topic.
Confirm whether you are entered for Core or Extended. Your school will tell you which paper you will sit. Each paper has a different duration and total mark, but the planning principle is the same: build knowledge first, then speed, then accuracy under exam pressure. Write a one-page topic checklist from the syllabus and date each topic when you first review it and again when you master it.
In the first week, take one complete past paper under timed conditions. Do not worry about the score. Mark the paper and classify every error into four types: concept gap, calculator mistake, misreading the question, or weak interpretation. This diagnostic gives you a baseline and shows you which type of practice will make the biggest difference.
For example, if you lose most marks because you used the population standard deviation when a sample was required, that is a concept and formula issue. If you used the correct method but entered the wrong list into your calculator, that is a calculator habit. Treating these differently is the essence of smart revision.
A 12-week plan fits most school calendars and gives enough time to cover every topic without panic. Weeks 1 and 2 cover data handling and representation; weeks 3 and 4 cover averages and dispersion; week 5 covers probability; week 6 covers correlation and regression; week 7 covers sampling and data collection; weeks 8 and 9 combine mixed-topic practice; weeks 10 and 11 are full past papers; and week 12 is error-log review plus a final mock.
Within each week, aim for four or five short study sessions of 45 to 60 minutes rather than one long session. A simple rhythm such as review the previous error log for 10 minutes, learn or practise one topic for 30 minutes, then attempt a short exam question for 15 minutes keeps revision active.
In Cambridge IGCSE Statistics, descriptive statistics and data representation usually carry the largest share of marks. Topics such as cumulative frequency, histograms, box-and-whisker plots, mean, median, quartiles, range, and standard deviation appear in almost every series. Probability and bivariate data are also frequent, especially tree diagrams, expected frequency, scatter graphs, correlation, and line of best fit.
Allocate about 60 percent of your topic-revision time to these core areas before spending time on less common topics such as index numbers or moving averages. Use the mark scheme from one recent paper to calculate how many marks each topic was worth. This turns your priority list from guesswork into evidence.
5. Master Calculator and Formula Skills | 熟练计算器与公式技能
Statistical papers reward fluency with your scientific calculator. Practise entering data lists, finding the mean and standard deviation, and checking whether your calculator displays the population standard deviation or the sample standard deviation. In many IGCSE Statistics questions, you need the sample standard deviation when data come from a sample, but always follow the wording of the syllabus and the question.
population standard deviation σ = √[Σ(x – μ)² ÷ n] | sample standard deviation s = √[Σ(x – x̄)² ÷ (n – 1)]
Know the meaning of each symbol so you do not mix them up. Use brackets carefully on your calculator, especially when dividing by n or by n – 1. A common mistake is typing the square root before the division when the whole fraction should be under the root. Practise with small data sets first, then check that your calculator gives the same answer as a hand calculation.
理解每个符号的含义,避免混淆。在计算器上要小心使用括号,尤其是除以 n 或除以 n – 1 的时候。一个常见错误是在除法之前输入平方根,而实际上整个分数都应该在根号下。先用小数据集练习,然后检查计算器的结果是否与手算一致。
6. Use Past Papers in Phases | 分阶段使用历年真题
Use past papers in three phases. In Phase A, work through questions untimed and open-book, sorting them by topic. Focus on method marks and correct working. In Phase B, do single questions or sections with a timer, aiming for accurate calculations and clear layout. In Phase C, complete full papers under exam conditions and mark them against the official mark scheme.
分三个阶段使用历年真题。阶段 A:不限时、开卷,按主题整理
Published by TutorHao | IGCSE 统计 Revision Series | aleveler.com
📚 IGCSE Cambridge Statistics: Key Points for Practical and Experimental Assessment | IGCSE 剑桥统计:实验/实践考核要点
This guide summarises the key assessment points for the practical and experimental side of Cambridge IGCSE Statistics (0479). Although the qualification is mainly examined through written papers, the questions are deliberately practical: they require you to plan a statistical investigation, collect or interpret data, choose suitable diagrams, carry out calculations, and evaluate findings. Mastering these skills is essential for high marks.
1. Understand the Statistical Enquiry Cycle | 理解统计探究循环
Practical work in statistics is not just about calculating numbers. Examiners expect you to follow the statistical enquiry cycle: pose a question, plan and collect data, process and present data, interpret results, and evaluate the whole process. When answering a practical-style question, first identify which stage is being tested.
2. Planning an Investigation and Stating a Hypothesis | 规划调查与陈述假设
A good practical task starts with a clear aim and a testable hypothesis, for example ‘Older customers spend longer in the shop than younger customers.’ Avoid vague aims such as ‘I want to find out about shopping.’ You should state the population, the variables to be measured, and how the data will be collected. If the task is experimental, also identify the independent, dependent and control variables.
Sampling is often a key mark area. You must know simple random, systematic, stratified, cluster, quota and convenience sampling. Be able to explain advantages and disadvantages, and justify a choice for a given context. For a representative sample, random or stratified methods are usually preferred; quota and convenience sampling are often biased but quicker and cheaper.
Divide into groups, then sample proportionally 分组后按比例抽样
Represents key groups 代表关键组别
More complex to organise 组织更复杂
Systematic 系统
Select every kth member 每隔 k 个抽一个
Simple to use 使用简单
May miss a periodic pattern 可能错过周期性模式
Cluster 整群
Randomly select whole clusters 随机选整个群体
Cheap for scattered populations 适合分散总体
High sampling variance 抽样方差较大
Quota 配额
Choose fixed numbers per category 每类定额选取
Quick and inexpensive 快速且成本低
Interviewer bias possible 可能产生访问员偏差
Convenience 便利
Choose easy-to-reach people 选容易接触的人
Very cheap 成本很低
Not representative 无代表性
4. Designing Data Collection Instruments | 设计数据收集工具
Questionnaires, observation sheets and experiments must be designed carefully. Questions should be clear, unbiased, and not leading. For example, avoid ‘Do you agree that the new service is excellent?’ Instead ask ‘How would you rate the new service: excellent, good, fair, poor?’ Include a pilot survey to identify problems before the main data collection. Closed questions are easier to process; open questions give more detail but are harder to analyse.
5. Types of Data and Levels of Measurement | 数据类型与测量层次
You must distinguish categorical (qualitative) data from numerical (quantitative) data, and discrete from continuous variables. Also know nominal, ordinal, interval and ratio levels, though IGCSE usually focuses on qualitative, discrete and continuous data. Choose a suitable chart based on data type: bar chart or pie chart for categorical data; histogram for continuous grouped data; scatter diagram for bivariate data.
After collecting raw data, you need to organise it into frequency tables, grouped frequency tables, or two-way tables. Diagrams must include clear titles, labelled axes, and a key if needed. For histograms, use frequency density, not raw frequency, when class widths are unequal. For cumulative frequency graphs, plot points at upper class boundaries.
Frequency density = frequency ÷ class width | 频率密度 = 频数 ÷ 组距
7. Descriptive Statistics: Averages and Spread | 描述统计:平均数和离散程度
You must be able to calculate the mean, median, mode, range, quartiles, interquartile range, and standard deviation where required. Choose the best measure: the median and IQR are resistant to outliers; the mean and standard deviation use all data but are affected by extreme values. Show working clearly; in practical questions, interpret what a large or small IQR tells you about consistency.
8. Bivariate Data: Correlation and Regression | 双变量数据:相关与回归
For paired data, plot a scatter diagram and describe the relationship by direction (positive or negative), form (linear or non-linear) and strength (strong, moderate, weak). Do not confuse correlation with causation. If a line of best fit is drawn, use it to make predictions only within the range of the data (interpolation); extrapolation outside the range is unreliable.
Probability questions often involve relative frequency from an experiment, expectation, sample space diagrams, tree diagrams, or two-way tables. When using experimental data, relative frequency is an estimate of probability and becomes more stable as the number of trials increases. For equally likely outcomes, use P(A) = number of favourable outcomes ÷ total number of outcomes.
10. Interpreting Results and Drawing Conclusions | 解释结果并得出结论
A conclusion must relate back to the original hypothesis or aim. State whether the data supports or does not support the hypothesis, and quote supporting figures, for example ‘The median waiting time for branch A was 4.2 minutes, compared with 6.8 minutes for branch B, which supports the claim that branch A is faster.’ Avoid overclaiming; sample results are evidence, not proof.
结论必须回扣原假设或目的。说明数据支持或不支持假设,并引用支持性数据,例如“A 分店等待时间中位数为 4.2 分钟,B 分店为 6.8 分钟,这支持 A 分店更快的说法”。避免过度断言;样本结果只是证据,不是证明。
11. Evaluating Limitations and Suggesting Improvements | 评估局限性并提出改进建议
High-scoring practical answers always evaluate the investigation. Consider sample size, sampling method, non-response, measurement error, and confounding variables. Suggest specific improvements, for example ‘Use a larger stratified sample by age group
Published by TutorHao | IGCSE 统计 Revision Series | aleveler.com
This guide explains how Cambridge IGCSE Statistics answers are marked and how to write responses that gain full credit. It covers command words, method and accuracy marks, rounding, diagram drawing, probability, and common errors.
Cambridge questions use command words such as “calculate”, “describe”, “compare”, “interpret”, “estimate”, “explain” and “justify”. Each word tells you how much working and what style of answer is expected. For example, “calculate” means show your method and give an exact or suitably rounded answer, while “describe” means state the trend, shape or features shown by a graph or data set.
You should underline the command word and any key conditions in the question. This prevents you from calculating when the examiner asks for a comparison, or describing when the examiner asks for a calculation. Misreading the command word is one of the most common causes of lost marks.
Questions that say “state” or “write down” require no working and often carry only a quick mark. Questions that say “show that” require every step so the examiner can follow your reasoning. Adjust the detail of your answer to the command word used.
2. How Marks Are Awarded: M, A, B and CAO | 评分方式:方法分、准确分与独立分
In a Cambridge IGCSE Statistics mark scheme, marks are usually split into method marks (M), accuracy marks (A) and independent marks (B). Method marks are earned for using a correct process, even if the final answer is wrong. Accuracy marks require the correct answer or a correct answer following an earlier error.
Some marks are labelled “cao”, meaning correct answer only. “ft” means follow through: if you use an earlier incorrect value in a correct way, you can still receive the mark. “oe” means or equivalent, “SC” means special case, and “isw” means ignore subsequent working. Knowing these codes helps you understand why some partially correct answers still score.
📚 IGCSE Cambridge Statistics: Full Syllabus Breakdown | IGCSE 剑桥统计:课程大纲全面解析
Cambridge IGCSE Statistics (0479) gives learners a practical introduction to collecting, presenting, analysing and interpreting data. This article breaks down the full syllabus, assessment structure and core skills to help you plan revision and focus on the areas that matter most.
Statistics is not just about numbers; it trains you to make decisions under uncertainty. The Cambridge IGCSE Statistics syllabus develops skills in data handling, graphical methods, probability modelling and critical interpretation.
The main assessment objectives are: AO1 knowledge and understanding of statistical techniques, AO2 application of statistical methods to problems, and AO3 interpretation and evaluation of statistical results. You should be able to choose the correct technique, perform calculations, and comment on reliability, bias and limitations.
Cambridge IGCSE Statistics is assessed through two written papers. Both papers allow calculators and cover the full syllabus, so there is no Core or Extended tier.
Both papers assess the same content, so you should not leave any topic out. Past paper practice is essential because the questions often combine two or three syllabus areas in one context.
You will study primary and secondary data, questionnaires, and sampling methods such as random, stratified, systematic and quota sampling. You must be able to judge reliability, bias and the suitability of a data source.
Primary data is collected first-hand for a specific purpose. | 一手数据是为特定目的直接收集的数据。
Secondary data already exists and may be cheaper but less controlled. | 二手数据已经存在,可能成本更低但控制更弱。
Stratified sampling keeps the same population proportions in the sample. | 分层抽样保持样本中总体比例不变。
Systematic sampling selects members at regular intervals from an ordered list. | 系统抽样从有序名单中按固定间隔选取成员。
Quota sampling is non-random and can easily introduce interviewer bias. | 配额抽样是非随机的,容易引入调查者偏差。
4. Data Representation and Diagrams | 数据表示与图表
Candidates should construct and interpret diagrams: pictograms, bar charts, pie charts, histograms, frequency polygons, cumulative frequency curves, stem-and-leaf diagrams, box plots and scatter diagrams.
For histograms with unequal class widths, frequency density is used rather than raw frequency.
对于组距不等的直方图,应使用频数密度,而不是原始频数。
Frequency density = Frequency ÷ Class width
You should also know how to read median, quartiles and percentiles from a cumulative frequency curve, and how to interpret box plots for skew and spread.
The mean, median and mode summarise the centre of a data set. Weighted mean and geometric mean may appear for grouped data, index numbers or rates of change.
For grouped data, use the midpoint of each class as x. The median is useful when data is skewed, while the mode is the only average for qualitative data.
Range, interquartile range, percentiles, variance and standard deviation measure spread. A small standard deviation means data is clustered close to the mean; a large one means it is widely spread.
σ = √(Σ(x − μ)² / n) for a population | s = √(Σ(x − x̄)² / (n − 1)) for a sample
Remember that the interquartile range covers the middle 50% of data and is resistant to outliers, whereas the range is strongly affected by extreme values.
记住四分位距覆盖中间 50% 的数据,不受异常值影响;而极差受极端值影响很大。
7. Probability Basics | 概率基础
Probability measures how likely an event is. You must handle mutually exclusive events, independent events, conditional probability, tree diagrams and Venn diagrams.
概率用于衡量事件发生的可能性。你需要掌握互斥事件、独立事件、条件概率、树状图和维恩图。
P(A ∪ B) = P(A) + P(B) − P(A ∩ B) | P(A ∩ B) = P(A) × P(B) for independent events | P(A|B) = P(A ∩ B) / P(B)
Mutually exclusive events cannot happen at the same time, so P(A ∩ B) = 0. Conditional probability questions often require you to reduce the sample space after an event has occurred.
互斥事件不能同时发生,因此 P(A ∩ B) = 0。条件概率题通常需要在事件发生后缩小样本空间。
8. Probability Distributions | 概率分布
A discrete random variable has a probability mass function. The binomial distribution models n independent trials with two outcomes; the normal distribution models continuous data with mean μ and standard deviation σ.
离散随机变量具有概率质量函数。二项分布对 n 次独立、两结果试验建模;正态分布对均值为 μ、标准差为 σ 的连续数据建模。
For a binomial distribution, mean is np and variance is npq. For the normal distribution, you must be confident using the standard normal table or calculator inverse normal functions.
Scatter diagrams show relationships between two variables. You may calculate Pearson’s product-moment correlation coefficient r and Spearman’s rank correlation coefficient, and use the least squares regression line y = a + bx.
散点图展示两个变量之间的关系。你可能需要计算皮尔逊积矩相关系数 r、斯皮尔曼等级相关系数,并使用最小二乘回归直线 y = a + bx。
r = Sxy / √(Sxx × Syy) | b = Sxy / Sxx | a = ȳ − bx̄
Correlation measures strength and direction of a linear relationship, but it does not prove causation. Extrapolation beyond the data range can be unreliable.
相关性衡量线性关系的强度和方向,但相关性不代表因果关系。超出数据范围的外推可能不可靠。
10. Time Series and Index Numbers | 时间序列与指数
Time series analysis includes trend, seasonal variation, moving averages and forecasting. Index numbers compare prices or quantities over time, often using a base period of 100.
Moving averages smooth out short-term fluctuations and help reveal the underlying trend. Seasonal variation can be estimated by subtracting the moving average from the actual value.
移动平均可以消除短期波动并揭示潜在趋势。季节性变动可通过实际值减去移动平均来估计。
11. Sampling Distributions and Inference | 抽样分布与推断
Advanced questions may involve the sampling distribution of the mean, standard error and confidence intervals for a population mean. This connects sample statistics to population parameters.
Standard error = σ / √n | 95% confidence interval for μ: x̄ ± 1.96 × σ / √n
A larger sample size reduces the standard error, so the confidence interval becomes narrower. You should interpret a confidence interval in terms of repeated sampling, not as a probability statement about one interval.
Show working clearly, label axes on diagrams, use exact calculator values during intermediate steps, and always check units. Common errors include using the ungrouped mean formula for grouped data, confusing independent and mutually exclusive, and misreading cumulative frequency scales.
📚 IGCSE CCEA Statistics: How UK University Entry Requirements Compare | IGCSE CCEA 统计:英国大学申请要求对照
Statistics is often treated as a supporting subject at GCSE/IGCSE, but it can strengthen a university application in data-rich fields. This article maps CCEA GCSE Statistics against common UK university entry requirements and explains how you can use it strategically in your application.
1. What is CCEA GCSE Statistics? | CCEA GCSE 统计学概览
CCEA GCSE Statistics develops skills in collecting, presenting and interpreting data. The syllabus includes averages, dispersion, correlation, probability, distributions, sampling and basic hypothesis testing.
A key feature of the course is its real-world focus: students learn how data are used in business, health, sport and government, rather than only manipulating algebraic expressions.
This formula for the sample mean is typical of the calculations CCEA Statistics students must interpret, not just compute.
这个样本平均数公式是 CCEA 统计学学生不仅需要计算、更需要解读的典型计算之一。
2. How UK Universities Treat GCSE Statistics | 英国大学如何看待 GCSE 统计学
Most UK universities do not list GCSE Statistics as a separate entry requirement. Their standard conditions usually specify GCSE Mathematics, and often GCSE English, with a minimum grade such as C, C* or B depending on the course and institution.
Statistics is therefore best understood as an additional qualification. It does not replace Mathematics, but it can reinforce a candidate’s quantitative profile.
因此,统计学最好被理解为一门附加资格。它不能替代数学,但可以增强申请者的定量能力背景。
3. The Difference Between GCSE Mathematics and GCSE Statistics | GCSE 数学与 GCSE 统计学的区别
GCSE Mathematics is generally compulsory and is used by universities to check core numeracy, algebra and problem-solving. GCSE Statistics is optional and focuses on data handling, probability and inference.
Because universities already require Mathematics, a high grade in Statistics is rarely a substitute for a low grade in Mathematics. It works best when it sits alongside a strong Maths result.
4. Subjects Where GCSE Statistics Gives an Edge | GCSE 统计学能带来优势的学科
Statistical thinking is increasingly important across many degree programmes. A strong CCEA Statistics grade can signal readiness for quantitative methods in the following areas:
Economics: data interpretation and econometric-style thinking
Psychology: research methods, significance testing and experimental design
Geography and environmental science: spatial data and climate statistics
Biology and medicine: clinical trials, risk and evidence evaluation
Business and management: market research, finance and decision-making
Data science and actuarial science: probability models and inference
经济学:数据解读与计量经济学式思维
心理学:研究方法、显著性检验与实验设计
地理与环境科学:空间数据与气候统计
生物与医学:临床试验、风险与证据评估
商业与管理:市场研究、金融与决策
数据科学与精算学:概率模型与推断
In these subjects, admissions tutors often view a good Statistics grade as evidence that you can handle numerical evidence rather than just abstract equations.
在这些学科中,招生导师通常认为良好的统计学成绩证明你能够处理数字证据,而不仅仅是抽象方程式。
5. Typical UK University GCSE Requirements by Subject Area | 英国大学各学科 GCSE 要求对照
Requirements vary by institution and year, so always check the specific university website. The table below gives a general guide to how GCSE Mathematics requirements and GCSE Statistics relevance compare.
This guide breaks down the essential terms in the CCEA IGCSE Statistics specification into quick, memorable clusters. Use the paired definitions and memory hooks to revise actively before your exam.
Population means the entire set of individuals or items that you want to study. A sample is a smaller group selected from the population. A census collects data from every member of the population.
总体指你想研究的全部个体或项目。样本是从总体中选出的较小群体。普查则收集总体中每一个成员的数据。
A parameter is a numerical summary of a population, while a statistic is a numerical summary calculated from a sample. Raw data are unprocessed values before being organised into tables or charts.
参数是总体的数值概括,而统计量是从样本计算出的数值概括。原始数据是尚未整理成表格或图表的原始数值。
Memory hook: ‘Population = whole pie, sample = one slice, census = eat the whole pie.’
记忆线索:’总体是整块饼,样本是一块切片,普查是吃掉整块饼。’
2. Types of Data | 数据类型
Qualitative data describe qualities or categories, such as colour, gender, or type of transport. Quantitative data are numerical and can be either discrete or continuous.
定性数据描述性质或类别,例如颜色、性别或交通方式。定量数据是数值型数据,可以是离散型或连续型。
Discrete data can only take certain values, usually counted, such as the number of cars in a car park. Continuous data can take any value within a range, usually measured, such as height or time.
Primary data are collected by you or your team for a specific purpose, such as a questionnaire, interview, or experiment. Secondary data are data that already exist, such as government reports, textbooks, or websites.
Common primary collection tools include questionnaires, interviews, observations, and experiments. Each has strengths: questionnaires reach many people quickly, while interviews allow deeper follow-up.
A pilot survey is a small trial run of a questionnaire used to identify unclear or biased questions before the main data collection.
试点调查是问卷的小规模试运行,用于在正式收集数据前发现不清晰或有偏差的问题。
4. Sampling Techniques | 抽样方法
Random sampling gives every member of the population an equal chance of selection, which helps reduce bias. Stratified sampling divides the population into groups called strata and samples proportionally from each group.
Systematic sampling selects every nth item after a random starting point. Cluster sampling selects whole groups or clusters at random. Quota sampling fills fixed numbers from subgroups but is not random.
系统抽样在随机起点后每隔 n 个抽取一个。整群抽样随机选取整个群体。配额抽样按固定人数从子群中选取,但不是随机抽样。
Convenience sampling uses people who are easy to reach, such as friends or people in the same street, and often introduces bias. A sampling frame is a list of all members of the population from which a sample can be drawn.
Mean is the sum of all values divided by the number of values. It is calculated as:
平均数是所有数值之和除以数值个数。计算公式为:
Mean: x̄ = Σx / n
Median is the middle value when data are ordered from smallest to largest. Mode is the most frequent value or category.
中位数是将数据从小到大排列后的中间值。众数是出现频率最高的值或类别。
For grouped data, the modal class is the class with the highest frequency, and the mean can be estimated using midpoints. The median can be read from a cumulative frequency curve.
对于分组数据,众数组是频数最高的组,平均数可用组中点进行估算。中位数可从累积频数曲线中读取。
The mean is sensitive to outliers, while the median is more robust. Choose the median when data are skewed or contain extreme values.
平均数对异常值敏感,而中位数更具稳健性。当数据偏斜或含有极端值时,应选择中位数。
6. Measures of Spread | 离散程度度量
Range = largest value – smallest value. It is quick to calculate but affected by outliers. Interquartile range (IQR) = upper quartile Q₃ – lower quartile Q₁, and it measures the spread of the middle 50% of data.
Percentiles divide ordered data into 100 equal parts. The lower quartile Q₁ is the 25th percentile, the median is the 50th percentile, and the upper quartile Q₃ is the 75th percentile.
This guide links the core skills in the CCEA Statistics specification to the demands of international mathematics and data competitions. It focuses on statistical reasoning, efficient calculation, and clear communication under time pressure.
1. Understand the CCEA Specification and Competition Overlap | 熟悉 CCEA 考纲与竞赛交叉点
CCEA Statistics tests data collection, averages, spread, charts, probability, bivariate data, and simple inference. International competitions rarely ask for definitions alone; they combine these tools in unfamiliar, multi-step contexts.
A good starting point is to list every CCEA topic and mark whether you can apply it to a modelling or puzzle question. If a topic only works in textbook exercises, practise it with competition-style follow-up questions.
International competition questions often value insight over calculation. For example, you might be given a misleading average and asked to explain why the median is better. This is exactly the kind of judgement CCEA exam questions reward.
Multi-stage tree and conditional logic / 多阶段树图和条件逻辑
Bivariate data / 双变量数据
Correlation vs causation arguments / 相关性与因果性论证
2. Master Data Types and Sampling Methods | 掌握数据类型与抽样方法
Competitions often hide a sampling error in a realistic scenario. You need to recognise whether data are categorical or quantitative, and whether quantitative data are discrete or continuous.
Know that random sampling reduces selection bias but does not remove non-response bias. A large sample does not automatically fix a biased sampling method.
要知道随机抽样能减少选择偏差,但不能消除无回答偏差。大样本并不会自动修复一个有偏差的抽样方法。
Stratified sampling keeps important groups represented in the correct proportion. The formula is used frequently in competition questions that ask for a sample allocation.
分层抽样能让重要群体按正确比例被代表。竞赛题中经常要求计算样本分配,公式使用频率很高。
Stratified sample from group = (group size ÷ total population) × total sample size
For example, if 120 of 600 students are in Year 10 and a stratified sample of 50 is needed, the Year 10 sample size is (120 ÷ 600) × 50 = 10.
例如,如果 600 名学生中有 120 名在 Year 10,需要抽取 50 人的分层样本,那么 Year 10 的样本人数是 (120 ÷ 600) × 50 = 10。
3. Descriptive Statistics: Centre and Spread | 描述统计:集中趋势与离散程度
The mean, median and mode measure centre. The range, interquartile range and standard deviation measure spread. Competition questions often ask which measure is most appropriate, not just how to calculate it.
Sample standard deviation s = √(Σ(x − x̄)² ÷ (n − 1))
A competition trick is to give raw data with an extreme value. The mean shifts toward the outlier, while the median stays stable. Use median and interquartile range for skewed distributions.
Always ask: is the variable skewed? Income, house prices and reaction times often need median and IQR. Symmetric data allow the mean and standard deviation to summarise well.
A good diagram communicates shape, centre, spread and outliers. A poor diagram hides them. In competitions, you may need to choose the best chart for a given data story or criticise a misleading graph.
Histograms use area for frequency. The height of a bar is frequency density, not frequency. This is one of the most common competition errors.
直方图用面积表示频数。条形的高度是频率密度,而不是频数。这是竞赛中最常见的错误之一。
Frequency density = frequency ÷ class width
Cumulative frequency diagrams give the median, lower quartile and upper quartile from the graph. Box plots then show the five-number summary and expose outliers.
累积频率图可以从图中读出中位数、下四分位数和上四分位数。箱线图则展示五数概括,并能揭示离群值。
Chart / 图表
Best for / 适用场景
Bar chart / 条形图
Comparing categories / 比较类别
Histogram / 直方图
Continuous grouped data / 连续分组数据
Box plot / 箱线图
Comparing distributions and outliers / 比较分布和离群值
Cumulative frequency graph / 累积频率图
Finding quartiles and percentiles / 求四分位数和百分位数
5. Probability for Competition Problems | 竞赛中的概率问题
Competition probability problems require careful sample spaces. Write down the sample space or draw a tree before applying formulas. This prevents double-counting and forgotten branches.
Independent events satisfy P(A ∩ B) = P(A) × P(B). Mutually exclusive events satisfy P(A ∩ B) = 0. Do not confuse the two ideas.
独立事件满足 P(A ∩ B) = P(A) × P(B)。互斥事件满足 P(A ∩ B) = 0。不要把这两个概念混淆。
Expected value is the long-run average. A fair game has expected value zero after the stake is included. Use the weighted formula:
期望值是长期平均结果。如果计入赌注后期望值为零,就是公平游戏。使用加权公式:
E(X) = Σx · P(X = x)
6. Bivariate Data and Correlation | 双变量数据与相关性
Scatter graphs show whether two variables move together. Correlation measures strength and direction, but it does not prove causation. A competition answer that claims causation without evidence will lose marks.
For a line of best fit, plot the mean point (x̄, ȳ) because the regression line passes through it. The regression equation has the form:
画最佳拟合线时,要标出平均点 (x̄, ȳ),因为回归线经过该点。回归方程的形式为:
y = a + bx
Spearman’s rank correlation is used when data are ranks or when the relationship is monotonic but not linear. Its formula is:
当数据是等级数据,或关系单调但非线性时,使用斯皮尔曼等级相关。其公式为:
rₛ = 1 − (6Σd²) ÷ (n(n² − 1))
Beware extrapolation: predicting far outside the data range is invalid. A strong correlation within the observed range does not mean the trend continues forever.
要警惕外推:预测远超数据范围的值是不可靠的。即使在观测范围内有强相关,也不意味着趋势会永远持续。
7. Statistical Inference and Margin of Error | 统计推断与误差范围
Competition questions may ask you to compare two groups from sample data. Always comment on both centre and spread, not just one number. A comparison based only on means can be misleading.
If a reported difference is smaller than the margin of error, it may not be meaningful. Competitions reward students who recognise this rather than overclaiming a result.