📚 Year 10 CIE Statistics: Key Points for Experimental/Practical Assessment | Year 10 CIE 统计:实验/实践考核要点
In the CIE IGCSE Statistics syllabus, the practical or experimental aspect of data collection is not just an exercise in gathering numbers; it is a structured process that requires careful planning, ethical consideration, and an understanding of how bias can arise. Mastering these key points will help you design reliable investigations, whether you are conducting a survey at school, observing traffic flow, or testing a hypothesis about consumer preferences. This guide walks you through the essential skills you need for your practical assessment.
在 CIE IGCSE 统计课程中,数据收集的实践或实验环节并不仅仅是收集数字的练习,而是一个需要周密计划、伦理考量并理解偏差如何产生的结构化过程。掌握以下要点能帮助你设计出可靠的调查,无论你是在学校进行问卷调查、观察交通流量还是检验有关消费者偏好的假设。本指南将带你梳理实践考核所需的关键技能。
1. Planning a Statistical Investigation | 规划统计调查
Every practical task begins with a clearly defined aim. You need to state exactly what you want to find out and, if the investigation is looking for a relationship, formulate a simple hypothesis. For example, ‘Students who study more hours per night get higher test scores’ is a testable hypothesis.
每一项实践任务都始于一个明确的目标。你需要准确陈述你想探究的问题,如果调查寻找某种关系,就要提出一个简单的假设。例如,“每晚学习时间更长的学生考试成绩更高”就是一个可检验的假设。
Next, identify the variables you will measure. The explanatory (independent) variable is what you change or group by, such as the number of study hours, while the response (dependent) variable is the outcome, like test scores. Decide whether these variables are categorical or numerical. This choice affects how you will collect and later analyse the data.
接着,识别你将测量的变量。解释变量(自变量)是你改变或分组依据的变量,比如学习时数;响应变量(因变量)则是结果,例如考试分数。判断这些变量是分类的还是数值型的,这一选择会影响你后续收集和分析数据的方式。
Finally, define your target population and, if possible, a sampling frame – a list of all members in that population from which you will draw your sample. Without a sampling frame, you risk omitting important subgroups.
最后,界定你的目标总体,如果可能的话,还要确定一个抽样框——即总体所有成员的名单,你将从该名单中抽取样本。没有抽样框,你就可能遗漏重要的子群体。
2. Designing a Questionnaire | 设计问卷
A good questionnaire uses a mix of question types. Closed questions offer a limited set of answers (Yes/No, multiple choice) and are easy to code, while open questions allow respondents to express opinions in their own words. Pre-coded response boxes speed up data entry but should always include an ‘Other’ option if the list is not exhaustive.
一份好的问卷会混合使用多种问题类型。封闭式问题提供有限答案(是/否、选择题),易于编码;开放式问题则允许受访者用自己的语言表达观点。预先编码的答题框能加快数据录入,但如果选项列表不完全,应始终包含“其他”选项。
Avoid leading questions such as ‘Don’t you agree that homework is beneficial?’ They push respondents towards a particular answer. Also avoid ambiguous wording and double-barrelled questions (e.g. ‘Do you enjoy maths and science?’) which combine two issues into one.
避免引导性问题,例如“你难道不认为家庭作业是有益的吗?”这类问题会诱导受访者偏向某一答案。还要避免模糊的措辞和双重问题(如“你喜欢数学和科学吗?”),它们将两个议题混为一谈。
Arrange questions in a logical order. Start with simple, non-sensitive questions to build trust, and place demographic questions towards the end. A pilot study will help you spot confusing items before the main data collection.
按照逻辑顺序排列问题。从简单、不敏感的问题入手建立信任,将人口统计类问题放在末尾。试点研究能帮助你在主要数据收集前发现令人困惑的题目。
3. Sampling Methods | 抽样方法
Simple random sampling gives every member of the population an equal chance of being selected, often using random number tables or a calculator’s random number generator. It is unbiased but requires a complete sampling frame and can be impractical for large, dispersed populations.
简单随机抽样让总体中的每个成员都有相等的机会被选中,常用随机数表或计算器的随机数生成器来实现。这种方法无偏向,但需要一个完整的抽样框,对于庞大、分散的总体可能不太可行。
Stratified sampling divides the population into distinct groups (strata) based on a characteristic such as age or gender, then takes a random sample from each stratum. The size of each stratum’s sample is proportional to its share of the population:
nₕ = (Nₕ ÷ N) × n
where Nₕ is the size of the stratum, N is the total population, and n is the overall sample size. This method ensures underrepresented groups are included.
分层抽样根据年龄或性别等特征将总体划分为不同的层,然后从每一层中随机抽取样本。每层的样本大小与该层在总体中所占比例成比例:
nₕ = (Nₕ ÷ N) × n
其中 Nₕ 为该层大小,N 为总体大小,n 为总样本量。这种方法能确保代表人数不足的群体也被包含在内。
Systematic sampling selects every k‑th member from a list, where k = N ÷ n. It is quick and easy but can introduce periodicity bias if the list has a hidden pattern. Quota and convenience sampling are non‑probability methods often used in market research but are prone to interviewer bias and cannot be used to make reliable statistical inferences.
系统抽样从名单中每隔 k 人抽取一个样本,k = N ÷ n。这种方法快速简便,但若名单存在隐藏的周期模式,则可能引入周期性偏差。配额抽样和便利抽样是常用于市场研究的非概率方法,但容易受访员偏差影响,且不能用来进行可靠的统计推断。
4. Avoiding Bias and Sources of Error | 避免偏差与误差源
Selection bias occurs when the sampling frame does not accurately represent the population. For example, using a telephone directory to sample households misses those without landlines. Always check your frame for under‑coverage and over‑coverage.
当抽样框不能准确代表总体时,就会出现选择性偏差。例如,使用电话簿抽取住户样本会遗漏没有固定电话的家庭。务必核查抽样框是否存在涵盖不足或过度涵盖的问题。
Non‑response bias is a serious issue in questionnaire surveys. If people who refuse to answer differ systematically from those who respond, your results will be skewed. Aim for a high response rate by making surveys short, sending reminders, and ensuring anonymity.
无回应偏差是问卷调查中一个严重的问题。如果拒绝作答的人与回答的人存在系统性差异,结果就会发生偏斜。应通过缩短调查长度、发送提醒和保证匿名来提高回复率。
Measurement error arises from poorly worded questions, inaccurate measuring instruments, or misunderstandings during interviews. Pilot testing and careful training of interviewers can minimise these errors.
测量误差来源于措辞不当的问题、测量工具不精确或访谈中的误解。试点测试和对访员的仔细培训可以尽量减少这类误差。
5. Data Collection in Practice | 实践中的数据收集
Face‑to‑face interviews give high control over the process, allow you to clarify questions, and often yield richer data, but they are time‑consuming and expensive. Telephone interviews are quicker but may have a lower response rate.
面访能让你对过程高度控制,可以澄清问题,并常常获得更丰富的数据,但耗时长且成本高。电话访谈速度更快,但回复率可能较低。
Self‑completion questionnaires, whether distributed on paper, by post, or online, are cheap and allow anonymity, but you cannot probe for explanations. Online surveys are popular but risk excluding those without internet access.
自填式问卷,无论是通过纸质、邮寄还是网上分发,成本低廉且能保护匿名性,但无法追问解释。在线调查很流行,但有将没有互联网的人排除在外的风险。
Observation is useful when you cannot question subjects, such as counting vehicles at a junction. Decide on systematic timing and precisely define what you are counting to avoid observer drift.
当无法向对象提问时,观察法就很有用,例如在路口统计车辆。要确定系统化的时间安排,并精确定义计数内容,以避免观察者漂移。
6. Conducting a Pilot Study | 进行试点研究
A pilot study is a small‑scale trial run of your data collection process. Its main purpose is to identify flaws in the questionnaire, check that instructions are clear, and estimate the time required. Skipping this step often leads to unusable data.
试点研究是对数据收集过程的小规模预演。其主要目的是发现问卷中的缺陷,检查说明是否清晰,并估算所需时间。跳过这一步常常导致收集到的数据无法使用。
Select a small sample that is similar to your target population but not part of the final sample. After collecting pilot data, review every question: are there any that everybody skips or misunderstands? Use the feedback to reword or remove problematic items.
选择一个与目标总体相似但不属于最终样本的小样本。收集到试点数据后,检查每一个问题:有没有所有人都跳过或误解的地方?根据反馈对存在问题的条目进行改写或删除。
Also test your data entry and coding scheme. If respondents frequently choose ‘Other’ on a multiple‑choice question, your answer categories may need expanding.
同时也要测试数据录入和编码方案。如果受访者频繁在选择题中选择“其他”,你的答案类别可能就需要扩充。
7. Recording and Organising Data | 记录与整理数据
Design a data collection sheet before you start. It should have columns for every variable, with clear headings and space for notes. Using a tally chart is an effective way to record categorical data as it is collected, and then convert tallies into frequencies.
开始调查前先设计好数据收集表。表中应为每个变量设置一列,标题清晰,并留有备注空间。使用划记表是在收集过程中记录分类数据的有效方式,之后再将划记转换为频数。
For numerical data, decide on appropriate class intervals if you plan to group them. Classes must not overlap and should cover the full range of possible values. Use 5–9, 10–14, not 5–10, 11–15 if you want to avoid gaps.
对于数值型数据,如果打算分组,就要确定合适的组距。组与组之间不得重叠,并应覆盖所有可能取值的范围。分组时可采用 5–9, 10–14 这样的写法,以避免出现空缺。
Coding involves assigning numbers to non‑numerical responses, e.g. 1 for ‘Male’, 2 for ‘Female’. This speeds up computer analysis but keep a codebook so you do not forget what each number represents.
编码是指为非数值回答分配数字,例如 1 表示“男性”,2 表示“女性”。这能加快计算机分析的速度,但要保留一份编码手册,以免遗忘每个数字所代表的含义。
8. Using Random Numbers for Sampling | 使用随机数进行抽样
To select a simple random sample, number every member of your sampling frame uniquely. Then use a random number table or a calculator: read across rows, and ignore numbers that are out of range or repeated.
要抽取一个简单随机样本,先给抽样框中的每个成员分配一个独一无二的编号。然后使用随机数表或计算器:横向逐行读取,忽略超出范围或重复出现的数字。
A typical approach is to decide on the number of digits needed (e.g. if your population is 200, use 3‑digit numbers from 001 to 200). Generate as many valid numbers as your sample size. Remember that random does not mean haphazard – you must follow a predetermined rule.
典型的做法是确定所需位数(例如总体为 200,则使用 001 到 200 的三位数)。生成与样本量相等的有效数字。请记住,随机并不意味着随意——你必须遵守预先设定的规则。
For systematic sampling, calculate the sampling interval
k = N ÷ n
and pick a random starting point between 1 and k. Then select every k‑th unit. This ensures spread across the whole list.
系统抽样时,计算抽样间距
k = N ÷ n
,并在 1 到 k 之间随机选一个起点。然后每隔 k 个单位抽取一个样本。这能确保样本遍布整个名单。
9. Ensuring Reliability and Validity | 确保信度与效度
Reliability refers to the consistency of your measurements. If you repeated the survey under the same conditions, you should get similar results. Standardising the wording of questions, training interviewers, and using calibrated instruments all improve reliability.
信度指的是测量结果的一致性。如果在相同条件下重复进行调查,应得到相似的结果。将问题措辞标准化、培训访员以及使用校准过的仪器,都能提高信度。
Validity is about whether you are measuring what you intended to measure. A question about ‘time spent reading’ might not capture true reading habits if students include social media reading. Check that your questions have face validity, and consider comparing with existing records if possible.
效度涉及你是否测量到了本来打算测量的内容。一个关于“阅读时间”的问题,如果学生把读社交媒体也算进去,就可能无法捕捉真实的阅读习惯。应检查问题是否具有表面效度,如有可能,可与现有记录进行对照。
In practical assessments, you should document every step you take to enhance reliability and validity. This shows the examiner that you understand the importance of rigorous methodology.
在实践考核中,你应当记录为提高信度和效度所采取的每一个步骤。这能向考官表明你理解严谨方法的重要性。
10. Evaluating the Data Collection Process | 评估数据收集过程
After completing your investigation, critically evaluate the entire process. Which sources of bias might have affected your data? Were there problems with the sampling frame, response rate, or measurement technique? Be honest about the limitations.
完成调查后,要对整个过程进行批判性评估。哪些偏差来源可能影响了你的数据?抽样框、回复率或测量技术是否存在问题?要诚实地面对这些局限。
Suggest specific improvements. For instance, if your response rate was low, you could recommend using incentives or changing the mode of delivery. If you used a convenience sample, acknowledge that you cannot generalise the findings to the wider population.
提出具体的改进建议。例如,如果回复率低,你可以建议采用激励措施或改变发放方式。如果使用的是便利样本,就要承认结论不能推广到更广泛的人群。
Finally, link your evaluation back to the original aim: did the data allow you to answer the research question reliably? A thoughtful evaluation often earns higher marks than a flawless but unexamined set of results.
最后,将评估与最初目标联系起来:数据是否让你能可靠地回答研究问题?深思熟虑的评估往往比一组完美但未经审视的结果更能获得高分。
Published by TutorHao | Statistics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply