📚 Sampling | 抽样
Sampling is a core topic in statistics that underpins how we gather data and make reliable inferences. In the Edexcel A-Level Mathematics specification, you are expected to understand various sampling techniques, recognise potential sources of bias, and select the most appropriate method for a given situation. This article provides a comprehensive overview of sampling, from basic definitions to practical exam tips.
抽样是统计学的核心主题,支撑着我们收集数据并做出可靠推断的方式。在Edexcel A-Level数学考试大纲中,你需要理解各种抽样技术,识别潜在的偏差来源,并为给定情况选择最合适的方法。本文提供了从基本定义到实用考试技巧的全面概述。
1. Populations and Samples | 总体与样本
A population is the entire set of individuals, items, or measurements that you want to study. A sample is a subset of that population, selected to represent the whole. For example, if you are investigating the average test score of Year 12 students in England, the population is all Year 12 students in England, and a sample might be 500 students chosen from various schools.
总体是你想研究的全部个体、项目或测量值的集合。样本是从总体中选出具有代表性的子集。例如,如果你正在调查英格兰12年级学生的平均考试成绩,总体就是所有英格兰12年级学生,而样本可能是从不同学校选出的500名学生。
The process of drawing a sample from a population is called sampling. Whether the sample is truly representative determines the validity of any conclusions you draw about the population.
从总体中抽取样本的过程称为抽样。样本是否真正具有代表性决定了你所得到的关于总体的任何结论的有效性。
2. Why Do We Sample? | 我们为什么要抽样?
Investigating an entire population is often impractical due to constraints of time, cost, and accessibility. Sampling allows researchers to collect data efficiently and still make reliable inferences, provided the sample is well chosen. Moreover, in some cases, testing destroys the item (for example, strength‐testing light bulbs), so sampling is a necessity.
由于时间、成本和可及性的限制,研究整个总体通常是不切实际的。抽样使研究人员能够高效地收集数据,并仍然做出可靠的推断,前提是样本选取得当。此外,在某些情况下,测试会破坏物品(例如灯泡的强度测试),因此抽样必不可少。
However, sampling introduces sampling error — the natural variation that occurs because only part of the population is observed. Careful design helps minimise this error.
然而,抽样会引入抽样误差——由于只观察到总体的一部分而产生的自然变异。精心的设计有助于将这种误差降至最低。
3. Sampling Frame | 抽样框
A sampling frame is a list or database of all the units in the population from which the sample is drawn. Examples include a school register, a customer mailing list, or a map of postcode sectors. The quality of the sampling frame directly affects the representativeness of the sample.
抽样框是列出总体中所有单位的一份清单或数据库,样本就是从该清单中抽取的。例子包括学校学生名册、客户邮寄名单或邮政编码分区地图。抽样框的质量直接影响样本的代表性。
If the sampling frame is incomplete — for instance, a phone book that excludes people without landlines — the sample will suffer from undercoverage bias and may not reflect the true population.
如果抽样框不完整——例如,电话簿排除了没有固定电话的人——样本就会存在覆盖不足的偏差,可能无法反映真实总体。
4. Simple Random Sampling (SRS) | 简单随机抽样
In a simple random sample, every member of the population has an equal probability of being selected, and each possible combination of n members has the same chance of becoming the sample. This is usually implemented by assigning a number to each population unit and using a random number generator or drawing lots.
在简单随机抽样中,总体的每个成员都有相等的被选中的概率,并且每个可能的n个成员组合都有相同的机会成为样本。这通常通过给每个总体单位分配一个号码,并使用随机数生成器或抽签来实现。
SRS is free from selection bias, but it can be hard to achieve when the population is large and geographically dispersed. It also requires a complete and accurate sampling frame.
简单随机抽样没有选择偏差,但当总体很大且地理分布分散时很难实现。它还需要一个完整且准确的抽样框。
5. Stratified Sampling | 分层抽样
Stratified sampling divides the population into mutually exclusive groups, called strata, based on a characteristic such as gender, age group, or income level. A simple random sample is then taken from each stratum, usually in proportion to the stratum’s size relative to the population.
分层抽样是根据性别、年龄段或收入水平等特征,将总体划分为互斥的群体(称为层)。然后从每个层中抽取简单随机样本,通常按照该层相对于总体的大小按比例进行。
If a population of 1000 students is 60% male and 40% female, and you want a sample of 100, you would select 60 males and 40 females randomly from their respective strata. The formula for the stratum sample size is nₕ = (Nₕ / N) × n, where Nₕ is the stratum size, N the total population, and n the total sample size.
如果一个有1000名学生的总体中60%是男生、40%是女生,而你要抽取一个容量为100的样本,那么你应该从相应的层中随机选择60名男生和40名女生。层样本大小的计算公式为 nₕ = (Nₕ / N) × n,其中Nₕ为该层大小,N为总体大小,n为样本总大小。
6. Systematic Sampling | 系统抽样
Systematic sampling selects units at regular intervals from an ordered list. First, a sampling interval k is determined by dividing the population size N by the desired sample size n (k = N/n, rounded if necessary). Then, a random starting point r between 1 and k is chosen, and every k-th unit thereafter is included.
系统抽样是从有序列表中按固定间隔选择单位。首先,用总体大小N除以所需样本大小n来确定抽样间隔k(k = N/n,必要时四舍五入)。然后,从1到k之间随机选择一个起始点r,此后每隔k个单位选取一个。
This method is easier and faster than SRS when a physical list is available, but it can introduce periodicity bias if the list order coincides with a pattern in the variable of interest.
当有实物清单可用时,此方法比简单随机抽样更简单快捷,但如果列表顺序与所研究变量的某种模式一致,可能引入周期性偏差。
7. Quota Sampling | 配额抽样
Quota sampling is a non‑probability method in which the researcher decides on quotas for specific subgroups (e.g., 40 males and 40 females) and then finds individuals to fill those quotas using convenience or judgement. It is often used in market research interviews on the street.
配额抽样是一种非概率方法,研究者确定特定子群体的配额(例如40名男性和40名女性),然后通过便利或判断找到个体来填满这些配额。它常用于街头的市场调研采访。
Quota sampling is quick, cheap, and does not require a sampling frame, but it is prone to interviewer bias because the choice within each quota is not random. The resulting sample may not be representative.
配额抽样快速、成本低且不需要抽样框,但很容易引入访问员偏差,因为每个配额内的选择不是随机的。因而得到的样本可能不具有代表性。
8. Opportunity (Convenience) Sampling | 机会抽样(便利抽样)
Opportunity sampling uses individuals who are readily available and willing to take part. For example, surveying the first 20 customers leaving a shop or asking classmates sitting nearby to complete a questionnaire.
机会抽样利用那些容易接触到并愿意参与的个体。例如,调查最先离开商店的20名顾客,或请坐在附近同学填写问卷。
It is the easiest sampling technique to carry out, but it is also the most susceptible to bias because the sample is unlikely to reflect the wider population. Conclusions drawn from such samples should be treated with extreme caution.
这是最容易实施的抽样技术,但也最容易受到偏差的影响,因为样本不太可能反映更广泛的总体。从这种样本中得出的结论应极其谨慎地对待。
9. Bias in Sampling | 抽样偏差
Bias occurs when a sampling method systematically over‑ or under‑represents certain parts of the population. Common sources include using a flawed sampling frame, self‑selection (voluntary response), convenience sampling, and non‑response, where selected individuals refuse to participate.
当抽样方法系统性地过度代表或不足代表总体的某些部分时,就会产生偏差。常见来源包括使用有缺陷的抽样框、自选(自愿回应)、便利抽样以及无回应——即被选中的个体拒绝参与。
Bias leads to estimates that are consistently too high or too low, and it cannot be reduced simply by increasing the sample size. Choosing a suitable probability‑based method and a good sampling frame is critical for reducing bias.
偏差会导致估计值始终偏高或偏低,而且不能仅靠增大样本量来减少。选择适当的基于概率的方法和良好的抽样框对于减少偏差至关重要。
10. Choosing the Right Sampling Method | 选择合适的抽样方法
The choice of sampling method depends on the research objective, available resources, time constraints, and whether a complete sampling frame exists. Probability methods (SRS, stratified, systematic) are preferred for inferential statistics because they allow the calculation of margins of error. Non‑probability methods (quota, opportunity) are sometimes used for exploratory studies or when a sampling frame is impossible to obtain.
抽样方法的选择取决于研究目标、可用资源、时间限制以及是否存在完整的抽样框。概率方法(简单随机抽样、分层抽样、系统抽样)更适合推断统计,因为它们允许计算误差范围。非概率方法(配额抽样、机会抽样)有时用于探索性研究或无法获得抽样框的情况。
| Method | Key Feature | Advantages | Disadvantages |
|---|---|---|---|
| Simple Random | Equal probability for all | Unbiased, easy to analyse | Needs full frame, may miss small subgroups |
| Stratified | Proportional selection from subgroups | Ensures representation, more precise | Requires knowledge of strata sizes, complex |
| Systematic | Every k‑th unit | Simple to implement, spread over population | Danger of periodicity, needs ordered list |
| Quota | Fills pre‑set numbers for subgroups | Quick, no frame needed | Non‑random selection, high bias risk |
| Opportunity | Whoever is available | Easiest and cheapest | Very biased, poor representation |
In an exam, you may be asked to recommend and justify a method. For instance, when studying dietary habits across different age groups, stratified sampling is often ideal because you can ensure each age band is properly represented.
在考试中,你可能会被要求推荐一种方法并给出理由。例如,在研究不同年龄组的饮食习惯时,分层抽样通常是理想的选择,因为它能确保每个年龄段都得到适当代表。
11. Sample Size Considerations | 样本大小的考量
A larger sample generally yields more precise estimates and reduces the standard error. However, doubling the sample size does not halve the standard error; the precision improves with the square root of n. Therefore, beyond a certain point, increasing the sample size gives only marginal gains and becomes more expensive.
较大的样本通常能产生更精确的估计值并减少标准误差。然而,样本量翻倍并不会使标准误差减半;精度随着 n 的平方根提高。因此,超过一定限度后,增加样本量只能带来微小的收益,并且成本更高。
When designing a study, the necessary sample size is calculated based on the desired margin of error, the variability in the population, and the confidence level (typically 95%). For Edexcel A‑Level, you are not required to perform these calculations, but you should understand the qualitative trade‑offs.
在设计研究时,所需的样本量是根据期望的误差范围、总体的变异性以及置信水平(通常为95%)来计算的。对于Edexcel A‑Level,你不需要进行这些计算,但应理解其中的定性权衡。
12. Summary and Exam Tips | 总结与考试提示
Sampling is about choosing a representative subset from a population so that valid conclusions can be drawn. Master the definitions, strengths and weaknesses of the five main sampling methods: simple random, stratified, systematic, quota and opportunity. Be prepared to explain why a particular method is suitable for a given scenario and how to avoid bias.
抽样就是从一个总体中选出一个具有代表性的子集,以便得出有效的结论。掌握五种主要抽样方法的定义、优点和缺点:简单随机抽样、分层抽样、系统抽样、配额抽样和机会抽样。准备好解释为什么某种方法适合某个给定场景,以及如何避免偏差。
When tackling exam questions, always state clearly why the chosen method reduces bias or is practical. Note if a sampling frame is available, mention it. If the question highlights practical constraints, non‑probability methods may be acceptable but you must acknowledge their limitations.
在应对考试题目时,始终清楚说明所选方法为何能减少偏差或为何切实可行。如果有抽样框,要提及。如果题目强调了实际限制,非概率方法可能可以接受,但你必须承认它们的局限性。
Finally, remember that no sample is perfect, but good sampling design gives us the best chance of inferring the truth about the population we care about.
最后,请记住,没有任何样本是完美的,但良好的抽样设计能为我们推断所关注的总体真相提供最好的机会。
Published by TutorHao | Mathematics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply