Populations and Samples: Clarifying Core Statistical Concepts | 总体与样本的基本概念辨析

📚 Populations and Samples: Clarifying Core Statistical Concepts | 总体与样本的基本概念辨析

In statistics, the terms ‘population’ and ‘sample’ form the foundation of all inferential reasoning. Yet these concepts are frequently misunderstood or used interchangeably by students, leading to errors in exam questions and real-world data analysis alike. This article aims to clarify these essential definitions, distinguish related terms, and explore how sampling works in practice.

在统计学中,”总体”与”样本”这两个术语构成了所有推断性推理的基础。然而,学生们经常误解或混用这些概念,导致在考试题目和实际数据分析中出错。本文旨在厘清这些基本定义、区分相关术语,并探讨抽样在实际中的应用方式。


1. Population: The Complete Set | 总体:完整的集合

A population in statistics refers to the entire set of individuals, objects, or measurements of interest in a particular study. It is not limited to people; it can be animals, plants, manufactured items, or even experimental outcomes. For example, if we wish to study the height of all students in a school, the population consists of every student enrolled at that school, not merely those we can conveniently measure.

统计学中的总体,是指在某一特定研究中所有感兴趣的个体、物体或测量值的完整集合。它不仅限于人,也可以是动物、植物、工业产品甚至实验结果。例如,如果我们想研究某所学校所有学生的身高,总体就包括该学校在册的每一名学生,而不仅仅是那些我们方便测量的学生。

An important distinction is that ‘population’ denotes a complete enumeration of every element that satisfies the study’s definition. The key word is ‘complete’. If any element is missing, the set no longer qualifies as the population. For example, studying the average lifespan of light bulbs produced by a factory on a specific day requires every bulb from that day’s production line.

一个重要的区别在于,”总体”表示对满足研究定义的每一个元素的完整列举。关键词是”完整”。如果缺少任何元素,该集合就不再符合总体的定义。例如,研究某工厂某一天生产的灯泡的平均使用寿命,需要当天生产线上生产的每一只灯泡。


2. Sample: A Representative Subset | 样本:有代表性的子集

A sample is a subset of the population that is selected for actual observation or measurement. Because examining an entire population is often impractical, too expensive, or even destructive, we rely on samples to draw conclusions about the population from which they are drawn. For instance, quality control engineers do not test every biscuit produced; they sample a few batches and infer the quality of the entire production.

样本是从总体中选取的、用于实际观察或测量的一个子集。由于检查整个总体往往不切实际、成本过高甚至具有破坏性,我们依靠样本来推断其所来自总体的特征。例如,质量监控工程师不会测试生产的每一块饼干,而是抽取若干批次进行检验,再推断整个生产批次的质量。

The critical requirement for a sample is that it must be representative of the population. A representative sample mirrors the essential characteristics of the population, such as proportion by gender, age distribution, or socioeconomic background, so that conclusions drawn from the sample can be generalised back to the population with confidence.

对样本的关键要求是:它必须能够代表总体。一个有代表性的样本应当反映总体的基本特征,如性别比例、年龄分布或社会经济背景,从而使从样本得出的结论能够可靠地推广到总体。


3. Individual and Sample Size: Precise Terminology | 个体与样本量:精准的术语

An individual (also called a sampling unit or elementary unit) is a single member of the population. In a survey of household income, each household is an individual unit. In a study of blood pressure medication, each patient is an individual. The term ‘individual’ does not necessarily mean a human being; it simply refers to the smallest unit being studied.

个体(也称为抽样单元或基本单元)是总体中的一个单一成员。在家庭收入调查中,每个家庭是一个个体单位;在降压药研究中,每位患者是一个个体。”个体”这一术语并不一定指人,它仅仅指被研究的最小单位。

Sample size, often denoted by the letter n, is the number of individuals in a sample. A common misconception is that a larger sample size guarantees representativeness. While larger samples reduce sampling error, they do not eliminate bias if the sample is not randomly selected. Conversely, a small but carefully chosen random sample can be more reliable than a large but biased one.

样本量(通常用字母 n 表示)是指样本中所包含的个体数量。一个常见的误解是:样本量越大就越有代表性。诚然,较大的样本可以减少抽样误差,但如果样本不是随机选取的,它并不能消除偏差。反之,一个精心挑选的小型随机样本,可能比一个庞大但带有偏差的样本更可靠。

Population → Sample → Sample Size (n)
总体 → 样本 → 样本量 (n)


4. Parameter vs. Statistic: Measuring the Right Set | 参数与统计量:测量正确的集合

A parameter is a numerical characteristic that describes a population. Parameters are usually denoted by Greek letters. For instance, the population mean is written as μ (mu) and the population standard deviation as σ (sigma). The true value of a parameter is fixed but unknown unless we conduct a complete census of the population.

参数是描述总体特征的数值,通常用希腊字母表示。例如,总体均值写作 μ(mu),总体标准差写作 σ(sigma)。除非我们对总体进行完整的普查,否则参数的真实值是固定的但未知的。

A statistic is a numerical characteristic computed from a sample. Statistics are typically denoted by Roman letters, such as x̄ (sample mean) or s (sample standard deviation). We use statistics as estimates of parameters. This distinction is fundamental: a parameter belongs to the population; a statistic belongs to the sample. Failing to distinguish between the two leads to confusion about what our calculations actually describe.

统计量是根据样本计算得出的数值特征,通常用罗马字母表示,例如 x̄(样本均值)或 s(样本标准差)。我们用统计量来估计参数。这一区分是根本性的:参数属于总体,统计量属于样本。如果混淆这二者,就会对我们的计算结果究竟描述了什么感到困惑。

Concept 概念 Symbol 符号 Source 来源
Population Mean 总体均值 μ Population 总体
Sample Mean 样本均值 Sample 样本
Population Standard Deviation 总体标准差 σ Population 总体
Sample Standard Deviation 样本标准差 s Sample 样本

5. Why We Sample: Practical Necessity | 为何抽样:现实必要性

There are several compelling reasons to use samples rather than conducting a full census. First, cost: collecting data from an entire population can be prohibitively expensive. Second, time: a complete enumeration may take so long that the results become obsolete before the process is finished. Third, destruction: in some quality tests, the very act of measurement destroys the item being tested, making a full census impossible.

使用样本而非进行全面普查,有着充分且现实的原因。第一是成本:从整个总体收集数据可能极其昂贵;第二是时间:完整的查点可能需要很长时间,以至于结果在完成之前就已过时;第三是破坏性:在某些质量测试中,测量行为本身就会损坏被测物品,使得全面普查根本无法进行。

Additionally, a well-designed sample can actually produce more accurate results than a poorly executed census. A carefully controlled sample allows researchers to devote more time and resources to ensuring measurement accuracy and reducing non-sampling errors, such as recording mistakes or respondent misunderstandings, which can plague large-scale censuses.

此外,一个设计良好的样本,实际上可能比执行不好的普查产生更准确的结果。精心控制的样本允许研究者投入更多时间和资源来确保测量的准确性,并减少非抽样误差(如记录错误或受访者误解),而这些误差往往困扰着大规模普查。


6. Random Sampling: The Gold Standard | 随机抽样:黄金标准

Random sampling is a method of selecting a sample in which every member of the population has a known and equal probability of being chosen. This principle eliminates selection bias and ensures, in the long run, that the sample is representative of the population. Simple random sampling can be implemented using random number tables, software generators, or lottery methods.

随机抽样是这样一种抽样方法:总体中的每个成员都有已知且相等的概率被选中。该原则排除了选择偏差,并确保在长期来看样本能够代表总体。简单随机抽样可以通过随机数表、软件生成器或抽签法来实现。

In exam contexts, students often need to describe how to take a simple random sample. The standard procedure involves: first, assigning a unique number to each member of the population; second, using a random mechanism to select numbers; and third, including in the sample the individuals corresponding to those selected numbers. It is important to sample without replacement to ensure each individual appears at most once.

在考试情境中,学生经常需要描述如何进行简单随机抽样。标准程序包括:第一步,为总体中的每个成员分配一个唯一的编号;第二步,使用随机机制来选取号码;第三步,将对应这些号码的个体纳入样本。我们应当采用不放回抽样,以确保每个个体至多出现一次。


7. Other Sampling Methods Compared | 其他抽样方法的比较

While simple random sampling is conceptually elegant, other methods are often used in practice. Systematic sampling involves selecting every k-th member of the population after a random starting point. Stratified sampling divides the population into homogeneous subgroups (strata) and then randomly samples from each stratum proportionally. Cluster sampling divides the population into clusters, randomly selects some clusters, and measures every member within selected clusters.

虽然简单随机抽样在概念上很优雅,但实践中常使用其他方法。系统抽样是在随机起点之后选取总体中的第 k 个成员;分层抽样将总体划分为若干同质子组(层),然后按比例从每一层中随机抽样;整群抽样则将总体划分为若干群,随机选取若干群,并对被选中的群内所有成员进行测量。

Each method has advantages and trade-offs. Stratified sampling ensures representation of all subgroups and yields smaller sampling error for the same sample size. However, it requires knowledge of the population structure. Cluster sampling is cost-efficient when the population is geographically dispersed, but it tends to produce larger sampling errors. Understanding these trade-offs is essential for exam questions that ask you to evaluate or choose an appropriate sampling method.

每种方法都有其优势和权衡。分层抽样确保所有子组都得到代表,在相同样本量下产生更小的抽样误差;但它要求了解总体结构。当总体地理分布分散时,整群抽样成本效益高,但它往往产生较大的抽样误差。理解这些权衡取舍,对于要求评估或选择合适的抽样方法的考试题目至关重要。


8. Bias and Representativeness: Avoiding Pitfalls | 偏差与代表性:避免陷阱

Sampling bias occurs when certain members of the population are systematically more likely to be included in the sample than others. A classic example is a telephone poll that excludes households without phones. Self-selection bias arises when individuals volunteer to participate. Convenience sampling, selecting whichever individuals are easiest to reach, is particularly prone to bias and should be avoided in rigorous statistical studies.

当总体中的某些成员比其他成员更有可能被纳入样本时,就会出现抽样偏差。一个经典例子是电话民意调查,它排除了没有电话的家庭。自我选择偏差在个体自愿参与时出现。便利抽样是指选择最容易接触到的个体,这种方法特别容易产生偏差,在严谨的统计研究中应当避免。

A representative sample is one that accurately reflects the characteristics of the population. However, even a perfectly random sample will exhibit some natural variation from the population, known as sampling error. This is not a mistake; it is an expected consequence of examining only a subset. Sampling error can be reduced by increasing the sample size, but it can never be eliminated entirely unless a full census is conducted.

有代表性的样本是指能够准确反映总体特征的样本。然而,即使是完全随机的样本,也会与总体之间存在一些自然差异,这被称为抽样误差。这不是错误;它是仅考察一个子集所产生的预期结果。抽样误差可以通过增加样本量来减少,但除非进行全面普查,否则永远无法完全消除。


9. Common Exam Mistakes | 常见考试错误辨析

One frequent error is describing the population incorrectly. For example, when studying the reading habits of teenagers in a city, the population is all teenagers in that city, not all people in the city. Another common mistake is confusing a parameter with a statistic: writing ‘the sample mean μ’ is incorrect; it should be written as x̄. The symbol μ is reserved exclusively for the population parameter.

一个常见错误是错误地描述总体。例如,在调查某城市青少年的阅读习惯时,总体是该城市的所有青少年,而不是该城市的所有人。另一个常见错误是混淆参数与统计量:写”样本均值 μ”是不对的,应当写作 x̄。符号 μ 仅专门用于表示总体参数。

Students also frequently misunderstand what a sample actually represents. A sample is a tool for inference, not the object of interest itself. If a question asks about the average height of all trees in a forest, and we measure 50 trees, our sample statistics describe the 50 measured trees, but our conclusions should be drawn about the entire forest. The wording must always distinguish between describing the sample and inferring about the population.

学生还经常误解样本实际上代表了什么。样本是推断的工具,而不是兴趣对象本身。如果一道题目询问森林中所有树木的平均高度,而我们测量了 50 棵树,那么我们的样本统计量描述的是这 50 棵被测量的树,但我们的结论应针对整片森林。措辞必须始终区分”描述样本”与”推断总体”。


10. Worked Examples: Applying the Concepts | 实例演练:应用概念

Example 1: A quality inspector wants to estimate the mean diameter of bolts produced in a day. There are 2,000 bolts total. She randomly selects 40 bolts and measures them.

例 1:一名质量检验员想估计一天生产的螺栓的平均直径。螺栓总数为 2,000 只。她随机选取了 40 只螺栓进行测量。

Here, the population is all 2,000 bolts. The sample is the 40 bolts selected. The sample size n = 40. If she calculates the mean diameter of these 40 bolts, she obtains a statistic x̄. The unknown true mean of all 2,000 bolts is the parameter μ. Her goal is to use x̄ to estimate μ.

在此,总体是全部 2,000 只螺栓;样本是被选出的 40 只螺栓;样本量 n = 40。如果她计算这 40 只螺栓的平均直径,得到的是统计量 x̄。全部 2,000 只螺栓的未知真实均值就是参数 μ。她的目标是用 x̄ 来估计 μ。

Example 2: A school has 800 students: 350 in Year 12 and 450 in Year 13. A researcher wants to survey 80 students, ensuring proportional representation from both year groups.

例 2:一所学校有 800 名学生:350 名在十二年级,450 名在十三年级。一位研究者想调查 80 名学生,并确保两个年级按比例得到代表。

Stratified sampling is appropriate here. From Year 12: (350 ÷ 800) × 80 = 35 students. From Year 13: (450 ÷ 800) × 80 = 45 students. Within each stratum, the researcher selects students at random. This ensures that the sample’s year-group proportions match the population.

此时宜采用分层抽样。来自十二年级:(350 ÷ 800) × 80 = 35 名学生;来自十三年级:(450 ÷ 800) × 80 = 45 名学生。在每个层内,研究者随机选取学生。这确保样本中年级比例与总体一致。


11. Symbol Summary and Memory Aids | 符号总结与记忆技巧

To keep these concepts clear, remember that Greek letters describe populations and Roman letters describe samples. A helpful mnemonic: ‘Greeks own the whole; Romans own a part.’ When you see μ or σ, think of the entire group. When you see x̄ or s, think of the selected subset.

为了保持概念清晰,请记住:希腊字母描述总体,罗马字母描述样本。一个有用的助记方法:”希腊人拥有整体;罗马人拥有部分。” 当你看到 μ 或 σ,想到的是整个群体;当你看到 x̄ 或 s,想到的是被选中的子集。

Another critical distinction concerns notation: the population variance is σ², while the sample variance is s². When computing sample variance, we divide by (n − 1) rather than n. This small adjustment, called Bessel’s correction, prevents the sample variance from underestimating the population variance, an important detail that frequently appears in examination marks schemes.

另一个关键区别涉及符号:总体方差写作 σ²,而样本方差写作 s²。计算样本方差时,我们除以 (n − 1) 而不是 n。这一小的调整被称为贝塞尔校正,它防止样本方差低估总体方差,这是一个在考试评分标准中经常出现的细节。


12. Conclusion: Building a Solid Foundation | 结论:奠定坚实基础

The distinction between population and sample is not merely a matter of vocabulary; it shapes how we define research questions, select data collection methods, and interpret results. A clear grasp of parameters versus statistics, and of representativeness versus bias, is essential for success in statistics examinations and for credible data analysis beyond the classroom. Every time you encounter a data set, ask yourself: Is this the complete population, or a sample drawn from it? The answer determines the meaning and validity of every conclusion you draw.

总体与样本的区别不只是词汇问题;它决定了我们如何界定研究问题、选择数据收集方法和解读结果。清晰掌握参数与统计量的区别、代表性与偏差的区别,对于统计学考试的成功以及课堂之外可信的数据分析至关重要。每当你面对一组数据时,问问自己:这是完整的总体,还是从总体中抽取的样本?这个答案决定了你所下每一个结论的意义和有效性。

Published by TutorHao | Mathematics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Exit mobile version