📚 Part 3 Skills and Fieldwork Investigation | 第三部分 技能与实地调查
In any real-world statistical investigation, the ability to design a study, collect data properly and analyse it with appropriate mathematical techniques is just as important as mastering the underlying theory. This part of the course focuses on the practical skills needed to carry out fieldwork and to make valid inferences from the data gathered. Whether you are estimating a population mean from a sample or testing a claim about a proportion, the methods introduced here will help you turn raw observations into reliable conclusions.
在任何真实的统计调查中,设计研究、正确收集数据并使用恰当的数学方法进行分析的能力,与掌握基础理论同样重要。本部分课程着眼于开展实地调查所需的实践技能,以及如何从收集到的数据中做出有效的推断。无论你是通过样本估计总体均值,还是检验关于某个比例的论断,这里介绍的方法都将帮助你从原始观测中提炼出可靠的结论。
1. Designing a Fieldwork Investigation | 设计实地调查
Every successful statistical investigation begins with a clear objective. You need to define the population of interest, decide what you are trying to measure, and formulate a research question or a pair of hypotheses. Without a well-defined aim, data collection can quickly become unstructured and bias may be introduced.
每一项成功的统计调查都始于一个明确的目标。你需要界定感兴趣的总体,决定要测量什么,并形成一个研究问题或一对假设。如果没有明确的目标,数据收集很快就会变得杂乱无章,而且可能引入偏差。
A pilot study is often a wise first step. By trying out your data collection instruments on a small scale, you can check whether your questions are understood, whether your measuring equipment is reliable, and whether your sampling frame is realistic. The pilot helps you refine your approach before committing to the full fieldwork.
预调查通常是明智的第一步。通过在小范围内测试你的数据收集工具,你可以检查问题是否被理解、测量设备是否可靠,以及抽样框架是否切合实际。预调查有助于在投入全面实地调查之前改进你的方法。
You also need to choose the type of data you will collect. Continuous variables such as height or time can be measured on a scale, while discrete variables like number of people in a household are counted. Categorical data, which fall into groups such as eye colour or type of transport used, require different summary techniques.
你还需要选择将要收集的数据类型。连续型变量(如身高或时间)可以用标尺度量,而离散型变量(如家庭人口数)则通过计数获得。分类数据(例如眼睛颜色或使用的交通方式)属于不同的组别,需要不同的汇总技术。
2. Sampling Methods for Fieldwork | 实地调查的抽样方法
It is rarely possible to measure every member of a population, so you must select a sample. Simple random sampling ensures that every individual has an equal chance of being chosen, and it removes personal judgement from the selection process. Random number tables or a calculator’s random integer function can be used to pick members from a numbered list.
要测量总体中的每一个成员通常是不可能的,因此你必须选择一个样本。简单随机抽样确保每个个体都有均等的机会被选中,并且排除了选择过程中的个人判断。可以使用随机数表或计算器的随机整数功能从编号列表中挑选成员。
Stratified sampling is particularly useful when the population divides naturally into subgroups, such as year groups in a school. By taking a sample proportional to the size of each stratum, you guarantee that the sample represents the structure of the population more precisely than a simple random sample might.
当总体自然地划分为若干子群时,例如学校中的年级组,分层抽样特别有用。通过按各层大小比例抽取样本,你可以保证样本比简单随机样本更准确地代表总体结构。
Systematic sampling chooses individuals at regular intervals from a list, such as every tenth house on a street. While easy to implement, it can introduce bias if there is an underlying pattern in the ordering. Cluster sampling, often used in large-scale geographical surveys, selects whole groups, but analysis must account for the clustered design.
系统抽样按固定间隔从列表中选取个体,例如一条街上每第十栋房屋。虽然实施简便,但如果排序中存在潜在模式,就可能引入偏差。整群抽样常用于大规模地理调查,选取整个小组,但分析时必须考虑整群设计。
Opportunity (or convenience) sampling, where you simply ask whoever is available, is common in exploratory fieldwork but may not give a representative sample. You must always be honest about the limitations of your sampling method when drawing conclusions.
机会抽样(或便利抽样)简单地询问当时在场的人,这种方法在探索性实地调查中很常见,但未必能提供有代表性的样本。在得出结论时,你必须诚实地说明所用抽样方法的局限。
3. Collecting Data and Minimising Bias | 收集数据与减少偏差
The design of a questionnaire or recording sheet can heavily influence the responses you obtain. Questions should be neutral, avoiding leading language that suggests a particular answer. Closed questions with fixed response options are easier to code and analyse, whereas open questions can provide richer detail but are harder to summarise numerically.
问卷或记录表的设计会严重影响你得到的回答。问题应当保持中性,避免使用暗示特定答案的引导性语言。带有固定选项的封闭式问题更易于编码和分析,而开放式问题可以提供更丰富的细节,但难以用数字概括。
Data must be recorded in an organised manner, ideally using a spreadsheet with clear column headings and consistent coding. In fieldwork, you might also need to record environmental conditions, the time of day, or the exact location, as these contextual variables can help explain variation in your measurements.
数据必须以有组织的方式记录,最好使用电子表格,并配有清晰的列标题和一致的编码。在实地调查中,你可能还需要记录环境条件、当天的时段或确切地点,因为这些背景变量有助于解释测量值的变异。
A digital checklist can help ensure that no required data are missed. When using instruments such as a stopwatch, a ruler or a light meter, take multiple readings and average them to reduce random measurement errors. Always include units on recording sheets to avoid confusion later.
数字检查清单有助于确保不会遗漏所需数据。当使用秒表、尺子或照度计等仪器时,应多次读数并取平均值,以减少随机测量误差。在记录表上务必标注单位,以免后期混淆。
4. Organising and Presenting Data | 整理与呈现数据
Once data are collected, they must be organised into a clear, analysable format. Tally charts and frequency tables are the simplest tools. For continuous data, you may need to group values into classes, making sure that class widths are equal where possible and that classes do not overlap.
数据收集完毕后,必须整理成清晰、可分析的格式。划记表和频率表是最简单的工具。对于连续数据,你需要将数值分组,确保在可能的情况下组距相等,且组与组之间不重叠。
Visual displays make patterns immediately visible. A bar chart is suitable for categorical data, while a histogram displays the distribution of continuous data, with area representing frequency. Frequency polygons and cumulative frequency curves (ogives) are useful for comparing several sets of data and for estimating medians and percentiles.
可视化展示让数据模式一目了然。条形图适用于分类数据,而直方图显示连续数据的分布,用面积表示频率。频率多边形和累积频率曲线(累积频数图)对于比较多组数据以及估计中位数和百分位数很有用。
Box plots (box-and-whisker diagrams) provide a five-number summary of the data: minimum, lower quartile, median, upper quartile and maximum. They are especially helpful for spotting outliers and for comparing the spread of several groups side by side. All graphs must be titled, with axes labelled and appropriate scales chosen.
箱线图(盒须图)给出了数据的五数汇总:最小值、下四分位数、中位数、上四分位数和最大值。它们对于发现异常值以及并排比较几组数据的散布情况特别有用。所有图表都应有标题,坐标轴要有标签,并选择合适的刻度。
5. Measures of Central Tendency and Spread | 集中趋势与离散程度的度量
The most common measure of central tendency is the mean, calculated as
x̄ = Σx / n
不过,当数据存在极端异常值时,中位数是更稳健的替代指标,它代表有序数据的中点。对于分类数据,众数(出现频率最高的类别)通常是唯一合适的集中趋势度量。
Spread is just as important as centre. The range (maximum − minimum) is quick to find but heavily influenced by outliers. The interquartile range (IQR = Q₃ − Q₁) gives the spread of the middle 50% of the data and is resistant to extreme values.
离散程度与集中趋势同等重要。极差(最大值 − 最小值)容易计算,但受异常值影响很大。四分位距(IQR = Q₃ − Q₁)给出了中间 50% 数据的散布范围,并且不受极端值的影响。
Variance and standard deviation measure the average squared deviation from the mean. The sample standard deviation s, calculated using
s = √[ Σ(x − x̄)² / (n − 1) ]
是用得最广的离散度量,因为它与原始数据具有相同的单位。在比较不同均值的变异性时,变异系数也很有价值。
6. Probability Distributions in Fieldwork Contexts | 实地调查中的概率分布
Many fieldwork measurements are approximately normally distributed, especially when they arise from a large number of small, independent influences. Recognising a bell-shaped histogram or a straight line on a normal probability plot helps you justify using normal distribution theory for confidence intervals and tests.
许多实地调查测量值近似服从正态分布,特别是当它们源于大量独立的小影响之和时。识别出钟形直方图或正态概率图上的直线,可以帮助你有理由使用正态分布理论来构建置信区间和进行检验。
The binomial distribution is relevant when you count the number of successes in a fixed number of independent trials, each with the same probability of success. For example, surveying whether a person recycles can be modelled binomially, providing the sample size is small relative to the population and the trials are independent.
当你在固定次数的独立试验中统计成功的次数,且每次试验的成功概率相同时,二项分布就派上了用场。例如,调查一个人是否进行回收利用,就可以用二项分布建模,前提是样本量相对总体较小,且试验相互独立。
For large sample sizes, the normal distribution can approximate the binomial distribution, provided np and n(1 − p) are both greater than 5. This approximation is extremely useful in fieldwork because it allows you to construct confidence intervals for a population proportion without using exact binomial tables.
对于大样本,只要 np 和 n(1 − p) 都大于 5,正态分布就可以近似二项分布。这种近似在实地调查中非常有用,因为它允许你不用查阅精确二项分布表就能构建总体比例的置信区间。
7. Confidence Intervals from Fieldwork Data | 基于实地调查数据的置信区间
A confidence interval provides a range of plausible values for an unknown population parameter. For a population mean μ when the population standard deviation σ is unknown, the interval is
x̄ ± t × (s / √n)
其中 t 值取决于所需的置信水平和自由度 n − 1。在实地调查中,95% 置信区间最为常用,它意味着如果重复抽样多次,大约有 95% 的区间会包含真实的总体均值。
When estimating a population proportion p from a large sample, the approximate confidence interval is
p̂ ± z* × √[ p̂(1 − p̂) / n ]
置信区间的宽度受样本大小控制:将 n 扩大到原来的四倍,可以使区间宽度大约减半。这在实地调查的规划阶段是一条关键的实用信息。
Always interpret a confidence interval carefully. It is not a probability statement about the specific interval you have computed, but a statement about the long-run behaviour of the method. In your fieldwork write-up, state the interval and discuss whether it is narrow enough to be useful for the decision at hand.
解释置信区间时务必要小心。它不是关于你具体计算出的区间的概率陈述,而是关于该方法长期表现的一种陈述。在你的实地调查报告中,要陈述区间并讨论它是否足够狭窄,能对当前决策有所帮助。
8. Correlation and Regression in Field Studies | 实地研究中的相关与回归
When two variables are measured in the same individuals or locations, a scatter diagram is the first tool to assess the relationship. The product moment correlation coefficient r, calculated as
r = Σ[(x − x̄)(y − ȳ)] / √[ Σ(x − x̄)² Σ(y − ȳ)² ]
取值从 −1 到 +1,量化了线性关联的强度和方向。在实地调查中,注意相关不代表因果:两者可能都受到某个潜在变量的影响。
If a linear relationship is plausible, the least squares regression line of y on x has equation ŷ = a + bx, where the slope b = r (sᵧ / sₓ) and the intercept a = ȳ − b x̄. This line can be used for prediction within the range of the original data, but extrapolation beyond that range can be highly unreliable.
如果线性关系是可信的,那么 y 对 x 的最小二乘回归线方程为 ŷ = a + bx,其中斜率 b = r (sᵧ / sₓ),截距 a = ȳ − b x̄。这条线可以在原始数据范围内用于预测,但超出该范围的外推可能极不可靠。
Residual plots should be examined to check that the assumptions of the regression model hold: linearity, constant variance (homoscedasticity) and approximate normality of the residuals. In fieldwork, patterns in residuals can often be traced back to missing explanatory variables or to non-linear relationships that would require transformation.
应检查残差图,以验证回归模型的假设是否成立:线性、等方差性(同方差性)以及残差的近似正态性。在实地调查中,残差图呈现出的模式往往可以追溯到遗漏的解释变量,或者需要变换的非线性关系。
9. Introduction to Hypothesis Testing | 假设检验导论
A hypothesis test is a formal procedure for deciding whether observed data provide enough evidence against a claim. The null hypothesis H₀ represents a default position (e.g., no difference or no effect), and the alternative hypothesis H₁ states what you are trying to prove. In fieldwork, you typically gather data to see whether they contradict H₀.
假设检验是一个正式的程序,用于判断观测数据是否提供了足够证据来反对某一论断。原假设 H₀ 代表一种默认立场(例如无差异或无效),备择假设 H₁ 则陈述你想证明的内容。在实地调查中,你通常收集数据来看看它们是否与 H₀ 相矛盾。
The p-value is the probability of obtaining a result at least as extreme as the one observed, assuming H₀ is true. If the p-value is smaller than the significance level α (commonly 0.05), we reject H₀ and conclude that the result is statistically significant. In write-ups, always report the p-value and justify your choice of α.
p 值是在 H₀ 为真的前提下,得到至少与观测结果一样极端的结果的概率。如果 p 值小于显著性水平 α(通常为 0.05),我们就拒绝 H₀,并得出结果具有统计显著性的结论。在报告中,务必报告 p 值,并说明你选择 α 的理由。
Errors can occur: a Type I error rejects a true null hypothesis, while a Type II error fails to reject a false null. Fieldwork investigations must balance these risks, often by increasing sample size to boost the power of the test.
可能发生的错误有两类:第一类错误拒绝了正确的原假设,第二类错误则未能拒绝错误的原假设。实地调查必须权衡这些风险,通常通过增加样本量来提高检验的功效。
10. Chi-Squared Tests for Categorical Fieldwork Data | 分类实地调查数据的卡方检验
When your fieldwork produces frequency counts in categories, the chi-squared test is the standard tool for hypothesis testing. The test for association asks whether two categorical variables are independent. The observed frequencies are entered into a contingency table, and expected frequencies are calculated under the assumption of independence.
当你的实地调查产生分类频数时,卡方检验是假设检验的标准工具。独立性检验用来考察两个分类变量是否相互独立。将观测频数录入列联表,并在独立性假设下计算期望频数。
The test statistic is
χ² = Σ [ (O − E)² / E ]
应当确保每个单元格的期望频数都足够大。
Degrees of freedom are (r − 1)(c − 1) for a table with r rows and c columns. The resulting p-value is found from a chi-squared distribution table or a calculator. A significant result indicates an association, but not its strength; a follow-up analysis of residuals or a measure such as Cramér’s V can help describe the effect size.
对于 r 行 c 列的表格,自由度为 (r − 1)(c − 1)。根据得到的 p 值,可从卡方分布表或计算器中查找。显著结果表明存在关联,但并不能说明关联的强度;对残差的进一步分析或像 Cramér’s V 这样的度量,有助于描述效应大小。
For a goodness-of-fit test, you compare an observed distribution to a theoretical one, such as testing whether the observed colour preferences of birds visiting a feeder match a 1:1:1 ratio. Always state the hypotheses clearly and check that the expected frequencies meet the usual criterion (≥5 for at least 80% of cells).
拟合优度检验则用于将观测分布与理论分布进行比较,例如检验来访食槽的鸟类颜色偏好是否符合 1:1:1 的比例。务必将假设陈述清楚,并检查期望频数是否满足通常的准则(至少 80% 的单元格≥5)。
11. Drawing Conclusions and Evaluating Fieldwork | 得出结论与评估实地调查
Your conclusions must be firmly anchored in the data analysis, not in wishful thinking. Start by revisiting the original research question and state whether the evidence supports the alternative hypothesis. Use the confidence intervals to quantify the uncertainty, and avoid overstating the certainty of your findings.
你的结论必须牢牢建立在数据分析之上,而非一厢情愿的想法。首先回顾最初的研究问题,并陈述证据是否支持备择假设。使用置信区间来量化不确定性,避免夸大学术发现的确定程度。
Every fieldwork investigation has limitations. Discuss possible sources of bias, such as non-response, measurement error or an imperfect sampling frame. Reflect on the sample size: could a larger or better-stratified sample have changed the inferences? Acknowledging these factors demonstrates mature statistical thinking.
每一项实地调查都有其局限性。讨论可能的偏差来源,例如无应答、测量误差或不完美的抽样框。反思样本量:一个更大或分层更精细的样本是否会改变推断?承认这些因素体现了成熟的统计思维。
Finally, suggest improvements for future studies. Could you use a different sampling strategy, a more reliable instrument, or a longitudinal design? By framing your evaluation constructively, you show that you understand statistics as a dynamic process rather than a set of rigid recipes.
最后,为未来的研究提出改进建议。你是否可以采用不同的抽样策略、更可靠的仪器,或者纵向设计?通过建设性地构建你的评估,你表明了自己将统计学理解为一个动态过程,而非一套僵化的配方。
12. Using Technology to Support Fieldwork Data Analysis | 使用技术辅助实地调查数据分析
Spreadsheet software such as Excel, Google Sheets or GeoGebra can greatly speed up routine calculations. Functions for mean, median, standard deviation, quartiles and correlation coefficients are built in. Pivot tables are especially useful for quickly producing frequency counts and cross-tabulations from raw fieldwork data.
电子表格软件,如 Excel、Google Sheets 或 GeoGebra,能极大提高常规计算的速度。均值、中位数、标准差、四分位数和相关系数等函数都是内置的。数据透视表尤其适用于从原始实地调查数据中快速生成频数统计和交叉表。
Graphical capabilities allow you to generate histograms, box plots and scatter diagrams with one click, making it easier to explore patterns before committing to formal tests. Many spreadsheet packages also offer add-ins that perform t-tests and chi-squared tests, but you must understand the output rather than just relying on the software.
其图形功能使你能够一键生成直方图、箱线图和散点图,从而更容易在进行正式检验之前探索数据模式。许多电子表格软件包还提供了能执行 t 检验和卡方检验的加载项,但你必须理解输出结果,而不是仅仅依赖软件。
In A‑level examinations, you are also expected to be familiar with the statistical mode of a scientific calculator. You should be able to enter lists of data, compute summary statistics, and find probabilities from the normal and binomial distributions efficiently. Practice switching between the standard and statistical modes until it becomes second nature.
在 A‑level 考试中,你还需要熟悉科学计算器的统计模式。你应该能够输入数据列表、计算汇总统计量,并能高效地从正态分布和二项分布中查找概率。练习在标准模式和统计模式之间切换,直到运用自如为止。
Published by TutorHao | Mathematics Revision Series | aleveler.com
Find Maths Textbooks on eBay UK
New, used and second-hand copies of textbooks and revision guides are often much cheaper than retail — check current listings and prices before you buy.
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply