📚 Gender, Age, Ethnicity and Region as Factors in Voting Behaviour: An AS Statistics Application | 性别、年龄、种族与地区作为投票行为因素的统计学分析
In election studies, gender, age, ethnicity and region are among the most commonly examined demographic variables because they often show clear patterns in voting behaviour and turnout. For AS Mathematics students, these real-world variables provide excellent contexts for applying statistical techniques such as sampling, data presentation, conditional probability, correlation and hypothesis testing. This article shows how the Edexcel AS Statistics toolkit can be used to analyse voting data without losing sight of the fact that correlation does not imply causation.
在选举研究中,性别、年龄、种族和地区是最常被分析的人口统计学变量,因为它们通常与投票行为和投票率呈现出清晰的模式。对 AS 数学学生而言,这些现实变量为抽样、数据表示、条件概率、相关性和假设检验等统计方法提供了极佳的应用背景。本文将展示如何利用 Edexcel AS 统计工具分析投票数据,同时始终牢记相关关系并不意味着因果关系。
1. The Statistical Question: Measuring Voting Behaviour | 统计问题:用数据刻画投票行为
Voting behaviour can be studied by recording whether an individual votes, which party or candidate they choose, and how this choice changes over time. Turnout is usually measured as the percentage of registered voters who cast a ballot in a given election. In statistical terms, turnout is a binary outcome for each individual: voted or did not vote. Aggregating this binary variable over groups defined by gender, age, ethnicity or region produces proportions that can be compared using probability and inference.
投票行为可以通过记录个人是否投票、选择哪个政党或候选人以及这种选择如何随时间变化来研究。投票率通常指在特定选举中实际投票的登记选民所占的百分比。在统计术语中,投票率对每个个体而言是一个二元结果:投票或未投票。将这一二元变量按性别、年龄、种族或地区分组汇总,可以得到可用概率和推断进行比较的比例。
For example, a survey may report that 68% of voters aged 65 and over turned out, while only 43% of voters aged 18 to 24 did so. The statistical question is whether this observed difference is large enough to be considered significant, or whether it could have occurred by chance in a sample.
例如,一项调查可能报告 65 岁及以上选民的投票率为 68%,而 18 至 24 岁选民的投票率仅为 43%。统计问题在于:这一观测差异是否足够大,以至于可以认为显著,或者它是否可能只是样本中偶然出现的。
2. Types of Data and Variables | 数据类型与变量
Gender, ethnicity and region are categorical variables, meaning they place individuals into non-numerical groups. Age can be treated as numerical if measured in years, or as categorical if grouped into bands such as 18–24, 25–34, 35–44 and so on. Turnout is a categorical variable with two outcomes, but when aggregated it becomes a proportion, which is numerical and continuous between 0 and 1.
性别、种族和地区是分类变量,即将个体归入非数值的组别。年龄如果以岁数计量,则为数值变量;如果按年龄段(如 18–24、25–34、35–44 等)分组,则为分类变量。投票率是只有两种结果的分类变量,但经过汇总后变成比例,即介于 0 和 1 之间的数值型连续变量。
In the Edexcel AS specification, understanding variable types is essential for choosing the correct diagram and summary statistic. Categorical data are summarised by counts and proportions, while numerical data can be summarised by the mean, median, standard deviation and interquartile range.
在 Edexcel AS 考纲中,理解变量类型对于选择正确的图表和汇总统计量至关重要。分类数据用频数和比例进行汇总,数值数据则可用均值、中位数、标准差和四分位距进行汇总。
3. Sampling Methods and Bias in Election Surveys | 抽样方法与选举调查中的偏差
To study gender, age, ethnicity and region as factors in voting, pollsters usually take a sample of the population. A simple random sample gives every eligible voter an equal chance of being selected, which reduces selection bias. However, election surveys often use stratified sampling to ensure that subgroups such as age bands, ethnic groups and regions are represented in proportion to the population.
为了研究性别、年龄、种族和地区对投票的影响,民调机构通常从总体中抽取样本。简单随机抽样使每个合格选民都有相等的被选中机会,从而减少选择偏差。然而,选举调查通常采用分层抽样,以确保年龄组、种族群体和地区等子群在样本中的比例与总体一致。
Common sources of bias include non-response bias, where certain groups such as younger voters are less likely to answer surveys, and social desirability bias, where respondents may not truthfully report whether they voted. When interpreting voting statistics, AS students should ask: was the sample random? Was it large enough? Who was left out?
常见的偏差来源包括无响应偏差(如年轻选民不太可能回答调查)和社会期望偏差(受访者可能不如实报告是否投票)。在解释投票统计时,AS 学生应当思考:样本是否随机?样本量是否足够大?哪些人被遗漏了?
4. Displaying Voting Data: Bar Charts and Segmented Bar Charts | 用条形图与分段条形图展示投票数据
Categorical voting data are often displayed using bar charts. A simple bar chart can compare turnout rates across age groups or regions. For example, the height of each bar might show the percentage of registered voters who voted in the last general election for each age band. This makes differences in turnout by age immediately visible.
分类投票数据通常用条形图展示。简单的条形图可以比较不同年龄组或地区的投票率。例如,每根柱子的高度可表示各年龄段在上次大选中投票的登记选民百分比。这样便能使年龄导致的投票率差异一目了然。
A segmented bar chart is useful when we have two categorical variables, such as gender and party choice, or ethnicity and turnout. Each bar represents 100% of one group, and the segments show the proportion of that group in each category. This allows visual comparison of conditional distributions.
当存在两个分类变量时(如性别与政党选择,或种族与投票率),分段条形图非常有用。每根柱子代表某一组的 100%,各分段显示该组在每个类别中的比例。这样可以对条件分布进行直观比较。
5. Conditional Probability and Two-Way Tables | 条件概率与双向表
Two-way tables are central to analysing the relationship between demographic factors and voting. Suppose a survey of 400 adults records gender and whether they voted in a recent election. A possible table is shown below.
双向表是分析人口统计学因素与投票关系的重要工具。假设对 400 名成年人进行调查,记录其性别以及是否在最近一次选举中投票。可能的表格如下。
| Voted | Did not vote | Total | |
| Male | 120 | 80 | 200 |
| Female | 135 | 65 | 200 |
| Total | 255 | 145 | 400 |
The conditional probability that a randomly selected male voted is P(Voted | Male) = 120 / 200 = 0.60. For females, P(Voted | Female) = 135 / 200 = 0.675. The difference of 0.075 suggests that in this sample, women were slightly more likely to vote than men. AS students can use such calculations to compare turnout across gender, age, ethnicity or region.
随机选取一名男性且其投票的条件概率为 P(Voted | Male) = 120 / 200 = 0.60。对女性而言,P(Voted | Female) = 135 / 200 = 0.675。两者之差 0.075 表明在该样本中,女性的投票概率略高于男性。AS 学生可以利用这类计算比较不同性别、年龄、种族或地区的投票率。
6. Correlation Between Age and Turnout | 年龄与投票率之间的相关性
Age and turnout often show a positive correlation: as age increases, turnout tends to increase. If we record the midpoint of each age band and the percentage turnout for that band, we can plot a scatter diagram and calculate the product moment correlation coefficient r.
年龄与投票率通常呈正相关:年龄越大,投票率往往越高。如果我们记录每个年龄段的组中值以及该年龄段的投票率百分比,就可以绘制散点图并计算积矩相关系数 r。
Suppose five age groups give data points (20, 43), (30, 52), (40, 61), (50, 68), (60, 74). A scatter diagram shows an upward linear trend. Calculating r using the Edexcel formula is likely to give a value close to 0.98, indicating strong positive correlation. However, AS students must remember that correlation measured on grouped data can be affected by the choice of age bands and does not prove that ageing causes higher turnout.
假设五个年龄组给出数据点 (20, 43)、(30, 52)、(40, 61)、(50, 68)、(60, 74)。散点图呈现上升的线性趋势。使用 Edexcel 公式计算 r 很可能得到接近 0.98 的值,表明强正相关。然而,AS 学生必须记住,基于分组数据计算的相关性可能受到年龄段划分的影响,并且不能证明年龄增长会导致投票率上升。
7. Binomial Models for Turnout and Hypothesis Testing | 投票率的二项模型与假设检验
Turnout can be modelled as a binomial random variable. If the national turnout rate is p = 0.65, then the number of voters X in a random sample of n people follows X ~ B(n, 0.65). This model allows us to test whether a particular subgroup, such as young voters, has a significantly lower turnout.
投票率可以建模为二项随机变量。如果全国投票率为 p = 0.65,则随机抽取 n 人时投票人数 X 服从 X ~ B(n, 0.65)。该模型可用于检验某个子群体(如年轻选民)的投票率是否显著偏低。
Example: In a random sample of 200 voters aged 18–24, only 108 voted. Let p be the true probability that a young voter votes. We test H₀: p = 0.65 against H₁: p < 0.65. Under H₀, X ~ B(200, 0.65). The expected number of voters is 130, and the observed 108 is much lower. If P(X ≤ 108) is less than a 5% significance level, we reject H₀ and conclude that turnout among young voters is significantly below the national average.
示例:在 200 名 18–24 岁选民的随机样本中,只有 108 人投了票。设 p 为年轻选民投票的真实概率。检验 H₀:p = 0.65 对 H₁:p < 0.65。在 H₀ 下,X ~ B(200, 0.65)。期望投票人数为 130,而实际观测值 108 远低于此。如果 P(X ≤ 108) 小于 5% 显著性水平,我们就拒绝 H₀,并得出结论:年轻选民的投票率显著低于全国平均水平。
8. Interpreting Regional Differences with Averages and Spread | 用平均数与离散程度解释地区差异
Regional turnout data can be summarised using the mean and standard deviation. If the turnout rates in 12 regions are given, the mean shows the typical turnout, while the standard deviation measures how spread out the regions are. A small standard deviation indicates that turnout is fairly uniform across regions; a large standard deviation suggests strong regional differences.
地区投票率数据可以用均值和标准差来汇总。如果给出 12 个地区的投票率,均值表示典型投票率,标准差则度量各地区数据的离散程度。标准差小表明各地区投票率较为一致;标准差大则表明存在明显的地区差异。
For example, regional turnout values of 62%, 65%, 63%, 66% have a smaller spread than 50%, 70%, 55%, 80%. AS students should also consider outliers, such as one region with exceptionally low turnout due to local factors, which can pull the mean downward.
例如,地区投票率数值 62%、65%、63%、66% 的离散程度小于 50%、70%、55%、80%。AS 学生还应考虑异常值,例如某个地区因地方性因素投票率异常偏低,这会拉低均值。
9. Limitations: Confounding Variables and Misleading Graphs | 局限性:混杂变量与误导性图表
When interpreting associations between demographic factors and voting behaviour, AS students must be alert to confounding variables. Age, income, education, social class and region are often related. For example, older voters may have higher turnout not simply because of age, but because they are more likely to be homeowners, have stable residence and feel a stronger sense of civic duty.
在解释人口统计学因素与投票行为之间的关联时,AS 学生必须警惕混杂变量。年龄、收入、教育、社会阶层和地区往往相互关联。例如,年长选民投票率更高,可能并不仅仅因为年龄,而是因为他们更可能拥有住房、居住稳定,并具有更强的公民责任感。
Graphs can also mislead. A bar chart with a truncated vertical axis can exaggerate small differences in turnout between genders or ethnic groups. Always read the axis labels, check the scale, and ask whether the sample size is shown before accepting a chart’s visual message.
图表也可能产生误导。纵轴经过截断的条形图会夸大性别或种族群体之间投票率的微小差异。在接受图表的视觉信息之前,务必阅读轴标签、检查刻度,并确认是否给出样本量。
10. Exam-Style Tips and Summary | 考试技巧与总结
In Edexcel AS Statistics exam questions, demographic voting data may appear in contexts involving two-way tables, conditional probability, binomial hypothesis tests or scatter diagrams. Always define your variables clearly, show your working, and interpret your results in the context of the question.
在 Edexcel AS 统计考试题中,人口统计学投票数据可能出现在涉及双向表、条件概率、二项假设检验或散点图的情境中。务必清晰定义变量、展示计算过程,并结合题目背景解释结果。
Key points to remember: gender, ethnicity and region are categorical variables; age may be numerical or categorical; conditional probability compares turnout between groups; correlation measures the strength of a linear relationship; and a binomial test can determine whether an observed turnout rate is significantly different from a claimed value. Most importantly, statistical evidence from observational data cannot establish that a demographic factor causes a particular voting pattern.
需要记住的要点:性别、种族和地区是分类变量;年龄可以是数值变量或分类变量;条件概率用于比较不同群体的投票率;相关系数度量线性关系的强度;二项检验可以判断观测到的投票率是否与某个给定值存在显著差异。最重要的是,来自观测数据的统计证据不能确定某个人口统计学因素导致了特定的投票模式。
Published by TutorHao | Mathematics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply