📚 Voting Behaviour and the Media: Statistical Models and Bias | 投票行为与媒体:统计模型与偏差
In A-Level Mathematics, statistics gives us the tools to analyse how opinion polls, media coverage and voting behaviour interact. This article examines sampling, bias, confidence intervals, hypothesis tests, correlation, regression and probability models through the lens of voting and the media.
在A-Level数学中,统计学为我们提供了分析民意调查、媒体报道与投票行为之间互动的工具。本文从投票与媒体的视角,探讨抽样、偏差、置信区间、假设检验、相关、回归和概率模型。
1. Sampling Methods in Opinion Polls | 民意调查中的抽样方法
A reliable opinion poll must use random sampling so that every voter has an equal chance of being selected. The simple random sample, stratified sample and cluster sample are common methods. Stratified sampling is often used to ensure key groups such as age, region and income are represented in proportion to the population.
可靠的民意调查必须采用随机抽样,使每个选民都有相等的被选机会。常用的方法包括简单随机抽样、分层抽样和整群抽样。分层抽样通常用于确保年龄、地区和收入等关键群体在样本中的比例与总体一致。
A sampling frame is a list of all eligible voters. If mobile-only households are missing from the frame, the frame is incomplete and coverage bias occurs. Quota sampling is non-random and can be quicker, but it does not allow a valid margin of error to be calculated because the selection probability is unknown.
抽样框是所有合格选民的名单。如果抽样框遗漏了仅使用手机的住户,抽样框就不完整,会产生覆盖偏差。配额抽样是非随机方法,速度较快,但由于选择概率未知,无法计算有效的误差范围。
- Simple random sampling – every member has an equal chance of selection. 简单随机抽样——每个成员都有相等的被选机会。
- Stratified sampling – population divided into groups, then random samples taken from each. 分层抽样——将总体分组,然后从每组中随机抽样。
- Quota sampling – interviewer selects a fixed number from each group, but not randomly. 配额抽样——调查员从每组中选取固定数量,但并非随机。
2. Bias in Media Reporting and Polling | 媒体报道与民调中的偏差
Bias can enter through question wording, response bias, non-response bias and leading questions. A poll asking ‘Do you agree that candidate X is too extreme?’ will not produce neutral data. Media outlets may also select polls that favour their editorial line, which is a form of publication bias.
偏差可能通过问题措辞、回答偏差、无回答偏差和诱导性问题进入。一项调查如果问“你是否同意候选人X过于极端?”,就不会产生中立数据。媒体机构也可能只选择有利于其编辑立场的民调,这是一种发表偏差。
- Selection bias – non-random samples give distorted estimates. 选择偏差——非随机样本会使估计失真。
- Non-response bias – certain groups are less likely to answer, so they are under-represented. 无回答偏差——某些群体不太可能作答,因此代表性不足。
- Leading question bias – wording pushes respondents towards a particular answer. 诱导性问题偏差——措辞会引导受访者给出特定答案。
- Publication bias – only favourable or dramatic poll results are reported. 发表偏差——只报道有利或戏剧性的民调结果。
3. Confidence Intervals for Vote Share | 投票率的置信区间
A confidence interval estimates the true population proportion of voters supporting a party. For a large random sample, the approximate 95% confidence interval for a proportion p is given by the formula below, where p is the sample proportion and n is the sample size.
置信区间用于估计支持某政党的真实选民比例。对于大样本随机抽样,比例 p 的近似95%置信区间由以下公式给出,其中 p 为样本比例,n 为样本量。
p ± 1.96 × √(p(1-p)/n)
For example, if a poll of 1,000 voters reports p = 0.52 for party A, the standard error is √(0.52 × 0.48 / 1000) ≈ 0.0158. The 95% confidence interval is 0.52 ± 1.96 × 0.0158, which gives approximately 0.489 to 0.551. Since the interval crosses 0.5, the poll cannot conclusively say party A has majority support.
例如,若一项1000名选民的调查显示A党支持率 p = 0.52,则标准误为 √(0.52 × 0.48 / 1000) ≈ 0.0158。95%置信区间为 0.52 ± 1.96 × 0.0158,约等于 0.489 到 0.551。由于区间跨过0.5,该民调无法确凿地说明A党获得多数支持。
4. Hypothesis Testing: Does Media Coverage Change Support? | 假设检验:媒体报道是否改变支持率?
We can test whether media coverage has changed support for a policy by comparing a sample proportion with a previously known value. The null hypothesis H₀ states that the true proportion is p₀, while the alternative hypothesis H₁ states that it is different, larger or smaller. The test statistic for a binomial proportion is calculated as shown.
我们可以通过将样本比例与已知值比较,检验媒体报道是否改变了对某项政策的支持率。原假设 H₀ 表示真实比例等于 p₀,备择假设 H₁ 表示真实比例不同、更大或更小。二项比例的检验统计量如下所示。
z = (p – p₀) / √(p₀(1-p₀)/n)
Suppose a party’s support was 0.40 before a media campaign. A later random sample of 500 voters finds 225 supporters, so p = 0.45. Testing H₀: p = 0.40 against H₁: p ≠ 0.40 at the 5% level, the z-value is (0.45 – 0.40) / √(0.40 × 0.60 / 500) ≈ 2.28. The critical value is 1.96, so we reject H₀ and conclude that support has significantly changed. However, this does not prove that media coverage caused the change.
假设某政党在媒体宣传前的支持率为0.40。之后一个500名选民的随机样本中有225人支持,p = 0.45。在5%显著性水平下检验 H₀: p = 0.40 对 H₁: p ≠ 0.40,z 值为 (0.45 – 0.40) / √(0.40 × 0.60 / 500) ≈ 2.28。临界值为1.96,因此拒绝 H₀,得出支持率显著变化的结论。但这并不能证明媒体报道导致了该变化。
5. Correlation vs Causation: Media Exposure and Voting Intention | 相关与因果:媒体接触与投票意向
A scatter diagram of media hours per week against intention to vote may show a linear correlation. The product moment correlation coefficient r measures the strength and direction of a linear relationship, with values between -1 and 1. A value of r = 0.62 suggests a moderate positive correlation, but correlation does not imply causation.
每周媒体接触小时数与投票意向的散点图可能呈现线性相关。积矩相关系数 r 衡量线性关系的强度和方向,取值范围为 -1 到 1。r = 0.62 表明存在中等正相关,但相关并不意味着因果。
r = Sxy / √(Sxx × Syy)
Even if high media use is associated with higher turnout, a third variable such as age or political interest may drive both. Media coverage may also reflect existing preferences rather than creating them. To claim causation, a controlled experiment or rigorous longitudinal design is needed.
即使高媒体使用与较高投票率相关,年龄或政治兴趣等第三变量也可能同时驱动两者。媒体报道也可能只是反映已有的偏好,而非创造偏好。要声称因果关系,需要进行对照实验或严格的纵向设计。
6. Regression Analysis of Campaign Effects | 竞选效应的回归分析
Linear regression can model how an independent variable, such as advertising spending, affects a dependent variable, such as vote share. The least squares regression line has the form y = a + bx, where b is the gradient and a is the intercept. The coefficient b estimates the change in vote share per unit increase in spending.
线性回归可以建模自变量(如广告支出)对因变量(如得票率)的影响。最小二乘回归直线的形式为 y = a + bx,其中 b 是斜率,a 是截距。系数 b 估计每增加一单位支出所带来的得票率变化。
b = Sxy / Sxx, a = ȳ – b x̄
For example, if b = 0.02 and x is measured in thousands of pounds, an extra £10,000 of advertising is predicted to increase vote share by 0.2 percentage points. The residual for each data point is the observed value minus the predicted value; large residuals may indicate outliers or omitted variables.
例如,若 b = 0.02 且 x 以千英镑为单位,则额外10,000英镑广告支出预计会使得票率增加0.2个百分点。每个数据点的残差等于观测值减预测值;较大的残差可能表明存在异常值或遗漏变量。
7. Conditional Probability and Bayesian Updating | 条件概率与贝叶斯更新
Voters update their beliefs when new media information arrives. Conditional probability is the probability of event A given B, written P(A|B). Bayes’ theorem connects prior probability, likelihood and posterior probability, which is highly relevant when a pollster or voter revises expectations after a debate or scandal
Published by TutorHao | A-Level Mathematics Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导