Statistical Analysis of Welsh Assembly and Government Data | 威尔士议会与政府数据的统计分析

📚 Statistical Analysis of Welsh Assembly and Government Data | 威尔士议会与政府数据的统计分析

In Edexcel A-Level Mathematics, the applied statistics content requires students to work with real-world data and understand how statistical techniques can inform decision-making. The Welsh Assembly (Senedd Cymru) and the Welsh Government publish a wide range of official data, including election results, budget allocations, health waiting times and education performance indicators. This article uses those data sets to illustrate key statistical concepts from the specification, such as sampling, measures of location and spread, probability distributions, hypothesis testing and regression.

在 Edexcel A-Level 数学中,应用统计部分要求学生使用真实数据,并理解统计技术如何为决策提供依据。威尔士议会(Senedd Cymru)和威尔士政府发布了大量官方数据,包括选举结果、预算分配、健康等待时间和教育绩效指标。本文利用这些数据集来说明考纲中的关键统计概念,例如抽样、位置与离散程度的度量、概率分布、假设检验和回归。


1. Types of Data and Sampling Methods | 数据类型与抽样方法

When analysing data published by the Welsh Government, the first step is to classify variables correctly. Qualitative or categorical data describe qualities, such as the political party of a Senedd member or the region of a local authority. Quantitative data are numerical and can be discrete, such as the number of seats won by a party in the 40 Senedd constituencies, or continuous, such as the amount of money allocated to health boards in millions of pounds. Sampling methods are essential because it is rarely possible to survey every resident in Wales. Simple random sampling gives every individual an equal chance of selection. Stratified sampling divides the population into groups, such as the 22 unitary authorities of Wales, and samples proportionally from each group. Systematic sampling selects every kth item from a list, while quota sampling is often used by opinion pollsters to ensure representative proportions of age, gender and region.

在分析威尔士政府发布的数据时,第一步是正确分类变量。定性或分类数据描述的是属性,例如 Senedd 议员的所属政党或地方政府的所在地区。定量数据是数值型数据,可以是离散的,例如某政党在 40 个 Senedd 选区中赢得的席位数;也可以是连续的,例如分配给各卫生委员会的金额(以百万英镑计)。抽样方法至关重要,因为几乎不可能调查威尔士的每一位居民。简单随机抽样使每个个体被选中的机会均等。分层抽样先将总体划分为若干组,例如威尔士的 22 个单一管理区,然后按比例从每组中抽样。系统抽样从列表中每隔 k 个抽取一个,而配额抽样常被民意调查机构使用,以确保年龄、性别和地区的代表性比例。


2. Measures of Central Tendency and Spread | 集中趋势与离散程度的度量

Measures of location summarise the centre of a data set. For example, if we collect the per-pupil education spending in each of the 22 Welsh local authorities, the mean is calculated as x̄ = Σx ÷ n, where Σx is the sum of all values and n is the number of values. The median is the middle value when the data are ordered, and the mode is the most frequent value. However, the mean alone does not tell us how spread out the data are. The variance and standard deviation measure dispersion. For a sample, the variance is s² = Σ(x − x̄)² ÷ (n − 1), and the standard deviation is its positive square root. A small standard deviation indicates that local authorities spend similar amounts per pupil, while a large one shows significant regional inequality.

位置度量概括了数据集的中心。例如,如果我们收集威尔士 22 个地方政府每名学生的教育支出,平均值计算为 x̄ = Σx ÷ n,其中 Σx 是所有数值的总和,n 是数值的个数。中位数是将数据排序后位于中间的值,众数是出现频率最高的值。然而,仅靠平均值并不能告诉我们数据的离散程度。方差和标准差用于度量离散程度。对于样本,方差为 s² = Σ(x − x̄)² ÷ (n − 1),标准差是其正的平方根。标准差小表示各地方政府每名学生的支出相似,标准差大则表明地区间存在显著不平等。


3. Presenting Data: Histograms and Box Plots | 数据展示:直方图与箱线图

Histograms are used to display the distribution of continuous data, such as the salaries of employees in Welsh Government departments. The area of each bar is proportional to the frequency, so the vertical axis represents frequency density. Box plots are useful for comparing different groups, for example social care spending in north and south Wales. A box plot shows the minimum, lower quartile Q₁, median Q₂, upper quartile Q₃ and maximum. Outliers can be identified using the rule that any value more than 1.5 × IQR below Q₁ or above Q₃ is an outlier, where IQR = Q₃ − Q₁. The table below gives a five-number summary for two fictional groups of local authorities.

直方图用于展示连续数据的分布,例如威尔士政府各部门员工的薪资分布。每个条形的面积与频率成正比,因此纵轴表示频率密度。箱线图有助于比较不同组,例如威尔士北部和南部的社会关怀支出。箱线图显示最小值、下四分位数 Q₁、中位数 Q₂、上四分位数 Q₃ 和最大值。异常值可以通过以下规则识别:任何低于 Q₁ − 1.5 × IQR 或高于 Q₃ + 1.5 × IQR 的值都是异常值,其中 IQR = Q₃ − Q₁。下表给出了两组虚构地方政府的五数概括。

Group Min Q₁ Q₂ Q₃ Max
North Wales 120 145 160 180 210
South Wales 100 130 155 170 230

By comparing the box plots, we can see that north Wales authorities have a smaller range but a higher minimum, while south Wales has a larger overall spread and at least one potential outlier at the top end.

通过比较箱线图,我们可以看到威尔士北部的地方政府极差较小但最小值较高,而南部整体离散程度更大,并且在上端至少有一个潜在异常值。


4. Probability Basics and Government Polling | 概率基础与政府民意调查

Probability is the foundation for understanding opinion poll results related to the Welsh Assembly. If we select one voter at random from the Welsh electorate, the probability that this voter supports a particular party is equal to that party’s proportion of support in the population. Venn diagrams can represent events such as “lives in Wales and is aged 18-24” and “supports independence”. The intersection of two events A and B is written as A ∩ B, and the union as A ∪ B. Conditional probability P(A | B) is the probability that event A occurs given that event B has already occurred. For example, we might calculate the probability that a randomly selected Welsh voter supports a given policy, conditional on them living in a particular region. The formula is P(A | B) = P(A ∩ B) ÷ P(B).

概率是理解与威尔士议会相关的民意调查结果的基础。如果我们从威尔士选民中随机选择一人,该选民支持某一特定政党的概率等于该政党在总体中的支持率。维恩图可以表示诸如“居住在威尔士且年龄在 18-24 岁”和“支持独立”等事件。两个事件 A 和 B 的交集写作 A ∩ B,并集写作 A ∪ B。条件概率 P(A | B) 是在事件 B 已经发生的条件下事件 A 发生的概率。例如,我们可以计算随机选择的威尔士选民在特定地区居住的条件下支持某项政策的概率。公式为 P(A | B) = P(A ∩ B) ÷ P(B)


5. Discrete Random Variables: Modelling Assembly Seats | 离散随机变量:模拟议会席位

The Senedd uses an additional member system with 40 constituency seats and 20 regional list seats. We can model the number of constituency seats won by a particular party as a discrete random variable X. The probability distribution of X lists each possible value x and its corresponding probability P(X = x). The expected value E(X) is the long-run average number of seats, calculated as E(X) = Σ x × P(X = x). The variance Var(X) is given by Var(X) = E(X²) − [E(X)]². If a party has an equal chance of winning each constituency and the constituencies are independent, then X follows a binomial distribution with n = 40 and p equal to the probability of winning a single constituency.

Senedd 采用额外议员制度,设有 40 个选区席位和 20 个地区名单席位。我们可以将某一政党赢得的选区席位数建模为离散随机变量 X。X 的概率分布列出了每个可能值 x 及其对应的概率 P(X = x)。期望值 E(X) 是长期平均席位数,计算公式为 E(X) = Σ x × P(X = x)。方差 Var(X) 由 Var(X) = E(X²) − [E(X)]² 给出。如果某政党在每个选区获胜的概率相同,并且各选区相互独立,那么 X 服从二项分布,其中 n = 40,p 等于在单个选区获胜的概率。


6. Binomial Distribution: Predicting Election Outcomes | 二项分布:预测选举结果

The binomial distribution X ~ B(n, p) models the number of successes in n independent trials, each with probability p of success. In the context of Senedd elections, suppose a party wins each constituency with probability p = 0.35. The probability of winning exactly k seats out of 40 is given by P(X = k) = C(n, k) × pᵏ × (1 − p)ⁿ⁻ᵏ, where C(n, k) is the binomial coefficient. To find the probability of winning at least 15 seats, we calculate P(X ≥ 15) = 1 − P(X ≤ 14), which can be evaluated using a calculator or statistical tables. The mean of X is np = 40 × 0.35 = 14, and the variance is np(1 − p) = 40 × 0.35 × 0.65 = 9.1. The assumption of independent constituencies is a simplification, because regional factors can create correlation, but the binomial model still provides a useful first approximation.

二项分布 X ~ B(n, p) 用于描述 n 次独立试验中成功的次数,每次试验成功的概率为 p。在 Senedd 选举的背景下,假设某政党在每个选区获胜的概率为 p = 0.35。在 40 个选区中恰好赢得 k 个席位的概率由 P(X = k) = C(n, k) × pᵏ × (1 − p)ⁿ⁻ᵏ 给出,其中 C(n, k) 是二项式系数。要计算至少赢得 15 个席位的概率,我们计算 P(X ≥ 15) = 1 − P(X ≤ 14),可以使用计算器或统计表求值。X 的均值为 np = 40 × 0.35 = 14,方差为 np(1 − p) = 40 × 0.35 × 0.65 = 9.1。各选区相互独立的假设是一种简化处理,因为区域因素可能产生相关性,但二项模型仍然提供有用的第一近似。


7

Published by TutorHao | A-Level Mathematics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Exit mobile version