一、计量经济学是什么:从数据到经济结论的桥梁 | What Is Econometrics? From Data to Economic Conclusions
计量经济学(Econometrics)是把数学、统计学与经济理论结合起来的学科。它用现实世界的数据去检验经济理论,回答”需求曲线真的向右下方倾斜吗””广告支出每增加一百万元,销售额会增加多少”这类具体问题。与纯理论不同,计量经济学强调的是用证据说话:任何理论都必须经过数据的检验才能被接受。
Econometrics combines mathematics, statistics and economic theory. It uses real-world data to test economic theories and answer concrete questions such as “Does the demand curve really slope downwards?” or “How much does sales revenue rise when advertising spending increases by one million yuan?” Unlike pure theory, econometrics emphasises evidence: any theory must pass the test of data before it is accepted.
在 A-Level 经济学的学习中,你不需要推导复杂的计量公式,但必须理解回归分析的基本思想:当我们观察到两个变量一起变动时,如何用一条直线或曲线把这种关系量化出来。回归分析是计量经济学的核心工具,也是本篇文章的主线。
In A-Level Economics you are not expected to derive complicated econometric formulas, but you must understand the basic idea of regression analysis: when two variables move together, how do we quantify that relationship with a straight line or a curve? Regression analysis is the core tool of econometrics and the main thread of this article.
计量方法在考试中通常以数据回应题(data response)的形式出现:题目给出一组统计数据,要求你画散点图、判断相关方向、解释回归结果,或者评价结论的可靠性。掌握本章内容,意味着你同时拿到了描述、解释和评价三类题目的分数。
In examinations, econometric methods usually appear in data-response questions: you are given a set of statistics and asked to plot a scatter diagram, judge the direction of correlation, interpret regression output, or evaluate how reliable the conclusion is. Mastering this material earns you marks in all three question types: describe, explain and evaluate.
二、相关性与因果关系:为什么相关不等于因果 | Correlation vs Causation: Why “Related” Does Not Mean “Caused”
相关(correlation)只描述两个变量一起变动的倾向:收入上升时消费也上升,这就是正相关;价格上升时需求量下降,这就是负相关。相关程度可以用相关系数 r 来度量,r 的取值在 -1 到 +1 之间。r 越接近 +1,正相关越强;越接近 -1,负相关越强;接近 0 则表示几乎无关。
Correlation only describes the tendency of two variables to move together: when income rises and consumption also rises, that is positive correlation; when price rises and quantity demanded falls, that is negative correlation. The degree of correlation can be measured by the correlation coefficient r, which takes values between -1 and +1. The closer r is to +1, the stronger the positive correlation; the closer to -1, the stronger the negative correlation; values near 0 mean there is almost no relationship.
因果(causation)则更进一步,说明一个变量的变化直接导致了另一个变量的变化。相关不等于因果,这是 A-Level 经济学考试中最常考的判断之一。经典的例子是:冰淇淋销量与溺水人数高度相关,但吃冰淇淋并不会导致溺水,真正的原因是夏天的高温同时推高了两者。
Causation goes further: it states that a change in one variable directly causes a change in another. Correlation does not imply causation – this is one of the most frequently tested judgments in A-Level Economics. The classic example: ice-cream sales and drowning deaths are highly correlated, but eating ice cream does not cause drowning; the real cause is hot summer weather, which raises both at the same time.
在分析回归结果时,必须警惕三类问题。第一是遗漏变量(omitted variable):真正起作用的第三个变量没有被纳入模型。第二是反向因果(reverse causation):也许不是广告带动销售,而是销售好的公司更有钱投广告。第三是虚假相关(spurious correlation):两个变量只是因为共同趋势而看起来相关。例如,研究教育对收入的影响时,如果不控制个人能力这个变量,教育变量的回归系数就会被高估。
When interpreting regression results, watch out for three problems. First, omitted variables: a third variable that really matters is left out of the model. Second, reverse causation: perhaps it is not advertising that drives sales, but profitable firms that can afford more advertising. Third, spurious correlation: two variables look related only because they share a common trend. For example, when studying the effect of education on income, if personal ability is not controlled, the regression coefficient on education will be overestimated.
三、散点图与拟合线:用图像识别变量关系 | Scatter Diagrams and Lines of Best Fit: Reading Relationships from Graphs
散点图把每一组数据画成图上的一个点,横轴是自变量(解释变量),纵轴是因变量(被解释变量)。通过观察点的分布形态,可以初步判断关系的方向(正或负)和强度(紧密或松散)。如果点大致沿从左下到右上的带状分布,就是正相关;沿从左上到右下的带状分布,就是负相关;如果点散成一团,则两者关系很弱。
A scatter diagram plots each pair of data as a single point, with the independent (explanatory) variable on the horizontal axis and the dependent (explained) variable on the vertical axis. The shape of the point cloud reveals the direction (positive or negative) and the strength (tight or loose) of the relationship. Points forming a band from bottom-left to top-right indicate positive correlation; a band from top-left to bottom-right indicates negative correlation; a shapeless cloud indicates a weak relationship.
拟合线(line of best fit)是一条尽可能靠近所有数据点的直线。画拟合线时不需要让线穿过每一个点,而是让各点到直线的垂直距离总体最小,线的两侧大致分布着差不多数量的点。拟合线的作用是把数据中的趋势提炼出来,方便我们预测和比较。
A line of best fit is a straight line that lies as close as possible to all the data points. It does not have to pass through every point; instead, the vertical distances from the points to the line should be as small as possible overall, with roughly equal numbers of points on each side. The line summarises the trend in the data so that we can predict and compare.
下面是一家公司最近五年的广告支出与销售额数据,我们用它作为全篇文章的工作例子(worked example)。
Below is five years of advertising spending and sales data for a company. We will use it as the worked example throughout this article.
| 广告支出(十万元)Advertising Spend (CNY 100,000) | 销售额(十万元)Sales Revenue (CNY 100,000) |
|---|---|
| 10 | 120 |
| 15 | 150 |
| 20 | 175 |
| 25 | 210 |
| 30 | 230 |
从表中可以看到,广告支出增加时销售额也随之增加,五个点大致沿一条从左下到右上的直线分布,说明两者之间存在正相关,而且关系相当紧密。这样的数据就适合用线性回归来建模。
The table shows that sales rise as advertising rises; the five points lie roughly along a straight line from bottom-left to top-right, indicating a positive and fairly strong correlation. Data like this is well suited to linear regression modelling.
四、简单线性回归模型:y = a + bx 的数学与经济学含义 | Simple Linear Regression: The Mathematics of y = a + bx
简单线性回归模型写作 y = a + bx。其中 y 是因变量(被解释变量),x 是自变量(解释变量),b 是回归系数(斜率),a 是截距。模型的基本假设是:在观测范围内,x 与 y 之间存在近似线性的关系,即 x 每变化一个单位,y 平均变化固定的大小。
The simple linear regression model is written y = a + bx, where y is the dependent (explained) variable, x is the independent (explanatory) variable, b is the regression coefficient (slope) and a is the intercept. The basic assumption is that, within the observed range, the relationship between x and y is approximately linear: each one-unit change in x is associated with a constant average change in y.
为什么经济问题常常可以近似为线性?因为在很多情况下边际变化相对稳定。例如,每增加 1 个单位的广告投入,销售额平均增加约 5.6 个单位;即使真实关系不是完美的直线,线性模型已经足够捕捉主要趋势,也便于理解和计算。
Why can economic problems often be approximated as linear? Because in many cases the marginal change is roughly constant. For example, each additional unit of advertising raises sales by about 5.6 units on average; even if the true relationship is not a perfect straight line, a linear model captures the main trend well enough and is easy to understand and calculate.
需要注意,回归线给出的是平均关系,而不是精确关系。预测值通常记作 y-hat,表示给定 x 时 y 的平均预期值;实际观测值会在预测值附近波动,这种波动正是残差(residual)的来源。理解”平均关系”这一点,是正确解读回归结果的前提。
Note that the regression line gives an average relationship, not an exact one. The predicted value, usually written as y-hat, is the expected average value of y for a given x; actual observations fluctuate around the prediction, and this fluctuation is the source of residuals. Understanding the idea of an “average relationship” is the prerequisite for interpreting regression results correctly.
五、最小二乘法:回归线是如何计算出来的 | The Least Squares Method: How the Regression Line Is Calculated
最小二乘法(ordinary least squares,简称 OLS)是计算回归线最常用的方法。它的目标是最小化所有残差的平方和。残差是每个实际观测值与回归线预测值之间的差,即”实际值减预测值”。OLS 找到的直线,是所有可能直线中残差平方和最小的那一条。
The ordinary least squares (OLS) method is the most common way to calculate a regression line. Its goal is to minimise the sum of the squares of all residuals. A residual is the difference between an actual observed value and the value predicted by the regression line, that is, “actual minus predicted”. The OLS line is the one with the smallest possible sum of squared residuals among all candidate lines.
为什么要对残差平方而不是直接相加?因为残差有正有负,直接相加会相互抵消,一条偏离严重的线也可能得到接近零的总和。平方之后所有残差都变成正数,而且远离直线的点会被赋予更大的权重,从而保证回归线不会被个别极端点过度拉动。
Why square the residuals instead of adding them directly? Because residuals are positive and negative, and direct addition would let them cancel out: even a badly fitting line could produce a sum close to zero. Squaring makes every residual positive and gives larger weight to points far from the line, ensuring that the line is not pulled too far by a few extreme points.
考试中不需要手算最小二乘法的完整公式,但需要知道两个结论:第一,斜率 b 等于 x 与 y 的协方差除以 x 的方差,它反映了 x 与 y 共同变动的强度;第二,截距 a 的取值使得回归线必定通过数据的均值点,即 x 的平均值和 y 的平均值的交点。
In the exam you do not need to calculate the full OLS formula by hand, but you should know two results. First, the slope b equals the covariance of x and y divided by the variance of x; it measures how strongly x and y move together. Second, the intercept a is chosen so that the regression line always passes through the mean point, the intersection of the average of x and the average of y.
用第 3 节的广告数据计算,可以得到回归线 y = 65 + 5.6x(数值为约数)。验证一下:当广告支出为 20(十万元)时,预测销售额为 65 + 5.6 x 20 = 177(十万元),与实际观测值 175 非常接近;当广告支出为 10 时,预测值为 121,与实际值 120 也几乎一致。这说明这条线对数据的拟合相当好。
Using the advertising data from Section 3, we obtain the regression line y = 65 + 5.6x (rounded figures). Let us check: when advertising is 20 (units of CNY 100,000), predicted sales are 65 + 5.6 x 20 = 177, very close to the actual value of 175; when advertising is 10, the prediction is 121, almost identical to the actual 120. The line fits the data quite well.
六、回归系数的解读:斜率与截距分别说明什么 | Interpreting Coefficients: What the Slope and the Intercept Tell Us
斜率 b 表示 x 每增加一个单位,y 平均变化 b 个单位。在广告与销售额的例子中,b = 5.6 意味着每多投入 1 个单位的广告费,销售额平均增加 5.6 个单位。如果换算成金额,就是每多花十万元广告费,销售额平均增加五十六万元。斜率的大小直接反映两个变量之间关系的强度,也常常是题目要求你解释的重点。
The slope b shows the average change in y when x increases by one unit. In the advertising-sales example, b = 5.6 means that each additional unit of advertising raises sales by 5.6 units on average. In money terms, every extra CNY 100,000 of advertising is associated with an average increase of CNY 560,000 in sales. The size of the slope directly reflects the strength of the relationship and is often the focus of exam questions.
截距 a 表示当 x = 0 时 y 的预测值。在例子中 a = 65,意思是即使广告支出为零,销售额仍预计为 65 个单位。这部分可以理解为品牌原有客户带来的基础销售,或者说企业在完全不投放广告的情况下依然保有的市场份额。
The intercept a is the predicted value of y when x = 0. Here a = 65, meaning that even with zero advertising, sales are predicted at 65 units. This can be interpreted as baseline sales from existing brand customers, the market share the firm keeps even without any advertising at all.
解读系数时必须注意单位与适用范围。斜率的可靠性只限于样本数据覆盖的 x 范围:用回归线外推(extrapolation)到样本之外的 x 值,例如把广告支出推测到样本区间十倍之外的规模,是非常危险的,因为真实关系可能不再是线性的,边际回报也可能递减。
When interpreting coefficients, pay attention to units and range. The slope is reliable only within the range of x covered by the sample: extrapolating the regression line to x values far outside that range, for example predicting sales at ten times the sample’s advertising level, is very risky, because the true relationship may no longer be linear and marginal returns may diminish.
七、判定系数 R²:模型的解释力有多强 | The Coefficient of Determination R2: How Strong Is the Model?
判定系数 R² 衡量回归线对数据的解释程度,取值在 0 到 1 之间。R² = 0.9 表示 y 的变动中 90% 可以由 x 的变动来解释,剩下 10% 来自其他因素和随机误差。R² 越接近 1,散点越紧贴回归线,模型的预测越可靠;R² 接近 0,则说明 x 几乎解释不了 y。
The coefficient of determination R2 measures how well the regression line explains the data, taking values between 0 and 1. R2 = 0.9 means that 90% of the variation in y can be explained by variation in x, with the remaining 10% coming from other factors and random error. The closer R2 is to 1, the tighter the points cluster around the line and the more reliable the predictions; an R2 near 0 means x explains almost nothing about y.
但高 R² 不等于因果关系成立,也不等于模型正确。两个经济上完全无关的变量,只要各自都有上升的时间趋势,把它们放在一起回归也可能得到很高的 R²,这就是前面提到的虚假回归。判断模型好坏,除了看 R²,还要结合经济理论、样本来源和数据的实际含义。
But a high R2 does not prove causation or model correctness. Two variables with no economic connection at all can still produce a high R2 if both have upward time trends; this is the spurious regression mentioned earlier. To judge a model, look beyond R2 and consider economic theory, the source of the sample and the real meaning of the data.
在考试中评价一个回归结果时,可以这样说:”R² 为 0.72,说明模型解释了约 72% 的变动,解释力较强;但仍需进一步检验是否存在遗漏变量,以及样本是否具有代表性。”这样的回答既展示了概念理解,又体现了批判性思维,正是评分标准中”评价”一档所需要的。
When evaluating a regression result in an exam, you could say: “The R2 of 0.72 means the model explains about 72% of the variation, showing reasonable explanatory power; however, we should still test for omitted variables and check whether the sample is representative.” Such an answer demonstrates conceptual understanding plus critical thinking, exactly what the “evaluation” band of the mark scheme requires.
八、回归分析的局限:异常值、样本量与虚假相关 | Limitations of Regression: Outliers, Sample Size and Spurious Correlation
异常值(outliers)是偏离整体趋势的极端数据点,它们会显著拉动回归线。例如,某一年因为一次性的促销活动导致销售额暴增,这个点会让斜率偏大,从而高估广告的长期效果。识别异常值的常用方法,是观察散点图中明显孤立于主趋势之外的点,并在分析中说明是否将其剔除。
Outliers are extreme data points that deviate from the overall trend, and they can pull the regression line noticeably. For example, a year of exceptional sales caused by a one-off promotion would make the slope too steep and overstate the long-run effect of advertising. A common way to spot outliers is to look for points that sit clearly apart from the main trend on the scatter diagram, and to state in the analysis whether they should be removed.
样本量过小会降低回归结果的可靠性。用 5 个数据点得到的回归线,其斜率远不如用 50 个数据点得到的斜率可信,因为个别点的影响在小样本中被放大。考试中常用的评价用语是:”样本量较小,结论可能不具有普遍性,需要更多数据来验证。”
A small sample reduces the reliability of regression results. The slope from a 5-point sample is far less credible than one from a 50-point sample, because each individual point carries more weight when the sample is small. A standard evaluation phrase in exams is: “The sample is small, so the conclusion may not be generalisable; more data is needed to verify it.”
虚假相关(spurious correlation)指两个变量因为共同趋势而看似相关,实际上并没有直接的经济联系。经典的例子是:儿童的鞋码与阅读能力呈正相关,但真正的原因是年龄,年龄同时让脚变大、让阅读能力变强。处理虚假相关的思路,是把真正起作用的第三个变量纳入分析,例如加入”年龄”这个控制变量。
Spurious correlation means two variables appear related because they share a common trend, without any direct economic connection. The classic example: children’s shoe size is positively correlated with reading ability, but the real cause is age, which simultaneously makes feet bigger and reading better. The way to deal with spurious correlation is to bring the truly active third variable into the analysis, for example adding “age” as a control variable.
此外,经济数据本身往往带有时间趋势(trend)、季节波动(seasonality)和测量误差(measurement error)。趋势会让两个无关变量显得相关,季节波动会影响短期数据的回归结果,测量误差则来自统计口径和数据收集过程。虽然这些属于进阶内容,但 A-Level 考生至少要知道它们的存在,并在评价数据质量时提及。
In addition, economic data often carries time trends, seasonality and measurement errors. Trends make unrelated variables look correlated, seasonality distorts regression results based on short-term data, and measurement errors come from statistical conventions and the data-collection process. These are advanced topics, but A-Level candidates should at least know that they exist and mention them when evaluating data quality.
九、考试实战:数据回应题中的回归问题答题框架 | Exam Practice: A Framework for Regression Questions in Data Response
在数据回应题中遇到回归问题时,可以按四步框架组织答案。第一步:描述(Describe)。描述数据的总体趋势,例如:”广告支出与销售额呈正相关,关系较强,散点大致沿一条直线分布。”
When facing a regression question in a data-response paper, organise your answer with a four-step framework. Step one: Describe. Describe the overall trend, for example: “Advertising spend and sales are positively and fairly strongly correlated, with the points lying roughly along a straight line.”
第二步:解释(Explain)。用回归系数说明经济含义,例如:”斜率 5.6 表明每增加一个单位的广告支出,销售额平均增加 5.6 个单位,说明广告投入对销售有显著的促进作用。”解释时要明确说出变量的单位,并联系题目背景。
Step two: Explain. Use the regression coefficient to state the economic meaning, for example: “The slope of 5.6 means that each additional unit of advertising raises sales by 5.6 units on average, showing that advertising has a significant positive effect on sales.” State the units clearly and link the coefficient to the context of the question.
第三步:应用(Apply)。用回归线做预测或比较,例如:”当广告支出为 25(十万元)时,预测销售额约为 65 + 5.6 x 25 = 205(十万元)。”计算时保持单位一致,并写出关键步骤,即使结果算错也能拿到过程分。
Step three: Apply. Use the regression line to predict or compare, for example: “When advertising is 25 (units of CNY 100,000), predicted sales are about 65 + 5.6 x 25 = 205 (units of CNY 100,000).” Keep units consistent and show the key steps, so that even a wrong final answer earns method marks.
第四步:评价(Evaluate)。讨论结果的局限,例如:”样本仅包含 5 个观测值,R² 未知,而且没有控制品牌口碑、季节、市场竞争等因素,因此结论应当谨慎使用,不能直接外推。”评价是拉开分数差距的关键,也是考官区分优秀考生与普通考生的地方。
Step four: Evaluate. Discuss the limitations, for example: “The sample contains only five observations, the R2 is unknown, and brand reputation, seasonality and market competition are not controlled, so the conclusion should be used with caution and cannot be extrapolated directly.” Evaluation is the key to separating top marks from average ones, and it is where examiners distinguish excellent candidates.
最后提醒两个细节。第一,读表时先看表头单位,是千元、万元还是百分比,代入公式时必须一致;第二,答案记得带单位,例如”平均增加 5.6 个单位”而不是只写”5.6″。这些细节看似简单,却是数据回应题中最常见的失分点。
Two final details. First, always check the units in the table heading – thousands, ten-thousands or percentages – and keep them consistent when substituting into formulas. Second, include units in your answer, for example “an average increase of 5.6 units” rather than just “5.6”. These details seem trivial, but they are the most common source of lost marks in data-response questions.
Summary | 总结
计量分析方法与回归模型是 A-Level 经济学的重要考点。核心要点可以概括为:区分相关与因果,会画会读散点图,理解 y = a + bx 与最小二乘法的基本思想,正确解读斜率与截距的经济含义,用 R² 判断模型的解释力,并掌握异常值、样本量、虚假相关等常见局限。
Econometric methods and regression models are an important part of A-Level Economics. The key points can be summarised as: distinguish correlation from causation, plot and read scatter diagrams, understand y = a + bx and the idea of least squares, interpret the economic meaning of the slope and the intercept, use R2 to judge explanatory power, and recognise common limitations such as outliers, sample size and spurious correlation.
考试答题时,按”描述-解释-应用-评价”四步框架组织答案,先稳稳拿下基础分,再用评价部分争取高分。理解回归的思想比记住公式更重要,因为它培养的是用证据说话的经济学思维方式,这正是整个 A-Level 经济学科想要教给你的核心能力。
In the exam, organise your answer with the four-step framework of describe, explain, apply and evaluate: secure the basic marks first, then chase top marks with evaluation. Understanding the idea of regression matters more than memorising formulas, because it builds the evidence-based way of thinking that the whole A-Level Economics course is designed to teach you.
更多咨询请联系16621398022(同微信)