Scatter Graphs and Correlation | 散点图与相关性

📚 Scatter Graphs and Correlation | 散点图与相关性

In mathematics, data is not just about single lists of numbers. Often we want to investigate the relationship between two different quantities, such as height and arm span, or hours spent studying and exam scores. This is called bivariate data, and one of the most powerful tools for exploring it is a scatter graph. By the end of this article, you will be able to draw scatter graphs, describe the correlation between variables, and use a line of best fit to make predictions. This is a key topic in the Cambridge KS3 mathematics curriculum, helping you develop essential data-handling and analysis skills.

在数学中,数据不仅仅是一串单独的数字。我们经常想研究两个不同量之间的关系,比如身高与臂展,或者学习时长与考试成绩。这便称为双变量数据,而探索它的最强工具之一就是散点图。读完本文,你将能够绘制散点图,描述变量之间的相关性,并使用最佳拟合线进行预测。这是剑桥 KS3 数学课程中的一个关键主题,能帮助你培养必要的数据处理和分析技能。

1. Introduction to Bivariate Data | 双变量数据简介

When we collect data on a single variable, such as the heights of pupils in a class, we call it univariate data. However, real-world questions often involve two variables measured on the same subjects. For example, you might record the temperature and the number of ice creams sold each day. Each day gives you a pair of values (temperature, ice creams sold). This is bivariate data. Scatter graphs are used to display bivariate data visually, with one variable on the x-axis and the other on the y-axis. By looking at the pattern of points, we can see whether there is a link or relationship between the two variables.

当我们收集单个变量的数据,比如班级学生的身高,我们称之为单变量数据。然而,现实世界的问题常常涉及在同一对象上测量的两个变量。例如,你可能每天记录气温和卖出的冰激凌数量。每一天给你一对数值(气温,冰激凌销量)。这就是双变量数据。散点图用于可视化展示双变量数据,一个变量放在 x 轴上,另一个放在 y 轴上。通过观察点的分布模式,我们可以看出两个变量之间是否存在联系或关系。


2. What is a Scatter Graph? | 什么是散点图?

A scatter graph (or scatter plot) is a diagram that uses Cartesian coordinates to display values for two variables for a set of data. Each data pair is plotted as a single point. The x-coordinate represents the value of the independent or explanatory variable, while the y-coordinate represents the value of the dependent or response variable. Unlike a line graph, the points are not connected by lines. The purpose is simply to show the overall trend, clusters, or any unusual points. Scatter graphs are particularly useful for qualitative analysis before more formal statistical calculations.

散点图是一种利用笛卡尔坐标来展示一组数据中两个变量数值的图表。每一对数据被绘制成一个单独的点。x 坐标代表自变量或解释变量的值,而 y 坐标代表因变量或响应变量的值。与折线图不同,这些点不用线相连。其目的只是展示整体趋势、聚类或任何异常点。散点图在进行更正式的统计计算前,对于定性分析尤其有用。


3. Plotting Points and Scales | 描点和标度

To draw a scatter graph, first decide which variable goes on each axis. Usually, the variable you think might influence the other is placed on the x-axis. For instance, if you are studying how hours of exercise affect resting heart rate, put “hours of exercise” on the x-axis and “resting heart rate” on the y-axis. Then choose suitable scales for both axes. The scale should cover the whole range of your data and be easy to read. Label each axis clearly, including the units. Plot each pair of values as a small cross or dot. Make sure each point is plotted accurately; a common mistake is to misread coordinates or plot points in the wrong order.

要绘制散点图,首先要决定哪个变量放在哪条轴上。通常,你认为可能影响另一个变量的那个放在 x 轴上。例如,如果你研究运动时长如何影响静息心率,则把“运动时长”放在 x 轴,把“静息心率”放在 y 轴。然后为两条轴选择合适的标度。标度应涵盖数据的整个范围,并易于读取。明确标注每条轴,包括单位。把每一对数值绘制成小十字或圆点。确保每个点绘制准确;常见的错误是看错坐标或以错误的顺序描点。


4. Types of Correlation | 相关性的类型

Correlation describes the direction and strength of a relationship between two variables. There are three main types: positive correlation, where as one variable increases the other also increases; negative correlation, where as one variable increases the other decreases; and no correlation, where no clear pattern exists. The strength of the correlation can be described as strong, moderate, or weak, depending on how closely the points follow a straight-line pattern. Perfect correlation means all points lie exactly on a straight line, which is rare in real data. Understanding these categories is essential for GCSE-level data handling and the Cambridge Checkpoint test.

相关性描述了两个变量之间关系的方向和强度。主要有三种类型:正相关,即一个变量增加时另一个也增加;负相关,即一个变量增加时另一个减少;以及无相关,即没有明确的模式。相关性的强度可根据点与直线模式的紧密程度描述为强、中等或弱。完全相关意味着所有点恰好落在一条直线上,这在真实数据中很少见。理解这些类别对于 GCSE 水平的数据处理和剑桥 Checkpoint 测试至关重要。


5. Positive Correlation Explained | 正相关解释

Positive correlation means that higher values of one variable tend to be associated with higher values of the other. The points on the scatter graph slope upwards from left to right. For example, there is usually a positive correlation between a person’s height and their shoe size: taller people tend to have larger feet. Another example is the relationship between the number of hours studied and the score on a test: more study time generally leads to higher marks. A strong positive correlation shows points clustered tightly around an upward-sloping line, while a weak positive correlation shows a more scattered but still upward trend.

正相关意味着一个变量的较高值往往与另一个变量的较高值相关联。散点图上的点从左到右呈上升趋势。例如,一个人的身高和鞋码之间通常存在正相关:较高的人往往脚更大。另一个例子是学习时数与考试分数之间的关系:更多的学习时间通常带来更高的分数。强的正相关显示点紧密地聚集在一条向上倾斜的线周围,而弱的正相关则显示更分散但仍然向上的趋势。


6. Negative Correlation | 负相关

Negative correlation occurs when higher values of one variable are linked to lower values of the other. The points slope downwards from left to right. A classic example is the relationship between the speed of a car and the time taken to travel a fixed distance: the faster you drive, the shorter the journey time. Another example is the number of hours spent watching TV and physical fitness scores; typically, more screen time correlates with lower fitness levels. As with positive correlation, the strength varies. If the points are almost exactly in a straight downward line, the correlation is strong negative; if they only roughly follow that pattern, it is weak negative.

负相关发生在一个变量的较高值与另一个变量的较低值相关联时。点从左到右呈下降趋势。一个经典的例子是汽车速度和行驶固定距离所需的时间之间的关系:开得越快,行程时间越短。另一个例子是看电视的小时数与体能分数;通常,屏幕时间越多,与较低的体能水平相关。与正相关一样,强度会有所变化。如果点几乎精确地落在一条向下的直线上,则是强负相关;如果它们只是大致遵循这种模式,则是弱负相关。


7. No Correlation | 无相关

When there is no obvious relationship between the two variables, we say there is no correlation. The points are scattered randomly, forming no clear pattern. For example, a person’s shoe size and their IQ would likely show no correlation. Similarly, the number of clouds in the sky and the number of pages in a book have no logical link. It is important not to force a relationship where none exists. Just because two variables change doesn’t mean one causes the other. Always look at the scatter graph and ask yourself: do the points show an upward slope, a downward slope, or just a random spread?

当两个变量之间没有明显关系时,我们说无相关。点随机散布,没有形成清晰的模式。例如,一个人的鞋码和智商很可能显示无相关。同样,天空中云的数量和一本书的页数没有逻辑联系。重要的是不要硬在不存在关系的地方找出关系。仅仅因为两个变量都变化并不意味着一个导致了另一个。始终看着散点图并问自己:点显示出上升趋势、下降趋势,还是仅仅是随机散布?


8. Line of Best Fit | 最佳拟合线

When a scatter graph shows a correlation (positive or negative) that is roughly linear, we can draw a straight line that best represents the data. This is called the line of best fit or trend line. It does not need to pass through all the points, and it often doesn’t pass through any exact point. Instead, it should have roughly the same number of points above and below it, balancing the distances. The line should follow the general direction of the data. To draw it, use a transparent ruler and position it so that the points are evenly distributed on both sides of the line.

当散点图显示出大致线性的相关(正或负)时,我们可以画一条最能代表数据的直线。这称为最佳拟合线或趋势线。它不需要穿过所有的点,而且往往不会精确穿过任何一个点。相反,它应该在其上下方有大致相同数量的点,使距离平衡。线条应顺应数据的总体方向。要画它,使用透明直尺,将尺子定位到使点均匀分布在直线两侧的位置。


9. Using the Line of Best Fit for Prediction | 使用最佳拟合线进行预测

One of the main uses of a scatter graph and its line of best fit is to make predictions. Suppose you have data on the number of hours of sunlight and plant growth. Once you draw the line of best fit, you can estimate the expected growth for a given number of sunlight hours, even if that exact value was not in your original data. To predict a y-value for a chosen x-value, go up from the x-axis to the line, then across to the y-axis. Always be careful when predicting beyond the range of your data; this is called extrapolation and may be unreliable because the trend might not continue.

散点图及其最佳拟合线的主要用途之一是进行预测。假设你有关于日照时数和植物生长的数据。当你画出最佳拟合线后,你可以估计某一给定日照时数下的预期生长,即使该确切数值不在原始数据中。要预测选定 x 值的 y 值,从 x 轴向上到达直线,然后横向读取 y 轴。在数据范围之外进行预测时要格外小心;这称为外推,可能不可靠,因为该趋势可能不会延续。


10. Interpreting Outliers | 解释异常值

An outlier is a point that lies far away from the general pattern of the scatter graph. It might be caused by a measurement error, a special circumstance, or it could be genuine but unusual data. When you see an outlier, you should consider whether to include it when drawing the line of best fit. Usually, a single outlier should not heavily influence the line. If the outlier is due to an obvious mistake, you might decide to exclude it, but always state your reason. Identifying outliers is a higher-order skill tested in Cambridge Checkpoint, as it requires critical thinking about data.

异常值是指远偏离散点图总体模式的点。它可能由测量误差、特殊情况造成,或者可能是真实但不寻常的数据。当你看到一个异常值时,你应考虑在绘制最佳拟合线时是否将其包含在内。通常,单个异常值不应该过于影响直线。如果异常值是明显错误造成的,你可以决定排除它,但要始终说明理由。识别异常值是剑桥 Checkpoint 中考查的高阶技能,因为这需要对数据进行批判性思考。


11. Real-life Applications | 实际应用

Scatter graphs and correlation are used in many fields. In economics, they help show the link between price and demand. In medicine, they can reveal relationships between a drug dosage and its effect. In sports, coaches use scatter graphs to analyse training load and injury rates. In geography, they plot rainfall against crop yield. Anytime you need to see whether two numeric variables move together, a scatter graph is the go-to tool. Mastering this topic not only prepares you for exams but also equips you with a valuable life skill for interpreting data in the news and your own projects.

散点图和相关性被用于许多领域。在经济学中,它们帮助显示价格与需求之间的联系。在医学中,它们可以揭示药物剂量与其效果之间的关系。在体育中,教练使用散点图分析训练负荷与受伤率。在地理中,他们绘制降雨量与作物产量图。每当你需要查看两个数值变量是否共同变动时,散点图便是首选工具。掌握这个主题不仅为考试做准备,还为你提供了一项宝贵的生活技能,用于解读新闻和自己项目中的数据。


12. Summary and Exam Tips | 总结与考试技巧

Always remember: a scatter graph displays bivariate data, correlation describes the relationship, and a line of best fit must be drawn with care, balancing points above and below. Use a sharp pencil and ruler. Label axes, include units, and give your graph a title. When describing correlation, use phrases like “strong positive correlation” rather than just “positive”. When asked to predict, show your working by drawing dashed lines to the line of best fit. Practice with different data sets until plotting and interpreting become second nature. With these skills, you will confidently handle any KS3 data question.

永远记住:散点图用于展示双变量数据,相关性描述关系,最佳拟合线必须用心绘制,使点上下均衡。使用削尖的铅笔和直尺。标注轴,包含单位,并给图表加上标题。在描述相关性时,使用诸如“强正相关”这样的短语,而不只是“正相关”。当要求做预测时,通过画虚线连接到最佳拟合线来展示你的步骤。用不同的数据集进行练习,直到绘图和解释成为第二天性。有了这些技能,你将自信地应对任何 KS3 数据问题。


Published by TutorHao | Mathematics Revision Series | aleveler.com

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading