Scatter Graphs and Correlation | 散点图与相关性

📚 Scatter Graphs and Correlation | 散点图与相关性

Scatter graphs are used to display the relationship between two sets of data. Each point on the graph represents a pair of values. By looking at the pattern of points, we can describe the correlation between the variables. Understanding scatter graphs helps us make predictions and analyse data in real life, such as comparing height and shoe size or temperature and ice cream sales.

散点图用来显示两组数据之间的关系。图上的每一个点都代表一对数值。通过观察点的分布模式,我们可以描述变量之间的相关性。理解散点图有助于我们在现实生活中进行预测和数据分析,例如比较身高与鞋码、温度与冰淇淋销量之间的关系。


1. What is a Scatter Graph? | 什么是散点图?

A scatter graph (or scatter plot) is a type of diagram that uses Cartesian coordinates to display values for two variables from a set of data. The data is displayed as a collection of points, each having the value of one variable determining the position on the horizontal axis and the value of the other variable determining the position on the vertical axis.

散点图(或散点图)是一种利用笛卡尔坐标来显示一组数据中两个变量数值的图表。数据以点的集合形式显示,每个点的横坐标由其中一个变量的值决定,纵坐标由另一变量的值决定。

Scatter graphs are particularly useful for spotting patterns or trends. They allow you to see at a glance whether there is a link between two sets of measurements. For example, a teacher might plot test scores in mathematics against test scores in science to see if students who do well in one subject also tend to do well in the other.

散点图特别适合用来发现模式或趋势。它让你一眼就能看出两组测量数据之间是否存在关联。例如,老师可以把数学考试成绩和科学考试成绩的数据绘制成散点图,看看数学学得好的学生是否科学也往往学得好。


2. Plotting Points on a Scatter Graph | 在散点图上描点

To plot a point, you need a pair of values (x, y). The first number is plotted along the horizontal axis (x-axis), and the second number along the vertical axis (y-axis). You place a small cross or dot where the two values meet. Repeating this for all data pairs builds up the scatter graph.

要描一个点,你需要一对数值 (x, y)。第一个数字沿着水平轴(x 轴)确定位置,第二个数字沿着垂直轴(y 轴)确定位置。在这两个数值交汇的地方画一个小叉或圆点。对所有数据对重复这一步骤,就构成了散点图。

It is important to choose sensible scales for the axes. The scale does not have to start at zero, but it should be easy to read and spread the points across the graph paper. Always label both axes with the variable name and unit if necessary.

为坐标轴选择合适的刻度非常重要。刻度不一定非要起始于零,但要容易读数,并且能让点分布在整张图上。记得给横轴和纵轴都标上变量名称,必要时还要标明单位。


3. Variables and Axes | 变量与坐标轴

In correlation studies, we often have an independent variable and a dependent variable. The independent variable is usually plotted on the x-axis. For example, if you are investigating how revision time affects test marks, revision time is the independent variable and goes on the x-axis, while test marks are on the y-axis.

在相关性研究中,我们通常有一个自变量和一个因变量。自变量一般标在 x 轴上。例如,如果你在调查复习时间对考试成绩的影响,复习时间就是自变量,放在 x 轴;考试成绩是因变量,放在 y 轴。

Sometimes it does not matter which variable goes on which axis, especially when you are simply looking for an association. However, once you decide, keep it consistent and label clearly.

有时候把哪个变量放在哪个轴上并不重要,尤其是在你只是单纯寻找关联的时候。但一旦决定了,就要保持一致并标注清楚。


4. Types of Correlation | 相关性的类型

Correlation describes the direction and nature of the relationship between the two variables. There are three main types: positive correlation, negative correlation and no correlation. Recognising these helps us understand the data.

相关性描述了两个变量之间关系的方向和性质。主要有三种类型:正相关、负相关和无相关。识别出属于哪种类型有助于我们理解数据。

We describe correlation by looking at the general slope of the points. If the points go upwards from left to right, it is positive. If they go downwards, it is negative. If the points are scattered randomly with no pattern, there is no correlation.

我们通过观察点群的整体倾斜方向来描述相关性。如果点从左到右呈上升趋势,就是正相关;如果呈下降趋势,就是负相关;如果点随机分布,没有任何趋势,就是无相关。


5. Positive Correlation | 正相关

Positive correlation means that as one variable increases, the other variable also tends to increase. The points on the scatter graph will generally slope upwards. For example, the more hours a student spends revising, the higher their test score is likely to be.

正相关意味着一个变量增加,另一个变量也倾向于增加。散点图上的点整体会呈现向上倾斜的趋势。例如,学生花在复习上的时间越长,他的考试成绩就可能越高。

Other real-life examples of positive correlation include: height and arm span, temperature and cold drink sales, age and vocabulary size in children. The relationship can be strong or weak, but the general direction is upward.

其他正相关的生活实例包括:身高与臂展、温度与冷饮销量、儿童年龄与词汇量。这种关系可能强也可能弱,但总体方向是向上的。


6. Negative Correlation | 负相关

Negative correlation means that as one variable increases, the other variable tends to decrease. The points on the scatter graph slope downwards. An example is the number of hours watching television and test scores: more TV time might be linked to lower scores.

负相关意味着一个变量增加,另一个变量倾向于减少。散点图上的点会向下倾斜。例如,看电视的小时数与考试成绩之间:看电视时间越长,分数可能越低。

Other examples include the distance a car has travelled and the amount of fuel left in the tank, or the price of an item and the quantity demanded. These inverse relationships are very common in science and economics.

其他例子包括汽车行驶的距离与油箱中剩余的燃油量,或者商品价格与需求量之间的关系。这类反比关系在科学和经济学中非常常见。


7. No Correlation | 无相关

No correlation occurs when there is no obvious pattern or clear relationship between the two variables. The points are spread out randomly, showing that changes in one variable do not predict changes in the other. For instance, a student’s height and their mathematics test score are likely to show no correlation.

当两个变量之间没有明显的模式或清晰的关系时,就是无相关。这些点随机分布,表明一个变量的变化无法预测另一个变量的变化。例如,学生的身高和数学考试成绩之间很可能不存在相关性。

It is just as important to recognise when there is no correlation as it is to find positive or negative correlations. Not all pairs of variables are linked, and trying to force a conclusion from random points is a common mistake.

识别出无相关和发现正相关或负相关同样重要。并不是所有变量对之间都存在联系,试图从随机分布的点中强行得出结论是一个常见错误。


8. Strength of Correlation | 相关性的强弱

Correlation is not just about direction; we also describe its strength. If the points lie very close to a straight line, the correlation is strong. If they are more spread out around a general trend, the correlation is weak. For example, height and weight usually show a moderately strong positive correlation in adults.

相关性不仅仅有方向,我们还会描述它的强弱。如果点非常靠近一条直线,这种相关性就很强。如果点围绕一个大致趋势分布得更散,相关性就较弱。例如,成年人的身高和体重通常表现出中等偏强的正相关。

We can use words like “strong positive correlation”, “weak negative correlation” or “moderate positive correlation” to be precise. A scatter graph can show the same direction but with very different strengths.

我们可以精确地使用“强正相关”、“弱负相关”或“中等正相关”这样的表述。同一个方向的相关性在散点图上可能呈现出截然不同的强弱程度。


9. Line of Best Fit | 最佳拟合线

When the scatter graph shows a reasonably strong correlation, we can draw a line of best fit. This is a straight line that goes through the middle of the points, following the trend. It should have roughly the same number of points above and below the line.

当散点图显示出相当强的相关性时,我们可以画一条最佳拟合线。这是一条穿过点群中间且顺应趋势的直线。线上方和线下方的点应该大致数量相等。

The line of best fit is drawn by eye and does not need to pass through the origin or any specific point. It is used to model the relationship and to make predictions. A line of best fit is also called a trend line.

最佳拟合线是目测画出来的,不需要通过原点或任何特定点。它用来建立关系模型并进行预测。最佳拟合线也称为趋势线。


10. Outliers | 离群值

An outlier is a point that lies far away from the general pattern of the other points on a scatter graph. Outliers can be caused by measurement errors, unusual circumstances or simply natural variation. For example, if one student sleeps 2 hours and scores 95%, that point might be an outlier if everyone else follows a clear trend.

离群值是指散点图上远离其他点整体分布模式的一个点。离群值可能由测量误差、异常情况或仅仅是自然变异造成。例如,如果其他学生都呈现明显趋势,但有一个学生只睡了2小时却得了95%,那么这个点可能就是离群值。

Outliers are important because they can affect the position of the line of best fit if not treated carefully. Sometimes it is sensible to disregard them when drawing the line, but you should always assess whether they are genuine data points first.

离群值很重要,因为如果不小心处理,它们会影响最佳拟合线的位置。有时候在画最佳拟合线时忽略它们是合理的,但你应该始终先判断这些点是否是真实的数据点。


11. Using the Line of Best Fit to Make Predictions | 使用最佳拟合线进行预测

Once the line of best fit is drawn, we can use it to estimate unknown values. Predicting within the range of the data is called interpolation. For example, if you have data for 10 to 20 hours of revision, you could use the line to estimate a score for 15 hours. Interpolation is usually reliable.

画出最佳拟合线后,我们可以用它来估算未知值。在已知数据范围内进行的预测称为内插。例如,如果你拥有10到20小时复习时间的数据,你可以用这条线估算复习15小时的分数。内插通常是可靠的。

Predicting outside the range of data is called extrapolation. If you extend the line beyond the data points to predict a value for 30 hours of revision, you are extrapolating. Extrapolation can be risky because the trend might not continue in the same way.

在已知数据范围之外进行的预测称为外推。如果你把线延伸到数据点之外,去预测复习30小时的分数,这就是外推。外推是有风险的,因为趋势可能不会以相同的方式继续下去。


12. Correlation Does Not Imply Causation | 相关性并不意味着因果关系

A very common and critical mistake is to assume that because two variables are correlated, one causes the other. This is not necessarily true. There might be a third hidden factor influencing both, or the correlation may be purely coincidental. For instance, ice cream sales and drowning incidents are positively correlated, but eating ice cream does not cause drowning. The hidden factor is hot weather.

一个非常普遍且关键的错误是,因为两个变量相关,就假定一个是导致另一个的原因。这不一定正确。可能存在第三个隐藏因素同时影响着两者,或者这种相关性纯属巧合。例如,冰淇淋销量与溺水事件呈现正相关,但吃冰淇淋并不会导致溺水。隐藏的因素是炎热的天气。

Always remember: correlation shows an association, not a cause-and-effect relationship. When interpreting scatter graphs, be cautious about making causal claims unless you have evidence from a controlled experiment.

永远记住:相关性显示的是关联,而不是因果关系。在解读散点图时,要谨慎做出因果断言,除非你有来自对照实验的证据。


Published by TutorHao | Mathematics Revision Series | aleveler.com

Find Scatter Graphs Correlation Textbooks on eBay UK

New, used and second-hand copies of textbooks and revision guides are often much cheaper than retail — check current listings and prices before you buy.

Browse on eBay UK →

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading