Random Variables and Probability Distributions | 随机变量与概率分布

📚 Random Variables and Probability Distributions | 随机变量与概率分布

A random variable is a fundamental concept in probability and statistics, representing a numerical outcome of a random phenomenon. Understanding random variables and their probability distributions allows us to model uncertainty, make predictions, and draw conclusions from data. In IB Mathematics, you will encounter both discrete and continuous random variables, learn to calculate their expected values and variances, and apply standard distributions such as the binomial, Poisson, and normal distributions to real-world problems.

随机变量是概率与统计中的基本概念,它表示随机现象的数值结果。理解随机变量及其概率分布使我们能够对不确定性进行建模、做出预测并从数据中得出结论。在IB数学中,你将学习离散和连续随机变量,学会计算它们的期望值和方差,并应用二项分布、泊松分布和正态分布等标准分布解决实际问题。

1. Introduction to Random Variables | 随机变量简介

A random variable, often denoted by a capital letter like X, is a function that assigns a real number to each outcome of a random experiment. For example, when rolling a fair six-sided die, we can define X as the number that appears face up. The value of X is determined by chance, but we can describe its behaviour using probabilities. Distinguishing between discrete and continuous random variables is essential, as the mathematical tools used for each differ.

随机变量通常用大写字母(如X)表示,它是一个将随机试验的每个结果映射为实数的函数。例如,掷一枚均匀的六面骰子时,我们可以定义X为朝上的点数。X的取值由机会决定,但我们可以用概率来描述它的行为。区分离散和连续随机变量至关重要,因为用于两者的数学工具各不相同。

2. Discrete Random Variables | 离散随机变量

A discrete random variable is one whose possible values can be counted, typically integers or a finite set of numbers. Examples include the number of heads when tossing three coins, the sum of two dice, or the number of defective items in a batch. The probability distribution of a discrete random variable specifies the probability for each possible value.

离散随机变量是其可能取值可数的变量,通常是整数或有限的一组数。例子包括抛三枚硬币时正面朝上的次数、两个骰子的点数之和,或一批产品中次品的数量。离散随机变量的概率分布给出了每个可能取值的概率。

For a discrete random variable X, the probability mass function (PMF) gives P(X = x) for each value x in the sample space. The sum of all probabilities must equal 1:

对于离散随机变量X,概率质量函数(PMF)给出了样本空间中每个取值x的概率P(X = x)。所有概率之和必须等于1:

∑ P(X = x) = 1

Moreover, each individual probability must satisfy 0 ≤ P(X = x) ≤ 1.

此外,每个单独的概率必须满足0 ≤ P(X = x) ≤ 1。


3. Probability Mass Function (PMF) | 概率质量函数 (PMF)

The probability mass function is the function f(x) = P(X = x). It completely defines the distribution of a discrete random variable. Often presented in a table, the PMF makes it easy to identify the most likely outcomes and to compute probabilities of events by summing individual probabilities. For instance, if X is the sum of two fair dice, the PMF can be listed for values 2 through 12.

概率质量函数是函数f(x) = P(X = x)。它完整地定义了一个离散随机变量的分布。PMF通常以表格形式呈现,使人容易识别最可能的结果,并通过将单个概率相加来计算事件的概率。例如,如果X是两个公平骰子的点数之和,则可以为2到12的值列出PMF。

When working with PMFs, remember that for any event A, P(X ∈ A) = Σ_{x∈A} P(X = x). This additive property is straightforward yet powerful for calculating probabilities of compound events.

在使用PMF时,请记住对于任何事件A,P(X ∈ A) = Σ_{x∈A} P(X = x)。这一可加性质虽然简单,但对于计算复合事件的概率非常有用。


4. Expected Value and Variance of Discrete Random Variables | 离散随机变量的期望与方差

The expected value, or mean, of a discrete random variable X, denoted E(X) or μ, is the probability-weighted average of its possible values:

离散随机变量X的期望值(或称均值),记作E(X)或μ,是其可能值的概率加权平均:

E(X) = μ = ∑ x · P(X = x)

It represents the “long-run” average outcome if the experiment were repeated many times.

它表示如果实验重复多次,得到的“长期”平均结果。

The variance, Var(X) or σ², measures the spread of the distribution around the mean. It is defined as the expected squared deviation from the mean:

方差,Var(X)或σ²,衡量分布围绕均值的离散程度。其定义为与均值偏差平方的期望值:

Var(X) = E[(X – μ)²] = ∑ (x – μ)² · P(X = x)

A computationally convenient formula is often used: Var(X) = E(X²) – [E(X)]². The standard deviation is σ = √Var(X).

通常使用一个计算简便的公式:Var(X) = E(X²) – [E(X)]²。标准差是σ = √Var(X)。


5. Binomial Distribution | 二项分布

A binomial distribution describes the number of successes in a fixed number of independent trials, each with the same probability of success p. If we let X be the number of successes in n trials, then X follows a binomial distribution, denoted X ~ B(n, p). The probability of obtaining exactly k successes is given by:

二项分布描述了在固定次数的独立试验中成功的次数,每次试验成功的概率相同为p。如果我们令X为n次试验中成功的次数,则X服从二项分布,记作X ~ B(n, p)。恰好获得k次成功的概率由下式给出:

P(X = k) = ⁿCₖ pᵏ (1 – p)ⁿ⁻ᵏ

where ⁿCₖ = n! / [k!(n – k)!] is the binomial coefficient. The distribution is symmetric when p = 0.5 and skewed otherwise.

其中ⁿCₖ = n! / [k!(n – k)!] 是二项式系数。当p=0.5时分布对称,否则偏斜。

The mean and variance of a binomial random variable are simple: E(X) = np, Var(X) = np(1 – p). These arise directly from the properties of sums of independent Bernoulli trials.

二项随机变量的均值和方差很简单:E(X) = np, Var(X) = np(1 – p)。这些直接源于独立伯努利试验之和的性质。


6. Poisson Distribution | 泊松分布

The Poisson distribution models the number of events occurring in a fixed interval of time or space, assuming events happen independently and at a constant average rate λ (lambda). If X represents the number of occurrences, we write X ~ Po(λ). Its probability mass function is:

泊松分布用于模拟在固定时间或空间间隔内事件发生的次数,假设事件独立发生且具有恒定的平均速率λ(lambda)。如果X表示发生的次数,我们写作X ~ Po(λ)。其概率质量函数为:

P(X = k) = (λᵏ e⁻λ) / k!

where k = 0, 1, 2, … and e ≈ 2.71828. The Poisson distribution is often used for rare events, such as the number of phone calls per minute at a call centre or the number of flaws per metre of fabric.

其中k = 0, 1, 2, …,e ≈ 2.71828。泊松分布常用于稀有事件,比如呼叫中心每分钟的电话数或每米布料上的瑕疵数。

For the Poisson distribution, the mean and variance are both equal to λ: E(X) = λ, Var(X) = λ. This unique property helps identify Poisson-like data. Additionally, the binomial distribution B(n, p) can be approximated by Po(np) when n is large and p is small.

泊松分布的均值和方差都等于λ:E(X) = λ, Var(X) = λ。这一独特性质有助于识别类似泊松的数据。此外,当n很大且p很小时,二项分布B(n, p)可用Po(np)来近似。


7. Continuous Random Variables | 连续随机变量

A continuous random variable can take any value within an interval on the real number line. Examples include the height of a student, the time taken to complete a task, or the temperature of a room. Because there are infinitely many possible values, the probability that X equals exactly any single value is zero; instead, we consider probabilities over intervals.

连续随机变量可以在实数轴上的某个区间内取任何值。例子包括学生的身高、完成任务所需的时间或房间的温度。由于有无限多个可能的值,X恰好等于任何一个精确值的概率为零;相反,我们考虑区间上的概率。

Thus, we describe continuous distributions using a probability density function (PDF) rather than a mass function. The area under the density curve over an interval gives the probability that X falls within that interval.

因此,我们用概率密度函数(PDF)而不是质量函数来描述连续分布。区间上密度曲线下方的面积表示X落在该区间内的概率。


8. Probability Density Function (PDF) | 概率密度函数 (PDF)

For a continuous random variable X, the PDF, denoted f(x), satisfies two conditions: f(x) ≥ 0 for all x, and the total area under the curve equals 1:

对于连续随机变量X,PDF,记作f(x),满足两个条件:对所有x,f(x) ≥ 0,且曲线下方的总面积等于1:

∫₋∞⁺∞ f(x) dx = 1

The probability that X lies in the interval [a, b] is the definite integral of the PDF over that interval:

X落在区间[a, b]内的概率是PDF在该区间上的定积分:

P(a ≤ X ≤ b) = ∫ₐᵇ f(x) dx

Note that P(X = c) = 0 for any single point c. This is a key difference from discrete distributions and means that including or excluding the endpoints does not change the probability of an interval.

注意,对于任何单点c,P(X = c) = 0。这是与离散分布的一个关键区别,意味着包含或不包含端点不会改变区间的概率。


9. Cumulative Distribution Function (CDF) | 累积分布函数 (CDF)

The cumulative distribution function, F(x), gives the probability that the random variable X is less than or equal to a particular value x: F(x) = P(X ≤ x). For a continuous random variable, the CDF is the integral of the PDF from –∞ to x:

累积分布函数F(x)给出了随机变量X小于或等于某一特定值x的概率:F(x) = P(X ≤ x)。对于连续随机变量,CDF是PDF从–∞到x的积分:

F(x) = ∫₋∞ˣ f(t) dt

The CDF is always non-decreasing and approaches 0 as x → –∞ and 1 as x → +∞. For any interval, P(a < X ≤ b) = F(b) – F(a), which provides a convenient way to compute probabilities without direct integration.

CDF总是非递减的,当x → –∞时趋近于0,当x → +∞时趋近于1。对于任意区间,P(a < X ≤ b) = F(b) – F(a),这提供了一种无需直接积分即可计算概率的便捷方法。

For discrete random variables, the CDF is a step function obtained by summing the probabilities of all values less than or equal to x. Both forms are essential in statistical inference and hypothesis testing.

对于离散随机变量,CDF是通过将所有小于或等于x的值的概率相加而得到的阶梯函数。这两种形式在统计推断和假设检验中都是必不可少的。


10. Mean and Variance of Continuous Random Variables | 连续随机变量的均值与方差

The expected value (mean) of a continuous random variable X with PDF f(x) is defined analogously to the discrete case but using an integral:

具有PDF f(x)的连续随机变量X的期望值(均值)的定义与离散情况类似,但使用积分:

E(X) = μ = ∫₋∞⁺∞ x · f(x) dx

The variance, which measures dispersion, is again the expected squared deviation from the mean:

衡量分散程度的方差同样是偏离均值的平方的期望值:

Var(X) = ∫₋∞⁺∞ (x – μ)² f(x) dx = E(X²) – μ²

Calculating these quantities often requires techniques of integration, but for well-known distributions like the normal distribution, formulas are standardised.

计算这些量通常需要积分技巧,但对于像正态分布这样的常见分布,公式是标准化的。


11. Normal Distribution | 正态分布

The normal distribution is the most important continuous probability distribution. A random variable X that is normally distributed with mean μ and variance σ² is written X ~ N(μ, σ²). Its probability density function is the famous bell-shaped curve:

正态分布是最重要的连续概率分布。均值为μ、方差为σ²的正态分布随机变量X记作X ~ N(μ, σ²)。它的概率密度函数是著名的钟形曲线:

f(x) = (1 / (σ√(2π))) e^[–(x – μ)² / (2σ²)]

The normal distribution is symmetric about the mean, and the empirical rule states that roughly 68% of data lie within 1σ, 95% within 2σ, and 99.7% within 3σ of the mean. Many natural phenomena, like heights, weights, and measurement errors, approximately follow a normal distribution.

正态分布关于均值对称,经验法则指出大约68%的数据落在均值的1σ范围内,95%落在2σ内,99.7%落在3σ内。许多自然现象,如身高、体重和测量误差,都近似服从正态分布。


12. Standard Normal Distribution and Z-scores | 标准正态分布与Z分数

To calculate probabilities for any normal distribution, we convert to the standard normal distribution Z ~ N(0, 1), which has mean 0 and variance 1. The transformation is done by standardising:

要计算任意正态分布的概率,我们将其转换为标准正态分布Z ~ N(0, 1),其均值为0,方差为1。转换通过标准化完成:

Z = (X – μ) / σ

The resulting Z-score represents how many standard deviations an observation is from the mean. Values of the cumulative distribution function for Z, denoted Φ(z), are widely tabulated or provided by calculators, allowing us to find P(Z ≤ z) and therefore probabilities for any normal variable X.

得到的Z分数表示一个观测值距离均值多少个标准差。Z的累积分布函数值,记作Φ(z),被广泛地列成表格或由计算器提供,使我们能够找到P(Z ≤ z),从而计算任何正态变量X的概率。

Mastering the use of Z-scores and the standard normal table is crucial for solving IB problems involving normal distributions, such as finding unknown means or standard deviations or performing inverse normal calculations.

掌握Z分数和标准正态表格的使用对于解决IB中涉及正态分布的问题至关重要,例如求未知的均值或标准差,或进行逆正态计算。


Published by TutorHao | Mathematics Revision Series | aleveler.com

Find IB Maths Textbooks on eBay UK

New, used and second-hand copies of textbooks and revision guides are often much cheaper than retail — check current listings and prices before you buy.

Browse on eBay UK →

更多咨询请联系16621398022(同微信)

Comments

屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导Cancel reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from aleveler.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Exit mobile version