📚 Key Points for CAIE AS Psychology Experimental and Practical Assessments | CAIE AS 心理学实验与实践考核要点
Mastering the experimental and practical assessment components in CAIE AS Level Psychology (9990) requires a firm grasp of research methodology, from formulating precise hypotheses to interpreting data ethically and accurately. This article outlines the essential skills and knowledge that examiners look for, helping you confidently tackle Paper 2 and related practical-style questions.
要在CAIE AS心理学(9990)的实验与实践考核中取得高分,你需要扎实掌握研究方法,从提出精确的假说到合乎伦理地准确解读数据。本文梳理了考官最看重的关键技能和知识点,助你自信应对Paper 2以及相关的实践类题目。
1. Crafting Testable Hypotheses | 构建可验证的假说
An experimental hypothesis must clearly state the expected effect of the independent variable (IV) on the dependent variable (DV), and it should be operationalised. For the practical exam, you will often need to write a directional (one-tailed) or non-directional (two-tailed) hypothesis depending on previous research. A null hypothesis, which states that any difference or correlation is due to chance, is also essential for statistical testing.
实验假说必须清晰陈述自变量(IV)对因变量(DV)的预期影响,并且需要操作化。在实践考试中,你常常需要根据已有研究写出定向(单尾)或非定向(双尾)假说。零假说同样重要,它声明任何差异或相关都源于偶然因素,是统计检验的基础。
A well-written hypothesis avoids vague language. Instead of ‘memory will be better’, use ‘participants who read a list with semantic encoding will recall significantly more words than those who read the list with structural encoding’. This specificity allows precise replication and measurement.
书写优秀的假说要避免模糊表述。与其写“记忆会更好”,不如写“接受语义编码的参与者将比接受结构编码的参与者回忆出显著更多的单词”。这样的具体性确保了精确的复制和测量。
Typical examination scenarios ask you to identify the IV and DV from a description, operationalise them, and formulate both experimental and null hypotheses. Practise rewriting loosely worded predictions into formal, testable statements.
典型考题会要求你从一段描述中找出IV和DV,将其操作化,并分别写出实验假说和零假说。多练习把松散的口语预测改写为正式、可检验的陈述。
2. Operationalising Variables | 变量的操作化
Operationalisation means defining exactly how a variable will be manipulated or measured in a way that is clear, replicable, and quantifiable. For the IV, you must specify how conditions differ (e.g., ‘listening to 70 dB white noise’ versus ‘silence’). For the DV, you need a concrete measurement (e.g., ‘number of words correctly recalled from a 20-item list’).
操作化是指以清晰、可复制且可量化的方式,精确界定变量的操纵或测量方法。对于IV,你要明确条件如何区分(如“聆听70分贝白噪音”对比“静音”)。对于DV,你需要具体的测量指标(如“从包含20个项目的列表中正确回忆出的单词数量”)。
Weak operationalisation leads to low internal validity. If you simply state ‘the DV is attention’, an examiner cannot tell whether you will measure reaction time, number of lapses on a sustained attention task, or a self-report rating. Always provide the unit of measurement and the precise task instructions.
操作化不当会导致内部效度低下。如果你只是写“DV是注意力”,考官无法判断你将测量反应时间、持续注意任务中的失误次数还是自评量表评分。务必提供测量单位和精确的任务指示。
In Paper 2, you may be asked to improve operational definitions. Look for ambiguity: replace ‘happiness’ with ‘score on the Oxford Happiness Questionnaire’ or ‘number of spontaneous smiles coded from video’. Remember that the DV should yield interval or ratio data whenever possible to allow parametric analysis.
在Paper 2中,你可能会被要求改进操作定义。留意模糊不清的地方:把“快乐”换成“牛津快乐问卷得分”或“从视频中编码的自发微笑次数”。请记住,DV应尽可能产生等距或等比数据,以便进行参数分析。
3. Choosing an Experimental Design | 选择合适的实验设计
The three main designs are independent groups, repeated measures, and matched pairs. Independent groups avoids order effects but requires more participants and is susceptible to participant variables. Repeated measures uses the same participants in all conditions, reducing individual differences but introducing order effects. Matched pairs attempts to control for participant variables by pairing similar participants, yet matching is never perfect.
三种主要设计是独立组设计、重复测量设计和配对组设计。独立组设计避免了顺序效应,但需要更多参与者且容易受到参与者变量的影响。重复测量设计让同一批参与者经历所有条件,减少了个人差异但引入了顺序效应。配对组设计通过匹配相似参与者来控制参与者变量,但匹配永远不可能完美。
Your choice of design should be justified by the nature of the task and the risk of demand characteristics. For example, in a memory study where knowing the aim would influence performance, an independent groups design is often preferred. If you select repeated measures, you must describe counterbalancing or randomisation of condition order to enhance validity.
你的设计选择需要根据任务性质和需求特征的风险来解释。比如,在一项记忆研究中,如果了解实验目的会影响表现,独立组设计往往更优。若选择重复测量设计,你必须说明如何采用平衡顺序或随机化条件顺序来提高效度。
Examiners will also expect you to recognise the limitations of each design. Write about how independent groups could be affected by individual differences in memory span, and propose a solution such as random allocation or a pre-test to check baseline equivalence.
考官也期望你能识别每种设计的局限。写出独立组设计可能如何受到个体记忆广度差异的影响,并提出解决方案,如随机分配或前测以检验基线等同性。
4. Controlling Extraneous Variables | 控制额外变量
Confounding variables threaten internal validity because they systematically change with the IV. Common confounds include environmental factors (noise, lighting), participant expectations, and experimenter bias. Technically speaking, situational variables should be held constant through standardisation, whereas participant variables are managed via design or random assignment.
混淆变量因为它们随IV系统性地变化而威胁内部效度。常见的混淆变量包括环境因素(噪音、光照)、参与者期望和实验者偏差。从技术上说,情境变量应通过标准化程序保持恒定,而参与者变量则通过实验设计或随机分配来管理。
A frequent Paper 2 task is to identify two extraneous variables from a scenario and suggest how they could be controlled. For instance, if an experiment on reaction time uses two different rooms, you could note that room temperature or disturbance might differ, and recommend using the same quiet laboratory for all participants.
Paper 2中常见的任务是让考生从情景中找出两个额外变量,并建议如何控制它们。例如,如果一项反应时实验使用了两个不同的房间,你可以指出室温和干扰可能不同,并建议所有参与者使用同一间安静实验室。
Standardised instructions, double-blind procedures, and consistent timing are your allies. When writing about control, be precise: instead of ‘control the time of day’, specify ‘all participants were tested between 9 am and 11 am to minimise circadian rhythm effects’.
标准化指导语、双盲程序和一致的时间安排是你的有力工具。在描述控制手段时要精确:与其写“控制一天中的时间”,不如写“所有参与者均在上午9点至11点间接受测试,以最小化昼夜节律影响”。
5. Sampling and Generalisability | 抽样与推广性
The target population must be clearly defined before you select a sampling method. Opportunity sampling is quick but likely biased; random sampling is ideal but often impractical; volunteer sampling can attract a particular type of participant. Stratified sampling ensures proportional representation of subgroups and is highly valued in exam responses.
在你选择抽样方法之前,必须清晰界定目标总体。机会抽样快捷但容易有偏差;随机抽样理想但往往不切实际;志愿者抽样可能吸引特定类型的参与者。分层抽样确保各子群体的比例代表,在考试答案中比较受重视。
Generalisability is about the extent to which findings can be applied to the target population. A sample of 20 university psychology students cannot easily generalise to all adults. Always evaluate the sample size, demographic diversity, and cultural context. When suggesting improvements, propose recruitment from multiple clubs, workplaces, or age brackets.
推广性指的是研究发现可以在多大程度上应用到目标总体。一个由20名大学心理学学生组成的样本难以推广到所有成年人。要严谨评估样本量、人口统计多样性和文化背景。在提出改进建议时,可以向来自多个社团、工作场所或年龄段的招募靠拢。
Examiners reward linking sampling choices to ethical and practical constraints. For example, you might acknowledge that while random sampling would be representative, a school-based study may have to rely on an opportunity sample for logistical reasons, and still discuss how to broaden the sample later.
考官喜欢看到考生将抽样选择与伦理和实际限制联系起来。例如,你可以承认虽然随机抽样更具代表性,但基于学校的研究可能由于后勤原因只能依靠机会样本,并依然讨论以后如何拓宽样本。
6. Ethical Considerations in Practice | 实践中的伦理考量
The British Psychological Society (BPS) guidelines for human participants must be embedded in your experimental planning. Core principles include informed consent, right to withdraw, confidentiality, and protection from harm. In the CAIE exam, you are often required to apply these to a proposed study, not just list them.
英国心理学会(BPS)关于人类参与者的指南必须融入你的实验设计中。核心原则包括知情同意、退出权、保密和避免伤害。在CAIE考试中,往往要求将这些原则应用到一个具体的研究计划中,而非简单罗列。
For example, if a study involves a mildly embarrassing task, you should describe how you would obtain fully informed consent through a detailed information sheet, reassure participants they can stop at any time without penalty, and arrange a debrief session to explain the true purpose and offer support.
比如,如果一项研究涉及一个稍显尴尬的任务,你应当描述如何通过详细的信息表获得完全知情同意,安抚参与者他们可以随时无责退出,并安排事后解说以解释真实目的并提供支持。
Confidentiality goes beyond anonymising data; it includes storing questionnaires and recordings securely and restricting access to authorised personnel only. In a practical scenario, you might allocate code numbers to names and keep the linking list in a locked cabinet. These operational details earn top marks.
保密不止于匿名化数据;还包括安全储存问卷和录音,仅限授权人员访问。在一个实践情景中,你可以为每个名字分配代号并将对应名单锁在柜子里。这些操作细节能赢得高分。
With animal studies, always justify necessity, minimise the number of animals, reduce suffering, and consider replacement alternatives. CAIE often includes a short animal ethics question; be prepared to discuss housing, anaesthesia, and euthanasia criteria.
对于动物研究,务必证明其必要性,减少动物数量,降低痛苦,并考虑替代方案。CAIE常包含简短的动物伦理问题;准备好讨论饲养条件、麻醉和安乐死标准。
7. Types of Data and Measurement Levels | 数据类型与测量层次
Data can be quantitative or qualitative. In experimental reports, quantitative data in the form of scores, reaction times, or recall counts is standard. However, you must also identify the level of measurement: nominal (categorical, e.g., recalled/not recalled), ordinal (ranked, e.g., rating scale), or interval/ratio (equal intervals and true zero, e.g., time in seconds).
数据可以是定量或定性的。在实验报告中,以分数、反应时或回忆数量等形式呈现的定量数据是常态。但你还必须识别测量层次:名义(分类,如回忆起/未回忆起)、顺序(排序,如评分量表)或等距/等比(等距间隔且有绝对零点,如以秒为单位的时间)。
The level of measurement dictates which statistical test and graph are appropriate. Nominal data usually requires a chi-square test and bar chart, while ordinal data from a repeated measures design might be analysed with the Wilcoxon signed-rank test and a median with interquartile range. Interval data allows parametric tests like the related t-test if assumptions are met.
测量层次决定了何种统计检验和图表是合适的。名义数据通常需要卡方检验和条形图,而来自重复测量设计的顺序数据可能用Wilcoxon符号秩检验以及中位数和四分位距分析。等距数据在满足假设时可以使用配对t检验等参数检验。
In the exam, you may be given a table of results and asked to justify a choice of descriptive or inferential statistics. Always link your justification to the experimental design, level of measurement, and whether the data are normally distributed. Memorising a decision tree for statistical tests is invaluable.
考试中可能会给你一个结果表,要求你为描述性或推理性统计的选择辩护。务必将你的理由与实验设计、测量层次以及数据是否正态分布联系起来。熟记统计检验决策树非常有价值。
8. Descriptive Statistics and Graphical Representations | 描述统计与图形呈现
Measures of central tendency – mean, median, and mode – summarise the typical score. The mean is the most sensitive but susceptible to outliers; the median is robust for skewed data; the mode is rarely used in AS experiments. Measures of dispersion such as range and standard deviation describe variability. Always pair a central tendency with an appropriate dispersion measure.
集中趋势度量——均值、中位数和众数——概括了典型分数。均值最灵敏但易受异常值影响;中位数对偏态数据稳健;众数在AS实验中很少使用。散布度量如全距和标准差描述了变异性。始终搭配使用一个集中度量和合适的离散度量。
Graphs must be titled clearly, with labelled axes showing the variable names and units. Bar charts are for categorical IVs, histograms for continuous data, and scattergraphs for correlations. In an experimental report, avoid 3D graphics and unnecessary colour effects; precision counts more than aesthetics.
图表必须有清晰标题,坐标轴标注变量名称和单位。条形图用于类别IV,直方图用于连续数据,散点图用于相关分析。在实验报告中,避免使用三维图形和不必要的色彩效果;精准比美观更重要。
A common exam requirement is to sketch a graph and then interpret it. If you plot a bar chart of two conditions, you might add error bars representing standard deviation to show variability and then note that the means appear different but overlap in spread, suggesting further inferential testing is needed.
常见的考试要求是绘制图表然后解读它。如果你画出一个代表两种条件的条形图,可以添加表示标准差的误差线来展示变异性,然后指出均值看起来有差异但分布有重叠,暗示需要进一步的推断检验。
9. Inferential Statistics and Significance | 推断统计与显著性
Inferential tests let you decide whether to reject the null hypothesis. The conventional significance level p ≤ 0.05 means there is less than a 5% probability that the observed effect is due to chance. You must compare the calculated statistic with a critical value from tables, taking into account the degrees of freedom (df), whether the test is one- or two-tailed, and the number of participants (N).
推断检验让你决定是否拒绝零假说。传统的显著性水平p ≤ 0.05意味着观察到的效应由偶然因素导致的概率小于5%。你必须将计算出的统计值与查表得到的临界值比较,并考虑自由度(df)、检验是单尾还是双尾,以及参与者人数(N)。
For AS level, the key tests include the chi-square, Mann-Whitney U, Wilcoxon signed-rank, and related/independent t-tests. You should be able to state which test is appropriate for a given design and data type, calculate it using a provided formula, and interpret the result. Understanding the difference between a Type I error (false positive) and Type II error (false negative) is also frequently assessed.
在AS阶段,关键检验包括卡方检验、Mann-Whitney U检验、Wilcoxon符号秩检验以及相关/独立t检验。你应能够针对给定设计和数据类型选择合适检验,用提供的公式计算,并解释结果。理解I型错误(假阳性)和II型错误(假阴性)的区别也是常见考点。
When the calculated value exceeds the critical value (for most tests), the result is significant. Then you must write a concluding statement that references the hypothesis and the statistical evidence, such as: ‘The calculated U = 18 was less than the critical value of 27 for a one-tailed test at p ≤ 0.05, therefore the null hypothesis is rejected, supporting the experimental hypothesis that participants in condition A rate images as more pleasant.’
当计算值大于临界值(对于多数检验),结果显著。此时你必须写出结合假说和统计证据的结论陈述,例如:“计算出的U = 18 小于单尾检验p ≤ 0.05下的临界值27,因此拒绝零假说,支持实验假说:条件A中的参与者对图片的愉悦度评分更高。”
10. Writing a Coherent Experimental Report | 撰写连贯的实验报告
The standard sections – Abstract, Introduction, Method, Results, Discussion, References – follow APA conventions. In the Method, you must detail the design, participants, apparatus/materials, and procedure with enough precision that someone else could replicate the study exactly. Use past tense and passive voice where appropriate (e.g., ‘Participants were randomly allocated to one of two conditions’).
标准部分——摘要、引言、方法、结果、讨论、参考文献——遵循APA格式。在方法部分,你必须详细说明设计、参与者、仪器/材料以及程序,足够让他人准确复制研究。适当使用过去时和被动语态(例如,“参与者被随机分配到两个条件之一”)。
The Results section presents descriptive and inferential statistics without interpretation. The Discussion interprets findings, relates them to the hypothesis and background research, acknowledges limitations, and suggests modifications. Weaker answers dump all information into one paragraph; stronger answers label clear subsections.
结果部分只呈现描述性和推断性统计,不加解读。讨论部分解读发现,将其与假说和背景研究关联起来,承认局限并提出修改建议。较差的答案把所有信息挤在一段;优秀的答案会标出清晰的子部分。
A practical exam might ask you to write a short abstract or a method section from a scenario. Here, conciseness is vital: an abstract typically includes the aim, hypothesis, brief method, key results, and conclusion in around 150 words. Practice this reduction exercise regularly.
实践考试可能要求你根据情景写一段简短的摘要或方法部分。此时简洁至关重要:摘要通常包含目的、假说、简要方法、主要结果和结论,大约150词。定期练习这种精简撰写。
11. Evaluating Experiments: Validity and Reliability | 评估实验:效度与信度
Internal validity asks whether the IV truly caused the observed change in the DV. Threats include confounding variables, demand characteristics, and social desirability bias. External validity concerns generalisability across people, settings, and times. Ecological validity (mundane realism) is often enhanced by using naturalistic tasks, though at the cost of control.
内部效度探询IV是否真正引起DV的观测变化。威胁包括混淆变量、需求特征和社会期许偏差。外部效度关注跨人群、跨场景和跨时间的普遍性。生态效度(世俗逼真性)通常通过使用自然任务来提高,但会牺牲控制。
Reliability refers to consistency. If the same experiment were repeated with the same design on a similar sample, would the results be similar? Strengthen reliability by standardising procedures, using calibrated instruments, and training researchers to follow a strict protocol. Split-half or test-retest methods can be mentioned for questionnaires used as DV.
信度指一致性。如果相同设计、类似样本下重复该实验,结果是否会相似?通过标准化程序、使用校准仪器并训练研究者严格遵循方案来提高信度。对于用作DV的问卷,可提及分半信度或重测信度。
A top-band evaluation does more than list strengths and weaknesses; it discusses trade-offs. For example, a highly controlled lab experiment has strong internal validity but may lack ecological validity. An examiner wants to see you propose a specific improvement, such as adding a field component while still controlling key confounds.
最高等级的评估不止罗列优缺点,而是讨论取舍。比如,高度控制的实验室实验内部效度强但可能缺乏生态效度。考官希望看到你提出具体的改进,如增加现场环节同时依旧控制关键混淆变量。
12. Common Pitfalls and Examiner Advice | 常见误区与考官建议
Avoid muddling the hypothesis with a null hypothesis or omitting the null altogether. Many students lose marks by writing ‘The hypothesis is there will be a significant difference’ without specifying direction or operationalising variables. Also, do not treat the null hypothesis as a negative prediction; it is a statement of no effect.
避免混淆实验假说与零假说,或完全遗漏零假说。许多学生因写“假设会有显著差异”却不指明方向或未操作化变量而丢分。同样,不要将零假说当作消极的预测;它是一个无效应的陈述。
Another slip is misidentifying the experimental design. If you say ‘repeated measures’ when the scenario clearly describes two separate groups, you will lose marks. Underline the key phrases in the stem that show whether the same participants experience both conditions. Practise with past papers to spot these traps.
另一个失误是错误识别实验设计。如果你说“重复测量设计”,而情景明显描述的是两组独立的参与者,你就会丢分。在题干中划出显示同一批参与者是否经历两种条件的关键短语。用历年真题练习识别这些陷阱。
In the data analysis section, always show your working even if the question does not explicitly command it. Label the columns, write the formula, and substitute values systematically. This helps you avoid careless arithmetic errors and allows partial credit. Lastly, manage time: allocate roughly 1 minute per mark in the exam.
在数据分析部分,即使题目未明确要求,也要写出计算过程。给各列加标题,写下公式并系统地代入数值。这有助于避免粗心算错,并能得到步骤分。最后,管理好时间:考试中大约每分对应一分钟。
Published by TutorHao | Psychology Revision Series | aleveler.com
更多咨询请联系16621398022(同微信)
屏轩国际教育cambridge primary/secondary checkpoint, cat4, ukiset,ukcat,igcse,alevel,PAT,STEP,MAT, ibdp,ap,ssat,sat,sat2课程辅导,国外大学本科硕士研究生博士课程论文辅导