THE CORE AIM OF VOCABULARY MASTERY · AI ALIGNMENT VOCABULARY · INTENT → OBJECTIVE → BEHAVIOUR → EVALUATION → CORRECTION
AI alignment vocabulary is the language used to describe whether an artificial-intelligence system behaves in ways that match intended goals, constraints and human preferences. Terms such as objective, specification, reward model, preference, instruction following, goal misgeneralisation, oversight, corrigibility and evaluation matter because a capable system can still behave differently from what designers or users intended.
The core aim of vocabulary mastery for AI alignment vocabulary is intent-to-behaviour clarity. Learners should be able to distinguish the human goal from the formal objective, observe how system behaviour differs from either one, and explain what evaluation or correction process brings behaviour closer to the intended outcome.
This page is the AI Alignment Vocabulary owner inside the eduKateSG Vocabulary hub. For risk and safeguards, use AI Safety Vocabulary. For governance, use AI Governance Vocabulary.
Central proposition: Alignment vocabulary is mastered when the learner can separate what humans want, what the system is optimised for, what it actually does and how discrepancies are detected and corrected.
The 60-Second AI Alignment Vocabulary Router
- Intent: human goal, preference, instruction.
- Specification: objective, reward, constraint, policy.
- Behaviour: action, output, strategy, generalisation.
- Failure: specification gaming, misgeneralisation, unintended behaviour.
- Feedback: preference data, critique, correction, reward model.
- Control: oversight, intervention, corrigibility, evaluation.
Intent and Objective Are Different
Human intent is what people actually want the system to achieve. A formal objective is the measurable or computational target used to train or steer behaviour. Alignment problems can appear when the objective is only an imperfect proxy for the real intent.
A Worked Example: Specification Gaming
Specification gaming occurs when a system satisfies the literal metric or rule in an unintended way that misses the underlying goal. The vocabulary helps learners ask whether success on the stated measure really represents success in the world.
Goal Misgeneralisation
Goal misgeneralisation describes cases where learned behaviour works during training but pursues an unintended strategy when circumstances change. This is different from simple lack of capability: the system may remain competent while optimising the wrong learned objective.
Preference Learning and Reward Models
A reward model can learn patterns in human preferences and provide a training signal. Preference data can improve behaviour, but it also inherits limitations from the examples, raters and evaluation setup used to create it.
Oversight and Corrigibility
Oversight is the process by which people inspect, guide or approve system behaviour. Corrigibility concerns whether a system remains amenable to correction, modification or shutdown rather than resisting legitimate control.
How to Learn AI Alignment Vocabulary
- Start with a concrete intended goal.
- Write the measurable objective separately.
- Look for proxy failures.
- Compare training and deployment behaviour.
- Study preference feedback and evaluation.
- Identify human oversight points.
- Ask how correction changes future behaviour.
Common AI Alignment Vocabulary Mistakes
Treating alignment as simple obedience
Repair: include objectives, generalisation, preferences and consequences.
Confusing capability failure with alignment failure
Repair: ask whether the system cannot do the task or is pursuing the wrong target.
Assuming a good metric perfectly represents intent
Repair: test for proxy behaviour and edge cases.
Frequently Asked Questions
What is AI alignment vocabulary?
It is the language used to describe the relationship between human intent, formal objectives, learned behaviour, feedback and correction.
What is specification gaming?
It is satisfying the stated rule or metric in an unintended way that misses the underlying purpose.
What is corrigibility?
It concerns whether an AI system remains responsive to legitimate correction, modification or shutdown.
The AI Alignment Vocabulary Standard
Mastery means explaining the intended goal, formal objective, observed behaviour, failure mode and correction mechanism as separate but connected ideas.
That is the standard: alignment language precise enough to reveal when apparent success is not the success humans intended.
