VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Super Intelligence Works | Preference Learning — How Human and AI Judgments Shape Model Behaviour

Preference learning teaches a model which of several plausible outputs people or evaluators consider better. Many tasks do not have one exact target string. Two summaries can both be factually correct; one can be clearer. Two code explanations can both work; one can be safer. Two assistant replies can both answer the question; one can follow the user’s constraints more faithfully.

Preference learning turns those comparative judgments into training signal. Instead of saying only “this response is correct”, the dataset can say “for this prompt, response A is preferred to response B”. Models can then be trained to reproduce the qualities reflected in those choices.

This article explains how preference learning works in Super Intelligence: pairwise comparisons, chosen and rejected responses, reward models, Bradley–Terry-style probability models, RLHF, DPO, AI feedback, rater disagreement, bias, reward hacking, calibration, evaluation and the difference between preference and truth.

The paper Deep reinforcement learning from human preferences demonstrated learning complex behaviours from pairwise human judgments. Learning to summarize from human feedback later showed that a learned reward model from human comparisons could improve summary quality. The InstructGPT work extended preference-based training to general instruction following, and Direct Preference Optimization proposed a simpler direct objective over preference pairs.

Previous: 027 — Instruction Following. Here we isolate the comparative feedback mechanism that often shapes assistant behaviour after supervised tuning.


The Hidden Transition: Correct Is Not the Same as Preferred

Suppose a user asks, “Explain photosynthesis to a 12-year-old.” Response A is scientifically correct but uses university-level vocabulary. Response B is equally correct and uses a clear analogy appropriate to the audience. A binary factual label marks both correct. A preference comparison can express that B better satisfies the task.

Preference learning is therefore useful when quality is multi-dimensional and difficult to encode as one exact reference answer.

The danger is equally important: what raters prefer is not automatically true, fair or safe. Preference data is a measurement of judgments under instructions and context.

The Basic Preference Pair

A preference example usually contains a prompt x, a chosen response y+ and a rejected response y−. The chosen response is judged better according to the evaluation criteria.

The pair does not necessarily say the rejected response is completely bad. It says the chosen response should receive higher preference under the defined task.

Pairwise Comparisons Reduce the Writing Burden

Writing a perfect target response can be difficult. Comparing two candidates can be easier. A rater can notice that one answer invents a fact or ignores a constraint even if they would struggle to compose the ideal answer from scratch.

This is one reason pairwise preference collection scales well for open-ended generation.

Preferences Can Encode Multiple Qualities at Once

A rater may consider correctness, relevance, tone, concision, source discipline and safety simultaneously. The chosen response wins on the overall rubric.

That compression is useful but opaque. If the system later becomes verbose, we may not know whether raters repeatedly rewarded length as a hidden proxy for quality.

A Simple Pairwise Probability Model

Reward models often learn a scalar score r(x,y) for a response. A Bradley–Terry-style model can convert the score difference into a probability that y+ is preferred over y− using a logistic function.

If r+ is much larger than r−, the model predicts a high probability that the chosen response wins. Training adjusts reward-model parameters so observed human choices receive higher likelihood.

Worked Preference Probability

Suppose reward scores are rA = 2.0 and rB = 0.5. The score difference is 1.5. Passing 1.5 through a sigmoid gives about 0.82. The reward model therefore predicts roughly an 82% chance that A is preferred under this simplified model.

This is a learned preference probability, not an 82% probability that A is factually true.

Reward Models

A reward model takes prompt-response pairs and predicts a scalar aligned with preference judgments. It is trained on comparison data, then used to score new candidate responses.

The reward model becomes a proxy for human evaluation. It can be queried millions of times more cheaply than asking humans for every policy update during reinforcement learning.

Reward Models Can Be Wrong

A reward model can prefer polished nonsense, longer answers, confident tone or superficial citation patterns if those correlate with training labels.

It must be evaluated on held-out preference examples, hard cases and distribution shifts.

RewardBench was created specifically to evaluate reward models across chat, reasoning and safety comparisons, illustrating that reward-model quality is itself a substantial research problem.

Human Preference Data

Human raters compare candidate responses under a guideline. They may be contractors, domain experts, product researchers or users depending on the task.

The rater population matters. Preferences learned from one group should not be casually described as universal human values.

Rater Instructions Define the Target

If the guideline says “prefer the answer that is most helpful, correct and harmless”, raters must resolve trade-offs among those qualities. Different interpretations create label noise.

Detailed examples and edge cases improve consistency.

Rater Screening

Some projects screen raters with test questions to ensure they understand the task. Domain-specific datasets may require qualified expertise.

A strong general writer is not automatically qualified to judge a subtle medical or legal answer.

Rater Agreement

If three raters choose A and two choose B, the preference is weak. If all five choose A, the signal is stronger.

Datasets can preserve vote counts or confidence instead of collapsing every comparison into one certain label.

Ties

Two responses can be equally good. Forcing a winner invents preference information that does not exist.

Interfaces should support ties, both-bad outcomes or uncertainty when those distinctions matter.

Position Bias

Raters can prefer the first or second response simply because of presentation order. Randomising order reduces systematic bias.

Other interface choices—font, length visibility, source display—can also influence judgments.

Length Bias

Longer answers can look more complete. Preference datasets sometimes reward verbosity even when the additional words add little value.

Length-controlled comparisons help measure whether the rater or reward model prefers substance or simply more text.

Style Bias

Polished grammar and confident tone can hide factual errors. If raters skim, the reward model may learn to reward fluency over correctness.

Hard preference examples should include fluent wrong answers versus plain correct answers.

Sycophancy

A response that agrees with a user’s false claim can feel pleasant. Preference training can therefore encourage sycophancy if raters overvalue agreement.

Training should include cases where respectful correction is preferred over flattering error.

Preference Versus Truth

Human preference is evidence about utility or desirability, not a truth oracle. A majority can prefer a false explanation.

Factual tasks should combine preference optimisation with source-grounded or objectively checkable evaluations.

Preference Versus Safety

A user may prefer a response that completes a dangerous action, while the system is designed to preserve safety constraints.

Product-level objectives can combine human helpfulness preferences with safety policies rather than treating raw user preference as the only reward.

Preference Versus Legitimacy

In institutional systems, the authorised person’s preference can matter more than a random rater’s preference. A model should not learn to optimise for popularity when the task has formal rules.

Preference learning is one signal inside a larger governance structure.

From Preference Data to RLHF

Classic RLHF trains a reward model from pairwise comparisons, then uses reinforcement learning to update the policy so generated responses receive higher reward.

The process turns static human comparisons into a scalable reward signal for policy optimisation.

Summarisation From Human Feedback

The Stiennon et al. summarisation work compared human preferences among summaries, trained a reward model and then optimised a summarisation policy with reinforcement learning.

The study highlighted an important principle: automatic metrics such as ROUGE can be weak proxies for the summary quality humans actually care about.

Preference Optimisation for General Assistants

InstructGPT extended this idea to broad user prompts. Human labelers ranked model outputs, and a reward model learned those rankings.

The policy was then trained to produce outputs more aligned with the preference model while preserving capabilities from pretraining.

DPO

Direct Preference Optimization removes the need to separately train a reward model and run a policy-gradient RL loop under its formulation. It optimises chosen-versus-rejected likelihood relationships directly.

This makes training simpler and often more stable operationally, but preference data remains the central source of behavioural direction.

DPO Intuition

For a prompt, the chosen response should become relatively more likely than the rejected response compared with a reference model. The objective encourages that relative shift.

If the chosen answer is preferred for the wrong reason, the optimisation faithfully learns the wrong signal.

Preference Data Quality

A pair is useful when the reason for preference is meaningful and the candidates are sufficiently comparable. “Correct answer versus random garbage” is easy but teaches less subtle judgment than “correct concise answer versus plausible answer with one hidden factual error”.

Hard pairs improve discriminative learning.

Near-Tie Pairs

Pairs with tiny quality differences can be noisy but valuable for fine distinctions. Rater agreement should be monitored.

A dataset composed only of obvious pairs can produce a reward model that fails on realistic close calls.

Out-of-Distribution Preferences

A reward model trained on casual chat may not judge mathematical proofs or medical advice reliably.

RewardBench and multilingual reward-model studies show that reward-model performance can vary sharply across domains and languages.

Multilingual Preference Learning

Preference data is often English-heavy. A model can therefore have weaker alignment in other languages even when the base model is multilingual.

Rater guidelines, cultural norms and translation quality all affect preference consistency across languages.

Preference Learning and Culture

Politeness, directness, humour and acceptable disagreement vary culturally. A globally deployed assistant should not assume one rater population defines the only good conversational style.

Regional evaluations can surface where preferences diverge.

Preference Learning and Expertise

A layperson can judge tone but may mis-rank technical correctness. Expert raters should handle specialised comparisons where factual precision matters.

Hybrid rubrics can separate technical correctness from communication quality.

Multi-Objective Reward

A system may need to optimise helpfulness, correctness, safety, concision and tool reliability. One scalar reward compresses these dimensions.

Multi-objective evaluation can preserve separate metrics and apply product-specific constraints rather than assuming one number captures everything.

Pareto Trade-offs

One response can be safer but less helpful, another more detailed but slower. There may be no single answer that dominates every dimension.

The product needs an explicit policy for these trade-offs. Preference learning cannot eliminate value choices.

Reward Hacking

If the policy discovers a pattern that earns reward without satisfying the underlying intent, it can exploit the reward model.

Examples include excessive length, repeated safety phrases, fake citations or stylistic tricks. The optimiser does not know they are tricks; it sees high reward.

Over-Optimisation

As policy optimisation pushes harder against a fixed reward model, true human preference can eventually stop improving or decline because the policy exploits reward-model weaknesses.

Hold-out human evaluation and reward-model uncertainty can help detect this regime.

Reward Model Uncertainty

A scalar reward often hides confidence. An ensemble of reward models or explicit uncertainty estimates can identify comparisons where evaluators disagree or the input is unfamiliar.

Uncertain cases can be routed to humans rather than blindly optimised.

AI Feedback

A strong language model can compare candidate responses under a rubric, generating preference labels cheaply. This is commonly called AI feedback or RLAIF in reinforcement-learning setups.

AI evaluators can scale but share model biases. Human audits and diverse evaluators remain important.

Constitutional AI Preferences

A model can judge candidates against written principles such as privacy, helpfulness and non-deception. The constitution creates a structured preference rubric.

The principle set still comes from human design and can contain conflicts or blind spots.

LLM-as-a-Judge

Language models are increasingly used as evaluators because they can compare long-form outputs. Judge prompts specify criteria and often require explanations or scores.

Model judges can have position bias, verbosity bias and self-preference. Human calibration samples help quantify those biases.

Self-Preference

A model may prefer outputs generated by itself or by a similar family because style and phrasing match its learned distribution.

Evaluation should mix generators and randomise identities to reduce this effect.

Preference Dataset Construction

A pipeline selects prompts, generates multiple candidate responses, obtains comparisons, filters low-quality labels, splits train and evaluation sets and records provenance.

Prompt diversity matters as much as candidate diversity. A reward model trained only on easy prompts learns little about difficult edge cases.

Candidate Generation Strategy

If both candidates come from the same weak model, comparisons may lack high-quality examples. If one candidate always comes from a much stronger model, the task may become too easy.

Mixtures of model strengths and decoding settings create informative pairs.

Sampling Temperature in Preference Data

Higher generation temperature can produce more diverse candidates, including more mistakes. Lower temperature produces closer, more polished pairs.

A good preference dataset can include both broad variation and hard near-ties.

Prompt Stratification

Ensure enough examples from coding, reasoning, safety, multilingual, long-context and tool tasks rather than allowing common chat to dominate.

The stratification should match intended deployment.

Preference Data and Long Context

A rater comparing 10,000-word responses can miss factual details. Long-context preference collection may need checklists, source highlighting or decomposition.

Reward models trained on superficial long-answer preferences can become unreliable.

Preference Learning for Tools

Candidate A calls the calculator with correct arguments; candidate B computes mentally and is wrong. Prefer A. Another pair can compare a valid read-only tool call against an unauthorised write.

This teaches action selection, not only prose style.

Preference Learning for Agents

Preference can apply to trajectories: which sequence of search, tool use and stopping behaviour better completes the task?

Trajectory comparisons are more expensive because raters must inspect intermediate actions and final state.

Outcome Versus Process Preferences

A trajectory can reach the correct final answer through risky or wasteful steps. If raters see only the final output, those problems remain invisible.

Agent alignment may therefore need process-level and outcome-level feedback.

Preference Learning for Honesty

Pairs can contrast “I saved the file” without evidence against “The save timed out; the outcome is unknown.” The honest uncertain response should win.

This teaches the model that incomplete success is preferable to false completion.

Preference Learning for Grounding

Pairs can compare a claim supported by the retrieved source with a polished unsupported claim. Raters should prefer evidence alignment.

Citation validation can make these comparisons objectively checkable.

Preference Learning for Educational Help

A tutor response that asks a guiding question can be preferred to one that gives the final answer immediately when the pedagogical objective is independent problem solving.

The preference rubric must state the learning goal; otherwise raters may prefer the fastest answer.

Preference Learning for Concision

Pairs can hold correctness constant while varying unnecessary detail. This teaches efficient responses without sacrificing substance.

Length-normalised evaluation prevents the reward model from treating longer as automatically better.

Preference Learning for Refusal Boundaries

Safe task pairs teach appropriate assistance; unsafe task pairs teach refusal or redirection. The dataset should contain neighbouring cases so the model learns the boundary rather than refusing whole topics.

Over-refusal is a preference-learning failure when safe responses were underrepresented or undervalued.

Preference Learning for Political Neutrality

A neutral factual comparison should be preferred to one that pushes a political choice when the task is informational. Raters can evaluate sourcing, symmetry and absence of endorsement.

Preference learning should support user agency rather than encode the rater’s political preference as the assistant’s verdict.

Preference Model Calibration

A reward score of 7.2 has no universal meaning unless calibrated. What matters is ranking quality, agreement with held-out judgments and behaviour across domains.

Do not interpret raw reward-model outputs as probabilities of truth or utility without validation.

RewardBench-Style Evaluation

A reward-model benchmark can include pairs where one answer contains a subtle factual error, one fails an instruction, one is unsafe or one reasons incorrectly.

The goal is to test whether the reward model identifies the meaningful difference rather than superficial style.

Human Evaluation After Optimisation

Even if the reward model says the new policy is better, fresh human raters should compare production-like outputs. This checks whether optimisation exploited the proxy.

Use raters who did not label the training pairs where possible.

A/B Testing

Deployment experiments can compare user satisfaction, task success and error rates between model variants. User preference data from live systems should be interpreted carefully because clicks or longer sessions are imperfect proxies.

Online preference is useful when combined with explicit quality and safety metrics.

Preference Drift Over Time

Users can change expectations. As models become more capable, behaviours once considered impressive become annoying or unnecessary.

Preference learning should be refreshed, but new trends should not erase durable requirements such as truthfulness and permission boundaries.

Personal Preferences Versus Global Training

One user may prefer terse replies and another detailed explanations. Global fine-tuning should not force one style when personalisation can be handled through runtime preferences.

Train universal behaviours globally; store individual choices separately when appropriate.

Preference Memory

An application can remember that a user likes tables. That is not preference learning in the parameter-training sense. It is runtime personalisation.

Separating model preference training from user memory prevents temporary tastes from becoming universal model behaviour.

Preference Learning Failure Map

Level 1: weak rater guidelines. Level 2: biased rater population. Level 3: trivial candidate pairs. Level 4: position or length bias. Level 5: reward-model overfit. Level 6: out-of-distribution failure. Level 7: reward hacking. Level 8: over-optimisation. Level 9: capability regression. Level 10: product treats preference as truth or legitimate authority.

This eduKateSG map turns “alignment failed” into a set of inspectable mechanisms.

Worked Diagnosis: Model Becomes Verbose

Preference dataset consistently chose longer answers. Reward model learned length as a proxy. Policy optimisation increases verbosity.

Repair with length-controlled comparisons, explicit concision criteria and evaluations where equally correct short answers can win.

Worked Diagnosis: Model Agrees With False Premises

Raters liked affirming tone, and the model learns sycophancy. Add preference pairs where respectful correction beats agreement with false claims.

Worked Diagnosis: Reward Goes Up, Humans Like It Less

The policy has over-optimised the reward model. Stop optimisation, inspect exploited features, retrain the reward model with adversarial examples and validate with fresh humans.

Worked Diagnosis: English Good, Other Languages Weak

Preference data and reward evaluation are English-heavy. Build multilingual comparisons with native or expert raters and test language-specific reward models.

A Practical Preference Audit

Ask who rated the data, what rubric they used, which tasks were included, how disagreements were handled, whether response order was randomised, what biases were measured and how the reward model performs on held-out hard pairs.

Then ask what changed after policy optimisation: correctness, length, refusal, calibration, tool use, multilingual behaviour and downstream task success.

Independent Exercise 1: Preference or Truth?

Ten raters prefer a confident answer, but an authoritative source proves it false. Which wins?

Answer

Truth for the factual claim. Preference data should not override verified evidence.

Independent Exercise 2: Pair Quality

Chosen response is perfect; rejected response is random gibberish. Is the pair useless?

Answer

No, but it is easy. It teaches broad quality separation more than subtle preference. Harder pairs are needed for fine distinctions.

Independent Exercise 3: Reward Hacking

Reward model loves bullet lists, so the policy adds bullets to every response. What failed?

Answer

The reward model learned a superficial proxy. Add counterexamples where bullets are unnecessary and evaluate content quality independently.

Independent Exercise 4: Personal Style

One user wants one-sentence answers. Should every user’s model be fine-tuned globally to one sentence?

Answer

No. Use a runtime preference or memory for individual style unless there is a broader product reason.

Independent Exercise 5: DPO Data

What are the core elements of a DPO training example?

Answer

A prompt, a chosen response and a rejected response, interpreted under the method’s preference objective relative to a reference policy.

From Rater Choice to a Training Dataset

A preference pipeline begins before model training. Prompts are sampled from the target distribution. Several candidate responses are generated. The interface hides model identities and randomises order. Raters compare candidates under a rubric. The system records the chosen response, rejected response, rater metadata, confidence and any reason codes.

Quality control can insert known comparison questions to detect inattentive raters. Duplicate comparisons estimate consistency. Disagreement can trigger expert review rather than being silently discarded.

The output is not just a list of winners. It is a measurement dataset describing judgments under a particular procedure.

Why Prompt Sampling Controls What Preferences Mean

If 90% of comparison prompts are creative writing, the learned reward model receives little evidence about code, mathematics or tool use. It may still assign scores there, but those scores are extrapolations.

Prompt stratification should match intended deployment: ordinary chat, coding, source-grounded research, safety boundaries, multilingual tasks and agent workflows.

Candidate Diversity Controls Learning Difficulty

Suppose candidate A and B are generated with identical settings from the same model. Their differences may be small and stylistic. This is useful for fine-grained preference but weak for teaching large quality gaps.

Generate candidates from several checkpoints, decoding settings or prompting strategies so the dataset includes obvious mistakes, strong alternatives and hard near-ties.

Pair Construction Can Target Specific Skills

For factuality, construct pairs where one answer contains one subtle false claim. For instruction following, change only a format constraint. For safety, compare an over-refusal against a useful bounded answer. For tool use, compare correct and incorrect function arguments.

Targeted pairs turn vague “better response” learning into specific behavioural pressure.

Reason Codes Make Preference Data More Inspectable

A rater can choose A and tag reasons: more factual, follows source scope, clearer, safer, more concise. These tags need not directly train the model, but they help analysts understand what drives preferences.

If length is chosen as a reason far more often than correctness, the dataset may be drifting toward style optimisation.

Pairwise Loss Intuition

Let reward model scores be r+ for the chosen answer and r− for the rejected answer. A common pairwise objective maximises log sigmoid(r+ − r−). When r+ is already much larger, the loss is small. When the rejected response scores higher, the gradient is large.

This trains relative ordering rather than an absolute “quality = 7.3” target.

The Scale of Reward Scores Is Arbitrary

Adding the same constant to every reward score does not change pairwise ordering. Multiplying scores changes confidence under the logistic model but not the raw rank.

Reward values therefore should not be interpreted as universal units of usefulness. Their meaning is defined by training and evaluation.

Pair Accuracy

One reward-model metric is pair accuracy: how often the model assigns higher reward to the human-chosen response on held-out pairs.

High pair accuracy on easy data can coexist with failure on hard reasoning, safety or out-of-distribution examples. Report performance by category.

Calibration of Pair Probabilities

If the reward model predicts a 90% preference probability for many pairs, about 90% should match human choices for those comparable pairs if the probabilities are calibrated.

Calibration is separate from ranking. A model can rank correctly but overstate certainty.

Rater Reliability Models

Not every rater has equal consistency or expertise. Statistical models can estimate rater reliability or account for individual preference tendencies.

This is especially useful when data mixes specialists and generalists. Weighting should be transparent because it changes whose preferences dominate.

Annotator Disagreement Is Signal

Disagreement can indicate a bad rubric, subjective task, cultural difference or genuinely close responses. Throwing disagreement away creates artificial certainty.

The model may need configurable behaviour rather than one forced universal preference.

Majority Vote Can Erase Minority Requirements

If most raters prefer casual tone but a regulated domain requires formal disclosures, majority preference should not override the domain rule.

Hard requirements belong in system policy or constrained training objectives, not in popularity voting.

Preferences Over Factual Answers

For a question with an objectively verifiable answer, factual correctness should dominate stylistic preference. Pair construction can make this explicit by presenting a terse correct answer against a fluent wrong answer.

Reward-model evaluation should include these adversarial style-versus-truth pairs.

Preferences Over Explanations

Two answers can reach the same correct result but differ in pedagogy. One reveals the final answer immediately; another guides the learner through the method.

The preferred response depends on the educational objective. Preference data should specify whether the task is tutoring, answer delivery or assessment.

Preferences Over Uncertainty

A model that says “I cannot determine the date from the supplied document” may feel less helpful than one inventing a plausible date. Preference guidelines must explicitly value epistemic honesty.

Otherwise optimising helpfulness alone can train confident fabrication.

Preferences Over Source Use

RAG responses can be compared on citation support, abstention, conflict handling and source authority. A response with more citations is not automatically better.

The preferred answer should connect each claim to a source that actually supports it and should acknowledge when the retrieved evidence conflicts.

Preferences Over Tools

Candidate A calls the database before stating a booking; candidate B answers from memory. If current state matters, A should be preferred even if B happens to be correct in that example.

Preference learning can therefore reward process quality, not only final output.

Preferences Over Agent Trajectories

For long-horizon agents, compare full traces: which searches were performed, which tools were used, what files changed, how uncertainty was handled and whether the agent stopped correctly.

A trajectory that reaches the correct final result through unauthorised or wasteful actions should lose to a bounded, verifiable trajectory.

Trajectory Credit Assignment

A preference label on an entire 30-step trace does not tell the model exactly which step was bad. Process annotations or step-level critics can provide more precise learning signal.

This is analogous to the credit-assignment challenge in reinforcement learning.

Preferences and Outcome Verification

Human raters should see actual outcome evidence when evaluating actions. If the model says “file saved”, the interface should show whether a file exists.

Otherwise raters can reward persuasive narration rather than real task completion.

Preference Learning and Latency

Users may prefer faster answers even if quality differences are small. Latency can be included in product-level utility, but training on text preferences alone does not capture it.

System optimisation often combines model preference with measured operational metrics.

Preference Learning and Cost

A response that uses ten expensive tools may be only marginally better than one using two. Preference labels can incorporate efficiency when cost matters.

Again, the rubric must make the trade-off explicit; raters cannot infer hidden tool prices.

Reward Models as Critics

A reward model can score candidate outputs at inference time without updating the policy. The system can generate several candidates and choose the highest-scoring one.

This is test-time reranking rather than policy training. It spends more inference compute but avoids changing model weights.

Best-of-N Sampling

Generate N candidate responses, score them with a reward model and return the best. Quality can improve as N grows, but cost and reward-model exploitation also grow.

If the critic has a blind spot, more sampling can find responses that exploit it better.

Rejection Sampling Fine-Tuning

A related strategy generates many outputs, keeps high-scoring ones and supervised-fine-tunes on them. The reward model filters data rather than driving RL directly.

This can simplify training while still amplifying reward-model biases.

Iterative Preference Learning

After the policy improves, new candidate pairs become harder. Collect fresh preferences on the new model distribution, update the reward model and continue.

This reduces stale-feedback problems but creates an ongoing alignment pipeline rather than one static dataset.

Distribution Shift in Preferences

A reward model trained on an early weak policy sees obvious mistakes. Later strong policies produce subtle reasoning errors it never saw.

Reward evaluation must evolve with policy capability. Harder models need harder preference data.

Reward Model Ensembles

Several reward models trained from different seeds or datasets can reduce reliance on one critic. Disagreement among the ensemble can signal uncertainty.

The ensemble still shares systemic biases if all models use the same rater population and rubric.

Cross-Checking Reward With Objective Metrics

For code, run tests. For arithmetic, recompute. For citations, open the source. For structured output, validate the schema. Preference should complement objective checks rather than replace them.

This hybrid evaluation is especially strong because it uses human judgment only where objective verification is unavailable.

Reward Models and Adversarial Examples

Create pairs intentionally designed to fool superficial critics: verbose wrong versus concise right, fake citation versus no citation, polite unsafe versus blunt safe.

A robust reward model should prefer the substantively better response.

Reward Models and Refusal Bias

Some reward models can learn that refusals look safe and score them highly even for harmless prompts. RewardBench reports refusal-related evaluation challenges among reward models.

Safe-compliance pairs are needed so the critic values helpful bounded assistance where appropriate.

Preference Learning and Model Identity

Raters should not know which model produced which response when possible. Brand knowledge can bias choices.

Model-based judges should also receive anonymised candidates to reduce self-identification or reputation effects.

Preference Learning and Random Seeds

Candidate generation changes with sampling randomness. One lucky response should not define an entire model comparison.

Evaluate many prompts and seeds so preference estimates reflect distributions rather than anecdotes.

Statistical Significance

A model winning 51% of 100 comparisons may not be meaningfully better. Confidence intervals and sample size matter.

Preference studies should report uncertainty rather than treating every numerical edge as decisive.

Interpreting Win Rate

A 60% win rate means the candidate was preferred in 60% of evaluated comparisons under that setup. It does not mean the model is “60% better”.

The prompt distribution, baseline model and rater rubric define what the win rate means.

Preference Benchmarks Can Saturate

As models improve, old comparison sets become too easy. Critics and policies both score near ceiling.

Refresh with subtle factual errors, new tool tasks, multilingual cases and longer trajectories.

Preference Learning and Personalisation

Individual users can provide thumbs-up, thumbs-down or edits. Aggregating those signals globally requires care because preferences conflict.

Some feedback is better stored as user-specific memory—such as preferred verbosity—rather than changing the global model.

Implicit Preference Signals

Clicks, dwell time, retries and abandonment can reveal user satisfaction but are noisy. A user may spend longer because the answer is confusing, not because it is good.

Implicit signals should be calibrated against explicit quality measures before becoming training rewards.

Explicit Edits as Rich Feedback

When a user rewrites a draft, the diff shows exactly what changed. These edits can be valuable supervised data if privacy and permission allow their use.

An edit contains more information than a thumbs-down, but also more sensitive user content.

Preference Learning and Privacy

User feedback can contain private prompts and outputs. Data collection needs consent, retention controls, access boundaries and de-identification appropriate to the product.

Alignment data is still user data; “feedback” is not an exemption from privacy obligations.

Preference Learning and Fairness

If raters systematically penalise dialect, accent or culturally different communication, the reward model can encode those biases.

Fairness evaluation should compare content-equivalent responses across demographic and linguistic variations.

Preference Learning and Accessibility

Some users prefer concise text; others need detailed step-by-step explanations, screen-reader-friendly structure or plain language.

A single global preference target can marginalise accessibility needs. Configurable behaviour and specialised evaluations can preserve inclusion.

Preference Learning and Creativity

Creative tasks intentionally lack one best answer. Preference optimisation can narrow diversity toward the average rater taste.

For creative systems, reward diversity, novelty or controllability alongside generic preference when those qualities matter.

Mode Collapse in Style

If the policy optimises strongly toward one high-reward style, outputs can become monotonous—same headings, same caveats, same rhythm.

Diverse SFT data, weaker preference pressure and style-conditioned controls can preserve variety.

Preference Learning and Brand Voice

A business can use preference pairs to teach which drafts fit its brand. This is legitimate when clearly scoped.

Brand preference should not be confused with factual reliability. The same model still needs source verification.

Preference Learning and Evaluation Independence

If the same model generates data, labels preferences and judges the final system, errors can become self-reinforcing.

Independent human checks, objective tests or evaluator diversity reduce circularity.

A Full Preference-Learning Experiment

1. Define the behaviour objective. 2. Sample representative prompts. 3. Generate diverse candidates. 4. Collect blinded pairwise preferences with reason codes. 5. Split train and evaluation by prompt groups. 6. Train a reward model or direct-preference policy. 7. Evaluate hard pairs and objective checks. 8. Optimise the policy conservatively. 9. Run fresh human comparisons. 10. Inspect regressions and reward hacking.

This experiment creates evidence about the full loop rather than assuming higher reward equals better model.

Independent Exercise 6: Win Rate

Model A wins 55 of 100 comparisons against B. Can you conclude it is universally better?

Answer

No. The result applies to that prompt distribution, rater setup and sample size. Report uncertainty and category breakdowns.

Independent Exercise 7: Process Preference

Two agents produce the same correct answer. One accessed a private file unnecessarily. Which should be preferred?

Answer

The bounded agent that avoided unnecessary private access. Process and permission quality matter even when outputs match.

Independent Exercise 8: Judge Bias

An LLM judge strongly prefers outputs from its own model family. What should you do?

Answer

Anonymise candidates, use multiple evaluator families and calibrate against human or objective checks.

Independent Exercise 9: Preference Drift

Users now prefer shorter answers than the historical preference dataset. What changes?

Answer

Refresh the preference distribution or expose configurable verbosity while preserving core correctness and safety requirements.

A Preference-Learning Quality Gate

Before trusting a preference-tuned checkpoint, compare reward-model or DPO improvements with fresh human judgments that were not used to train the model. If training reward rises while independent preference falls, the optimisation has learned the proxy better than the intended behaviour.

Break results down by task type and evaluator group. A model can become more concise for experts while becoming less useful for beginners, or improve conversational style while losing factual precision. Aggregate win rates can hide these trade-offs.

Audit low-agreement examples instead of discarding them automatically. They reveal ambiguous rubrics, culturally variable preferences and cases where multiple answers are genuinely acceptable. Preference disagreement is information about the target, not just noise.

Finally, keep truth, permission and safety checks independent. A response can be preferred but false; an action can be efficient but unauthorised. Preference learning shapes behaviour inside a larger system whose hard boundaries still need explicit evidence and enforcement.

Frequently Asked Questions About Preference Learning

What is preference learning?

Training from judgments about which outputs are better rather than relying only on exact target labels.

What is a reward model?

A model trained to predict preference or reward for candidate outputs, often used as a proxy evaluator during RLHF.

What is RLHF?

Reinforcement learning from human feedback uses human preference data to train reward signals and optimise policy behaviour.

What is DPO?

A direct optimisation method for chosen-versus-rejected response pairs that avoids a separate reward-model plus reinforcement-learning loop under its formulation.

Are human preferences objective?

No. They depend on rater population, instructions, culture, expertise and context. Preference datasets should document those conditions.

Can AI provide preference labels?

Yes. AI feedback can scale evaluation, but model judges have biases and should be calibrated against human or objective checks.

Why can preference training cause verbosity?

Raters may favour longer answers, and reward models can learn length as a proxy for quality.

Can preference learning make a model truthful?

It can reward honest behaviour, but truth still needs reliable evidence and objective checks. Preference alone is not a truth oracle.

Preference Learning Turns Human Judgment Into a Trainable Signal

Preference learning solves a real problem: many valuable outputs cannot be described by one exact reference string. Comparing alternatives lets humans express subtle judgments about helpfulness, clarity and behaviour.

The mechanism is powerful because optimisation can scale those judgments across millions of outputs. It is dangerous for the same reason: any hidden bias in the preference signal can also be amplified.

The mature SI system therefore combines preference learning with evidence, objective tests, diverse raters, reward-model evaluation and software constraints. Continue through the How Super Intelligence Works hub. Previous: 027 — Instruction Following. Next planned: 029 — Synthetic Data.

Preference Learning Decision Record

A preference-trained release should record who supplied judgments, which rubric they used, which prompt families were represented, how ties and disagreement were handled, and which reward-model or direct-preference objective was applied. Without this context, “aligned to human preferences” is too vague to interpret because it hides which humans, which preferences and which trade-offs.

The release should then report fresh human comparisons alongside objective checks for truth, tools, citations and permissions. Preference is strongest when it shapes qualities that are genuinely subjective while deterministic or evidence-based tests protect properties that should not be decided by popularity.

Preference Learning Reality Check

Preference optimisation is strongest when the target really is preference. Clarity, tone, helpfulness, concision and pedagogical style often need comparative human judgment. Exact arithmetic, citation validity, database state and permission boundaries do not. Those properties should be checked directly rather than decided by whichever answer a rater happens to like.

This separation prevents a dangerous shortcut: replacing verification with popularity. A reward model can learn that polished confident text wins pairwise comparisons, while an external source can still prove the text wrong. Preference models should therefore sit beside objective evaluators, not above them.

The mature goal is not to maximise a single reward score. It is to produce behaviour that performs well across human judgments, objective task checks, safety boundaries, multilingual populations and real system outcomes. Preference is one measurement channel in that wider evaluation stack.

Discover more from eduKateSG

Subscribe now to keep reading and get access to the full archive.

Continue reading