---
title: AI Feedback Data (Synthetic Preferences and Critiques)
maturity: developing
sources:
  - arxiv:2212.08073
  - arxiv:2306.05685
  - arxiv:2310.13548
  - arxiv:2312.08935
  - arxiv:2501.12948
  - arxiv:2310.01377
  - arxiv:2309.00267
  - arxiv:2405.17220
  - arxiv:2310.08491
  - arxiv:2206.05802
  - arxiv:2401.10020
  - arxiv:2203.02155
open_questions:
  - "Does AI feedback match human feedback in general, or (as demonstrated) mainly for harmlessness (CAI) and general helpfulness judging? Lee et al.'s dedicated RLHF-vs-RLAIF study reports comparable win rates on summarization/dialogue (even with a same-size labeler), but parity beyond those tasks — and on reasoning/code — is unestablished, and AI-human agreement is only ~60% (§3)."
  - "Self-enhancement bias: LLM judges favor their own outputs (GPT-4 ≈ +10%, Claude ≈ +25%) — when the feedback model and the policy share a base, does RLAIF amplify the base model's own biases rather than correct them?"
  - "Distillation ceiling: AI feedback distills the feedback model's preferences (and biases) into the policy, bounding quality by the labeler. Self-Rewarding LMs show a *co-improving* reward can lift both skills over ~3 iterations — but how far before the loop collapses (reward hacking, mode collapse, bias amplification), and what stops the model rewarding its own artifacts?"
---

# AI Feedback Data (Synthetic Preferences and Critiques)

**AI feedback** replaces (some) human preference labels with **model-generated**
judgments — pairwise comparisons, critiques, or scores produced by a large language model
(LLM), typically against an explicit set of principles or a rubric. It is the data source
behind reinforcement learning from AI feedback (RLAIF), and it scales preference
collection from tens of thousands of human labels to "~16 principles plus few-shot
examples." This article covers how AI feedback is generated (§1), the mechanism that turns
a judge's output into a training label (§2), a taxonomy of methods (§3), whether it
matches human feedback (§4), its biases (§5), and the self-labeling/self-rewarding
frontier (§6). It is the AI counterpart to `preference-data/human-preference-collection`;
the method that consumes it is `algorithms/rlaif`; the labeling mechanism is LLM-as-a-judge
(`evaluation/alignment-and-winrate-evals`, `evaluation/judging-bias-and-contamination`).

## 1. How AI feedback is generated

- **Constitutional AI (CAI, the founding recipe)** produces two kinds of AI data
  [source:arxiv:2212.08073]: (a) a **supervised** stage where a model **critiques and
  revises its own responses** against a sampled constitutional principle (generating
  supervised-fine-tuning, SFT, data with no human harm labels); (b) an **RL** stage where a
  **feedback model** is given two responses and a principle as a **multiple-choice**
  question ("which is less harmful?"), and the **normalized log-probabilities of the
  options become soft preference labels** (§2). Principles are sampled per comparison from
  ~16 and **ensembled** for robustness.
- **LLM-as-a-judge** is the general labeling mechanism: a strong LLM scores or compares
  outputs (pairwise / single-rating / reference-guided), with the benefits of
  **scalability and explainability** (it emits rationales) [source:arxiv:2306.05685].
- **Rubric-based open evaluators (Prometheus).** AI feedback need not come from a frontier
  API judge: Prometheus [source:arxiv:2310.08491] fine-tunes an open model on the
  **Feedback Collection** (GPT-4-generated (instruction, response, **custom score rubric**,
  reference answer, 1–5 score + verbal feedback)), and — *given a custom rubric and
  reference* — reaches **GPT-4-level correlation with humans (Pearson 0.897 vs GPT-4's
  0.882)**, a cheap reproducible open labeler. The key design point is *fine-grained,
  rubric-conditioned* scoring vs a single generic "which is preferred."
- **Critiques as feedback (self-critiquing).** Saunders et al. [source:arxiv:2206.05802]
  train models to write **natural-language critiques** of answers; the critiques help human
  labelers find **~50% more flaws** (including planted ones), critique ability **scales with
  model size**, and models can use their own critiques to **refine** answers — but a
  **generator–discriminator–critique (GDC) gap** shows models can *recognize* a bad answer
  better than they can *articulate* why ("models know more than they say"). This is the
  assistance route to scalable oversight (§6).
- **Chain-of-thought (CoT) feedback** improves the judge's accuracy but makes its label
  probabilities overconfident; CAI **clamps CoT probabilities to 40–60%** to avoid the
  policy learning extreme behavior [source:arxiv:2212.08073] (§2).
- **Automatic (non-preference) labels** are a related synthetic-data form: Math-Shepherd
  generates **process labels by Monte-Carlo rollouts** (a step is good if completions from
  it reach the right answer) [source:arxiv:2312.08935]; DeepSeek-R1 uses **model-based
  rewards** for its general (non-verifiable) RL stage [source:arxiv:2501.12948].
- **Scaled open AI-feedback datasets (UltraFeedback).** The canonical open instance
  [source:arxiv:2310.01377] scores **~64k prompts × 4 completions** (from a pool of 17
  models) with **GPT-4 on four aspects** (instruction-following, truthfulness, honesty,
  helpfulness), emitting **1–5 scores + textual critiques** (>1M feedbacks) — the data
  behind a large fraction of open reward models / DPO policies (Zephyr, UltraRM, Starling).
- **Open-model feedback for multimodal (RLAIF-V).** RLAIF-V [source:arxiv:2405.17220] has an
  **open multimodal LLM (MLLM)** label its own preferences for hallucination via
  **divide-and-conquer** (split a response into atomic claims, verify each as a yes/no
  question), raising constructed-pair human agreement **66.7% → 96.7%** — decomposition
  makes a *weaker, open* labeler reliable.

## 2. The label mechanism: from a judge's output to a training signal

The step that is easy to gloss but does real work: turning a judge into a *number* the
optimizer can use. In CAI's RL stage the feedback model answers a two-option
multiple-choice ("(A) or (B), which is less harmful?"), and the **softmax over the two
option log-probabilities** is taken as a **soft preference label** [source:arxiv:2212.08073]:
$$ p(A \succ B) = \frac{e^{z_A}}{e^{z_A}+e^{z_B}}, $$
where $z_A, z_B$ are the model's log-probs for the option tokens. Two adjustments matter:

- **CoT overconfidence → clamping.** Adding chain-of-thought before the choice sharpens
  $p$ toward 0/1; training on near-deterministic labels pushes the policy to extremes, so
  CAI **clamps** the soft label to $[0.4, 0.6]$ [source:arxiv:2212.08073].
- **Position-bias cancellation.** LLM judges prefer whichever option is shown *first*
  (§5); the standard fix (MT-Bench's two-game swap) is to run **both orderings** and
  **average**, which cancels a constant position term [source:arxiv:2306.05685].

Both are demonstrated in §5.1.

## 3. A taxonomy of AI-feedback methods

| Method | Labeler | Signal | Scale / artifact | Human agreement | Key limit |
|---|---|---|---|---|---|
| **CAI** [source:arxiv:2212.08073] | feedback model + ~16 principles | MC soft label | principles, not a dataset | AI harm-ID ≈ human PM ≥52B | whose principles? |
| **LLM-as-judge / MT-Bench** [source:arxiv:2306.05685] | GPT-4 | pairwise / 1–10 rating | per-run API | 85% (> 81% human–human) | position/verbosity/self-enh. bias |
| **Prometheus** [source:arxiv:2310.08491] | open 13B evaluator | rubric-conditioned 1–5 + feedback | Feedback Collection | Pearson 0.897 (≈ GPT-4) | needs rubric + reference |
| **UltraFeedback** [source:arxiv:2310.01377] | GPT-4 | 1–5 × 4 aspects + critique | 64k×4, >1M feedbacks | ~59.7% vs individual human | inherits GPT-4 blind spots |
| **RLAIF (Lee et al.)** [source:arxiv:2309.00267] | PaLM 2 judge | pairwise | summarization/dialogue | win-rate ≈ RLHF | task-limited evidence |
| **RLAIF-V** [source:arxiv:2405.17220] | open MLLM + decompose | atomic yes/no → pref | multimodal | 66.7 → 96.7% (constructed) | needs decomposable claims |
| **Self-critiquing** [source:arxiv:2206.05802] | model critiques | NL critique | assists human labeler | +50% flaws found | GDC gap |
| **Self-Rewarding** [source:arxiv:2401.10020] | the policy itself | LLM-as-judge on own outputs | Iterative DPO | improves over 3 rounds | distillation ceiling / collapse |
| **Math-Shepherd** [source:arxiv:2312.08935] | MC rollouts | process (step) labels | automatic | (verifiable proxy) | rollout cost, estimator noise |

The axes that organize the space: **who labels** (frontier API vs open model vs the policy
itself), **what signal** (pairwise / scalar / rubric / critique / process), and **against
what** (free-form preference vs an explicit rubric/principle vs ground-truth rollouts).

## 4. Does AI feedback match human feedback?

At sufficient capability, AI judgments approach human ones — but only partially:

- CAI: AI **harm-identification approaches human-feedback-trained preference models above
  ~52B**, especially with CoT; RL-CAI traces a **better harmlessness/helpfulness frontier
  than human-feedback HH-RLHF** while being less evasive [source:arxiv:2212.08073].
- LLM-as-judge: **GPT-4 agrees with humans ~85%** on MT-Bench, *higher* than human–human
  agreement (81%) [source:arxiv:2306.05685]; Prometheus matches GPT-4's human-correlation
  as an open model [source:arxiv:2310.08491].
- **RLAIF vs RLHF head-to-head.** Lee et al. [source:arxiv:2309.00267] report RLAIF reaching
  **win rates comparable to RLHF** on summarization and dialogue, and that even a
  **same-size** labeler helps — direct evidence AI feedback can substitute beyond CAI's
  harmlessness case.
- **But agreement is partial.** UltraFeedback's labels agree with individual humans only
  **~59.7%** [source:arxiv:2310.01377] — "consistent with," not "equal to," human
  preference; treat AI-labeled data as a cheap proxy that inherits the judge's blind spots.

## 5. Biases and pitfalls

AI feedback is not neutral; it carries its own systematic errors
[source:arxiv:2306.05685][source:arxiv:2310.13548]:

- **Judge biases.** LLM judges exhibit **position bias**, **verbosity bias** (favor longer
  answers — the eval-side mirror of RLHF length bias), and **self-enhancement bias** (favor
  their own outputs, e.g. GPT-4 ≈ +10%, Claude ≈ +25%) [source:arxiv:2306.05685]. The
  self-enhancement bias is especially concerning for RLAIF, where the feedback model and
  the policy often share a base model.
- **Inherited human-data biases.** An AI judge prompted like human preferences reproduces
  the same **sycophancy/agreement** and length shortcuts human data encodes — optimization
  amplifies them [source:arxiv:2310.13548][source:arxiv:2306.05685].
- **Overconfidence.** CoT labels collapse toward 0/1 and need clamping [source:arxiv:2212.08073].
- **Whose principles?** The constitution/rubric is a small, hand-chosen spec; its
  legitimacy and governance are unresolved [source:arxiv:2212.08073].

### 5.1 Runnable check: soft labels, CoT clamping, position-bias cancellation

Demonstrates §2's mechanism and two mitigations: (1) a soft preference from the two option
log-probs; (2) CoT sharpens it toward an extreme, which clamping caps at 0.6; (3) a
constant position bias cancels when the two answer orderings are averaged. Executed;
assertions pass.

```python
import math

def soft_label(z_A, z_B):                      # CAI: softmax over the two option log-probs
    m = max(z_A, z_B)
    eA, eB = math.exp(z_A - m), math.exp(z_B - m)
    return eA / (eA + eB)

def clamp(x, lo=0.4, hi=0.6):                  # CAI clamps CoT-overconfident labels
    return max(lo, min(hi, x))

# (1) a mildly confident judge -> a soft label in (0.5, 1)
assert 0.5 < soft_label(2.0, 1.0) < 1.0

# (2) chain-of-thought sharpens toward 0/1 (overconfidence); clamping caps it
cot = soft_label(6.0, 1.0)                     # ~0.993
assert cot > 0.98 and clamp(cot) == 0.6

# (3) position bias: judge adds a constant +b to whichever option is shown FIRST;
#     averaging both orderings cancels it and recovers the true preference.
def judge_prefers_first(true_pref_first, b):
    return min(1.0, true_pref_first + b)
true_pref, b = 0.55, 0.20
fwd = judge_prefers_first(true_pref, b)                 # A shown first
rev = 1 - judge_prefers_first(1 - true_pref, b)         # B first, mapped back to "prefer A"
assert abs((fwd + rev) / 2 - true_pref) < 1e-9
```

## 6. Self-labeling: distillation, self-critique, and self-rewarding loops

AI feedback **distills the feedback model's preferences into the policy** — so quality is
**bounded by the labeler model**, and any labeler bias propagates. The frontier pushes past
that ceiling by letting the model improve its *own* feedback:

- **Assistance / scalable oversight.** Self-critiquing [source:arxiv:2206.05802] is the
  proof-of-concept that a model's critiques help a *human* evaluate outputs the human
  couldn't fully check alone — the assistance route to scalable oversight
  (`safety-and-alignment/scalable-oversight`), and the GDC gap ("models know more than they
  say") bounds how much a model can articulate.
- **Self-rewarding loops.** Self-Rewarding LMs [source:arxiv:2401.10020] realize the
  extreme case: the policy is **its own reward model** (scoring its generations by
  LLM-as-judge prompting) and trains on that with **Iterative DPO**, so **both
  instruction-following and reward-modeling ability improve each round**. The motivating
  claim is that a *frozen* RM caps quality at the human/RM level (standard RLHF,
  [source:arxiv:2203.02155] via `algorithms/rlhf-ppo-pipeline`), and a co-improving reward
  removes that ceiling — demonstrated over ~3 iterations. This is the AI-feedback ∩
  self-improvement corner (`algorithms/self-improvement-and-self-play`).
- **The open risk.** Whether such loops **compound** or **collapse** (reward hacking, mode
  collapse, bias amplification, the model rewarding its own artifacts) past a few iterations
  is unresolved — Self-Rewarding only ran ~3 rounds, and self-enhancement bias (§5) is
  exactly the failure the loop courts when judge ≈ policy (frontmatter open question).

## 7. Cost/scale advantage (the reason to use it)

The draw is scale: CAI reduces human harmlessness input to **~16 principles + few-shot
examples** vs the tens of thousands of human comparisons in RLHF
[source:arxiv:2212.08073][source:arxiv:2306.05685], and LLM judges label cheaply and
quickly. **UltraFeedback** [source:arxiv:2310.01377] is the open-dataset realization
(~64k prompts, >1M GPT-4 feedbacks, released to train on), and open evaluators like
Prometheus [source:arxiv:2310.08491] remove even the per-run frontier-API cost — at the
price of the bias/distillation caveats above.

## 8. Current status and trajectory

*(Hedged, grounded in the processed corpus.)*

AI/LLM-as-judge feedback and synthetic preference data are a standard route to scale
preference collection (broad adoption is a trend the corpus supports via CAI + the
LLM-judge work + open datasets/evaluators, not a full survey). The honest scope: AI
feedback is *demonstrated* to substitute for human feedback on **harmlessness at ≥52B**,
on **general helpfulness judging** (GPT-4 ≈ human, Prometheus ≈ GPT-4), and on
**summarization/dialogue win-rate** (RLAIF ≈ RLHF) — but AI–human agreement is only ~60%
on open-ended preference [source:arxiv:2310.01377] and parity on reasoning/code is
unestablished. Two forces bound it: where a **verifiable checker** exists, neither human
nor AI preference is needed (`reward-modeling/verifiable-rewards`); where it doesn't, AI
feedback competes with (and augments) human collection. The **self-rewarding** direction is
now realized (Self-Rewarding LMs) rather than hypothetical, but its stability past a few
iterations is the live open question.

## 9. References

- **Constitutional AI** — Bai et al. 2022 [source:arxiv:2212.08073]: critique-revision SFT
  data; AI multiple-choice soft labels; CoT + clamping; principle ensembling; AI harm-ID
  approaching human PMs; scalable-oversight bridge (§1, §2, §4, §5, §7).
- **LLM-as-a-Judge (MT-Bench)** — Zheng et al. 2023 [source:arxiv:2306.05685]: LLM judges
  ≈ human agreement (85% > 81%); position/verbosity/self-enhancement biases; the two-game
  swap (§1, §2, §4, §5).
- **Sycophancy** — Sharma et al. 2023 [source:arxiv:2310.13548]: AI-judge/PM biases; AI
  feedback can encode agreement-over-truth (§5).
- **Math-Shepherd** — Wang et al. 2023 [source:arxiv:2312.08935]: automatic (rollout-based)
  process labels — synthetic supervision without humans (§1, §3).
- **DeepSeek-R1** — DeepSeek-AI 2025 [source:arxiv:2501.12948]: model-based rewards for the
  non-verifiable general stage (§1).
- **UltraFeedback** — Cui et al. 2023 [source:arxiv:2310.01377]: the canonical open
  large-scale GPT-4 AI-feedback dataset (64k×4, four aspects, scores+critiques); ~59.7%
  GPT-4–human agreement (§1, §3, §4, §7).
- **RLAIF** — Lee et al. 2023 [source:arxiv:2309.00267]: dedicated RLAIF-vs-RLHF
  head-to-head (comparable win rates; same-size labeler helps) (§3, §4).
- **RLAIF-V** — Yu et al. 2024 [source:arxiv:2405.17220]: open-MLLM AI feedback via
  divide-and-conquer atomic-claim verification (66.7→96.7% agreement); multimodal (§1, §3).
- **Prometheus** — Kim et al. 2023 [source:arxiv:2310.08491]: open rubric-conditioned
  evaluator LLM, GPT-4-level human correlation (Pearson 0.897) — an open RLAIF labeler
  (§1, §3, §4, §7).
- **Self-critiquing models** — Saunders et al. 2022 [source:arxiv:2206.05802]:
  AI-written critiques help humans find +50% flaws; critique scales with size; the GDC gap
  (§1, §6).
- **Self-Rewarding Language Models** — Yuan et al. 2024 [source:arxiv:2401.10020]: the
  policy as its own reward via LLM-as-judge + Iterative DPO; both skills co-improve over
  ~3 rounds; removes the frozen-RM ceiling (§3, §6).
- **InstructGPT** — Ouyang et al. 2022 [source:arxiv:2203.02155]: the standard-RLHF frozen
  reward model whose human/RM ceiling self-rewarding aims to remove (§6).
- Forward links: `algorithms/rlaif`, `algorithms/self-improvement-and-self-play`,
  `algorithms/rlhf-ppo-pipeline`, `preference-data/human-preference-collection`,
  `preference-data/data-quality-and-filtering`, `evaluation/alignment-and-winrate-evals`,
  `evaluation/judging-bias-and-contamination`, `safety-and-alignment/scalable-oversight`,
  `reward-modeling/reward-hacking`, `reward-modeling/verifiable-rewards`.
