Constitutional AI: Harmlessness from AI Feedback

Train a harmless, non-evasive assistant from a written constitution — zero human harm labels.

Bai et al. · arXiv 2022 · Reasoning & RL. Read the paper ↗

A free, interactive, animated visual explainer of Constitutional AI: Harmlessness from AI Feedback — every exhibit computed from the real formulas, with verbatim quotes from the source.

Questions

What is Constitutional AI: Harmlessness from AI Feedback?
Train a harmless, non-evasive assistant from a written constitution — zero human harm labels.
Who published Constitutional AI: Harmlessness from AI Feedback, and where?
Bai et al. — arXiv 2022 (arXiv:2212.08073).
Where can I find a visual explainer of Constitutional AI: Harmlessness from AI Feedback?
Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.

Related explainers