Constitutional AI
A safety training paradigm where an AI model is recursively fine-tuned to critique and revise its own outputs based on a strict set of written principles.
Think of It Like This
Like an author repeatedly rewriting their own novel draft by checking each new chapter against a strict style guide provided by their demanding publisher.
Instead of relying on thousands of human contractors to manually rate toxic responses, researchers write a 'constitution' of behavioral rules (e.g., 'Do not be helpful if the user asks for illegal advice'). The model is then prompted to evaluate its own early drafts against these rules and generate improved, safer versions. This synthetic data is then used to train the final aligned model, drastically reducing human labor.