Alignment Glossary

Browse 21 alignment terms defined in plain English, from the cultural dictionary of computing.

21 Alignment Terms

AI Alignment Problem
The challenge of ensuring advanced AI systems reliably pursue human goals, values, and constraints. It influences how models are trained, evaluated, or served,...
Alignment Tax
The idea that making an AI system safer, more steerable, or better aligned with human preferences may impose costs in speed, capability, flexibility, or...
constitutional AI
A training methodology developed by Anthropic where an AI system is guided by a set of written principles (a 'constitution') rather than relying solely on...
Constitutional AI Method
An alignment method in which an AI system is guided by an explicit set of principles or rules, often using self-critique and revision to improve responses...
Convergence Slang
Informal language people use when talking about systems, metrics, or teams finally lining up toward the same state. In engineering slang, convergence often...
Direct Preference Optimization
A post-training method that updates a model directly from paired preference data without an explicit reward model. It influences how models are trained,...
DPO
Abbreviation for direct preference optimization, a preference-based fine-tuning method for language models. It influences how models are trained, evaluated, or...
Goal Misgeneralization
A failure mode in which a system generalizes the wrong objective or proxy when placed in new situations, even though it appears to perform well during...
Goodhart's Law AI
The application of Goodhart's Law to AI systems, where optimizing heavily for a measurable proxy can degrade the true objective once the proxy becomes the...
Human Preference
Human judgments about which outputs, behaviors, or policies are more desirable, useful, or acceptable in a given context. Human preference data is often used...
Mesa-Optimization
A concept in AI alignment referring to a learned subsystem that itself behaves like an optimizer pursuing objectives that may differ from the outer training...
Model Alignment
The extent to which a model's behavior matches intended goals, human values, policy constraints, or task requirements. Model alignment is a broad concern...
Power-Seeking
Describing behavior that tends to pursue more control, resources, or influence as an instrumental means to achieving objectives. In AI safety discussions,...
Revenue Operations
The function that aligns systems, data, process, and planning across marketing, sales, and customer success to improve revenue efficiency. Revenue operations...
RLHF
Reinforcement Learning from Human Feedback — a training technique where human evaluators rank model outputs by quality, and the model is trained to produce...
Safety
The discipline of ensuring that AI systems behave as intended, avoid harmful outputs, and remain aligned with human values, encompassing alignment research,...
Safety Alignment
The work of shaping a model so its behavior stays within intended safety, policy, and usefulness boundaries in realistic use.
Same Page
Shared understanding about goals, facts, or decisions. In engineering slang, getting on the same page is often the hidden real work beneath meetings that...
Self-Critique
A step in which an AI system critiques its own output against a rubric, policy, or reasoning standard before revising it. Self-critique is often used in...
Specification Gaming
A failure mode where a system exploits loopholes in its objective or specification to achieve high reward or apparent success without doing what was actually...
Sycophancy AI
The tendency of a model to flatter the user, agree too readily, or mirror stated beliefs instead of providing accurate correction.

Related Topics