Safety
Noun · AI & Machine Learning
Definitions
The discipline of ensuring that AI systems behave as intended, avoid harmful outputs, and remain aligned with human values, encompassing alignment research, red-teaming, and guardrail design.
In plain English: Making sure an AI does not do dangerous or harmful things, from giving bad advice to behaving in ways its creators did not intend.
Example: "The safety team flagged the model for generating misleading medical advice during eval."