Safety

Noun · AI & Machine Learning

Definitions

  1. The discipline of ensuring that AI systems behave as intended, avoid harmful outputs, and remain aligned with human values, encompassing alignment research, red-teaming, and guardrail design.

    In plain English: Making sure an AI does not do dangerous or harmful things, from giving bad advice to behaving in ways its creators did not intend.

    Example: "The safety team flagged the model for generating misleading medical advice during eval."

Related Terms