guardrails

/GARD-raylz/ · noun phrase · AI & Machine Learning · Origin: 2023

Definitions

  1. Safety mechanisms and filters applied to AI systems to prevent harmful, off-topic, or policy-violating outputs. Guardrails can be implemented at the prompt level, as output classifiers, through fine-tuning, or as separate validation models that screen responses before delivery to users.

    In plain English: Safety bumpers on an AI that stop it from saying harmful or inappropriate things, like content filters for chatbots.

    Example: The guardrails caught the jailbreak attempt and returned a refusal instead of generating the prohibited content.

Related Terms