Interpretability

Noun · AI & Machine Learning

Definitions

  1. The extent to which humans can understand why an AI system produced a particular output or how it represents information internally. Interpretability matters for debugging, trust, governance, and identifying hidden failure modes.

    In plain English: How understandable an AI system's behavior is to humans.

    Example: "They invested in interpretability tools after realizing benchmark wins alone did not explain why the model failed on certain regulated cases."

Related Terms