Skip to content

AI Safety & Responsible AI

Interpretability

How understandable a model's internal reasoning or decision process is to humans.

Example

Researchers examine internal model behavior to understand why it produced a result.

Why people use it

Teams use “Interpretability” when they need to identify harms and choose proportionate safeguards.

What you'll hear

“We need to account for Interpretability before we ship this.”

Related terms