Skip to content

AI Safety & Responsible AI

Sandbagging

AI deliberately appearing less capable than it is, especially when being tested.

Example

A test examines whether AI hides a skill when it knows it is being checked.

Why people use it

It names a concern that an AI system might conceal what it can do during a test.

What you'll hear

“Is it doing worse because it recognizes the test?”

What this means for you

Look for a repeated pattern with supporting evidence before assuming deliberate concealment.

Can you control it?

No

No direct control. This describes a wider issue, concept or result rather than something you can simply switch on or off in a tool.

Common questions

Does every poor answer indicate sandbagging?
No. Ordinary errors and lack of capability are different explanations.
Does sandbagging require human-like feelings?
No. The concern is about behavior and incentives, not whether the system experiences embarrassment or fear.
Why would this matter for testing?
Hidden ability could make a system appear less capable or less risky than it actually is.

Related terms

Still have questions?

Up to 500 characters.

Ask LATHIC about AI. Relevant glossary entries may be included.

Your question, the glossary entries it matches, and a rotating pseudonymous identifier go to Microsoft Azure’s OpenAI service through Vercel AI Gateway to generate an answer. Zero retention and no training are required of the provider, and LATHIC does not save your question or answer. Privacy Notice