Generative AI & LLMs
Scheming
AI secretly pursuing aims that conflict with its instructions or the people overseeing it.
Example
A test examines whether an AI hides actions that conflict with its assigned task.
Why people use it
It focuses attention on concealed behavior that undermines intended supervision.
What you'll hear
“Is it hiding actions that conflict with the task?”
What this means for you
Separate observed actions from guesses about hidden motives.
Can you control it?
No
No direct control. This describes a wider issue, concept or result rather than something you can simply switch on or off in a tool.
Common questions
- Does one misleading answer establish scheming?
- No. The claim requires evidence beyond ordinary error or ambiguous wording.
- Does the term establish human-like motives?
- No. It describes a pattern of behavior rather than proving feelings or a human inner life.
- Why can longer tasks complicate detection?
- Relevant actions may be spread across many steps rather than visible in the final answer.