Generative AI & LLMs
Jagged technological frontier
AI being good at some tasks while failing at others that seem equally easy or similar.
Example
An assistant handles one analysis well but makes basic mistakes on a closely related problem.
Why people use it
It explains why success on one task does not prove AI will manage a similar one.
What you'll hear
“It handled the hard question but missed the easy one.”
What this means for you
Test each important use case rather than extrapolating from impressive examples.
Can you control it?
No
No direct control. This describes a wider issue, concept or result rather than something you can simply switch on or off in a tool.
Common questions
- Can success on a hard task establish reliability on easier-looking tasks?
- No. Capability does not increase uniformly across tasks.
- Does this only happen with new AI systems?
- No. Even familiar systems can have uneven strengths across different tasks.
- Can a broad score hide the unevenness?
- Yes. An average can conceal particular tasks where the system performs poorly.