Scientists identify the main weakness of modern AI models

15:01    15 June, 2026

An international group of researchers has conducted an experiment that revealed a serious and previously understudied problem in leading language models. It turns out that when the length of a task requiring sustained attention increases, the accuracy of AI drops sharply — to the point of virtually refusing to follow instructions.

The study results were published in the journal PNAS Nexus.

A classic test in new conditions

To test this, the scientists used the famous Stroop test — a psychological experiment first developed in 1935. The subject is shown words denoting colors ("red," "blue," "green") but written in a color that does not match the meaning of the word. The task is to name the color of the ink while ignoring the word itself.

Humans handle this relatively consistently even with long lists, although they experience cognitive conflict. The brain successfully suppresses the automatic reading response.

Researchers led by Suketu Patel adapted the test for AI and tested several leading models on it:

Shocking results

With short lists (5 words), all models showed high accuracy. However, as task length increased, the results deteriorated dramatically:

The models gradually "forgot" the original instruction and began simply reading the written word — that is, they reverted to the strongest pattern on which they had been trained.

A fundamental difference from humans

Unlike humans, who are capable of sustaining voluntary attention and suppressing automatic reactions over long periods, modern AI demonstrates extremely low tolerance for prolonged cognitive loads. In essence, the longer the task, the more pronounced this fundamental flaw becomes.

In brief

Scientists have discovered a serious weakness in modern language models: as the length of a task requiring concentration increases, their accuracy drops sharply. The Stroop test showed that GPT-4o and Claude 3.5 Sonnet, even at moderate list lengths, virtually stop following instructions. This fundamental difference between AI and human thinking underscores the need for further research into the area of sustained attention in artificial intelligence.



© NEWS.am Medicine