Recent studies by Epoch AI and Anthropic have revealed that current AI models, including GPT-5.6 Sol and Claude Fable 5, significantly overstate their research capabilities. Despite their ability to conduct experiments, these models are still far from achieving true autonomy in scientific research.

The research highlighted that, at best, GPT-5.6 Sol achieved only 15 percent of the human reference score during evaluations. This score was primarily derived from established methods that researchers were already familiar with, underscoring the models’ reliance on existing knowledge rather than innovative thinking.

One of the critical weaknesses identified in these AI systems is their lack of genuine self-criticism and creative thought processes. As AI technology continues to evolve, these findings raise important questions about the current state of AI research and its implications for future developments in the field.


Compiled automatically by the Tech AI Newsdesk from public AI-news sources and summarised in our own words.