HyperAIHyperAI

Command Palette

Search for a command to run...

2 months ago
LLM
Generative AI

AI outperforms humans in classic Turing test

A new study from the University of California San Diego provides the first empirical evidence that modern artificial intelligence can successfully pass the classic Turing test. Published in the Proceedings of the National Academy of Sciences, the research demonstrates that advanced large language models (LLMs) can mimic human conversation so convincingly that people often fail to distinguish them from real individuals. This marks a historic shift, as it is the first time AI has been judged as human as frequently as actual humans within the framework established by Alan Turing in 1950. The experiments involved 500 participants divided into two groups: UC San Diego undergraduates and online users recruited via Prolific. Participants engaged in brief, text-based conversations with two parties simultaneously—one being a human and the other an AI—and were tasked with identifying the real person. The study tested four models, including state-of-the-art systems like GPT-4.5 and LLaMa-3.1-405B, alongside older baselines like GPT-4o and the 1960s chatbot ELIZA. The results revealed a significant performance gap. When provided with specific persona prompts instructing them to adopt a human character with specific tone, humor, and even fallibility, GPT-4.5 was incorrectly identified as human 73% of the time, outperforming the actual human participants. LLaMa-3.1-405B was judged human 56% of the time, a rate statistically indistinguishable from the human baseline. In stark contrast, baseline models without such detailed prompting or older versions like ELIZA and GPT-4o were selected as human less than 25% of the time. Cameron Jones, the study's corresponding author and a UC San Diego PhD graduate, noted that while LLMs are known for their knowledge base, this test proves they can also convincingly display social behavioral traits. Ben Bergen, a co-author and cognitive science professor, explained that the models are no longer succeeding through raw intellectual power but by exhibiting humanlike imperfections. The study found that without explicit instructions to act like a human, the models failed to adopt these traits, suggesting they possess the ability to appear human but require human guidance to understand how to do so effectively. The implications of these findings extend beyond academic curiosity to real-world concerns regarding trust and deception. The researchers warn that the ease with which AI can now pass as human raises significant risks for online interactions, where prolonged conversations can lead to genuine deception. Bergen highlighted the potential for bad actors to use these sophisticated bots to manipulate users into sharing sensitive information, influencing voting behavior, or driving sales. The study challenges the traditional interpretation of the Turing test. Originally designed to measure machine intelligence equivalent to human intellect, the test now increasingly measures humanlikeness. As AI systems become better at simulating human error and personality, society must reconsider the safeguards needed to protect against automated deception. The researchers have made the test interface available online to help the public better understand these capabilities and the potential for AI to be indistinguishable from people in digital spaces.

Related Links