New Study Suggests ChatGPT Is Getting Dumber—But, Is It?

A new study released on Tuesday by researchers from Stanford University and University of California, Berkeley has ignited debate within the AI community regarding the performance of OpenAI’s GPT-4 language model. 

The paper, titled “How Is ChatGPT’s Behavior Changing over Time?” and published on arXiv by Lingjiao Chen, Matei Zaharia, and James Zou, investigates changes in GPT-4’s outputs over a span of a few months, suggesting a potential decline in coding and compositional task abilities.

The study utilized API access to test GPT-3.5 and GPT-4 versions from March and June 2023 on various tasks, including math problem-solving, answering sensitive questions, code generation, and visual reasoning. 

Source: Chen, Zaharia, Zou

Notably, the research found a significant drop in GPT-4’s ability to identify prime numbers, plummeting from an accuracy of 97.6 percent in March to just 2.4 percent in June. Surprisingly, GPT-3.5 displayed improved performance during the same period.

This investigation comes amidst growing concerns expressed by users who have observed a subjective decline in GPT-4’s performance over the past few months. Speculations about the reasons behind this decline abound, including OpenAI’s possible distillation of models to enhance efficiency, fine-tuning to mitigate harmful outputs, and unfounded conspiracy theories suggesting a reduction in GPT-4’s coding capabilities to promote GitHub Copilot usage.

OpenAI has consistently denied the alleged decrease in GPT-4’s capabilities. Peter Welinder, OpenAI’s VP of Product, recently took to Twitter to counter the claims, asserting that each new version of the AI language model is more advanced than its predecessor. He posits that intensive usage may lead to heightened awareness of these perceived issues.

But the company, unlike its large language model, isn’t completely closed to the possibility. Logan Kilpatrick, OpenAI’s head of developer relations, confirmed on Twitter that the team is aware of the reported regressions and is actively investigating the matter.

Arvind Narayanan, a computer science professor at Princeton, also pointed out some problems in the study, arguing that the research’s findings do not definitively prove a decline in GPT-4’s performance and could align with fine-tuning adjustments made by OpenAI. 


Information for this story was found via Twitter, Gizmodo, Ars Technica, and the sources and companies mentioned. The author has no securities or affiliations related to the organizations discussed. Not a recommendation to buy or sell. Always do additional research and consult a professional before purchasing a security. The author holds no licenses.

One Response

  1. This issue isn’t even dumber or smarter – the issue is the performance is changing radically even in a short time. Inconsistency directly correlates to unreliability, and this isn’t a GIGO situation so much as GO potentially happening any time for any reason.

Video Articles

Gold and Silver May Be Ready for Another Run | Shawn Khunkhun – Contango Silver & Gold

Silver Is Strong Again, and This Producer Is Ramping Up | Arturo Prestamo – Santacruz Silver

Gold Giant Agnico Eagle Makes a Critical Minerals Bet | Avenir Minerals x Fox River

Recommended

Altamira Gold Extends Maria Bonita Porphyry System Westward With 70.6 Metres At 0.51 g/t Hit

Antimony Resources Reports 13.9% Antimony in Latest Drill Core at Bald Hill

Related News

Elon Musk Isn’t Happy About Apple Partnering with OpenAI, Says It’s ‘Creepy’

Elon Musk has threatened to ban Apple (Nasdaq: AAPL) devices at his companies if the...

Tuesday, June 11, 2024, 07:55:59 AM

OpenAI Doesn’t Expect to Be Profitable Until 2029

OpenAI‘s backers are a long-ish way from making money. Recent financial projections indicate the AI...

Thursday, October 10, 2024, 01:23:00 PM

Florida AG Targets OpenAI And ChatGPT in Probe of 2025 FSU Mass Shooting

Florida Attorney General James Uthmeier has launched a formal investigation into OpenAI, targeting its ChatGPT...

Thursday, April 9, 2026, 01:51:42 PM

Musk Reveals ChatGPT Competitor, xAI Grok, And It’s Already Provided Fake Information

Elon Musk on Friday night unveiled what is referred to as xAI’s Grok system, which...

Sunday, November 5, 2023, 09:07:00 AM

Nvidia to Invest Up to $100 Billion in OpenAI Under Massive Infrastructure Partnership

Nvidia Corp. (Nasdaq: NVDA) announced Monday it intends to invest up to $100 billion in...

Tuesday, September 23, 2025, 12:55:00 PM