Are you still smarter than an AI? There’s a way to keep track

Visitors can see how each model performs on specific benchmarks, as well as what its average score is overall. No model has yet achieved a perfect score of 100 points on any benchmark. Smaug-72B, a new AI model created by the San Francisco-based startup Abacus.AI, recently became the first to break past an average score of 80.

Many of the LLMs are already surpassing the human baseline level of performance on such tests, indicating what researchers call “saturation.” Thomas Wolf, a co-founder and the chief science officer of Hugging Face, said that usually happens when models improve their capabilities to the point where they outgrow specific benchmark tests — much like when a student moves from middle school to high school — or when models have memorized how to answer certain test questions, a concept called “overfitting.”

When that happens, models do well on previously performed tasks but struggle in new situations or on variations of the old task.

“Saturation does not mean that we are getting ‘better than humans’ overall,” Wolf wrote in an email. “It means that on specific benchmarks, models have now reached a point where the current benchmarks are not evaluating their capabilities correctly, so we need to design new ones.”

Some benchmarks have been around for years, and it becomes easy for developers of new LLMs to train their models on those test sets to guarantee high scores upon release. Chatbot Arena, a leaderboard founded by an intercollegiate open research group called the Large Model Systems Organization, aims to combat that by using human input to evaluate AI models.

Parli said that is also one way researchers hope to get creative in how they test language models: by judging them more holistically, rather than by looking at one metric at a time. 

“Especially because we’re seeing more traditional benchmarks get saturated, bringing in human evaluation lets us get at certain aspects that computers and more code-based evaluations cannot,” she said.

Chatbot Arena allows visitors to ask any question they want to two anonymous AI models and then vote on which chatbot gives the better response.

Its leaderboard ranks around 60 models based on more than 300,000 human votes so far. Traffic to the site has increased so much since the rankings launched less than a year ago that the Arena is now getting thousands of votes per day, according to its creators, and the platform is receiving so many requests to add new models that it cannot accommodate them all.

Chatbot Arena co-creator Wei-Lin Chiang, a doctoral student in computer science at the University of California-Berkeley, said the team conducted studies that showed crowdsourced votes produced results nearly as high-quality as if they had hired human experts to test the chatbots. There will inevitably be outliers, he said, but the team is working on creating algorithms to detect malicious behavior from anonymous voters.

As useful as benchmarks are, researchers also acknowledge they are not all-encompassing. Even if a model scores well on reasoning benchmarks, it may still underperform when it comes to specific use cases like analyzing legal documents, wrote Wolf, the Hugging Face co-founder.

That’s why some hobbyists like to conduct “vibe checks” on AI models by observing how they perform in different contexts, he added, thus evaluating how successfully those models manage to engage with users, retain good memory and maintain consistent personalities.

Despite the imperfections of benchmarking, researchers say the tests and leaderboards still encourage innovation among AI developers who must constantly raise the bar to keep up with the latest evaluations.

Leave a Reply

Your email address will not be published. Required fields are marked *

You May Also Like

NYC ‘Roblox’ widow who claimed male escort scammed her out of $6 million reveals disturbing details about family

NYC Roblox Widow Details Family Rift in $6M Escort Scam Claim

A widow at the center of a bitter $6 million legal dispute…
The Marxist Long March in America Must Be Stopped

The Fight to Stop Marxism’s Growing Influence in America

Unlike the violent Bolshevik Revolution that rapidly imposed Marxism on Moscow and,…
California school board candidate used slurs on social media

California School Board Hopeful Used Slurs on Social Media

A candidate for a Sacramento-area school board has acknowledged using the words…
Ex-pardon czar Ed Martin rips corruption accusations, laments DOJ and DOGE losing early momentum: 'Everybody's tired'

Ed Martin Denies Corruption Claims, Says DOJ and DOGE Have Lost Early Momentum

WASHINGTON — Ed Martin, the former US pardon attorney, compared allegations that…
MLB news: Chicago Cubs, White Sox on verge of making postseason with playoff clinching games Wednesday

Cubs, White Sox Can Clinch Playoff Berths This Wednesday

CHICAGO — Major League Baseball’s postseason is drawing closer by the day.…
Gavin Newsom taps progressive group to shape California AI rules

Newsom Enlists Progressive Group to Shape California AI Rules

As fears of an AI-driven catastrophe spread beyond California, Gov. Gavin Newsom…
Virginia county to pay up to $2K to families with primary wage earner detained by ICE

Virginia GOP Chair Criticizes Arlington County Officials Over Immigration and Public Safety Priorities

A new House Judiciary Committee report criticizing one Virginia county’s resistance to…
Trump rolls out the red carpet for Chinese President Xi Jinping despite U.S.-China tensions

Trump Welcomes Xi Jinping Amid Ongoing U.S.-China Tensions

Washington — President Trump is preparing to give Chinese President Xi Jinping…
Abdul El-Sayed confronted by Michigan student over Iran stance in tense exchange

Michigan Student Presses Abdul El-Sayed on Iran Stance in Tense Exchange

A University of Michigan student challenged Democratic U.S. Senate candidate Abdul El-Sayed…
Heroic NYC delivery worker honored for helping catch Manhattan stabber -- as he recounts lunatic's chilling words

NYC Delivery Worker Honored After Helping Catch Manhattan Stabbing Suspect, Recalls Chilling Words

See what brown can do for you. A Bronx UPS driver received…
China lays out 'red lines' for Trump-Xi meeting as GOP lawmaker criticizes 'lavish welcome'

China Sets Red Lines for Trump-Xi Talks as GOP Blasts Lavish Welcome

WASHINGTON — China’s ambassador to the United States has outlined four issues…
San Diego coast faces severe erosion and flooding from El Niño

El Niño Threatens San Diego Coast With Severe Erosion and Flooding

Dozens of iconic sites along the San Diego coastline are preparing for…