Visitors can see how each model performs on specific benchmarks, as well as what its average score is overall. No model has yet achieved a perfect score of 100 points on any benchmark. Smaug-72B, a new AI model created by the San Francisco-based startup Abacus.AI, recently became the first to break past an average score of 80.

Many of the LLMs are already surpassing the human baseline level of performance on such tests, indicating what researchers call “saturation.” Thomas Wolf, a co-founder and the chief science officer of Hugging Face, said that usually happens when models improve their capabilities to the point where they outgrow specific benchmark tests — much like when a student moves from middle school to high school — or when models have memorized how to answer certain test questions, a concept called “overfitting.”

When that happens, models do well on previously performed tasks but struggle in new situations or on variations of the old task.

“Saturation does not mean that we are getting ‘better than humans’ overall,” Wolf wrote in an email. “It means that on specific benchmarks, models have now reached a point where the current benchmarks are not evaluating their capabilities correctly, so we need to design new ones.”

Some benchmarks have been around for years, and it becomes easy for developers of new LLMs to train their models on those test sets to guarantee high scores upon release. Chatbot Arena, a leaderboard founded by an intercollegiate open research group called the Large Model Systems Organization, aims to combat that by using human input to evaluate AI models.

Parli said that is also one way researchers hope to get creative in how they test language models: by judging them more holistically, rather than by looking at one metric at a time. 

“Especially because we’re seeing more traditional benchmarks get saturated, bringing in human evaluation lets us get at certain aspects that computers and more code-based evaluations cannot,” she said.

Chatbot Arena allows visitors to ask any question they want to two anonymous AI models and then vote on which chatbot gives the better response.

Its leaderboard ranks around 60 models based on more than 300,000 human votes so far. Traffic to the site has increased so much since the rankings launched less than a year ago that the Arena is now getting thousands of votes per day, according to its creators, and the platform is receiving so many requests to add new models that it cannot accommodate them all.

Chatbot Arena co-creator Wei-Lin Chiang, a doctoral student in computer science at the University of California-Berkeley, said the team conducted studies that showed crowdsourced votes produced results nearly as high-quality as if they had hired human experts to test the chatbots. There will inevitably be outliers, he said, but the team is working on creating algorithms to detect malicious behavior from anonymous voters.

As useful as benchmarks are, researchers also acknowledge they are not all-encompassing. Even if a model scores well on reasoning benchmarks, it may still underperform when it comes to specific use cases like analyzing legal documents, wrote Wolf, the Hugging Face co-founder.

That’s why some hobbyists like to conduct “vibe checks” on AI models by observing how they perform in different contexts, he added, thus evaluating how successfully those models manage to engage with users, retain good memory and maintain consistent personalities.

Despite the imperfections of benchmarking, researchers say the tests and leaderboards still encourage innovation among AI developers who must constantly raise the bar to keep up with the latest evaluations.

Leave a Reply

Your email address will not be published. Required fields are marked *

You May Also Like
Lansing news: Former employee Devon Johnson charged in deadly shooting of Andrew Coleman at Nippon Paint Automotive Americas

Lansing: Former Nippon Paint Automotive Americas Employee Devon Johnson Charged in Fatal Shooting of Andrew Coleman

LANSING, Ill. (WLS) — A former employee has been charged with murder…
'Happy Face' killer warns fellow serial killer Rex Heuermann could be 'tossed to the wolves' in prison

‘Happy Face’ killer says accused serial killer Rex Heuermann could face danger in prison

Keith Jesperson — the “Happy Face” serial killer who has been corresponding…
Gilgo Beach serial killer joins infamous group of monsters as he opens ghoulish mind to FBI

Judge gives Rex Heuermann maximum sentence in Gilgo Beach serial killings case

RIVERHEAD, N.Y. — Rex Heuermann, the Long Island serial killer who confessed…
Austin tech leader Joshua Baer identified as victim of Texas plane crash after jet caught fire along highway

Austin Tech Leader Joshua Baer Killed in Texas Plane Crash After Jet Catches Fire on Highway

Joshua Baer, founder of Capital Factory and one of Austin’s most prominent…
Pixar's new curly hair technology in 'Toy Story 5' advances diversity in the animation space

Toy Story 5’s New Curly Hair Technology Marks a Major Leap for Diversity in Animation

LOS ANGELES — Pixar is once again pushing its animation tools forward,…
Man dies after carriage horse gets loose in New York City's Central Park, crash

Central Park Carriage Horse Crash Leaves Man Dead After Runaway Incident in NYC

NEW YORK — An 18-year-old man died after being critically injured in…
Smiling suspect stands out as authorities release mugshots of 5 accused in alleged White House UFC attack plot

Authorities release mugshots of five suspects in alleged White House UFC attack plot, with one image drawing attention

New details emerge on alleged UFC terror plot targeting White House Authorities…
Chicago, Illinois weather forecast: Tornado Watch issued for parts of Chicago area | Radar

Chicago Weather Alert: Tornado Watch Issued Across Parts of the Chicago Area — Live Radar Updates

Severe weather is expected to impact the Chicago area on Wednesday, with…
New Mexico seeks massive penalty from Meta after jury found tech giant liable for endangering children

New Mexico Demands Massive Meta Penalty After Jury Finds Facebook Parent Liable for Endangering Children

New Mexico’s Department of Justice is pushing to make Meta pay far…
Man killed after horse-drawn carriage bolts and flips near popular New York City tourist destination

Man Dies After Horse-Drawn Carriage Flips Near Central Park in New York City

An 18-year-old tourist from India was killed Wednesday after a horse-drawn carriage…
Florida couple sues fertility clinic after allegedly giving birth to someone else's baby

Florida Couple Settles With Biological Parents in Alleged IVF Embryo Mix-Up Case

A Florida couple who say a fertility clinic mistakenly implanted the wrong…
Chicago crime: Suspect Merlin Lu, 21, charged with hate crime, arson for burning cross in Grant Park, police say

Chicago Police Charge 21-Year-Old Merlin Lu With Hate Crime, Arson After Cross Burning in Grant Park

CHICAGO (WLS) — A 21-year-old Chicago man is facing a series of…