New study finds AI’s abilities may be exaggerated by flawed tests

NBC Universal, Inc.

Researchers behind a new study say that the methods used to evaluate AI systems’ capabilities routinely oversell AI performance and lack scientific rigor.

The study, led by researchers at the Oxford Internet Institute in partnership with over three dozen researchers from other institutions, examined 445 leading AI tests, called benchmarks, often used to measure the performance of AI models across a variety of topic areas.

Stream Philadelphia News for free, 24/7, wherever you are with NBC10. WATCH HERE

AI developers and researchers use these benchmarks to evaluate model abilities and tout technical progress , referencing them to make claims on topics ranging from software engineering performance to abstract-reasoning capacity . However, the paper, releas

See Full Page

Interests (0)

Settings

New study finds AI’s abilities may be exaggerated by flawed tests

West Virginia University Statler College launches online master’s program for cybersecurity

Wyoming Cows Go High-Tech With Electronic Collars And ‘Virtual Fencing’

Computer Chips in Our Bodies Could Be the Future of Medicine

Unlock Peak Efficiency with the Best Productivity Apps of 2025: Learn to Boost Your Productivity Fast

New app to help residents navigate parking rules in San Francisco

Kim Kardashian's Bold Selfie Holds a Cheeky Surprise

The 'Shock' of a lifetime for a special Steelers fan

'Remain strong': Democrats celebrate 'repudiation' of Trump after election wins

Tommy Tuberville slammed over 'war zone' comments

Scammers use real apartments, agent IDs to con prospective renters

House speaker Nancy Pelosi won’t seek re-election to Congress – NBC10 Philadelphia

Alumni sue University of Pennsylvania after cybersecurity breach

Second Trump-designated holiday is next week: What's the significance?

US ends protected status for South Sudanese nationals

Supreme Court Questions Legality of Trump's Tariffs

Trump pressures GOP senators to end the government shutdown, now the longest ever

Jonathan Bailey calls People's Sexist Man Alive title 'the honour of a lifetime'

Food Distribution Event Draws Large Crowd Amid SNAP Shutdown

Democrat Jay Jones wins race to be Virginia attorney general despite texts endorsing violence

Dick Cheney, powerful former US vice president who pushed for Iraq war, dies at 84

Pittsburgh Man Sets Up Front Yard Food Pantry

Democrats Achieve Significant Wins in Key Elections