Skip to main content

Share story

AI

Study: Platforms that rank the latest LLMs can be unreliable

Study: Platforms that rank the latest LLMs can be unreliable Image: Primary
MIT researchers found that platforms ranking large language models can be skewed by a small number of user interactions. Their study shows that removing a tiny fraction of crowdsourced data can change which models rank at the top. The researchers developed a fast approximation method to test these platforms and pinpoint the individual votes most responsible for shifts in rankings. In one case involving more than 57,000 votes, dropping just two altered the top model. A separate platform that uses expert annotators and higher quality prompts required removal of 83 out of 2,575 evaluations, or about 3 percent, to flip the results. Tamara Broderick, an associate professor at MIT and senior author of the study, said the platforms proved more sensitive than expected. She noted that if the top ranked model depends on only two or three pieces of feedback out of tens of thousands, users cannot assume it will consistently outperform others when deployed. The work will be presented at the International Conference on Learning Representations. The researchers suggest platforms gather more detailed feedback, such as confidence levels for each vote, to reduce the impact of noise or user error. They also propose using human mediators to assess crowdsourced responses. The study was funded in part by the Office of Naval Research, the MIT IBM Watson AI Lab, the National Science Foundation, Amazon, and a CSAIL seed award.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from MIT News and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
Infrastructure Capital
Infrastructure Capital

Nscale secures $3.36 billion in financing ahead of planned IPO

British AI data center developer Nscale said it secured $3.36 billion in convertible-note financing ahead of a planned US stock-market listing. The Third Point-led deal makes $2.36 billion available immediately, while a further $1...

Infrastructure Products
Infrastructure Products

Tower and Japan plan $4 billion optical chip expansion

Tower Semiconductor and Japan's government plan to invest a combined $4 billion in Japanese factories that make chips for optical connections, Tom's Hardware reports. Tower plans to contribute $3 billion and Japan's Ministry of Ec...

Capital AI
Capital AI

Enveda raises $311 million to advance AI-assisted drug candidates

Enveda has raised $311 million in Series E financing at a $2 billion valuation as it moves drug candidates found through its AI-assisted search of natural compounds into human testing. Catalio Capital Management led the round, wit...

Security Infrastructure
Security Infrastructure

Researchers find personal data exposed in around 16,000 Supabase databases

Security firm UpGuard found around 16,000 databases hosted by Supabase with some personal data exposed to the public web, TechCrunch reports. The accessible information included names, addresses and phone numbers; a smaller number...

Security Policy
Security Policy

CISA sets deadlines for agencies to address four exploited software flaws

CISA has added exploited vulnerabilities affecting WSO2 products, Adobe Commerce, Microsoft SharePoint and Mikrotik RouterOS to its Known Exploited Vulnerabilities catalog. Federal agencies using the affected products must apply r...

AI Products
AI Products

S&P Global Energy opens governed data queries to customer AI agents

S&P Global Energy says it has built a way for customers' AI agents to query its structured energy data in natural language. Its domain experts organize datasets into focused Databricks Genie Agents, each with business definitions ...