LMSYS launches Chatbot Arena, a leaderboard built from crowdsourced pairwise votes

LMSYS launched Chatbot Arena on 2023-05-03. Users chat with two anonymous models, vote for the better answer, and votes are converted to Elo ratings; the first leaderboard had nine models and about…

Date
3 May 2023
Who
LMSYS (UC Berkeley, UCSD, CMU and others)
People
Lianmin Zheng, Ying Sheng, Wei-Lin Chiang and others
Confidence
High
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): LMSYS (UC Berkeley, UCSD, CMU and others) · People: Lianmin Zheng, Ying Sheng, Wei-Lin Chiang and others · Confidence: High LMSYS launched Chatbot Arena on 2023-05-03. Users chat with two anonymous models, vote for the better answer, and votes are converted to Elo ratings; the first leaderboard had nine models and about 4,700 votes, with Vicuna-13B on top. The authors said their own GPT-4-based evaluation in the Vicuna launch gave no scalable, incremental way to rate models, which motivated the Arena (LMSYS blog). The follow-up paper introduced MT-Bench and tested LLM-as-a-judge, finding that strong judges such as GPT-4 reached over 80% agreement with controlled and crowdsourced human preferences, the same level as agreement between humans, while showing position, verbosity and self-enhancement biases (arXiv:2306.05685, v1 2023-06-09). It belongs in this chapter because the Arena is a preference-learning interface with the labor outsourced to volunteers, and it became the de facto yardstick for the RLHF-style assistants this chapter describes, with the length and style biases that RLHF critiques predict (B05-38b; B23). Sources: LMSYS Arena post · Zheng et al.

Read it in the deep dive