LMSYS launches Chatbot Arena, a leaderboard built from crowdsourced pairwise votes
LMSYS launched Chatbot Arena on 2023-05-03. Users chat with two anonymous models, vote for the better answer, and votes are converted to Elo ratings; the first leaderboard had nine models and about…
- Date
- 3 May 2023
- Who
- LMSYS (UC Berkeley, UCSD, CMU and others)
- People
- Lianmin Zheng, Ying Sheng, Wei-Lin Chiang and others
- Confidence
- High
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): LMSYS (UC Berkeley, UCSD, CMU and others) · People: Lianmin Zheng, Ying Sheng, Wei-Lin Chiang and others · Confidence: High LMSYS launched Chatbot Arena on 2023-05-03. Users chat with two anonymous models, vote for the better answer, and votes are converted to Elo ratings; the first leaderboard had nine models and about 4,700 votes, with Vicuna-13B on top. The authors said their own GPT-4-based evaluation in the Vicuna launch gave no scalable, incremental way to rate models, which motivated the Arena (LMSYS blog). The follow-up paper introduced MT-Bench and tested LLM-as-a-judge, finding that strong judges such as GPT-4 reached over 80% agreement with controlled and crowdsourced human preferences, the same level as agreement between humans, while showing position, verbosity and self-enhancement biases (arXiv:2306.05685, v1 2023-06-09). It belongs in this chapter because the Arena is a preference-learning interface with the labor outsourced to volunteers, and it became the de facto yardstick for the RLHF-style assistants this chapter describes, with the length and style biases that RLHF critiques predict (B05-38b; B23). Sources: LMSYS Arena post · Zheng et al.