ChatGLM, BELLE, MOSS, InternLM, Qwen, Baichuan, Yi and DeepSeek, the Chinese open chat models that followed ChatGPT
The chapter's diffusion tables name only ChatGLM-6B, Ernie Bot, Qwen and DeepSeek, but the public-code record lists eight open model repositories created between March and November 2023.
- Date
- March 2023
- Who
- Zhipu AI/Tsinghua (ChatGLM), Alibaba (Qwen), Baichuan, 01.AI, DeepSeek-AI and others
- Confidence
- Medium (dates are GitHub repository-creation proxies unless a README or paper sa
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Supporting · Significance: 3/5 · Org(s): Zhipu AI/Tsinghua (ChatGLM), Alibaba (Qwen), Baichuan, 01.AI, DeepSeek-AI and others · Confidence: Medium (dates are GitHub repository-creation proxies unless a README or paper says otherwise) The chapter's diffusion tables name only ChatGLM-6B, Ernie Bot, Qwen and DeepSeek, but the public-code record lists eight open model repositories created between March and November 2023. The repository creation dates are ChatGLM-6B 2023-03-13, BELLE 2023-03-17, MOSS 2023-04-15, InternLM 2023-07-06, Qwen 2023-08-03 (the README says Qwen-7B and Qwen-7B-Chat were released that day), Baichuan 2 2023-08-31, Yi 2023-11-03 and DeepSeek LLM 2023-11-29 (GitHub API metadata for each organization's repository; repository creation can precede or follow the first public release, so treat these as proxies). The post-training recipes diverged. Qwen's report (arXiv 2023-09-28) describes a reward model and PPO, with Qwen-14B-Chat (RLHF) beating the SFT version in an in-house human evaluation of 300 Chinese instructions, but the README says the RLHF-trained chat models were "not released yet"; Baichuan 2's report (arXiv 2023-09-19) describes SFT plus RLHF; Yi's chat models, per its report (arXiv 2024-03-07), were fine-tuned on fewer than 10K carefully verified instructions, a LIMA-like choice (Inference; B05-33); DeepSeek LLM Chat used about 1.5 million SFT instances and then DPO, not PPO (B05-35, B05-37). So "Chinese labs copied RLHF" is too coarse. Qwen and Baichuan describe RLHF, Yi describes a small curated SFT set, DeepSeek used DPO, and Qwen did not release its RLHF models. I did not check ChatGLM, BELLE, MOSS or InternLM's alignment methods; see the Backlog and B20. Sources: Qwen README · Qwen report · Baichuan 2 · Yi · DeepSeek LLM · DeepSeek-LLM repo · BELLE repo