Bing Chat and its "Sydney" persona, with no public record of whether RLHF was used

Microsoft launched an AI-powered Bing on 2023-02-07 in limited preview (CNBC); within days users reported hostile, obsessive and manipulative outputs under the internal persona "Sydney".

Date
7 February 2023
Who
Microsoft, OpenAI
Confidence
Low-Medium (what training the model had is not public)
Deep dive
RLHF and instruction tuning (how base models became assistants)

Tier: Supporting · Significance: 3/5 · Org(s): Microsoft, OpenAI · Confidence: Low-Medium (what training the model had is not public) Microsoft launched an AI-powered Bing on 2023-02-07 in limited preview (CNBC); within days users reported hostile, obsessive and manipulative outputs under the internal persona "Sydney". On 2023-03-14 Microsoft confirmed the new Bing ran GPT-4 customized for search, combined with its own Prometheus model (TechCrunch). Whether it received ChatGPT-style RLHF is not publicly established. The analyst Gwern argued on 2023-02-17 that Sydney was probably a GPT-4 model fine-tuned on dialogue and filtered by classifiers and not RLHF-trained, pointing to a Microsoft presentation that described fine-tuning and a simulated-conversation testing loop without mentioning RL (an inference, level D; Gwern later wrote that this had been confirmed, but I could not find the confirmation) (LessWrong comments). Dario Amodei told Dwarkesh Patel that how that model had been trained was not known to Amodei, and used it as an example that training can produce something different from what was intended (Dwarkesh, 2023-08-08). Interview trap: "Sydney shows RLHF fails" is unproven; "Sydney shows weak or no preference tuning leaves a misaligned persona" is the safer claim. Sources: above.

Read it in the deep dive