TIME's Kenya report and the outsourcing vendors behind RLHF labeling
RLHF needs people to write, rank and rate model outputs at scale; TIME and The Verge showed that much of this runs through outsourcing vendors, often at low pay and high secrecy, and TIME documented…
- Date
- 18 January 2023
- Who
- OpenAI, Sama, Scale AI (Remotasks), Surge AI, Upwork, Amazon Mechanical Turk, Lionbridge
- People
- Billy Perrigo (TIME), Josh Dzieza (The Verge), Edwin Chen (Surge AI)
- Confidence
- High (TIME's documents and company statements); Medium (pay ranges, which are di
- Deep dive
- RLHF and instruction tuning (how base models became assistants)
Tier: Landmark · Significance: 4/5 · Org(s): OpenAI, Sama, Scale AI (Remotasks), Surge AI, Upwork, Amazon Mechanical Turk, Lionbridge · People: Billy Perrigo (TIME), Josh Dzieza (The Verge), Edwin Chen (Surge AI) · Confidence: High (TIME's documents and company statements); Medium (pay ranges, which are disputed or anonymous) Primary sources: TIME, 2023-01-18 · The Verge/New York Magazine, 2023-06-20
One-liner. RLHF needs people to write, rank and rate model outputs at scale; TIME and The Verge showed that much of this runs through outsourcing vendors, often at low pay and high secrecy, and TIME documented one case where the most disturbing labeling went to low-paid workers.
Why it happened. The preference loop needs a lot of labor. InstructGPT used about 40 contractors in a close, researcher-supervised relationship (B05-13); Anthropic about 30 "select" workers plus a larger per-task pool (B05-14); Llama 2 relied on vendor annotators and noted that vendors differ markedly in downstream quality (B05-37). Vendors existed because image annotation had already built an industry. The Verge traces it to ImageNet's use of Mechanical Turk and describes Scale AI, founded in 2016 and valued at $7.3 billion in 2021, as selling labeled data to OpenAI and the US military (Verge). OpenAI used Scale for labels as early as the 2019 GPT-2 work (B05-05). Early papers named their contractors in the acknowledgments (about 40 labelers in InstructGPT); the workers in the 2023 reporting below are anonymous.
What TIME reported. Beginning in November 2021, OpenAI sent tens of thousands of text snippets, many describing sexual abuse, violence, hate, self-harm and similar content, to Sama, an outsourcing firm with workers in Kenya, Uganda and India, to label for a toxicity detector. Per documents TIME reviewed, there were three contracts worth about $200,000 (OpenAI said about $150,000) and roughly three dozen workers in three teams, with an hourly billing rate of $12.50 versus take-home pay of about $1.32 to $2 per hour. Three workers said they were expected to label 150-250 passages per nine-hour shift (Sama said 70 and that pay ranged from $1.46 to $3.74), and four workers called the wellness counseling unhelpful (Sama disputed this). Sama canceled the work in February 2022, eight months early, after a separate image-collection pilot that included illegal categories; in January 2023 Sama announced it would exit all NLP and content-moderation work to focus on computer-vision annotation, with the exit to be complete as of March 2023 (TIME). The Sama story did not end there, and B05-43d covers the 2026 layoffs. OpenAI confirmed that Sama staff contributed to a tool to detect toxic content that was built into ChatGPT, and said this also helped remove toxic data from training sets.
What The Verge reported. Remotasks is the worker-facing subsidiary of Scale AI, with a different name by design (Scale cited customer confidentiality); workers in Kenya, the Philippines and elsewhere were organized by anonymous project code names; task instructions grew to dozens of pages. Chatbot-training work ("chatbot trainer") was among the better-paid categories. The reporter found a Texas worker paid about $14 an hour to chat with a DeepMind model code-named "Dolphin", later identified as Sparrow (B05-16); an annotator whose task instructions were nearly identical to OpenAI's, and who therefore was likely training ChatGPT, said they earned about $3 an hour; Surge AI workers reported $15 to $30 an hour, and Surge's CEO said it had about 100,000 annotators and that requests surged after ChatGPT; specialist annotation could pay $50 or more per hour. OpenAI, Microsoft, Meta and Anthropic declined to comment on annotator numbers, pay or location; DeepMind's Geoffrey Irving said Sparrow annotators were paid at least a local living wage (Verge).
How it spread.
| Lab | Evidence of vendor or worker model | Date |
|---|---|---|
| OpenAI | Scale AI (2019); Upwork and Scale contractors (2022); Sama, Kenya (2021-22 safety labeling) | 2019 → 2023-01 |
| Anthropic | Upwork and US MTurk workers; Surge AI affiliations on a 2022 paper | 2022-04, 2022-12 |
| DeepMind | Remotasks project "Dolphin" for Sparrow (Verge) | reported 2023-06 |
| Meta | Unnamed "vendor-based" annotators; named internal annotation leads in acknowledgments | 2023-07 |
| Surge relabeled a Google emotion dataset previously labeled in India (Verge, per Surge) | reported 2023-06 |
The common pattern is outsourcing at several removes, strict confidentiality, and little public pay information; the shift from "crowd" to "expert" labor after 2023 is in B05-42.
Why it mattered. It put a human face on "alignment" and moved labor conditions into the AI-ethics debate. It also showed the economics, since labs treat labeling as a procurement item while their own papers (InstructGPT, HH-RLHF) describe it as a close, researcher-supervised relationship with a small group.
Nuance, controversy and myths. (1) The Sama work was safety/toxicity labeling, not RLHF preference ranking; the popular summary "Kenyan workers did RLHF for ChatGPT" overstates what TIME documented (B05-20). (2) The pay figures conflict. Sama disputed parts of TIME's account, OpenAI said it set no productivity targets, and the Verge's pay figures come from anonymous workers. (3) TIME itself separates Sama's Facebook moderation contract from the OpenAI work, and the Kenyan court cases against Meta and Sama (183 moderators; ruled maintainable in April 2023) concern Facebook content moderators, not OpenAI, so they should not be conflated (TechCrunch, 2023-04-20). (4) Self-reported satisfaction in InstructGPT came from 19 survey respondents (B05-13).
Interview kit.
- 30-second version: RLHF depends on large amounts of human labeling, much of it outsourced. TIME found OpenAI's safety-labeling contractor in Kenya paid roughly $1.32-$2 an hour; The Verge showed the vendor layer, such as Scale's Remotasks and Surge, and the secrecy around it.
- Likely follow-ups: Did Kenyan workers train ChatGPT with RLHF? → They labeled harmful text for a toxicity detector that was built into ChatGPT, per TIME and OpenAI. Who are the main vendors? → Scale AI, Surge AI and platforms like Upwork and MTurk; later Mercor and others (B05-42).
- Common mistake: Quoting one pay figure as fact.
- Connect it to: B05-13, B05-14, B05-42, B18.
Sources. 1. TIME, Perrigo, 2023-01-18 · 2. The Verge, Dzieza, 2023-06-20 · 3. Ouyang et al. · 4. Bai et al. · 5. Touvron et al. · 6. TechCrunch on the Kenya Meta/Sama case