The invisible workforce: data labour and the Global South

Lesson 5 of 5 in The UK, Canada, Brazil — and the Missing Map: Africa, the Middle East, Latin America.

Every frontier model you have studied in this curriculum was built, in part, by workers you have never heard of. Reinforcement learning from human feedback needs humans. Safety classifiers need people to read the worst text on the internet and label it. Autonomous-vehicle vision needs millions of hand-annotated frames. This is data labour — annotation, content moderation, model evaluation — and it is overwhelmingly performed in the Global South: Nairobi, Manila, Caracas, Gurugram, often through layered outsourcing chains that keep it off the AI industry’s books and outside most AI governance.

The case that made it visible: in January 2023, a Time investigation by Billy Perrigo revealed that OpenAI, via the outsourcing firm Sama, paid Kenyan workers roughly $1.32 to $2 per hour to label textual descriptions of child sexual abuse, violence and self-harm — the training data for the safety filters that made ChatGPT deployable. Workers described lasting psychological trauma; counselling provision was disputed; Sama terminated the contract early and later exited content-moderation work entirely.

The pattern repeats beyond moderation. Annotation platforms recruited heavily in Kenya, the Philippines and crisis-era Venezuela, where currency collapse made piece-rate dollar work viable; when Scale AI’s Remotasks platform abruptly exited Kenya and several other markets in March 2024, thousands of workers lost income overnight, some with pay disputes unresolved — no notice period, no severance, no local employer to pursue. Researchers — the Oxford Internet Institute’s Fairwork project, the ILO’s digital-platform studies, Milagros Miceli’s data-work research, the DAIR institute — have documented the same structural features across the industry: piece rates and opaque task pricing, no occupational-health provision for psychologically hazardous work, sudden platform exits, and contracts routing disputes to foreign law.

Scholars Nick Couldry and Ulises Mejias call the wider arrangement data colonialism: value extracted from the South as raw material — data, labour, minerals — refined into products in the North, and sold back. You do not need to adopt the term to see the governance point the AU strategy makes with its data-sovereignty language: the AI value chain has a human layer, and almost no AI law reaches it.

Why doesn’t the EU AI Act cover these workers?

The Act regulates AI systems and models — their risks to users and affected persons — not the labour conditions of the people who build them. A provider can be fully EU-compliant while its annotation subcontractors work in conditions no EU employer could lawfully offer. Labour law is territorial; the work was deliberately placed where protections are weakest and enforcement thinnest.

What legal hooks exist today?

Four, all partial. Local labour and constitutional law — the Kenyan cases show host-country courts can pierce the outsourcing veil. Supply-chain due-diligence law — the EU’s CSDDD-style duties could, in principle, force large AI firms to assess labour conditions in their data supply chains, the same logic applied to cobalt and cotton. Procurement leverage — governments and enterprises can demand fair-work certification (Fairwork-style scoring) from AI vendors. Platform-work regulation — the EU Platform Work Directive gestures at algorithmic management, but its reach stops at EU borders.

Isn’t this just outsourcing economics — why is it an AI-governance issue?

Because the harms are produced by the AI production process itself: exposure to traumatic content exists so that safety filters can exist; annotation piece rates are set by the model-training pipeline’s economics. A governance discipline that audits training data for bias but not for the conditions under which it was labelled has drawn its system boundary to exclude the humans inside the system. Expect this boundary to be contested in due-diligence law, ESG disclosure and procurement long before it appears in AI statutes.

What would better look like?

Concrete proposals on the table: living-wage floors and trauma-informed limits written into annotation contracts; disclosure of data-work supply chains in model documentation; recognised collective representation (the moderators’ union model); host-country enforcement capacity; and buyer-side standards — several researchers argue a ‘fair-trade data’ certification is the most realistic near-term instrument, because it prices labour conditions into the one thing AI firms respond to: procurement.

Key terms: data labour, data colonialism, RLHF, vendor AI risk

Interactive checkpoint quiz (2 questions) — open this page in a browser to take it.