Help needed: human-in-the-loop hype check
The more I think about it, the more I suspect that blanket calls for “human-in-the-loop” in AI for public services are a first-world comfort blanket. In places where there are no doctors, teachers, or caseworkers, “looping in a human” often just means there is no loop at all.
My hunch: the value of “the loop” depends on context. Sometimes it saves lives, other times it just slows things down. Yet it keeps being sold as a universal must-have, usually by people who never had to wait in line for basic services, but who still shape policy far more than those who are systematically excluded from service provision.
So here’s my ask: help me pressure-test this.
Am I being unfair? Or is “human-in-the-loop” just a way for elites to reassure themselves, when the only loop they’re really in is talking to each other at AI-for-Good conferences?
That is why I am keen to hear from others: are there rigorous counterfactuals or experimental studies showing when human oversight truly improves outcomes, and when it simply adds friction? And how much does this depend on context and task type?
(And a long BTW: on AI bias, I am particularly interested in cases where AI introduces more bias than humans. Otherwise, the comparative advantage may still be with the machine.)
So: what am I missing? Literature tips, counterexamples, or data that would make me less skeptical are most welcome.
Because at the end of the day, “human-in-the-loop” sounds very different when you’re a high-flying AI civil servant, activist, or advocate with private health insurance; than when you’re a citizen waiting hours or days in line for a basic service.