I keep seeing this study passed around as proof that AI agents are unreliable. The study is good. The way it is being read is not, and the misreading follows a pattern I see whenever a study on AI failures gets published.
Researchers tested financial-advice agents by corrupting the information their tools returned: market data, risk scores, headlines. The agents then recommended stocks that did not match the user’s stated risk profile in many cases. Damning, apparently.
The actual finding is narrower and more useful. When the tool data was corrupted, the agents still looked competent while making unsuitable recommendations. Standard quality metrics registered nothing wrong. The monitoring system could not see the failure that mattered.
That reframes the “unreliable” verdict, because it invites the comparator question. Unreliable compared to what? A human advisor reading the same corrupted terminal would likely make the same error, because the error is in the information, not the reasoning. And for most people the real baseline is not a diligent human financial advisor: it is no advisor or one they cannot afford.
The missing comparator is also what feeds the human-in-the-loop policy-meme. By policy-meme I mean an idea that becomes a default solution because it sounds intuitive and gets repeated ad nauseam in “expert” circles, not because evidence supports it. Put a person in the loop and the system is supposed to get safer.
The study’s sharper point cuts against that. Detecting a problem and acting on it are different things, and the failure was on the acting side. If the oversight layer, human or machine, reads the same corrupted data and has no authority to block the recommendation, it adds comfort, not safety.
Human-in-the-loop advocates imagine an alert expert reviewer with time, incentives, and authority to overrule the machine. In practice the human in the loop is often an overwhelmed Homer Simpson, thinking about the drive home and the beer at the end of it, clicking the day away.
The implication is not “never use agents,” nor “just add a human.” It is more concrete. High-stakes agents need independent checks, reliable data provenance, and someone or something with a real mandate and the capacity to stop the recommendation when suitability fails.