Here’s a super interesting study published in Nature Medicine this month: an RCT in China (n = 2,069) tested a patient-facing LLM chatbot that conducts the intake interview before a specialist visit and generates a structured pre-consultation summary.
Efficiency and quality in service delivery are often treated as a zero-sum game: you either speed things up by cutting corners, or you protect the experience by accepting a slower, more expensive service. This trial suggests that some of this “trade-off” is less a law of nature than a design failure. The results were robust, including similar effects whether patients used the tool independently or with staff support: about a 29% reduction in specialist consultation time, alongside a measurable increase in patient-reported ease of communication with doctors.
But from my standpoint, the real contribution here is methodological. It is a direct challenge to the “local data” fetishism that shows up over and over again in the AI-for-service-delivery narrative: the idea that raw local data is automatically authentic and therefore a virtuous training target.
When the researchers fine-tuned a model on recordings of real primary-care visits, the model did not “improve” the system. It encoded its constraints. It skipped guideline-recommended questions, mirrored an unfriendly tone, and treated the pathologies of an overstrained workflow as the target behavior. The co-designed version performed better because it was not a passive reflection of the status quo. It was an attempt to specify what “good” looks like.
Often, I see people rolling their eyes when I talk about participatory, user-centric design, as if it were a normative luxury. But no: it is a functional requirement. And if your training data is a record of a system under stress, fine-tuning on it will scale friction and burnout, not fix them. There’s a chance you aren’t “learning” the local context, no matter how good it sounds at AI-do-good conferences. You are just automating its pathologies.
You cannot fix a process by sampling from its failure modes. You have to name the norm you want, with your users, because the observed reality, and the local data that describes it, may often be just a transcript of constraints that end-users are trying to escape.