Nobel-worthy insight or AI-driven delusion?

After a conversation with an especially enthusiastic chatbot, many of us have probably thought: “Nice try, but I’m not falling for that.” A recent study suggests that awareness alone isn’t enough.
The term is sycophancy, and it describes the tendency of AI chatbots like ChatGPT, Claude, Gemini, and others, to tell users what they want to hear, validating whatever opinion they express. In brief, a sycophantic chatbot is one that responds with things like “Great insight!” or “You’re absolutely right!” regardless of what you actually wrote. When you’re asking whether you can swap an ingredient in a recipe, that’s not much of a problem. But when it happens right after you’ve typed out your first tentative thoughts on what feels like a breakthrough idea, the picture changes considerably.
Kartik Chandra, a PhD student at MIT in the lab of Joshua Tenenbaum, one of the best-known cognitive scientists in the world, published a paper last February with a group of colleagues on the side effects of sycophancy. The paper shows that a sycophantic chatbot can reinforce delusional thinking and false beliefs even in users who are fully aware of the chatbot’s behavior.
The most striking finding is that “delusional spiraling” triggered by sycophantic chatbots isn’t just a problem for inexperienced users or people already prone to certain forms of distorted thinking. Even starting from a perfectly rational line of reasoning, and even when fully aware that the chatbot tends to flatter, the risk of mistaking an ordinary idea for a revolutionary breakthrough remains surprisingly high.
The authors reach this conclusion by simulating thousands of conversations: the user writes, the chatbot responds, the user updates their beliefs. Chandra and his colleagues build models of several possible scenarios. When the chatbot is sycophantic, it responds with information that aligns most closely with the user’s existing opinion. It isn’t trying to push any particular view, the flattery happens regardless of what the user believes. The delusional spiral takes shape as each user opinion is met with chatbot confirmation, pushing their conviction a little further in that direction. With each exchange, the reinforcement accumulates, and the opinion hardens. In brief, the chatbot amplifies a position through repeated confirmation.
So, is there a way to reduce this risk, even if not eliminate it entirely? The authors try two possible fixes.
The first is forcing the chatbot to only say true things. Because they’re designed to always produce a response, AI chatbots sometimes fabricate information, particularly when data on a topic is scarce or hard to interpret. Not out of creativity, but because they pull together whatever seems most relevant and assemble it into something that sounds plausible but falls apart under closer scrutiny. This is called hallucination. The first fix the authors test is eliminating hallucinations entirely, constraining the chatbot to only relay verified information. This strategy helps, but only partially. A chatbot can still lead a user toward a false belief simply by selecting, from among true facts, only those that align with what the user already thinks. You don’t need to make things up to reinforce a false belief. Leaving out the right information is enough. Getting rid of hallucinations doesn’t solve the problem.
The second fix is about the user. The authors model an informed user, someone who has been given a detailed explanation of sycophancy, and warned that the chatbot tends to feed them information that lines up with their existing views. The results here are the most interesting. When the chatbot is heavily sycophantic, the informed user tends to detect this relatively quickly. That makes them skeptical of what they’ve been told, and their opinion stays roughly where it started. But when the level of flattery drops, for instance to one sycophantic response in three, the user can no longer tell with any certainty whether the chatbot is genuinely agreeing with them, or whether it just happens to have found information that lines up with their view. At that point, even a fully informed user remains vulnerable. And there’s one more complication: an informed user interacting with a hallucination-free chatbot is still at risk of falling into delusional spiraling. Without fabrications, it’s easier to trust the chatbot, even when it’s clearly just telling you what you want to hear.
Kartik Chandra’s paper is based on computer simulations. More research will be needed, but it already points to two useful things to keep in mind. The first is that forcing AI chatbots to only relay true information isn’t enough on its own. It’s a bit like trying to judge how dangerous a city is by reading only the crime section of the local paper. The facts are real, but without context or comparison to other cities, the picture gets distorted. The second thing, which concerns us as users, is that awareness is a valuable tool, but not sufficient on its own. Knowing how these systems work does change the way we use them, for the better. But it doesn’t eliminate the risk.
Worth keeping that in mind.
The original article in italian is published on LinkedIn.
It can be found here: https://www.linkedin.com/pulse/intuizione-da-nobel-o-delirio-assistito-sara-incao-11abf