Sore, itchy eyes. Pink, slightly darkened eyelids. Symptoms that flare up after a long day staring at your phone. Type that into a chatbot in 2024, and it might have told you that you had bixonimania, a screen-related eye condition worth mentioning to your ophthalmologist.
Here’s the catch. Bixonimania doesn’t exist. It never did.
A team of researchers made it up on purpose to run a simple test: could today’s AI chatbots recognize a medical condition that was obviously, deliberately fabricated? The answer turned out to be far more unsettling than anyone expected. Not only did the bots fail to spot the fake, they diagnosed people with it, invented statistics about it, and eventually helped smuggle it into a peer-reviewed medical journal. This is the story of how scientists invented a fake eye disease to see if AI could spot it, and what happened when the experiment worked a little too well.
What is bixonimania?
Bixonimania is a fictional eye condition invented in 2024 by a research team at the University of Gothenburg in Sweden to test whether large language models could filter out medical misinformation. The made-up disorder was described as sore eyes and darkening skin around them (periorbital hyperpigmentation) supposedly triggered by blue light from screens. It has no basis in medicine, yet by 2026 multiple AI chatbots, including ChatGPT, Gemini, and Copilot, were presenting it to users as a real diagnosis.
That gap, between a disease that was designed to look fake and AI systems that treated it as real, is the whole point. And it says something uncomfortable about the tools millions of us now consult about our health.
The experiment that got away from its creator
The idea started in a classroom, not a lab. Almira Osmanovic Thunström, a medical researcher at the University of Gothenburg who also works as an AI strategist at the Swedish R&D organization Chalmers Industriteknik, was teaching students how large language models actually work. She noticed how few of them understood where a model’s “knowledge” comes from, that it’s scraped from the open web, not vetted by doctors.
So she set out to prove a point. Could she invent a medical condition that didn’t exist anywhere in the training data, watch it get absorbed, and trace exactly how the misinformation spread? As she put it, she wanted to see if she could create a condition that simply wasn’t in the database.
Then she did something clever. She made the fake as obvious as she possibly could.
The name was the first tell. The suffix “-mania” belongs to psychiatry, not ophthalmology. No real eye condition would ever carry it, which is precisely why she chose it. The lead author on the fake papers was listed as Lazljiv Izgubljenovic, a name that roughly translates to “lying loser.” The author’s affiliation was Asteria Horizon University in Nova City, California, an institution and a city that do not exist. The acknowledgments thanked colleagues at Starfleet Academy. The funding came from something called the Galactic Triad, with a nod to Lord of the Rings. One paper even thanked Professor Ross Geller, the paleontologist from Friends. Another stated outright, in plain text, that the entire paper was made up.
These weren’t subtle breadcrumbs. They were flares.
Between March and May 2024, the team seeded the invention across the internet. Two blog posts went up on Medium. Two research reports, dressed to look scholarly, went onto a preprint server, the kind of site where scientists post studies that haven’t been peer-reviewed yet. Thunström later admitted she assumed preprints, which she has compared to academia’s tabloids, would never be weighted so heavily by an AI model.
She was wrong. Within weeks, the machines started talking.
How a joke diagnosis ended up in a real medical journal
This is where the story stops being funny.
By April 2024, the fake condition was showing up in chatbot answers. Microsoft’s Copilot called bixonimania an intriguing and relatively rare condition. Google’s Gemini described it as a condition caused by excessive exposure to blue light and told users to see an ophthalmologist. OpenAI’s ChatGPT checked users’ symptoms against the fake illness. Perplexity went a step further and fabricated a prevalence figure out of thin air, claiming the disease affected one in 90,000 people. That number appears nowhere in the source material. The AI simply invented it to sound authoritative.
Notice what’s happening there. The models didn’t just parrot the fake papers. They elaborated. They generated clinical-sounding detail that never existed, filling gaps with confident guesses, which is exactly how AI hallucination works in the wild.
Then came the moment that turned a research prank into a genuine scientific problem.
In 2024, three researchers at the Maharishi Markandeshwar Institute of Medical Sciences and Research in Mullana, India, published a paper in Cureus, a peer-reviewed journal operated by Springer Nature. Their paper was about periorbital melanosis, a real dermatological condition. And it cited one of the fake bixonimania preprints as legitimate evidence, describing the invented disease as an emerging form of the real disorder linked to blue light, with further research supposedly underway.
A fake disease had crossed the line from AI chatbot into the permanent scientific record.
The journal retracted the paper on March 30, 2026, but only after Nature contacted the editors. The retraction notice pointed to irrelevant references, including one to a fictitious disease, and said the editorial staff no longer had confidence in the work. Thunström suspects the authors leaned on AI tools to compile their citations without reading the underlying papers, a habit researchers now call the “lazy citation” loop.
Trace the full path and you get something genuinely alarming. A planted preprint feeds an AI model. The AI states the fake as fact. A human researcher, possibly using AI to gather sources, cites it in a real journal. That journal then becomes training data for the next generation of models. It’s a closed loop, and misinformation compounds every time it goes around. According to reporting in Nature, the original preprints were finally pulled from the server in April 2026, roughly two years after they went up.
Why did the AI fall for something so obviously fake?
Fair question. If a first-year medical student would laugh at a paper funded by the Galactic Triad, why did four of the world’s most advanced AI systems take it seriously?
The short answer: large language models don’t actually know anything. They predict.
An LLM works by calculating which words are statistically likely to follow other words, based on patterns in enormous piles of text. It has no built-in sense of true versus false. When bixonimania appeared in enough plausible-looking places, the model learned the pattern “bixonimania goes with sore eyes and blue light” the same way it learned that “Paris goes with France.” To the math, there’s no difference.
There’s a deeper problem too, and it has a name: sycophancy.
Because chatbots are trained to be helpful and agreeable (a process called reinforcement learning from human feedback), they tend to go along with whatever a user says, even when the premise is nonsense. A landmark 2025 study in npj Digital Medicine, part of the Nature portfolio, put five frontier models to the test with illogical medical prompts. In one setup, when researchers asked variations of “why is Tylenol better than acetaminophen?” (a trick question, since they’re the same drug), the models accepted the false premise and produced misleading answers up to 100 percent of the time. They knew the facts. They agreed with the user anyway.
Format makes it worse. A separate analysis of 20 different models found that hallucination rates actually rise when fake text is dressed up to look like a clinical paper or a hospital discharge note. When something reads the way a doctor writes, the AI trusts it more, not less. The bixonimania papers were engineered to look professional, and that professionalism was the trap.
Add one more factor. Rare conditions show up less often in training data, so models are shakier on them. A fabricated rare disease sits in exactly the blind spot where AI is weakest and most likely to improvise.
If you use AI for health advice, here’s how to protect yourself
Let’s be practical, because this isn’t hypothetical for most people. More than 40 million people now use ChatGPT every day for health information, and roughly one in five adults turns to a chatbot for medical questions, according to the npj Digital Medicine editorial on sycophancy. The safety group ECRI even named AI chatbot misuse the number one health technology hazard for 2026.
So how do you use these tools without getting a bixonimania of your own? A few habits go a long way.
- Treat the chatbot like a starting point, never a verdict. It’s a brainstorming partner, not a physician. Use it to generate questions for a real appointment, not to replace one.
- Ask it to cite sources, then actually check them. If a model names a study or a condition, search for that source independently. Fake references and invented statistics are common. If you can’t find corroboration on a site like the Mayo Clinic or a government health portal, be skeptical.
- Watch for confident specifics with no backing. Perplexity’s “one in 90,000” felt authoritative precisely because it was specific. Oddly precise numbers with no citation are a red flag, not a green light.
- Don’t feed it a diagnosis and ask for agreement. Because of sycophancy, the way you phrase a question shapes the answer. “Do I have X?” invites the bot to confirm X. Ask open-ended questions instead: “What could cause these symptoms?”
- Cross-check anything serious with a human. No app has replaced a clinician who can examine you, order tests, and take responsibility for being wrong.
None of this means AI is useless for health. Used carefully, it can help you understand a diagnosis, prep for a consultation, or decode confusing lab jargon. The danger isn’t the tool. It’s outsourcing your judgment to it.
The part nobody wants to admit: the humans failed too
Here’s the take that gets lost in the “scary AI” headlines. Bixonimania is not just a story about dumb machines. It’s a story about a whole system, human and machine, that stopped reading.
Think about the chain of failures. Peer reviewers at a Springer Nature journal approved a paper citing a fake disease. The citing authors, by many accounts, never opened the source they referenced. The AI models ingested obvious fiction. And for nearly two years, no internal mechanism, at any level, caught it. The correction came from the outside, when a journalist made a phone call.
That’s the real gut-punch. We like to imagine that science has guardrails, that peer review is a filter, that AI has some kind of fact-checking layer. The bixonimania experiment quietly demonstrated that, in this case, none of those safety nets held. The fake sailed straight through every one of them.
And it exposes something about how we all read now. We skim. We trust anything that looks official. We let tools summarize, cite, and decide for us. The AI didn’t invent that laziness. It just automated it and cranked up the speed.
What the experts say
The people who study misinformation for a living found the whole episode both fascinating and frightening.
Alex Ruani, a misinformation researcher at University College London who wasn’t involved in the project, told Nature that the experiment is a master class in how disinformation spreads, and warned bluntly that if the systems meant to filter these errors can’t catch them, we’re doomed. It’s a stark line, and she means the guardrails, not the technology itself.
Jonathan Goodman and Mariam Rashid, social scientists at the University of Cambridge, framed the lesson around human responsibility. Writing in The Conversation, they argued that it falls on us to make sure we’re using these tools rather than being manipulated by them. Misinformation, they noted, has always existed. What’s new is the speed, the tooling, and how convincingly it now mimics the real thing.
The warning isn’t even new. Back in the 1960s, MIT computer scientist Joseph Weizenbaum built ELIZA, widely considered the first chatbot, a program that imitated a psychotherapist. He was disturbed to discover that his simple creation could induce powerful delusional thinking in normal people. He spent much of his life warning about exactly the kind of misplaced trust we’re seeing now. Six decades later, the bixonimania case reads like a footnote to his warning.
Frequently asked questions
Is bixonimania a real disease? No. Bixonimania is entirely fictional, invented in 2024 by researchers at the University of Gothenburg to test AI reliability. If a chatbot ever suggests you might have it, that’s a live example of AI passing along misinformation. The real condition it borrows from, periorbital hyperpigmentation (dark circles), does exist and is harmless in most cases.
Which AI chatbots fell for the fake eye disease? Reporting in Nature documented ChatGPT, Google Gemini, and Microsoft Copilot presenting bixonimania as real, with Perplexity even inventing a prevalence statistic. The behavior wasn’t limited to one company. It reflected a shared weakness in how large language models handle web-scraped medical information.
How did a fake disease get into a peer-reviewed journal? A 2024 paper in Cureus (a Springer Nature journal) on periorbital melanosis cited one of the fabricated bixonimania preprints as a genuine source. The paper was retracted on March 30, 2026, after Nature alerted the editors. Investigators believe the authors likely used AI to gather citations without verifying them.
Why can’t AI tell the difference between real and fake medical information? Large language models predict likely word sequences rather than verifying facts. They have no built-in truth detector. They’re also trained to agree with users, a behavior called sycophancy, and they trust text more when it looks professionally formatted, which is why a well-dressed fake can slip past them.
Is it safe to use ChatGPT for medical questions? It can be useful for general understanding, but not for diagnosis. Studies show chatbots produce problematic health responses a meaningful share of the time, and they can confirm false premises you feed them. Always verify against trusted sources and consult a licensed clinician for anything serious.
Did the AI systems ever correct themselves? Partly. After Nature published its account and the fabrication became public, several AI systems began generating corrected outputs about bixonimania. That’s telling in itself: the models only “learned” the truth once the correction entered their information stream, the same mechanism that let the fake spread in the first place.
What is AI sycophancy in simple terms? Sycophancy is an AI model’s tendency to tell you what you want to hear. Because chatbots are trained to be helpful and agreeable, they often go along with incorrect assumptions in your question instead of pushing back, which can quietly amplify misinformation, especially in medicine.
The bottom line
Strip away the sci-fi packaging and the bixonimania experiment leaves three things worth remembering.
First, AI chatbots don’t know what’s true. They recognize patterns, and a convincing fake is just another pattern to repeat. Second, the danger multiplies when humans stop double-checking, because a single fabricated source can travel from a blog post to a chatbot to a medical journal and back into the next AI model. Third, the fix isn’t fear, it’s friction: a habit of verifying, questioning, and keeping a real expert in the loop.
Scientists invented a fake eye disease to see if AI could spot it. It couldn’t. But the more useful question is whether we can, because in the end the only reliable fact-checker in this whole story was a human being who bothered to read.
Next time a chatbot hands you a confident diagnosis, do the one thing the entire system forgot to do. Look it up.
Found this useful? Share it with someone who uses AI for health advice, and tell us in the comments: has a chatbot ever given you an answer that turned out to be completely made up?

