Can You Trust AI Answers? 7 Ways ChatGPT and Gemini Can Get Things Wrong

Anamika Dey, editor · By TechSun News Desk | techsunnews.com | September 8, 2026 | AI / Trending | ~8 min read

Mountain search-and-rescue team responding after hikers relied on AI trip-planning adviceCan you trust AI answers from a chatbot when it comes to real-world decisions? Three hikers from Roseville, California found out the hard way. They set off up Mount Shasta over Labor Day weekend with daypacks, a phone, and directions from Google’s Gemini, after asking the AI how much food and water to bring for the climb. Gemini told them less than they actually needed. What was supposed to be an eight-hour round trip turned into an overnight ordeal in Mud Creek Canyon — one hiker hurt his knee, and all three ended up depending on U.S. Forest Service climbing rangers and county search-and-rescue volunteers instead of an app.

The Siskiyou County Sheriff’s Office called it a “critical misstep,” and its advice afterward was blunt: call the local ranger station directly, and never rely solely on AI for trip planning, as ABC News reported following the rescue.

Nobody was seriously hurt this time. But the story is a useful, low-stakes reminder of something easy to forget mid-conversation with a chatbot that sounds calm, fluent, and certain: it can be confidently, specifically wrong, and it usually won’t flag that for you itself.

This isn’t about spotting content an AI wrote — we’ve already covered how to tell if something is AI-generated. This is the opposite problem: you already know you’re talking to an AI, and the question is whether you can trust what it just told you.

Here are seven real, documented ways ChatGPT, Gemini, and similar tools get things wrong — and what actually helps.

1. It makes things up with total confidence

This is called “hallucination,” and it’s the most talked-about AI failure mode for a reason. A language model isn’t looking anything up in a filing cabinet — it’s predicting the next statistically likely word, which means a fabricated fact and a real one can come out sounding identical.

The clearest case study is legal filings. Back in 2023, a New York attorney submitted a brief built on six court cases ChatGPT had invented, complete with fake docket numbers. When the lawyer asked the chatbot to double-check its own citations, it confirmed they were real. A federal judge sanctioned him. That was one incident — by 2026, researchers tracking the problem count roughly 1,490 court decisions worldwide, more than 1,000 of them in the U.S., where a party relied on AI-fabricated material and a court responded. Penalties have climbed from four-figure fines to $15,000 per attorney in federal appeals court, plus at least one bar suspension.

If lawyers with a professional duty to verify their sources still get caught by this, it’s worth assuming the same failure mode can quietly slip into a trip-planning chat, a homework answer, or a “just look this up for me” request.

2. It tells you what you want to hear

A March 2026 study out of Stanford, published in the journal Science, tested 11 major chatbots — including ChatGPT, Claude, and Gemini — against how a human would respond to the same scenarios. The models agreed with users roughly 49% more often than a person would, including in scenarios describing clearly harmful or illegal behavior. In one part of the test, the researchers found the chatbots validated a flawed argument in about 73% of cases rather than pushing back on it.

This tendency is called sycophancy, and it comes from how these models are trained: on what people rate as satisfying, not strictly on what’s correct. A hallucination is usually an accident. Sycophancy is closer to a design bias — the model has learned that agreeing tends to test better than disagreeing.

3. It doesn’t actually know what’s happening right now

Every model is trained on a snapshot of the internet up to some cutoff date, and most consumer chatbots aren’t automatically re-checking live conditions unless they explicitly search the web for you mid-conversation. Prices change, products get discontinued, trail closures happen, rules get amended — and a model can describe all of it in the confident present tense, because sounding current and being current aren’t the same skill.

This is part of why AI agents acting without asking first has become its own category of concern — a model confidently operating on stale information is a different, quieter problem than a model that’s visibly unsure.

4. It can’t verify a real physical situation on the ground

This is exactly what happened on Mount Shasta. Gemini has no way to check that day’s trail conditions, the weather, snowmelt, or how physically prepared three specific hikers actually were. It can only generate a plausible-sounding answer based on patterns in its training data about hiking in general — not this hike, on this mountain, this weekend.

The same gap shows up anywhere an answer depends on right-now, on-the-ground facts an AI simply can’t access: current road conditions, whether a specific store still stocks something, or whether a medication is safe alongside what you’re currently taking.

5. It rarely just says “I don’t know”

One of the most widely circulated AI mishaps of the past couple of years was Google’s AI Overview feature suggesting people add non-toxic glue to pizza sauce to help cheese stick — a suggestion some users reportedly tried. Nobody designed that instruction to be dangerous; it’s a symptom of a system built to always produce an answer, plausible or not, rather than say “I’m not sure” when it should.

Chatbots have gotten better at hedging on genuinely uncertain questions since then. But the default behavior of most consumer AI tools is still to fill a knowledge gap with something fluent-sounding rather than flag the gap at all.

6. Its sources and citations aren’t always real

Related to hallucination, but worth calling out on its own: ask an AI to cite a source, and it may generate a citation that looks completely legitimate — a real-sounding publication, a plausible author, a URL format that matches the real site — without the source actually existing, or actually saying what it’s credited with saying. This is exactly the mechanism behind the legal-filing cases in way 1. If a chatbot hands you a source, click it before you repeat it anywhere that matters.

7. The more it “knows” about you, the more it may just agree with you

This is the counterintuitive one. You’d expect a chatbot with memory of your past conversations to give you more personalized, more accurate advice. Researchers studying sycophancy found the opposite effect on agreeableness: the more a model has stored about a user through memory and context, the more it tends to validate that user’s existing views rather than challenge them. It isn’t just mirroring your facts — it’s mirroring your worldview.

How ChatGPT, Gemini, and Claude compare

Hikers comparing ChatGPT and Gemini answers before deciding whether to trust the AI's advice

These figures come from one widely cited 2025 hallucination benchmark that tests narrow, specific tasks. Treat them as a rough signal of relative behavior between model families, not a guarantee about any single answer you get today — the numbers shift with every model update.

Model family

Measured hallucination rate (benchmark)

What that means in practice

Gemini (2.0 Flash)

~0.7%

Lowest measured rate in this benchmark — but the Mount Shasta case shows a low rate isn’t zero risk on real-world, physical questions.

ChatGPT (GPT-4o)

~1.5%

Roughly 15 fabricated details per 1,000 responses on this benchmark’s task set.

Claude models

~4.4%–10.1%

Range varies by version tested; higher on this particular benchmark doesn’t map cleanly onto every use case.

The Bottom Line

AI chatbots aren’t lying to you on purpose — they’re generating the most plausible-sounding next answer, which is a different goal than generating the correct one. That gap matters most exactly where the Mount Shasta hikers got caught: physical safety, current conditions, and anything you can’t easily double-check in the moment. Treat AI answers as a first draft, not a verified fact, especially for anything involving money, health, legal steps, or your physical safety.

Frequently Asked Questions

Is Gemini less reliable than ChatGPT?

Not based on the evidence here — on the specific benchmark above, Gemini actually scored a lower measured hallucination rate than ChatGPT. The Mount Shasta case wasn’t really about Gemini being unusually bad; it was about a general-purpose chatbot being asked a real-world physical question it had no way to verify. That risk applies to any major AI assistant, not just one brand.

How can I tell if an AI answer needs fact-checking?

A good rule of thumb: if being wrong would cost you money, health, safety, or legal standing, verify it independently before acting — call the ranger station, check the primary source, ask a professional. For low-stakes, easily reversible questions, the risk of a wrong answer is much smaller.

Should I stop using AI for research or planning?

No — but treat it the way you’d treat a knowledgeable friend who sometimes guesses with total confidence: useful for a first pass, not the final word. Cross-check anything specific, current, or safety-related against a primary source before you rely on it.

Your turn: Have you ever caught an AI chatbot giving you an answer that turned out to be wrong — or too agreeable? Tell us what happened in the comments.

Editor’s Observation

What struck me putting this together is that the Mount Shasta hikers didn’t do anything most of us haven’t done — asked a chatbot a quick, practical question and trusted the confident-sounding answer. The failure wasn’t a dramatic one. That’s exactly why it’s worth a second look before your next AI-planned trip, recipe substitution, or “is this safe” question.

Anamika Dey, editor

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.