Student Stories

I Asked Three AIs the Same Question. I Got Three Answers.

One student's experiment comparing ChatGPT, Claude, and Gemini, and what the differences taught them about not trusting the first reply.

One student's experiment comparing ChatGPT, Claude, and Gemini, and what the differences taught them about not trusting the first reply.

No headings found

I'd always treated AI answers as roughly interchangeable, ask one, get the truth. Then I asked ChatGPT, Claude, and Gemini the exact same question in the exact same wording and got three noticeably different answers. That broke something for me, in a good way.

THE EXPERIMENT

For a class project I needed to fact-check a historical claim I wasn't fully sure about. Instead of asking one AI and moving on, I decided to ask all three tools I had access to, copy-pasting the identical question into each, and compare the responses side by side before writing anything down.

THE QUESTION

The question itself was straightforward enough to seem like it should have one clean answer. It wasn't a trick question or an edge case, just a specific factual claim with a date attached to it.

THREE DIFFERENT ANSWERS

One tool answered confidently and was mostly right but got a date wrong by a couple of years. Another hedged more, gave a broader range, and turned out to be the most accurate. The third gave an answer that sounded just as confident as the first but was wrong in a different way, it had conflated two similar events into one.

None of them said "I'm not sure." All three sounded equally certain. That was the unsettling part.

WHY THEY DISAGREED

Each model was trained on different data, at different times, using different processes, so they don't all "know" the same things even when a fact is genuinely fixed and unambiguous. None of them are looking anything up live unless the tool explicitly tells you it searched the web. Without that, you're getting a prediction of what a correct-sounding answer looks like, not a guaranteed lookup of a fact.

WHAT I LEARNED ABOUT TRUSTING THE FIRST REPLY

The scary part wasn't that they disagreed, disagreement is normal and even useful. It's that I almost didn't check. If I'd only asked one tool, I would have submitted a wrong date with total confidence, because the wrong answer read exactly as convincingly as the right one. There was no visual cue, no warning sign, nothing to make me pause.

Since then my rule is simple: for anything that actually matters, a name, a date, a statistic, I check it against at least one other source, AI or otherwise, before it goes in my work. The tools are useful. Just not more trustworthy than the effort I put into verifying them.

Ready to think like a reasoning engineer?

2Skills helps young people learn how to identify meaningful problems, choose the right AI tools and prompt them clearly to build real solutions.

Task list of learning how to use AI tools

Ready to think like a reasoning engineer?

2Skills helps young people learn how to identify meaningful problems, choose the right AI tools and prompt them clearly to build real solutions.

Task list of learning how to use AI tools

Ready to think like a reasoning engineer?

2Skills helps young people learn how to identify meaningful problems, choose the right AI tools and prompt them clearly to build real solutions.

Task list of learning how to use AI tools

Ready to think like a reasoning engineer?

2Skills helps young people learn how to identify meaningful problems, choose the right AI tools and prompt them clearly to build real solutions.