Skip to content

Why AI often gets it wrong

  • 6 min read
  • GenAI

Have you ever taken an AI-generated answer at face value, only to later question whether it was actually right? In this episode, we build on something we mentioned in a recent ‘Using AI in extension’ workshop about why tools like ChatGPT and Copilot can sound convincing even when they’re wrong, and what that means for us as enablers of change.

In that online workshop, many participants shared how quickly generative AI has become part of their everyday work, whether drafting communications, summarising research, or preparing for farmer discussions. Tools like ChatGPT and Copilot are helpful because they provide us with a generally well-structured response in seconds, while giving the impression of expertise. However, that fluency can mask a key limitation. That is, these systems don’t know information in the way we do, and they’re not designed for accuracy.

Gen AI systems work by predicting the next word in a sequence, based on patterns learned from vast amounts of data. Each response is built by selecting what is statistically most likely to come next given the prompt and the words already generated. This means the system is optimised for producing plausible language, not for facts, which helps explain why outputs can appear coherent and confident even when they contain errors.

Research by Kalai et al. (2025) showed that these hallucinations are not random glitches in the matrix but a natural outcome of how language models are trained and evaluated. Like students facing an exam question, these systems tend to guess when they are unsure, because producing an answer is rewarded more than admitting uncertainty. In other words, the system has been shaped to respond, not to pause and check.

This matters for us working in advisory and extension roles, where trust and credibility are central to our work. We’re increasingly using tools like ChatGPT and Copilot to support decision-making, communication, and learning, often under time pressure. But if these tools are designed to prioritise quick answers over uncertainty, we need to be more deliberate in how we interpret and use their outputs.

Let’s bring this into a few everyday scenarios. Imagine a farm advisor preparing a session on nutrient management and using AI to generate key talking points or recommendations. The response is clear, structured, and confident, which makes it easy to reuse, but as the research mentioned highlights, generating a plausible answer is actually easier for the model than determining whether that answer is factual. This creates a risk of passing on information that sounds credible but isn’t. And the AI won’t be fact checking!

Now consider a practitioner summarising farmer feedback from a series of workshops using AI to speed up the process. The tool can identify patterns and themes efficiently, which is valuable, but it may also smooth over nuance or introduce interpretations that were not present in the data. These are not obvious mistakes but distortions that arise from how models generalise patterns.

Finally, think about exploring a new topic, such as biodiversity markets or carbon frameworks, where our own expertise is still developing. AI tools can act as a useful starting point, helping us summarise the information available so we can orient ourselves quickly. But if the knowledge is sparse or uneven, the model is more likely to generate confident guesses. This reflects what researchers call epistemic uncertainty, where the system simply does not have enough reliable information but still produces an answer.

Another practical consideration we touched on in the workshop is choosing the right AI tool for the task. Choosing between tools is less about which tool is better, and more about understanding the trade-off between control and flexibility in our specific context. For example, when working with a set of documents such as reports, project data, or workshop notes, tools like NotebookLM can be useful because they only use the material we provide, rather than drawing on information from other sources such as the web. This can reduce the risk of hallucination because it’s relying on our data and makes it easier to track where the information came from. 

But a more practical step is to change how we prompt and engage with AI. Instead of asking for the answer, we can ask for multiple perspectives, limitations, or areas of uncertainty, which encourages more useful outputs. We can also build simple verification habits into our use of AI tools, such as cross-checking key claims, reviewing cited sources, and comparing outputs across tools or prompts. These actions don’t need to be time-consuming, but they do require us to be more than just passive users of AI.

It’s worth acknowledging that these research results are already shaping how tools like ChatGPT and Copilot are being developed, particularly in how they handle uncertainty and avoid overconfident errors. However, these improvements are likely to emerge gradually rather than as a single fix, and hallucinations will not disappear entirely because they stem from how these systems work. 

So we think this means we’re not moving towards AI that never makes mistakes, but towards AI that is better at signalling when it might be wrong. For those of us working in extension and advisory roles, this doesn’t mean waiting for perfect tools, but instead focusing on building the skills and habits that allow us to use these tools confidently in our work. Maybe the question isn’t about why AI gets things wrong, but how we respond when it does. 

We’ve shared our thoughts, now it’s your turn! Drop a comment below with your experiences, tips, and ideas about these AI hallucinations and how you check and minimise them. We’d love this to be a conversation—your insights are invaluable to our community. 

Thanks for reading this Enablers of Change episode! Don’t forget to subscribe to our newsletter to catch future episodes, and share this post with friends who might enjoy the discussion. 

Until next time, all the best and happy prompting!

Resources

Kalai, A. T., Nachum, O., Vempala, S. S., & Zhang, E. (2025). Why language models hallucinate. OpenAI & Georgia Tech. Available online.

Ji, Z., et al. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12).

5 2 votes
Article rating
Subscribe
Notify of
guest
0 Comments
oldest
newest most voted
0
We would love to hear your thoughts, so please leave a comment!x
()
x