Listen to this post

0:00
When AI answers differ each time

The AI Feedback Lottery

Large Language Models judge your work differently every time you ask. That is not just a quirk of the technology - it is a fundamental challenge to how we create, validate, and trust ideas.

Imagine you have spent weeks on a business plan. You show it to an AI assistant, and it comes back glowing: “This is a compelling, well-structured concept – you are ready to move forward.” Energized, you close the chat window and get to work. A few days later, curiosity gets the better of you. You open a fresh conversation, paste in the exact same plan with the exact same prompt, and wait. The verdict this time? “There are significant structural weaknesses here. The market analysis is underdeveloped, the financial projections rest on untested assumptions, and the value proposition needs a substantial rethink.”

Same document. Same question. Two completely different realities.

This is not a bug. It is one of the most underappreciated structural truths about working with Large Language Models – and it carries real consequences for anyone using AI as a creative partner, a sounding board, or a validator of their ideas.

LLMs can’t give you a consistent opinion

A Large Language Model does not hold stable opinions the way a human collaborator does. Each response is generated anew from probabilities shaped by the current conversation context. Clear the session, and the model approaches your work from scratch again – which means the same prompt can produce materially different judgments across sessions, especially for subjective or creative work.

Both responses come from the same model. Both are coherent, well-reasoned, and written with conviction. Neither is lying. But they cannot both be right – and yet the model will defend either one with equal fluency if you ask it to. Technically, the model is not retrieving a fixed evaluation stored somewhere internally. Even with the identical prompt, each generation involves probabilistic token sampling, where multiple plausible continuations compete, and small statistical differences cascade into different overall responses. In creative domains – where there is no single correct answer – those branching probabilities can lead the model toward enthusiasm in one session and skepticism in another.

J.K. Rowling and the luck of the draw

When J.K. Rowling finished the first Harry Potter manuscript, she received twelve verdicts from professional editors at different publishers, and every single one of them was a rejection. The story that would become one of the most successful children’s book series of the past decades was, by the consensus of expert opinion at the time, not worth publishing.

After those rejections, J.K. Rowling sent out her manuscript another thirteenth time, and it opened the doors to the successful launch of Harry Potter and the Philosopher’s Stone in 1997.

The editors were making genuine, experienced judgments – but those judgments were shaped by the mood of the reader, the publisher’s current preferences, the editor’s personal sensibility that day, and a thousand other unknown variables. Creative evaluation has always been, at its core, a probabilistic process. AI has simply made that probabilistic chaos instantaneous, frictionless, and disconnected from reality.

Three ways to deal with LLM feedback

When creators come across inconsistency in AI feedback, there are three ways to respond, and only one leads to a meaningful contribution to the evaluation process. It requires knowledge about what the tool is – and what it isn’t.

Rowling had no choice but to treat twelve verdicts as twelve draws from the same deck. With AI, you can run the draws yourself – the question is what you do with them.

How to use AI as a creative partner?

None of this means AI feedback is worthless. A single session can surface blind spots, identify unclear passages, suggest alternatives you had not considered, and push you toward a more refined version of your idea.

The real leverage, though, comes from using the statistical distribution to your advantage. Ask the model the same question three times in three fresh sessions, and you get something no single answer can give you: a distribution. Where all three runs agree – in praise or in criticism – you are probably looking at a genuine property of your work. Where they contradict each other, you have found the places where reasonable judges could disagree, which is exactly where your own judgment is needed. If you want, you can hand the comparison work back to the model and let it map the agreements and contradictions for you.

The lottery doesn’t stop being a lottery. But once you know that’s what it is, you stop mistaking a single ticket for a verdict – and start reading the odds. Your business plan deserves at least three draws.