When AI Answers differ each time
The AI Feedback Lottery
Large Language Models judge your work differently every time you ask. That is not just a quirk of the technology — it is a fundamental challenge to how we create, validate, and trust ideas.
Imagine you have spent weeks on a business plan. You show it to an AI assistant, and it comes back glowing: "This is a compelling, well-structured concept - you are ready to move forward." Energized, you close the chat window and get to work. A few days later, curiosity gets the better of you. You open a fresh conversation, paste in the exact same plan with the exact same prompt, and wait. The verdict this time? "There are significant structural weaknesses here. The market analysis is underdeveloped, the financial projections rest on untested assumptions, and the value proposition needs a substantial rethink."
Same document. Same question. Two completely different realities.
This is not a bug. It is one of the most underappreciated structural truths about working with Large Language Models - and it carries real consequences for anyone using AI as a creative partner, a sounding board, or a validator of their ideas.
LLMs can't give you a consistent opinion
About this post
This post explores a structural characteristic of Large Language Models and its practical implications for creators, entrepreneurs, and anyone using AI as a thinking partner.
Key takeaways
If you want to use AI for feedback, do it with caution and don't take everything it says for granted. And do multiple runs as if you would ask multiple people.
The original ideas, structure, and much of the language are human-created, but AI was used to develop, enrich, or rework portions of the content — for example, researching sources, rewriting sections for clarity, or expanding on arguments - please click for more information
A Large Language Model does not hold stable opinions the way a human collaborator does. Each response is generated anew from probabilities shaped by the current conversation context. Clear the session, and the model approaches your work from scratch again - which means the same prompt can produce materially different judgments across sessions, especially for subjective or creative work.
Session One
"This is an excellent foundation. The core idea is strong, the structure is sound, and the execution is ready for prime time. I would move forward with confidence."
Session Two
"This has potential, but the concept has a number of flaws that would need to be addressed. I would recommend a substantial rework before proceeding."
Both responses come from the same model. Both are coherent, well-reasoned, and written with conviction. Neither is lying. But they cannot both be right - and yet the model will defend either one with equal fluency if you ask it to. Technically, the model is not retrieving a fixed evaluation stored somewhere internally. Even with the identical prompt, each generation involves probabilistic token sampling, where multiple plausible continuations compete, and small statistical differences cascade into different overall responses. In creative domains - where there is no single correct answer - those branching probabilities can lead the model toward enthusiasm in one session and scepticism in another.
J.K. Rowling and the luck of the draw
When J.K. Rowling finished the first Harry Potter manuscript, she received twelve verdicts from professional editors at different publishers, and every single one of them was a rejection. The story that would become one of the most successful children's book series of the past decades was, by the consensus of expert opinion at the time, not worth publishing.
The editors were making genuine, experienced judgments - but those judgments were shaped by the mood of the reader, the publisher's current preferences, the editor's personal sensibility that day, and a thousand other unknown variables. Creative evaluation has always been, at its core, a probabilistic process. AI has simply made that probabilistic chaos instantaneous, frictionless, and disconnected from reality.
What changes with AI is not inconsistency itself, but our tendency to mistake the confident, well-structured prose of an AI response for a stable, reliable, and authoritative verdict. The deeper issue is that all creative judgment is inherently inconsistent: human taste shifts, markets fluctuate, editorial fashions change. Yet AI’s extraordinary fluency conceals this instability. A response generated by a Large Language Model carries the tone of a considered opinion because it has been trained on text written by humans expressing considered opinions. The surface markers of certainty are all present. The underlying stability is not.
Rowling’s rejection letters, at least, arrived with the implicit understanding that they reflected human judgments - fallible, contextual, one among many. An AI response, by contrast, arrives wrapped in the aesthetic of authority, when in reality it is nothing more than a stochastically generated synthesis of past human opinion as captured on the internet.
After those rejections, J.K. Rowling sent out her manuscript another thirteenth time, and it opened the doors to the successful launch of Harry Potter and the Philosopher's Stone in 1997.
Three ways to deal with LLM feedback
When creators come across an inconsistency in AI feedback, there are three ways to respond, and only one leads to an efficient evaluation process. It requires the skill to know what the tool is - and what it is not.
The Deflated Creator
A self-critical person may respond to this kind of AI feedback discrepancy with discouragement. If the AI loved your work this morning and dismissed it this afternoon, what does that say about the work? For many people, especially those already prone to self-doubt, a single harsh AI verdict can be enough to shelve an idea entirely. The cruelty is that the harsh verdict may simply have been a statistical outlier: the equivalent of one editor at one publisher having a bad day. But without the context that it was only one draw from a probabilistic deck, it reads as definitive judgment while also devaluing the earlier praise.
The Overconfident Creator
The second response is the mirror image: unfounded confidence. A confident creator shops sessions until they get the enthusiastic verdict they were hoping for, takes a screenshot, and proceeds as though the AI has validated their work. This is a perfectly natural psychological response to an uncertain situation. But it is dangerous, because the encouragement carries no more weight than the discouragement. It is simply a lucky draw, not an informed assessment. The creator moves forward with conviction built on sand.
The Sophisticated Creator
The third response is the only genuinely useful one: treating the inconsistency as information in itself. A creator who asks the same AI the same question multiple times - or who deliberately varies the framing to stress-test an idea - is not gaming the system. They are doing something close to what a good editor or creative director does when they seek multiple opinions before committing to a direction. The variance in the responses tells them something real: where the idea is genuinely strong (consistent praise across sessions), and where it is genuinely fragile (inconsistent or conflicting feedback).
How to use AI as a creative partner?
None of this means AI feedback is worthless. A single AI session can surface blind spots, identify unclear passages, suggest alternatives you had not considered, and push you toward a more refined version of your idea.
A good way to go about this is to use the statistical distribution to your advantage. If you ask the LLM the same question three times with the same prompt, you will likely get a solid distribution of praise and criticism. And if you want, you can additionally outsource the task to find the similarities and differences in these evaluations to your LLM.
One challenge in this context is “persistent memory”, where the Large Language Model stores the content of previous conversations. If this is included in the second and third evaluations, it dilutes the effect of giving you a fresh perspective and will continue to justify previous statements. Turn it off!
One caveat: The more novel and outside-of-the-box your idea is, the less a LLM will be able to provide you with content feedback. It is built on past knowledge and patterns, and not well equipped to detect valuable novelty that is beyond extending what already exists. Thus, you can still spell-check and check for completeness, but don’t expect a solid evaluation of the substantive content.