{"id":227644,"date":"2026-04-27T14:33:00","date_gmt":"2026-04-27T14:33:00","guid":{"rendered":"https:\/\/www.9senses.ai\/ns-lab-the-ai-feedback-lottery\/"},"modified":"2026-08-27T08:05:37","modified_gmt":"2026-08-27T08:05:37","slug":"the-ai-feedback-lottery","status":"publish","type":"post","link":"https:\/\/www.9senses.ai\/fr\/the-ai-feedback-lottery\/","title":{"rendered":"The AI Feedback Lottery"},"content":{"rendered":"<p>Imagine you have spent weeks on a business plan. You show it to an AI assistant, and it comes back glowing: &#8220;This is a compelling, well-structured concept &#8211; you are ready to move forward.&#8221; Energized, you close the chat window and get to work. A few days later, curiosity gets the better of you. You open a fresh conversation, paste in the exact same plan with the exact same prompt, and wait. The verdict this time? &#8220;There are significant structural weaknesses here. The market analysis is underdeveloped, the financial projections rest on untested assumptions, and the value proposition needs a substantial rethink.&#8221;<\/p>\n<p>Same document. Same question. Two completely different realities.<\/p>\n<p>This is not a bug. It is one of the most underappreciated structural truths about working with Large Language Models &#8211; and it carries real consequences for anyone using AI as a creative partner, a sounding board, or a validator of their ideas.<\/p>\n<h2>LLMs can&#8217;t give you a consistent opinion<\/h2>\n<p>A Large Language Model does not hold stable opinions the way a human collaborator does. Each response is generated anew from probabilities shaped by the current conversation context. Clear the session, and the model approaches your work from scratch again &#8211; which means the same prompt can produce materially different judgments across sessions, especially for subjective or creative work.<\/p>\n<aside class=\"ns-obj ns-obj--aside ns-obj--full ns-obj--c2 ns-obj--split\" style=\"width:72%\"><div class=\"ns-obj-body\"><div class=\"ns-obj-col ns-obj--fmt-normal\" style=\"background:rgba(10,123,17,0.25)\"><h3>Session One<\/h3>\n<p>&#8220;This is an excellent foundation. The core idea is strong, the structure is sound, and the execution is ready for prime time. I would move forward with confidence.&#8221;<\/p><\/div><div class=\"ns-obj-col ns-obj--fmt-normal\" style=\"background:rgba(234,6,6,0.25)\"><h3>Session Two<\/h3>\n<p>&#8220;This has potential, but the concept has a number of flaws that would need to be addressed. I would recommend a substantial rework before proceeding.&#8221;<\/p><\/div><\/div><\/aside>\n<p>Both responses come from the same model. Both are coherent, well-reasoned, and written with conviction. Neither is lying. But they cannot both be right &#8211; and yet the model will defend either one with equal fluency if you ask it to. Technically, the model is not retrieving a fixed evaluation stored somewhere internally. Even with the identical prompt, each generation involves probabilistic token sampling, where multiple plausible continuations compete, and small statistical differences cascade into different overall responses. In creative domains &#8211; where there is no single correct answer &#8211; those branching probabilities can lead the model toward enthusiasm in one session and skepticism in another.<\/p>\n<h2>J.K. Rowling and the luck of the draw<\/h2>\n<aside class=\"ns-obj ns-obj--aside ns-obj--right ns-obj--c1 ns-obj--fmt-inverted\" style=\"width:54%\"><div class=\"ns-obj-body\"><h3>AI vs. human inconsistency<\/h3>\n<p>What changes with AI is not the inconsistency itself. All creative judgment is inherently inconsistent: taste shifts, fashions change, humans have good and bad days that influence their judgment. What changes is the packaging. Rowling\u2019s rejection letters arrived with the implicit understanding that they reflected single human judgments &#8211; fallible, contextual, one among many. An AI response carries the tone of a considered opinion because it was trained on millions of them &#8211; yet it arrives with <a href=\"\/fr\/a-confident-confabulator\/\">an authority it has not earned<\/a>: a stochastically generated synthesis of past human opinion as captured on the internet, unrelated to the work in front of it.<\/p><\/div><\/aside>\n<p>When J.K. Rowling finished the first Harry Potter manuscript, she received twelve verdicts from professional editors at different publishers, and every single one of them was a rejection. The story that would become one of the most successful children&#8217;s book series of the past decades was, by the consensus of expert opinion at the time, not worth publishing.<\/p>\n<p>After those rejections, J.K. Rowling sent out her manuscript another thirteenth time, and it opened the doors to the successful launch of Harry Potter and the Philosopher&#8217;s Stone in 1997.<\/p>\n<p>The editors were making genuine, experienced judgments &#8211; but those judgments were shaped by the mood of the reader, the publisher&#8217;s current preferences, the editor&#8217;s personal sensibility that day, and a thousand other unknown variables. Creative evaluation has always been, at its core, a probabilistic process. AI has simply made that probabilistic chaos instantaneous, frictionless, and disconnected from reality.<\/p>\n<h2>Three ways to deal with LLM feedback<\/h2>\n<p>When creators come across inconsistency in AI feedback, there are three ways to respond, and only one leads to a meaningful contribution to the evaluation process. It requires knowledge about what the tool is &#8211; and what it isn&#8217;t.<\/p>\n<aside class=\"ns-obj ns-obj--aside ns-obj--full ns-obj--c3 ns-obj--split\"><div class=\"ns-obj-body\"><div class=\"ns-obj-col ns-obj--fmt-normal\"><h3>The Deflated Creator<\/h3>\n<p>A self-critical person may respond to negative or mixed AI feedback with discouragement. If the AI loved your work this morning and dismissed it this afternoon, what does that say about the work? For many people, especially those already prone to self-doubt, a single harsh AI verdict can be enough to shelve an idea. If they were aware that the harsh verdict was merely a statistical outlier &#8211; the equivalent of one editor at one publisher having a bad day &#8211; they might react differently. Without that context, it reads as definitive judgment.<\/p><\/div><div class=\"ns-obj-col ns-obj--fmt-normal\"><h3>The Overconfident Creator<\/h3>\n<h3>\n<p style=\"font-weight: 400\">The second response is the mirror image: unfounded confidence. A confident creator shops sessions until they get the enthusiastic verdict they were hoping for, takes a screenshot, and proceeds as though the AI has validated their work. This is a perfectly natural psychological response to an uncertain situation. But it is dangerous, because the encouragement carries no more weight than the discouragement. It is simply a lucky draw, not an informed assessment. The creator moves forward with conviction built on sand.<\/p>\n<\/h3><\/div><div class=\"ns-obj-col ns-obj--fmt-normal\"><h3>The Sophisticated Creator<\/h3>\n<p>The third response is the only genuinely useful one: treating the inconsistency as information in itself. A creator who asks the same AI the same question multiple times &#8211; or who deliberately varies the framing to stress-test an idea &#8211; is not gaming the system. They are doing something close to what a good editor or creative director does when they seek multiple opinions before committing to a direction. The variance in the responses tells them something real: where the idea is genuinely strong (consistent praise across sessions), and where it is genuinely fragile (inconsistent or conflicting feedback).<\/p><\/div><\/div><\/aside>\n<p>Rowling had no choice but to treat twelve verdicts as twelve draws from the same deck. With AI, you can run the draws yourself &#8211; the question is what you do with them.<\/p>\n<h2>How to use AI as a creative partner?<\/h2>\n<p>None of this means AI feedback is worthless. A single session can surface blind spots, identify unclear passages, suggest alternatives you had not considered, and push you toward a more refined version of your idea.<\/p>\n<aside class=\"ns-obj ns-obj--aside ns-obj--right ns-obj--c1 ns-obj--fmt-inverted\" style=\"width:62%\"><div class=\"ns-obj-body\"><h3>Practitioner advice<\/h3>\n<p>One challenge in this context is \u201cpersistent memory\u201d, where the Large Language Model stores the content of previous conversations. If this is included in the second and third evaluations, it dilutes the effect of giving you a fresh perspective and will continue to justify previous statements. When shopping for independent feedback, turn it off!<\/p>\n<p>One caveat on topics: The more novel and outside-of-the-box your idea is, the less an LLM will be able to provide you with content feedback. It is built on past knowledge and patterns, and not well equipped to detect valuable novelty that is beyond extending what already exists. Thus, you can still spell-check and check for completeness, but don\u2019t expect a solid evaluation of the substantive content.<\/p><\/div><\/aside>\n<p>The real leverage, though, comes from using the statistical distribution to your advantage. Ask the model the same question three times in three fresh sessions, and you get something no single answer can give you: a distribution. Where all three runs agree &#8211; in praise or in criticism &#8211; you are probably looking at a genuine property of your work. Where they contradict each other, you have found the places where reasonable judges could disagree, which is exactly where your own judgment is needed. If you want, you can hand the comparison work back to the model and let it map the agreements and contradictions for you.<\/p>\n<p>The lottery doesn\u2019t stop being a lottery. But once you know that\u2019s what it is, you stop mistaking a single ticket for a verdict &#8211; and start reading the odds. Your business plan deserves at least three draws.<\/p>","protected":false},"excerpt":{"rendered":"<p>Large Language Models judge your work differently every time you ask. That is not just a quirk of the technology &#8211; it is a fundamental challenge to how we create, validate, and trust ideas.<\/p>","protected":false},"author":15,"featured_media":225034,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"ns_references":"","ns_references_title":"","n9tr_seo_title_de_DE":"","n9tr_seo_description_de_DE":"","n9tr_seo_title_fr_FR":"","n9tr_seo_description_fr_FR":"","footnotes":""},"categories":[47,50],"tags":[],"class_list":["post-227644","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-blog","category-homepage"],"acf":{"tag_line":"When AI answers differ each time","about":"This post explores a structural characteristic of Large Language Models and its practical implications for creators, entrepreneurs, and anyone using AI as a thinking partner.","tldr":"If you want to use AI for feedback, do it with caution and don't take everything it says for granted. And do multiple runs as if you would ask multiple people.","ai_support":"The text was written by a human and submitted to an AI system for final review, such as checking grammar, typos, or logical consistency."},"_links":{"self":[{"href":"https:\/\/www.9senses.ai\/fr\/wp-json\/wp\/v2\/posts\/227644","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.9senses.ai\/fr\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.9senses.ai\/fr\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.9senses.ai\/fr\/wp-json\/wp\/v2\/users\/15"}],"replies":[{"embeddable":true,"href":"https:\/\/www.9senses.ai\/fr\/wp-json\/wp\/v2\/comments?post=227644"}],"version-history":[{"count":28,"href":"https:\/\/www.9senses.ai\/fr\/wp-json\/wp\/v2\/posts\/227644\/revisions"}],"predecessor-version":[{"id":228162,"href":"https:\/\/www.9senses.ai\/fr\/wp-json\/wp\/v2\/posts\/227644\/revisions\/228162"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.9senses.ai\/fr\/wp-json\/wp\/v2\/media\/225034"}],"wp:attachment":[{"href":"https:\/\/www.9senses.ai\/fr\/wp-json\/wp\/v2\/media?parent=227644"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.9senses.ai\/fr\/wp-json\/wp\/v2\/categories?post=227644"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.9senses.ai\/fr\/wp-json\/wp\/v2\/tags?post=227644"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}