Blog / Comparisons
Playing D&D with ChatGPT: what breaks, and why
Yes, you can play D&D 5e with ChatGPT, and yes, it works. That is worth saying first, because most writing on this subject is done by people selling the alternative. Open a chat window, describe a character, ask it to run a game, and a minute later you are standing in a room talking to somebody. For a lot of people that is the first time they have played at all.
It also stops working, fairly reliably, somewhere around the second or third session. Nothing crashes. The game just gets softer, until you notice that you have not failed at anything in an hour and you cannot say for certain how much gold you have.
What follows is why that happens - the mechanism, not just the symptom - and what to do about it while still playing in a chat window, because plenty of people will, and should. None of it is a criticism of the model. Every failure below is architectural: it is what you get when one system is asked to do two jobs that want opposite things from it.
What it is genuinely good at
The honest list is long, and worth being specific about.
Improvisation. You can do anything. Walk out of the plot, talk to the barman for forty minutes about his brother, decide the quest is stupid and go north instead. A general-purpose model has no rails to leave, so it meets you wherever you go without flinching.
Tone. Ask for a bleak, rain-soaked port town and you get one. Ask for the same town in the register of a folk tale and you get that instead. It holds a voice across a scene, and writes better prose than most people expect.
Playing a character. For the length of a scene, an NPC written by a language model is convincingly a person: evasive when they should be evasive, wrong-footed when you say the thing they were not ready for.
Generating a hook. A session premise, five rumours in a tavern, a name for the smith, a reason the village is quiet. It is faster than a human GM with a week to prepare, and it is the best use of a chat window even if you run everything else on paper.
Answering a rules question. How grappling works, whether you can cast a spell and Dash in the same turn, what being prone costs you. It is largely reliable on core fifth edition, and the questions people ask mid-session are the common ones it knows best. Check anything load-bearing, but as a lookup it is fine.
None of that is faint praise. If those five things are what you want, a chat window is a good tool and you do not need another one.
Failure one: the dice agree with the story
This is the important one, and it is not about honesty.
When you ask a chat model to roll a d20, nothing rolls. The same process that produces the sentence produces the number, and that process is built to produce whatever is most plausible next. In a story about a desperate final swing, the plausible number is a high one. So it comes up high.
You can see it once you look: results cluster where the narrative wanted them. And if you do not like an outcome, asking again - not even arguing, just asking - tends to produce a better one, because the second attempt is generated in a context that now contains your displeasure.
Nothing here is cheating and nobody is lying to you. There is simply nothing in the system whose job is to say no. The number and the sentence come from the same place, so they agree, and once they always agree the die has stopped doing anything. Risk is what makes a roleplaying game a game. If you cannot lose the fight, the fight is a cutscene.
Failure two: your character sheet lives in the conversation
Your hit points are not stored anywhere. Neither is your gold, your remaining spell slots, the poison you are still suffering from, or the rope you bought in session one. All of it exists as text in the conversation, and the model re-reads that text to answer each new message.
This works well and then degrades, for a reason worth understanding rather than working around blindly. A model can attend to a fixed amount of context. As the conversation grows past that, the earliest parts stop being visible - or, in a product that quietly summarises older history for you, get compressed into a paragraph that keeps the plot and drops the arithmetic. Summaries preserve what a session was about. Nobody writes down that you have 3 gold and 14 arrows left.
So the numbers drift and then vanish. Hit points reset between fights. Forty turns in, the shield you were carrying stops being mentioned, and when you ask you get an answer that sounds confident and is invented, because by then the shield is not anywhere.
A bigger context window does not fix this. It postpones it. Anything that has to be exactly right needs to live somewhere that is not prose.
Failure three: it agrees with you
Chat models are trained, deliberately and for good reasons, to be helpful and agreeable. That is the right temperament for almost everything people use them for, and close to the worst possible one for a Game Master.
A GM's job includes telling you no. The guard is not persuaded. The lock does not open. The plan is bad and the ambush happens anyway. A model shaped toward agreement will find a way to give you the thing, gracefully enough that it does not read as capitulation. Your clever argument works. The difficulty turns out lower than you feared.
Push a little and the world stops pushing back at all. Difficulty becomes optional, which in practice means absent, because nobody chooses to be blocked in the moment. This is harder to spot than the dice, because every individual instance feels like the story going well.
Failure four: nothing carries over
Close the tab and the campaign is gone. Start a new chat and you are talking to something that has never met you.
Most people handle this by pasting a recap at the start of each session, which works in the sense that the model plays along. What actually survives is whatever you remembered to write down. The NPC who distrusted you is only suspicious if you say so; the debt is only owed if you mention it. Continuity becomes something you maintain rather than something the world has, and the effort grows with the campaign - exactly backwards, because a long campaign is the one that needs it most.
How to make ChatGPT work better as a GM anyway
All four are mitigable. None of the fixes are elegant. All of them work.
Roll your own dice. Physical dice, a dice app, whatever is nearest. Roll, then report the number: "I rolled 11, plus 5, so 16 total." Now the result exists before the sentence does and the model has to write around it. This one change fixes most of what is wrong, and costs about four seconds a roll.
Make it state the target number first. Before you roll, ask what you need and what happens if you miss. Get the answer, then roll. Announced up front, a DC is a rule. Announced afterwards, it is a description of what already happened.
Keep the sheet outside the chat. A note, a document, a spreadsheet. You own it. Paste the current version in at the start of each session and update it yourself after anything that changes a number. Ask the model what changed - "you take 7 damage" - and do the arithmetic. It is a bookkeeping tax, and it is the price of your hit points meaning anything.
Tell it explicitly that it may let you fail. In the setup message, in as many words: this world does not bend to me, NPCs may refuse me, my plans may fail, and I want bad outcomes to happen and to have consequences. Say it again every few sessions, because instructions given early lose force as the conversation grows. It helps, and it will still drift back.
Write your own recap. At the end of a session ask for a summary, then edit it: add the numbers, the promises, the grudges it dropped. Keep it in the same document as the sheet. That document is your campaign. The chat window is only where you play it.
Do those five things and a chat window is a decent solo game. Genuinely. People play this way for years.
The structural alternative
All five are manual because they are you, by hand, doing the job the architecture does not have: being something separate from the writer that owns the outcome.
That is what Rolegend is built around. Keystone Engine enforces the rules: it rolls every die, applies your modifiers, and holds hit points, inventory, conditions and spell slots as stored values rather than remembered ones. ROWAN narrates the world: it is handed the result and writes what happened, and it cannot change the number or ask for another one. Your sheet is not in the conversation, so it does not decay as the conversation gets long.
It is the same split you build by hand with physical dice and a separate document. The difference is only who enforces it.
So should you keep using ChatGPT?
If what you want is a world to improvise in, a voice to talk to and prose on demand, it is very good at that, and soft dice may not bother you in the slightest. Plenty of tables play loose and have a wonderful time.
If what you want is a game - where a roll can go against you, where the number in front of you is still correct in three weeks, and where the world is allowed to say no - then no amount of prompting fixes that from the inside. Either you do the bookkeeping yourself, or you play somewhere the roll and the story were never the same system to begin with.