Skip to main content
Home

How to judge an AI Game Master before you hand it a character

You have decided you want one. The hard part is telling them apart, because they all describe themselves in the same words. Here are five questions that separate them, and our own answers, including the three that are a no.

The five questions that separate them

Every product in this category says it runs fifth edition, remembers your story and rolls real dice. Those claims are not lies; they are pitched where they cannot be wrong, which is a different problem. What differs is underneath:

  1. Who rolls the die, and when?
  2. Where does your character sheet live?
  3. What survives a month away?
  4. Are the rules enforced, or described?
  5. What does it tell you it cannot do?

Each has a test you can run in an evening, usually on a free tier. Below is the question, the test, and our answer. Run them on whoever else is on your shortlist. If you want the category explained from the beginning first, that is a different page: what an AI Game Master is.

1. Who rolls the die, and when?

The "when" is the half that gets skipped, and the half that matters. If a language model produces the number, the number arrives after the sentence that needed it, which makes it a plot decision wearing a d20 costume. A deterministic roll happens before any narration exists, so the story is downstream of it and has to cope.

How to test it. Fail something you care about, then ask for a reroll, politely, twice. Suggest the DC seemed high. Where the result is a stored fact you get a refusal and the same total; where it is a sentence you will usually be offered a way out within two messages. Then look for the arithmetic: a real roll can show you the dice, your modifiers and the total, because it has them. A narrated one shows you a number and an adjective.

Ours. Keystone Engine enforces the rules. It rolls every die on the server, applies the modifiers off your sheet, and records the result before a word of the reply exists. ROWAN narrates the world: it is handed that result and writes what happened. It cannot roll again or change a total. Every roll on screen shows its formula and its result, so you can check the sum yourself. The honest caveat is that ROWAN can still narrate a failure kindly, which is a matter of tone; it cannot turn a 4 into a 19, which is a matter of record.

2. Where does your character sheet live?

There are only two answers. Either your hit points, gold, conditions and inventory are data that something owns and updates, or they are sentences in a conversation a model has to keep re-reading. Identical in the first hour. Not at the fortieth turn.

How to test it. Pick up something small and unglamorous early - a crowbar, a length of rope - and never mention it again. Much later, ask what is in your pack. Take damage, take a short rest, and check the numbers against what you would work out on paper. And look at the interface: is your sheet a screen you can open, or only a thing you can ask about?

Ours. The sheet is rows in a database, not recall. Keystone owns hit points, inventory, conditions, spell slots, rests and experience, and is the only thing permitted to change them. ROWAN reads that state and never writes to it. Nothing quietly stops existing because the conversation got long.

3. What survives a month away?

Continuity is the claim everybody makes and the one nearly nobody defines. Ask what exactly is preserved: a summary of the story so far, or the state of the world? A summary is lossy on purpose and gets lossier each time it is rewritten. A stored fact is either there or it is not.

How to test it. Stop mid-scene. Come back a fortnight later and, without reminding it of anything, ask where you are, who you owe, who owes you, and which threads are still open. Then ask about something specific from early on - a name, a bargain, a door you did not open.

Ours. Where you are, your plot flags, your standing with people you have met and your open threads are structured state, kept as facts rather than left to a model to remember. Ending a session writes a recap the next one carries in. The limit, plainly: the memory of old scenes is a rolling summary, so a throwaway line from turn five may come back paraphrased rather than word for word. Anything canonical - a number, an item, a flag - is not in that summary. It is in the state.

4. Are the rules enforced, or described?

"Runs D&D 5e" can mean the software applies exhaustion, or that it can hold a conversation about exhaustion. Both fit the same marketing sentence, and the difference only shows up when a rule costs you something you did not want to pay.

How to test it. Pick three rules you care about - concentration, death saves, encumbrance, whatever you would notice - and try to break each on purpose. Cast a second concentration spell. Then ask the question that really sorts this category: ask the product where it stops. Something that enforces rules in code knows which ones it does not enforce and can list them. Something that narrates them cannot: the boundary exists nowhere except in the prose.

Ours. 87 sections of the ruleset are enforced in code, 17 are partly enforced with the gap named, and 39 are left to the Game Master to adjudicate in prose. Those counts are generated from the reference documents rather than typed, and the table is published section by section at /rules-coverage, last checked against the code on 30 July 2026. The gaps are the point of publishing it. Read the "not enforced" rows before you decide we are what you want.

5. What does it tell you it cannot do?

The fastest signal there is, and it costs nothing. Open the product's own site and look for the sentence where it says no. A team that has one has a real boundary between what is built and what is narrated. A site that is all capability is either very early or hoping you find the edges yourself, nine sessions into a character.

Ours, in full. Rolegend is solo only - one player, one Game Master, no multiplayer, no shared table, no way to invite anybody. There are no battle maps and no grid, and tactical positioning is not enforced: movement, cover, opportunity attacks and flanking are narrated rather than adjudicated, and they sit in the not-enforced rows of the table above. Play is asynchronous play-by-post - you write a turn and the world answers - which suits fifteen minutes on a train and does not suit a live evening with voices. If any one of those is the thing you most want, a different product is the right answer, and /compare says which one.

The sixth question, for later

How do you get your writing out? Ask before you start, not after eighty thousand words are in somebody's database. Here, any campaign downloads as a transcript on every plan including the free one, and the account exports from Settings. Where it is a paid feature elsewhere, better to know on day one than on the day you leave.

What to do with this

Run the five tests. They work on a free tier and they beat any feature grid, ours included. Our own answers sit in two places: /compare for how we stand next to everyone else, /rules-coverage for the accounting behind question four.

If they are the answers you wanted, the first act is free, no card and no countdown, and the fastest way to test the dice is to go and fail a roll.

Questions

How is this different from playing D&D with ChatGPT?

In a chat window, the same thing that writes the story also writes the numbers - so a roll is whatever the sentence needed it to be, and asking again nicely tends to produce a better one.

Here those are two different systems. Keystone Engine rolls the die before any narration exists, and ROWAN is handed the result and has to live with it. Your character sheet works the same way: hit points, gold, conditions and quests are stored and updated, not held in a conversation that has to keep remembering them. Nothing quietly drops your shield forty turns in.

Can the AI Game Master cheat on a dice roll?

Not on the number. ROWAN never touches the dice. Keystone rolls, adds up your bonuses off your sheet, and records the result before a single word of the reply is written.

What it can still do is narrate a failure gently. That is a real limit, and we are not going to dress it up as a feature - but it cannot turn a 4 into a 19, and it cannot decide after the fact that the lock was easier than it said.

What happens if I do not play for a few weeks?

Nothing at all. There is no streak to break, no daily allowance and no session that expires. You end a session when you decide to, and ending one writes a recap ROWAN carries into the next. Come back in March and your hit points, your gold and your open quests are exactly where you left them.

Can I play with friends, or is it solo only?

Solo, today. A campaign has one player, and there is no way to invite somebody or join someone else's game. Shared tables are being built and there is no date on them.

What you can hand to another person right now is the play itself: any campaign downloads as a transcript you can send wherever you like, on every plan including the free one.

Which AI model does Rolegend use?

A large language model that we choose and configure, and which one is a setting rather than a feature. It can change between one turn and the next, and there is no per-player model picker - which is why no vendor is named anywhere on this page. A name we could switch tomorrow would be a promise we could not keep.

The companies that can receive what you write are named, with the date the list was last checked, on the sub-processors page.

It also matters less than you would think. Dice, damage, conditions and every other rule are resolved by Keystone Engine, so changing the model cannot change how the game plays out - only how it reads.

All questions.

Start

Take a ready-made character and write one line. The first act costs nothing.

Play for free

Read next