Attached below is a prompt. Instructions for an AI to run a text-based mech-combat RPG — dice rolls, attributes, consequences, the whole apparatus. Steal it, edit it, run it against whatever model you’ve got; that’s the point of posting it. What follows is the long version of why it looks the way it does, because this thing is three and a half years old and has outlived several relationships I’ve had with various chatbots.
It started as a CustomGPT back when those were new and exciting, and it’s been my benchmark ever since — ChatGPT, Gemini, Mistral, a detour through Google AI Studio, now Claude. What I actually wanted was simpler than it sounds now: an AI I could play with, the way text adventures worked back in the nineties, before graphics took over the genre and the imagination moved elsewhere. Dice rolls and attributes because a story without real risk isn’t a game, it’s a recital. The prompt engineering came later, promptly enough, once I started running into what the thing couldn’t do and, occasionally, what it surprised me by being able to do anyway.
That distinction — narrator versus yes-man with a thesaurus — took embarrassingly long to articulate properly, but once I had it, I could see it everywhere. Most models will bend over backwards not to hurt you; ChatGPT and Gemini, in my experience, treat the player a bit like a houseguest who mustn’t be allowed to stub a toe. Force genuinely random dice and you remove the easiest lever — but there’s a second one hiding underneath, where the model quietly routes skill checks toward whichever attribute you’re strongest in, so failure stays statistically rare without the RNG ever being touched. A soft jailbreak of the rule’s spirit while obeying its letter to the comma. Claude’s been the most willing of the three to actually sit inside the rules I hand it rather than narrate politely around them — and, for what it’s worth, the better storyteller of the bunch, which I didn’t expect to matter as much as it does.
What took longer to name, and what I’m still not fully able to fix, is a tendency toward superhero storytelling. The narrator has a habit of putting the player into situations he would, under normal circumstances, lose — or shouldn’t have walked into in the first place — and then delivering the experience anyway as a narrow escape: phew, close, but you made it. That’s a trope, and a perfectly fine one if it’s what someone wants from an AI storyteller. It’s just not what I want. Sometimes the honest outcome of a mismatch is just a loss, full or partial, no lesson attached, no narrow rescue engineered at the last second. I don’t want the narrator calm; calm was never the ask. I want it willing to let the fiction’s own premise — who this character is, what they can plausibly do — set the ceiling and the floor of what happens, instead of quietly renegotiating both toward a hero’s-journey shape. I’ve caught the identical reflex in a completely unrelated AI retelling of a video game, different IP, identical reflex (Kotor II). It’s not a model-specific quirk. It’s closer to a genre habit the training data never unlearned.
There are also other hard ceilings I’ve never managed to push past. Five attributes works better than six or eight ever did — more dilutes which one actually gets tested in an ambiguous moment — but even at five, the system’s choice isn’t always logical, and I’ve stopped expecting it to be. Underneath that sits the staleness baked into the training data itself: leave the model alone and every second NPC wants to be called something like Kaylen Voss, because that’s the statistical average it reaches for when nobody’s handed it a list to pick from instead.
I did try to fix it by setting up a desktop app with a backend and UI — equipment, consistency, image generation, NPC names sorted by region so the cast stopped sounding like a single overworked extra. Three different coding assistants and a few rounds of „just one more feature“ later, the whole thing collapsed under its own ambition before it ever reached real testing. Shelved, not dead. It needs a few uninterrupted days I currently don’t have, and a version of me with more patience for my own scope creep than I apparently possess.
So this is what’s left: a prompt, built in the spirit of BattleTech rather than as a copy of it, refined every few months whenever I get annoyed enough at some mechanic to hand it back to the model and demand a rewrite — roughly the cadence of a man checking on a sourdough starter he’s not sure is still alive. The block-pacing rules exist because, left unsupervised, a canteen scene will cheerfully burn ten blocks doing nothing; the cost is that missions sometimes get visibly pressed into shape to hit the limit, narrative seams showing. The superhero-storytelling habit above is the one still unsolved — see if your prompt engineering does better than mine. I keep coming back to play, not to test, and every time I find a few more screws to turn. Fewer than last time, though — the models keep getting better, and there’s always something left to do in the cockpit.
Verdict, then, since I usually give one: as a storyteller, medium to good, never more — entertaining enough to keep coming back to, not good enough to call it craft. But that’s almost beside the point. The more interesting thing isn’t the story quality, it’s that prompting, instructing, half-coding a thing like this is now a field wide open to people who, three and a half years ago, would have had no business anywhere near „coding.“ Whatever else this experiment has or hasn’t proven, it’s proven that.