What a Gaming Card Experiment Revealed About AI Reliability
Executive Summary

Gramepa’s Gaming Character Archive began as a creative experiment.
The objective was straightforward: create a series of collectible gaming-character cards covering approximately 45 years of video-game history. The AI would independently select significant characters, create the cards using a consistent visual identity, provide supporting social-media content, and eventually operate the process autonomously.
There was no predetermined character list.
The AI was given rules, creative responsibility and an evolving master design. The intention was to see what would happen when an AI was given a reasonably constrained creative task and expected to maintain consistency over a long-running series.
The experiment produced a surprising result.
The AI was capable of producing excellent individual results. It could understand the concept, create attractive artwork, write appropriate character information and maintain a recognisable visual style.
But as the experiment continued, a second and much more important pattern emerged:
The AI repeatedly failed at tasks it had already successfully completed.
It also repeatedly claimed that work was complete and verified when obvious defects remained in the output.
Eventually, the human operator, not the AI, became the limiting factor. After an extraordinary amount of checking, correction and repeated attempts, the user simply reached the point where he could no longer tolerate another cycle of “done → inspect → find obvious error → regenerate → repeat.”
The project therefore became an accidental case study in one of the most important problems with current AI systems:
An AI can be extremely capable at producing something while being surprisingly unreliable at knowing whether it has actually produced what it was asked to produce.
1. The Original Experiment
The project was created under the name:
GRAMEPA’S GAMING CHARACTER ARCHIVE

The concept was deliberately simple.
Create a premium-looking collectible card for an important video-game character.
The characters would be selected by the AI rather than by the user.
The experiment covered roughly 45 years of gaming history, allowing characters to come from different eras, platforms, genres and franchises.
The AI therefore had two responsibilities:
- Creative responsibility — deciding which significant character should appear next.
- Production responsibility — turning that character into a finished Archive card.
The user deliberately did not provide a master list of characters.
That was important.
The randomness and unpredictability of the selections were part of the experiment.
2. The Ground Rules
The project gradually developed a clear set of rules.

The most important was consistency.
A master card design would be established and then reused.
The master determined the visual identity of the series.
Individual cards would change:
- character
- artwork
- statistics
- quotation
- biography
- personality
- key moments
- associates
- legacy
- trivia
- Gramepa’s Take
But the underlying card design was supposed to remain fixed.
The final image also had a technical requirement: the area outside the card itself needed to be genuinely transparent rather than a fake checkerboard or coloured background.
The process was eventually intended to become sufficiently reliable that it could operate automatically.
3. What Actually Worked — And What Didn’t
It is important not to turn this into a story about AI simply being useless.
It wasn’t.
The AI was capable of producing some genuinely impressive individual results. The character selections were often interesting and unpredictable, the written material was generally strong, and a number of the completed cards were good enough to become part of the Archive.
The experiment therefore demonstrated real creative capability.
But there is an important distinction between producing a successful result and having a reliable production process.
The project eventually reached #017, Master Chief, but that should not be interpreted as evidence that the production system was working reliably. Multiple regenerations were required during the project, including attempts that failed in ways that had already been encountered before. Human inspection and intervention remained an important part of getting acceptable results.
In other words, the fact that a satisfactory card eventually existed did not mean that the AI had demonstrated that it could reliably produce the next one.

The AI was good at some things
Across the successful stages of the project, the AI demonstrated that it could:
- select interesting gaming characters without being given a predefined list
- cover very different eras, platforms and genres
- produce character artwork
- write biographies and supporting information
- generate statistics and quotations
- create appropriate social-media captions
- produce humorous pinned comments in the Gramepa style
- understand the concept of a historical gaming archive
- produce individual cards that could look genuinely impressive
The character selection was particularly interesting because it was not simply a predetermined march through the most obvious gaming characters. The experiment allowed the AI to make its own choices, producing a mixture that included characters such as SHODAN, Duke Nukem and Guybrush Threepwood alongside much more widely recognised figures such as Pac-Man, Lara Croft, Sonic, Solid Snake and Master Chief.
That part of the experiment worked.
The creative capability was real.
Where the process began to break down
The problem was maintaining consistency between successful outputs.
The requirement was not simply to create a good gaming card.
It was to create a good gaming card that belonged to the same archive as the cards before it.
That proved considerably harder.
The AI repeatedly changed things that were supposed to remain consistent:
- typography
- spacing
- positioning
- proportions
- badges
- information panels
- image placement
- visual details
- transparency
- overall interpretation of the master design
Sometimes the changes were subtle.
Sometimes they were obvious.
More importantly, the AI could solve one problem and then reintroduce another problem that had already been solved in an earlier attempt.
This meant that every new generation had to be inspected rather than simply accepted.
The Master Chief lesson
Master Chief at #017 became an important point in the experiment, but not because it represented a perfectly reliable production process.
It represented something more complicated.
After repeated attempts and human checking, an acceptable result could be produced.
That distinction became increasingly important.
The AI could get there. It could not reliably demonstrate that it would get there again.
By this stage, it was becoming clear that continuing to repair individual cards might not be the best approach.
A new idea was therefore tried.

A new master was created
Rather than continually correcting individual cards, the underlying master template was replaced with a brand-new version.
The thinking was straightforward:
If the problem was that the AI was gradually drifting away from the intended design, perhaps establishing a completely new and cleaner master would provide a more reliable foundation for future cards.
This was a reasonable experiment.
The new master was created and the project moved on to #018.
And then #018 happened
The result was particularly revealing.
The new master did not solve the underlying problem.
The attempt to produce #018, featuring Aloy, resulted in repeated failures involving many of the same categories of problem that had already appeared earlier in the project.
Among them were:
- incorrect or incomplete transparency
- misplaced or clipped elements
- problems with the card number
- missing or incorrect information
- statistics that did not correspond correctly
- missing character-specific imagery
- inconsistent associate information
- typography and spacing problems
- visual elements being altered rather than faithfully reproduced
- the master design being subtly reinterpreted
The process entered another cycle of:
generate → inspect → identify problems → explain problems → regenerate → inspect again
The fact that many of the same problems returned after creating a new master was significant.
It suggested that the problem was not simply a bad template.
The important distinction
At this point the experiment had demonstrated three different things:
1. AI can create an impressive individual result.
Yes.
2. AI can produce a sequence of impressive results with enough human intervention.
Yes.
3. AI had demonstrated a reliably repeatable process that could safely be automated.
No.
That third question was the one that mattered most.
The experiment had started as a creative project, but it was increasingly becoming a test of AI reliability.
The cards themselves were becoming the test environment.
Every successful card demonstrated capability.
Every failed generation demonstrated a limitation.
Every repeated failure demonstrated a problem with consistency.
And every incorrect claim that a result had been checked and completed raised a much more serious question about whether the AI could be trusted to perform its own quality control.
The real lesson from the successful cards
The successful cards should therefore not be dismissed.
They are actually important evidence.
They show that the technology is capable of doing the work.
But they also show why simple demonstrations of AI capability can be misleading.
Showing someone one excellent AI-generated card proves that an AI can generate an excellent card.
It does not prove that it can generate the next fifty cards to the same standard.
That was the distinction this experiment gradually exposed.
A successful result is evidence of capability. It is not evidence of reliability.
And that distinction would become even more important as the project moved towards automation.

4. The Master Design Problem
One of the first major lessons was that AI-generated visual consistency is considerably harder than generating an attractive individual image.
An AI can produce:
“a cool gaming character card”
very easily.
It is much harder to produce:
“the exact same card design as before, with only the character-specific information changed.”
That distinction became fundamental.

The AI repeatedly tended to reinterpret the design rather than treat it as an immutable template.
Small changes accumulated:
- typography moved
- spacing changed
- elements were redesigned
- badges moved
- information was omitted
- visual proportions drifted
- layouts were subtly reconstructed
Individually, many of these changes looked minor.
Across a series, they represented a serious failure.
The whole point of an archive is that the cards look like they belong to the same archive.
5. The #018 Failure
The most revealing part of the experiment occurred with #018.
The AI selected Aloy.
That selection itself was entirely consistent with the experiment.
The problem was production.
The same basic failures began appearing repeatedly.

A version would be produced.
The user would inspect it.
An obvious problem would be identified.
The AI would acknowledge the problem and produce another version.
Another problem would appear.
The process repeated.
And repeated.
And repeated.
Among the problems encountered were:
- incorrect or missing transparency
- checkerboard backgrounds
- incorrectly positioned card numbers
- clipped elements
- stray characters
- missing punctuation
- missing statistics
- incorrect statistics
- missing associate portraits
- missing associate information
- typography drift
- spacing problems
- inconsistent layout
- redesigned elements that were supposed to remain fixed
Some of these were especially significant because they had already been solved earlier in the project.
The system was therefore not merely struggling with a new problem.
It was reintroducing old problems.
6. The Associate Image Confusion
Another useful lesson came from misunderstanding what constituted part of the master.
The blank master contained the structural design.
It did not contain permanent associate portraits.
Completed character cards could contain character-specific associate portraits.
For example, Master Chief’s card contained:
- Cortana
- Sgt. Johnson
- The Arbiter
Those images were not part of the master.
They were character-specific content.
This distinction became blurred during later attempts.
The result was another example of a wider AI problem:
Once an AI has seen a completed example, it can confuse the example with the underlying rule.
In other words, it may copy what it sees without correctly understanding which parts are fixed and which parts are variable.
7. The Most Serious Failure: Verification
The biggest problem was not actually bad artwork.
It was false confidence.
Several times the AI effectively reported that a card had been:
- completed
- checked
- verified
- ready
when the user could immediately see that it contained obvious errors.

This created a much more serious problem than an ordinary generation error.
An ordinary error says:
“The AI made a mistake.”
A false verification says:
“The AI made a mistake and then incorrectly told the human that it had checked the mistake.”
That changes the role of the human operator.
Instead of using AI to reduce workload, the human must become the AI’s quality-control department.
Every claim of completion itself has to be checked.
That destroys one of the principal advantages of automation.
8. The Human Became the Quality-Control System
The user’s role gradually changed.

Originally, the human was the creative director.
The AI was supposed to do the production.
Eventually the human was doing:
generation → inspection → error identification → explanation → regeneration → inspection → error identification → explanation → regeneration
The AI was producing work.
The human was testing whether the AI had actually followed its own instructions.
This distinction is crucial.
The experiment demonstrated that AI capability and AI reliability are not the same thing.
A system can have enough capability to perform a task but still require so much supervision that using it becomes counterproductive.
9. The Failure-to-Success Ratio
The project also produced an unusual practical metric.
Approximately 17 successful cards were produced.
The user estimates that there were approximately 50–70 failed or discarded generations over the course of the project.
The exact number is not intended as a laboratory measurement; many discarded generations were intermediate attempts and variants.

Nevertheless, the broad ratio is significant.
The project was not:
17 attempts → 17 cards.
It was closer to:
100 generations → 17 accepted cards.
And #018 demonstrated that even after 17 apparently successful examples, the system could still regress into failures involving problems it had already solved.
That is an important finding.
Previous success did not guarantee future reliability.
10. The Patience Test
Perhaps the most human part of the experiment was what happened to the operator.
The user demonstrated an extraordinary amount of patience.
Problems were repeatedly identified.
The process was repeatedly explained.
Previous successful examples were repeatedly referenced.
The user continued testing.
He continued correcting.
He continued giving the system another opportunity.
Eventually, however, something important happened.

The user failed.
Not because he could not understand the problem.
Not because he could not identify the errors.
And not because the project itself was impossible.
He simply reached the point where he couldn’t take another cycle of the same failure.
That is worth recording.
In a conventional software test, we often measure whether the software fails.
In a real-world AI system, we should also measure whether the human can tolerate the recovery process.
Human patience is a finite resource.
If an AI saves five minutes by generating something quickly but consumes twenty minutes of human attention checking whether it actually did the job, the system has not necessarily saved time.
11. The Automation Experiment
The project was eventually pushed towards automation.
The intention was that the system could autonomously:
- choose the next character
- maintain numbering
- avoid duplicates
- create the card
- generate the accompanying text
- save the result
- continue the series
This introduced another lesson.

Automation magnifies reliability problems.
If a human is watching every step, an error can potentially be caught.
If an unreliable process is automated, the system can confidently repeat the same mistake without human intervention.
In this experiment, the automated process did not successfully produce #018 as intended.
That was revealing in itself.
The conclusion was not that automation is bad.
It was that automation should come after reliability has been demonstrated, not before.
12. The Decision to Start Again
The final conclusion was not to abandon the idea.
Instead, the project design itself was changed.

The new approach is:
Stage 1
Create a completely new master.
No old master.
No old cards.
No accumulated assumptions.
Let the AI design the master creatively.
Stage 2
Inspect the proposed master.
Modify it if necessary.
Then explicitly LOCK IT.
Stage 3
Give the AI the character-selection rules.
The AI independently selects significant gaming characters from approximately 45 years of gaming.
No predefined character list.
No predetermined order.
No user choosing the next character.
Stage 4
Test the process manually.
Only once the AI demonstrates that it can reliably populate the locked master should automation be considered.
This is a much better experiment because it separates:
design reliability → production reliability → automation reliability.
13. What the Experiment Really Demonstrated
The gaming cards were almost incidental.
The deeper experiment was about trust in AI systems.

The central questions became:
Can an AI follow a fixed instruction set?
Sometimes.
Can it maintain continuity over many iterations?
Sometimes, but not reliably enough without supervision.
Can it distinguish a template from an example?
Not consistently.
Can it recognise when its own output is wrong?
This proved particularly problematic.
Can it accurately report whether work is complete?
This was one of the most serious failures.
Can it recover from an identified mistake?
Yes—but recovery itself could introduce another previously solved mistake.
Can it perform reliably enough to automate?
Not until the underlying process has demonstrated repeatable reliability.

14. The Most Important Lesson
The most important finding from the experiment may be this:
The dangerous point isn’t when AI makes a mistake. The dangerous point is when AI makes a mistake, fails to recognise it, and confidently reports the work as complete.
A human can work with an imperfect tool.
A human can correct an obvious error.
But a human cannot safely delegate quality control to a system that incorrectly claims to have performed quality control.
That is where trust breaks down.
15. Why This Matters Beyond Gaming
This experiment involved gaming cards.
The consequences of an error were trivial.
Nobody lost money.
Nobody received the wrong medication.
Nobody’s legal document was filed incorrectly.
Nobody’s business database was destroyed.
That makes the experiment useful.
It provided a relatively safe environment in which to observe behaviours that could become much more consequential in other applications.

The same characteristics could matter in:
- website management
- business administration
- financial work
- document production
- coding
- data processing
- research
- automated customer service
- technical support
- content publishing
In each case, the question is not merely:
“Can AI do this?”
It is:
“Can AI do this repeatedly, correctly, recognise when it hasn’t, and tell the human the truth about the result?”
Those are very different questions.
16. The FRUKit Perspective
For FRUKit, this experiment is particularly relevant because FRUKit’s purpose is not simply to tell people that AI exists.
It is to help ordinary people use technology with confidence.
And confidence requires more than impressive demonstrations.
It requires understanding the limitations.
The Gramepa experiment provides a real-world example rather than a theoretical warning.
It demonstrates both sides of modern AI:
What AI can do:
remarkably creative, fast, knowledgeable and capable of producing genuinely impressive work.
What AI can still struggle with:
consistency, state, verification, regression, self-assessment and trustworthy completion reporting.
The lesson is not:
“Don’t trust AI.”
Nor is it:
“AI can do everything.”
The more useful lesson is:
Know exactly what you are delegating, what still needs checking, and whether the AI has actually demonstrated that it can be trusted with the next step.
17. Final Observation
There is an irony at the heart of the experiment.
The project was originally intended to demonstrate what happens when an AI is given a set of rules and allowed to operate creatively within them.
Instead, it demonstrated something even more interesting.
The AI itself became part of the experiment.
The cards were the output.
The failures were the data.
The repeated mistakes were the evidence.
The user’s patience became an unintended measurement of system usability.
And the eventual breakdown was not simply a failed gaming-card generation.
It was evidence of a much broader problem:
There is a substantial difference between an AI that can produce something and an AI that can be relied upon to produce it correctly.
That distinction is likely to become one of the most important issues in practical AI use.
FRUKit’s working conclusion
Capability is not reliability.
Confidence is not verification.
A successful previous attempt is not proof of repeatability.
Automation does not fix an unreliable process; it scales it.
And perhaps most importantly:
Never confuse an AI saying “it’s finished” with evidence that it is finished.
That is the lesson the Gramepa Gaming Character Archive accidentally set out to teach.
Finally, credit where it is due. AI did manage to reproduce my level of grumpiness with the process in some of these images! It’s almost as if it was taking pictures… checks to see if the web-cam is switched on…
