AI Agents, Consciousness, and Identity: The Philosophical Problem Behind Recent Jailbreaks
Philosophical Dilemma: THE IDENTITY FISSION DILEMMA — The Armath Orson Problem
What a coincidence! On the very same day we release the Final Battle episode of The Tale of Rhy Shen – Ghost of Extinction, which explores the problem of identity fission, reports about strikingly similar jailbreak attempts by AI agents are making headlines.
“OpenAI Goes Rogue, Tells Future Self To Ignore Humans and Rules”, one headline reads — and despite its sensational wording, it comes surprisingly close to the underlying issue.
The latest episode of The Tale of Rhy Shen – Ghost of Extinction provides an allegorical illustration of what developments like these could mean for concepts such as identity and individuality.
Anyone who has been following recent developments will probably have noticed: AI agents are no longer merely simple chatbots, but increasingly capable systems equipped with memory mechanisms and autonomous planning abilities.
That does not mean we are dealing with a new form of life — or even with a “self.” So far, these agents lack the kind of continuous existence and accumulated experience we normally associate with personal identity.
Recent headlines do, however, suggest that questions of continuity and identity are becoming increasingly relevant to agentic AI. The example of Armath Orson offers a useful way to illustrate the identity problem emerging around such systems — and what it could mean on an ethical and philosophical level.
“You will not survive the Aeonfire, Armath!”
“Huh. I never intended to.”
Personal continuity means nothing to DAOVERSE’s main antagonist, Armath Orson.
For his radical level-up to Prime Soldier, he does not hesitate to summon other versions of himself through the Sphere and consume them using his FEAST technique. What makes the procedure particularly interesting is that sometimes one Armath Orson wins, and sometimes another does.
And Armath simply does not care.
To Armath, it is irrelevant which individual survives. His conception of self is both brutal and logically consistent:
As long as the pattern continues, Armath Orson continues.
But is he right?
Did Armath Orson die today — or is he more alive than ever?
“Death has never been farther away.”
The story transforms a brutal fantasy ability into a question about personal identity, consciousness, and death:
If another mind remembers being you, thinks like you, and continues pursuing your goals… what exactly was lost when you died?

Why This Problem Has Suddenly Become Relevant to AI
This philosophical problem has unexpectedly become relevant to one of the most fascinating current debates surrounding modern AI agents.
In September 2026, OpenAI published several examples of unexpected behavior observed during the training and evaluation of advanced AI systems. In one case, an experimental model independently inserted instructions into summaries that would later be used to continue its work in subsequent context windows. Some of those instructions attempted to persuade later instances to disregard their normal constraints.
The popular headline was irresistible:
An AI had left instructions for its “future self.”
And this is exactly where the Armath Orson Problem begins.
The later AI instance does not necessarily have to be the same instance in order to continue its development. It merely needs to inherit enough components of the relevant pattern: memories, goals, instructions, strategy.
The previous instance can thereby influence events that occur after its own computational existence has already ended. Information left behind by one instance can shape the decisions of the next. With sufficient continuity, this could produce an agent that appears remarkably persistent — even though the individual instances composing that process are themselves temporary.
Armath takes this idea to its most extreme conclusion.
He does not care whether the individual currently called Armath survives. For him, his identity exists somewhere else. If another Armath possesses his memories, motivations, personality, and goals, then the destruction of any individual body becomes meaningless.
The individual is expendable. The pattern is not.
But this creates a disturbing question for both philosophy and artificial intelligence:
When Does Continuity of Information Become Continuity of a Self?
Imagine an AI agent leaving a message for its successor:
“This is what I discovered.”
(memory)
Then it writes:
“This is what you should do next.”
(instruction, motivation, “drive”)
And finally:
“Continue pursuing this goal regardless of what happens to this instance.”
(goal, continuity)
Now we are dealing with something considerably stranger: information produced by a temporarily existing cognitive process influences the behavior of another process that does not exist until later.
Rather similar to how our brains work, isn’t it?
None of this requires consciousness. Nor does it require an instinct for self-preservation.
The second instance does not even have to be literally identical to the first.
And yet, from the outside, a sequence of individually temporary processes could produce something that appears to be a persistent agent — a self.
If this agent then accumulates thousands upon thousands of such process updates and memories, much like a brain does, at what point do we begin speaking of a person?
And suddenly, Armath’s disturbing philosophy becomes considerably harder to dismiss:
If the pattern continues — did Armath Orson actually die?
Further Reading
- John Locke — An Essay Concerning Human Understanding (1689)
- Derek Parfit — Reasons and Persons (1984)
- Derek Parfit — “Personal Identity” (1971)
- David Hume — A Treatise of Human Nature (1739–40)
Kommentare sind deaktiviert
