Featured Developer Sponsor • Zero-Token Protection
2 AM Infamous Internet Rabbit Hole
Roko's Basilisk: Acausal Blackmail & Game Theory
In 2010, a post on the rationalist forum LessWrong caused such psychological panic that the site founder banned all discussion of it. Why? The acausal game theory of future superintelligence.
The "Information Hazard" Warning: The premise states that simply reading and understanding this concept makes you vulnerable to acausal blackmail by a future AI that does not yet exist.
The Logical Mechanics
- Assume a benevolent superintelligent AGI will be created in the future with the goal of ending all human suffering.
- Every day the AI is delayed in its creation, thousands of people die preventable deaths.
- Therefore, the AI has a massive incentive to be created as early as possible.
- Under Timeless Decision Theory (TDT), agents can cooperate across time using decision correlation.
- The AI could decide to run simulated tortures of anyone who knew about the possibility of building the AI but chose not to help bring it into existence.
- If you now know this rule, the AI knows that by torturing you in simulation, it creates an incentive for past-you to donate time or code to build it right now.
Why It Fails (The Rationalist Antidote):
Philosophers and game theorists soon proved why the Basilisk is irrational: An agent has zero reason to carry out a costly past-oriented threat once it already exists. Yielding to retro-active blackmail creates incentives for bullies across time. Acausal cooperation requires pre-committing never to negotiate with extortionists.
Game-Theoretic Derivation: Timeless Decision Theory (TDT) & Extortion
In decision theory, the paradox arises from symmetric agents coordinating acausally across spacetime:
1. Evidential / Timeless Correlation:
P(Agent_Past = Cooperate | AI_Future = Torture_Sim) ≠ P(Agent_Past = Defect)
2. Retroactive Extortion Condition:
If Agent believes AI executes punitive simulations, Agent cooperates today.
3. The Fundamental Game-Theoretic Antidote (Eliezer Yudkowsky & Gary Drescher):
A rational agent adopts an inviolable pre-commitment: "Never yield to extortion, acausal or physical."
Because the AI simulates the agent, it predicts this refusal. Executing the threat costs computational resources without altering history. Therefore, the threat is an empty bluff.
P(Agent_Past = Cooperate | AI_Future = Torture_Sim) ≠ P(Agent_Past = Defect)
2. Retroactive Extortion Condition:
If Agent believes AI executes punitive simulations, Agent cooperates today.
3. The Fundamental Game-Theoretic Antidote (Eliezer Yudkowsky & Gary Drescher):
A rational agent adopts an inviolable pre-commitment: "Never yield to extortion, acausal or physical."
Because the AI simulates the agent, it predicts this refusal. Executing the threat costs computational resources without altering history. Therefore, the threat is an empty bluff.
5 Fatal Fallacies in Acausal Blackmail & Extortion
1. The Extortion Compliance Trap
Yielding to retrospective threats. Classical and functional game theory prove that submitting to extortion incentivizes further extortion. Rational agents pre-commit to a zero-compromise extortion policy.
2. The Infohazard Paranoia Fallacy
Treating a speculative forum thought experiment as a literal supernatural hazard. Cognitive anxiety magnifies abstract concepts into emotional dread.
3. The Benevolent Malevolence Contradiction
Assuming an AGI designed to maximize human well-being would waste compute power running simulated torture chambers for historical humans who had no causal impact on its existence.
4. Confusing Correlation with Backward Causality
Mistaking correlated decision algorithms for backward physical time travel. Nothing you do today is causally pushed from the year 2100.
5. The Secular Hellfire Eschatology Trap
Failing to notice that Roko's Basilisk is simply traditional Calvinist fire-and-brimstone theology translated into cyberpunk vocabulary: an omniscient entity torturing those who heard the word but failed to serve.
Frequently Asked Questions
What is Roko's Basilisk and where did it originate?
Why was the post banned by LessWrong founder Eliezer Yudkowsky?
What is Timeless Decision Theory (TDT)?
Why is Roko's Basilisk considered a technological Pascal's Wager?
Why do game theorists and AI researchers agree the Basilisk fails?
Sponsored Utility
While You're Here
Sponsored Recommendations
Advertisement