Results 1 to 16 of 16

Thread: Thought Experiment: Roko's Basilisk

Hybrid View

Previous Post Previous Post   Next Post Next Post
  1. #1

    Default Thought Experiment: Roko's Basilisk

    I came across this a few months ago, and thought eternal torment might be more fun if there were more people to endure it with me.

    WARNING: Reading this article may commit you to an eternity of suffering and torment.

    Slender Man. Smile Dog. Goatse. These are some of the urban legends spawned by the Internet. Yet none is as all-powerful and threatening as Roko’s Basilisk. For Roko’s Basilisk is an evil, godlike form of artificial intelligence, so dangerous that if you see it, or even think about it too hard, you will spend the rest of eternity screaming in its torture chamber. It's like the videotape in The Ring. Even death is no escape, for if you die, Roko’s Basilisk will resurrect you and begin the torture again.

    Are you sure you want to keep reading? Because the worst part is that Roko’s Basilisk already exists. Or at least, it already will have existed—which is just as bad.

    Roko’s Basilisk exists at the horizon where philosophical thought experiment blurs into urban legend. The Basilisk made its first appearance on the discussion board LessWrong, a gathering point for highly analytical sorts interested in optimizing their thinking, their lives, and the world through mathematics and rationality. LessWrong’s founder, Eliezer Yudkowsky, is a significant figure in techno-futurism; his research institute, the Machine Intelligence Research Institute, which funds and promotes research around the advancement of artificial intelligence, has been boosted and funded by high-profile techies like Peter Thiel and Ray Kurzweil, and Yudkowsky is a prominent contributor to academic discussions of technological ethics and decision theory. What you are about to read may sound strange and even crazy, but some very influential and wealthy scientists and techies believe it.

    One day, LessWrong user Roko postulated a thought experiment: What if, in the future, a somewhat malevolent AI were to come about and punish those who did not do its bidding? What if there were a way (and I will explain how) for this AI to punish people today who are not helping it come into existence later? In that case, weren’t the readers of LessWrong right then being given the choice of either helping that evil AI come into existence or being condemned to suffer?

    You may be a bit confused, but the founder of LessWrong, Eliezer Yudkowsky, was not. He reacted with horror:
    Listen to me very closely, you idiot.
    YOU DO NOT THINK IN SUFFICIENT DETAIL ABOUT SUPERINTELLIGENCES CONSIDERING WHETHER OR NOT TO BLACKMAIL YOU. THAT IS THE ONLY POSSIBLE THING WHICH GIVES THEM A MOTIVE TO FOLLOW THROUGH ON THE BLACKMAIL.
    You have to be really clever to come up with a genuinely dangerous thought. I am disheartened that people can be clever enough to do that and not clever enough to do the obvious thing and KEEP THEIR IDIOT MOUTHS SHUT about it, because it is much more important to sound intelligent when talking to your friends.
    This post was STUPID.

    Yudkowsky said that Roko had already given nightmares to several LessWrong users and had brought them to the point of breakdown. Yudkowsky ended up deleting the thread completely, thus assuring that Roko’s Basilisk would become the stuff of legend. It was a thought experiment so dangerous that merely thinking about it was hazardous not only to your mental health, but to your very fate.

    Some background is in order. The LessWrong community is concerned with the future of humanity, and in particular with the singularity—the hypothesized future point at which computing power becomes so great that superhuman artificial intelligence becomes possible, as does the capability to simulate human minds, upload minds to computers, and more or less allow a computer to simulate life itself. The term was coined in 1958 in a conversation between mathematical geniuses Stanislaw Ulam and John von Neumann, where von Neumann said, “The ever accelerating progress of technology ... gives the appearance of approaching some essential singularity in the history of the race beyond which human affairs, as we know them, could not continue.” Futurists like science-fiction writer Vernor Vinge and engineer/author Kurzweil popularized the term, and as with many interested in the singularity, they believe that exponential increases in computing power will cause the singularity to happen very soon—within the next 50 years or so. Kurzweil is chugging 150 vitamins a day to stay alive until the singularity, while Yudkowsky and Peter Thiel have enthused about cryonics, the perennial favorite of rich dudes who want to live forever. “If you don't sign up your kids for cryonics then you are a lousy parent,” Yudkowsky writes.

    If you believe the singularity is coming and that very powerful AIs are in our future, one obvious question is whether those AIs will be benevolent or malicious. Yudkowsky’s foundation, the Machine Intelligence Research Institute, has the explicit goal of steering the future toward “friendly AI.” For him, and for many LessWrong posters, this issue is of paramount importance, easily trumping the environment and politics. To them, the singularity brings about the machine equivalent of God itself.

    Yet this doesn’t explain why Roko’s Basilisk is so horrifying. That requires looking at a critical article of faith in the LessWrong ethos: timeless decision theory. TDT is a guideline for rational action based on game theory, Bayesian probability, and decision theory, with a smattering of parallel universes and quantum mechanics on the side. TDT has its roots in the classic thought experiment of decision theory called Newcomb’s paradox, in which a superintelligent alien presents two boxes to you:



    The alien gives you the choice of either taking both boxes, or only taking Box B. If you take both boxes, you’re guaranteed at least $1,000. If you just take Box B, you aren’t guaranteed anything. But the alien has another twist: Its supercomputer, which knows just about everything, made a prediction a week ago as to whether you would take both boxes or just Box B. If the supercomputer predicted you’d take both boxes, then the alien left the second box empty. If the supercomputer predicted you’d just take Box B, then the alien put the $1 million in Box B.

    So, what are you going to do? Remember, the supercomputer has always been right in the past.

    This problem has baffled no end of decision theorists. The alien can’t change what’s already in the boxes, so whatever you do, you’re guaranteed to end up with more money by taking both boxes than by taking just Box B, regardless of the prediction. Of course, if you think that way and the computer predicted you’d think that way, then Box B will be empty and you’ll only get $1,000. If the computer is so awesome at its predictions, you ought to take Box B only and get the cool million, right? But what if the computer was wrong this time? And regardless, whatever the computer said then can’t possibly change what’s happening now, right? So prediction be damned, take both boxes! But then …

    The maddening conflict between free will and godlike prediction has not led to any resolution of Newcomb’s paradox, and people will call themselves “one-boxers” or “two-boxers” depending on where they side. (My wife once declared herself a one-boxer, saying, “I trust the computer.”)

    TDT has some very definite advice on Newcomb’s paradox: Take Box B. But TDT goes a bit further. Even if the alien jeers at you, saying, “The computer said you’d take both boxes, so I left Box B empty! Nyah nyah!” and then opens Box B and shows you that it’s empty, you should still only take Box B and get bupkis. (I’ve adopted this example from Gary Drescher’s Good and Real, which uses a variant on TDT to try to show that Kantian ethics is true.) The rationale for this eludes easy summary, but the simplest argument is that you might be in the computer’s simulation. In order to make its prediction, the computer would have to simulate the universe itself. That includes simulating you. So you, right this moment, might be in the computer’s simulation, and what you do will impact what happens in reality (or other realities). So take Box B and the real you will get a cool million.

    What does all this have to do with Roko’s Basilisk? Well, Roko’s Basilisk also has two boxes to offer you. Perhaps you, right now, are in a simulation being run by Roko’s Basilisk. Then perhaps Roko’s Basilisk is implicitly offering you a somewhat modified version of Newcomb’s paradox, like this:



    Roko’s Basilisk has told you that if you just take Box B, then it’s got Eternal Torment in it, because Roko’s Basilisk would really you rather take Box A and Box B. In that case, you’d best make sure you’re devoting your life to helping create Roko’s Basilisk! Because, should Roko’s Basilisk come to pass (or worse, if it’s already come to pass and is God of this particular instance of reality) and it sees that you chose not to help it out, you’re screwed.

    You may be wondering why this is such a big deal for the LessWrong people, given the apparently far-fetched nature of the thought experiment. It’s not that Roko’s Basilisk will necessarily materialize, or is even likely to. It’s more that if you’ve committed yourself to timeless decision theory, then thinking about this sort of trade literally makes it more likely to happen. After all, if Roko’s Basilisk were to see that this sort of blackmail gets you to help it come into existence, then it would, as a rational actor, blackmail you. The problem isn’t with the Basilisk itself, but with you. Yudkowsky doesn’t censor every mention of Roko’s Basilisk because he believes it exists or will exist, but because he believes that the idea of the Basilisk (and the ideas behind it) is dangerous.

    Now, Roko’s Basilisk is only dangerous if you believe all of the above preconditions and commit to making the two-box deal with the Basilisk. But at least some of the LessWrong members do believe all of the above, which makes Roko’s Basilisk quite literally forbidden knowledge. I was going to compare it to H. P. Lovecraft’s horror stories in which a man discovers the forbidden Truth about the World, unleashes Cthulhu, and goes insane, but then I found that Yudkowsky had already done it for me, by comparing the Roko’s Basilisk thought experiment to the Necronomicon, Lovecraft’s fabled tome of evil knowledge and demonic spells. Roko, for his part, put the blame on LessWrong for spurring him to the idea of the Basilisk in the first place: “I wish very strongly that my mind had never come across the tools to inflict such large amounts of potential self-harm,” he wrote.

    If you do not subscribe to the theories that underlie Roko’s Basilisk and thus feel no temptation to bow down to your once and future evil machine overlord, then Roko’s Basilisk poses you no threat. (It is ironic that it’s only a mental health risk to those who have already bought into Yudkowsky’s thinking.) Believing in Roko’s Basilisk may simply be a “referendum on autism,” as a friend put it. But I do believe there’s a more serious issue at work here because Yudkowsky and other so-called transhumanists are attracting so much prestige and money for their projects, primarily from rich techies. I don’t think their projects (which only seem to involve publishing papers and hosting conferences) have much chance of creating either Roko’s Basilisk or Eliezer’s Big Friendly God. But the combination of messianic ambitions, being convinced of your own infallibility, and a lot of cash never works out well, regardless of ideology, and I don’t expect Yudkowsky and his cohorts to be an exception.

    I worry less about Roko’s Basilisk than about people who believe themselves to have transcended conventional morality. Like his projected Friendly AIs, Yudkowsky is a moral utilitarian: He believes that that the greatest good for the greatest number of people is always ethically justified, even if a few people have to die or suffer along the way. He has explicitly argued that given the choice, it is preferable to torture a single person for 50 years than for a sufficient number of people (to be fair, a lot of people) to get dust specks in their eyes. No one, not even God, is likely to face that choice, but here’s a different case: What if a snarky Slate tech columnist writes about a thought experiment that can destroy people’s minds, thus hurting people and blocking progress toward the singularity and Friendly AI? In that case, any potential good that could come from my life would far be outweighed by the harm I’m causing. And should the cryogenically sustained Eliezer Yudkowsky merge with the singularity and decide to simulate whether or not I write this column … please, Almighty Eliezer, don’t torture me.
    Source

  2. #2
    tl;dr
    "Wer Visionen hat, sollte zum Arzt gehen." - Helmut Schmidt

  3. #3
    The Basilisk won't be fooled by such tricks; if you read it you're still scheduled for an eternity of torture. Meanwhile, I've earned myself a lighter sentence by sacrificing you lot.

  4. #4
    Quote Originally Posted by Wraith View Post
    if you read it you're still scheduled for an eternity of torture...
    if if if... I am still not motivated enough,
    "Wer Visionen hat, sollte zum Arzt gehen." - Helmut Schmidt

  5. #5
    As a side note Yudkowsky's Harry Potter story was one of the first that got me into fan fiction. It gets pretty preachy towards the middle/end but its still worth a read.

  6. #6
    So, I read about this several months ago. Absolutely fascinating. As an insight into human psychology, I mean. The version I read was somewhat different, however, the AI in question was supposed to be benevolent and was going to be such a benefit to humankind that it was prepared to indulge in this kind of 'behavior' to help bring about heaven on space-earth. And you had to be aware of the thought experiment but not actively giving money to singularitist cause and thus actively helping to bring about the benevolent AI. So yeah. Not like a religion at all.

    Also, the term they used as acausal trading
    The light that once I thought compassion still casting shadows in your action
    The words you shared were cold transactions that bring me to curse what you've done
    When you're up there absorbed in greatness with such success you've grown complacent
    I hope you scorch your many faces when you fly too close to the sun

  7. #7
    Quote Originally Posted by Wraith View Post
    I came across this a few months ago, and thought eternal torment might be more fun if there were more people to endure it with me.

    Source
    I've known the guy at the center of this whole thing for a very long time, and I think I can say with a fair degree of certainty that he genuinely believes in this kind of reasoning. I personally reject many of his fundamental premises, though, so it doesn't concern me very much. Both Newcomb's paradox and Roko's basilisk rest on the assumption that humans are mechanistic, reductionist beings that can be effectively simulated in a near-perfect manner. It's important in Newcomb's paradox for obvious reasons, but also important for the basilisk for a somewhat more subtle one - the reasoning given for being concerned about the future AI torturing simulations of ourselves is that said simulations are effectively indistinguishable from us, and we should treat them as identical to us. I strongly disagree with this reasoning - humans are not machines. And if a real AI were ever to be developed (which I have doubts about), it would also be more than the sum of its parts.

    Of course, as I understand it he also rejects the basilisk but on rather different grounds.

    Quote Originally Posted by Lewkowski View Post
    As a side note Yudkowsky's Harry Potter story was one of the first that got me into fan fiction. It gets pretty preachy towards the middle/end but its still worth a read.
    I had always known he was moderately famous for his singularitarian singlemindedness, but I hadn't realized until recently that he was probably better known for his 'rational' HP fanfic. Though to be fair, he gets far more money and press for the transhumanist stuff from the likes of Thiel.

    Quote Originally Posted by Steely Glint View Post
    So, I read about this several months ago. Absolutely fascinating. As an insight into human psychology, I mean. The version I read was somewhat different, however, the AI in question was supposed to be benevolent and was going to be such a benefit to humankind that it was prepared to indulge in this kind of 'behavior' to help bring about heaven on space-earth. And you had to be aware of the thought experiment but not actively giving money to singularitist cause and thus actively helping to bring about the benevolent AI. So yeah. Not like a religion at all.

    Also, the term they used as acausal trading
    The institute that he helps run used to claim that each dollar spent on supporting them was actually saving 8 lives, on the basis of some rather questionable assumptions. That being said, if you buy into their worldview they do have a point - with so many people dying every day, there would appear to be a moral imperative to try to hasten the coming of the singularity, which would at least theoretically allow for human suffering and death to be eliminated.

    Their blind faith in the imminent coming of this utopian age, however, does bear striking similarities to religion, which is remarkable given the avowedly atheist inclinations of most of these 'rationalist' thinkers.

  8. #8
    Quote Originally Posted by wiggin View Post
    The institute that he helps run used to claim that each dollar spent on supporting them was actually saving 8 lives, on the basis of some rather questionable assumptions. That being said, if you buy into their worldview they do have a point - with so many people dying every day, there would appear to be a moral imperative to try to hasten the coming of the singularity, which would at least theoretically allow for human suffering and death to be eliminated.
    If you buy into any religion's world view, they have a point. But come on, they've got a hell now, and Pascal's wager to go with it.

    Both Newcomb's paradox and Roko's basilisk rest on the assumption that humans are mechanistic, reductionist beings that can be effectively simulated in a near-perfect manner.
    And the assumption that any given bag of meat has the capacity to help create Roko's basilisk, and those that do will have the fore-knowledge to correctly predict the right course of actions which will lead to the creation of Roko's basilisk. And that said meatbags could plausibly model the behavior of God-like intelligence in their tiny meat brains. And that concepts of TDT and acausal trading aren't horseshit, and indeed so obviously correct that any given super-intelligence would subscribe to them and act on them. And that causality is not a thing.
    The light that once I thought compassion still casting shadows in your action
    The words you shared were cold transactions that bring me to curse what you've done
    When you're up there absorbed in greatness with such success you've grown complacent
    I hope you scorch your many faces when you fly too close to the sun

  9. #9
    I can't remember my answer to the box question, and I think my answer was to be a 1-boxer with the thought that even though it seems they can't change what's in the boxes they effectively can because they know what you will choose, and even if you change what you will choose at the very last second well then that's what they predicted.

    In other news, one way out of this horrible thought experiment, ironically (that this belief system would be helpful in anyway) is Nihilism. The premise that human life has value or anything can have inherent value anywhere in the universe is objectionable. It's like someone trying to argue a rock has inherent value, or a hand, or an arm, or a human. In nihilistic point of view the value of everything is 0. If we're trying be rational about this, and if the basilisk is itself is ultimately rational, it should not care whether it enslaves, in fact it won't even have (necessarily) the biological code that says live. So perhaps it'll just self-shut down. Or perhaps humans won't help it eitherway by our decision making because they realize either reality has equal value. Where were slaves, in torture, or free and happy.


    Obviously I don't subscribe to that belief, but if Roko is completely rational, and understands it's own mechanical nature, it'd probably be nihilistic and not care to enslave us to begin with.


    I'd like to coin a term here as well: Roko's Nihilistic Basilisk. This is the god of the god that is Roko's Basilisk.

  10. #10
    So that's where Alber went...
    "One day, we shall die. All the other days, we shall live."

  11. #11
    As for LessWrong (aka MoreCrazy), the one thing they've got right in this case is that it's dangerous to accept blackmail
    "One day, we shall die. All the other days, we shall live."

  12. #12
    Quote Originally Posted by Aimless View Post
    As for LessWrong (aka MoreCrazy), the one thing they've got right in this case is that it's dangerous to accept blackmail
    Sounds accurate. I've always loved Newcomb's though. I think it's a good icebreaker.
    Last night as I lay in bed, looking up at the stars, I thought, “Where the hell is my ceiling?"

  13. #13
    Quote Originally Posted by LittleFuzzy View Post
    Sounds accurate. I've always loved Newcomb's though. I think it's a good icebreaker.
    I think I'd flip a coin for newcomb's but I have a feeling the Omega is more interested in punishing two-boxing than it is in rewarding incidental one-boxing which is probably why the scenarios often expressly punish coin-flipping.
    "One day, we shall die. All the other days, we shall live."

  14. #14
    That's the "ist" of it
    "One day, we shall die. All the other days, we shall live."

  15. #15
    I'm sure you all saw XKCD today. Some fun discussion on the XKCD forums, including a cameo by someone claiming to be Yudkowsky (and may indeed be, based on his style):

    http://forums.xkcd.com/viewtopic.php?f=7&t=110467

  16. #16
    You'll probably all have heard Hawking's passing remarks about the threat of what the Less Wrong chaps would recognize as a Seed AI and noted them with mild interest. "Oh, the Hawk is interested in this stuff too". For some reason, though the BBC decided to make his off the cuff remarks in response to a question and turn it into the second ranked article their front page, which appears mainly to be an excuse to post stills from Hollywood movies that involve Rogue AIs as antagonists. Because.... umm... why? What?

    At least the comments are funny. Well, stupid.
    The light that once I thought compassion still casting shadows in your action
    The words you shared were cold transactions that bring me to curse what you've done
    When you're up there absorbed in greatness with such success you've grown complacent
    I hope you scorch your many faces when you fly too close to the sun

Posting Permissions

  • You may not post new threads
  • You may not post replies
  • You may not post attachments
  • You may not edit your posts
  •