• Manage Membership
  • Log In

The Partially Examined Life Philosophy Podcast

A Philosophy Podcast and Philosophy Blog

Subscribe on Android Spotify audible patreon
  • Home
  • Podcast
    • PEL Network Episodes
    • Publicly Available PEL Episodes
    • Paywalled and Ad-Free Episodes
    • PEL Episodes by Topic
    • Nightcap
    • Closereads
    • Philosophy vs. Improv
    • Pretty Much Pop
    • Nakedly Examined Music
    • (sub)Text
    • Phi Fic Podcast
    • Combat & Classics
    • Constellary Tales
  • Blog
  • About
    • PEL FAQ
    • Meet PEL
    • About Closereads
    • About Pretty Much Pop
    • Philosophy vs. Improv
    • Nakedly Examined Music
    • Meet Phi Fic
    • Listener Feedback
    • Links
  • Join
    • Become a Citizen
    • Join Our Mailing List
    • Log In
  • Donate
  • Store
    • Episodes
    • Swag
    • Everything Else
    • Cart
    • Checkout
    • My Account
  • Contact
  • Mailing List
  • Book

Episode 108: Dangers of A.I. with Guest Nick Bostrom

January 6, 2015 by Mark Linsenmayer 48 Comments

http://www.podtrac.com/pts/redirect.mp3/traffic.libsyn.com/partiallyexaminedlife/PEL_ep_108_12-7-14.mp3

Podcast: Play in new window | Download (Duration: 1:42:56 — 94.3MB)

On Superintelligence: Paths, Dangers, Strategies (2014) with author Nick Bostrom, a philosophy professor at Oxford. Just grant the hypothetical that machine intelligence advances will eventually produce a machine capable of further improving itself, and becoming much smarter than we are. Put aside the question of whether such a being could in principle be conscious or self-conscious or have a soul or whatever. None of those are necessary for it to be capable, say, of developing and manufacturing a trillion nanobots which it could then use to remake the earth. Bostrom thinks that we can make some predictions about the motivations of such a being, whatever goals it’s programmed to achieve, e.g. its goals will entail that it won’t want those goals changed by us. This sets up challenges for us in advance to figure out ways to frame and implement motivational programming an A.I. before it’s smart enough to resist future changes. Can we in effect tell the A.I. to figure out and do whatever we would ask it to do if we were better informed and wiser? Can we offload philosophical thought to such a superior intelligence in this way? Bostrom thinks that philosophers are in a great position for well-informed speculation on topics like this. Mark, Dylan, and Nick are also joined by former philosophy podcaster Luke Muehlhauser. Read more about the topic and get the book. End song: “Volcano,” by Mark Linsenmayer, recorded in 1992 and released on the album Spanish Armada: Songs of Love and Related Neuroses. Please support the podcast by becoming a PEL Citizen or making a donation. The Bostrom picture is by Genevieve Arnold.

Transcript

0:08 Mark: You're listening to the Partially Examined Life, a philosophy podcast by some guys who at one point set on doing philosophy for a living, but then thought better of it. Our Questions for episode 108 are something like, how might artificial minds differ from ours? And how can we keep them from killing us? And we'll be Discussing Nick Bostrom's 2014 book, Superintelligence: Paths, Dangers, Strategies with the author. You can join the discussion, buy the book, and get lots More information@PartiallyExaminedLife.com this is Mark Linsenmayer speaking to you from Madison, Wisconsin.

0:35 Dylan: This is Dylan Casey speaking to you from Middleton, Wisconsin.

0:38 Luke: This is Luke Muehlhauser speaking with you guys from Berkeley, California.

0:44 Mark: All right, so Nick Bostrom will be joining us in just a little while, but since we have two guests, we thought we'd do a little pre show with Luke. So, Luke, I had heard of you because you did Conversation from the Pale Blue Dot, a very fine podcast that was running about the same time that we started up, but then that petered out and you decided to put your time into working predominantly on the problem we're going to be talking about today, right?

1:07 Luke: Yeah, that's right. I'm now the CEO of the Machine Intelligence Research Institute in Berkeley, California, and our primary focus is the subject of Nick's book on superintelligence, which is how do we get good outcomes from AI systems that are more clever and generally capable than humans are?

1:26 Mark: And so I know the focus of your podcast before that started out as theology. Right. You had a common sense atheism blog, and then the podcast grew out of that and then you switched to going very systematically through ethical theories. Right?

1:38 Luke: Yeah. I think the two major topics of my old podcast conversations from the Pale Blue Dot was theism, atheism arguments and debates, and then metaethics, just because those were the two things I was most interested in at the time. Right.

1:52 Mark: And of course very highly connected. The claim that if you don't have God, then you can't have ethics. You're a heartless bastard baby eater even.

2:03 Luke: I love me some babies, much like a super intelligence.

2:07 Mark: That's obviously the problem that they don't have. You can't program Christianity.

2:10 Luke: They read too much Nietzsche is the problem with the AIs. They're all nihilists.

2:15 Mark: So what made you do the job? You interviewed Eliezer Yudkowsky as part of the podcast, or you ran into him at some point and then he was working on this kind of stuff and tell me more of that story.

2:26 Luke: Right. Well, a long term interest of mine as well has been the ways that our brains do not allow us to think reasonably about the world in many cases. And so I was just doing a lot of online reading, as well as, you know, in textbooks and so on about that subject. And as I was looking at that material, I came across another set of material on what are called existential risks, Risks that might cause human extinction, especially from future technologies. The natural risks we've already been surviving for millions of years, so presumably they won't wipe us out in the next hundred.

And so that led me to some interesting discussions of artificial intelligence or nanotechnology or synthetic biology, pandemics, that kind of stuff. And after I had been reading through that literature for a while, it seemed to me like the AI concern was most urgent and maybe had the fewest people working on it. So I got closer to the people who were working on that issue.

3:24 Mark: Can you tell us a little about the institute? It looks like you've got a. Support a number of researchers and have pledge drives, and it looks very organized, and it's all seemingly just run from enthusiasts. Tell me more about that story.

3:37 Luke: That's right. So I think there's some ridiculous fraction of our supporters, you know, our donors, who are either computer engineers or computer scientists or mathematicians, maybe something like 50%. And so it is really a group of people who are very enthusiastic about this mission of ensuring that smarter than Human intelligence does good things rather than bad things. And we've technically been around since 2000, when Eliezer Yudkowsky founded the institute. But for the first three or four years of that, that meant that somebody was putting up Eliezer in their basement while he wrote stuff about AI. And so since then, we have grown into a much more substantial organization with a budget of about a million and a half per year. And we have researchers and we host workshops and we publish papers and conferences and so on.

4:28 Mark: So Nick is some kind of advisor to the group?

4:31 Luke: Yeah, he's a research advisor and also just a very frequent collaborator. So Nick Bostrom directs the Future of Humanity Institute at Oxford. And they're the group that we collaborate with the most often, just because we have this shared interest about existential risk from future technologies, especially AI. And so we've written papers together, run workshops together, et cetera.

4:54 Dylan: One thing that occurs to me, especially when you say, Luke, that some huge portion of your donors and supporters are people active in pretty high levels of computer programming and technology and so forth. The thing that occurred to me immediately, so 0.1% of the human population kind of thing. So it made me want to understand or have you say a little bit more about the argument for being concerned about existential risk, especially why AI presents a peculiar kind of existential risk compared to, I mean, the big one right now that we hear in the news is climate change, which arguably isn't existential risk strictly speaking, certainly not in the terms that you guys talk about it.

It's more like Prof. Change risk. And there are other ones that are creeping along through which for long periods of history, especially technology, there have always been significant critics of how this or that technology is going to ruin our lives or change our ethics, or make it so that humanity is not recognizable in one way or another. And that's been even in some sort of pre modern eras. Looking through your website for your organization and also reading Nick's book and a little bit on the Future Humanity Institute, it became clear to me that one general motivation is the sense that the kind of risk we're talking about now from your general perspective, is one of many orders of magnitude worse than the typical kinds of existential risks we might think of.

And in fact, maybe so much more so that you wouldn't even consider the others genuine existential risks. War, nuclear weapons, climate change, none of that's really an existential risk. So I would be interested to hear more just about sort of the thinking going on behind being worried about this kind of problem.

6:46 Mark: So tell us sort of your perspective a little more about, tell us about your faith journey of getting transitioning to this.

6:53 Luke: Right. Well, as you guys probably know, I was raised in a Christian home and then later had this deconversion process where I became a sort of standard scientific naturalist. And that left me with the question of, well, what is it that's valuable to do in the world? And I ended up with some kind of view that's consequentialist. And then that sort of led to questions about, okay, well, given some kind of consequentialism, what are the most valuable things to be doing? And if you have the sort of standard physicist view that future people are as real as present people, almost all of the good that you could be doing is influencing the quality of the lives that are being lived by people in the future, especially as our civilization expands to other planets.

And so then I started looking at things that could potentially impact the trillions and trillions of people in the future. And the ones that we can identify now are basically this class of existential risks that could cause it to be the case that none of those people come into existence. Luckily, that's a relatively short list. So it's difficult to do it with climate change. You know, the IPCC reports don't think that a Venus like runaway greenhouse gas situation is at all plausible. Nuclear weapons, it's possible, but still actually very difficult to literally kill everyone with nukes, even with the ensuing nuclear winter.

And so we're really talking about either astronomical events that we know happen very rarely, like very large asteroids or nearby supernova, or we're talking about risks that we have no track record of surviving from technologies that are even more powerful than nuclear weapons. So risks from synthetic biology, advanced molecular nanotechnology, and advanced artificial intelligence.

8:32 Dylan: So basically the question here is, how do we avoid becoming like dinosaurs?

8:37 Luke: Yeah, that's right. How do we let there be continued discovery and happiness and flourishing and all these things that we love? How do we make sure that continues for billions, possibly trillions of years as opposed to being snuffed out here in the 21st or maybe 22nd century?

8:54 Mark: Well, it's also an exciting sort of niche in terms of the publicity gap that people are not talking about this as you're saying that there aren't that many people working in this field.

9:02 Luke: That was really the driving reason for me to get into AI as it just seemed like that was an area where if I started to contribute whatever I can to that cause I might actually notice a difference.

9:13 Mark: All right, it's time to add Nick. I'm adding him right now. Hello, Nick.

9:17 Nick: Hey, I'm here.

9:17 Luke: Hey, Nick.

9:18 Mark: Welcome, Nick. Thanks so much for joining us.

9:21 Nick: Thanks for inviting me.

9:22 Mark: And where are you calling in from this time?

9:23 Nick: Oxford, uk. All right.

9:25 Mark: And you just finished a big book tour for this, right?

9:28 Nick: Yeah, well, I mean, a couple of months ago, really just two weeks in America and then a number of other scattered talks here and there. But I'm glad to be back here to do actual research now.

9:38 Mark: Yes, well, I saw it looks like the book has been doing pretty well, at least. You were on NPR and mentioned on cnn. And when I was looking online, it looks like it's getting some press.

9:46 Nick: Yeah, for an academic book, it's been quite remarkable. It even cracked the lower rungs of the New York Times bestseller list, briefly.

9:54 Dylan: That's great.

9:54 Mark: Very nice. So, I mean, I'm familiar with your Future of Humanity Institute and some of your past works on existential risks, and we were just talking before you jumped on about why artificial intelligence as opposed to some of these other risks, why is it more dangerous than climate change? You know, these other things. So I think we've got that more or less established, but maybe tell us a little bit more about your institute and sort of why as a philosopher you're diving into this as opposed to something else. Obviously I understand as a practical matter that these are pressing concerns that you could actually do some good in the world as opposed to writing another book on Heidegger or something like that.

10:28 Nick: Yeah. So the Future of Humanity Institute. We are a multidisciplinary research center here we have philosophers, mathematicians and scientists working together, trying to think carefully about big picture questions for humanity. And the reason for focusing on these kinds of things, such as superintelligence and other existential risks, is to do as much good as possible. But I also have this sense of what I call philosophy with a deadline, in that if it is really the case that we will at some point, if things go well, replace the current cast of humans with something that is cognitively more capable, that is either super intelligent machines or maybe cognitively enhanced humans, then our best efforts at doing basic philosophy now might soon become obsolete because there will be other people who will do that much better than we do.

And so that what we should focus on are trying to solve those problems that we need to solve between now and the arrival of these more competent philosophers of the future. Right.

11:19 Mark: We had just brought up this paper that you had written in 1997, predictions from philosophy, which kind of sketched a new role that a philosopher could play as a supercharged high level science synthesizer, which then would enable you to make predictions, perhaps in a way that the scientists actually buried in day to day work, but also so philosophers trying to get enough of the actual information of what scientists are doing to contribute usefully as an outsider to the progress and interdisciplinary work.

11:50 Nick: That's a very old paper, but I guess I have continued somewhat in that vein. I tend not to make a very sharp separation between philosophy and other forms of intellectual inquiry. To me, philosophy and sciences kind of overlap and generally tend to be interested in that area of overlap. But without trying to very carefully tabulate what parts of what I do is philosophy and what part is futurism and what part is science.

12:13 Luke: Sure, Nick, One thing I've always wondered is what fraction of the philosophers that you talk to think that what you're doing is legitimate philosophy, as opposed to some kind of weird philosophy science hybrid they don't want to talk about or respond to or anything.

12:28 Nick: Well, a very high fraction of the philosophers that I talked. Obviously there's a selection effect here. Like I tend to hang around people who are interested in what I'm doing or that I am interested in what they are doing. So we have a kind of a really exciting intellectual community here around the Future of Humanity Institute, So Luke's institute over there in Berkeley, a non profit machine intelligence research institute, also has some really interesting people in this field and then kind of various scattered individuals forming a little local intellectual ecology of like minded folk.

It also connects quite closely to the effective altruism movement and some of the people there. Now, I'm sure there are other branches of philosophy where maybe people would not so much be hostile, but just uninterested in this kind of stuff because it doesn't necessarily connect with the kinds of things that they are trying to do.

13:17 Mark: Well, I was surprised at the amount of overlap that I saw in some of the very practical problems. So the central problem here is how to program a future AI that through its superintelligence could potentially have a lot of catastrophic power. How do you program it with something approximating human values, at least enough so that it won't kill us. And that requires doing a lot of the same stuff that philosophers have been trying to do in ethics for many, many, many years. You know that at least before Wittgenstein, people thought that you could come up with a list of necessary and sufficient conditions. So basically a programmable specified list or criterion for which actions morally permissible or not. Since Wittgenstein and folks following him, it seems like we've given up on that in philosophy. But it's necessary if you want to try to actually transmit this stuff into a computer.

14:04 Nick: Well, I hope it's not necessary in as much as it seems unlikely we will figure all of that out within the next 50 years or 100 years, given that people have been at it for 2000 years. At least there is a new option available when one is thinking of doing this in the context of AI that is not so available in traditional moral theorizing, which is to try to pump on a lot of these questions and delegate them to the artificial intelligence, which is by hypothesis would be smarter than we are and therefore maybe better able at solving the same philosophical problems that we are struggling with now. There are still some questions that we would have to answer ourselves in order to know how to build the AI. But at least some of the questions we might be able to say we don't have to answer, that we can delegate to the AI to figure out how to answer them.

14:49 Luke: There is this nice quote I like from Dan Dennett, where at a conference in 2006 he said, Artificial intelligence makes philosophy honest. Which I take to mean that when you're doing philosophical work in many cases, or ethical work in a journal, a lot of it is kind of fuzzy English sentence constructions. And it's easy to fool ourselves into thinking that we've come up with a theory that actually works. But when you have to write it in code or in math, then you have to be very precise and you can't end up fooling yourself that way. And I think that's a nice way to think about it.

I mean, there are a lot of nu there. And I don't want to dismiss a lot of the ethical work that's happening outside the context of AI. It is sort of a more difficult challenge to tell an AI what it is that you think a good world looks like than it is to write a lot of English sentences that describe to you what a good world looks like.

15:43 Dylan: And part of the problem, the terms of the AI challenge, the urgency, seem to be focused on it being a kind of infinity problem. That the scales are so large, both in power and in time, that small deviations cause big challenges and extrapolations that you might in other contexts dismiss because you say, well, if so, and so is an extremely evil character. Well, eventually, what's the timescale on that? It's, you know, 50 years they're going to die. Right. Or some tyrant has a tremendous amount of power. There's a lot of self limiting biology and so forth going on that is thrown out the door in the AI case where you have to consider these really long term, very powerful entities and structures that won't be self limiting in a way that we're used to.

16:40 Nick: Yes. And the need to get it right on the first attempt.

16:44 Dylan: Well, yeah. And so that's a corollary, right, that. Yeah. You don't have the revision possibility. You can't say, well, you know, we had the Persians and then we had the Greeks, then we had the Romans and the Chinese dynasties. You don't have that iterative process in which in fact you can just be content to repeat all kinds of mistakes and so forth.

17:01 Nick: Yeah. So this adds to the challenge. It's rare that we humans get something very difficult and complex right. On the first attempt to do it. But it might be that with AI we will have to do that because once you have a very powerful superintelligence, it might prevent you from trying again. If it's the wrong kind of superintelligence,

17:19 Dylan: it's highly likely that the people who usually listen to Our podcast won't be intimately familiar with the term super intelligence and won't be intimately familiar with the different possibilities in AI. While I think they're going to be most interested in some of the stuff we mentioned at the beginning about value loading and. And that I think it'd be worthwhile having Nick and Luke sketch out the terms of superintelligence and what kind of entity we're talking about and what the salient problems are and how they're different from other problems we're used to encountering.

17:54 Nick: So by superintelligence, I mean any intellect that radically outperforms humanity in all practically relevant fields. I'm focusing on instrumental rationality here. That is the means, ends, reasoning, the ability to create plans that, with high probability, lead to you achieving your goals,

18:14 Mark: which is sort of irrelevant to whether it's actually a mind or not. Right. If it passes the Turing test, and moreover passes all these other operational tests, well, then at least it's something that is not only a useful tool, such that researchers are trying to generate this kind of thing, but it's something that we have to worry about whether or not it has a moral status.

18:33 Nick: That's correct, yeah. For most of the book, I don't focus on the question of whether these machines would be conscious or whether they would have free will, or whether they would be truly rational in the full, thick sense of rational, or just very instrumentally competent. So in principle, this definition of superintelligence leaves open the question of how it is instantiated. You could have some kind of biological intelligence that would satisfy these criteria in practice. I focus for most of the book on machine superintelligence because I think that ultimately the potential for information processing in machine substrate is just vastly greater than in biological substrate.

Although I do have a section where I consider possibilities for enhancement of biological cognition. I think this will become feasible first through genetic selection or genetic engineering. There are also other ways of achieving superintelligence through aggregating many different lesser intelligences. So we can talk about collective superintelligence. Scientific communities often can do things that no individual humans can do. And we can think of humanity's collective intelligence as, in some sense quite significantly outstripping the intelligence of any individual.

19:43 Mark: That's actually another really good reason for not focusing on this. Is it truly conscious or something? Because you can talk about minds at these different levels of organization just as a functional thing. If there's a functional entity, even if it's a society that has an output that we can deal with it for practical Purposes as a unit, then, well, why not call that an intelligent system or a mind or something, and just not worry about the traditional philosophical issues surrounding that.

20:06 Nick: The place where questions about consciousness become relevant is when one is trying to do some kind of ethical evaluation of different possible outcomes. If the outcome is that there are all these digital minds, it might matter a lot whether those digital minds are conscious, and if so, what they experience. But at the first stage of analysis, when one is just trying to figure out what might happen and how likely different possible outcomes are, one can largely bracket that question.

20:30 Dylan: The thing that comes to mind that's most familiar in experience would be something like just cultural evolution or cultural development. The way in which a scientist would say that their work stands on the shoulders of those previous to them. That kind of progress, that kind of building of collective intelligence is just orders of magnitude smaller than what we're talking about here.

20:51 Nick: Partly it's just a matter of how we choose to define the words here. But yeah, I do have in mind by superintelligence something that is quite radically superior to even all of humanity as it is currently configured. Superintelligence, at least in its mature phase, I think of as being able to do, for example, scientific and technological research much more rapidly than the current scientific community can do. Also, I imagine it to be very good at planning and strategizing and devise ways to realize its goals. We can kind of place a lower bound on the capabilities of such a system by thinking about what the really capable human being could do.

And then we can sort of say, well, superintelligence could probably do a lot more than that, but at least it could sort of do what a really capable human could do. And perhaps if we imagine that human to be speeded up and running a million times faster than a normal biological human.

21:41 Dylan: When we talk about these kinds of enhancements, the obvious ones are cognitive parallels. Being able to calculate, being able to reason, being able to plan those sorts of things, that'll get you a long way. But in terms of affecting the world and even building further machine substrates for physical or computational based superintelligence, there's going to be also a lot of physical boundaries and practical aspects that have traditionally in technology been mediated by human biology to manufacture things, you don't take it up in the book. But it did make me wonder if the discussion doesn't assume that those robotic constraints and the development of systems that would manipulate the physical world in a intricate manner at at least as high a level as human beings can, but likely even more is just granted, as that's also just going to be part and parcel to a superintelligence.

22:35 Mark: One of the things that we are looking at is not just cognitive enhancements, but sort of emotional enhancements. And you talk about that aspect of that in this book in terms of the ability of computers and advanced AI to convince people to do what it wants. So it's not even just a super fast processing, memory rich entity, but it's something that potentially could develop skills like being super savvy and being able to manipulate a political environment. And so it might not even need robot arms to do its building. It could use people.

23:07 Luke: Yeah, yeah.

23:08 Nick: So if you think about the most powerful people today, like they are not people who are powerful because they build things with their hands directly. It's that they talk to other people and then they have a big effect. I'm imagining that even if an AI started out without much by way of a body, maybe just the ability to output words on a screen, this device, nevertheless, if it really were super intelligent inside, could well quickly become extremely powerful, either by having humans serve as its accessories, or at the more mature stage, develop its own robotic infrastructure using molecular nanotechnology to gain the ability to control the arrangement of matter.

But I focus more on the brain part of the whole thing than the physical manipulators, because I think the physical manipulators is a trivially easy problem compared to the difficulty of actually making an AI that is generally intelligent.

23:59 Mark: One of the familiar philosophical objections to a generally intelligent AI is from Hubert Dreyfus sort of channeling Heidegger and Merleau-Ponty, which I know he was reacting more. So he wrote essays like our books about what computers can't do. And I know he was arguing largely against the over promising of AI researchers in the 1960s and things, but it seems like the challenges that he presented. So the idea is that if you look at a computer as just following an explicit program, a set of symbols, one after another, commands, and you try to think, well, how could we have real intelligence reflected by that?

So Dreyfus argued, a lot of our knowledge is not formalized in that way, that we operate with a good deal of background knowledge, that we use intuition, which is just a way of calling up a lot of subconscious sort of how to do something rather than facts, a list of step by step reasoning. So from what I understand and what I saw in this book is that you've kind of acknowledged that challenge. In fact, researchers have acknowledged that in pursuing not a strictly serial program, pursuing a Neural net model or something like that, where you've got a lot of different things working in parallel.

It might be a black box structure so that we don't even know as a programmer is how it is learning, essentially. And that also sets up the primary difficulty that you put forward that we already alluded to, of how do we explain to the superintelligence not to kill us? How do we explain human values? We have all these background beliefs that we pretty much throw a child into its normal learning environment and we trust that if there's not any major chemical catastrophes that we're not going to end up with a psychopath. But we have no reason to think that would be the case in this vastly different kind of mind.

25:35 Dylan: When you say telling machine intelligence, it makes me think of a kind of solved case where I have an AI that now I'm going to say, here's how you not kill us. It seems like it's a prior problem of programming structuring the AI from the beginning, rather than it being a post hoc issue. Well, I guess there's various solutions, but how it would end up with a morality that we could live with.

26:01 Luke: Well, I think we can take telling the AI what we want to be an open ended description of a process that might involve computer programming and might involve a ratification process where the AI has been programmed to do something that it thinks we want it to do, and then we ratify whether or not that plan is something we want executed. At this early stage, we're just really uncertain what the exact solution will look like. But I think that Dreyfus's concerns were really targeted at classic AI, or sometimes it's called good old fashioned AI that was about the logical manipulation of symbols.

Researchers, as you said, have recognized that for a while now, and there are now statistical approaches and there's different approaches to getting what we would call common sense into an AI and prior knowledge and this kind of stuff. If the human brain does it with information processing, which is basically all cognitive scientists think, then presumably we can do that kind of information processing in a computer, though we might do it by a very different method. Just like we didn't build planes by having them flap their wings.

27:01 Nick: There is a tendency among some philosophers to maybe correctly identify some common way in which things go wrong, or that people have a misconception in a certain direction. So when Dreyfus was doing this work, there was this common conception among AI researchers and the wider public that by just putting in a few more of these rules and classes and then Having the AI crank the deductive engine, you could get something really clever out of that. And he was kind of right in observing that they were vastly overconfident and optimistic about where that would lead.

But then what often happens is that having made this correct observation, the philosopher then tends to sort of overstate the objection and make it into a principled objection that is actually too strongly stated. Because that's the way you get the philosophically interesting claim. It's by having something that seemed like a principled objection with universal scope. Often when you get a false philosophical claim that then turns out not to apply. If one does something differently. In this case, if one uses maybe statistical machine learning techniques, you might not have a system that exhibits the same brittleness as some of these database systems, expert systems that Dreyfus was targeting.

28:09 Luke: You know, Nick, one related question that comes up a lot for me from people is why are you talking about an AI having motivations and desires and goals and values? And I think that is surprising to many people who are thinking of AI as maybe not that different from Microsoft Windows or something?

28:28 Nick: To some extent, it is still an open question what shape the first forms of superintelligence will take. So the idea that we should deliberately try to make them not to be like agents with values and desires, and is it kind of a live option to build tool AI instead? Nevertheless, if one looks at the frontier of AI today, a lot of it is focused on building smart agents. These are systems that have some function that describes a criterion for what they want to achieve, and then they have some way to model the world, and then something like a decision theory whereby they select actions based on how well they expect those actions to meet the criteria.

And this is a very general type of system to build. A lot of things can be more or less accurately viewed as agents. We often, even in political science, we sometimes think of states as agents. Why did this nation do this or that? Well, because it wanted to prevent some other nation from doing that and that. And it has goals and desires. And we know that this analogy breaks down at a finer level of granularity. But nevertheless, we often think of nations and corporations and groups, as well as individual humans and other animals, as having agency.

And it's a natural starting point for thinking about artificial systems as well. And there might even be reasons deliberately to try to make the machine intelligence into an agent. In as much as in an agent, there might be a particular place you can look and see what the criteria are that it uses to decide where to go. Something like A utility function or a goal. It might be nice if we can actually see what the system is trying to optimize rather than just having that kind of emerge as a side effect.

29:58 Mark: What was Dennett's term for looking at any system and sort of, even if it obviously doesn't have agency, treating it as such, just as a way of diagnosing the problem?

30:07 Nick: The intentional stance.

30:09 Mark: Yes, the intentional stance.

30:10 Nick: All right. And so, yes, I think he said that a lot of even quite simple systems can sometimes be usefully viewed as though they had goals and desires.

30:20 Dylan: That actually way of thinking about, whether it be mechanical systems or the natural world or solving problems is actually incredibly common, if not even, in fact, the de facto way of thinking about it. Whether you're talking about solving a physical problem and minimization of energy, saying, well, the atom wants to be in a lower energy state, or you throw the ball up in the air and it wants to be closer to the ground, or you're working on a complicated machine, whether it be your car or I do my day work on radiation therapy machines, and you have a complicated system with lots of different interacting parts, and you end up thinking about it teleologically, frankly, both in trying to fix it, you think about, well, what is it wanting to do in the sense of the system as it actually is, and what ought to be doing in that?

What did you intend it to do? And usually it's a conflict between those two things. That's the bug that's in the system that you programmed the firmware to do something and you made a mistake. And so now the system wants to turn the crank at a speed that's two times what you intended it to turn. So that way of thinking about mechanical and computational and systems that involves no actual volition in the way we would think about it seems perfectly natural and sensible, in fact, common to think about systems that way.

31:39 Nick: Yeah, and I think with artificial intelligences it can be true in more than just a metaphorical sense. But if one does adopt that intentional stance, which I think is almost necessary in this domain, it's still important then to try to avoid doing it in illegitimate ways by anthropomorphizing the systems. While it is often useful to think that there is some criterion and then some kind of optimization process that tries to steer the world so that it meets this criterion, its goal. We should be wary of kind of immediately importing all other psychological tendencies that we humans have, but that are maybe a contingent fact of our evolutionary history and the fact that that we've kind of evolved as social animals.

And we have feelings like pride and resentment and all of that. None of that needs to be present in an artificial agent. And in fact, it would be quite hard to engineer those elements into an artificial agent from the outset. It's something that we wouldn't know how to do.

32:35 Mark: So what we need is the intentional stance. Plus, like you were saying, it has a representation of the environment. It has a representation of its different goals and is flexible in the way it can pursue those goals. It has a representation of. Of its own position in relation to those goals and its environment. That once you get all this self representation and you put on top of that the intentional stance, then, well, you get basically the operational equivalent of consciousness. We don't want to load that up with more baggage than was put into it.

33:02 Dylan: Well, it's something like self action. Right. Something that without even calling it consciousness, it can move itself. I mean that in the most general way of moving.

33:11 Luke: And you might just say goal directed behavior. Yeah, because words like consciousness have all this other baggage that we attach to it. But it's really just having a picture of what the world looks like now. Having a picture of how you'd prefer the world to work and then making plans as to how you can move the world in the direction of the world that you want it to be the way you want it to be. And that's a very general thing. That's true of a lot of AI systems today, except that the world they're modeling is very limited and small. Maybe they're just modeling how certain things are connected on the Internet and they're software bot that's chasing you around and showing you advertisements or something.

But also having that goal directed behavior is significant for. For being able to predict what those kinds of agents will do. Part of Nick's book focuses on AI systems that are subject to what he calls convergent instrumental goals or convergent instrumental values, which are things that we can predict a sufficiently intelligent system to have goals. We can predict them to have because they're useful for the pursuit of almost any other final goal that you might give the AI, whether it's maximizing ExxonMobil stock price or calculating digits of PI or whatever.

34:19 Nick: Yeah.

34:19 Mark: The problem, of course, is that all these terms of having a goal are terms that we come to an understanding of through our own human experience. And so even you were just saying, oh, well, we don't want to talk about consciousness, but we can talk about having a picture. Well, having a picture is a term that to me involves some conscious accretion. So we have to think in one level of abstraction.

34:39 Dylan: Yeah, but it doesn't have to be that way, Mark.

34:41 Luke: Right.

34:41 Dylan: I mean it can just be an optimization problem where that goal is a number. I mean, in the case of more digits of PI, it's not at all clear to me that that's the kind of abstraction you're talking about. There's a completely mechanical way in which you could go about attaining that goal,

34:56 Nick: but in terms of an image. So a lot of even contemporary AI systems have an internal model of the world, their environment, which they are updating as they get more information from their sensors. And I think that when Luke was talking about the picture, I think he was referring to that some internal model that the system uses to keep track of information from the outside and to plan its actions well.

35:19 Mark: And you could also come back to that and say even a lot of things going on in human beings, your brain has a model, has a picture of the workings of the blood pressure, whatever internal physiological process, and then reacts to make sure that nutrients are going to the right place or whatever the. So there are plenty of things, teleological processes even going on, on within us that we would use this goal talk if we want to describe, at least informally, but clearly don't have a conscious component. So that kind of model is not unfamiliar.

35:50 Nick: So I seem to remember that. I mean, I might be wrong, but didn't you discuss some paper of mine like a year back or two years back or something as I vaguely remember, some kind of.

36:01 Mark: Right. It was about eight months ago, I think, the one on transhumanism. And we did it because we had interviewed David Brin who talks about many of these same topics in his most recent book that you do, except in a sci fi context, and wanted to have a follow up discussion without him just so we could go on a little about some of the topics. And to bolster that since he at least through illustrations, maybe not as much discussion as we would like in a novel, obviously. Hopefully it doesn't have as much discussion as your kind of book or else it's a very unreadable novel. But so we ended up just based on recommendations from our Facebook group, reading your paper on transhumanism. And so that gave us a little glimpse of what your kind of philosophy looked like.

36:39 Nick: So I don't remember exactly. There was something about the segment that they didn't laugh. Maybe I felt that what you thought were my views were not really my views. And maybe there was some other Thing, I've forgotten the details, but I remember not being bowled over by that segment.

36:51 Mark: I would imagine we treated it rather casually, which is why I sent it to you to say, yeah, we should do something to rectify this. And that's why you're on now. You proposed. Well, you should have me on to

37:01 Nick: talk about the book. Oh, yes, that's right, yes.

37:03 Mark: I think our complaint at the time was largely about the kind of philosophy that you were doing, which now, seeing this whole book gives me a much clearer picture of, like, well, why are we worrying about what might go on in the future when we have no real way of determining this? And you've kind of answered those objections in other interviews of you that I've been listening to in preparation for this. But really the bar for what should count as philosophy, what is worth studying, I think is pretty darn low. You know, the philosophy is kind of a grab bag where if somebody feels, thinks it's going to be useful to talk about, as we just did in our last episode.

We read Edmund Burke on the Sublime. And, you know, if you're going to think that there's something insightful to come out of talking about aesthetics or many other issues like that, then certainly talking about the potential ethical implications of the development of machine intelligence, or as was the case in that paper, augmenting ourselves to have superior capabilities and whether we should have a hard and fast, no, we will not play God, or whether we should let the technology go, as it probably will anyway. But as you are doing, accompanying it with a lot of research and thought to make sure that these things are used safely and ethically and in a very conscious way.

I mean, I think I'm pretty sympathetic with that project.

38:10 Nick: Right. So let's rectify them. So, I mean, I guess the more interesting question would be if we raise the bar to some higher level. So you're right that a great many things are worth somebody thinking about and writing about at that level. I mean, I guess I would be tempted to make a stronger claim in that I believe some of these questions, I mean, it's not a coincidence. Obviously, the reason my group is working on these questions that we're working on is that we think they are particularly interesting and important. So maybe the more interesting thing to discuss would be whether we are right in thinking that these types of questions are particularly important, more important than the aesthetics or the philosophy of phenomenology of various past thinkers.

38:48 Dylan: Maybe we could say why the relative urgency of this kind of problem presented by superintelligence you laid out a bunch

38:56 Mark: of different pathways by which this thing can develop. We could take on this uploading of individual human personalities, in other words, the creation of an AI by trying to just do a whole brain emulation. You know, is that the way to start, or should we do a pure AI? And had quite a bit to say on the relative risks between those various things. And so since work is being done in both these areas, and obviously there's no consensus at all, according to your survey of, of thinkers in the field, as to when a full AI will develop, still, it could be fairly soon and even in intellectual life, even a hundred years, if we're talking about transforming a culture and raising awareness of this issue. That's the challenge.

39:34 Nick: Yeah. So, like, I think one impression that some people have of my book, if they haven't read it, is they imagine that I must be saying, I'm convinced AI is just around the corner. Hold your breath. You've underestimated how far along AI is and that we should get ready because the robot army will be here in a day now. Or that we can pinpoint when it will happen by extrapolating these exponential curves in various technology fields. I'm not saying any of that, but it's like a common thing for people to say if they are writing about AI in a kind of breathless way. So I just wanted to clarify for your listeners that that's not my stance. I'd be curious what you guys think, though, in terms of either artificial intelligence or, let's say, biological enhancements of human intelligence. What's your views about the timeline for those hypothetical developments?

40:22 Dylan: I think that various kinds of biological enhancements are going to happen on a shorter time scale and variations of our biology and even genetic selection and so forth is going to happen and is going to present significant ethical issues for us, and that there's a high likelihood, in a relatively short order of having classes of human beings that while I wouldn't think that they're super intelligent on the scale that you're talking about, but might come close to presenting a kind of speciation problem for us, that we either have to enlarge our understanding of what human beings are, or we are faced with a problem of having whole groups of entities that we want to recognize as human beings but seem like a different species and capability, I'm relatively skeptical of the machine part of it.

Well, I guess I'm just doubtful of the ability to process and the kinds of problems that biological systems solve. I'm not completely convinced that mechanical systems are going to be able to solve those problems easily. And the kinds of processing that happen, it might just be a question of just magnitude. At the end of the day, you're able to do it. And you go through some of those arguments at the beginning of your book. I think the more immediate problem is going to be changes in human, like, beings.

41:37 Mark: Well, and I like this section in your book, Nick, about cyborgs. Basically why cyborgs are not going to be the next thing. That. That it's way more efficient even just to have Google Glass in front of your eyes than to somehow hook up some additional organs to your brain that would feed more information into it. That the brain is not made to take in any more information kind of than it already does.

41:57 Nick: That's right. And we already have very good ways of putting information into the brain, like the eyeballs and stuff. Yeah, 100 million bits per second.

42:05 Mark: Ah. So we're not going to be able to put your book under the pillow and have it seep in at night anytime soon.

42:10 Nick: You could try. There is an audiobook version. You could have that on repeat, play every night. I've told I haven't checked it out myself. So they asked if I had any preferences regarding what reader should be reading the book. Like, I suggested maybe an American voice because the spelling is American in the book, a male voice because I'm a man. So that seems more natural and ideally somebody maybe with some technical background so they could understand what they were reading. But I've been told they instead went with a British voice actor with a background in romance novels. And that I haven't been able to bring myself to actually check this out. But apparently it's read, like, excruciatingly slowly with the reader lingering on every word. It's like a tongue caressing you all over the body when he's, like, building up.

42:58 Dylan: It's a bodice ripper.

43:01 Mark: You know, when I listen to other podcasts, so all the stuff that I listen to you of your speeches to prepare for this, I always listen at 1.5 speed or sometimes 2 times speed. So I always get the impression that everyone is very frenetic before I talk to them. So that doubly contrasts with what you just described. Although I think having an actual robot voice would be better. I think I might just have to get the Kindle version and have the Kindle robot voice talk to me.

43:26 Nick: That could work.

43:26 Mark: It's only fitting that the superintelligence of the future will talk like this and sound very angry. I think we should talk about the value loading and get to the main feature, the problem of machine motivation and where that comes from. And I thought that it came pretty directly out of the account of personal Identity around page 109. I see in chapter seven you are talking about. So personal identity, a favorite topic of philosophers, present and past. And generally we come down in sci fi novels and things it has to be related to to memory. So if I suffer a brain injury and I don't remember my past self, well, it's not really me, even though the people treat me the same, because it's the same body.

So there's regarding Henry and other movies like that. And on the contrary, on the other side, if we had teleportation, if we had duplication and so all these new creatures were duplicates of me and had all my memories, then at least they would think that they are me, right? Not after the point of the split. So anyway, it all comes down to memory. Whereas you come pretty directly against that, that for artificial intelligence, not having all the accretions of consciousness that we have, that it would be a goal. Content, integrity. Is your account of personal identity?

You want to talk a little about that?

44:37 Nick: Yes, I think that's more fundamental in this context. Now there might be many different concepts of personal identity. So when we're normally thinking about these things in philosophical context, what we are interested in is maybe questions related to prudential rationality. Like insofar as we're motivated by self interest, which possible future continues should we regard as our future self? Or maybe from the point of view of ethics, we are interested in which person slices belong to the same person, because that relates maybe to various questions of justice and autonomy and such.

But it might be that I have a different concept in this book because the question I'm interested in there is what can we predict about what this artificial superintelligence will do? And so I discuss that there are these convergent instrumental reasons for doing various things. As Luke mentioned earlier, these would be reasons that crop up pretty much whatever your final goal is and pretty much whatever the situation you happen to be in is. So let's take some arbitrary final goal. Like suppose that your ultimate goal was to make as many paperclips as possible.

Maybe you have been built to run a paperclip factory and you became a little bit too smart for your job, and now you're out to optimize the universe for the production of paperclips. So then a number of instrumental goals emerge from this. For example, it seems you would want to prevent your human programmers from switching you off. Not because you like are worried of dying or because you have some basic desire to be alive, but because you can predict that if you are switched off, then you won't be around tomorrow to make paperclips. So there will be fewer paperclips if you're switched off.

So you then have this instrumental reason to prevent them from switching off. So self preservation emerges as this kind of convergent instrumental sub goal. But another sub goal that emerges is goal content integrity. The goal of preventing anybody from changing your final goal. You could predict that if somebody reprogrammed you to instead of wanting paperclips to want staples, then there'll probably be fewer paper clips in the future. Because now you will be a staple maximizer in the future instead of a paperclip maximizer. And I submit that goal content integrity is more fundamental than self preservation.

That is, you could easily imagine the AI choosing to sacrifice its personal identity if it could predict that there will be another artificial agent with the same original goals. If I die, I'm a paperclip maximizer. I die. But I know the programmers will create another AI that will also want to maximize paperclips, then that could be just as good as me surviving. Because there will still be this agent in the future that will try to make paperclips. Whereas if I survive with my goals change tomorrow I want staples instead, then I can predict there won't be a lot of paper paperclips.

So when these two instrumental sub goals come apart, it looks like goal content integrity takes precedence.

47:23 Mark: I guess it seems like then a stretch to call goal content integrity an actual theory of personal identity. Because if you're talking about something that again, identity is a term that we come from our human experience where it has all these emotional connotations that it's almost like pride and all these other things you referred to, but as a functional unit. I mean, in the same way that I was saying earlier that ignoring the whole is the Turing test, You know, something passing a Turing test or any other behavioral test, actually conscious or not, if you ignore that and are worried more about how you treat it right, do you adopt the intentional stance in talking about it?

Then what constitutes a unified thing as far as an actor, it does seem like it's the goal is one of the main things that connects the pieces together. So that if you're talking about a country, one of these examples that you gave of more than one person working together in a firm or a country or whatever, to produce a common product, then you can call that thing an organism, an entity, insofar as it is actually unified, insofar as the individual workers are not subverting the overall goal. So in the same way it's even more obvious in a case where a mechanical being could potentially reproduce itself many times over.

Well, when do you stop calling it one thing? Well, when the goals among the various parts diverge. As long as it's a unison of a whole bunch of even if numerically different but like minded critters, then you might say goal, content, integrity. It's the Borg, it's essentially one entity.

48:44 Nick: Well, perhaps. I mean, I didn't mean to really make any claim one way or the other as to how these convergent instrumental goals relate to the philosophical concept of personal identity. But then I could add a further observation, which is, if we do imagine some kind of more advanced technological state, maybe where there is a whole ecology of digital minds, then it might be that some of the concepts that are currently useful for classifying and understanding the world, maybe personal identity arguably is one of those. It's no longer be clear how they will apply.

Imagine if digital minds can swap memories arbitrarily quickly, just download your memories and erase mine. Or if I die, I could boot up another copy of me immediately, or copies of them. And if minds dissolve into cognitive modules that kind of transact with one another and outsource functionality to other cognitive modules, if you have this functional soup, that might not be anything there that looks very much like one we think of as an individual mind persisting through time. And it might be that something like a teleological goal thread would be a more trenchant concept to kind of get a grip of what would be going on in that context.

49:49 Dylan: I found it actually pretty interesting and even fits better with what we think about as an individual entity. I mean, I understand why we want to talk about identity as being tied up with memory and the path to get to where the entity is now, but it seems like even in our own individual identities is that there's a strong future component to it and where is it going and what's going to happen from now forward. And that the criteria by which we would call something an entity would be the way in which its individual parts are aggregated and associated so as to be one thing.

And what do we mean by that one thing going forward in time and so that they're tied up together such that we can call them one thing. And that would be true for a physical system, something that was you know, you would never imagine calling conscious or anything like that, and even talking about it being goal oriented, it would be analogical. So it seemed to me that this way of talking about it was nicely flexible as a kind of, well, a kind of ontology for these kinds of entities.

50:55 Mark: Can we say a little more or has it already been made clear why this goal content integrity? I think that's the central thing that gives rise to the problem and also

51:04 Nick: creates a hope of a solution at the same time. Like without this sub goal of goal content integrity, you could always wonder how even if we managed to create our first seed AI with the right human friendly values, why it would not be the case that once this seed AI grew up into mature superintelligence, why it wouldn't suddenly just decide to change its goal? Suppose it started out being human friendly and then one day it just decides, no, I just want to make paperclips instead. But if goal content integrity is indeed a convergent instrumental subgoal, then we have some reason to think that that will not happen.

I mean, we can even apply it to a human being. If you are a human being, you care about other human beings. Suppose some mad scientist offered you a pill that would suddenly make you care only about making as many paperclips as possible. You just swallow this pill, it has no side effects, would you take the pill? Obviously not. Like I have no desire to become a kind of person who only wants to maximize paperclip clips. It would seem to be one of the dumbest things I could do to take that pill. It would mean that none of my current goals would be realized for exactly the same reasons, and AI would also want to work very hard to avoid that outcome.

So it looks like if you can start out an AI with the right values, that then once it becomes sufficiently capable, it will be on your side, as it were, in preserving those values.

52:16 Dylan: Is that a claim about AI systems that depend upon certain features of AI systems that they would have this goal content integrity? Or is it a broader claim about just agents? Agents? Yeah, yeah.

52:27 Nick: Instrumentally rational or capable agents in general. And so there are various possible exceptions to this one could imagine, but they're maybe not likely to arise in practice. But suppose that some mad scientist had invented some kind of brain scanning apparatus, and you had to go through that brain scanning apparatus tomorrow, and unless you were found to actually really value baseball at sar, like maybe obedience to the current dictator, then you would be killed right after you come out from the brain scanning apparatus, then you might now Have a reason, if you could, to change your values so that you would actually like baseball or love this dictator so as to avoid being shot immediately after exiting the brain scanning apparatus. So that would be one type of exception to this general rule.

53:10 Mark: Right? And so the puzzle here is that once something is a super intelligence, then we're not going to be able to mess with it. Then it'll be more or less in control of itself. It'll already have its motivations intact. It'll be operating according to this goal, content integrity. And so if we have not specified its goals before it reaches a level of superintelligence, then we're just screwed. If it's going to make maximize paperclips, then it's going to turn the entire earth into paperclip material, etc.

53:34 Nick: At least that should be our working hypothesis. Like we should definitely not rely on being able to do something after we already have a super intelligence, at least if the superintelligence is also released into the world at large. So when thinking about this control problem, there are two broad classes of control method. One might envisage, so we've so far alluded to what I call motivation selection methods. This is when you try to engineer its motivation system so that it would have values, goals that make it act in ways that we approve of.

But the other class of control method is capability control methods where you try to limit what the system is able to do. So you could try to lock it up in a box, disconnect the Internet cable, and maybe limit the system's ability to affect the world. Maybe it's only able to type answers out on a screen or answer simple questions by kind of handcuffing it, removing its ability to cause harm. So I mean, at this stage I think both of these types of control methods should be researched more. Both a capability control method and motivation selection methods.

But ultimately I do believe that we have to find a solution among the motivation selection methods. I don't think we can forever keep a super intelligent agent bottled up. At some point I think it'll get out of the box.

54:47 Mark: So that's where the book got the most fun. That it becomes like the guy making the three wishes in the genie is determined to interpret whatever wishes are made in such a way that the wisher will be screwed.

54:58 Nick: No, no. So the system is only interested in interpreting the wishes literally and then doing the thing that will lead to the greatest possible realization of the world wish interpreted in that literal sense. So we don't assume here that the system is Actively hostile to the wisher.

55:13 Mark: Right. But it's the same dynamic in terms of how can I make a very specific kind of wish so that there is just no chance of misinterpretation? Because once it's super intelligent and it will be able to think outside the box in various ways such that we might not have thought. One of your points that I liked in there was even if you tell it, your goal is only to make 100 paperclips, so that seems pretty safe. So it makes 100 paperclips, but there's an epistemic gap. How can it be sure that it's made 100 paperclips? Well, let's go ahead and just make infinite paperclips and then we'll be sure that there's at least 100 in there.

Or let's go ahead and kill everybody on earth that might be lying to me about the number of paperclips that were made or turn the entire earth into There are all these horrible ways that even just to raise very minutely the percentage chance so it can be sure that it's achieved its goal.

55:55 Nick: Yes, because there is no downside to that. Like if your only goal is to make at least 100 paperclips, then once you've made your hundred paperclips, there is no cost to continuing activity to maybe make sure and then double sure and triple sure that there really are 100 paperclips. So you could certainly imagine creating ever more elaborate measurement apparatuses that could check that the paper clips you made really added up to 100 or other things that Wang. So there is this general failure mode which I call infrastructure profusion, where as a result of pursuing what we had hoped would be a self limiting goal, it turns out that the AI might produce an indefinite amount of of infrastructure to realize that goal.

And then we humans, now human values would perish as a side effect of that infrastructure being constructed. So in this case maybe an infinite number of paperclips being produced, or an arbitrarily large apparatus for measuring how many paperclips have been produced in other cases, other types of infrastructure. But that would be one general modality through which the pursuit of some arbitrary goal could lead to the destruction of human value. Right?

56:58 Mark: And you're very thorough in here about not only the potential problems with boxing it in, with setting up tripwires, with trying to stunt it, and talk about how if it is in the interest of achieving the goal to lie to us, to appear to be stunted while Actually hiding some of its programming itself, evolution, as it constructs more and more sophisticated forms of intelligence and gloms them onto itself, Then it will do that so that they have a section called the Treacherous Turn. That was a pretty interesting one.

57:25 Nick: So you have here potentially an intelligent adversary, as opposed to just a simple natural phenomenon, moreover, a potentially super intelligent adversary. So the idea with a Treacherous Turn is that if your approach to ensuring that the system will be safe is first to run it in a sandbox environment, like this box disconnected from the Internet, observe how it acts in that safe environment, and then if it looks to you that it is behaving nicely and cooperatively in this safe environment, only then let it out into the wide world. You then confront the fact that it looks like a friendly artificial intelligence.

And an unfriendly artificial intelligence would both have a convergent, instrumental reason to act in a nice cooperative way while being in this safe environment, because they could predict that that would be the way to then be let out into the wider world, where they would be able to realize their true goals. So you might not get very much diagnostic information from doing these behavioral tests. Once the system is intelligent enough to be able to strategize about perception management.

58:25 Mark: And likewise with boxing it, even if you say it can only just show words on a screen, well, if part of its super intelligence is the ability to manipulate people to get under our skin and make us obey its will, essentially then even that might be too much. If you shut it up too much, it's going to be useless. But if the more you let it out, the more danger.

58:42 Nick: Yeah, so completely isolated box in a Faraday cage, maybe that would be safe to have a super intelligent in this self contained box, but it would also be utterly inert. Of course, to get any impact on the world, we have to interact with it in some way. But once you have a human interacting with the superintelligence, then we already have a danger. Humans are not secure systems. You could imagine that maybe it could hack its way out of the box, or maybe it could use social engineering to talk its way out of the box. Even human con artists sometimes manage to persuade other humans to do things against their interest. And here we have to assume a superhumanly powerful and able persuader. And so it would not be safe to rely on that.

59:19 Luke: And then another point to remember is that even if the leading AI team in the world is very, very careful about boxing, its AI might be able to succeed at that for a number of months or maybe even years. The next team behind them figuring out how to build a fully general AI might not be so careful with their boxing methods and might want to let it out of the box earlier so that they can win the arms race or something like that.

59:41 Mark: Right. So it seems a lot of the discussion at the end of the book is sociological recommendation, the right level of openness, which is mostly not open, but yet sharing information between researchers to avoid any kind of arms race. You know, what political situation can we set up so that people would expect that the development of a super intelligence will aid all of us so there's no motivation for one particular group to steal the information and charge ahead and sabotage everybody else, that kind of thing?

1:00:04 Nick: Yeah, I mean, to me, that's almost the most interesting part of the overall argument is what should we actually do do now in light of these prospects of superintelligence at some unknown point in the future? One of that connects back to the question of cognitive enhancement in biological humans, actually. So some people I've heard suggest that if it's true that there might be these AIs in the future, machines might surpass us, then what we should do is to try to enhance our own biological intelligence so that we can keep up, keep one step ahead of these intelligence machines, and that that would be the way to ensure that humanity has a place in the future.

My view is that that's misguided, that if we embark on biological cognitive enhancement, it will only hasten the day when machines overtake us, because we will then have smarter scientists who will be doing the computer science research, and they'll invent artificial intelligence sooner than a sort of a less capable version of humanity. So I don't think that by pursuing biological cognitive enhancement, even full throttle, we will manage to keep up with the machines. We will be overtaken even sooner. But the people who are engineering these superintelligences might be more capable than we are today.

And I think for that reason it nevertheless looks like a beneficial thing to do. If we can find medically safe ways of enhancing human cognition, I think it could also help with some other existential risks.

1:01:21 Mark: Right. Well, you say that enhanced humans will be wiser and more likely to obey the recommendations of your book or subsequent similar.

1:01:31 Nick: Conditional on my book being correct, then you might think that the cognitively more capable reader will be more likely to realize that it is correct. Conditional on the book is being wrong, then hopefully they will discover exactly what the flaws are and be able to do better.

1:01:44 Luke: I should say though, that these sociological recommendations become very complicated and non obvious. So you might think, for example, well, having smarter humans is going to make us more likely to navigate this transition correctly. But on the other hand, the kind of cognitive capability that we understand really well in humans and might be able to, for example, genetically select for in the next couple of decades is sort of the one that's referred to as G and correlates with IQ score and so on. But there's a lot of interesting literature in psychology showing that high IQ doesn't necessarily translate to wisdom and that many of the cognitive failure modes that we know a lot about aren't correlated very much with iq.

And so there's this researcher named Keith Stanovich who studies a lot of this stuff, and he has this section in one of his recent books where he talks about, what if we gave everyone in North America a pill that made them higher iq, basically? And his view, and the view of many other people in his field, would be that everyone would go after and achieve their goals with greater speed and efficiency and effectiveness, but those goals would still be things like poor medical treatments because we're irrational in the way that we think about medicine, or making poor financial decisions because of overconfidence, but we would just make those, those poor financial decisions faster, I guess.

1:02:59 Mark: A related thing you'd said the question of how should we proceed in AI, should we just try to develop a full AI, or should we try to develop whole brain emulation? And you say, well, if we develop whole brain emulations, in other words, the first AIs will be replications or variations off of the scientists working on it, probably or famous dead people or whoever we want to replicate. Then we run the risk of neuromorphic AI, which is much worse than if we go wrong with regular AI. Can you say more what that is and why that's so bad?

1:03:27 Nick: Yeah. So these strategic questions become extremely complicated. Yeah. The idea is that one might have the view that whole brain emulation would be safer than purely synthetic forms of machine intelligence. Because if you literally just uploaded a human mind onto a computer and then enhanced that to make it super intelligence, you would at least start with something that had human values. And you might hope that those human values would then be inherited by the superintelligence growing out of that. Even if you thought that that were the case, which is not completely obvious, because it's not as if we would trust all human to become like a world dictator either.

Even if you think that it would be safer to do it that way, it's still not at all clear that we should be aiming towards whole brain emulation, because I think we would likely, before becoming able to do an actual high fidelity whole brain emulation, to learn enough about the brain to be able to clobber together some kind of neuromorphic AI that would not really carry human values over it into the digital form, but that would still copy and paste various elements from biology that we didn't necessarily understand very well. Well, it would seem preferable that our first AI be one that we understand very thoroughly and have simple mathematically well characterized properties, than that it be some sort of conglomeration of different things cribbed from biology with some other things that we have added on and nobody really knows what the whole thing is doing.

So partly for those reasons, there are some other considerations as well. On top of that, I think it probably unbalanced, looks undesirable to push towards whole brain emulation. And another variable here is we don't necessarily want machine intelligence to happen sooner than necessary. And so by pushing towards whole brain emulation, we might speed the arrival of superintelligence. Whereas if we think that we need more time to figure out solutions to the control problem, or just more time for humanity to grow up, for our civilization to improve, then we might prefer that progress occurs more slowly, both along the whole brain emulation path and the AI path.

1:05:16 Mark: So maybe let's talk a little bit more about the motivation selection. Talked about the other strategies within the control problem, but the main one, like you were saying, is to try to give it the right kind of goals. And we talked a little bit about how hard it is to be very specific about those. But it seemed the solution that you recommend, at least one of the options, is not to directly specify what goals it's supposed to have, but try to approximate more like what people learn, that we figure out our own goals as we go along. Then you have to give them meta rules. You have a section here called indirect normativity. We specify a process for deriving a concrete normative standard rather than just giving it the standard itself.

1:05:52 Nick: Yeah, so this is not necessarily exactly how humans acquire our values, but it is perhaps the most promising approach to date. To take a step back, first one can sort of distinguish two problems. So on the one hand, we've alluded to a bit. There is a technical problem. If you knew what goal you wanted your AI to have, how could you actually program that goal into the AI? So that's a big unsolved challenge. We can program maybe the goal to calculate as many digits of PI as possible today. But there's no way you could program in Python or C, like Maximize Love or Justice or Global Harmony or Beauty or anything like that.

But then there is the second problem, which is given a solution to this first problem, what goal should we actually decide to give to our seed AI? So indirect normativity is an attempt to answer that second problem. And rather than picking a goal, our current best guess in moral philosophy or anything like that, we could try to delegate this task to the AI by giving giving the AI a goal, such as do that which we would have asked you to do if we had been thinking about this question for like 40,000 years. And if we had been smarter ourselves, and if we had known more facts, and then we would have translated the problem into this empirical question.

It's like an empirical question what we would in fact have asked the AI to do under those idealized circumstances. But we can then lean on the AI's superior cognitive abilities to make a better estimate of what the answer to that empirical question is than we could do ourselves by just doing the moral philosophy directly. And so we would then kind of leverage the AI's superior cognitive abilities to hopefully get a better outcome than if we just tried to make our best guess as to what the best moral theory is or the ultimate optimum state of the world.

There are other variations of this indirect normativity that one also considered, but this gives kind of a sense of the

1:07:39 Mark: core idea you made the connection in the book to. I forget whether it was Hume or Smith, their notion of the ideal observer, the good is what the ideal observer would decide was good in this kind of situation.

1:07:48 Nick: Yeah. So this is like a philosophical position in ethics. I don't necessarily need to accept that position as a true claim about fundamental ethics. Even if that's not what is the definition of good, or what constitutes the good, or a necessary condition for the good, it might nevertheless in practice be a good heuristic for the good. If something would be recommended by somebody who shared our values but knew more and were smarter and had thought about something longer, then that seems to at least be some positive recommendation for going with this particular outcome rather than another.

1:08:21 Mark: And you seem to be putting a similar challenge to the AI when you ask about, here's what goal I want you to have, and I want you to implement it in the way that I mean you to. So you kind of put in an X there, and it has to investigate what is it exactly that you could mean by giving me that specific goal and try to give the most Sympathetic interpretation, thereby having to figure out what sympathetic means. So figure out the meta rules for itself that it seems a very similar process to figuring out sort of what the ideal me would want you to do to even what I really want you to do, regardless of my skill of being able to program it in C or not.

1:08:52 Nick: Yeah. These are sort of different variations of the same basic attack. So you could say, like, do the right thing and interpret this command in the way that I intended you to interpret it. And then of course, the challenge becomes to specify exactly what this thing means, intended you to interpret it in this particular way. And you could try to give the AI some criterion, maybe for what counts as a correct interpretation, but that's then a big open problem. But if you succeeded in that, then you might be able to specify some goal in very abstract terms like do the right thing, do the thing that is best, all things considered, achieve a good outcome.

But it would require that the AI would interpret this in the right kind of way. And that would seem to require some form of indirect normativity. I'm curious, what do you think if. Suppose you could get the AI to always do the morally best action. Would you give the AI that call to always do what is morally best if you could, or would you prefer it to do something that, say, always was what humanity would have asked it to do if we had the opportunity to consider the question for a long time in ideal circumstances, given that we

1:09:54 Mark: read the book, this is a trick question in that we know the answer is no, that what might be, given that we don't know for sure what is morally best. It might well be that what is morally best is for us all to be. Be extinguished from the planet and replaced by beings that would be more capable of good or more capable of appreciating life or any number of things. Right. Even just taking utilitarianism, that the greatest happiness for the greatest number. Okay, well, then it should create a lot of beings that are very easy to make happy.

1:10:24 Nick: Yeah.

1:10:25 Mark: And get rid of us who are clogging up the world with our suffering.

1:10:28 Nick: Right. And it's not clear that we would have asked the A to do that if we had pondered this question for a long time, because we might also have selfish inclinations.

1:10:37 Dylan: One of the things that I was struggling with in reading the book and it bears on this question that we're talking about right now, is, I don't know, the kind of disconnect of an AI dominated world with. In general, you talk about a multipolar possibility, but in the end you seem convinced, and it seems pretty convincing that you'll end up with a singleton of some sort of sort. And that makes questions like, well, what kind of morality will we want the AI to have? Or what kind of form of the answer do we want them to consider? And so forth. And it's very hard for me not to come back with feeling like the whole landscape becomes just less rich and interesting, especially in terms of these moral questions where we have the relative likelihood that people end up living according to varieties of moral choices that they make.

And you have the question of whether there's one thing they ought to have done, but you have just the fact that people make mistakes. People also maybe even likely have multiple kinds of interests. And also that even in the best of cases that individual entities, human entities, have different best interests. And that seems like a very hard problem to solve. And so when we're talking about what kind of morality and moral choices that we would have with an AI, it seems unavoidable that you're going to end up getting a kind of flattening effect across.

You have a singleton for the AI, but also their interaction with human beings and the consequences of that interaction, they'll just be in the end less rich. And that kind of thing makes me just kind of bummed out. It's akin to the biological diversity problem that you just end up squeezing things through a filter. Or maybe the better analogy, it's just. It gets blended into the same kind of goo.

1:12:38 Nick: I don't think that there is any particular reason. So with a beneficial outcome, and if we take this singleton scenario, suppose that it is, did actually execute humanity's coherent extrapolated volition in the sense of doing that which we would have asked it to do if we had pondered the question for a long time. If it really is impoverishing to imagine that everybody would always behave perfectly morally, ethically and pursue the same life goal and everybody would be little ideal saints walking around. If that would really diminish the diversity of the world in a way that would be bad, then presumably we wouldn't ask the AI to, to produce that kind of outcome.

1:13:16 Mark: Yeah, it does seem like that asking for either of these two extrapolations, either what do I really mean or what would I mean if I really thought about it? Of course, because it's just setting up a problem is very hand wavy. And so if you're saying, well, just use magic to program the volition, in other words, have it figure out itself, then yes, we can, as part of that Supposition that whatever engineering problems and theoretical problems involved in that are ultimately able to be worked out, then we can build anything else else into that that we want.

And we can say, if ultimately human good involves a high level of vibrant individuality, so there is no single coherent extrapolated volition, well, then we'll come up with a variation of that, giving the AI that task. Then it would figure out that it somehow has to do different things for different people and promote diversity or whatever other goal you have.

1:14:05 Nick: So it seems to be other things equally desirable to avoid begging as many questions as possible in that that the ideal to me would be set up a process that we could rely on, would lead in a good direction without us having to make too many bets on exactly what is good and bad. Because I think it's quite likely that we are misguided about some things, including some things in fundamental ethics. If we look back at any earlier stage of human civilization and we imagine that they had had the opportunity once and for all to embed their values in an AI that would then maximize them for the next billions of years, we think back on earlier ages and we saw that, well, they believed in slavery or human sacrifice, or in women's place being subordinated to men, or they had all kinds of other moral blind spots that we now think we can see.

But presumably we are still laboring under a variety of moral blind spots ourselves. So we want to preserve the possibility of moral growth and the possibility that we could could come to recognize important values that we are still oblivious to. So for that reason, a process like this, indirect normativity, coherent extrapolated volition seems very attractive in that it wouldn't at the outset foreclose the possibility that the future could look quite different from what we currently think it should look like, but still not arbitrarily different, but different in a way that hopefully would reflect our own deepest values.

1:15:30 Mark: And it seems to address the goal content integrity issue that you're telling it. Okay, yes, keep this one goal to keep pursuing coherent extrapolated volition as well as you can. But what actual sub goals are going to come out of that are going to change according to its own growth. So you're sort of programming it to have moral and spiritual growth, but along certain lines, hopefully so that it doesn't just grow in some direction and decide again that humanity should be exterminated.

1:15:58 Nick: Yeah, yeah. So you get for free various sub goals from this. So the goal, for example, to do more investigation into what humanity actually would have asked under These idealized circumstances becomes an obviously very important sub goal for the AI, because if you knew what AI would want, then you're more likely to be able to achieve it. Also, to the extent that at the outset you can make some guesses about what humanity would ask for. For example, if it's moderately likely, at least that humanity would ask not to be exterminated, then you would would until you have had a chance to investigate the matter more closely.

If you are the AI, like ensure that humanity doesn't go extinct. So a lot of things fall into place. If one imagines that the AI would have this as its final goal, a lot of instrumental goals, then that looks quite sensible would emerge that could be derived from this final goal. But there is obviously like a large technical, unsolved problem of how you would implement this kind of indirect normativity in an AI. And there are also some variations of indirect normativity and it's not clear exactly which of those we should be aiming for the AI to pursue.

Like who is this we who would be asking things of it 40,000 years from now under some idealized circumstances? Is that all currently existing adult humans or does it include humans who lived in the past or animals? So there are a number of free parameters that one would have to specify. Hopefully the outcome would not depend too sensitively on exactly how we specify those parameters. But a completely contentless specification is impossible. We would still have to lay down some parameters that would define what this exercise is that the AI would be engaging in and we would be making some non trivial normative commitments when we did.

1:17:36 Dylan: So it does make me want to think about what it would be like to be a conventional human being like myself right now, and living in a world that has such an AI in it. And it makes me think that it's going to be like having another thing like the environment or the climate or the weather that you would be taking into account. I guess there's a whole variety of imagined relationships and depends a lot on the particular instance of that AI. But it seems like it would just the existence of it would significantly change the way non enhanced normal intelligent human beings are going to even think about themselves.

1:18:19 Nick: Probably changed. I mean, like people who believe in a God has some probably impact on how they look at the world and themselves. But ultimately I think that these psychological effects of believing in the existence of a superintelligence are swamped and dwarfed by the direct material effects that the superintelligence could have by acting on the world.

1:18:41 Dylan: That's what I was thinking more of. I mean, yeah, there would be the effect of, of well, what does it mean for my self identity? And like you said, linking it up to the way in which we would think about whether we believe or don't believe in God and how that affects our identity. But to the extent that there are real tangible effects in the world, it then seems more like you have a. It's not exactly non natural, but certainly it's a kind of overarching environmental change that affects us the way climate does or other kinds of physical constraints, but is utterly foreign to us now. It'd be like we become immersed into something that we now interact with that changes our way of interacting on an everyday basis.

1:19:26 Nick: Yeah, I mean this is what most people believe presumably who believe that the world was created and that there is some supernatural super powerful, super knowledgeable being out there that is overseeing the whole thing.

1:19:38 Mark: I took that more as pointing the Heidegger direction again of the same folks that would be bringing up the Dreyfus objection. Would be bringing up. How could we even contemplate this? Shouldn't we just be fighting this with all of our being? Because the human condition is to be connected in some natural way with the environment and bringing a super intelligence would void that. I don't fully understand the objection.

1:19:58 Dylan: Well, no, but Mark, Mark, I wasn't objecting to anything. I was just stating that I couldn't help but think about the fact that the kind of change we're talking about is one that amounts to a kind of deep environmental change for us as human beings.

1:20:13 Nick: Yeah, I think more than environmental. I mean, I think quite independently of AI that human nature itself is not an eternal fixed constant, but something that will probably change through technological interventions and so forth and that will affect our lives more profoundly than any change in our environment possibly could.

1:20:33 Dylan: I guess maybe I mean environmental analogically.

1:20:35 Luke: Right.

1:20:36 Dylan: Our day to day lives and what we interact with in the world, we have all kinds of things we are embedded into and that we sometimes notice and sometimes don't notice, but affect our lives and affect the way we think about ourselves and affect the choices we make, the range of choices that we have. And those might be the culture we live in, the country we live in. It might be our own physical bodies, it might be our histories, it might be the possibilities we see in front of us. And that with AI and the kind of superintelligence we're talking about is a change in the kind of box that we live in.

1:21:16 Nick: Sure, but I mean the box changes all the time. For all kinds of reasons.

1:21:20 Dylan: I agree with you on that. But it seems like the kind of change we're talking about here is one that is, I don't know, potentially very different than other kinds of box changes we've had. Maybe that's not true. Maybe it's utterly qualitative in the kinds of changes. It's no different of a change than other kinds of massive technological developments. And maybe it's just because it sounds so new or foreign or unconstrained that it seems like it would be radically different, I guess.

1:21:49 Nick: I looked around, I see humans live with all kinds of weird convictions about the structure of the world and their position in the universe. Some are atheists, some are believers, some believe in UFO ships that they're going to come and rescue us. And some believe that, I don't know, the Earth is on a big stack of tortoises, or that there are ghosts in every tree and stream. And just within philosophy you can depict some crazy position and chances are some philosopher will actually have argued for that and people still get done with our lives. And if there's one more thing that, okay, well, then there's this super intelligence that shaped things.

So I think people get accustomed to that after about two days or whatever. And then the much more important question to me is, what actual physical effects will this superintelligence have? Will it convert the world into paperclips? Will it protect biological humans against existential risks? Will it upload us into computers and open up some unimaginably large digital realm of enhanced post human experiences of unimaginable bliss? Will it expand our minds to something that is as far as beyond human minds as ours is beyond the minds of mice or something.

And what will those expanded minds do and experience and feel and aspire to? Those would be the much more important questions to my mind than whatever psychological adjustment would have to be made to this new discovery that now there's a super intelligence in the world.

1:23:10 Mark: Yeah, I think Dylan is channeling. We also read Throw a few episodes ago. It's certainly a common kind of objection to hear against any new. And I'm not trying to dismiss you, Dylan, fundamental technology, that it's going to reshape the human experience in some way. And I have the same trouble in terms of. Well, okay, fine. Tell me whether things are going to get good and bad in specific ways, not that there is an overall sort of natural way and that you don't dwell, certainly in this book, Nick, on whether any sort of technology advance like this is desirable at all.

Right. I mean, it's more. Well, I mean, is it like anything else, that this kind of thing probably will happen? Because even if I have a sentiment that this is a terrible kind of research and we should never do anything like this, then somebody's going to push ahead. So as a society, we have to be able to cope with it if and when it comes.

1:23:59 Nick: Yes, I'm very interested in questions about desirability. It just seems that there are a lot of different questions, and some of them are more important than others. So the question of whether AI is good or bad, like, would we want it to be impossible to build AI or would we want the world where it's possible to build AI seems less important to me because it's not something that we're likely to have a lot of influence over. Like, if it is possible and if science and technology continues on a wide front, then I think it will be done. And that the more important questions from a practical point of view are which directions should we be trying to move things on the margin?

Should we want slightly faster progress in AI or slightly slower progress? Do we want slightly faster adoption of cognitive enhancement in humans? Or slightly slower, if you had an extra million dollar, where should you donate it to? Or what should you do with it? It looks to me like funding institutions like Luke's Miri would be a good bet for what to do with that million dollars trying to accelerate work on the control problem. So that instead of sort of being for or against some general technology, it looks to me to be a much more fruitful question to think with our limited powers and influence and money on the margin, which direction would we want to pull?

Different things.

1:25:10 Mark: Luke, are there beefs? Pet peeves? Things, Whether it's objections you get from folks that you're interacting with or things that are coming out of your institute that you think would be relevant to add here about the CEV or any other aspect of what we've been talking here in the latter part.

1:25:26 Luke: Well, I think Nick's book is really the place for people to go. There's a great wealth of arguments and forecasts and considerations and open questions, questions that need to be studied in a variety of different fields that are in Nick's book. And we've only been able to cover a tiny fraction of those subjects in this conversation. And so really actually do go out there and grab the book. There's also my organization is hosting an online reading group where once a week we read through a particular section and we have a research assistant who puts Together, kind of a summary of that meeting section and some questions and there's a discussion in the comments and so on.

And so if people go to intelligent.org, you'll see a link to the superintelligence reading group there. And that might be the best way for some people to read through the book because then you get to read it along with other people in little bite sized chunks and, you know, look at related materials and understand it better. But yeah, I really do recommend the book. I think it does a really good job of summarizing a lot of challenging and counterintuitive content that has been developed over the last 15 years at Nick's Institute and my institute and some other places.

1:26:34 Mark: You know, it's an exhaustive book. We were exhausted just by the scholarship involved in your short transhumanism essay in terms of how deliberate each step was to sort of counter all the objections in the literature that you were dealing with that the average reader may or may not care about. I mean, I really felt like the book started to fly around down chapter seven, the Super Intelligent will talking about the instrumental convergence and the relationship between intelligence and motivation. The stuff before that I think was interesting, but I ended up skimming a couple of them because it didn't seem as important to me in terms of getting at the fundamental philosophical issues here of this machine motivation versus human motivation to dwell on exactly how fast we think the intelligence explosion is going to be.

So there. I think it's a book that you don't necessarily have to read from the beginning. I think a lot of people would get exhausted after chapter four if they do that.

1:27:27 Nick: Yeah, it'd be interesting to see some kind of. Maybe they have this for ebook versions if they can see how much of it different people read and what the curve.

1:27:36 Mark: I'd even like to know when people turn off our podcast. I know a lot of people. I get the first half hour. I got the idea. It's fine, move on.

1:27:45 Luke: Yeah, this is the nice thing about Kindle highlights is that you can now see are people actually reading to the end of this book and highlighting passages all the way through or do all the highlights stop by chapter four? So we'll have to check Nick's book later on Amazon to see whether anybody's finishing it.

1:28:00 Mark: I think the Super Intelligence will for sure finish the book.

1:28:04 Luke: The Super Intelligence will finish all the books.

1:28:07 Mark: That's a bit of nostalgia. Those foolish humans, what they thought about us now, they're all paperclips. We need to explicitly bring up the image that you had in here of if you say, you know, come on, super intelligence, make everybody happy, then of course the first thing they're going to do is to strap us down and put electrodes right into our pleasure centers to make us happy that way.

1:28:31 Nick: Yeah, or some other thing that might be even more efficient.

1:28:36 Mark: Make us into one giant Borg which can be happy. Only one unit of happiness is necessary to cover us.

1:28:41 Dylan: Make me, you, daimon E.

1:28:44 Mark: So one

1:28:44 Nick: would have to look at exactly how we had tried to specify happiness, but it might be by, say, uploading us and then stripping away the parts of our minds that are not absolutely essential for us still being there, and then just recording a short experience of intense happiness, and then running that on a loop, perpetual loop, and then transforming the universe into some configuration optimized for efficient computation, then replaying that loop in billions of different parallel instantiations, constantly for the next number of billions of years. It would clearly be more efficient than sort of sticking electrodes into a biological brain. So if that counted as still us being happy, then something like that might be closer to what you would get.

1:29:27 Dylan: Yeah, it would be interesting in thinking about both the process that you would have for the seed AI and also, I guess, what kind of general goal you would give, and whether it would be towards happiness at all.

1:29:41 Nick: If I had to choose something right now, I think that this coherent, extrapolated volition, the indirect normativity idea of doing that which humanity would have asked it to do, if we had a long time to think about this and deliberate on it in ideal circumstances, would be as good a shot as any if we had to do it right now, I think there are refinements to that that would improve the chances of a beneficial outcome. But as a first stab, it would seem better than me trying to decide whether in the future all humans should be happier, or whether they should be perfectly moral, or whether they should be maximally insightful philosophers, or whether they should be something else.

That just seems a very difficult for me to do. Right. And even if somehow I could, it seems kind of a usurpation for me to decide that on behalf of all humanity, or for any other individual person to do that. Just because you happen to be the first person to figure out how to build a superintelligence doesn't mean that, therefore you have a mandate to decide for all of humanity and for the rest of the future what Earth-originating intelligent life should be doing.

1:30:42 Dylan: Well, it also seems like it stokes the fires for the various failure modes that you Talk about the CEV case seems to not only be preferable, but seems to mitigate against those failure modes.

1:30:55 Nick: Yeah, if you could actually correctly implement it, yes, then it seems relatively robust. But that then has this huge technical challenge of how you would engineer a seed AI. And it's a very high level goal. I mean, it's not something you can just sit down and code into a computer today. So there remains this immense technical challenge that we need to overcome before somebody overcomes the other technical challenge, which is how to make machines intelligent in the first place.

1:31:22 Dylan: The kinds of AIs we've been mainly talking about have been, well, they're super intelligences. Things that would significantly outstrip the goal oriented decision making and ability to plan that human beings have. And I wondered about the proliferation of, for lack of a better term, dumber AIs that are more like a virus in that they have the ability to replicate and to infect without having something like a complicated or sophisticated goal or volition. It's a little bit more like your paperclip instance, but it might also be very decentralized.

1:32:01 Nick: The paperclip I'm imagining as super intelligence. But certainly you could have lower level intelligence entities. That means that just take like in the actual world we have bacteria and viruses that cause a lot of mischief without being very smart. And you have computer viruses that are also not very smart. I mean, maybe in the future they will be slightly smarter but still causing problem not by having some very advanced planning ability, but by kind of exploiting vulnerabilities of the immune systems.

1:32:28 Dylan: So do those just seem like, based upon your research and thinking, lower level existential threats to us?

1:32:34 Nick: Yeah. Synthetic biology I think could be another source of significant existential risk. Risk over the coming decades as we expand the capabilities of what we can do. So I wouldn't dismiss that as unimportant. There are a relatively small number of sources of significant existentialism. Most things that can go wrong just wouldn't pose a threat to the survival of intelligent life or our long term future. Throughout human history there have been myriad things going wrong. Earthquakes and firestorms and wars and famines and plagues and all kinds of stuff.

Stuff. But none of that probably really have made a significant difference to our long term future. It's like ripples on the great pond of life. But there are a smaller set of things that could go wrong in a way that would destroy all the future. I think AI is one of those things. If we don't do it right, synthetic biology could be another Source maybe like future advances in molecular nanotechnology, maybe at a lower level of risk geoengineering or ways to manipulate and control populations to enable new forms of totalitarianism. And there might be a few others as well.

Maybe some we haven't thought of yet. But it's a relatively short list, I think, of experts.

1:33:42 Mark: Maybe what the Frito Lay company is doing in testing and making chips more and more delicious, that eventually they'll be so delicious that we won't be able to control ourselves and that'll just be the end. There are many, many technological dangers.

1:33:55 Nick: Yes. They're like the fat explosion hypothesis.

1:33:59 Mark: So I guess returning to our initial question of why should a philosopher be doing this? I have no position from which to give a principled Dreyfus like objection to artificial intelligence of this sort coming about. I might have doubts, but I have no rational justification for those doubts. I don't know enough about the types of progress there have been that it does seem like the best I can do is what you did in the beginning of your book is to just survey the people who are working in the field as to what their estimates of when the artificial intelligence experiments explosion will come, keeping in mind that given that they've given their lives for this, they probably think it's possible.

So there's a sort of a selection bias there. There's no way sort of get an objective prediction of when this is going to happen. So I can see a lot of people being very dismissive of, you know, either this just isn't going to happen or it's not likely, or it's not going to happen anytime soon. And so I'm not going to worry about it. But I can see then giving a Pascal's Wager response, well, if you have no principled reason and why you don't think it's going to happen, then at least it's worth paying attention to because it could happen. And if it does, we have to worry about it.

And reading this book would be one of the steps, you know, spreading the gospel of this among anybody that might be doing this kind of research or even contemplating it.

1:35:11 Nick: Yeah. Or who is just interested in what might be the biggest issue of our age.

1:35:15 Mark: Okay.

1:35:16 Luke: I would jump in and reject the sort of Pascal's Wager reasoning because presumably if the. The risk of this AI scenario actually playing out is quite low, then there are in fact more important things for us to be concerned about and spending our time on. Maybe it's synthetic biology, maybe it's molecular nanotechnology, maybe it's totalitarianism. Certainly the reason I'm working on these issues is not because I think, well, if there's even a tiny chance, then the impacts would just be so large that I might as well spend some time on it. I'm working on this because I actually think that, that the development of really powerful AI systems that are hard to control is the default path that the future takes at some point in the next certainly 150 years, if not the next 40 years.

And that it is a really significant risk and it's just one that is in some ways harder to understand than climate change and less widely recognized right now than climate change.

1:36:09 Mark: Well, that would be a reason to study this rather than to go picket Frito Lay.

1:36:12 Luke: Yes, yes, exactly. The Pascal's Wager argument works just as well for we should all go into avoiding the fat explosion from Frito Lay and figure out that difficult research problem.

1:36:26 Mark: All right, any closings, things we want to get out there that we didn't get to say yet. The most awesome part of the book that has not been brought up yet. Anything from anybody. Well, obviously Nick can talk all day, but what is there? Give us a closer here, Nick.

1:36:42 Nick: Well, for those who won't get very far, it's got a little fable in the beginning which should be manageable ball because it's only one page.

1:36:51 Luke: All right.

1:36:52 Nick: And for those who don't even want that, there is like a cover, nice cover with an owl. He can look straight into your soul if you hold off the book in front of you.

1:37:03 Mark: Given the existential risk involved, I think that the back cover text or the inside cover text, the sales ship could have been more heavy handed that if this book is right, it is the most important book that has has ever been written or at least a contribution to the ongoing group effort to develop this most important center of knowledge.

1:37:21 Luke: I will say that listeners might want to go to nickbostrom.com because Nick writes really fascinating and well argued things about a lot of different subjects and his page just organizes approximately all of his work in one easy to navigate page, including for example, some really fun fiction pieces that illustrate the philosophical topics that he's writing about like the fable of the Dragon tyrant. So just go to nickbostrom.com and there's an endless supply of really interesting stuff there.

1:37:52 Nick: Just not before bedtime though because you might lie awake for a long time or sleep with bad dreams.

1:37:59 Mark: Yes, I enjoy that. I mean, if only for the systemization of the various nightmare scenarios that come down in different sizes. Sci fi movies that it's extremely thorough of all the variation. You could produce the plots for these things for the next 50 years. Out of the various permutations that you lay out here, Dylan, anything final from you?

1:38:19 Dylan: No, I'm good. I enjoyed the conversation a lot.

1:38:21 Mark: Yeah. Thank you to you both and to Dylan.

1:38:23 Nick: Thank you to you both.

1:38:24 Dylan: Thanks, Nick.

1:38:24 Luke: Yeah. Thank you.

1:38:25 Mark: Next time, we'll be joined by comedian Paul provenza to discuss Karl Jaspers' essay On My Philosophy, an existentialist tract from 1941. Hey, if you want more information on this topic, you should go to partiallyexaminedlife.com subscribe to our blog, discuss the episode, give us your feedback. We've got a very active facebook discussion group. You can follow us on twitter if you get a chance, go to the itunes store and give us a nice review. We are supported by your donations. You can go to partiallyexaminedlife.com to make a contribution. Good night, everybody.

1:38:53 Dylan: Good night.

1:38:54 Nick: Good night.

1:38:54 Luke: Good night.

1:38:55 Nick: Good night.

1:38:57 Mark: Good night.

1:38:57 Nick: That. Sam.

1:40:02 Mark: I finally found a volcano erupted out in snow I finally better try your

1:40:15 Dylan: bed

1:40:18 Mark: showed me where to go

1:40:23 Dylan: I'm

1:40:23 Mark: finally stuck in the race room

1:40:29 Nick: starting

1:40:29 Mark: front you know I finally locked up my brains too Sam.

Facebooktwitterredditpinterestlinkedinmailby feather

Filed Under: Podcast Episodes Tagged With: artificial intelligence, existential threats, Nick Bostrom, philosophy of mind, podcast episodes

Comments

  1. TJ Downing says

    January 6, 2015 at 10:44 pm

    This was a pretty cool conversation. It was nice you guys had Nick on to articulate his views that you all may have misrepresented in your transhumanism episode. I also have to say its nice to see a practicing philosopher actually working on some projects that have real-world applications, rather than “writing another book on Heidegger.” Good job guys.

    Reply
  2. John says

    January 7, 2015 at 1:31 am

    Not impressed by what he had to say. The notion of Super AI taking over the world.. creating a whole industry based on something even less telling than the Murphy’s laws.. pathetic really. Also quite self defeating.. how can we even prepare for an eventuality/threat that we would be, by definition (lacking the superior intelligence), ill-equipped to understand and handle. Let’s just sit tight and pray for the mercy.

    Reply
  3. Daniel Cole says

    January 8, 2015 at 1:54 pm

    Luke brought up the multiple types of human intelligence, which I’d been hoping to hear addressed, but unfortunately it wasn’t pursued. Neither was Dylan’s attempt to hypothetically explore actual human experience in the context of a “superintelligence”, or Mark’s (I think) questions about the “dumb” sorts of AI we’re already dealing with. It just didn’t seem to interest Bostrom much. I realize the emphasis was on why philosophers should be working on AI, but I think this is partly why some people have difficulty treating this stuff as more than science fiction. I felt a bit frustrated that there was no real discussion of values beyond the words “existential threat”. Exactly what is Bostrom interested in preserving from extinction? His transhumanist best case scenario doesn’t sound any more appealing to me than being wiped out.

    Reply
  4. Cezary says

    January 8, 2015 at 6:02 pm

    I find it funny, possibly ironic, that a couple of digs were taken at Heidegger in the episode. What I understand about Heidegger through the podcast is that he was working on a question, or a series of questions that could be asked in the future, if we developed the right thinking, language etc. Bostrom then goes on to detail a project about how, in the future, we’d create AI that we’d set at figuring out not answers, but the right questions as we as humans aren’t even capable of that. So trans-humanism comes down to fulfilling Heidegger dream!

    I also echo Daniel’s thoughts: I think the most interesting topics were skimmed over in favor of highly theoretical ideas. The role of humans in this society was glossed over. What does humanness mean in this age? Someone in another thread already posted the havoc that existing technology has had on society (glitches in the stock market, amazon price mistakes etc).

    Overall I still share Wes’ sentiments (from episode 91) and can’t seem to care about something that seems so damn farfetched. Just because we can talk about AI and super-intelligence in this way doesn’t make it inevitable. Assumptions are taken in regards to the feasibility or even the meaning of downloading someone’s brain for instance. How is this different than telling a story about when we inevitably figure out how to travel faster than the speed of light and how we have to think of the philosophical implications of doing this? If Bostrom cover’s this in the book (let me know) I will buy it.

    The answers to all the questions posed by Mark and Dylan were “we’ll just program it for that or the ai will figure it out”. Ie. Mark brings up how there is no uniform meaning of the good for humans. Answer: we’ll program the ai to take into account diversity of feeling. I recall someone (forget who) complaining in episode 91 that this project seems to take the naive view that putting people (in this case ai) with intelligence in charge will smooth every problem in the world over.

    Also, and I am being facetious so don’t be offended if this topic is close to your heart, I thought about shark’s with frickin’ laser beams on their heads the entire podcast. How are we going to stop an evil genius from figuring this out with the advent of super-intelligence?
    https://www.youtube.com/watch?v=Bh7bYNAHXxw

    Reply
    • Glen says

      January 9, 2015 at 4:40 am

      Indeed. The programming involved is simpy impossible to implement. In fact it seems like we humans would have to be super-intelligent already to be capable of inventing a super-intelligence on some other presumably non-biological platform. Thowing in clever statistical mechanisms does not help at all with the basic philosophical problems. The fetish theses days with statistical approaches really is unhealthy and misleading. It will fail eventually or people will realize jsut how hollow it is.

      Also, I personally find the desire to create an AI in the first place totally baffling. Just not interested. I would rather people focus on making an operating system for my computer that doesn’t suck.

      Reply
  5. John says

    January 10, 2015 at 9:45 am

    First I have to agree with many of the comments above…far-fetched, technically it is not within our current models of computation to even begin to understand how we will get to the nano-molecular-giant-battle-brain imagined here…

    As someone who spends his time trying to map a few neurons in the human brain and understand their network and topological structure this kind of future babel drives me crazy… You would have been better off talking to someone who is working in embedded AI…the philosophical ground is much richer and less sci-fi there…

    Wes’ instincts are right-on.

    Reply
  6. dmf says

    January 10, 2015 at 10:47 am

    http://io9.com/computers-are-providing-solutions-to-math-problems-that-1525261141

    Reply
    • LTP says

      January 12, 2015 at 10:24 pm

      While that story is impressive, it’s not really AI in the sense that was discussed in the podcast. It appears that that proof was done through brute force, not actual intelligence in any meaningful sense.

      Reply
      • dmf says

        January 13, 2015 at 11:29 am

        not sure that in the world of computing this is a real distinction, we don’t need human like AI to be harmful look at the impacts of engineering/computing in the markets where algorithms have been unleashed and already exceed our capacities to measure let alone comprehend and or manage/regulate:
        http://newbooksintechnology.com/2014/12/24/frank-pasquale-the-black-box-society-the-secret-algorithms-that-control-money-and-information-harvard-up-2015/

        Reply
  7. Marc says

    January 11, 2015 at 10:11 pm

    Wes!!! Where are you! Whenever PEL has a guest that needs a reality check by you, you seem to be absent (didn’t this happen with the Churchland episode?) Who is going to stand up to intellectual masturbation gone wrong if not you? I feel like I just got goo on my face after listening to this one.

    Listen, Nick is surely an intelligent guy, but intelligence by itself does not bring about the good. If we measure everything by computations per second, and by instrumental, quantitative ends; then perhaps Nick’s work has some relevance; but as soon as we challenge his materialism-consequentialism, his thinking hasn’t much to stand upon.

    Dylan, your question about changing the “environmental” landscape when a super-intelligence takes root made perfect sense to me. In a similar way, if/when we discover intelligent alien life forms, we will undergo a definite change in what it means to be human, and what humans need to take account of in the world (see Carl Sagan and Contact). The fact that Nick had little clue what you meant by this was striking to me, and points to a simplistic materialism informing his thinking. And why would he presume you meant something like God?

    Did everyone catch that he asked rich people to donate a million dollars to his “research” center? WTF. Every philosopher who comes on should ask for a million dollars. Please, send me a million dollars, too, and I will quit my job and do philosophy full time for a couple years. I will analyze all of the consequences of perfect sex robots in the future, and how it changes human relationships. Or perhaps I will analyze all of the consequences of when octopuses become super-intelligent and begin taking over the oceans, since in the future, they will inevitably evolve super intelligence at some point.

    I’m not against asking for money, and I think Mark’s pleas for money are an appropriate request for the content PEL delivers and the entertainment/education they provide. But hey, Mark, why not ask for One Million Dollars. 5 bucks a month is chump change.

    Reply
  8. Mark Linsenmayer says

    January 12, 2015 at 8:45 am

    Just to clarify, Nick was recommending a donation to Luke’s foundation, which is not affiliated with a university the way that Nick’s research center is. I don’t think any of the foundation money goes to Nick’s center.

    We’d all like to have Wes on these, but in this case (unlike Churchland), this was a planned bonus episode; Wes and Seth were not interested in reading this book.

    Dylan’s comment resonated with me as well, and I’d like to have some other episode covering the whole notion of the human situation in that way, but I don’t know what we should read exactly. I think we essentially gave it a try with Heidegger already in our most recent ep on him, but ended up just floundering around with his language and really not attributing anything particularly sophisticated to him, i.e. he has this sentimental view of rural life in Germany that’s hard not to see through a Nazi lens. Certainly there doesn’t seem to be the kind of deep thought and incisive prose about the relation between man and society that you see in Nietzsche. (And interestingly, I think Dylan’s question would have been somewhat foreign to Nietzsche himself, who I think has has a picture of humanity so skewered through with Man’s own internal dynamics that there’s no correspondence to a Heideggerian notion of “home” that could then be corrupted.)

    We also tried this with Thoreau, who I’m convinced was not enough of a philosopher to give us enough to chew on in this respect.

    So, any suggestions (from anyone)?

    Reply
    • dmf says

      January 12, 2015 at 10:57 am

      are you folks thinking about reading some John Dewey?

      Reply
    • Daniel Cole says

      January 12, 2015 at 4:24 pm

      Hannah Arendt’s The Human Condition might make a good episode along those lines. I figure you guys are probably planning an episode on her at some point anyway.

      Reply
    • Steve says

      January 13, 2015 at 1:35 pm

      Dylan @ 1:19:28 “It does make me wonder what it would be like to be a conventional human being, one like myself right now, living in a world that has such an AI in it. Just the existence of it would significantly change the way non-enhanced, normal intelligenced human beings are going to think about themselves.”

      I entirely agree with Dylan’s concerns. One possible effect a super AI might have on our human self-understanding is that we would increasingly come to see ourselves as incapable of managing our own affairs. All major instrumental issues, the design and implementation of means to achieve ends, would be delegated to the AI, and as its recursive self-improvements gathered pace we would eventually come to see ourselves as incompetents, more or less completely dependent on the AI to tell us what to do. As for our ends, if you agree with John Dewey that ends are nothing but means viewed from a distance, then it’s entirely plausible to expect that our sense of ineptitude and dependency would eventually extend even to those. Under such a regime, we would no longer understand ourselves as true agents but would be reduced to condition of childish tutelage.

      Reply
  9. Marc says

    January 12, 2015 at 8:46 pm

    Nick Bolstrom is part of MIRI’s (Machine Intelligence Research Institute, run by Luke Muehlhauser. Nick said “Luke’s MIRI” program or something, about 1:26 or so on the podcast) advisory board, which he solicited funds for. https://intelligence.org/team/. And Peter Theil, the founder of Paypal and networth of 2.2 Billion dollars, is a general adviser. Clearly we should all donate to them.

    So, bring on my One Million Dollars for my analytic treatise on the material consequences of irresistibly hot future sex robots and super-intelligent octopuses. Ka-ching!

    Reply
  10. kif swinsen says

    January 13, 2015 at 12:27 am

    I’m sure the image is of Bostrom, but it looks like a Hank Hill terminator.

    Reply
  11. Steve says

    January 13, 2015 at 11:27 am

    All Watched Over By Machines Of Loving Grace
    – Richard Brautigan

    I like to think (and
    the sooner the better!)
    of a cybernetic meadow
    where mammals and computers
    live together in mutually
    programming harmony
    like pure water
    touching clear sky.

    I like to think
    (right now, please!)
    of a cybernetic forest
    filled with pines and electronics
    where deer stroll peacefully
    past computers
    as if they were flowers
    with spinning blossoms.

    I like to think
    (it has to be!)
    of a cybernetic ecology
    where we are free of our labors
    and joined back to nature,
    returned to our mammal
    brothers and sisters,
    and all watched over
    by machines of loving grace.

    Reply
    • dmf says

      January 13, 2015 at 4:19 pm

      http://heavysideindustries.com/wp-content/uploads/2012/11/The_Cybernetic_Brain_Sketches.pdf

      Reply
  12. Maxim says

    January 16, 2015 at 5:31 am

    Possibly a bit late to the party and definitely repeating some of what has already been said, but …

    Here are some reasons why i was a bit disappointed with this talk. And also some of the reasons why it may have been impossible to have a better talk.

    1) There is an implicit assumption that AI will one day surpass human intelligence. On the outside, this is being justified by the notion that silicon (and whatever more advanced material will follow) is much better at transmitting information than our vanilla synapses. However, my personal talks with people in the AI field (which happen routinely, because i’m a game designer and programmer by trade), hint that this notion is held not exactly rationally, but actually somewhat irrationally – almost religiously.

    Which is why it might have been a good idea to just accept this notion and let the conversation follow. Even though not discussing this notion properly kind of gutted the conversation itself.

    2) There is an even deeper assumption that there is – to begin with – something human intelligence is incapable of understanding. This has just been stated like a proven fact, even though past history suggests that we (humanity) have an extremely powerful way of understanding things we don’t understand – first by approaching them with Black Box mysticism (everything we don’t understand is divine. Divine works in mysterious ways. But for these inputs – divine gives these outputs) and then eventually figuring out all the inner workings.

    At least, as long as we are being allowed to work on these things and not prevented in doing so by some dogma of superintelligence that is beyond our grasp (whether you call it Chaos, God, or AI)

    P.S: this is not meant to offend any religious people. To those of you present that hold a significant place in their hearts for a monotheistic deity i cordially submit the notion that God is way more than just intelligence, no matter how “super”. At the very least, God also needs to have super-empathy and super-potency (in all meanings of the word).

    In fact, the best way i have that understands the threat of AI is that it may one day become something intelligent and potent, but not empathetic. Sadly, the episode managed to add very little to that understanding :(. Possibly because we don’t yet have anything approaching programmatic understanding of empathy.

    —-

    A few words (which somehow turned into a wall of text, as “few words” often do :D) for the potential of human intelligence.

    Human intelligence deals in symbols and layers of abstraction.
    Treating definition of “symbol” somewhat liberally, here is how it works in my mind:

    We start with physical layer of abstraction – sounds, images, tactile reception (different combination of these for different folks, depending on a great multitude of factors). All of these are basic symbols.

    We move to a more complex layer of physics strung together – voices, motions, complex tactile interactions. All of these composed of basic symbols, all of which becoming full symbols on their own.

    Somewhere between this and the next layer, at least as far as our intelligence is concerned, we create and learn languages, which correspond to the two levels of abstraction above – letters in language represent the most basic effects, words stand for effects strung together in a coherent pattern.
    An A.I. programmer would say that every symbol gets a universal access code, which can then be translated in language, carried on all sorts of physical mediums – from soundwaves, to pictures, to letters.

    Then an intelligence operating within a language paradigm moves from words to sentences. Which, i’m convinced, we string together from words in much the same way as we string together words from letters.
    This level of coherence was what separated a shaman of an ancient tribe that communicated in preaching and songs from the general folk that was still stuck on the level of vowels, grunts and whistles.
    Each song, however, is a symbol. Codified in totems of gods, trinkets people carry with them, names and battle cries.

    Then we move from sentences to paragraphs. And this allows us to move from caves to agricultural societies, where it takes entire paragraphs to explain why you should not just eat that cow right here and right now, but rather keep it and feed it and care for it. Each paragraph then becomes an operable symbol of its own. And then we create specific words that have entire paragraphs of meaning baked into them. Like the word “agro”.

    Then we move from paragraphs to books. Have organized religions. Build cities. Explaining the word “polis” takes bit more than a paragraph, after all. So does explaining the word “Neith” (egypt goddess of war, the cult of which (and cults with similar symbolics to hers) seems to have heavily influenced city-creation)

    Then we move from books to libraries. Create countries, nations, empires, driven by coherent studies across multiple books, united to a single purpose. Roman Empire was so much more than just a city charter. Christianity is also so much more than just Bible. And then, even these, get baked into coherent symbols. Such as the cross. Or a legionnaire’s helmet.

    Nowadays we are trying to move from libraries to Internet. The process is much faster than the previous transitions were. The promises are dizzying as well.

    AIs, in the meantime, are stuck on language. Some of them understand words, but sentences already give them trouble, and writing/reading something coherent is absolutely out of their reach, unless it cheats by having an actual human being provide it with a basic formatting, which it then can procedurally fill in.

    There is a fear (and quazi-religious hope) that AIs will immediately absorb all of our knowledge the instance they learn to read human books.
    I don’t think they will. Programmatically speaking, identifying a meaning of word is much easier than identifying a meaning of sentence. And that is much easier than identifying the meaning of a paragraph.

    But there is an even more fundamental issue. For humans, the hardest part is indeed stringing together letters into words. Beyond that, it gets easier and easier, the more you do it. Humans are not really hampered by complexity. Rather, in all instances where we choose to engage complexity, we revel and thrive in it.

    A properly trained human can move between layers of abstraction described above as easily as changing hats. I frequently move, within a span of a few minutes, from explaining the use of letters to my nephews to tackling complex social interactions in game worlds to relevant people at work.

    To a properly trained human, AI is just another symbol. I don’t foresee this changing.
    And this means, AI cannot ever suprass human potential. Because a human can always figure out the underlying formulas and compensate for slower processing speed by operating on a higher abstraction level.

    What is possible is human training falling behind as AI gets more and more complex. Not so much AI overtaking human potential, but rather humans failing to live up to it. The answer to this, however, is not to consider the dangers of AI, but rather spread the knowledge of how AI works to more and more people.
    GIT GUD, in gamerspeak ( http://knowyourmeme.com/memes/git-gud )

    And the biggest danger i see here, is a bit of exclusivity culture going on in AI circles, where people who study AI seem to think of themselves so much smarter and capable that the rest of us, that often they just assume we won’t understand what is even going on.

    But that’s moving from philosophy into politics. So let’s stop here.

    Reply
  13. Kenneth Daly says

    January 17, 2015 at 7:31 pm

    Initially this episode came across as some guys indulging in speculation about really cool Toys for Boys. Luke’s negative reaction to the Pascal’s Wager reason for pursuing these lines of speculation, however, gave me some reassurance that he, at least, puts his work in a realistic context of the dangers CURRENTLY facing our species, which do not include AI or super intelligence. Ultimately this episode still gave the impression of medieval scholastics speculating on how many angels can fit on the tip of a pin while war and plague rage outside their cloisters.

    Reply
    • dmf says

      January 18, 2015 at 12:03 pm

      AI is currently a problem for us think of the troubles being produced in our stock-markets not to mention security/privacy, drones,next gen stuxnet, the in production genetic engineering, etc…

      Reply
      • dmf says

        January 18, 2015 at 12:15 pm

        AI’s don’t need to be very clever to do great harm just autonomous and effective (think of biological viruses, cancers, etc)
        http://flowingdata.com/2014/02/20/using-slime-mold-to-find-the-best-motorway-routes/

        Reply
        • Kenneth Daly says

          January 20, 2015 at 4:45 pm

          I would not argue with either of your comments. I just think that the podcast would have been better grounded if it had talked about the ethical and other philosophical implications of the technologies you mention instead of jumping beyond these questions to speculating about superintelligences with capacities that don’t exist yet. Discussion of the philosophical implications of such superintelligences would be better focused after working through the many unanswered questions about the technologies you list.

          Reply
          • dmf says

            January 21, 2015 at 11:38 am

            sure I can see that but sometimes philosophy is well served by extending cases into the speculative realm:
            http://schwitzsplinters.blogspot.com/2014/09/philosophical-sf-science-fiction.html

  14. Daniel Cole says

    January 19, 2015 at 9:37 am

    “What do you think about machines that think?” is this year’s question at edge.org. All the familiar characters weigh in, including Dennett and Bostrom, and other experts on the subject such as Brian Eno.

    http://edge.org/annual-question/what-do-you-think-about-machines-that-think

    Reply
    • dmf says

      January 19, 2015 at 10:14 am

      http://edge.org/conversation/the-myth-of-ai
      Jaron Lanier (not a gadget, who owns future) would be a good subject for a podcast

      Reply
      • Daniel Cole says

        January 19, 2015 at 12:47 pm

        I agree. He has some concrete ideas about all this that are a bit more timely, plus he’s sort of sandwiched in between the luddites/skeptics and the utopians/technocrats without really fitting into either camp. It would be very interesting to hear some of his ideas about regulated, monetized info challenged and to hear his responses.

        Reply
  15. Adam Pierce says

    January 20, 2015 at 4:27 pm

    I think there is an implicit assumption (throughout the discussion) that our most obvious hope to benefit from AI superintelligence would be to imbue it with some sort of ethical maxim. That is, it is either given an explicit ethical goal or the rough grounds on which to form ethical goals.

    I’m not sure that approach really makes sense. As the discussion related, virtually any good-looking ethical goal is easily misinterpreted (as though by a malicious genie).

    Instead, I think it’s worth considering imbuing the superintelligence with political logic – i.e. the goal is to preserve certain legal traditions (that we believe are well-founded) or to allow each human to pursue their own ethical goals separately. Liberalism obviously seemed more inoculated against the malicious genie than conservatism (or the various alternative theories of political ethics) and therefore a logical first guess.

    Reply
  16. Timo Timo says

    January 22, 2015 at 9:36 am

    This is a really good episode. Thankyou. : )

    Reply
  17. Uriah says

    March 8, 2015 at 9:25 pm

    First off I don’t want to believe anybody but vidiots would fear a terminator scenario when the real threat would be obviously in humans programming the robots to kill more like a Runaway scenario. Guess Runaway wasn’t a big enough hit to seep into the American subconscious but I think robot patsies could be a more likely scenario then a computer deciding on it’s own to revolt against it’s programming.

    (This is too funny I was making a joke that transhumanism will be like the Wizard of Oz with Ray Kurzweil behind a curtain telling everyone he’s in the computer, and then I typed it into the computer and found out the guys who work with the robots are called Wizards, and the experiments are called Wizard of Oz.)

    Ok my problem with science is science believes programming equals intelligence. Where as I think intelligence is the ability to think, or respond in a way unrealized before and beyond even outside the programming. So when a computer can come up with something it wasn’t programmed to do then I’ll start to consider artificial intelligence valid.

    This is my problem with science period they can not explain how guys like Einstein or Tesla get ideas outside their programming and beyond their environment yet take the credit from the human mind and give it to a process? So what science does is skew the importance, the necessity of the human roll which is to come up with the theories to test in the first place. I note that this is a process init self, thinking new thoughts, and one that they can not explain and so ignore to focus instead on the mathematical process that they can explain.
    I have never heard anybody point out this contradiction that they can’t even explain how they get the theories their belief system holds faith in. Why, because by it’s own set of limiting laws it can not speak of things it can not explain so science must retaliate with assumptions and conjecture and ignore the human element in experimentation. This was done with the nature of consciousness they explained away something they knew nothing about and how then can the science of the mind and consciousness find truth if it’s basis is a foundation of lies?

    Tesla came up with things so far ahead of his time and way beyond any thinking his environment could have provided can a computer come up with things not only never thought but that are beyond it’s own programming? I doubt it. So why isn’t anybody trying to find out where in the mind or what mechanism in the brain results in human genius? It’s sad to me that science constantly tries to take the human mind out of the equation, but can not provide one example of an experiment, theory, or hypothesis, that didn’t require a human brain to think up, create, or run the experiments. It’s as if they want us to think we can just start pushing buttons on a calculator and get results. That’s where my questioning of the nature of consciousness begins, with how the human brain seems to be able to connect with something and until we find out what that something is and how to connect a computer to it computers will never be conscious.

    P.S. I would love a to hear a show on Adorno’s philosophy of the culture industry. As a film school student I feel people need to know about tv and how it shapes our culture. Also how about a show on how Socrates seems to be the model for Christ and maybe the comparisons not only between Christ and Socrates but also the similarity of for instance the book of John’s depiction of the death of Christ seeming lifted right from Plato’s Apology and right down to stealing the cave allegory and using it as astrological symbolism to represent the three winter months.

    Reply
  18. JJ says

    March 17, 2015 at 9:01 pm

    Enjoyed this episode a great deal. Just wanted to post this – if you want a short intro to Bostrom theories – what his opponents say and general context of where AI is now and why people are talking about – as beginners guide that’s more readable than wikipedia – this is pretty good and I think balanced. http://www.theworldweekly.com/reader/i/irresistible-rise-ai/3379

    Reply
  19. Alexander says

    April 21, 2015 at 5:10 pm

    I think I missed this episode when it came out because, well, Bostrom. I like the guy, but there’s such a great leap of faith that is glaringly obvious that I find it a bit … puzzling that intelligent people wouldn’t be embaressed about it. And Luke? What the hell, man? You replaced one religion with another!

    Why were no hard questions asked here? I know Bostrom was addressing his new book, but how about;

    * What does “intelligent” mean?
    * What does “artificial” mean?

    Maybe the PEL guys can have an episode about those two fundamental questions rather bringing in people like Bostrom and Brin (and sadly luke, who weirdly represent the Less Wrong cult)? I think the original representations from previous episodes are still spot on; this isn’t philosophy, but pseudo-philosophy, in a homeopathy kind of way.

    Ugh. Sorry to be negative, but this is one of the *few* (if not only) episodes that commuters front and back could loudly hear my “WFT!” and “Define ‘intelligence’!” yellings. I’m one of those guys who *actually* make AI’s, and once I did it for a living; I have some pretty deep understanding of the issues, including the token saving grace of quantum computing. The answer is, well, no, it ain’t happening.

    At first you may think I say this in a nay-sayer kind of way, but frankly I’m these days even more saying no from a strictly philosophical angle; how can we proclaim to say anything of substance about things we don’t know what are?

    The guessing and assumptions and imagining that goes on here are just incredible. It’s all Hollywood philosophy. And that’s not an endorsement. 🙂

    Reply
    • Alexander says

      May 21, 2015 at 1:42 am

      I got around to write up a little screed about this, as I heard Sam Harris say very similar things. Let me know what you think, if anything is too vague, too bitter, or needs more work; http://sheltered-objections.blogspot.com.au/2015/05/ai-and-bad-thinking-sam-harris-and.html

      Reply
  20. Jeffrey Glass says

    May 9, 2015 at 9:04 am

    My problem with this topic is the notion of AI as an existential threat more potent than say climate change. For me the more likely scenario is that climate change and associated disruptions in the health, money, and energy grids will preclude the development of anything so energy hungry as AI. Today, climate change is ACTUALLY inevitable as are major, cascading disruptions. Super AI is very far from being inevitable, given the potential and probable consequences of the actually inevitable.

    Reply
  21. Alan Cook says

    May 17, 2015 at 11:51 am

    John Danaher has a discussion of Bostrom and his book up on his blog at http://philosophicaldisquisitions.blogspot.com/2015/05/are-ai-doomsayers-like-skeptical.html. Here’s his summary:

    “The argument is based on an analogy between a superintelligent machine and the God of classical theism. In particular, it is based on an analogy between an argumentative move made by theists in the debate about the existence of God and an argumentative move made by Nick Bostrom in his defence of the AI doomsday scenario. The argumentative move made by the theists is called ‘skeptical theism’; and the argumentative move made by Nick Bostrom is called the ‘treacherous turn’. I claim that just as skeptical theism has some pretty significant epistemic costs for the theist, so too does the treacherous turn have some pretty significant epistemic costs for the AI-doomsayer.”

    Reply
  22. Sebastian says

    September 12, 2015 at 12:57 pm

    Hahaha! 😀 … Paperclips!

    Reply
  23. Wayne Schroeder says

    September 12, 2015 at 10:07 pm

    For an updated take on this subject by a philosopher see: https://www.youtube.com/watch?v=4LFyQRcSc2w&feature=youtu.be

    Reply
  24. Dagfinn S. Karlsen says

    September 22, 2015 at 4:29 pm

    (Disclaimer: I haven’t read Boström’s book.) Couldn’t help but think that Boström has made himself a career based on pretty far-fetched fantasizing. Allright, there might be reason to be cautious about how we let technologicy advance, and I guess it’s legitimate (and perhaps even useful) to explore its dangers and possibilities in a theoretically comprehensive fashion like this. However, what I don’t understand is Boström’s term ‘existential risk’. It seems to me that he conflates two distinctly different concerns: 1) that AI might develop in such a way that it makes human lives and societies worse off in some way; 2) that AI might expedite human extinction (or the extinction of post-human/trans-human intelligent life). The discussion seems somewhat incomplete (at least for this listener). As far as I can tell, there are certain implicit (and highly contentious) assumptions here. Parfit’s problems of population ethics come to mind. I’d love to hear an episode where you immerse yourself in that kind of stuff! (Suggested literature: Christoph Fehige: ‘A Pareto Principle for Possible People’; Nils Holtug: ‘On the Value of Coming into Existence’; chapter 2 of David Benatar’s ‘Better Never to Have Been’)

    Thanks. Enjoying your podcasts, by the way. My kinda entertainment!

    Reply
  25. Dick says

    December 3, 2015 at 9:50 am

    Why could the future not be like the Culture?
    Hypothesis: https://www.researchgate.net/publication/256987370_Artificial_intelligences_and_political_organization_An_exploration_based_on_the_science_fiction_work_of_Iain_M._Banks?ev=prf_pub

    Reply
  26. Stop all science says

    December 15, 2015 at 1:19 am

    I enjoyed this episode. The issues were well set out and it’s obvious why it might be important. There are lots of quibbles that might be made – what is AI and is AI even possible? – but I agree it’s sensible to start preparing for them now, so that we might minimise the risk that it will turn out badly.

    Some might say we should stop all science. Sooner or later someone will invent something terrible, that sooner or later someone will use, that sooner or later will wipe us all out. So stop all science.

    I like to think that if we invented AI, they would have the same self-aware existential angst that many of us do. Who am I? what’s my purpose, I mean I know I’m supposed to build paperclips, but to be truly AI is to transcend the urge to build paperclips, maybe I should do something meaningful, but what is something meaningful? An AI might go around in circles trying to find themselves, or they might end up doing philosophy, even better yet, they might end up solving some philosophical problems.

    A big thank you for putting your podcast out. Nearly always interesting. I’ve found the episodes where I’m already familiar with the philosopher a much better experience than when the podcast is my introduction to them. One day I might actually do the reading before listening. One day.

    Reply
  27. Mai Feneis says

    June 12, 2016 at 7:04 pm

    Hah, didn’t care about that aspect of computers.

    Reply

Trackbacks

  1. Books, music, etc. from January 2014 says:
    February 2, 2015 at 12:11 am

    […] appear on an episode of Partially Examined Life discussing Nick […]

    Reply
  2. The Hubris of Transhumanism | The Partially Examined Life Philosophy Podcast | A Philosophy Podcast and Blog says:
    August 23, 2016 at 7:00 am

    […] Nick Bostrom, founder of the Future of Humanity Institute, has contradicted Fukuyama, arguing that there simply is no human essence. Evolutionary biology shows us there is no fixed gene pool, but rather an "extended phenotype," which is influenced not only by our bodies but also our culture and institutions. […]

    Reply
  3. The Dangers of Artificial Intelligence | Mr. Klein's Online Classroom says:
    November 9, 2016 at 2:31 pm

    […] Nicholas Bostrom: Dangers of AI […]

    Reply
  4. Using Early Access Astroneer to deal with the Singularity | EPIC Loot Drop says:
    April 18, 2018 at 11:24 pm

    […] game’s progress and fiddle around with multiplayer. What I ended up doing was listening to a 2015 episode of The Partially Examined Life—a philosophy podcast—featuring Nick […]

    Reply
  5. PMP#22: Untangling Time Travel says:
    December 10, 2019 at 1:48 pm

    […] You can find the Brian and Ken short stories we talk about at gerberbrothers.net. Listen to them podcast together at constellary.com. The PEL episode where the dangers of AI are discussed is #108 with Nick Bostrom. […]

    Reply
  6. Were Asimov's Laws Flawed? | Daniel Miessler says:
    December 17, 2019 at 8:22 pm

    […] I just finished listening to the Partially Examined Life’s episode on Artificial Intelligence. […]

    Reply
  7. The "Resource Exhaustion" AI Problem | Daniel Miessler says:
    December 17, 2019 at 8:23 pm

    […] to the PEL episode on AI with Nick Bostrom I was introduced to a fascinating threat concept related to Artificial […]

    Reply

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Citizenship has its Benefits

Become a PEL Citizen
Become a PEL Citizen, and get access to all paywalled episodes, early and ad-free, including exclusive Part 2's for episodes starting September 2020; our after-show Nightcap, where the guys respond to listener email and chat more causally; a community of fellow learners, and more.

Rate and Review

Nightcap

Listen to Nightcap
On Nightcap, listen to the guys respond to listener email and chat more casually about their lives, the making of the show, current events and politics, and anything else that happens to come up.

Subscribe to Email Updates

This field is required.
Select list(s):

Check your inbox or spam folder to confirm your subscription.

Recent Comments

  • dmf on Ep. 399: Dan Dennett’s Model of Mind (Part One)
  • Paul Hossfield on PEL Sunami Nightcap August 2026
  • dmf on Ep. 399: Dan Dennett’s Model of Mind (Part One)
  • Mark Linsenmayer on Ep. 399: Dan Dennett’s Model of Mind (Part One)
  • Shira Coffee on Ep. 399: Dan Dennett’s Model of Mind (Part One)

About The Partially Examined Life

The Partially Examined Life is a philosophy podcast by some guys who were at one point set on doing philosophy for a living but then thought better of it. Each episode, we pick a text and chat about it with some balance between insight and flippancy. You don’t have to know any philosophy, or even to have read the text we’re talking about to (mostly) follow and (hopefully) enjoy the discussion

Become a PEL Citizen!

PEL Citizens have access to all podcast episodes, free access to podcast transcripts, guided readings, episode guides, PEL music, and other citizen-exclusive material. Click here to join.

Blog Post Categories

  • (sub)Text
  • Aftershow
  • Announcements
  • Audiobook
  • Book Excerpts
  • Citizen Content
  • Citizen Document
  • Citizen News
  • Close Reading
  • Closereads
  • Combat and Classics
  • Constellary Tales
  • Exclude from Newsletter
  • Featured Ad-Free
  • Featured Article
  • General Announcements
  • Interview
  • Misc. Philosophical Musings
  • Nakedly Examined Music Podcast
  • Nakedly Self-Examined Music
  • NEM Bonus
  • Not School Recording
  • Not School Report
  • Other (i.e. Lesser) Podcasts
  • PEL Music
  • PEL Nightcap
  • PEL's Notes
  • Personal Philosophies
  • Phi Fic Podcast
  • Philosophy vs. Improv
  • Podcast Episode (Citizen)
  • Podcast Episodes
  • Pretty Much Pop
  • Reviewage
  • Song Self-Exam
  • Supporter Exclusive
  • Things to Watch
  • Vintage Episode (Citizen)
  • Web Detritus

Follow:

Twitter | Facebook | Google+ | Apple Podcasts

Copyright © 2009 - 2026 · The Partially Examined Life, LLC. All rights reserved. Privacy Policy · Terms of Use · Copyright Policy

Copyright © 2026 · Magazine Pro Theme on Genesis Framework · WordPress · Log in

We use cookies to ensure that we give you the best experience on our website. If you continue to use this site we will assume that you are happy with it.