The idea would be that in the same way that we have laws that say do not drive in these particular ways, you could have sets of numbers that say do not generate or propose any molecules that bind to these receptors any more strongly than these other molecules we know of. Don't bind to hemoglobin more than carbon dioxide. That's an important thing. There's a very meaningful way these preferences and these applications are cultural, personal, intersubjective.
You can't define toxic in a meaningful way unless you pick some very narrowly skilled thing. Like, oh, did they kill 50% of the mice at that dosage? And that's what we typically use. But it would be really nice to have something that was much more granular for people who are trying to treat terminal diseases.
So the hope is that you could have a tool that would highly constrain but also be human readable and understandable and contribute to our understanding of if an AI system proposes a molecule to do a thing. This is why we think it does the thing. This is why we can justify that we think it won't kill you. Welcome to the 20th episode of Humans on the Loop.
I'm your host, Michael Garfield. And this week we ask, are we doing AI alignment wrong? Game designers forced Emil and Gavin Valentine to find games as having meaningful decisions, certain outcomes, and measurable feedback. If any one of these breaks, the game breaks.
And we can think about tech ethics through this lens as well. Much of tech discourse is about how one or more of these dimensions has broken the game of life on Earth, the removal of meaningful decisions, the mathematical guarantee of self-determination through unsustainable practices, and or the decoupling of feedback loops. AI alignment approaches tend to converge on a strategy of restoring meaningful decisions by getting rid of uncertainty, but it's a lost cause. It's futile to encode our values into systems we can't understand.
To the extent that machines think, they think very differently than we do. And characteristically interpret our results in ways that reveal the assumptions we are used to making based on shared context and understanding with other people. We may not know how a black box AI model arrives at its outputs, but we can evaluate those outputs. And we can segment processes like this so that there are more points at which to review them.
One of this shows major premises is that the design and use of AI systems is something like spellcraft, a domain where precision matters because the smallest deviation from a precise encoding of intent can backfire. Magic isn't science. Inasmuch as we can say that for spellcraft, mechanistic understanding is frankly beside the point. However you may think of it, spellcraft evolved as a practical approach for operating in a mysterious cosmos.
Westernized modernity dismisses it because enlightenment-area thinking is predicated on the knowability of nature and the conceit that everything can and will eventually bend to principled rigorous investigation. But this confused accounting just reshuffled its uneradicable remainder of uncertainty back into a stubbornly persistent reel that continues to exist in excess of language and mechanistic frameworks. Economies, AI, and living systems guarantee uncertain outcomes. And in accepting this, we have to re-engage with magic in the form of our machines.
The more like us they become, the more mystery and open-ended co-improvisation loom back over any goals of final knowledge and control. In his 2016 essay, Danny Hillis called this the age of entanglement. It is a time that calls for an evolutionary approach to technology. Tinkering and re-evaluating, we find ourselves one turn up the helix in which quantitative precision helps us reckon with the new built wilderness of technology.
When we cannot fully explain the inner workings of large language models, we have to step back and ask, what are our values and how do we translate them into measurable outputs? How can we break down the wicked problem of AI controllability into chunks on which it's possible to operate? How can adaptive oversight and steering fit with existing governance processes? In other words, how can we properly task the humanities with helping us identify meaningful decisions and the sciences with providing measurable feedback?
Giving science the job of solving uncertainty or defining our values ensures we'll get as close as we can to certitude about outcomes we definitely don't want. But if we think like game designers, then interdisciplinary collaboration can help us safely handle the immense power we've created and keep the game going. This week's guest is my friend Evan Miyazono, CEO and Director of Atlas Computing, a tech nonprofit committed not to the false God of perfect alignment but to plausible strategies for provable safety. Focusing on community building, cybersecurity and biosecurity, Evan and his colleagues are working to advance a new AI architecture that constrains and formally specifies AI outputs with reviewable intermediary results, collaborating across sectors to promote this radically different and more empirical approach to applied machine intelligence.
After completing a PhD in Applied Physics at Caltech, Evan led research at Protocol Labs, creating the research grants program and leading the special projects team that created hyper certs, funding the commons, gov, forget and keep parts of discourse graphs and the initial open agency architecture protocol. In our conversation today we talk about regulatory scaling problems, specifying formal organizational charters, the specter of opacity and the quantification of trust, all in some sense measures of game design at the nexus of science, art, magic and religion in our entangled relationship to the fundamental uncertainty of our world. Before we begin, I want to announce a special perk. My recent Weirdosphere course, How to Live in the Future, is now recorded online and available in its entirety for members of the Founders tier on sub-stack, Patreon and every dot org.
Nearly 20 hours of lectures in discussion and an extensive reading list on everything from reconciling relativistic and thermodynamic perspectives on time, self-organization and the evolution of intelligence, the design of human-computer interactions that actually work for humans, the emergence of planetary culture, the evolution of human selfhood, it's like a grab-back of everything I've talked about for the last nine years of this show, synthesized and explored together in an incredibly intelligent community setting. I'll share this list later this week along with all of those recordings to Founders. If you want to go as deep as possible into the biggest ideas we've ever handled here, I'm delighted and very glad that I get to share it with you. Also open invite to all members, we're hosting monthly community hangouts, I've got a bunch of new essays coming soon and excellent episodes in the pipe, including conversations with George Pore, Rufus Pollock, Hyungen Gold, Terence Southern, Andrea Ferris, Amanda Scott and many more.
If you believe in the value of this work, consider becoming a member at humansontheloop.com. Thanks everyone for your support and now enjoy the show. Thank you. Yeah, the reality of it is that by the time this episode comes out, all of us will be living in a computer simulation if we're not already and so the fact that we're recording in 720 won't matter because GPT-9 or whatever will just automatically boost you up to stereoscopic 8K.
Have you seen the papers on the diffusion-based video compression where they send like three images and then some like tiny bitstream and the entire thing is just like reconstructing the image from like near loss look like, at least it looks nearly losslessly from this tiny bitstream? You know, it's so funny because this is actually a good, I don't want to get ahead of ourselves, but whatever we already are. You know, one of the premises of this show and of future fossils dating all the way back is that, you know, we can obsess with production quality as much as we want for contemporaneous audiences, but we're going to just be able to reconstruct everything to whatever degree we want in the future. So really, you can give the future garbage and it won't really be that big of a deal, I think.
I think that the weirdest thing is that you can, I agree that you'll be able to take something that's garbage and put it into the future and there people will be able to do ridiculous things with it. Like make it look like it's not garbage, but people in that future will still know that it came from garbage because the way in which you can add dimensionality to things is going to be something that we're not going to be able to predict right now. In the same way that you can take a black and white photo and turn it into a video, but you and I look at it and we're like, oh, it's a black and white photo because there's no color there. And if you like artificially color it, that might still look like an artificially colored black and white photo.
And maybe the gap between those still gets smaller, but like I think that we'll probably, we or maybe future generations will be able to distinguish between like, oh, this is like a model that was trained on all the great grandpa's writing versus podcasts versus something that was like, oh, this is like a high-fidelity brain scan upload. Yeah. I know it's funny though, because like, you know, you and I have both lived through the argument. Like I, so again, to the both sides of this, you know, I remember a lot of my high-fidelity audio friends being like, you're never going to get as good as vinyl.
Like you never will. And then on the other end, I kept being like, God, who wants to be the MP3 version of themselves? Who considers it satisfactory that your widow would not consider it, even the high-fidelity brain scan is missing the extended self. It's missing, you know, that Gregory Bateson kind of mined it large, extended phenotype, you know, all the epigenetic interactions are not captured in that scan.
And so except no substitutes. But then it's funny because like in the last 20 years, not only has digital audio gotten so good that even vinyl purists realize that it's just an affectation, but also most people just genuinely kind of don't care anymore that they're listening to 256 kilobytes per second MP3s. We're just glazing over the top of it. So I'm curious, there's a paradox in there.
And this is somewhat off the topic of what I wanted to speak with you today about. But like I guess it does matter to the degree that right now you and I care deeply about auditability and the question of like, are your kids going to care about whether they exist in a sort of reactive, per minuteic, like back foot relationship to opaque decision making structures? Like they just may not mind as much as we do. So yeah.
I've definitely had like a strong get off my lawn feeling of like, maybe I'm just getting old, but like I feel like humans should have a shot at understanding. At least collectively what's happening in the world. And I think that that is a trend that is currently breaking. That's not the direction of moving the world.
We're moving in the opposite direction for a while now. And like it's not clear to me, fighting for some of these things that I am like really strongly fighting for like, optimality, human review, adherence to laws and like, why it's like, I don't think that that's like, there are many, many things that are like, here's that thing that we should, that I think we should move forward and what it should look like. And like the long archifistory is like, go on the other way. Well, let's get there.
Let's get there. But first, Evan, you're on the loop. Welcome. Thanks.
This is years and coming where I would like to start so that we can unpack some of this philosophical stuff and then the technical aspects of your project and why this stuff might matter in practical terms, even if auditability does not matter in any kind of intrinsic aesthetic or moral way, right? Like there are probably still good reasons to care about this stuff, even if for all intents and purposes, we don't constitute those values in the same way that we do today. But yeah, I would love to just start with you telling a little bit about who you are, where you come from and why it is that you care about the stuff that you're working on now. Just I think knowing why you left protocol labs to pursue this, you know, why this matters to you so much and where that started in your biography would be really great.
Before I jump into that, I assume you can't hear traffic. My window is currently open because doors are closed up. Not your phone. Would it be poorly distracting if I walk on the treadmill while I talk?
No, actually walking on the treadmill while you talk, I've come to expect that. It's like a weird that you're standing still. So I have all these keep it a little bit slower. Yeah, it's often fun to reflect on how I got to where I am because it's such a weird path that I don't think I could give people the tools to replicate, even if they wanted to, even though I'm not sure if they would have that option.
Definitely was a child of the late 90s and believed that technology would solve problems and ended up at Stanford for undergrad doing material science. And I was largely driven by having trouble picking a major, but also the fact that it seemed like materials were a place of huge promise with lots of added manufacturing printing type stuff. And so I was like optical cloaking and levitation and room temperature superconductors. And this is kind of like the marketing take that material science is giving out early to mid-aughts.
And I wanted to be a part of that. I wanted to build the Star Trek future. And by the time I finished undergrad, I realized I needed more fundamentals, ended up applying to pure only physics and applied physics programs, ended up doing a PhD in experimental quantum optics, trying to build a provably secure quantum internet and made a little bit of progress after a lot of work and decided that I should probably not be an academia and be places where I'm a little bit closer to people that I'm trying to help with things that I'm building and when looking for hardware startups, a buddy of mine from college asked me to help out with an unconventional fundraise he was doing as I was looking for jobs. I asked everyone I could think of, genuinely, for me to interesting ambitious hardware startups.
And he introduced me to the most, but also said, like, if you have some time, could use your help. And that team was probably the most ambitious and capable group of humans that I'd seen assembled to date and they were gathered around the goal of embedding human rights to digital infrastructure. And I still considered this to be a very worthwhile and noble goal. And protocol labs, the entity has changed a lot since I joined back in 2017.
But that's still a general goal. And my role there slowly morphed from funding science and managing and supporting scientists to, and that was like at an individual level, to systemically, how do you try to increase investment and support for public goods generally at more of a systemical? And as you go from like supporting a researcher, that starts with like a lot of, I need to understand this field and these topics and these problems. And as you go more systemic, it ends up crossing through like economics and social choice theory and political philosophy.
And you end in this place that is very much regularly facing political philosophical questions and governance questions. And the last thing I did at protocol labs was build a venture studio that included some multiple companies and nonprofits that were building tools to help people effectively invest in public goods. And one of those, it was really like, how do people coordinate at scale to achieve outcomes that people collectively want? One of the projects was a project to have a human coordinate at scale with generally intelligent AI systems.
And I became convinced that this was a compelling project, like difficult problem that more people should be working on. And so when we spun out all the companies from the venture studio, I decided that I'd been talking with the lead researcher on that project, David Dalrypl who's now at ARIA in the UK. And he and I have been talking about having a nonprofit that exists to complement that effort. And it took a while before I raised my hand and said, we haven't found anyone.
I think this is important enough that I do it, but I have no background in AI. I have no background in formal verification, which is a key component to this plan. And if I start this, it would be like an interim thing that I would do based on primarily writing checks back by your reputation, David, you said, but he was okay with that. So I when started this nonprofit, that was a little over a year ago.
And in the meantime, we've been popularizing and finding kind of short term use cases for what I'm calling specification based AI, where instead of saying we're going to control this AI system by taking our values and putting our preferences and our beliefs and our utility functions and encoding them into this AI system, we should instead have a system where you don't need to trust that the black box is going to make the decisions that you want. You'll have some scaffolding around it. You can have very high confidence in that scaffolding and you'll have high confidence in it because you've stated objective requirements of what comes out of that system. And everything that comes out comes out with a proof that is compliant.
Sorry, very long winded. No, not at all. Some of these go up to half an hour. So I guess what it's interesting because it sounds to me like some of the people I talk to can really trace the work that they're doing now back to some sort of incitatory event very early in life.
And for you and I, it sounds more like a random walk, but I'm curious, like, do you? It's like a three year correlation, like, I think. Yeah. Like, if you told me, oh, yeah, you're going from running a medicines team to a formal verification, like, I, I spec generation nonprofit, that would have sounded insane.
But if you said, like, oh, yeah, Evan from like 2020, like, I can definitely recall, like, generally when I became convinced that technology does not solve societal problems, socio-technical solutions solve societal problems. And what I mean by that is specifically is that for any society problem that technology pairing, you could imagine finding a different society that has the same problem for us that technology would not actually solve the problem. I was going to say, you know, maybe the, in a poetic sense, maybe even the more fundamental thing is this notion of optics of like being able to see, being able to understand how you see and what you see is something that run through this. Another thing, linear algebra runs through this and the idea that there are these incredibly high dimensional spaces and you care actually only about a small narrow sunset, but infinitesimally narrow sunset of these high dimensional spaces.
And the question becomes, how do you identify, localize and optimize over these things? I mean, there's, there's also like, to the extent that everything is information, you can take information theoretic approach and say, like, oh, yeah, information ties all of these things together. It's quantum information. I don't know how to liberally if you apply to a computer perspective, but it's not, like, it's like, it's, it's like, it's like, yeah, like, it's cool, it's like, yeah, it's, it's just a different, a different way to take information out of this thing.
So there's a type of, it's like, give it further. that they were making in this is that the media environment that we are living in right now is one in which we recognize ourselves not merely as nodes in a network but as networks as adaptive systems and that the media that we engage with is also sensing us in some way and that like you're talking about linear algebra and storytelling. It's like actually I think there's an equivalence there. You know like Terrence McKenna used to say that as storytellers we know that there is no real beginning or end except that we choose them as matters of convenience and you know I think a lot about the various low-dimensional encodings the representations that we use to form the basis of identity that we use to navigate the world that we use to coordinate behavior.
I mean money is one of these things right I really appreciate it in in your pitch deck for Atlas Computing. You start with the second page on the problem is that as soon as we're actually replacing employees with AI agents suddenly we have more actors than we have supervisors right and so like you automatically get into these this sort of exponentiation of the difficulty with any kind of social regulation at scale. It's not a social regulation like it's insane to me that this is not something that more people are like I feel like the maybe you've all heard already we'll write about this in a couple of months and or the published that he's been writing about because I feel like he and I are crossing the medics paths in an unsettling way frequently but I feel like when this is something that like all of the major media like and like becomes common dinner table conversation and like NPR stories about it that will be the right amount of attention that I think that the review problem deserves but AI systems right now are only creating more review burden and you can either take that burden upon yourself or you can pass it along to the next point in the system but there are particular bottlenecks and choke points of human review these are like the court system these are like management boards financial reviewers crypto do compliance checks like journals and proposals like school applications all of these things are already pretty burdened and most of them had at least some gatekeeping based on like knowing the language like it was pretty easy for a judge to throw out a frivolous lawsuit because the wording was wrong if I tried to write it to sue random person but now I can use language law and you can't scale the review without trusting the reviewer and without making that review system accountable we have no tools to make a system accountable for failures in their review and so I don't know how this doesn't get very well I only have one idea of how this doesn't get very ugly and it's a lot of level it's not a solvable solution yeah so that's I want to try and tie this in the point that you make about the fact that review doesn't scale is you point to a specific case or it doesn't scale with with AI right right well that I want to generalize this I'm just like you say the resulting arms race from using profit as proxy for good destabilize the system this automatically links into the you know cosmuchelizi and Henry Farrell's comments that language models are a familiar kind of creature that you know the creature is easy to amplify or as well that it's it's oh and also like the reductive high modernist that it is the way of shooting down sampling choosing the variables that then become the coordinating factor in human behavior in the same way that the limited liability corporation or the democratic polity has been these things already that they're saying that basically the AI in whatever sense is an institution and if we can understand it as an institution and understand an institution as a kind of organism from an information theoretic generalization then understanding what it is that they're doing and the nature of these sort of problems with oversight you know it's like to me I could be wrong here it sounds like what you're saying is that we have a difference in the degree but not the kind of the problem that we are already facing in two related areas one of which is regulatory oversight of technological innovation generally and then the other is parenting right like that's a big theme on this show is like if you have one kid and you've got a whole village watching the kid then it's kind of easy to steer that behavior if your kids are proliferating super exponentially then you know you end up in like a similar kind of trap and if those kids are operating on an incentive landscape where the productivity or impact is the governing factor in their decisions because there is an opportunity to bring the lead bullets to bring this more sort of nuanced perspective in like do you feel like I'm getting this in a way that yeah yeah I wanted to pull in the quantity as a quality on its own yeah quote I do agree that it does feel like it's a difference in quantity where like I get less difference between the one kid and two kids versus zero to one kid but also like I would assume one to 500 kids would be the biggest and like the labor force had you think about like if something became very popular or very useful as a career it would usually take a couple of years for people to start being able to but for like the labor force actually modify itself at scale to be able to accommodate that and we're probably seeing that shrink down meaningfully to being the rate of technological development in some instances so I guess just to call the Sean the ambition here when you say everyone wants human level AI agents no one can define good behavior for humans good AI is harder subjective of moving target I'm hoping that you and I can whether or not you intend to do this in the prospectus of this organization I think that the points that you're making about how to address these issues actually yield insight into a much broader category of problem which is the socioeconomic encouragement of human flourishing and responsible institutions that like I think that actually the approach that you're suggesting for human reviewable AI starts to give some inspiration into how we can treat these other things because you're actually addressing in some respect what will probably be the much harder problem and then the question of how we apply this to old school LLCs or intentional human communities or so on becomes a much more addressable yeah I think there's like probably worth saying like what I'm proposing is that humans should have squishy subjective human laws for humans and we should have some like stricter more objective laws for AI systems you can rationalize this very clearly by the notion that like part of what it means to be human is that you deserve a presumption of innocence if you do something wrong then it should be shown that you like you definitely did that thing and maybe even also had some intent that was malicious in order to be punished for that we do not have that same level of scrutiny or belief around like a medication or like that this will not explode or that like this house will not collapse those things have to be presumed unsafe until proven safe and it's very unclear where one should draw the line for things like software I think there are types of software like those that run airplanes where they're very much presumed unsafe and not allowed on planes until they are proven safe and what I'm concretely proposing is that we really shift that line back to like these are types of critical industries these are types of systems these are use cases where like if you're putting it in front of a child then there are certain things that we should be able to say these hold and in the same way that like the bridge doesn't have to hold exactly the weight of the car it should have useful warning labels and you should have ways to understand as in the case of a bridge as a driver whether or not your car is too tall or too heavy to go over that bridge you should have ways of intuiting understanding what constraints are being put on AI systems that we're putting into different parts of the world whether that's AI systems building software AI systems designing drugs designing hardware creating entertainment content for children like these are things where there's a challenging example of Supreme Court case I know it when I see it with regards to pornography there was no way to encapsulate what that means it's a lot easier to identify and specify what for instance nudity is and say like maybe we just want to draw like a harder line that's more conservative but we can have that be like a baseline understanding and have these things be much more legible when you have users in different use cases let's actually give people because you know we're already up here let's give people I really appreciate your concrete example on organizing a competition to create tools to predict bioactivity and toxicity so like the biochemistry spec language specifically because I mean obviously yeah again you're talking about a very different problem if you're talking about the use of like complexity economics simulations but like eventually you can imagine something like this being used to simulate and model interpersonal interactions and then start to suggest if not enforce reasonable constraints on sort of like workplace interactions it's kind of where I was getting with this and I'd love for you to just like lay out how you actually imagine this in practice in this case I think we know what like there's enough information in the youtube watch X then walked Y graph to be able to identify what radicalization is and like radicalization looks like engagement so that's one problem but you could imagine a system that says okay I personally like Evan wants to make sure that I do not end up going down rabbit holes but youtube has a model of me that's like oh if Evan watches this video then he's more likely to watch this next video and like I could probably just pick a bunch of videos and say like do not give me videos that are likely to meaningfully push me more towards those videos like there's nothing that's going to get me interested in someone hard pitching flatters and like any video that moves me even one step on a chain of 10 toward that direction I'm probably not going to want to watch that might not be clear from my previous engagement videos but like this is what I think of it like a concrete example um a specification and like model I should briefly note that the spec base AI stuff that I'm talking about is part of a category that includes what people are calling guaranteed safe AI where your system has four components there's an anti-system that proposes like here's the thing that you should do given this input there's a model of the world that lets you evaluate what is likely to happen if the AI system does in fact if you do act based on the AI suggestion there's a specification language and a tool for saying in the models possible states of the world these are things that are bad that we don't want to enter and there's a last component that is if you take the rest of these this will generate a computational proof you could think of this as like a check sum or a zero-knowledge proofs or like a formal verification like existence proofs that given an input you will not have the model achieve a state that you specified as bad and so that's guaranteed safe AI I would say that the kinds of things we're doing very aligned with that but example I gave with the video recommendation also like very clearly fits into this because the model is if Evan has watched x will he watch y the specification is the thing that doesn't exist where like I should it would be great to have a tool where it says like I can go in hundreds of videos and say never get any closer to these things or maybe these are just like natural language labels that are on some videos that I don't need to like I don't actually like to watch the videos that I want to personally avoid but you could imagine the kinds of things that would be and then some certificate that like here your recommendations are now compliant would be like a great thing to have this is kind of fanciful because I don't control you too but what I can do is I can propose or something like small molecule pharmaceutical intervention type things you could have an AI system that recommends here's a molecule that might reduce headaches you'd have a system that simulates molecular dynamics and says here's how it interacts with everything else that we know of in the human body at a molecular level and I think that there needs to be a way of expressing boundaries on this is binding too tightly this is what I mean by binding too tightly to this particular receptor this is what I mean by reacting too much with this metabolite and then you need some system that would generate proofs over this but the idea would be that in the same way that we have laws that say do not drive in these particular ways you could have sets of numbers that say do not generate or propose any molecules that bind to these receptors any more strongly than these other molecules we know of don't bind to hemoglobin more than carbon dioxide that's an important thing there is a very meaningful way these preferences and these applications are cultural in personal intersubjective like you can't define toxic in a meaningful way unless you pick some like very narrowly snow thing like oh did the kill 50 percent of the mice at that dosage and that's what we typically use but it would be really nice to have something that was much more granular for people who are trying to treat terminal diseases and you don't really care if there's like a really bad side effect necessarily and so the hope is that you could have a tool that would highly constrain but all be human readable and understandable and contribute to our understanding of if an AI system proposes a molecule to do a thing this is why we think it does the thing this is why we can justify that we think it won't kill you and reduce in theory the number of clinical trial like initial review burden type things all starts things like that you know one of the places that this gets really interesting for me because you also talk about you know updating software so you want to make sure that updates don't break the code especially if it's existentially important code or like one of your foresight talks you know you're talking about aircraft autopilot stuff it's like okay so when you're talking about chemical interaction specs you know my mind automatically goes to in a similar way we don't have one single useful definition of species but we can more yield with a super interesting scientific discipline but if we get really granular and we are as careful as you're suggesting to be about specifications then this stuff doesn't just apply to drug discovery it also could theoretically be used in cases like modeling ecological interventions for the cultivation of biodiversity or I think about you know like to the degree that we can come together and there's a sort of endless regress here when you're talking about your earlier work in coordination to the degree that we can come together on a carefully specified set of social values and we can use a system like this at one layer to decide what we are and are not willing to accept as externalities in some respect then the approach in the chemical interaction system locks in like I'm glad you brought up to this one because then it actually becomes part of the specifications for thinking about the way that different organismal species interact in an ecosystem or it becomes the way that you think about the different waste products produced by different industries in economic relationships yeah and going back to stuff you were saying even earlier there is an extension to which all of these abstractions and ontologies are lossy we have a collaborator who made the claim that technology is reified abstractions and abstractions do violence to values and I thought did it together those that's sort of like all profit is theft right it's like well there's a conservation of energy and matter so like yeah at some point you're leaving something off the books but you have to you have to accept lossy compression I think yeah but I think that there does become a prescriptive improvement that you can make from this claim which is we should be able to be fluid in our abstractions yeah and if you're aware of your perspective and that other perspectives differ that empowers you to try to take on other perspectives change understand them etc and I think that that very much holds true for these specifications where making them legible makes them more legitimate because of the public nature of them the ability to set them like I'm very uncomfortable with the notion of like AI alignment being a long-term solution to problems partly because I think that as usually framed it is unclear if that's going to be done in time for it to matter but like unclear if any alternative will be done too so I wouldn't work less on it but also I'm not entirely convinced that if technical alignment is solved like if I gave you a box that you could put a person in and a language model like a general intelligent AI system into and like after 40 hours of work the person comes out and the AI system is now aligned with them and like the person would agree like yes that's how I would have wanted it done unclear what you do with that do you build like billions of copies of this and give it to everyone do you align against one person do you align against like US Congress do you align to say well then and then just say you're done I don't know if there's an answer to this that I feel very uncomfortable with and part of the reason is that what it means to be aligned is that you have put your values into a black box then you don't have to open because it's aligned and I think that having transparency about values preferences specifications and having a clean interface to existing social institutions to be able to modify those whether it's a government or corporate policy or like even some stupid and user licensing agreement like as long as it's in there it's better than having it just be like did you align it to a person so this is actually where I feel like for better words I am very deeply aligned with the ethos of your project right like I've been in the box for 40 hours with this which is you know Hakim Bay in temporary autonomous zone makes a very similar point but one generation's utopia is the next generation's evil empire right that this is a problem in standardized testing anywhere you look in in adaptive systems if you're not taking into consideration the feedback between that encoding of regularity and the metrics that you're trying to track and then the behavior of entities in that system then you end up with some sort of weird destructive distortion and so again to bring K.Allana McDowell back into this we had a really interesting conversation at one point about how this relates to learning organizations and learning societies I'm curious your thoughts on the more that human consortia are by default chimeric or centauric like the more an LLC or a 501c3 or a democratic republic or whatever involves this coordination between humans and AI the more I feel like guarantee safe AI as an output of this constant test and adjust investigate oversight kind of process the more it applies to pretty much every kind of human activity and one of the areas that I think is I'd like to hear you speak to even though it's maybe a little outside of the stated objective of this project is how this kind of approach seems to get us closer to the problem of the gap that occurs between the stated ethics of an organization and the people that work for it and the actual behavior of that organization you know I feel like so many of the conversations I've been having with people for this project are touching on in some way such and such a company that I work for used to do a really good job at upholding some sort of ethical principle and now it doesn't or it's lost the plot or that it has become you know cancerous in some way like it exists in order to exist rather than in order to serve its charter this actually goes back to the my immediate response to like each generation's revolution is the next generation's evil empire to someone asked me why is it that great research organizations seem to always go stale like is there something about them like is there something about the dynamics and like my initial response was like maybe what it means to be a great researcher is to be transient like why do we assume that I think that there's like a default Darwinistic belief that because it lasted it was good and the answer is like it was good at surviving I think that maybe my cleanest answer is create research or just go stale or like build an ages end because the things that optimize for long-term longevity are not often the things that you want out of that system and I think that this like sort of ties into some of the governance stuff that I'd previously done but I became strongly convinced that every governance system is some lossy compression of aggregated preferences and if you're going to say and people have preferences over these options and you want to say the group has preferences over these options you've just done a compression by a factor of m and there's no way around that compression that said there are many ways you can do it you could say the majority want like you could say we randomly selected one person or the people elected a representative or we did quadratic voting like a bunch of different things all lose different information but I do think that there's like interesting stuff about how like humanity has governance structures as metal organisms and you do have human consortia that take out a life of their own or like companies have life cycles they're born they eat they die they consume they have desires and those desires aren't necessarily the desires of any individual human there's the collaborator I love talking to who has also asked occasionally like if I say if I'm speaking and it unclear it's my personal opinion or like the organization will ask are you speaking as an individual separate from the organization as the leader of the organization or as a member of the organization or and I think I would add it as like proxy for the combined desires or like telos of the organization and these are all separate things they're like meaningful but they're often very illegible and encoded in human preferences and relationships that people have to each other like somewhere I guess there's a great post about humanity finding equilibrium other weird that claims that there's no system of morals that would claim that Las Vegas is a moral good or that there was a moral imperative to build Las Vegas given what it is what it represents where it is the like economic impact is better up and yet it exists so the question was the question could be where did the desire to build it exist and I think it has to exist both in individuals and between individuals and arguably maybe also in the systems that those humans used to coordinate so actually you know we have an opportunity to loop back around to the beginning of the conversation and that you know one of the themes that I pursue on the show a lot is this question of like well why are we so concerned about opacity in decision-making algorithms given the fact that we the human nervous system and human society are already predominantly opaque in this kind of way that it's like I'm really fond of the argument that most of us are running on autopilot most of the time or that most of us are ideologically possessed and unaware of the fact that our personal values are emergent from or reinforced by structural concerns in this way and so again like the question of a matter of degree a matter of kind to plug this into the scientific work it's like we haven't lived in a scientific regime in which one person can hold state-of-the-art knowledge about every domain of human research for centuries right like we haven't even in the course of Alexander von Humboldt's life he was already starting to outsource his research to a whole like nimbus of disciplinary experts and I'm curious like where you find this like faulty or inadequate. See like in some respects it would be useful for human civilization to accept that to actually embody the values of the rational enlightenment to actually ensuring the reason that we have to start from a position of accepting that most of what we know in this enormous abstracted modern world we are taking on faith somewhere else locally somebody might be able to show their work and demonstrate the chain of reasoning that resulted in a particular data point yeah right yeah so if you ask me like why am I concerned of opacity because transparency is what you need it when you have lost trust if you have trust you don't need transparency it's fine and like there's trust in the institutions there's trust in the people there's trust in the technologies any of those would probably be sufficient for me not to be worried about opacity but I would claim that I am still on some level as scientists I don't do a whole lot of research but I don't think that we should accept that what we know we're taking on faith and end there we should accept that what we know we're taking on faith and that part of that faith is that we or someone else if they wanted to could justify anything that is part of this of knowledge at least at some point by no longer need to work on the specification based AI I'm very interested in starting a youtube channel where I show like modern physics demonstrations you can do for about ten dollars I think they're much really cool things that you can do yourself and there is a point where no a person cannot do that cheaply and it takes like an entire particle accelerator to verify this but for a lot of the things that are really meaningful connection between like does part of your like actually depend on this knowledge and can you validate that knowledge in some way at least for like experimental physics chemistry I think like that to the extent that if it matters then there is sensitivities that you in leverage and observe and so I think that there's like a trust but verify element of reason in terms of knowing what we know I'm glad you brought it here because trust comes up on the show all the time unsurprising yeah it's like the you've heard yeah ask the question like how would you quantify that I imagine you get them like really interesting answers given the people you talk to you with like everyone's a wild people will say like trust comes up a lot and it's like how would you quantify trust you have intuition on that this comes from like I would claim that the more objective your criteria are the more dammable they are but you know if it's subjective you should still quantify it whether it's short to be pretty for a conference was quantifying subjective or something which seemed really dumb at the time because you know with 2021 to the issue of quantifying the subjective you know I come from Francisco Verrell of on phenomenology this book I read in grad school where they basically said this is like kind of the framework that I've used for this is like there is no such thing as purely subjective or objective but all knowledge starts from an independent observation some kind of experience and then basically there is no such thing as a pure hallucination because you may or may not be able to verify your experience with others but that doesn't give you any kind of absolute certainty as to whether or not that experience is verifiable because you never actually query it against all possible other perspectives what you can do is if you look at the history of knowledge production over the history of humankind you can say we go from like my truth to our truth to what we've been calling the truth but in the sort of post-modern world the truth is also revealed to be bound situational you know it's contextually specified in some way and so this is where I you want to talk about like plurality institute and at some point you know getting I guess there's like a fun but two by two of like my truth versus our truth exists doesn't exist I would say we went kind of from like our truth exists to my truth exists to like my truth doesn't exist and are we like I guess a metamodern general trend is towards the existence and validity of multiple traits I mean this is why I like who's all hovers work non transferable tokens and the ability to decouple verification and credentials from the formal institutional production of proxy credentials right because at that point you can say in a system like that where you can bridge inter-subjective agreement gradually up into something that becomes asymptotically more objective and it's not restricted to the way that this is handled within a particular discipline right so like where I think about like quantifying trust and so on it's that now we can say things like in my graduate advisor Sean S.
Burener-Hargans looked at over 200 different disciplines and their perspectives on ecology and looked at the ways that all of these different disciplines with their different kinds of validity claims and the ways that they uniquely establish expertise and verify things like the different ontologies of data represented by for instance behaviorist science versus you know hermeneutics versus deep ecology this kind of complex systems science that all of these they may not be fungible in terms of the ways that they actually perform measurements or the ways that they actually establish trust within disciplines but you get something by giving up the idea that there is like an ultimate ground to knowledge that you can rest in perfectly forever but then accepting that you can continue to run the inquiry and substantiate perspectives by articulating them with ever more foreign perspectives I think that this ties into everything that you're talking about in as much as that it's coming from the same position of you never earn complacency you never earn like this is true but you yeah it's only like no spec should be presumed sufficient to be timeless right right or like your system ties into the thinking that i've been doing with people on like the reproducibility problem is that people are not being transparent about some aspect of the methodology or of the epistemology of the researchers they're not being transparent about some kind of shared bias that they all possess I told you my proposal for how you should be able to predict your reproducibility as part of the peer review process I want to find a journal which try this because I think it'd be ridiculous but also potentially really interesting when you submit a paper editor should say like okay here's two questions I have about the methods that I think would be meaningful like with the answer might differ across people and you might end up with a meaningfully different result based on the answer all of the peer reviewers do that too and then they all have to answer each other's questions in terms of if they're the author what they actually did or if they're a reviewer what they think the author did and I would say this won't give you an upper bound on the likelihood that it won't replicate but I think it gives you a lower bound if like three different peer reviewers all agree this is a question that is impactful on the methodology and all three of them give different answers for what they think the author did because if they that's what they think the author did that's how they would go and replicate it and so what you're really probing is the tacit and latent knowledge that sits underneath the paper and I think it's like you could probably just say like oh yeah like on the five questions answered there was agreement on all of them this paper gets like five reproducibility check marks maybe like maybe they go down to like zero it doesn't mean that it's not going to replicate but it means that you should probably have a lot more information in the methods about what it was they actually did if you want someone else to be able to show the same result yeah so this actually ties back to in one of your presentations you quoted Alfred North Whitehead I love this quote civilization advances by extending the number of important operations which we can perform without thinking about them it's like there's it's funny because like I feel like in everything that you're saying scalability jeopardizes trust it depends on abstraction it requires us to double down on clarification about the ways that we formalize things the decisions that we make about specification what we gain from this is the ability to automate huge swaths of our activity but maybe the paradox that I'm holding here and I think would be an interesting place to end this conversation is sounds like the pressure or the burden on us to think more carefully about all of this is are we in some sense just redistributing degrees to which we have to think about things like are we actually gaining any kind of convenience from having to pay so much more care to the ways that we raise AI to talk about the way that Dikai talks about it I guess like the most concrete example I can give is back when I used to work at SFI I really wanted an intern and my manager Jenna Marshall made she rest in peace so she always said to me it's not actually less work it's not actually less work to have an intern because you have to like teach them to do stuff that you're used to doing without thinking about it so that's exactly what you're talking about like making the tacit knowledge you never have an intern it's like yeah so now we have immensely cheap super proliferate interns but if they're that easier so I actually I think that the human lifetime feels like the most finite resource to the extent that you get one hour per hour that's it for the rest of your life and I think this ties very closely into my answer to the question I occasionally get you probably occasionally get to if I have child nibling neighbor what have you who is like picking a major or trying to decide what to study in school or some way what job should they have in the future it looks so very different and my current take is that in the course of humanity you can look at how much of time was spent deciding what to do and how much time was spent doing it and I think that it would be a good future in which humans humans still have a lot of input on what gets done and that like if done could be what the AI systems are doing it could be how much potential we pay to what the AI systems are doing or like on that question of how do we balance our time if we end up in the day what we say is coming doesn't future where like you get it's all just on autopilot that's fine but I think that automation has increasingly taken over the doing it side of things to increasing degrees of abstract and so at some point that goes from like get the water from here to there to water the crops to like get water from out of the ground to water the crops to like oh actually just go harvest all of these things for me to at some point you might just be like make sure everyone has food and like logistics gets solved too and like it would be great to have that kind of automation and not have to think about that and be able to spend more time about what it means for people that have food like what we mean when we say that because we don't just mean like you have food being like good food that's affordable and supportive of the habits that you want to have and empowering kinds of activities you want to pursue lots of things like this there are a lot of places where humans make statements about what should be done and I you can imagine that the fraction of humans jobs like that increasing where like even within the companies there's a lot of like just do the thing but a lot of the most valued people and a lot of the most valuable contributions come from people who are saying like oh you said do this thing but we should be doing this and that should coming from a preference probably from some shared value metric for the org that is partly embodied in that individual and we'll get you to orgs that are doing different things and like even if it's like people talk about having AI scientists who decide what science gets done there's some level on which a human is probably going to decide like nope we care more about like longevity and neuroscience than we do about making this moss live longer like what creates aging and this moss versus in humans or like maybe it's materials versus aerospace versus pharmacology like these decisions have to be made I think that in most instances while control sits in human institutions humans when making those decisions and I hope that the way that they express those decisions is through specifications so that you have automated validation that the actions are meeting us next. It sounds to me like what you're saying is basically we can quantify an improvement over the history of the last 20 years by saying that at least in certain areas more people are concerned with what is ultimately the original scholastic or leisurely pursuit of the life of the mind which is to ask these questions about human flourishing we get to spread more evenly over the surface area of what might be worthwhile what questions might be worth asking is that ring true. I would say the good future I imagine is one in which everyone has more capacity to engage with other questions. Awesome.
Dude it's a pleasure always. Yeah. Good chat. Thanks and good luck.
Thank you very much. Thanks for having me. Thanks again for listening. If you liked this episode do the like subscribe comment dance and join us in the wisdom and technology discord server linked in the show notes.
Humans on the Loop is made possible thanks to the support of listeners like you. If you care about the promotion of wisdom in our age of magical technologies become a member at humansontheloop.com. Make tax deductible donations at every.org. humansontheloop or email me to work together at humansontheloop at proton.me.
In the next episode we talk with techno shaman George core founder and director of research at future how about collaborative hybrid intelligence for human and societal flourishing until then take care and remember attention is our greatest natural resource but also maybe resource is the wrong word.