>

Base Layer EP 06: Ari Shamash on Your Agent Shouldn't Have Your Password

[Listen now]
EP 06

Transcript

Your Agent Shouldn't Have Your Password

Ari Shamash · AuthZed

52:51

This transcript was generated automatically and lightly corrected. It may contain errors, so check the episode audio before quoting it.

00:00Cold open: there was no magic wand

Ari Shamash00:00

It's not like we had a magic wand or some divine intuition or something like that that told us this is the right thing to do. It's like standard engineering applies. We're tackled with a large problem, no idea how to go about solving it. We run lots of different trials and experiments. Yeah. The vast majority of them fail.

Jake Moshenko00:16

Welcome everybody to the Base Layer Podcast. I'm Jake, one of the co-founders and the CEO at AuthZed. All right. Welcome, everybody, to the first and maybe only fully in-person version of the Base Layer podcast. I'm your host, Jake Moshenko, co-founder and CEO of AuthZed, where we do authorization infrastructure. And today I'm joined by Ari Shamash. Ari is a longtime technology industry veteran. He got his career started way back, but some of the highlights of what he's done is he was a director of engineering over at Sun, where he was one of the first people to make compute available via the internet, which is what we all call cloud nowadays. Then he spent 16 years, was it 16 years? Yeah, 16 years at Google. Some of the highlights of his time at Google, and part of the reason we're talking today, is that Ari was the manager of the Zanzibar team at Google. Just a reminder, Zanzibar is the system that Google uses to do authorization for all of the internal services and customer-facing services, like Google Drive, Google Cloud, YouTube, et cetera. Ari is the reason that the world, or one of the reasons that the world even knows about Zanzibar. So Ari was in charge of the team when the paper was created and published. So that's the reason we're, as a community, even able to talk about this thing. And after his tenure on the Zanzibar team, Ari went on to do AI safety work at Google. Is that right?

01:14Managing the Zanzibar team at Google

Ari Shamash01:53

Yes. AI governance, AI safety.

Jake Moshenko01:55

AI governance, AI safety. I can't think of a better person to talk to about authorization in the age of AI, AI agents, how all these things come together, and how they're going to evolve and how they're going to work in the future. So it is with my great, great pleasure that I welcome Ari to the podcast.

Ari Shamash02:16

Thank you for having me. I'm excited to be here.

Jake Moshenko02:18

Yeah. Is there anything about your background that I missed or that you want to elucidate on?

Ari Shamash02:24

It hasn't been that long. You said I've been in the industry for a long, long time. I don't think it's been that long, though. OK.

Jake Moshenko02:30

So not that long, but a little while. Long enough to be at Google for 16 years. All right. Why don't we start with a quick warm-up question? You've been involved in AI, AI safety, those kinds of things. What is the coolest thing that you've done with AI to date?

Ari Shamash02:45

Yeah, that's actually a great question. So I heard an analogy a while ago that's very applicable to my answer. And the analogy was that if you imagine all the data that's available to us today, given that we're in Manhattan, I'll use a Manhattan analogy. It's like having a bunch of different skyscrapers all over the place. But the skyscrapers are not connected. Like to go from one data silo, you have to go from the 95th floor of one all the way down, walk across the street, and then get into another skyscrapers. So if you have two pieces of data, one on the 95th floor and this one, and one on the 58th floor of that one, connecting those two pieces of data is incredibly hard, right? Yes, some buildings have bridges, but those bridges were built at human speed, right? So it might take years and years and years to connect these data silos. Whereas all of a sudden, we have this AI stuff that's able to consume all the data across all of this and build virtual bridges amongst all these different data skyscrapers. skyscrapers to make it look like one piece of data. My father passed away recently. He was 98 years old. He lived a beautiful life. Several years before him, my grandfather passed away. He was also in his late 90s. They've lived long, long lives across many different countries. And as kids, we didn't really understand or know the full scope of their lives. And we had a hard time finding out information about them because all the information was sitting in thousands of different places. Some of it was handwritten on paper. Some of it was written in non-English languages. The hardest thing that we had was non-English languages handwritten on paper. Sitting in a vault somewhere, some other country. So all of a sudden, all this information became online. People started scanning in all the paper because the paper was going away. And we were able to cross correlate all these different private silos of information and start asking interesting questions. And we learned so many details about my father and grandfather's lives that we just didn't know. For example, one of them is that my grandfather was actually a policeman at some point in his life. The only reason we knew that is because we found his police ID card that was written in a foreign language that we would have never found if it wasn't for ability to cross corally um lots of different information new image recognition across you know miscellaneous uh photos or whatever um so i think that's the kind of coolest thing is that you can take information from all over the place and do instant analysis on it that in the past would have taken years and

03:44Reconstructing his father's and grandfather's lives

Jake Moshenko05:24

Years and years if if ever so did you have like an ai agent do that is that you put all the data in some place and said draw some conclusions or walk us through how that actually happened yeah so

Ari Shamash05:34

So a lot of the data owners have actually started putting up systems in front of it that use their own AI stuff in order to correlate all of these things. We did build agents to go scrape all of this information and kind of answer questions on our behalf. So we were able to direct questions to one place, and the agents went out and collected all the information and gave us information on our behalf.

Jake Moshenko06:01

That is super cool. I've done quite a few of these at this point. Every other answer that I've gotten has been, I built something cool for myself, some piece of software with AI. And I think yours is the first one where we've gotten like a data tapestry answer, where it's like, let's pull in all this data and let's use AI to correlate and build a complete picture. I think that's super neat.

Ari Shamash06:23

Yeah, like automating my life is something that I've been doing for a long, long time, right? I guess it's a lot easier to automate these days because I can build things that auto-learn APIs or whatever and auto-integrate them. But that's not interesting to me. Being able to analyze all the different data that's out there and cross-correlate data in different formats, in different languages, in different whatevers, and be able to synthesize it far faster than I'm able to. That, to me, is super interesting. Cool.

Jake Moshenko06:52

Well, I'm super interested to hear about your time on the Zanzibar team at Google and then maybe follow up a little bit about the AI safety stuff that you worked on. I understand we'll keep this within the limits of what you're allowed to talk about. We're both ex-Googlers, so both bound by our continued NDAs and whatnot. So if you ever need to, just say like, hey, I can't answer that. But as much as you can share would be greatly

Ari Shamash07:19

Appreciated. Yeah. Thank you for giving me the opportunity to talk about this kind of stuff. So I joined Zanzibar pretty early on in its evolution. It was already up and running as a system within Google when I joined the team. The adoption was pretty low, I would say, in that not everybody within Google was adopting it. And the query rates that were coming into Zanzibar were also on the lower side. Over my tenure of that particular team, Zanzibar adoption grew significantly, Both in terms of number of customers, certainly the amount of data that was inside of Zanzibar and also the number of queries that was coming into Zanzibar. We also had additional friends of Zanzibar, I'll call them, other systems that grew significantly over time. So, for the example, the ability to do authorized search on behalf of other systems at Google. Those systems started coming on board and started using Zanzibar, which drove some of the Zanzibar traffic, but also drove some of the systemic complexity around it that had to be dealt with in order to continue providing users both real-time responses that were low in latency that were true relative to the data. It would be horrible if, for example, you went to Google Drive and you did a search and all of a sudden, either you got too many documents that you clicked on and then you don't have access. One would argue that's a data breach because you're able to see the title of a document that you don't have access to.

08:49A stale ACL in Drive search is a data breach

Jake Moshenko08:54

Yeah, acquisition plans or super secret war plans, right? Stuff like that. People shouldn't be able to see things.

Ari Shamash09:01

Yeah. On the other hand, if you have documents that you should be able to see and they don't come up in search results because the data is stale, that's also a failure mode. But keeping all of that information up and running across the entire corpus of Google Drive documents and the constantly changing permission models was a systemic challenge that we had to deal with.

Jake Moshenko09:24

Yeah, I noticed in the Zanzibar paper, there's actually a reference to the Zanzibar-aware search system, but just a reference and no elucidation beyond that. Do you think we'll ever get a paper about the search system?

Ari Shamash09:38

You know, I will leave it up to my colleagues at Google to continue pushing this agenda. We, the engineers, have wanted to write this paper for many, many, many years. The same challenges that we had in writing the Zanzibar paper, this is a piece of work. Writing a paper takes a long, long time. Getting the right language in there, making sure we are not giving away company secrets, but also still keeping the paper interesting to others. what is the right balance between exposing information and not exposing too much information, getting the right level of approvals within a company. That took a significant amount of time. If I remember right, it took a year or two wall clock time to do it. And during that time, we had a lead engineer that was working on this. The lead engineer could have been working on something else. So finding the right person who had experience writing these kind of papers, who was motivated to write these kind of papers while also justifying the engineering time working on a paper versus all the other priorities that we had going on. For example, building out features, building out scalability, reliability, reducing the cost of Zanzibar, so it was a cost-effective service within Google, and so on. These were all competing priorities, and we had to fight for the paper to get done. My opinion, it was one of the best things that we ever did, both for Google, the company benefited from it. We got feedback that this is yet one more seminal paper along the other papers that Google has written about file systems and databases and whatnot.

09:55Getting the Zanzibar paper published took years

Jake Moshenko11:19

Yeah, Google had a reputation for writing these papers. And the analogy that I heard was that Google's living 10 years in the future and sending us notes back from the future in the form of these papers.

Ari Shamash11:31

Oh, that's an interesting analogy, yeah.

Jake Moshenko11:33

But I think Zanzibar might be the last big paper that I've seen come out of Google. Have they done, I mean, obviously the transformer work that happened in Google is the backbone of everything that's happening with LLMs and whatnot. But in terms of infrastructure papers, has there been anything since Zanzibar that you're aware of?

Ari Shamash11:50

I'm not aware of any.

Jake Moshenko11:51

Yeah.

Ari Shamash11:51

And of course, you asking me this question makes me wonder why we were the last ones to think that this was important.

Jake Moshenko11:58

You either did so well that no one thought they could top it, or you did so poorly that Google was like, never again, or done with papers. I would hope it's not one of those two.

Ari Shamash12:08

I don't think we did so well that we set the bar incredibly high. I would hope we weren't the worst ever.

Jake Moshenko12:14

Well, certainly the paper is not the worst. Talk to me about the process of getting it approved, if you can. How did you pitch the benefits to Google? You said there would be benefits, or there are benefits.

Ari Shamash12:26

Sure.

Jake Moshenko12:26

How did you pitch that?

Ari Shamash12:27

Yeah. So within engineering, we didn't really have to justify the benefits. It was well understood that writing a technology paper or writing a paper about Zanzibar was a net positive. It was really within us as a team when we stared at ourselves and said, we have a certain amount of work to do, a pile of engineering work to do. Do we honestly believe that we can spend some percentage of our time also writing this paper and it would benefit the product, the Zanzibar, it would benefit Google and it would benefit us as individuals and as a team? I think the hardest sell was for us rather than for the rest of Google. Later on, when we had versions of the paper ready for review, we had to fight for resources within legal to review it. There was competing resources inside of that. And the folks who did paper publications and press publications, whatnot, the outbound folks also had to review it. And we had to find the right person who had the capacity to review it and actually help us push it out. Those were later challenges, but initially it was us as a team saying, do we want to invest the engineering resource, our limited engineering capacity and engineering resources towards this paper, or are we better off focusing on something else?

Jake Moshenko13:50

Right.

Ari Shamash13:51

And we decided to do it. In retrospect, I think it was one of the best decisions we ever made. Is the legal team at Google, like, when they're reading through the technical matters of the paper, are they, I mean, they must be, like, aware enough of what's going on to know what's a company's secret, what can be shared? So, legal is a really broad term at Google. There are lots of different lawyers, and my experience is a wide variety of technical skills. So, we had to find the right person to go review it who was deeply technical. Doing legal support for higher level products requires a very different legal skill set than doing legal support for infrastructure products that are not directly accessible by the outside world. So we had to find the right person who had the right combination. That's cool, though, that you were even able to find that resource. Yeah. We also got some executives within engineering who had enough credibility with a legal organization to weigh in and say, we think this is okay. And that also raised the credibility of the paper, raised the confidence that we're doing reasonable things. At the end of the day, the biggest issue that we had to decide, and I put this under the legal umbrella, though technically it's not purely a legal issue, is how much information did we want to divulge? We as engineers wanted to maximal information. The company wanted minimal information. So finding the right balance for that. And for that, we really had a handful of executives, and they all had to agree that we found the right balance.

15:05Engineers wanted maximum detail, the company wanted minimum

Jake Moshenko15:32

Yeah. It was really interesting to me reading through the paper because sometimes I was able to read between the lines, right? Like, this is what Google said or was allowed to say or had to say. And sometimes it was like, this is a decision that makes sense in the context of Google infrastructure. It might not make sense in the context of someone else's infrastructure, et cetera. So I actually went through and made a version of the paper that has my annotations on it. You can find it at zanzibar.tech, as long as we keep running that thing. But yeah, obviously, huge, huge benefit to the industry. Huge benefit, like Gott said, probably wouldn't exist if you hadn't been able to write that paper. When you were working on the paper, did you have any inclination that there would be multiple open source projects, multiple companies founded on the concepts of the paper? Or you were just like, oh, this is a neat internal service and I'd like to tell the world about it? It was more the latter than the former.

16:14Nobody expected open-source Zanzibars

Ari Shamash16:31

I was actually surprised just how many instances of Zanzibar were put together directly as a result of the paper, right? For us, it was just a system that we built inside of Google that meant the need for Google, right? We never really envisioned this particular system being incredibly useful outside of Google. It was also at an age where authorization was not yet ubiquitous as a term outside in the industry, right? Like everybody needs file systems, everybody needs databases, everybody needs compute clusters or whatever, right? So some of the other, you know, open source projects or papers that were written filled a need, right? Like explaining to people how to build Spanner, for example, or worldwide distributed databases with transactional integrity. That's a hard computer science problem that there was a need for in the world. I'm not convinced that 10, 15 years ago, people really saw authorization as a huge need inside of the industry. Everybody solved authorization in their own way. The use cases for cross-industry authorization didn't exist. So as long as it was solved in whatever silo people were living in, it was good enough for them.

Jake Moshenko17:45

I mean, speaking only for myself, in this very office, we filled multiple whiteboards trying to figure out how to do scalable, flexible authorization that could describe a service that wasn't inherently coupled to the authorization model. So it's pretty standard or pretty easy to just build something in that works for a little while. And as time goes on, you run into scale challenges with those things and you run into flexibility challenges, right? Like when we originally wrote it, it didn't support teams and now we need to support teams or teams of teams and things like that. So, you know, I've spoken about this at length, but like we had certainly recognized this during our core OS days because we were trying to provide a Kubernetes-based cloud platform where you would bring in all of the higher level services through the use of operators. And so there it was like, how do you create an equivalent to the IAM service that exists in cloud providers without knowing what the things that you're authorizing are, right? It was a really hard problem. And so when I read the paper, I went, wow, this relationship-based access control is really cool, really flexible. And the paper really went, or not the paper, but the service, and then the way the paper talks about the service really went the extra mile to figure out how to make that scale. Because the earlier iterations of relationship-based access control are all kind of like graph database-based and just try to resolve everything in a single graph walk. And so the work that you all did to make it decomposable and to break the problems down and to cache the solutions to sub-problems, that's the hard work. So the other thing that I saw when the Zanzibar paper first came out is I saw a bunch of companies say, oh yeah, where's Zanzibar 2? And when you looked at it, it was just they were relationship-based access control, but hadn't done the hard work to solve the engineering challenges. So kudos. I mean, I read the paper in this office and literally turned to Joey, my second time co-founder now, and said, we need to go do this. This is the thing that we've been looking for. So again, kudos. Thank you. That's very cool. I mean, I kind of want to

18:48ReBAC is the easy part; making it scale is the work

Ari Shamash19:55

Reiterate, there was no magic behind, like, we were just a bunch of engineers that had the same exact engineering skills as anybody else out there in the industry who were faced with a particular problem. And it's not like we had a magic wand or some divine intuition or something like that that told us this is the right thing to do. Standard engineering applies. We're tackled with a large problem, no idea how to go about solving it. We run lots of different trials and experiments. The vast majority of them fail. Every now and then we find a nugget that worked and And we're like, oh, something is working in this particular direction. Let's go explore that direction more until we figure out what the right answer is for.

Jake Moshenko20:34

Well, you probably can't see it because you're on the outside. But I think the special sauce that made it such a great solution is that you weren't able to kind of cobble something together at small scale and just keep patching it as the scale grew because you were working at Google scale right out of the gate. Yeah. And the other thing is you also had availability, you know, you had access to Google budgets, Google engineers, headcount and infrastructure board and slicer and all these things. Right. So like it is kind of a unique area, unique environment where you can tackle things in these huge, huge scale ways and come up with like kind of a solution that's not infinitely scalable necessarily, but scalable much farther than other businesses might have done in the same position. So you say we're similar to every other engineer doing the same thing, solving a problem, but there's a little bit of magic there. Yes, that's fair.

Ari Shamash21:29

We had a tremendous amount of access to Google infrastructure. For example, we had Spanner, and Spanner was already deployed. We had all these data centers, and they were deployed. And when we needed extra data center capacity, yes, we had to justify the cost of that particular capacity, But it's not like we had to go and find a plot of land in the building and bust out an excavator and bring in power. Yes.

Jake Moshenko21:54

Yeah.

Ari Shamash21:54

And we were also lucky because we had clients who demanded that the service be available at a certain reliability, at a certain latency response time and stuff like that. So we had constraints to meet, but they weren't artificial constraints. They were constraints that were driven by, you know, the YouTubes of the world demanding certain response times that we had to go meet as an engineering team. But at the end of the day, we're just humans, engineers facing the same exact or what I want to say is using the same exact engineering techniques that everybody else uses. Like, we have to functionally decompose the problem into smaller problems and then go figure out how to solve those problems, possibly by functionally decomposing them as well, right? We had limited human resources, just like we had limited data center resources. So we had prioritization discussions about what is the biggest thing. We had clients who were demanding everything because, you know, we wanted to create a culture where clients would come to us and ask us whatever was on their mind. It was up to us to say no to them rather than pre-filtering on the client side. So, yeah, we had to exercise extreme caution in what we committed to because we wanted a reputation as an engineering team that if we commit to something, we're actually going to deliver it. And that human aspect of building the service, I feel like, was as important as the engineering or the technical reliability, right? If a client came up to us and said, I need XYZ that you don't provide today, and they're betting their entire bet, and by client, I mean internal to Google clients. If they were coming to us and saying, we're going to build this brand new feature, we're expected to launch this feature in six months. we need ABC from you, otherwise we're not going to be able to launch our thing. Can you do ABC? I would rather say no and have them figure out an alternative that enabled them to go and be successful rather than committing and under-delivering and then being the scapegoat for why the rest of the company was not able to deliver. So we were very careful in terms of what we committed to, and then we made sure we actually delivered relative to that. But that's hard. And that has nothing to do with authorization, has nothing to do with Zanzibar. This is a classic engineering problem that everybody faces.

Jake Moshenko24:19

Yep. I've spoken sort of at length about why it's important for companies to have an authorization service instead of one of the other ways to do authorization, which might be like homegrown RBAC, coupled to your database, or some sort of library or policy engine. In your words, why do you think it's important for a company to have an authorization service?

Ari Shamash24:43

So two examples come to mind. I'm not going to name the specifics for non-disclosure purposes, but I can talk in generalities, right? As a service, the first example is as a service, you never know when you're going to need scale. For example, during COVID, there was a massive shift in the industry between how people did work. And we had certain services within Google who went from usage at level X before COVID to overnight requiring significantly more than X and leave it at that, right? I mean, my own kids ended up spending a lot of time on Google Classroom. I don't know if that's related, but I'm not going to mention anything. Like, all of a sudden, things go, you know, I'm assuming many of the technology providers that all of a sudden became household names during the COVID era face similar scalability issues, right? It's nice to make scalability somebody else's problem.

Jake Moshenko25:38

Yep.

Ari Shamash25:39

All right. And I'll give you an example of how scalable Zanzibar was. There was one time that, and again, I won't mention names because it's internal details, but somebody triggered a load test and they were expecting the load test to run against their staging environment, but something was misset, something, something, something, and all of a sudden they're running a load test against their production environment. And that load test ended up generating, if I remember right, something like 10 million QPS extra traffic to Zanzibar. We didn't notice. You know, yes, we looked at the graph and there was a spike up in terms of traffic, but the system scaled. Nobody got paged, the system did not go down, like nothing bad happened on the Zanzibar side. And there was tremendous surprise as a result of that, both by the customers. Like, how did you not, how were you able to handle all this extra traffic? You just did this, you said. And, you know, from our perspective, you know, we also increased our confidence in the service that we knew there was headroom. Of course, the immediate questions that came back were why you have 10 million QPS extra capacity. Maybe you should reduce your production footprint and give you some machines back to the rest of the company. You know, these are always the tradeoffs, right? So at least we knew where we stood relative to that. So I think that's one particular aspect of it, is if I'm building an application, I want to really think about what my core value is of the application and go focus on that and outsource to people who know and have proven systems. And I would put authorization in that. Like, unless you're in the authorization business, you should go to somebody who knows how to build an authorization system and use their experience.

25:44The accidental load test: 10 million extra QPS

Jake Moshenko27:29

Yeah, focus on your core competency.

Ari Shamash27:31

Focus on your core competency. The second example that became true over the last five years is as the tech industry is getting regulated, proper authorization is a critical part of proving to regulators that you are indeed meeting their various different regulations, right? And while it's possible that every single, let's assume you have an enterprise where every single application is rolling out its own authorization scheme, It's possible that you're compliant. But in order to really prove that you're compliant, you have to go pull every single one of them, collect data. Every single authorization system is going to use a different vocabulary, different logging mechanism, different whatever.

Jake Moshenko28:12

So, okay, you go and audit them.

Ari Shamash28:14

But now you have to do this on a yearly basis to produce a compliance report and that becomes a really expensive operation, right? If all of this is centralized, you build one mechanism to prove that the entire company is compliant with whatever regulation and you're done. So the value above and beyond just the scalability, reliability, whatnot, in this new world that the tech industries is entering, the regulated world.

Jake Moshenko28:41

Yeah, both great examples. Your earlier accidental scale test example reminds me of an issue, not an issue, but an early customer of ours at AuthZed. They were doing a gradual rollout to move their traffic over to AuthZed. and they were like, we're going to do 1% at a time until we get to the full rollout. And they did 1%, 2%. And I don't know if it was an accident or someone just said YOLO, but it went from 2% to 100%. Oops. But the system worked, right? And that turned them into an advocate for us because they said, oops, but hey, it worked. Nice. Actually, I think latencies went down when they did that because they started getting such a better cache hit ratio.

Ari Shamash29:23

That was one of the reasons why we survived with 10 million QPS as the cache hits went up and we were able to absorb it as a result.

Jake Moshenko29:32

Yeah, super cool. Yeah, while you were in charge of Zanzibar at Google, it went from very low adoption to very high adoption. I don't know if you can mention specific numbers, but if not, what do you think was the biggest unlock? What do you think helped teams see the value or adopt the software or what?

Ari Shamash29:52

Yep. So it's several different things. We went from lower adoption to higher adoption in several different metrics. And the three that really come to mind are discrete number of users, queries per second, and amount of data that was stored in Zanzibar. And all three of those went up. As a team, we were very opportunistic about acquiring new customers. We spent a lot of time really understanding our customers. We would meet with them regularly. We would try to understand their perspectives. We would try to understand their pain points. And pushing Zanzibar onto a team that had a functioning authorization model that was not problematic to them was not going to work because every single team at Google was resource constrained. Everybody had a bigger backlog in terms of things that they wanted to do, as opposed to the number of human beings that were available to actually do all of that work. So if I was sitting in a, I imagine this, right, if I was sitting at a customer and they had a bunch of features that they wanted to deploy in service of their customers, if their authorization was working and some random people from Zanzibar came knocking on the door and said, you should really be using Zanzibar, they would turn around and say, to what end?

Jake Moshenko31:09

And we were like, because Zanzibar is great. They're like, everybody comes and tells me their service is great.

Ari Shamash31:16

Right. So we spent a lot of time really analyzing and finding customers that were having problems and brought solutions to them on how it solved the problem. But we also communicated it in their terminology. So, for example, during COVID, when all of a sudden people needed scalability, we were able to show them that they could actually achieve the scalability that they were wanting for their application by switching over to centralized authorization. There's cost in transitioning to a system like Zanzibar. I guess that's what I'm trying to say, right? And that cost had to make sense to... There's cost and there's loss of control. And both of those things had to make sense to the teams that were adopting Zanzibar. So we were very... We did not use a corporate mandate that must do anything. I find that those are rarely effective, right?

Jake Moshenko32:10

I make reference to that quite often. I say, well, I'm not allowed to tell people they have to use it because I'm not Google. And it turns out at Google, they didn't tell people they had to use it either. So I guess I'll take that out of my talking points.

Ari Shamash32:22

You know, I honestly don't remember. There may have been a mandate. Like if we were to go to the mandates book within Google, does it say you must use Zanzibar? I honestly don't know. It was not a stick that I wanted to use. using the carrot. One of my mentors at Google said, it is our job to walk a mile to save an inch for our customers. And that's something that I really believed in as an engineering discipline. Like, it's up to us to do the heavy lifting in order to make Zanzibar super easy to use and integrate into customer environments. So they only have to walk an inch in order to use it. It doesn't matter if it's a disproportionate amount of work on our side. So, like, being strategic like that, finding the customers, working with them, identifying their pain points. And again, no magic.

32:41Walk a mile to save an inch for your customers

Jake Moshenko33:07

Yeah. I'm looking for a silver bullet, and you come back with, we just did standard product management. We went and talked to a lot of people and asked them what their problems were.

Ari Shamash33:16

Yes. I feel like this is an industry norm.

Jake Moshenko33:19

You'd be surprised how many people just build and don't bother talking to their customers or figuring out what their problems are. So, yeah. No, I mean, kudos.

Ari Shamash33:29

I mean, that was the other thing that we did, is we built with customers. So when customers are ready to go shift over, we would actually co-develop integration with them, like a very high touch model across the board until it was ready to disengage and let them fly on their own. Like it was not a bunch of, here's a recipe or here's a bunch of documentation, go RTFM and figure it out. you're wrong with that was also not this style like our largest customers we had recurring meetings with them to continue to understand how the service was working when we built features we built it in collaboration with a customer because the last thing I wanted to do is build a feature and have nobody use it right right so we built features with the idea that feels horrible yeah I mean otherwise why are you building it right yeah build this thing and nobody's using it right So I always believe in solving the first customer problem by building something, knowing that a customer is going to use it, building it collaboratively with that customer. And then that customer goes and tells the second customer, this is a really cool thing, you should use it, and then it kind of snowballs from there.

Jake Moshenko34:33

So that was our strategy.

Ari Shamash34:34

You know, at the same time, we had to scale the things that we were constantly focusing on performance improvements. We were, as one of Google's largest systems, it was also constantly facing scrutiny in terms of costs. Did we really need to be as tight on our SLOs as we were? Or could we increase the latency by a handful of milliseconds and save ourselves whatever dollars in infrastructure? All of these things were constantly tested. For example, we had, when there was a particular data center crunch, I forget exactly what year it was, but there was a significant shortage of equipment. And the question was, can Zanzibar survive at double latency, whatever the average latency was at the time? Could Google survive if Zanzibar was responding at six milliseconds instead of three milliseconds at 95th percentile, whatever the numbers were at the time? So we actually started artificially increasing the Zanzibar's response time to see what would happen. There were a bunch of people who, we as engineers, had a sense that our clients were going to get very upset at this. I think there were a bunch of people who were secretly hoping that nothing would happen. We could actually take away a bunch of Zanzibar. Yeah. My advice to the engineering team at the time is let's go get our buckets of popcorn.

35:17Could Google survive Zanzibar at double the latency?

Jake Moshenko36:00

Let's go watch.

Ari Shamash36:01

We will see how this movie ends. If it ends with our customer screaming, we will buy ourself some ammunition, if you will, relative to the capacity people. If not, if nobody screams, then, you know, as shareholders in the company, we should be giving all these resources back. It's absolutely the right thing to do. So lo and behold, as the latency started going up, we started getting paged by all of our different clients saying Zanzibar is broken, Zanzibar is down, Zanzibar is this, Zanzibar is that, Zanzibar is this other thing. So we just redirected those customers to the capacity planners and the capacity planners were able to use that as justification for giving us the resources that we needed. That said, we've always invested resources in making Zanzibar more efficient from a hardware perspective. We're always looking to reduce the cost per query for Zanzibar. That's a never-ending engineering challenge that every large system faces.

Jake Moshenko37:05

I think these kinds of systems are also like the testing that you were doing by artificially inflating latency is probably way easier than organically inflating latency by taking away resources. Because these kinds of systems tend to have non-linear responses. When you-- like if you took away half of the compute, it wouldn't double the latency. It might 10x the latency. Because things start queuing up. Things just go awry when you try to do that. So it was good or interesting that you did it artificially and got the right answer to go and keep the service strong and fast and good for your customers.

Ari Shamash37:42

And as again, we were lucky that Google had the infrastructure to do this. So we didn't have to engineer solutions or do something destructive like taking resources away in order to test out our hypotheses.

Jake Moshenko37:54

Yeah, I mean, testing the hypotheses in that way might have actually resulted in an outage.

Ari Shamash37:59

Yes, 100%, especially for non-determinists. non-deterministic is the one where I'll go with complex systems like Zanzibar where tweaking some knob doesn't necessarily have predictable...

Jake Moshenko38:14

I mean, as soon as things start queuing up, you hit a tipping point, and it's bad times after that. Great. Were all of the teams that you were bringing on board, were they all clean, relationship-based access control implementations, or did people ask you for sort of attribute style stuff?

Ari Shamash38:35

Oh, so interesting question. People always asked us to do more than what Zanzibar had, right? There are always use cases that people had that pushed the model, right? Whether today's, so there's one particular example that I'm thinking about, And I'm not going to be able to share it for non-disclosure purposes, but there was one particular use case that one particular team had that Zanzibar didn't have natively built into it, and they asked us to build it. We thought about it and said, no, it doesn't make sense to build it into Zanzibar. Or using the Zanzibar model, you know, another analogy that I really loved. When I was in high school, I was taking one of these standardized test prep classes. And our instructor had symbolism, like he would take his right hand and scratch his left ear with his right hand. And all of us were like, what does this mean? Like when you see me do this, it means you're tackling a particular problem, but you're tackling it the hard way. Usually you use your right hand to scratch your right ear and your left hand to scratch your left ear, but you would not really do this. But yet, when you're dealing with solving a particular problem, you're so much in the weeds, you might be taking the hard approach rather than the easy approach. This one particular use case, one customer wanted us to do something that was really that, that we're using Zanzibar and a corner capability of Zanzibar to solve a particular application problem that would have been much easier to solve at the application later than Zanzibar. It was easier for them to just say, we're going to use this one feature, even though it was a lot more expensive. So we gently explained to them that they're better off building it in this other way. Over the last five years, especially as a response to regulation, we started building systems within Google that added attributes not to Zanzibar, but to layers above Zanzibar and next to Zanzibar that used attributes to calculate policy decisions rather than just the mechanisms that Zanzibar has. For example, Zanzibar can't do things like if this user is in Europe, then allow authorization, otherwise not, because that's based on a characteristic that the query has rather than a characteristic that or where... It could be if this user lives in Europe. If it's static.

40:25Attributes live above and beside Zanzibar, not in it

Jake Moshenko41:10

Right, if it's static. But if it's relative to the query, we needed alternative mechanisms. Makes sense.

Ari Shamash41:20

So the regulation, the regulatory world that we're entering, I think, is going to drive a lot of these attribute-based policy engines. And Google has built up numerous of these, but that's outside the scope of Zanzibar properly.

Jake Moshenko41:33

Gotcha. Cool. I think it's pretty indisputable that we're in the age of AI or at the very least an AI bubble. What do you think, is Zanzibar more relevant or less relevant in the age of AI? And, you know, can you expound on whichever one?

Ari Shamash41:52

Yeah. So I may be, well, it's the right word. Naive is the wrong word. I'm struggling for a word. But anyway, AI is taking distributed systems and increasing it by orders of magnitude because everybody can go write their own agents. But when I scroll back 10 years, 15 years ago, when distributed systems started becoming commonplace, especially with cloud providers, whatnot, anybody can go build some code, or people who knew how to build some code could build some code and have that code act on their behalf. So it was no longer human beings logging into systems and doing things. It was now automation or code on behalf of users. And we never considered then giving human credentials to these systems because I didn't want to take my own username and password and hard code it into code, right? Like, we're always told not to hard code passwords, right? So we needed some sort of authorization mechanisms for all of these services, and we built that into the framework for distributed systems, right? So if system X is going to call system Y, system Y have to trust that system X is allowed so it can build it or whatever. I don't understand why in the AI world we think it's completely reasonable to, here's my Amazon password agent, go do whatever you want with it. I trust you to only tell me when the price changes or when you find something interesting that I should buy rather than buying it on my behalf. So, like, I wonder if we're just one disaster away from realizing that we need strong authorization for all of these AI agents. Or maybe that has already happened.

43:10So why hand an agent your Amazon password?

Jake Moshenko43:47

I mean, you can see stories all over the Internet where people give one of his personal agents access to their email and it deletes them all. Yeah. You're like, wait, it's not like that. Yeah. You know, I've heard many stories about data being deleted, you know, source code repositories being deleted. Yeah, the database and then the backups. Like, oh, I helpfully cleaned up all the backups for you, too.

Ari Shamash44:09

Yeah, I guess spring cleaning has come early. We just need to get a hold of it and get out of the way. No, I'm a big believer that the world is learning, right? We're entering a brand new phase of, and it's actually really beautiful that this is happening. Like, I feel like we're going through yet another major disruption in the industry. I've been lucky in my career that I've lived through several of them. The first one being the browser, getting you big access, and then e-commerce being another one. And then mobile, getting everybody internet access from wherever they happen to be. These were major disruptions that were pretty seismic. This is another one. Whether it's bigger or smaller, same size, like, that doesn't worry me. It's just that we're going through another one. Every single one in the past has eventually landed and realized that some of the basics of security, authentication, authorization, reliability, whatnot, eventually data correctness, all of that comes back full circle. So it's only a matter of time, I feel like, before this AI disruption realizes that we need some of the basics that we've been talking about as an industry since the beginning. We just need to figure out what that looks like for the new world. I know there are models out there, and I certainly have opinions on how to do this. But we as an industry have to realize that the need is there, and then the solutions can fit the problems. Did all of those past paradigm shifts, did they all have a wild west phase before they eventually? Yeah.

Jake Moshenko45:45

Yeah. In my opinion, yes, right?

Ari Shamash45:47

Like I remember when the browser first came out, there were huge arguments about do we need to encrypt data on the wire or not? Can we just use regular HTTP? And then there was the whole HTTPS versus SHTP debate on what is the right encryption mechanism. And eventually we settled. We landed on one, and we ran with it. I feel like the original cell phones were unreliable with spotty network coverage. And the browsers were terrible. The browsers were terrible. The cell phones crashed left and right. You know, updates were pushed that break the device, whatever. And for the early adopters of cell phones, the screen was tiny and had 10 pixels by 10 pixels.

Jake Moshenko46:35

But you could tell it was the future, though. But you can tell it was the future.

Ari Shamash46:38

Yeah. Without a doubt, you can tell it was the future. But then, you know, things settle down, right? Like, I would have never given non-technical people access to some of those early devices. And yet now, the world, everybody uses one. You can't live without one. Yep. So same story with AI. I feel like it has to go through the same adoption curve before it becomes safe mainstream for everyone.

Jake Moshenko47:02

Okay. That brings me to my last question. And you may have hinted at your answer a little bit already, but we have these doomers who tell us that AI is going to kill us all, ruin the world, whatever. And increasingly, those doomers, the call is coming from inside the house. So now we've got Sam and Dario saying, hold me back, bro, hold me back. Humanity is on a potentially extinction-level path. Do you believe it? Are you an AI doomer or are you an optimist when it comes to this stuff?

Ari Shamash47:41

You know, let me answer this question in two different ways, and they're both analogies, and I'll come back to this whole thing. I love Calvin and Hobbes, and one of my favorite strips of Calvin and Hobbes, Calvin asks his dad, how do you know how much weight you can put on a bridge? And his dad answers, well, they go and build the bridge and then they take a truck, they drive it over the bridge. If the bridge doesn't crash, they take a bigger truck, they drive it on the bridge until the bridge crashes. And then they know, hey, that's how much is too much. They go rebuild the bridge and then they post the song that says 20 tons. 20 tons. And then you hear the mother in the background saying, if you don't know the answer, just tell them you don't know. This is so my brother is also an engineer. He happens to be involved in building buildings. And I imagine, you know, when they first started building skyscrapers, people immediately said the doomsdayers came out like we can't build a building that's 20 stories high or 10 stories high. How are people going to get out? How is this? How is that? And yes, there were disasters in the beginning, like buildings fell over because they weren't built properly. Because at the end of the day, we're building with complex systems that don't have answers on the right way to build them. We have to build them through experimentation. We're not just going to go build a building and then have it fall over. Yes, we're going to put a lot of thought into it. But at the end of the day, any early adoption comes with risk. And we have to build it. Let's put all of our best thoughts behind it. But let's go build it in such a way that we think it's not going to fall down. I don't know that we're using... And one of the things that... Let me take a step back. One of the things that I always found interesting about the software industry is that we don't learn from the other engineering disciplines. Like, I feel like building a building has become something that the human race knows how to do. Like, there's tremendous planning. You always think about the support structure. There's a couple buildings here in New York that would beg to differ. Yes, okay.

49:31Software refuses to learn from the other engineering disciplines

Jake Moshenko49:49

Some of them are leaning.

Ari Shamash49:50

Some of them are falling. I'm not saying it's a perfect science. But all in all, like, a lot of thought goes into it. Whereas for whatever reason, the software industry, maybe because the cost from idea to implementation is so low and shrinking that we think we can just jump to a prototype and put it out there in the wild and let it rip without really thinking about all the guardrails around it. I'm an optimist. I think this is a revolution that's going to make the world a whole lot better. We do have to go back to the basics and think about what does security look like? What does identity look like? What does authorization look like? What does all the different controls that we need to put in place look like? We're going to get it wrong, right? We have gotten it wrong as an industry. It's caused some damage already. We're going to continue. The key is to learn from our mistakes and continue building technology, but do it in a safer way. I do think we need to regulate this as a human race. We do need to regulate any technology, any one of them. I don't think, you know, some of the regulation that we're talking about is extreme oversight by governments. I don't know that I will be that heavy handed. I think the answer is somewhere in the middle between no wild, wild west and extreme oversight. There's some answer in the middle that is correct. I think we do have to do some deep soul digging or we have to have the real conversations amongst all of us as an industry in order to figure out what the right balance is. But I do think I'm definitely on the optimistic side. I do think we have the engineering disciplines to get this right. We just have to apply those.

Jake Moshenko51:40

Yeah, great. Do you write about this stuff anywhere? Is there anywhere people who are listening and might say like, oh, I like that Ari guy. I want to hear more of what he has to say.

Ari Shamash51:56

Yeah, that's a great question. It's something that I should have been doing a lot more throughout my career. Maybe I will take this as a commitment and I'll say here that I'm going to start talking about these kind of things a lot more transparently.

Jake Moshenko52:09

Cool. Cool. Great. And if people want to get in touch with you, LinkedIn.

Ari Shamash52:14

LinkedIn. Perfect.

Jake Moshenko52:15

All right, LinkedIn. Awesome. Well, thank you so much. Or maybe they could just sign up for your classes. I forgot to mention that Ari is a teacher, computer science teacher. I teach computer science at a local college here in November. Yeah.

Ari Shamash52:26

Very excited about that, too. I've been doing that for 10 years, and it's one of the most rewarding things that I've done in my career.

Jake Moshenko52:30

So if you want to talk to Ari, LinkedIn, or if you really want to hear a lot, sign up for one of his classes. Amazing. Thank you. All right. Yeah. Thanks for joining us.

Ari Shamash52:39

Thank you.