Ole Lensmar (00:00) Hello, and welcome to the Cloud Native Testing Podcast. I'm super happy and thrilled to be joined by Alex Buccino. Is that right? Buccino?
Alex Buccino (00:07) Buccino, like cappuccino.
Ole Lensmar (00:09) Like cappuccino. Okay, Buccino. Great to have you. Alex, how are you?
Alex Buccino (00:13) I'm doing well. Thanks for having me.
Ole Lensmar (00:16) Great. You and I have talked off-podcast about things, but why don't you please give a short introduction, and then I'd love to home in on the topic that you are the expert in.
Alex Buccino (00:27) Sure. I've been in the tech industry for about 30 years. I started off doing mainly front-end work in the early days of the web. I had the privilege of building the original missionimpossible.com, which I heard Tom Cruise said was, quote, "pretty cool," which I will accept.
Then I went on to doing QA work for a consultancy, working with lots of different clients in different industries. That led me to an interest in configuration management, which led to CI/CD pipelines—although at the time that terminology was not used too much—and then eventually to Google.
There, I was managing the team that did the testing for Maps data infrastructure, which was a major challenge. It was the second implementation of Spanner. The testing challenges there led to us creating a bunch of new tools for large-scale integration testing. Those tools were wonderfully designed to be general purpose, and so they moved from just testing the Maps infrastructure to testing most of the systems at Google. They are being used by tens of thousands of engineers there.
Most recently, I was working at Squarespace as their Director of Developer Experience, working with their testing teams, CI/CD pipeline, tech writers, and developer tools to improve the tooling and development stack there.
Ole Lensmar (02:07) Awesome. That's an amazing resume. I'd love to go back to what you did at Google. The scale of testing there, I'm sure, is pretty unusual, or something that most people never have to grapple with. You mentioned creating some tools. What can you share about the challenges there and what that tooling was focused on?
Alex Buccino (02:32) Sure, absolutely. At the time, there was a tendency to have essentially one production environment and maybe some very limited pre-prod environments, but mainly the software was getting tested on people's local machines. This was not very similar to the production environment.
The team that I was directing had a brilliant tech lead who came up with the idea of modeling these systems under test in a way that could be reusable and composable, and could be deployed on Google's internal cloud system called Borg. So he created a system called Sandman, which stood for Sandbox Manager. The idea was that every developer could get a mini version of Google Maps with all of the data and the services required to run a Maps instance in their own little sandbox, and they could play around with it and deploy new code.
One of the critical features of it, compared to some of the solutions that existed already in that space, was the separation of service and state components. State components, such as Bigtable or Spanner, were modeled as first-class citizens within the environment. You could bring up your own data source either empty or from a snapshot of production or a backup, and do your testing against a known good data source or a known good set of services with test data. You were able to do a kind of mix-and-matching that you were never able to do prior to that.
Ole Lensmar (04:25) And were those stateful services always ephemeral, or could they be shared and reused by a team over the span of a project?
Alex Buccino (04:38) Inherently, Sandman didn't have much of a concept of whether the instance was shared or not, or what the lifespan of the instance should be. We did later add TTLs (Time to Live) that you could apply to these environments.
There were pre-prod environments managed by Sandman that were long-running, but the more common case was ephemeral instances scoped to a single developer, or maybe shared by a small team. Inherently, there was no limitation on how many people could use it or how long it would live. You could always extend the time to live on an environment if you needed it around for longer.
Ole Lensmar (05:34) Would Sandman spin up all the services in the target application, or was there intelligence to just spin up the services you needed for the tests you were running?
Alex Buccino (05:48) I wish it had gotten to the point where it could automatically detect what services you needed for a particular test. That was certainly on the roadmap. It's possible in the intervening years that things have developed along those lines. But at the time I was at Google, you had to manually specify what services were in a sandbox.
That said, we always had this concept of the "horizon." If you looked at the fully connected graph of all the services within Google, you would find that eventually, all services were connected to each other, usually because of some connection to the user identity stores. Those are always at the center of the universe of any system graph. So you had to have some concept of the part of the fully connected graph of services that you were interested in for the context of this set of tests.
You would design a sandbox so that it had all the services you were likely to need based on your knowledge of how they talk to each other. For anything beyond that, you could either have it talk to an already existing production or pre-prod instance outside of the scope of your sandbox, or you could use a fake service. Sometimes we called them "lightweight reals," which had most of the business logic of the service but wasn't connected to everything.
Sometimes, if you knew for certain that you weren't going to need a particular interface, it was possible to just say, "Connect to this null service." It meant the service required a connection address, but in fact, nothing was behind that address. It was just empty space because we knew we weren't going to actually call it within the context of the tests.
Ole Lensmar (08:10) We've had previous guests whose tools are used to stub out or mock external services. If you're standing up a system under test and you have dependencies you don't want to spin up, you inject mocks, which could be based on recorded traffic or just null services. Did Google build a convention around how to create these mocks, or was it up to each team to fill in that gap?
Alex Buccino (08:46) It was really up to each team to do that. That was another one of those innovations that I really wish we had gotten to. But Google at the time was far too heterogeneous an environment to bring up anything that was generically useful in that regard. I think in many ways, the outside world is probably in much better shape in that regard, with far more standardized tools that can be used for this kind of work.
Ole Lensmar (09:20) Would you use these sandboxes only for functional testing, or would you also do load testing, performance testing, or chaos engineering?
Alex Buccino (09:31) There was some of that going on. They certainly were not specifically designed for any one type of testing. As far as performance testing goes, it would probably be small benchmarks. I don't know of any cases where somebody brought up a massive scale job with many replicas to try to impersonate a fraction of real-world traffic. It was much more about benchmarking performance based on synthetic benchmarks or collections of pre-recorded traffic. Teams definitely used them for that purpose.
I am not off the top of my head aware of anyone doing chaos testing with them, but there's no reason why you couldn't. With tens of thousands of engineers, I certainly wasn't aware of all the use cases people had for it. Often, we would only find out what people were doing with it if we happened to know someone on their team, or if they came to us because something wasn't working as intended.
Ole Lensmar (11:02) Because they had a bug they didn't want to fix themselves! To me, the concept of provisioning testing environments seems to have so much potential. If you're standing up a system, you can instrument it to do traces or collect data and logs specific to the test execution, without the noise coming from elsewhere, and then aggregate that for root cause analysis. Were you doing any of that?
Alex Buccino (11:45) There was some of it. The technology of tracing has evolved quite a bit since I was at Google. It was one of those things where we wished the tooling was better. There were theoretical things we wanted to do in that regard, and we did collect all the logs from these environments so you could do analysis. But it was often disappointingly hard to filter out the noise of the environment versus the actual test signal. Noisy environments and noisy logs were a massive problem there that I wish we had gotten better about.
Ole Lensmar (12:37) Another question that often comes up is around data. If you're importing snapshot data from a production environment, did you have to scrub or anonymize that data? Was that built on custom tools or processes? How would that have been injected into the pipeline?
Alex Buccino (12:58) One of the principles with Sandman was that we wanted things to be as declarative as possible. The basics were built on an internal DSL (Domain Specific Language) that was declarative. We tried whenever possible to elevate common patterns into something you could specify declaratively within the language.
The initial functionality around state components was either bringing up an empty instance of a database or copying a database from a running instance or a backup. The really nice data stores were the ones that had copy-on-write semantics, so the copy was instantaneous, and then you just diverged as you wrote. Those were great systems for testing.
Later on, there were tools around transformations you could do declaratively, based on another set of data management tools Google had. It was basically a hook that said, "As you are loading this data, perform this transformation on it." Prior to those hooks, we told people to do that as their export. They would export from a database, run it through their data cleaning process, save that off as a snapshot, and then say, "For the purposes of this test, load the latest sanitized snapshot." You could load that any number of times for any number of tests until the data got stale enough that you wanted to do a new export. For almost all use cases, that was perfectly sufficient.
Ole Lensmar (15:33) That makes sense. Another thing is managing dependencies—how do you manage the order in which you start things? Was that also managed declaratively in Sandman, to say "wait for this to start before you do this," or "run these in parallel"?
Alex Buccino (16:00) Our basic principle was that we wanted to start as much in parallel as possible. We conceptually had two different types of dependencies: startup dependencies and runtime dependencies. Obviously, everything within the sandbox was essentially a runtime dependency of each other.
But startup dependencies were system specific. At Google, the general pattern was that services expected their state components or databases to be available when they started up. There was something we considered to be a real anti-pattern: if a service didn't find its database during initialization, it would do a hard quit. Because that was so common at Google, the easiest thing for us to do was make it a base rule in Sandman that all state components had to be created before you started creating your service components.
Many services could be started in parallel. In the rare cases where a service depended on another service being live before it would initialize, we put in a parent dependency. That just said, "Service A has to be healthy before Service B can begin its initialization."
But we discouraged people from designing their systems that way. Ninety percent of the time when people said Sandman was slow, it turned out they had put in unnecessary parent dependencies. The services could just come up and then poll for the existence of their neighbor. Having to wait to even begin to start a system until another system is saying "I'm ready to receive traffic" is a ridiculous way to design distributed systems. Whenever we found that, we would talk to the service owners and tell them they were slowing down their ephemeral instance creation unnecessarily because of an assumption they made back when things ran in production or on a local machine.
Ole Lensmar (19:28) Fascinating. Thanks for sharing all that. I'm guessing most people who listen to this don't have the capacity of Google to build their own system. If someone approaches you and asks how to manage the system under test while using Kubernetes, do you feel like that's a good head start? What recommendations would you give them?
Alex Buccino (19:59) Absolutely. Being in various types of cloud-native environments is incredibly important. The elasticity to bring up lots of jobs in a declarative way is really the critical piece. Languages that allow you to specify the components of your environment declaratively are very important. We tried to impart that knowledge to the teams developing Google's external cloud. I don't know off the top of my head any specific piece of technology I would suggest as a foundation, though.
Ole Lensmar (21:02) We've seen people using tools like vCluster or Argo to provision ephemeral namespaces and then use Helm to deploy their applications as a testing environment. It's representative, but it doesn't really manage the state aspect, and there's still a lot of wiring to do. Is that similar to what you've seen, or have you seen other approaches based on open-source tools?
Alex Buccino (21:35) That is generally what I've seen. I wish the tooling looked a little more declarative. As I said with Sandman, every time we found a common use case where people had to leave the declarative world to do something imperative—like performing a specific order of operations that couldn't be statically analyzed—we would step in and ask, "What is the common pattern here, and how can we make it declarative?" I don't know all the tools people might be using today, but off the top of my head, I haven't seen an exact parallel to this.
Ole Lensmar (22:38) When do you think a dedicated tool is necessary, versus just getting away with a Helm deployment in a dedicated namespace? Is it when you need to manage test datasets, or when you have hundreds of microservices and you just want to deploy subsets of those for specific tests?
Alex Buccino (23:23) I think it's a matter of convenience and how much pain you're willing to go through when you need to make changes. The issues people run into usually happen when they try to do some sort of migration. They want to make a consistent change across all of their environments.
The more imperative the process is for creating that environment, the harder it is to make sure that the change is consistent. With procedural scripts, the end result is the process of running through those steps. If you want to apply a change across multiple environments, you have to figure out exactly where in that list of procedures you need to make the change. When that becomes painful is going to vary dramatically based on how many environments you have and how many different procedures you're running to create them. That's the core problem declarative configurations solve.
Ole Lensmar (25:21) Great. Fascinating. Thank you so much, Alex, for sharing. This is super interesting to learn about all the intricacies of systems under test and environment provisioning. It was great to have you on the podcast, and thank you so much for joining.
Alex Buccino (25:34) Thank you very much for having me. I really enjoyed it. Bye.
Ole Lensmar (25:37) Okay, great. Take care. Bye.