The Enterprise AI Show
The Enterprise AI Show explores the AI journey for Enterprise companies around the world. [formerly The Cloudcast]
As the AI revolution moves from experimentation to execution, The Enterprise AI Show provides the clarity needed to lead. Join Aaron Delp and Brian Gracely as they explore the intersection of generative AI, enterprise systems, and global business strategy. Each episode features clear-headed conversations with the people making actual decisions—founders, investors, and practitioners—focusing on the technical architectures and business models that drive real-world ROI.
New shows every Wednesday and Sunday.
Topics: Enterprise AI strategy · The AI Economy · LLMs in production · AI leadership · Agentic AI · Digital Sovereignty · Machine Learning · AI startups · Cloud Computing
The Enterprise AI Show
How Open-Source is Reshaping the AI Infrastructure Stack
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Aaron interviews David Aronchick, CEO @ Expanso (former PM lead for Kubernetes, Kubeflow co-founder, and open-source ML leader at Azure) about how open source is reshaping the AI infrastructure stack. Aronchick recounts his path from early Linux and enterprise work to launching Kubernetes and GKE, then creating Kubeflow in 2017 to orchestrate end-to-end ML workflows on Kubernetes. The discussion centers on gaps in AI infrastructure, especially reproducibility and determinism across hardware, drivers, OS, packages, and data lineage, arguing Kubernetes alone can’t fully solve it. They contrast open weights with true open-source models, noting that real openness would require reproducible training data and infrastructure. They explore “AI-native” enterprise architecture, the role of open-source harnesses/wrappers to add deterministic controls, and growing edge/distributed compute needs driven by governance, compliance, bandwidth, and hybrid deployment realities.
SHOW: 1061
SHOW TRANSCRIPT: The Enterprise AI Show #1061 Transcript
SHOW VIDEO: https://youtu.be/kpQg3YIIUL8
SHOW LINKS:
- Expanso homepage
- TechArena, "Expanso's David Aronchick on Data Gravity and Pipeline Debt":
- Open at Intel podcast, "Data Privacy and Efficiency with Bacalhau Compute Over Data"
SHOW SPONSORS:
- NordLayer - Use ENTERPRISE10 for 10% off
- Nasuni - Activate your data for AI and request a demo
SHOW TOPICS:
You have a super interesting background (First managing PM for Kubernetes, Co-founded Kubeflow, led open-source ML at Microsoft Azure). Give everyone a brief introduction and how you became so involved in open-source and the Enterprise
OSS topics:
- Back when we were The Cloudcast, we covered K8s in depth, but I’m not sure we ever did a show on Kubeflow. Kubeflow tried to bring Kubernetes-style orchestration to ML workflows. Looking back, what did that generation of open-source AI infrastructure get right, and what did it miss that the current wave (agents, inference at the edge) is now having to solve for again? Oh, and maybe give a quick intro to Kubeflow as well for those that aren’t familiar
- Zooming out - open source shaped your whole career, from Kubernetes to Kubeflow to Bacalhau. Where do you think open source has the most leverage in the AI infrastructure stack right now, and where do you think it's losing ground to closed, vendor-controlled platforms?
- What are your thoughts on “OSS models”? Today, OSS really means open weights. Do you think there will ever be a truly OSS model? What would it take? Thoughts on the state of the industry?
A couple of Enterprise “grab bag” questions for you on a few different topics while we have you:
- "AI-native" gets used a lot and means different things to different people. What does AI-native actually mean for enterprise architecture in your view, and how is it different from just bolting AI onto an existing cloud or data stack?
- Regulatory and data residency pressure keeps coming up across industries (telecom, healthcare, financial services). How much of the edge/distributed compute push is being driven by AI performance needs versus governance and compliance requirements? Which one is the bigger driver right now?
CLOSING: If anyone is interested, what’s the best way to get started?
FEEDBACK?
- Email: show @ the enterprise ai show dot com
- Bluesky: @TheEntAIShow.bsky.social
- Twitter/X: @TheEntAIShow
- Instagram: @TheEntAIShow
Good morning, good evening, wherever you are, and welcome back to the Enterprise AI Show. This is your host, Aaron, and we are back to our interviews again after a short summer break. Today's topic is how open source is reshaping the AI infrastructure stack. We talk about Kubernetes, Kubeflow, and then dig into a variety of topics at this intersection of AI infrastructure and open source. Now, this one runs a little longer than usual, but it's worth it, and all of that comes up right after this quick break Today's show is sponsored by Nasuni. There's a growing gap in AI right now between what's possible in theory and what successfully works at scale inside an enterprise. The difference comes down to unstructured file data. Many AI initiatives struggle because the file data they depend on is scattered, unstructured, and disconnected from where and how work actually happens. Nasuni changes that. It brings your unstructured file data into a single secure foundation so AI, both generative and agentic, can access it with the context, governance, and performance it needs in production. Bring AI to where your unstructured data lives. See what it takes to activate your data for AI and request a demo at nasuni.com/ai. What does the outside world already know about your company? Cybercriminals could be seeing leaked credentials, compromised session cookies, exposed infrastructure, and even impersonations of your brand and executives. NordLayer Intelligence by NordStellar gives security teams visibility into those threats across the deep and dark web, data breaches, attack surfaces, and brand impersonation all in one platform. Find out what attackers know before they can use it. Visit nordlayer.com/intelligence/enterpriseai and use code enterprise10 for ten percent off your NordStellar plan And we're back, and as we promised, we're getting back into our interview series. And to kick things off we have, uh, somebody who's been in the industry a good bit but, uh, newly introduced to myself, and, uh, certainly we've had some great conversations here recently. We are gonna be talking about open source today, and we've talked about open source a good bit on the show before, but we wanna talk specifically about how open source is reshaping the AI infrastructure stack. And for that, we have David Aronchick, CEO at Expenso. David, how you doing, man?
David:Good. How are you?
Aaron:I'm good. I'm good. So first of all, you have super interesting background. You know, I a lot of kind of accomplishments here. First managing PM for Kubernetes, co-founded Kubeflow, led open source ML at, at Azure. Give everyone a brief introduction in how you became so involved in open source and the enterprise.
David:You know, it's kind of fascinating. I really fell into it more than anything. I, um, you know, was originally like a, a pretty hardcore Linux nerd way back in the day. I installed whatever, Slackware and Debian, you know, when you used to get them on CDs. Not DVDs, CDs. Actually, I should actually probably have a, a boot on a one, you know, 1.44 floppy somewhere. But, um, uh, so I, uh, I, uh, but then I, I, I left the, you know, the just the open source and I, and one of my first two startups and, and went to work at Microsoft, and that's where I really got into, like, enterprise software, and I was doing a bunch of, like, Linux stuff for them ev- even then, um, as they were building out, you know, their own Linux strategy, and so many folks were working on it. Then went off to do my other uh, my, my third startup and, and, uh, came back out and, uh, I joined Amazon who was very, very big in, in Linux at the time, and then Chef, who obviously has an enormous, uh, enormously successful, uh, open source platform. I, you know, I was always a big fan of it because the idea of, like, how you bring people together is so powerful around software, and software is, one of the most collaborative things you can, experiences in the world. Like, you can take a pull request from anyone in the world and it, it's just as good. Um, all that matters is the code, the quality, you know, whether or not it achieves the goals and so on. So all of that was out there and then I just absolutely fell, luckily more than anything, into to Kubernetes. Craig McLuckie, one of the co-founders had identified me and, and we started talking and he actually had to convince me to take on the Kubernetes job mostly because at the time people real- knew how deeply involved Google was in open source, right? They were core contributors to the Linux kernel. Uh, they had released a, a number of different amazing papers, you know, the things that became Hadoop and so on. But they really hadn't developed a Open source project first thing from inside Google. You had, you had Angular and you had, um, uh, not, um, I'm blanking. Android, of course. But those were either come from external or they were very, very localized. So they, they didn't have the kind of reputation that they do now. And Kubernetes came along and it felt like a science experiment, right? Uh, you know, Docker was, at the time, and, and someone will probably correct me, maybe six months old. Um, so I was like, "Ah, you know, containers, like LXC sounds interesting," and obviously, you know, zones and th- there's been a lot of, uh, history to this. But anyhow, so Kubernetes is there. I see a lot of value in just the idea of this distributed orchestrator. Craig's like, "Come on, you know, I got six other projects I gotta go work on. You manage this, uh, for me for Google Cloud and launch the GKE project," which I did. And then, uh, and then, you know, just obviously, uh, the train, the rocket ship left the station and, and it was just immediately obvious that I was never gonna look back. After that, I was like, "Well, you know, Kubernetes is great, but this AI sounds like a thing. Somebody should work on that." Nobody wanted to, and I know you're like, "Really?" Yeah. In 2017, nobody wanted to work on AI. We'd been through 14 AI winters. Uh, everyone's like, "AI is dead." I was like, "I think there's gonna be an AI thing for Kubernetes. I'm gonna invent it with three, two other co-founders from Google and a bunch of people in the community." And so we created Kubeflow. And then, you know, obviously that, uh, did a bunch of work on that. Then I went over to Microsoft, lead OpenAI strategy. And then while I was there, OpenAI, uh, the, has the ChatGPT-3, uh, moment and, um, everything transformed from there. So, you know, I, I wish I could say I, I drew myself into it, I, I… This was a plan. I'm telling you, it was just, uh, random luck
Aaron:Yeah, well sometimes that's the best, right? And I, I distinctly remember, yeah, the early days of Kubernetes when, yeah, to your point it was like Docker Kubernetes, and, and Mesos were all like- Absolutely… k- competitors for a hot moment until-
David:I mean, we, we-
Aaron:Kubernetes kind of took it all
David:over … you know, i- in the team we genuinely would've been totally okay if Docker or Mesos had picked up Kubernetes. And in fact, we offered it to both of them, right? We said, "You take this. You lead it. Go." And, uh, neither wanted it. Now, again, now don't look at this as some kind of like, oh, they made this terrible mistake. There were lots of strategies there, and there was lots of good reasons to not pick it up. Oh, we didn't, you know, it doesn't really like perfect fit for our platform, not this, that, and the other. Those are perfectly valid reasons. Who knows what along the way caused Kubernetes to win versus, uh, these other ones? Unclear. But you know, it wasn't a bad decision, it was just like, look, we have our own thing and we like it and we control it. It- those were also open source. Swarm and Mesos were both open source as well, so it's not like it, one was open source and one wasn't, it just happened that one won.
Aaron:No, I agree. So let's, let's kind of, 'cause we've talked about Kubernetes extensively over the years back when this was the Cloud Cast. Yeah. Um, but I don't know that we ever did a show on Kubeflow. And maybe that was exactly what you're talking about, like Kubeflow was almost maybe a couple years ahead of its time, if you will. But- Oh, yeah … but tell- For sure … tell me a little bit about that. Like, okay, the way I understand it is Kubeflow was like Kubernetes style orchestration for specifically for ML workflows, right? But-
David:That's exactly it. That's exactly it … anyway. And, uh, sorry, I'm doing a little,
Aaron:No, no, please go
David:ahead. There's a light reflection that I'm just catching right here. No, you're, you are exactly it. So the, the year is 2017. I have now been leading Kubernetes and GKE, uh, you know, having a great time, love the team, and I'm like, look, I think there's a space here. One, one of the things that Kubernetes really lacked was you would go in and I loved giving this demo. I would go and I'd spin up whatever, 100 nodes on GKE in, in a minute for a customer, and they're like,"This is great. What do I do with it?" I was like whatever you want. It, you're great. You now have this. Go, have a good time." And the- that gap of like, okay, I have this general orchestrator, what do I do on top of that, is still, I mean, arguably still there, right? And so what I wanted to do was pick a workload that I thought was going to be interesting valuable, demo-able, all that kind of good stuff. And at the time, um, I don't think Attention Is All You Need had come out. I think that came out, like, at the end of 2017. But certainly, ML was already enormous in particular areas. A lot of people say, oh, you know… I, I love to ask this question. What's the first platform that has ever had a billion daily users using ML? And people are like, "I don't know, ChatGPT?" And I'm like, "Nope. The answer is Google," right? Google Search in 2012 was all… their indexing was all ML-based, right? And they had lightweight e-e-inference models and things like that, that allowed you to, present that, those, those 10 blue links. Uh, nobody knew about it. It had this wonderful interface, but they had this amazing infrastructure called MLX that, that coordinated all these things together, and they started releasing these as public papers. So go back and look at TFX from Google about, uh, using TensorFlow, a way to string together all the steps of a machine learning pipeline. You wanna shape your data, you wanna control it, you want to, um, normalize it, you wanna split it into training and inferent-- or training and test hold back. You wanna tune your hyperparameters, et cetera, et cetera. There are all these, like, many steps involved Anyhow, and I--
my thought was like:"Look, this is a perfect fit for Kubernetes. It's multi workload, it requires loose coupling between the steps of the, the workloads. It, uh, is very hard to set up on your own. There is no simple way. It's like independently upgradable. Each component can be upgraded on its own, so on and so forth. This is a Kubernetes workload, and it's a complicated one that will enable us to take Kubernetes to the next layer up." So that was the idea. I had a lot of trouble finding the right people inside of Google to help work on it, and outside of Google. Most internal people at Google thought that nobody was going to have their own ML pipelines at all. And that's where they went down some very, very innovative stuff that still is out there today around, uh, auto training and, and abstracting a bunch of things fine-tuning and so on and so forth. But I still felt there was a self-hosted thing. I talked to a bunch of customers and they're like: "Yeah, we actually have some self-hosted problems. We would love this." And so I ended up finding, uh, Jeremy Louie, um, who's now at OpenAI who had created a, uh, what's called a CRD, that's a custom resource definition for Kubernetes that allowed you to spin up a small TensorFlow component.
And I was like:"Oh, what would this look like if we expanded this a little bit?" Where it wasn't just that one component, I think it was just training. That one component, but we, we tied it to a notebook, a self-hosted notebook or a self-hosted inference. What would that look like? And he had some thoughts on it, and we started like iterating on it. I then pulled in uh, Vishnu, uh, Kannan from the, uh, uh, Kubernetes team, who had done a ton of work on, um, GPUs and other accelerators. He's the third co-founder of Kubeflow. And by September, we had like something running on our machines.
We're like:"Look, this actually feels like it's there." I had gotten a talk accepted to KubeCon. That's back in the day when, when I could get t-co- talks accepted to KubeCon. Um-
Aaron:I believe- And- I believe talks in KubeCon and acceptance is a running joke in the industry now.
David:Yeah, exactly. Go ahead. Go ahead. Um, and, and look, I, I was there. I gave the first, whatever, four keynotes at KubeCon, and I cannot get my talks accepted. So if you're out there, like it happens a lot. Anyhow, so we, we had a talk. I changed the title of the talk. I have no idea what the original talk was. I changed the title, said, "I'm gonna make this about ML." I think it was called Hot Dog or Not on Kubernetes, uh, which was a reference to the Silicon Valley show. And, uh, by that December, we had it launched. We had, uh, several cloud providers signed up for it. Red Hat was signed up for it. Bloomberg was signed up for it. They were all testing it out. They were all using it, uh, and we got going. But it really was, Kubeflow was almost nothing at the start. It was a, uh, uh, y- a notebook, a, a, a training CRD, and if I remember correctly, an inference CRD, and those were very, very loosely coupled. It wasn't until the spring that the k- the, uh, Cloud AI team at Google submitted a very large PR, pr- uh, pull request on, uh, pipelining them together to make it very easy to take your data from… Like, once your notebook finished auto-triggering a training and triggering auto-training the, uh, triggering the inference. Uh, now obviously there are many, many more elements. Uh, K-Kubeflow has just graduated from CNCF and so on. I still, by the way, think that this declarative rollout of pipelines is still a unsolved problem in a lot of ways, but it's obviously much further along, and Kubeflow has both achieved enormous success, of which I have very, very little to do. But also it's inspired a lot of other people to build, uh, really inspiring stuff as well.
Aaron:So, uh, let me ask you this,'cause I, I wanna both relate it back to our topic and then ask a follow-up question to it all at the same time. So I would say, at least as far as I'm aware Kubeflow was, was one of those early days, like you were saying, like it's even pre-GPT and, you know- Yeah,
David:yeah
Aaron:the big, the big aha moment, right? And so, like how everything has advanced forward and, and specifically talking about AI infrastructure and, and open source within AI infrastructure, tell me a little bit about like that generation, that first generation of infrastructure. And like obviously we talked about what it got right, but like what did it… I won't say miss, because it's hard to predict where this has gone. But like where'd, where is the current wave, like meaning like agents and inferencing at the edge and all of this, like where is it a little lacking? Like where is the gaps and where's the growth potential?
David:Yeah, I mean, I so my biggest thing, my bugaboo, and, and everyone's got their thing, is I want repeatability, right? I'm a big declarative infrastructure guy. I wanna reduce the things that can be varied, imperative and envi- runtime specific as much as possible, because those are the things that are almost impossible to debug. They're impossible to roll out reliably and so on, and, uh, that becomes a really big problem. Now, what I was trying to do at the time was solve that problem via pure Kubernetes, right? Because that's what it- Kubernetes provides, right? It says, "You give us a docker with a, a, a specific hash and a specific deployment contract, and it will reliably, it declaratively roll that thing out for you." And that is good. But what I think we… One of the things we missed is that wouldn- was not gonna be enough, right? Somehow you needed a language that spoke below that as well as above that. So below, I mean, like, what CUDA drivers are you on? What does your memory bandwidth actually look like? All, you know, what is the OS version? All that kind of stuff. How do you declaratively name that? That was often a different tool, and there were a lot of people who were like,"Okay, well, we'll do this in Terraform, and Terraform will take care of this on this cloud, but it will hand off to another layer." And I was like, "No, man, y- we gotta wrap this up together."'Cause you- even today, right? I just saw yesterday, Kubernetes or, uh, NVIDIA released a, uh, a cluster readiness um, toolkit, and I gotta dig in but, uh, some very, very smart people have been working on that. Uh, like some former Kubernetes, some former Mesos folks working on it, so I do trust them that they are- they're gonna get it. But it's a perfect example where it's like, you know, if you go to cloud A, you go to cloud B, you go to on-prem C and you're like, "I… Okay, they're all running whatever, Ubuntu 24.02, and da, da, da, da, da, da," right? They… We should be the same, and it turns out they're not. And just declaratively naming that makes things really hard. Why do people care? Well, one, a lot of your folks are just gonna run into big challenges because the the, the… They're gonna be like, "Why did this perform this well here or that well there?" Or maybe just longitudinally over time, right?"Hey, you know, we ran this, this training set last week, and it w- you know, here was the results, and we ran this training set today, and here are the results, and the results are different, but nothing changed as near as we can tell," and it turns out some random ass Python package is different. You're like, "Oh, crap," right? So that's a problem. If you want my vision for the world- This is how you get past things like scientific re- reproducibility. This is how you get to actual intelligence being shared. Go and talk to any scientist in the world, and they do this weird archeology when they pick up a paper. Even if it's got everything listed in there about, all the details of the experiment in incredible detail, it's something is wrong. They didn't measure the humidity or, oh, they didn't realize that a truck drove by, or whatever. There's just so much non-reproducibility. And so I'm not saying that we're gonna get there, but if you, if you started to build, to build things like this, if you said, "Hey, in order to run this scientific paper, this is specifically the entire environment that I used to set up," you can reproduce that. And if you get different answers, which you will 'cause the world is not reproducible, at least you can narrow down where the reproducibility failures have occurred. And I just, I have such a huge inspiration for that. But the, going all the way back, I thought th- to your question, what did we miss? We missed that reproducibility is still a massive unsolved problem, that it spans far beyond what Kubernetes does. But I would argue that we're just as bad today, right? Like, and maybe even worse because so many solutions, endpoints out there right now are just black boxes. You don't own the data, you don't own the training, you don't own the version. I, I love, you know, I'm u- I use these, uh, AI tools all the time, and they update every day. Every single day they update, right? I come in, I click that button every single time. Oh, update required, just restart. Every single time. You're telling me that I'm gonna get the same results as I gave yesterday? Of course not, right? Because there's no, like, transparency, and there's no way to inject determinism and things like that. So I we're still facing that today. I thought we could get it done then. We d- couldn't. And the failure there was twofold. One, it was hard to detail everything, but two even if you could detail out everything, there wasn't a unified tool to take a bunch of metadata and say,"Okay, now I know how to blow this out." Today, we're running into that problem even worse because so much of the data is hosted, so much of the models are hosted, you know, and, and, you know, there we have even less, less reproducibility available.
Aaron:Yeah. Well, okay, so this is, uh, super fascinating because, coming from an infrastructure background, and one of the biggest things I see is AI infrastructure, to your point, like a lot of thing, things you just mentioned, AI infrastructure today done properly is hard on-prem. Mm. Um, you know, and, uh, I'll go back to actually it has a, a, a Kubeflow tie-in. Uh, my first, my days in Nutanix, you know, we did a Nutanix AI, you know, obviously on Nutanix converged infrastructure kind of AI thing, and it was based off of Kubeflow and open models and things like that. Basic. But the biggest p- problem we had was Even with convergent infrastructure, damn, it was hard to get that stuff going. And the amount of people- Absolutely … that have the top to bottom expertise is hard. And so, like, the, I guess my question to you is, like, okay, where is the biggest opportunity for all of this? Because in my mind, uh, you know, open source and, on-prem or not is kind of losing ground right now because of the closed vendor control platforms because of the simplicity. Like that click to reload, if I use that term, is both a blessing and a curse.
David:Yeah. Look, I, I gotta be honest with you- And so, like, tell
Aaron:me a little bit about that. Like, where's your
David:head at around that? Yeah. You're putting your finger exactly on it, right? You know, I w- I, I love Linux. I've used it for many, many, many years and FreeBSD and all these things. They have never, never gotten set up and reliable, like ease of use right. And I don't know why. I mean, like, arguably, like the affordance, the, the interface that they're looking for is different, right? I want to go in. I remember the first time I ac- uh, I, I have something in the, uh, uh, Linux, uh, core tree, uh, in, uh, 19, uh, 98, something like that. I had a cheapo, you know, $40, which was incredibly cheap at the time, uh, Ethernet card, and the version of Linux that I was using was wrong. Or, and the, the driver was wrong. It kept crashing when I compiled it. I went in, I found the code, I debugged it, and submitted it as a patch. It was just some stupid thing. And, um, it's submitted, and I'm still very proud of that today. I don't know. It's probably long, long, long, long gone. It's like two lines of code. That is a-- So that was incredibly useful for me 'cause I'm a giant nerd. If you, if I actually had a job, that would've been incredibly miserable, right? That's a terrible user experience. I just wanna click a button and have it upgrade. So I would argue that, that a lot of these things they're trying to serve two masters. They're, they're trying to serve, again, and this goes to not just Linux, right? Windows and Mac, and everyone has the same problem. Some people are just like, "I have a job I wanna get done," and some people are like, "No, I want to solve this and make this optimized," and so on and so forth. And by serving those two masters, you run into this, this challenge. Now, to your point about the, the, um, you know, click to upload and things like that, look, 99.9% of people have a job to do and these things are simply functions of that job."Oh, I'm gonna use AI for this, and I'm gonna use Excel for that. I'm gonna, you know, use whatever Sed or, I don't know, R Studio. I have no idea how to do this other thing." Those are just tools in their the pursuit of their overall goal. The failure is that we ask them to go and understand what CUDA driver they're running, right? And so it's that mixing that I think ends up being this, this incredibly hard challenge. And If we wanted to make progress against this, it would be for those people, giving them the ease of use to, have that, "Hey, I'm- I have other jobs to do. I'm just gonna work on this." But also give them the visibility or the transparency or even the voice to say, "Hey, I now want you to dip this environment that I'm using in Amber. And I want you to be able to reproduce this thing 100 times over because I need this. I can't have a single thing change here, 'cause if it does, then that means I need to go back and redo everything." Having them be able to do those two simultaneously, that is the era, the, the problem for the era. And what I'll say is to just to up the level of challenge, and again, this is what we're working on Expanso, but it's just generally a problem don't think about it as, people may have heard the term SBOM, Software Bill of Materials. I- there's something I'm really pushing on, another open source project that I'm helping to work on called the, uh, a Data Bill of Materials or, or other mechanism for doing data lineage. Like, how do you not just detail this stack of the stuff you're running on, but detail all the input steps of the data that came in, right? How do you know what was in your training data? How do you know what was in your, uh, holdback data? How do you know, "Oh, hey everyone, we're gonna convert this from, I don't know dollars to euros at this stage, on this date." We're gonna re- record that somewhere. Those are all inputs that happen all, even before you get your infrastructure stack lined up. That's also part of the determinism that we need to do. So it really is like all of the above. You need to think about holistically, how did I create this artifact at the end that allows me to do my job, and what were all the things that came in? And some of them will be reproducible and detailable, and some of them won't, and that's okay. We just need to accept that some of them were not. Do not try and get to 100%, you will fail. If we got to 50%, I'd be ecstatic. Right now we're at, like, 2%. So there you go. I, uh, I hope that answered- Yeah,
Aaron:no, it's… Co- completely agree with your assessment and it, it mirrors a lot of the conversations I've been having. But, uh, let me kind of move into, I'm gonna call this kind of the grab bag section, if you will. I'm gonna ask you a bunch of random questions here.
David:Go ahead.' Aaron: Cause we have a, a somewhat common your thoughts on some of these things that it may not, may not be as related to the Y- AI infrastructure, but it's kind of related to everything we're talking about and doing in our day to days. So first of all, what's your thoughts on, like, OSS models? And what I mean by that is today a lot of us say, "Oh, this model's OSS," but it's really not. It's open weights, right? And do you think there will ever be a truly OSS model, like, run by Linux Foundation or CNCF? Like, what would it take and what is your thoughts on, like- Something like that ever happen? So that is an excellent, excellent question. I know I talk long, so just, I don't know, scratch your nose or something- All good … when I, when I go- All good … when I go into a… What I will say is this, you have put your finger on something enormous, right? We have open weights, we have open source. I would argue even o- open source is not enough. Because think about what you're doing here, right? So open source means I have access to the… go look at Stallman's, like, original thing, right? I have access to the code, I can use it according to this license, so on and so forth, right? There's a bunch of things involved there. But even if I have every line that, uh, Mistral or OpenAI or Anthropic used to train a model, every code, every hyperparameter, every everything, that's not enough. That's not the model, right? The model is that I ran it, and I ran it against this training data. So you could have an open source model that is totally useless because you don't have one byte of data, or maybe you only have sample data or something like that, so I'm not even able to create it. I think a true open source model, and I think this is something we need to define, but, like, this is true, would have the data and the full stack of infrastructure and the compute or the, the code necessary to do that, that allows someone to reproduce the open weights that were released as part of the model. That's what open source modeling would look like. I think that will be very, very, very challenging for people to actually do, right? Because ultimately you are talking about petabytes of data that, that, that went into training. Now again, it may be not na- not directly, maybe it was a subset or sampling, things like that. But somehow petabytes of data produced that, those weights at the end. I think the best that we could get to reasonably would be maybe hosting some of that petabytes of data, uh, that allowed a third party to audit it, and allowed someone to hash, create a hash of some other artifact from that data to say, "Hey, in this model, this is the hash of the data that went in." It's still not gonna allow you to produce the open weights, but at least you have transparency into what, that the, the hashes that went in produced something on the oth- on the other side. I don't know. But the, in the spirit of open source, I think we will have, you know, in the name, in the letter of the law, I think we will have open source, but I don't think that will be what we need if we really wanna have the spirit of open source be realized.
Aaron:Yeah. Agreed. Agreed.'Cause I had a conversation with somebody, actually, you and I were, uh, at an event, that's where we got to know each other. A- and at that event, actually, it was a side conversation, and it was interesting of like, "Hey, what would it take to do something like this?" And it was like- Yeah … it would, it would almost be like the Kubernetes moment all over again, of like, somebody would have to dump a whole bunch of stuff, in. Uh, but to your point, even then, I don't think that's enough, right?
David:No, totally
Aaron:agree. Um-
David:I mean, to- So, okay … to your point, Oh, sorry. Go ahead.
Aaron:No. No, go ahead.
David:No. No, no. I'm just gonna say, when you, when you think about the values of open source, when you think about, like, what revolutionized, what that unlocked, it was the idea, right? Like, um, you know, uh, Linus comes out and he releases the kernel, and then there are, whatever, 14 different distributions, and each one of those distributions was somebody sitting around in their basement or their whatever office and being like, "Oh, I think we could do this. So I'm gonna, like, take this. I'm gonna fork it. I'm gonna do something else with it." And again, that's the spirit of open source, right? If we were gonna get there with models, that's what it would look like. It would be like, "Oh, here's all this raw data. Here's this code. Here's the infrastructure. I'm gonna take that. I'm gonna fork it. I'm gonna make some tweaks in one or maybe all three of those things, and then I'm gonna do this other thing with it." That's what open source would look like. That's where I think things get really, really, really hard because one or more of those things may not be available. Like, uh, you could say, like, "Hey, here's all the code for this model," but you need an NV-whatever 72, running for three weeks in order to get it done. Well, it's effectively not open source, right? It is. I can go look at it, but I can't reproduce it, I can't compile it, I can't, you know, whatever. It's challenging and, and I just don't know that we're there.
Aaron:Yeah. Agreed. Agreed. So let me, uh, I'm gonna jump onto the next one here. AI native, and I'm using air quotes around it. You just, uh, you can't really, you know, get that as much on the, the au- at least the audio version of all of this, right? Like, AI na- native, like, it gets, it gets tossed around more and more, but, like- For an enterprise, what does AI native actually mean from like the architecture point of view? And like, how is it different than just bolting AI onto the existing Kubernetes cluster or, you know, whatever else- Yeah we've got going o- inside the enterprise today? Like, what's your thoughts on AI native?
David:Oh, gosh. I, there, there's another long one. I mean, you know, at the end of the day, the, to be AI native, I mean, it was like being cloud native, it was being whatever. It means that the majority, if not all of your workloads, your jobs, your everything that you do every day has the ability to use AI, right? Without having to re-architect everything. And in that you say, "Okay, well, what does AI do?" Like truly, what does it do? What's the magic? The magic is it takes poor or unstructured input and it produces a meaningful output, right? And for ChatGPT, this was like taking my text and saying like,"I wanna go, you know, make my, email sound like Freddie Mercury," and then it could do that, right? So that's poor unstructured produces something that is like reasonably structured at the other end. So that's what AI native is to me. Now that said what that means is yes, having infrastructure that supports this, having your own training, but even if you didn't do anything like that, it also means two other things. One, that the data that you're producing today is ready to be used by AI downstream. And that means having native APIs that know how to, like, provide limited information, so you're not like, you know, revealing everyone's Social Security number when you're just querying, like, what employees are working where, that kind of stuff. It means adding provenance, lineage of tracking of the data as it moves through your system. So if I have 400 wind turbines in Iowa every one of them, when they leave, you know, it's not just a bunch of comma-separated values. It actually has all the information associated with each piece of metadata so that downstream AI can make better decisions, 'cause it's now got all that data painted. And it means, as you pass it through the system, it's all recorded so that in six months' time when you realize you need this, it's there and available for you. Not storing everything. Storing everything pr- Or storing the stuff properly, so if you wanna go back and recreate it, that's kind of… So that's a big thing, right? Really working on your data story even before you get to AI, that's part of being AI native. The other side of it is how do you build the systems to watch these non-deterministic systems a, you know, results, AI produce stuff. You know, if I go and, uh, query a web server six million times, I'm probably gonna get the same result six million times, right? Every so often, maybe it got a half a question or whatever and it's gonna get something wrong. Fine, that happens. But my s- it's not gonna generally, like, come back with, you know, the, the text of the Iliad. AI has the opportunity to do that, and what that means is that what we need to do, what you need to do to be AI native is say, "Hey, you know what? We're gonna build in the reliability to our system to say, 'If you don't- match the output we expect, we're gonna have a way to fail elegantly. Or, and, or, you know, just say, "All right, ignore. Moving on," or re- re-request. It, just that is also incredibly important as well. That's very different than the systems we built for many, many, many years. Uh, where, y- sure, if my, if the brakes on my car failed it would know how to, fail that. But like I said, if it, if it, uh, you know, produced a picture of, I have no idea, Tom Brady it would not know how to respond to that, and maybe it would crash the system. Being AI native means having a way to fall back elegantly when stuff goes sideways. So it really is, it's about where you can build these stacks, how you use them intelligently, but also what are the, the, the dependencies, either before it lands on AI and after it, it lands on AI, in order to handle this new and interesting system that we've built. That's what AI data means to me.
Aaron:Now, okay, a follow-up. I, I do, I wanna close this out on a, on an edge question, but a follow-up on that before we get to that. Based off of what you just said, and kind of putting a whole bunch of threads together here, if we wanted, uh, obviously something that's, you know, LLM-based, but LLMs are non-deterministic, so we want something that is more deterministic, and we want something we have better visibility and control over and is m- like, you know, I'm gonna say is, is open source, you know, in all of this- Is there potentially a future in open source harnesses that then wrap around the model and help control the model? Like I go, for instance, I mean, you, you look at all the different models that are make- or excuse me, all the different harnesses that are making a lot of the models better. You look at Claude Code and, you know, that notoriously, I think, kicked everything off. But like, I wanna say it was Qwen. Qwen has a pretty interesting harness that they open sourced recently.
David:Yeah.
Aaron:And there's, you know, we're starting to see more and more of that going on. Like, what is your thoughts on how to solve the problem with open source? Is it a, an, a harness open source wrapper?
David:Oh, absolutely. No, no. So these tools are you know, th- they're, they're all these models are just the start of what we're needing to do here, right? The wrappers of them, the harnesses, are what begin to form the criteria, the deterministic criteria of inputs and outputs. And so for sure, no matter how good your model is, it will never be anything more than an endpoint for input and output. And um, arguably the majority of the work today and this is pretty arguable, so I'm not gonna take too spicy a take here. But I would argue that any model that we have today could be used to create 95 % of the value for the next five years, like really. It's about how you structure it, how you structure input, how you structure the output, how do you iterate on it, how do you update it, how do you cache it, all those kind of things. And all that stuff is pure distributed systems, deterministic, um, you know, ar- around the edges.
Aaron:Yeah. Makes sense. Makes sense. All right, so I wanna close this out on a final topic and this also kinda goes to a little bit more to your day job kind of thing here with, uh, Expanso. So I'm seeing a lot more both regulatory concerns, data residency concerns, really across a lot of industries, but certainly the highly regulated indus- industries, healthcare, financial services, you know, things like that. Yeah. How much of, like, edge and distributed compute is being driven by whether they're AI performance needs, governance, compliance? Like, tell me a little bit about that space right now and specifically the drivers behind them.
David:Yeah, I think it's a really interesting thing and you hear so much about this with sovereign AI and, and, um, sometimes you're getting government regulations doing it, sometimes you're just getting individuals caring about this. It really spans the map. Again it generally is gonna come back to who thinks they're gonna get fired or sued or, um, fail in some, like, bad way. And edge is a component of a solution for any of those, right? And the component of the solution is it allows you to execute portions of your job in, in a more specific way. Um, and again, this is our day job. This is what we do with Expanso, uh, distributed compute and distributed data pipelines and things like that, allowing you to run A- AI at the edge. We don't build any models. In fact, we have, uh, Mistral, we have Quinn, we have all these kind of examples of running on-prem 'cause all we wanna do is help you run what you already know and love in those locations. But- What it enables is those people who have these problems to say,"Oh, there is a solution here." It doesn't just mean handing all my data, handing all this, like, proprietary stuff to an endpoint in the cloud and trusting that they have a ZD- a zero deten- uh, data retention policy. I gotta get out of the acronym, uh, world. Uh, or they trust that they're not using my models to train, uh, their own stuff on it. The fact that I can now get control of this is great. Now, I wanna be really clear. I don't think that means stopping it, right? I don't think that means not using a hosted endpoint. No. What I think it means is being thoughtful about what occurs in each of the steps. What do I need to do on-prem to obfuscate, to clean, to whatever? What do I need to execute over AI locally because of bandwidth or regulatory concerns or whatever? And then what makes most sense for me to execute against that B300 in the cloud, right? Those things span the spectrum. Right now, there's only the one, right? There's only the, "Well, geez, I guess I gotta use this thing in the cloud 'cause that's the only one that's any good." Recently, arguably, you know, in the past six months, you've seen phenomenal results by folks with on-prem or other solutions, uh, excuse me, um, uh, self-hosted, uh, models, maybe not open source, I'm, I'm not gonna, like, fall into that tar pit, but certainly self-hostable models that have started to change that. Again, it's not a question of whether or not they are better or worse across every benchmark. That is not required But if they can do some components of the things you need locally, and maybe entirely for some scenarios, then great. And people now have that as part of the option. So what I generally hear, you know, the answer to your question is, I hear people being excited about this, adding it to their portfolio of solutions, but still in the same way they're not want- excited about being all in on the cloud, they're also not excited about being all in on, on-prem. They're happy to figure out the hybrid that, that fits best.
Aaron:That makes sense. That makes sense. I think that's a great place to close us out as well. So, so David, if there's anyone out there that's kinda interested in that topic specifically, what's the best way they can get started?
David:Uh, I, uh, I'm infinitely online, unfortunately. Uh, I don't know how to break it. Uh, you can reach me on, uh, Blue Sky, Twitter. Uh, you can reach me on, uh, LinkedIn, uh, or so on. It's always uh, just about every place, it's just my last name is the easiest way to find me. Uh, there aren't a lot of us out there. Uh, Aronchick at, uh, all these places. Although on Blue Sky, I'm Iron Yuppie. Outside of that I mean, take it-- You can email me. Uh, again, it's just my last name. It's uh, same, same everywhere.
Aaron:Love it. Thank you so much for your time today, David. And everyone out there- Thank you so much… thank you very much for listening. We certainly appreciate your time this week as well. So if you could please, wherever you get your podcasts, if you enjoy the show, uh, please leave us a rating and a review as well. And we're always open to feedback through all of our socials as well. L- Both feedback on our episodes as well as ideas for guests as well. And so for Brian, who wasn't able to make it, uh, this week, and myself, thank you everyone out there for listening, and we will talk to everyone next week. Thanks for listening. Check us out at theenterpriseaishow.com for past shows, newsletters, and all things enterprise AI
People on this episode
Podcasts we love
Check out these other fine podcasts recommended by us, not an algorithm.
Software Defined Talk
Software Defined Talk LLC
Dithering Preview
Ben Thompson and John Gruber
Everyday AI Podcast – An AI and ChatGPT Podcast
Everyday AI
Prof G Markets
Vox Media Podcast Network
Acquired
Ben Gilbert and David Rosenthal
Decoder with Nilay Patel
The VergetheCUBE
SiliconANGLE, Media