Your data warehouse still thinks a human is on the other end of the query. That's a problem MotherDuck CEO and co-founder Jordan Tigani knows from the inside, having spent years as a founding engineer on Google BigQuery. This week on Dev Interrupted, he joins Andrew to make the case that one warehouse per agent beats the distributed systems he used to work on. He lays out Watertown, his riff on Steve Yegge's Gastown that recasts the data engineer as the manager of a fleet of agents, and walks through how MotherDuck's Flights, Dives, and just-launched Guides keep pipelines, visualizations, and business context as code reviewed in GitHub.
Show Notes
- MotherDuck: The DuckDB-powered data warehouse built for small data and agents at motherduck.com
- Water-Town: Read Jordan's take on the agent swarm data stack, his riff on Gastown, at motherduck.com/blog
- Flights: Learn how MotherDuck turns data ingest into scheduled Python scripts your agents can write at motherduck.com/blog
- Dives: See how MotherDuck replaced its BI tool with LLM-written TypeScript visualizations at motherduck.com/blog
- Guides: Read the docs on MotherDuck's new Markdown context layer for agents at motherduck.com/docs
- MotherDuck Blog: Keep up with all things MotherDuck at motherduck.com/blog
- MotherDuck Community: Join the community Slack at community.motherduck.com
- Connect with Jordan: LinkedIn | X
Transcript
(Disclaimer: may contain unintentionally confusing, inaccurate and/or amusing transcription errors)
[00:00:00] Andrew Zigler: But I'm really excited to have this chat with you, Jordan. You know, you're the CEO and co-founder of a pretty cool company right now, in my opinion. There's a lot of really innovative things coming out of MotherDuck, and your background as a founding engineer at Google BigQuery is, you know, just like a driving force, I think, in all of this change in narrative, and it makes so many folks look to you to understand, like, how data science is transforming.
[00:00:22] Andrew Zigler: So we're really excited to have this conversation today, talk about some of, like, the ancient history, even back in, like, the early 2020s. I know that's, like, a long, long, long time ago. But even right now, how MotherDuck, I think, sits at the end, uh, center of an argument that the practitioners of the data stack are AI agents now.
[00:00:41] Andrew Zigler: And data scientists and the way that we work with data and the tools that we use and communicate through with our agents on, they need to transform to meet that new challenge. Uh, so we're super excited to dive into all of this today. Jordan, welcome to Dev Interrupted.
[00:00:56] Jordan Tigani: Yeah. Well, thanks, thanks for being here. This is, you know, these are exciting times [00:01:00] and, uh, it's good to be in the middle of things
[00:01:02] Andrew Zigler: It really is. And you, you couldn't really be more in the middle of things. I wanna, like, back up for a second and talk about, you know, MotherDuck and how we got to be here, because, you know, your background was in big data. You, you built BigQuery and one of the biggest distributed warehouses, like, that does data, right?
[00:01:19] Andrew Zigler: And so we're talking about a fundamentally different approach from how MotherDuck works now, which I wanna dive into with you, no pun intended, but also understand, like, how that change came to be in your mind. Like, what did you understand about what big data was and what it couldn't be and what it needed to change into that, you know, you were maybe one step ahead of on the insights?
[00:01:42] Jordan Tigani: So, you know, we used-- When I worked on BigQuery, we often used BigQuery to understand what was happening in BigQuery, I remember doing some, like, analysis, you know, trying to figure out, okay, well what... If somebody was using Snowflake, what size Snowflake instance would they, would they need? Um, and so that we could kind of do some comparisons, [00:02:00] 'cause the way BigQuery worked under the, under the hood was very different than Snowflake. and so I was looking at kind of the size of data that people had and the size of queries that people were running, and w-- I-- Well, the thing I found was just like it was dramatically less than I had expected, and, you know, that like something like ninety percent of people, you know, of queries were a hundred megabytes.
[00:02:24] Jordan Tigani: Um, and, you know, megabytes is something that, you know, at, at Google people, people sneeze at. Actually, they don't even bother to sneeze at it. They sneeze at like Um, but like, um, you know, and just something you wouldn't think you'd need, you'd need sort of a massively parallel engine to, to do.
[00:02:40] Jordan Tigani: And like kind of as I, you know, I was looking like, wow, like most of our customers are actually using tiny data. And then even the, even the giant customers like, um, you know, we had like HSBC, and Home Depot, and Walmart, and like like some of the biggest co-companies in the world, like the stuff that they were doing wasn't that [00:03:00] big.
[00:03:00] Jordan Tigani: And, and I think, you know, internally was a lot of energy around, okay, well, how do we make bigger stuff work and make it, you know, faster to do tens of terabytes and hundreds of terabyte queries? Um, but the things that people were complaining about were the things that like their hundred, hundred megabyte query wasn't any faster than their hundred gigabyte query.
[00:03:19] Jordan Tigani: And like, and so actually the, the things that people cared about was like, "Hey, I've got a human waiting for this, for this result." Like, um, if you add a second to every query be, you know, because of all this distributed mechanism, then that's actually gonna make the experience of your users worse, um, than being able to handle, um, some giant, giant thing faster.
[00:03:44] Jordan Tigani: 'Cause even that-- those giant, giant things are expensive, and so people tend to not do them very often. so kind of like I filed that, you know, away and, um, it took me a couple of years. You know, actually work-- I worked at, um, SingleStore for, for [00:04:00] a couple of years working on their, their SaaS offering and, and kept asking for smaller instances.
[00:04:06] Jordan Tigani: And, um, and we, you know, we had one that was half the size of Snowflake and then the smo-- Snowflake's smallest instance, then people were like, "What if we did a smaller one?" Like, and so it was like, it seemed like there really was a need for, you know, scale to zero, um, serverless, um, kind of smaller instances as well as, as well as large.
[00:04:25] Jordan Tigani: And then the other thing that, you know, you just sort of realize is that, you know, these A lot of these technologies that we're, we kind of have built the data infrastructure around, were designed when machines were small. Like, you know, if you had a 100 gigabyte query, well, 100 gigabyte wasn't gonna fit in memory.
[00:04:42] Jordan Tigani: It wasn't gonna like, like, you needed... You know, if you wanted 100 cores, like, you know, you, you had to, you had to combine a whole bunch of machines to get 100 cores. And, uh, and nowadays, like the just basic machines that are on EC2, like the physical machines [00:05:00] you know, um, hundreds of cores, terabytes of RAM. And so, like you don't actually need these distributed, distributed systems. And if you don't use the distributed systems, you can just do things, you know, much more easily. you can be faster, lower latency, like a whole bunch of, a whole bunch of important, important things. Um, and, and so yeah, it was like, hey, if, if we're gonna do something... If you were gonna, if you were gonna build these things now, you would do them differently. And that, to me, that was a, it was an opportunity
[00:05:36] Andrew Zigler: Yeah, that's-- There's a really interesting, I think, parallel in there to unpack of, like, the idea that, you know, you even said, like, the, the hardware was smaller, it was older, and you couldn't even get some of those, like, mid-sized queries, uh, and those things to live in the RAM of the machine. It was always thought that we were moving towards having these bigger, monolithic kind of like services that could handle all of that for us because it was just so hard [00:06:00] to, uh, really visualize and understand how quickly the tools and the technology we were using were going to evolve and change and become more capable underneath us.
[00:06:08] Andrew Zigler: And so, like, we make bets on, like, what we wanna build and what we need for the future, but then sometimes we have to make the bets to take things away and to simplify them. And there's a lot of things I think we're learning about that even right now with, you know, how everyone's workflows are transforming with agents, you know?
[00:06:24] Andrew Zigler: There's a lot of temptation to add a lot of complexity, a lot of stuff on top because you think you can orchestrate it in this kind of way. But there's, uh, a, a much bigger opportunity sometimes in taking away stuff you took for granted or that isn't necessary anymore, and you can only discover by starting to hack at it, you know?
[00:06:41] Andrew Zigler: And so there's like a, a, I think, a lot of lessons there that have parallels about how we work with technology and how our expectations change. And it's really important to call out too that, like, data too became more widespread and just more accessible to more people. Like, it was-- it kind of moved out of the domain of just being solely something that the [00:07:00] data scientist or the data researcher would do at scale on these huge amounts of data.
[00:07:04] Andrew Zigler: It became something that was more approachable and embedded into more tools and delivered reports in all sorts of ways to where working with data became more democratized. And because of that, we needed smaller, more customizable, bite-size, on your device ways of working with that data to meet the demands, right?
[00:07:24] Andrew Zigler: So, like, understanding how to throw away something that you took for granted before and, like, trust a new definition of good. Like, part of doing that too is understanding what you think the next level of good will look like. So for you, was that a vision of the access of these kinds of tools just being more widespread?
[00:07:45] Andrew Zigler: Um, and, and why-- Uh, like, what are the opportunities that you still see for the tool to continue to evolve for now its, like, more agentic consumers?
[00:07:54] Jordan Tigani: You know, I think a lot of, in a lot of things in, in, uh, in, in technology, [00:08:00] um, it's just sort of de-- You know, you can have two things that, that solve, solve a similar problem. If they have a different architecture, they're gonna be able to grow in different ways. And, um, and I think that kind of some of these older, these like, these older systems, um, the ways that they can grow is, are reasonably constrained, and they move very slowly.
[00:08:20] Jordan Tigani: I mean, it's just... I remember some of the optimizations that we were work-worked on in, uh, in BigQuery would just take months and months for something relatively simple. Um, and, you know, working on this single-node system, you know, we can just-- Those things can be done in, know, in a week or weekend. Um, and you can just move much, much faster.
[00:08:43] Jordan Tigani: And so because kind of the system has been shaped a little bit differently, the way it can expand and extend and react to new, to new changes is, uh, is, is different. And I think, you know, uh, you mentioned sort of the, you know, all the changes that are going on with AI. Like, I [00:09:00] think like if you have these sort of, you know, giant warehouses, um, or distributed warehouses, they've just from a perspective, they feel less, you know, the right, the, uh, the right, the right shape for like, for, for, for AI and agents, where you might have a lot of agents that are sort of hammering on different things and like, and having a single-node system where you basically, you know, one, one warehouse, one agent, um, you know, for as many agents as you, as you have, like those...
[00:09:37] Jordan Tigani: That seems to sort of fit the problem better than, um, than some of these, these older, older systems.
[00:09:44] Andrew Zigler: Yeah, absolutely. I feel like we're trying to make all of the bits that an agent touches as atomic as possible so that we can isolate them and put them in these, like, more sandbox things. The idea of having agents at scale and at real time all touching a same kind of data [00:10:00] space is, I think, scary for most people who would want to protect their data.
[00:10:03] Andrew Zigler: So, um, I do think that becomes, like, even, like, a strategic wedge in, like, how the tool is adopted and who its consumers are. And I wanna talk a little bit about who those new, like, users are because, you know, it are, it is the agents, but it is also, um, the still, the, the engineer, the data scientist, the, the data engineer, uh, you know, figuring out how to orchestrate them.
[00:10:25] Andrew Zigler: And that's wh- where Motherduck, I think, is really leading in terms of teaching the future of how the data scientists and the data engineers are going to work with their tools. Like, we really loved, uh, your, your entire coverage on Watertown. We've talked a lot about Gastown here on the show, and we've covered it since the top of the year, all of the many things from Steve Yegge and how all of those different narratives have evolved.
[00:10:48] Andrew Zigler: We've actually adopted a lot of them here on our engineering team and on the show. And so, like, when to see the same parallel apply towards, uh, data engineers and, like, the ways that th- they work with their data and how it will [00:11:00] transform, I think was really powerful and really smart too because it gave, uh, it gave us language to talk about the levels of AI fluency as you got better at using those tools.
[00:11:09] Andrew Zigler: I wanted to dive into your head, again, no pun intended, about that and learn a little bit about, um, you know, how long it took to develop that idea after thinking about, you know, what Gastown was doing for engineering and, like, what, how, uh, how innately did all of that come to you?
[00:11:26] Jordan Tigani: Um, so I, I think that we're in, you know, we're in this like transitional moment right now. Like I think, um, I had a college professor in, in, um, you know, who, who like was the guy who invented punctu- or named punctuated equilibrium, you know, in, in evolution. Um, and, and uh, and I feel like sys- software technology, like you often have something similar.
[00:11:53] Jordan Tigani: W- meaning like sort of at this equilibrium state and then some change will happen, and then [00:12:00] there's sort of rapid transition to a new equilibrium. And, um, and so I think we had the modern data stack was this sort of equilibrium state where you had like, okay, you have these, these set of tools you ingest, transform, analyze, visualize, uh, and those are each different, different companies and everybody had their swim lane, swim lanes and everybody was happy. And then along comes like and agents and all of a sudden like That, you know, world is gonna change. And so trying to predict exactly the way that it's going to change is sort of, um, it's a recipe for, you know, spectacularly bad predictions, but that doesn't mean, you know, you shouldn't, you shouldn't try. so I, you know, it's also when it's sort of super exciting 'cause that's when the biggest, the biggest opportunities happen when you are in these, these sort of transitional, uh, transitional times. Um, so we c- you know, started to notice that like, hey, text to SQL [00:13:00] getting it-- you know, it works. It's not necessarily the way we had thought it was gonna work, you know. Uh, but, you know, with an agent, you basically can get very high, um, high fidelity. So, so it's like, okay, that's... Okay, so that's, that's happening, and then you realize, okay, this, um, Claude and the agents can, can really do very good visualizations. Uh, and they're on- those are only gonna get, only gonna get better.
[00:13:22] Jordan Tigani: So, like, um, you know, if you can run a query and then you can visualize it, like that's a lot of what you have in a BI tool. So it's like, okay, well, BI, really of an opportunity to just sort of like change how you're doing, how you're doing BI, On the other side of things where, you know, you, you know, you're bringing data in, you're transforming the data, that also seems highly vibe codable.
[00:13:47] Jordan Tigani: Like it's, it's, you know, if you ask Claude, if you point Claude at Salesforce and say like, "Hey, you know, pull my data in," um, uh, it's gonna be able to figure out how to do that.
[00:13:58] Andrew Zigler: Mm-hmm.
[00:13:58] Jordan Tigani: um, and [00:14:00] it's going to get better. It's going to, you know, like the, the, the... You know, how long it takes it to do today, um, it's, you know, less than it'll take in, in the future. so you kind of think about the data stack and what people are doing, and what data engineers do, and what data analysts do, and sort of like, okay, how does that job change in, you know, when building pipelines can be done from a prompt? Building, you know, visualizations can be done from a prompt. Uh, you know, transformations can be done from a prompt.
[00:14:28] Jordan Tigani: Well, what-- but once you do that, what are the things that are gonna go wrong? And, you know, I think, you know, while people are... You know, the agent's gonna build a bad, bad schema, or they're not gonna have the right context, or they're not gonna have the right... You know, then there becomes like this, um, these other jobs that, you know, humans are going to have to do to, um...
[00:14:49] Jordan Tigani: But they're almost like, know what I mean? Just like, uh, you know, they say software engineers become managers of agents. Like, I think a data engineer is really gonna be sort of a [00:15:00] manager of a, fleet of agents that are gonna be doing, you know, doing data, data work. And really the things that they have to do are gonna be the, you know, reacting to, to changes.
[00:15:11] Jordan Tigani: And I think one of the ways that data engineering is different than software engineering is that it's so much easier for things to break, um Uh, through no fault of your own. It's like, because you have this data that's coming in, data may be changing, um, you know, there's schema changes or, you know, something may be stalled upstream or there may be a bug upstream or, um, uh, the, the, you know, the format or the, the distribution of the data changes and that, you know, changes, you know, the, how you have to visualize the da- There's just all these things that can, that can sort of go wrong with...
[00:15:49] Jordan Tigani: Even if you've built the right, the right pipeline, that I think humans are gonna have to, you know, somebody's gonna have to, um, handle that. And then you think, well, okay, well [00:16:00] But it-- can some of that be done automatically? It's like, hey, if, if the schema changed from a, you know, big int to a float, you know, can, can Claude, you know, go and chase that through?
[00:16:11] Jordan Tigani: Or if a, a field gets renamed, can Claude chase that through all of the, all of your visualizations? And chances are probably yes. so then you, you think about, okay, what do those agents look like? And anyway, so that was sort of the idea behind, you know, I, I called it Watertown after, after, you know, Steve Yegge's, you know, Ga- you know, Gastown.
[00:16:29] Jordan Tigani: It's a little bit, a little bit tongue in cheek. and so exactly what are those agents gonna be and what are they gonna do and what are the jobs gonna be? Who knows? I mean, just sort of like I think when, you know, in Gastown, you know, is do you need the, the deacon and the, like, all of these specific roles?
[00:16:46] Jordan Tigani: Like, you know, maybe not, you know, maybe, you know, but I think you're gonna need some, some way to sort of handle, okay, there's a bunch of tasks to do. There's a queue. There's, there's like, there's somebody you know, and then there's like a [00:17:00] mechanism to surface things to, um, to, to humans as well. And then, um, it's a- again, exciting, exciting times.
[00:17:09] Jordan Tigani: Stuff is changing, stuff is moving, stuff is moving quickly. But I think we're already starting to see some of these stuff, some of these things, you know, some of these things happen, uh, and some things, some things changed. Like, um, can't remember the last time I've written SQL, and I used to write the most SQL of anybody at the company just 'cause I, I, I love to poke at stuff and to sort of be like, "Okay, well, this is happening, but, like, what's actually happening?"
[00:17:34] Jordan Tigani: And, you know, and but now it's just a, it's a, it's a, you know, Claude session and,
[00:17:40] Andrew Zigler: I know
[00:17:41] Jordan Tigani: uh, and so I think that's gonna be happening through the, you know, through the rest of the data, data stack 'cause it's, um... And it's exciting, um, and it's not just-- I think it's not just, like, gonna people less, know, useful.
[00:17:55] Jordan Tigani: Like, I think you're still gonna need people. It, but it also brings, [00:18:00] you know, we call them at, at MotherDuck, we call them NTDs for non-technical ducks, which is like the people who would otherwise feel like they can't do this, this stuff. They're like, "I'm not smart enough. I'm not good enough. I don't have the knowledge or the background to be able to, to do these things, to be able to, you know, ask, answer these questions about what's going on in the business or their sales book or their, you know, their marketing campaigns." And now all of a sudden they can, they can do those and, uh, and I think
[00:18:26] Andrew Zigler: Right
[00:18:27] Jordan Tigani: that's really powerful and liberating.
[00:18:30] Andrew Zigler: Yeah, exactly. It's, it goes again to like the whole democratizing of the tool. Now all they need to d- do is know what to ask, which is its own challenge, by the way, understanding what you're looking for and what the query needs to be, what you're really asking for. But that becomes the new barrier, right?
[00:18:45] Andrew Zigler: Which is exciting because, um, text to SQL, like you said, just really transformed. It's gotten amazing, and it's gotten a lot better. Like I remember back in October of last year, I was at, um, I was at a hackathon. It was a data hackathon [00:19:00] actually for data agents. Uh, Brian Bischoff of Theory Ventures hosted this really wacky hackathon called America's Next Top Modeler, where we had to, uh, basically decide or figure out how we wanted to sort over like 10,000 unsorted documents and Parquet files and all sorts of just like unstructured and unstructured data for a fictional company that had had several mergers.
[00:19:24] Andrew Zigler: And then like adding cruelty to it, he even had like a dusty binder with like old things in it that contr- contradicted what was in our digital files, and we'd have to consult it in like the real world. So it was like mind-boggling. And as part of that, I was struggling with, I remember back then with, with Claude trying to get some basic interactions with SQL where I was comfortable with, where I was like able to actually understand what was going on in the data.
[00:19:47] Andrew Zigler: And now, um, I just feel like whenever I do work with, um, you know, mo- modern models to work with data, uh, it's really, really simple and conversational. Like a, a personal story for me is like [00:20:00] I, um, whenever I work out, I use an app called Heavy, and it pushes a web hook for after my workout of all of the different things of what I did.
[00:20:07] Andrew Zigler: And I actually have that hit a doc database where, uh, it might have an agent that then just looks over it and helps give me like a daily understanding of my workout and how I'm trending. And it even gives me nudges of like, "Hey, you haven't like increased that weight in a while. I see you're like you know, holding back or something.
[00:20:24] Andrew Zigler: And I've found that really fun and interesting to experiment with the data, but that's just something that's become really accessible to me as like an, I guess, a non-technical duck or a non-data engineer duck. You know, I'm like technical, but I'm not of the data science world. But before I would've never really even thought about trying to do that.
[00:20:42] Andrew Zigler: Um, but as agents get better at making, uh, uh, those SQL queries, like all of a sudden now you just have a rise in these huge or new or ad hoc SQL queries. You get the needs for having these like pipelines that can run it and create it. And, you know, I understand that y'all are tackling [00:21:00] that problem too, uh, with Flights, and that's what pipelines are, where agents build and they schedule and deploy things because now like you, they can build up that big SQL query and push it, right?
[00:21:10] Andrew Zigler: Um, so like you're understanding like the platform that they need to do their distributed work on. I have a question for you though, in that world where you have agents deciding what SQL to run and then to put it up into a pipeline to run it, where do the new gates and the checks fall? Because agents can make so many queries or even mutate the data, and so like in engineering and in the LinearB world where we talk about like the bottlenecks and the places where the humans put down the gates, that's like code review.
[00:21:40] Andrew Zigler: That's like a PR, right? But like there's no PR for like a data, or there's not one that's like immediately easy to visualize. So like what is that, what is that like on your side?
[00:21:49] Jordan Tigani: Um, uh, I, I think it, I think at some, to some extent these things have to be code. These have to be like, and treated like code. I mean, obviously they're code because they, you know, they run, they run stuff. [00:22:00] But like, um, but they have to be treated like code and they have to be, you know, I think there's gonna be, you know, review processes and, know, check, you know, people can, um, you know, submit, submit PRs against, against them.
[00:22:14] Jordan Tigani: Like, um, you know, w- whether it's going to be, you know, uh, AI, you know, agents or, or humans r- you know, writing the, writing the PRs or, um, or reviewing the P- PRs, like, um, you know, I think that's, uh, certainly, you know, an, an o- an open question and things, and things are gonna be, you know, and that part, that part is gonna be changing. you know, I think it, you know, you mentioned it, you know, that we, we built this thing called Flights at MotherDuck. And so what Flights are is like, know, really they're just a Python script, and we'll be adding Node as well.
[00:22:51] Jordan Tigani: But, um, you know, TypeScript, um, that runs on a schedule. Uh, and, you know, we have some like hints around [00:23:00] here's how to, here's how to do good things with, with MotherDuck, here's how to, here's how to connect to certain data sources. Um, you know, these sort of set of, set of guides that make that easy. Maybe a set of templates that, you know, um, that the, uh, you know, basically the AI will be able to the template or, or be able to write these, write these scripts.
[00:23:19] Jordan Tigani: But at the end of the day, it's, uh, it's Python scripts. And, you know, kind of I mentioned, you know, when we saw that, um, you know, a lot of the data pipelines and data ingestions is sort of heavily, highly, highly vibe codable. We said, "Well, how do we..." Our goal really is to make it easy for people to get data in because, you know, people would start using MotherDuck and they're like...
[00:23:41] Jordan Tigani: And we say to them, "Okay. Well, you know, what you need to do is you need to go and, uh, sign up for Fivetran and, you know, set MotherDuck as your destination and pull in a bunch of data, and then you can start using MotherDuck." And it's like, you know, um, uh, that's a lot of steps. Uh, that's a lot of things you have to, you have to [00:24:00] get right.
[00:24:00] Jordan Tigani: And, and so one of the things this lets us do is we, you know, just lets you sort of start from In medias res and everybody's watching the, the Odyssey, so it starts in the middle of things. Middle of things is, "Hey, I have a thing I wanna do with my data." And then okay, well, then we're gonna go back and we're gonna, we're gonna pull-- we're gonna fi-figure out where that data is, and we're gonna figure out how to, how to get it in. Um, I just did that this morning, actually. I was, um... I'm working on, like, a new intensity metric and, you know, with, with Claw trying to figure it out, and, like, it's quite a complicated, um, query. And so in order to get, you know, rather than having the, the visualizations compute that, you know, I wanted to sort of, precompute for, for our historical data. And so, you know, it basically kicked off one of these flights and creates one of these flights, and then did a backfill, and then also it's gonna be running every, every, every hour, so they'll be able to sort of keep that, keep that up to date. But I didn't start out to sort of build a data pipeline. The data pipeline was sort of pulled in by, okay, well, there's this thing that I need to do, this like-- [00:25:00] or maybe this data that's missing or this, this, like, this computation I, I, uh, I need to do that, you know, uses some external data or is, uh, is, uh, is, is highly intensive. And I think that tends to be how people, how people work is that, you know, the, the, the, the end goal isn't the data pipeline. The data pipeline is sort of this necessary thing in order to achieve this other, this other goal. Um, you know, I think getting back to the, you know, getting back to the gates, I do think that, you know, so the visual-- o-on the visualization side, we have this, you know You know, our, our visualization tool is called Dives, which is, which is similar.
[00:25:34] Jordan Tigani: It's basically the, you know, it's a TypeScript, uh, thing that can be created by, by, uh, you know, your favorite LLM. You know, I, I use Claude as sort of the generic LLM, uh, moniker.
[00:25:47] Andrew Zigler: too.
[00:25:49] Jordan Tigani: Um, so you use Claude, create this TypeScript file, um, which is a data visualization, and Claude is really good at it, as, you know, as is Gemini and, you know, and, um, ChatGPT, et cetera. [00:26:00] and, uh, and so you kinda have these, like, two things. You have this, this, these dives, which are TypeScript visualizations. You have these, um, flights, which are Python, uh, you know, Python scripts, you know. But those are both code. Those are both, you know, things that, like, you know, the way we use them internally is we have a GitHu- GitHub repo, and we, um, send PRs against the GitHub repo, and then we basically a commit hook that, like, that synchronizes that, that to, um, to Motherduck. and, and so, you know, for us, GitHub is the wor- is the, is the, um, is the source of truth. Um, but then, you know, that makes sure that it, it, it, you know, we can have, we can do reviews, we can-- anybody can see what, you know, what they are. We can clone them, we can tweak them. but then also, know, those get pushed into the tool so that it's, you know, easy to use from the, from the, uh, from the tool.
[00:26:57] Jordan Tigani: And I, I, I feel like that's a good [00:27:00] of stable way of, of sort of making that stuff work, 'cause it, it's all code, and, uh, you know, treated like code and, um, and you-- so you can use the same mechanisms that you would use for anything, which are, you know, also being geared towards, you know, agents and, and, and AI and, and those things are getting better. I think just the l- last thing I'll say on this s- topic, which I, I realize I'm being a little, little long-winded, one of the things that we're trying to do is make it so that these a- as these agents and these models get better, the features also get better. And, uh, and that, which is tricky because often like, you know, people work on something and as the agent gets better or the LLM gets better, it's like, oh, well, I don't even need to do that.
[00:27:44] Jordan Tigani: Like, that whole, like, that whole thing goes away. um, but I think if you can build things in the right shape, then as the models get better, these things just get better. So for example, the visualizations in our dives, like they just get better as Claude gets better [00:28:00] at building visualizations and, and gets better at writing, at writing SQL. And the flights, the same thing, as, you know, it gets better at like, you know, pulling in data from different sources and, you know, building, writing Python. Like those, those just get better. And synchronizing it to GitHub, like, well, you know, the GitHub, the whole like developer tools are only gonna get better at, at working with these, uh, with these agents.
[00:28:19] Jordan Tigani: Um- And it's, it's-- it can be tempting to just like, "Oh, well, the model doesn't do this, and so I'm gonna fill this gap." but, you know, chances are the model is gonna-- is, you know, as it improves, it's, you know, that gap is gonna, is gonna go away pretty quickly
[00:28:36] Andrew Zigler: Yeah, you gotta take stuff away. The shape is gonna change, and you have to just be comfortable reconfiguring it over and over again. It goes back to, like, the beginning of, like, just the paradigm of, like, expectations change. The technology evolves really quickly, faster than you think, and so you have to be comfortable with, you know, changing and evolving that way.
[00:28:53] Andrew Zigler: And there's a lot of smart things in there to unpack, like how ultimately boiling it down to a Git forge kind of source [00:29:00] of truth that's shared and everyone can have visualizations on this also allows you to take advantage of those same kinds of gates and checks like AI review on the types of queries and things that are getting shared out with folks, and you get, you get the collaboration, right?
[00:29:13] Andrew Zigler: And, and that's a really important part of creating and sharing the data too because I think the temptation in now is, like, with all of these capabilities, and Flights is, like, certainly the right shape because it's, it's, it's compatible with the atomic idea of the small and the versatile and the sandbox, right?
[00:29:31] Andrew Zigler: But it also allows for the sudden scaling of the thing you need. Like when you said you didn't set out to make a pipeline, you just got a pipeline for free just by nature of how the tooling is now shaped. And so, um, that becomes, like, the new opportunities to look for, and sometimes that does mean simplifying.
[00:29:47] Andrew Zigler: Another part of this too that I, uh, that keeps bouncing around in my head is when you have a lot of people with a lot of data needs and abilities, and you have a lot of agents that are doing these things at scale, it [00:30:00] becomes really easy for people to just reinvent things over and over again, for collaboration or communication to break down, for people to reinvent the same stuff.
[00:30:08] Andrew Zigler: And that's not economical for a company if they have 10 really agentic engineers and they're all reinventing the 10 same things. So, like, what do you think has to evolve to, um, help those organizations actually distribute the gains of everyone being able to work that way? Um, what has to, what has to evolve?
[00:30:27] Jordan Tigani: Um, you know, I think that's a good question. There's sort of, you know, the one way of thinking about it is, is a content management problem. Like, how do you make sure that if something-- if somebody has done something that becomes visible to other people so that they don't, you know, they can either, they can either use it directly, they can, they can, or they can riff off of it versus like, versus generating it directly.
[00:30:51] Jordan Tigani: And so we are building some, you know, bui- building that stuff into, uh, into MotherDuck. Then there's the other, the o- whole idea of, um, of context [00:31:00] and, um, you know, context meaning, you know, in the, you know, in the data world, like everybody's got a definition, different definition for how they compute revenue.
[00:31:09] Jordan Tigani: Like different companies have, you know, like different ways of, of do-- you know, they have different, um, you know, names for their regions or like they, you know, there's just a bunch of, there's a bunch of sort of business logic that is, you know, tends to be in people's heads or tends to be in a doc somewhere. um, I think one of the reasons that people have been skeptical of, uh, of sort of some of the and data is because they're like, "Well, how are they gonna, how is the AI gonna learn about, you know, this stuff that only I know about my data?" And, um, you know, it turns out it can actually start poking around and looking at things and like it can figure out like, you know, is this a, is this milliseconds or seconds, um, uh, just by looking at it and it's like, is this reasonable?
[00:31:52] Jordan Tigani: The same way a human would, would think that it's reasonable. Um, but I think there's also a, a need for, you know, a, [00:32:00] a context layer, semantic layer, you know, something that, uh, teaches, teaches the LLM about your business, about the specific schemas and, you know, things that you, things that you care about. Sometimes, sometimes it isn't necessary, but it's, um, uh, it's a way of creating shortcuts. You know, one thing
[00:32:22] Andrew Zigler: Mm-hmm.
[00:32:22] Jordan Tigani: find is that like, you know, tokens are expensive and so while the, you know, maybe an, maybe an LLM can, can figure out, you know, what your, um, uh, uh, fiscal quarter is, you know, like, uh, if it doesn't have to figure that out, if you just tell it, like it's, it's a way of, of getting, you know, getting to the answers you want much more quickly and less, and less expensively. so we, we, we just launched actually today, yesterday, um, we call it Guides, which is- All they are is sort of markdown, um, [00:33:00] documents that you can, that you can create, that describe, you know, aspects of, of, uh, of, you know, your, your data, your database, your organization, your schemas. they can be... They can even be like style guides.
[00:33:15] Jordan Tigani: Like, they just sort of describe how this is-- these are, uh, kind of like skills for, for
[00:33:21] Andrew Zigler: Yeah
[00:33:23] Jordan Tigani: Um, and it's pretty s- it's pretty simple. It took us a long time to, long time to build 'cause we, we, we were starting out with something much more complicated. we wanted it to-- Like, my belief was it should be sort of self-driving.
[00:33:34] Jordan Tigani: You should be able to-- We should be able to glean this from, from what individual users are doing, and then be able to combine them across users and be able to sort of like, And that turns out to be, you know, uh, very hard to, to do well. Um, and turns out what people actually just wanted was they, "Oh, I, I just, I know what I want.
[00:33:53] Jordan Tigani: Can I just write it?" And like, and so, like, now we're letting you write, you know, write these documents. I think the, the next [00:34:00] step of that is the, uh, is, is sort of like, okay, do you want like... 'Cause I think actually you can get very, very far from just a markdown doc that describes, that describes something.
[00:34:10] Jordan Tigani: Yeah, maybe you have little SQL snippets in it that describe how to do, do some, some computation. But, um, know, LLMs are very, very good at understanding English, um, and/or understanding whatever, you know, other human natural languages. Um, so I think that they're gonna continue to get better at, you know... they're gonna know that better than they are gonna know, like some, you know, metric flow or some, some semantic modeling language. and, uh, and the, the additional rigor involved just makes it harder to write. It doesn't actually, um, make it, make the, uh, LLM do, do a better, a better job. is the, the argument that I've heard made, um, that, well, the reason you need modeling, semantic modeling meaning like [00:35:00] actually something that en- if, if effectively enforces every time we do this calculation, we do this calculation exactly the same way.
[00:35:08] Andrew Zigler: Right
[00:35:09] Jordan Tigani: and I think that those, uh I don't know whether those are really going to be needed. Um, I think as the models get better, that's gonna be less important. If you describe it in English or like, um, you know, or in some, some way of, of being, being clear about it. Um, you know, I, I, I may be wrong, but I think the other, the other way of doing this is essentially the, um, you know, sort of trust but verify, where you describe it and then you run-- you have evals.
[00:35:39] Jordan Tigani: 'Cause I, you know, I think evals are, are important in the, um, in this world of, of agents and AI where you say like, well, when I say, you know, "Tell me about the revenue in February, you know, 2025," the number should be [00:36:00] this. Like, there is a right answer. And, um, and you can detect whether the, the LLM or your agents can, you know, are generating, generating that right number. um, you know, sort of you can use that versus having this sort of more fancy semantic, semantic model. But that's, that's my belief. Um, um, but I'm, I've been wrong about a lot of these things before, and we'll see, you know, we'll see what happens with, with that one
[00:36:27] Andrew Zigler: Well, trust, trust and verify is definitely the, the, the dual approach that's really important. I actually gave a talk called that recently at the Checkmarx AppSec Summit about that same as I think for sec- like, uh, engineers that are, uh, using, uh, agents to produce code of like you can, uh, trust and understand you can have these guardrails, but then you also need to have these verification systems in place to, to look at the other end.
[00:36:51] Andrew Zigler: It's about like measuring and understanding the inputs, but then also measuring and understanding those outputs on the other end as well. And I think that's been a big [00:37:00] challenge for engineering teams. It's just like a, a... In the last year, there's been like a big mandate, just use AI as much as possible and just like, you know, use your tokens, pick up tools.
[00:37:08] Andrew Zigler: We'll try every tool on the market. And then there's been a lot of shifts recently with people trying to pull back on their inference budgets and, and companies burning through all of their tokens that are available to them just in the first few months of the year. And it speaks to an, an inability to pick the right tool for the right problem.
[00:37:25] Andrew Zigler: Engineers maybe are, uh, you know, they always want to use the best for everything. And so there's still a challenge ahead of us as like engineering leaders of optimizing in those costs and reducing those kind of like duplicated work, and I think that's gonna be a big challenge. And it sounds like guides are like one step for kind of g- uh, putting that, those kinds of roadmaps in a shared place where agents and humans can understand them.
[00:37:49] Andrew Zigler: There's also things like, uh, understanding, like just to get a d- a, a number instead of having to crunch it every time and us having to really get the agents into that [00:38:00] kind of, uh, motion because they're, they are just, you know, apt to crunch. They love to use their tools. And so, um, uh, uh, there's another part of this too about like in that world where, uh, people are just using a lot more tokens.
[00:38:14] Andrew Zigler: We've been talking, uh, even about like the token maxing leaderboards and stuff. Like what are the signals that you look for as an engineering leader that tells you that like your AI usage and your AI productivity is actually giving you value to your engineering team or to your data scientists?
[00:38:32] Jordan Tigani: I, I think that's an un-unsolved, unsolved question. Um, you know, the, uh, um, you're right, everybody wants to use the latest model. It's just like if you've got a, you know, Ferrari in your garage, you're gonna wanna, you're gonna wanna drive it. Um, you know, maybe you won't wanna drive it in traffic, but like, you know, you're gonna wanna, you're gonna wanna use it.
[00:38:55] Jordan Tigani: And if you've got, you know, Fable, you know, uh, you know, [00:39:00] a-available to you and it's gonna get things right more than, you know, Flash, like, um, then you're gonna, you're gonna wanna use that instead. And so, uh, so far we have not put any sort of limits on people. I mean, we do have, um, you know, we do have, we do have lim-- you know, per user limits in Claude that we just, we sort of when people run into them, we, we increase them and, um, we just sort of want to have some sort of, you know, uh, visibility into what, what's happening.
[00:39:33] Jordan Tigani: And, um, the goal isn't to sort of find, uh, uh, try to restrict what people are doing. The, you know, the goal is just, you know, it does, it can cost money and we wanna prevent, uh, you know, prevent run-runaway, um,
[00:39:49] Andrew Zigler: you don't want a runaway bill. Like, you-- And also it helps you and the employee have, like, the check-in about, like, the AI usage and what you're doing with it, and it helps you know, like, what are you using the tokens for? We'd love to-- We, you know, [00:40:00] let's talk about that productivity. It's an opportunity even to share it with others.
[00:40:03] Andrew Zigler: I think that's actually a, a, a really strategic idea for distributing, you
[00:40:08] Jordan Tigani: But we
[00:40:08] Andrew Zigler: the abilities
[00:40:09] Jordan Tigani: you know, one of our engineers, like, I mean, I just, you know, said, "Hey, can you bump my Claude code again?" And it was like, and it was, it was $1,000, $1,000 a month. And, uh, you know, um, but it's like, you know, this person is incredibly productive engineer, like one of our top, top people, like him, you know, 10% more productive is certainly worth, um, you know, it worth, uh, $1,000 a month.
[00:40:36] Jordan Tigani: Um, or, you know, I, we bumped it to, to, to something higher. Um, and so that was an easy, that was an easy call to make. You know, at some point it won't be though. At some point, you
[00:40:45] Andrew Zigler: Right
[00:40:45] Jordan Tigani: like, um, I'm, I'm nervous for those, for those days where it's sort of like it starts to get, you know, people are spending, you know, you were talking about Gastown, you know, people are, you know, running dozens of agents and they're [00:41:00] just like, um... I mean, if you're running dozens of agents and you have, you know, if you're running, you know, the, uh, frontier models, then like, yeah, that's, that's probably gonna be more than, more than you, uh, are go- are gonna be able to or want to, or want to spend. I, I, I, I don't know. I don't... When, when that world arrives, I, I, um, I don't know how we're gon- how we're gonna deal with it.
[00:41:25] Jordan Tigani: Like I've heard, uh, I've heard other founders talk about like, oh yeah, we're, you know, like when you sign up, when you, you know, take a new job, like part of your, you know, you're gonna get a, a, uh, a token budget and
[00:41:36] Andrew Zigler: A token budget. Yeah
[00:41:38] Jordan Tigani: and that's gonna decide whether or not you're gonna take the j- take the job because like this, you know, it's like this, this company lets me drive a Ferrari to work, and this, this one I have to drive a, you know, a, Kia. Um, you know, they're, those are different, different experiences and, um, uh... But yeah, that, um, know, they, we're, uh, [00:42:00] transition- transitional s- states are, are, are exciting 'cause, and, and that we're certainly in one for now
[00:42:05] Andrew Zigler: Do you think this is one of those where it'll be like a quick equilibrium, like where it gets up to that new equilibrium, like going back to what you said earlier?
[00:42:12] Jordan Tigani: I think, I mean, there will-- there certainly will be an equilibrium about like, okay, this is just the way people do software engineering and like... 'Cause, 'cause right now, like the token budgets and the token usage and the, um, AI usage and agent usage is like, is just sort of rocketing, rocketing up, and we're just sort of figuring out what are the levers to, uh, you know, levers to constrain it, and when does it make sense to do it, and when does it make sense to not do it.
[00:42:36] Jordan Tigani: And, you know, um, and it does impact how many engineers you can hire. It impacts like, um, you know, a whole bunch of things. And, um, so I think we ha- you know, we, we will need some more time to get to the next, the next equilibrium. And, yeah. Well, I think a lot of things are gonna break before, before then, which is, you know
[00:42:57] Andrew Zigler: Absolutely. It's like there's so much that [00:43:00] buckles right now, I think, in modern engineering orgs underneath, like the workflows and honestly, the, the just the raw outputs that are coming out of some engineering orgs. And sometimes when you zoom in and it's one team or one engineer or it's one fl- swarm of agents, and it's something that's cracked the code, and then it, then, and then, uh, the, the leadership and the teams they wanna emulate and replicate this, but then it quickly gets into like a, a huge ballooning amount of cost.
[00:43:28] Andrew Zigler: And then there's a, a huge drive to prove the value. You know, the CFO is in these rooms now, it, for all, for all of these engineering teams because those token budgets are real. They're parts of compensation packages at some of like the biggest tech companies now. And you're so right that modern engineers, like they wanna work where they have the most, uh, what they consider like intellectual commoditization available to them to do their job at scale, because that's, that is how engineering's done now.
[00:43:56] Andrew Zigler: I think it's, uh, still a lot for us to learn, and there's [00:44:00] definitely gonna be more, uh, uh, like involve- like, uh, developments on this scene for sure. Uh, you know, I, I think that like in this conversation, we've talked a lot about how our expectations of how technology is, like shifts really dramatically underneath us, and we have to be challenged to create new versions and reinvent things that we took in- ad-advantage of or took for granted the day before.
[00:44:24] Andrew Zigler: And I think that we're in a world now where we have to manage the outputs of things that are running atomically and at scale and 24/7, and we have to have better guardrails and understanding of like how those things are, are working. But also too, critically, we need a fluency layer of data and, uh, and, and intent between workers and humans and, and their agents and, you know, things like, things like MotherDuck, things like DuckDB, those are, those are part of that language, right?
[00:44:55] Andrew Zigler: They allow the agent and the human to, to work more closely [00:45:00] together and in a way that leaves artifacts that are shareable and can be distributed and everyone can benefit from. So I think that's like the future of how people will work with their, their knowledge, and it's been really cool to, to kinda dive into that, again, no pun intended, with you, Jordan, today.
[00:45:15] Andrew Zigler: But as we wrap up, I wanna know, like where can our listeners go to keep up with you and MotherDuck and, and everything that's going on in your world?
[00:45:23] Jordan Tigani: I think probably the best place is, uh, you know, you know, follow, follow Motherduck on LinkedIn, um, or, or, or me. that's probably where we're most, most active. Um, and we know we do have a Motherduck, Motherduck blog where we talk about, um, we talk about all things Motherduck. We have a Motherduck community Slack for people who are using community Slack. There's also DuckDB Discord if you're, uh, if you're a DuckDB fan, that's also quite, quite an active place
[00:45:48] Andrew Zigler: Amazing. We'll, we'll get all of those links in our show notes. And to you listening, if you've made it this far, then you're obviously a data nerd, and you obviously loved today's conversation, so please give it a like or a comment wherever you're [00:46:00] listening or watching this, and join us on Substack or LinkedIn as well to read the whole newsletter that's accompanying that this episode as it comes out.
[00:46:08] Andrew Zigler: And if you have any thoughts about today's, uh, episode or the things that Jordan and I talked about, come find us on LinkedIn and let us know. You can drop us a comment, and we'd love to continue the conversation with you there. And Jordan, thanks again for coming on the show. It was a pleasure talking with you
[00:46:25] Jordan Tigani: Thank you. This was fun.