Mapping the data lineage and data provenance in the enterprise with Rich Miller

Introduction

This week I’m joined by Rich Miller, the CEO of Telematica.

We start off the conversation talking about the difference between data lineage and data provenance…and why they are critical for the enterprise. Rich discusses what data provenance is and how touching the data makes a difference. Rich further discusses the importance of data pedigree and data genetics. As part of our discussion, we talk about raspberries and how this applies.


Speaker Profiles

Rich Miller Twitter: https://twitter.com/rhm2k

Rich Miller LinkedIn: https://www.linkedin.com/in/richmiller/

Telematica: http://www.telematica.com


Audio File


#CIO #Leadership #Data #Lineage #Provenance #DigitalTransformation #CxOitk


Podcast Transcript

Tim Crawford 0:01
Companies are looking for new ways to transform their business. Technology plays a critical role in this transformation. Speed and innovation in both technology and thinking are key to this shift. Hello, and welcome to the Cxo in the Know podcast, where I take a provocative but pragmatic look at the intersection of business and technology through the lens of leading CxO executives. I’m your host Tim Crawford, a CIO and strategic advisor at Avoa. This week, I’m joined by Rich Miller, the CEO of Telematica. We start off the conversation talking about the difference between data lineage and data provenance, and why they are so critical for the enterprise. Rich discusses what data provenance is and how touching the data makes a difference. Rich further discusses the importance of the data pedigree and data genetics. As part of our discussion, we talk about raspberries and how this applies. Rich Miller, hey, how are you today?

Rich Miller 1:02
I’m fine, Tim. Nice to talk to you again.

Tim Crawford 1:06
Likewise, so Rich Miller, CEO of Telematica, Chairman and co-founder of Provident Data, and today we’re going to talk about data lineage and data provenance. And you know, I’ve heard these terms used together as though they’re one and the same, but then I’ve also heard where they’re separate. So maybe you can kind of help us set the stage, maybe a foundation for the rest of the conversation about what is data lineage and what is data provenance. Are they the same? Are they different?

Rich Miller 1:37
Well, they are definitely different, and part of the problem that you’re dealing with is that data lineage means a lot of different things to different people, and it’s been incorporated very often into some of the formal methodologies for data architecture and enterprise architecture, like TOGA and things like that, data lineage is the term you’ll hear more often, and one of the ways in which it’s defined is as a representation of the path along which data flows from its point of origin to its point of usage, and by the path, I’m really not talking about the geographic path so much as what is the location architecturally of where the data is, what processing is being done to the data at different points along the way, and data lineage is, for the most part, thought about the kind of information that’s used to design and describe the processes of a data transformation and data processing flow. And a matter of fact, very often some of the definitions don’t do a very good job of distinguishing a data flow from data lineage, and this is caused for a lot of confusion. But it’s perhaps the best way to think about it is how you use it, and data lineage is usually the kind of information that you record, and you represent a set of interlinked components, such as data elements, business processes, IT systems, specific applications, data controls, and these components are usually presented at different levels of abstraction and with different levels of detail. But what you want to be able to do with data lineage is actually take your record of the path that data has taken and the processes that have been applied to it, so that just on the basis of the starting point of the dataset and all of the instructions in the data lineage, you could actually recreate the final product after you went through a data flow. So, in other words, it’s like saying I’ll start with the first instance that I’ve seen this data set. I’ll have all the instructions and description of all the functions and all the processing that’s been done to it. And at the end of replaying those, I should have exactly the same result every time. If I don’t, then I’ve got a problem somewhere.

Tim Crawford 4:44
So, in some ways, if I were to try and play that back, I’m essentially able to understand what the data was when it started, and all of the different things that touch the data along the path, and how they touch that data along the path.

Rich Miller 4:59
It’s. What they did to the data along the path-it doesn’t necessarily, as a matter of fact, often doesn’t talk about the who or what or with what responsibility they touched the data. That is a great question because what that then does is starts to talk about what data provenance

Tim Crawford 5:24
is? Okay,

Rich Miller 5:25
data provenance in our parlance here is the record of who or what process had responsibility-that is, ownership or stewardship of the data at some point in time, and particularly at the point in time when some transformation or something was done to the data, so the lineage would be independent of who or who had possession of the data in the course of its path through the process. Lineage just kind of says these are all the things that got it from its origin to its endpoint. Provenance is basically a description of the chain of responsibility, the chain of who or under what conditions somebody had ownership or use of the data and actually applied one of these transformations. So we like to think about it as a combined concept for the tim being that I sometimes describe as the data pedigree, because in a sense, if I have the data lineage, it’s like having the genetic map or the genetic history of a data set. Here are all of the ancestors of my data set, all the transformations that went into changing it from one state to the next. Provenance is like saying, who had possession of it, and under what conditions was the data being held and processed? The term provenance is very often used in the art world when somebody asks, “What’s the provenance of that work of art, in whose hands, who had ownership or control of a work of art from the time the artist completed it and sold it or gave it away to somebody to the current time, and what people often do is use the history or the documentation of provenance to do a pretty good job of authenticating whether a work of art is a real one or has been faked. So there are two different things. One is how we got there, how we got from start to finish, and the provenances who had responsibility for all of that along the way.

Tim Crawford 8:06
When I think about these two, and I have a much better understanding of the differences between the two, and in full transparency to the audience, I mean, you and I have had conversations about these topics for a number of years now, and the example that you and I have used were raspberries. We talked about raspberries, and so you know, as I think about specific industries that might be going, okay, so I get what data lineage is, I get what data provenance is, but I’m in the business of raspberries. Why are these two things important? Use the art example, but maybe take that raspberry example and let’s pick it apart. Or if there’s another example you’d like to use, let’s pick it apart and help people understand why these are important in a practical way.

Rich Miller 8:54
Sure. Well, if we’re talking about a physical object like a pallet of raspberries, you want to know where that pallet went, and all the things, and all of the temperature situations under which it was kept. Whether it was doused with anything like a disinfectant or an antibacterial or an insecticide. Those are the kinds of things that have actually been done to the raspberries from the point of origin to the point where it’s a in a little green plastic container in the supermarket. The provenance. What I want to know is who had responsibility for that basket of raspberries at a certain point, perhaps because I’m worried about contamination, or there’s been some report of a problem with

Tim Crawford 9:55
food safety or

Rich Miller 9:57
food safety.

Tim Crawford 9:58
Sure,

Rich Miller 9:58
and what I need to. Not is just where it’s been or what’s been done to it. I also want to know who had responsibility for

Tim Crawford 10:10
it,

Rich Miller 10:11
because to a great measure, what we’re looking to do is find out and establish the chain of custody for that pallet of raspberries and the chain of responsibility or stewardship for that basket of raspberries. Now those are two different things. I can find out when it happened, where it happened, perhaps be able to determine the specific source of some sort of contamination. But the other thing I want to be able to do is be able to go back and say who really had responsibility for it. Now, rather than a pallet of raspberries, let’s talk about an actual data set. When I am in an enterprise and I want that data that I’ve purchased, perhaps from a Thomson Reuters or a Bloomberg or a marketing dataset from Axiom or whatever. I want to know how that data was gathered. I want to know what was done to it to anonymize it or process it in an appropriate fashion, and I also want to know: Did anyone in that chain, in the lineage, also follow or not follow the licensing requirements under which they had the data? Because at the end of the day, if I am using it for my business, and somehow I have use of data that wasn’t properly licensed, or I’m using data in some sort of an analysis and decisions that I’m making that don’t fall under the licensing terms for which the data was released. I can be liable. I can have compliance problems, and I don’t want that. So,

Tim Crawford 12:06
but let me kind of jump in here because you’re talking about how companies have used or relied on these two aspects of data in the past a bit. But what’s different today than say 10 years ago or even 20 years ago? I mean, we’ve always had data. We’ve always had supply chains. We’ve always had customer data that we’ve had to contend with. What’s different about today?

Rich Miller 12:29
Well, I think there are a number of things that we could point to. I mean, one of the things that we absolutely all have to recognize is that we are so much more driven by data, our processes are much more sensitive to and in almost in real time gathering data in order to be optimally operated. So we rely on a lot more data. We rely on it for everything from the just-in-time supply chain to sentiment analysis from a Twitter feed or things like this. So, our consumption of data for the purposes of running our businesses is much more pronounced today than it has ever been, and by really some enormous factors. That’s one aspect. There’s another one, and what’s different is also there’s a lot more attention being given to privacy preservation and security. The way in which data is treated and the responsibility that every steward of data has to mind the rules, be compliant with regulation. This is the basis on which we now have the GDPR in Europe or CCPA in California, the whole issue of managing and stating unequivocally the responsibility of whoever holds the data to the persons about whom that data is a representation, the consumer.

Tim Crawford 14:20
Okay,

Rich Miller 14:20
we’re both heavier consumers, and along with all that heavy consumption, we’re also under a lot more of the kind of attention and constraints and regulation to protect the sources of that information to begin with.

Tim Crawford 14:40
And I want to get into the regulatory compliance and privacy aspects in a second. One of the pieces that I’m curious about: what’s changed now versus 10 years ago or 20 years ago? Where does technology fit into this? Is technology a factor in why we’re having this conversation? Much more than we did maybe 10 or 20 years ago.

Rich Miller 15:04
Well, I think that’s pretty clear that it has. I mean, technology that we’re using 20 years ago, the architectures that we utilized for data were interesting, and we thought of them as being quite liberating, but by comparison, today they’re a little bit on the pedestrian side. We had data marts, we had BI, we had specific data being gathered and structured and placed in data markets, data marts of very, very rigid, structured form, so that we could answer specific questions, generate specific reports, and do so very cost effectively, very efficiently, with the increased use of streaming data with the increased use of and dependence on unstructured data to extract information by which we run our businesses. We’re now looking at a much broader range of architectural infrastructure. that we use for the storage or containment of our data. So we’ve got data lakes, we’ve got buckets of object-oriented data sitting in the cloud, not just in our data centers. We’re moving parts of those data sets or those data collections, and moving them around to different processes for different kinds of actionable information. So,

Tim Crawford 16:53
yeah,

Rich Miller 16:53
if you want to talk about the range of things that we do with data and places we keep it, and the form with which we keep it today, I think we’re by comparison to what we were doing 20 years ago, or even 10 years ago, there’s almost no comparison.

Tim Crawford 17:13
So, if I think about that data pedigree, and again from an enterprise perspective, I’m looking at the data pedigree, which is both the lineage and the providence. What happens if I continue, or I’m making that assumption? Maybe I shouldn’t. If I’m ignoring that, what happens? And then maybe you could weave in your thoughts on how regulatory compliance and privacy kind of tie into that.

Rich Miller 17:41
Happy to. Well, first of all, I think one has to think about data as not exactly a source of energy or fuel for today’s enterprise. I know that people like to bandy or this idea that data is the new oil. I don’t think I quite buy into that, but let’s put it this way: the data that you either generate in house or that you obtain from an outside source is a classic ingredient to the business, and just the way you would purchase a physical set of resources-a boxcar of nuts and bolts, where you need to know what the different sizes are, what the materials are that have been used in the construction of those nuts and bolts, what the test results for stress is on them. You really need kind of a data sheet, a description that you would like to have about the data set that you’re about to consume and use as the basis for your decision making. If you want to think about it. This would be like buying a raw resource or buying livestock. If you’re a cattle rancher, you really do need to know what is the pedigree, what is the backstory for that raw resource that I am utilizing for my business, if I don’t know that, I could mistakenly assume that it was one thing, not the other. I could assume that the cattle I just bought had been vaccinated, but they hadn’t been. Sure. Or that basket of strawberries had been kept at the appropriate temperature all the way through the supply chain. If it hadn’t, then I need to take some sort of different action when I take possession of it. So you have to take a look at what you’re getting,

Tim Crawford 19:59
but there. Two aspects here that I think I want to get teased apart. Which one is there’s a value to your enterprise, a value to the product or service that you are offering. There’s a value to that data and the accuracy of the data. But then separately, there’s a risk to mishandling it or misusing it, both because it starts to erode that value, but then also from a regulatory compliance and privacy requirement perspective. And I, I want to briefly have you share your thoughts on where regulatory compliance and privacy kind of hits in here, and then we’ll go on.

Rich Miller 20:39
Well, let’s just think about it this way: there are all kinds of technical solutions that we now have available to us for tracking and maintaining the pedigree. And to the degree that that’s available to us, that also means anybody that is doing compliance and auditing us also is going to expect us to have more information and more details about the quality and the validity of the data that we’re using as the basis for our businesses, and they should. But along with that, there have to be ways in which, if there is a problem, you want that same information of lineage and provenance to help you, the enterprise or some regulatory organization, to go back in time to turn back the clock, if you will, and go back upstream to find the source of a problem, a root cause for bad data. So what you’re really doing is using this as a way, not just of protecting someone’s rights or looking at data security, but you’re really trying to continuously improve the quality of the data that you’re consuming, you’re using, and quality comes in many dimensions that we could talk about.

Tim Crawford 22:15
We could probably take quality and make it its own episode.

Rich Miller 22:19
Oh, without a doubt, but let’s just for the purposes of our discussion here, let’s talk about two in particular. Was the data that is used in a data set that I’m basing my decisions on did it come legally from its source? Was it licensed appropriately? And did everybody in the chain of custody, chain of responsibility between the source of that data and me follow the rules and adhere to the terms of the license? If they didn’t, I want to know about it because if something gets shown up to be a problem with the data I’m using, and it wasn’t my fault. I was assuming I didn’t test or I didn’t get any attestations from my sources that this was good data or quality data or well licensed and documented. I don’t want to be liable for it. I want to be able to point back to the source of the problem and say, “No, this is the culprit. That’s extremely important. The other is, in today’s GDPR, today we have the possibility that a consumer invokes the right to be forgotten. There are situations in which a consumer can say, “You, sir, have marketing data about me. I have made a legitimate request to the authorities to have that data erased, or at the very least, there has to be a guarantee that it will not be used and it will not be further disseminated. How do you prove that you’ve done that? How do you prove that you have used the most authoritative dataset that is up to date with respect to all of the consumers who have elected to make that choice to be forgotten. Without something that tracks the lineage and also the provenance, you won’t be able to do it, and therefore you’re putting your business at risk for violation of compliance, and no one wants to do that knowingly.

Tim Crawford 24:48
Of course, of course. I mean, that’s something that is all of our responsibility, regardless of where you are in the organization, is good data husbandry. You know, when I think about the enterprise. Though and think about the data management aspects and the fact that data is everywhere. We talked about this on a prior episode talking about dark data and the complexities of that. So definitely go back and if you’re listening to this podcast, go back and listen to that episode to learn a little more about dark data. But who within the enterprise is responsible for insuring this, and and we only have a few minutes left for this episode. But who’s responsible for insuring this, and where do you get started? This seems like just a really complicated problem to tease apart.

Rich Miller 25:35
Yeah. Well, let’s talk about your first question: Who’s responsible, or what happens? And my answer to that is, it depends, and it really depends on the use to which the dataset is being put. If you start at the top, there’s the executive whose decisions and actions are really dependent upon not only having the data that’s fit for its purpose, which you understand by understanding the lineage and the quality, but it’s the data that comes with this pedigree, this assurance that the enterprise has the right to use that data, and that the enterprise is a steward who is compliant with regulation and legal constraints in order to follow the law. So let’s start right at the very top. There is a financial and economic and legal issue that needs to be attended to at the Cxo level.

Tim Crawford 26:40
Okay,

Rich Miller 26:42
for the analyst, that is whether it’s a business analyst or a very skilled data scientist. Good data husbandry really means that they don’t have to spend a lot of their precious time and their valuable time in the unproductive business of cleaning up data sets that are unhygienic. Let’s put it that way.

Tim Crawford 27:07
Sure.

Rich Miller 27:08
So it’s important that the analyst at the level of a group of people who have been asked specific questions by the executives, by management, who are under the gun to follow the processes established by the enterprise, they want to be able to do their job exceptionally well. They don’t want to be spending their time doing cleanup on the raw resource that they have to use as the basis for their job, and that goes down then to the first order technical responsibility for data pedigree, which really is the responsibility of the data engineering personnel. Whether you’re being active in the data ops movement and very agile and avant-garde, or very conventional reliance on BI and data marts or data cubes, what have you. The job of data engineering is to know about the pedigree, about the quality, and to assure the analysts who are going to use it and the executives who are going to use the results of that analysis that we have used the appropriate kinds of data with the necessary quality. So, where do you start? Good question. What you really want to do is start at the beginning, and that is where do you first encounter or capture the data. What is exceptionally important is that when you first encounter the data and onboard it or capture it into some sort of data architecture, that is the first place at which data quality, data validity, and the check on data lineage and provenance needs to be made.

Tim Crawford 29:10
Okay,

Rich Miller 29:10
I know there’s been a lot of talk about data lakes and how you can just throw data into the data lake and

Tim Crawford 29:17
or data swamps.

Rich Miller 29:18
How well somehow magically you’ll be able to find it when you need it. And to your point exactly, too often they do become data swamps, and they become expensive. And you don’t know what you’ve got in the data lake. You don’t know how much replicated data you’ve got. You don’t know which is authoritative and which is out of date. It’s important that you have at the outset that kind of application of data quality checking, things like deduplication.

Tim Crawford 29:55
Okay,

Rich Miller 29:56
and this is done using modern tools. Like data catalogs or metadata management solutions, software.

Tim Crawford 30:05
Very cool, Rich. There’s a lot to unpack here. We could have taken each of these concepts and delved deeper into it, but it’s a critical space for the enterprise to really think about, especially as they think about how data becomes a business driver, a core business driver, and it’s not straightforward, as you’ve articulated just in this episode alone. So, Rich, thanks so much for taking the time today to join me for this episode.

Rich Miller 30:33
It’s been a pleasure, as always, Tim.

Tim Crawford 30:36
For more information on the Cxo in the Know podcast, visit us online@cxoinththeknow.com You can also find us on Apple Podcasts or wherever you listen to your podcasts. Please subscribe and thank you for listening.


Discover more from AVOA

Subscribe to get the latest posts sent to your email.

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from AVOA

Subscribe now to keep reading and get access to the full archive.

Continue reading

Discover more from AVOA

Subscribe now to keep reading and get access to the full archive.

Continue reading