Data catalogs, a category of product in the broad field of data governance, are emerging in popularity.
That popularity has been brought on by the twin enterprise mandates of complying with data regulations and herding the growing number of repositories in the corporate data estate.
But data catalogs are a legacy product category too, originally stemming from simple data dictionaries – essentially table layouts with plain-English descriptions of tables and fields.
Today’s data catalogs have grown in capabilities, importance, and integration with other tools.
In a nutshell, data catalog platforms help organizations inventory their data by documenting data set content, location, and structure; and aligning business and technical metadata.

This organization yields control, and having control helps enterprises:

Achieve compliance with data protection regulations, through documentation and inventory.
End users know where to get data and will avoid duplicating it.
Organizations can control access to entire data sets where necessary and can better enforce role-based access to data subsets within them.
The EU’s General Data Protection Regulation (GDPR) is in effect now, with very strict fines for non-compliance.
The GDPR’s companion ePrivacy Regulation (ePR) may come into effect in 2019, and the California Consumer Protection Act (CCPA) has been passed and goes into effect on January 1, 2020.
These regulations demand the structure and controls that data catalogs provide.

Improve data lake ROI by making data within the lake more discoverable and increasing the lake’s usability in general.
A well-organized, searchable data catalog makes it easy to find relevant data, analyze it, derive insights, and make decisions with greater speed and conviction.
These are the very reasons most enterprises built their data lakes in the first place.

Unify the data landscape by creating a consolidated volume of information covering data (pool) lake, data warehouse, and operational databases.
Implemented correctly, data catalogs integrate these components through a shared abstraction, helping customers derive new value from older warehouse and operational database assets.

Bring data and the business closer together, by mapping business entity definitions onto data sets and columns within them.
A great data catalog provides a business glossary that helps business users find the data they need within the context of their own concepts, taxonomies, and vocabulary.
The summary, then, is that catalogs protect enterprises from regulatory jeopardy and benefit them by delivering more value from existing assets.
Today’s data catalogs enable collaboration between custodians of the data (“data stewards” in contemporary parlance) and business users by mapping out the organization’s data, which makes it more usable for analysis, and thereby benefits the organization.
The vendors discussed in this report all provide baseline functionality (discussed in the Definition section, below) and each has its own emphasis.
Broadly speaking, the products break down into those that are more governance-focused, and those that have a penchant for enabling self-service analysis in the organization by data enhancing data discoverability and usability.
Within those two broad categories are sub-emphases, detailed in the diagram below.

Figure 1: Data Catalogs: Categories and Priorities
Each of the above designations will become clearer through the course of this report.
Source: gigaom.com
Just three weeks into 2019, Veeam announced a $500M funding round.
The company is privately held, profitable, and with a pretty solid revenue stream coming from hundreds of thousands of happy customers.
But, still, they raised $500M!
I didn’t see it coming, but if you look at what is happening in the market, it’s not a surprising move.
Market valuation of companies like Rubrik and Cohesity is off the chart and it is pretty clear that while they are spending boatloads of money to fuel their growth, they are also developing platforms that are well beyond traditional data protection.

Backup is one of the most tedious, yet critical, tasks to be performed in the IT space.
You need to protect your data and save a copy of it in a secure place in case of a system failure, human error, or worse, like in the case of natural disasters and cyberattacks.
But as critical as it is, the differentiations between backup solutions are getting thinner and thinner.
Vendors like Cohesity got it right from the very beginning of their existence.

It is quite difficult, if not impossible, to consolidate all your primary storage systems in a single large repository, but if you concentrate backups on a single platform then you have all of your data in a single logical place.
In the past, backup was all about throughput and capacity with very low CPU, and media devices were designed for few sequential data streams (tapes and deduplication appliances are perfect examples).
Why are companies like Rubrik and Cohesity so different then?
Well, from my point of view they designed an architecture that enables them to do much more with backups than what was possible in the past.
Adding a scale-out file system to this picture was the real game-changer.
Every time you expand the backup infrastructure to store more data, the new nodes also contribute to increasing CPU power and memory capacity.
With all these resources at your disposal, and the data that can be collected through backups and other means, you’ve just built a big data lake … and with all that CPU power available, you are just one step away from transforming it into a very effective big data analytics cluster!
Starting from this background it isn’t difficult to explain the shift that is happening in the market and why everybody is talking more about the broader concept of data management rather than data protection.
Some may argue that it’s wrong to associate data protection with data management and in this particular case the term data management is misleading and inappropriately applied.
But, there is much to be said about it and it could very well become the topic for another post.
Also, I suggest you take a look at the report I recently wrote about unstructured data management to get a better understanding of my point of view.
Now that we have the tool (a big data platform), the next step is to build something useful on top of it, and this is the area where everybody is investing heavily.
Even though Cohesity is leading the pack and has started showing the potential of this type of architecture years ago with its analytics workbench, the race is open and everybody is working on out-of-the-box solutions.
In my opinion, these out-of-the-box solutions, which will be nothing more than customizable big data jobs with a nice and easy-to-use UI on top, will make data management within everyone’s reach in your organization.
This means that data governance, security, and many business roles will benefit from it.
As mentioned earlier, Cohesity is in a leading position at the moment and they have all the features needed to realize this kind of vision, but we are just at the beginning and other vendors are working hard on similar solutions.
Rubrik, which has a similar architecture, has chosen a different path. They’ve recently acquired Datos IO and started offering NoSQL DB data management.
Even though NoSQL is growing steadily in enterprises, this is a niche use case at the moment and I expect that sooner or later Rubrik will add features to manage data they collect from other sources.

Not long ago I spoke highly about Commvault, and Activate is another great example of their change in strategy.
This is a tool that can be a great companion of their backup solution but can also live alone, enabling the end-user to analyze, get insights and take action on data.
They’ve already demonstrated several use cases in fields like compliance, security, e-discovery, and so on.

Getting back to Veeam … I really loved their DataLabs and what it can theoretically do for data management.
Still not at its full potential, this is an orchestrator tool that allows to take backups, create a temporary sandbox, and run applications against them.
It is not fully automated yet, and you have to bring your own application.
If Veeam can make DataLabs ready to use with out-of-the-box applications it will become a very powerful tool for a broad range of use cases, including e-discovery, ransomware protection, index & search, and so on.
These are only a few examples of course, and the list is getting longer by the day.
Data management is now key in several areas.
We’ve already lost the battle against data growth and consolidation, and at this point finding a way to manage data properly is the only way to go.
With ever-larger storage infrastructures under management, and sysadmins that now have to manage petabytes instead of hundreds of terabytes, there is a natural shift towards automation for basic operations and the focus is more on what is really stored in the systems.

Furthermore, with the increasing amount of data, expanding multi-cloud infrastructures, new demanding regulations like GDPR, and ever-evolving business needs, the goal is to maintain control over data no matter where it is stored.
And this is why data management is at the center of every discussion now.
Originally posted on Juku.it
Source: gigaom.com
Many organizations start their Master Data Management (MDM) journey in support of a specific sponsor business need.
Often hard-fought concessions are made to allow MDM techniques to support the data needs of a particular function.
Now, since almost every application needs your customer master, it’s time to expand the scope of MDM, replace substandard methods of accessing customer data in the enterprise, and establish the MDM hub for use by new applications across all customer business processes.
Crossing the chasm to the enterprise customer 360 can be daunting. During this 1-Hour Webinar, GigaOm analyst William McKnight will share some strategies for moving from application to enterprise customer MDM and will discuss the ruggedization that must exist in MDM to be ready.
William McKnight has advised many of the world’s best-known organizations.
His strategies form the information management plan for leading companies in various industries.
He is a prolific author and a popular keynote speaker and trainer.
He has performed dozens of benchmarks on the leading database, streaming, and data integration products in the past year.
William is the lead data analyst for GigaOM and is the #1 global influencer in data warehousing and master data management and he leads McKnight Consulting Group, which has placed on the Inc. 5000 list in 2018 and 2017.
Source: gigaom.com
The pressure to leverage data as a business asset is stronger than ever. Enterprises everywhere are eager to devise sound data strategies that are realistic and achievable, based on available budget, and sensitive to in-house technology skill sets.

For a while, it looked like open source, specialized big data compute frameworks, including Hadoop and Spark, were the way to go.
Enterprise organizations found them compelling for reasons of novelty, economics, and the apparent prudence of a future-looking technology.
But those frameworks are at a bit of a crossroads: the hype around them has subsided and — while things are improving — the success rate of enterprise projects involving them has been modest.
Meanwhile, the data warehouse (DW) which, for decades, has been a key technology platform for enterprise analytics, never went away.
Yes, DWs struggled and incumbent DW platforms still do, but recent advances in storage costs and compute scalability, especially in the cloud, have addressed the most important challenges faced by DW platforms.
As a result, we are in a DW renaissance period.

Problems with DWs have largely been solved, petabyte-scale data volumes no longer defeat them and the familiarity and ease of use that kept them viable all this time are now helping them face their open source big data competition and, in many cases, emerge victorious.
Still, if DW platforms have changed, what should enterprises do to build an analytics strategy that integrates them?
Even the most DW-loyal shops will need to look at how DW platforms have evolved and adjust their strategies accordingly.
Organizations that have committed to open source analytics technologies will need to take a second look at DW platforms and consider a strategy that combines both, essentially bringing the data warehouse together with the data lake.
The trick to adapting to the new world of data warehousing is understanding that it is not only the technology that has changed but the applications and use cases for DW technology as well.

DW platforms do not need to be used exclusively for Enterprise DW implementations.
The platforms are now more versatile and can be used for use-case-specific workloads and even exploratory analytics.
In a sense, the DW isn’t just a DW anymore.
Even the juxtaposition of DW and Data Lake has shifted – DW platforms today are increasingly able to ingest raw, semi-structured data, or query it in place.
This means warehouse and lake technology can be used in combination and, sometimes, lake technology will not be necessary.
It is not just the Data Lake and its workloads that are becoming more integrated into the warehouse.
Streaming data, machine learning, and AI are onboarding as well.
In addition, data governance and data protection are starting to enter the DW orbit.
While the familiarity of the relational model, dimensional design, and SQL are back; the application of DW technology now covers territory that may be less familiar.
Moreover, new vendors who have championed or been born in the cloud are emerging in leadership positions.
The new capabilities, new use cases, and new vendors make the space exciting but also difficult to navigate (for newbies and veterans alike).
Enterprise customers will need to understand how DW platforms have morphed and shape-shifted; how best to use, deploy and implement them; combine them with other applications; and understand the key differences between the vendors and their offerings.
Without this knowledge, Enterprise buyers will be frozen in indecision.

With it, they will be armed to leverage today’s DW platforms to their fullest and cherry-pick other technologies that can augment and optimize them.
The end result is that organizations in the know will be ready to analyze all their data, in technologically familiar environments, for full competitive and operational advantage.
In this report, we map out today’s DW landscape in the context of where the technology has been, where it is, and where it is going.
Source: gigaom.com
The enterprise data protection market still leverages a traditional approach to solving the increasingly complicated enterprise requirements.
Solutions range from on-premises traditional software to specialized appliances to cloud-based solutions.
Some solutions support a hybrid cloud to meet a specific enterprise’s requirements.
Most of the solutions take a traditional approach to backup and recovery while a few are looking to leverage newer data sources and technologies.

Source: gigaom.com
Master Data Management (MDM) is an approach to the management of golden records that have been around for over a decade only to find a growth spurt lately as some organizations are exceeding pain thresholds in the management of common data.

Blockchain has a slightly shorter history, coming aboard with bitcoin, but also is seeing its revolution these days as data gets distributed far and wide and trust has taken center stage in business relationships.
Volumes could be written about each on its own, and given that most organizations still have a way to go with each discipline, that might be appropriate.
However, good ideas wait for no one and today’s idea is MDM on Blockchain.

Thinking back over our MDM implementations over the years, it is easy to see the data distribution network becoming wider.
As a matter of fact, master data distribution is usually the most time-intensive and unwieldy part of an MDM implementation anymore.
The blockchain removes overhead, costs, and unreliability from authenticated peer-to-peer network partner transactions involving data exchange.
It can support one of the big challenges of MDM with governed, bi-directional synchronization of master data between the blockchain and enterprise MDM.

Another core MDM challenge is arriving at the “single version of the truth”.
It’s elusive even with MDM because everyone must tacitly agree to the process used to instantiate the data in the first place.
While many MDM practitioners go to great lengths to utilize the data rules from a data governance process, it is still a process subject to criticism.
The consensus that blockchain can achieve is a governance proxy for that elusive “single version of the truth” by achieving group consensus for trust as well as a full lineage of data.

Blockchain enables the major components and tackles the major challenges in MDM.
Blockchain provides a distributed database, as opposed to a centralized hub, that can store data that is certified, and for perpetuity.
By storing timestamped and linked blocks, the blockchain is unalterable and permanent.
Though not for low latency transactions yet, transactions involving master data, such as financial settlements, are ideal for blockchain and can be sped up by an order of magnitude since blockchain removes the grist in a normal process.
Blockchain uses pre-defined rules that act as gatekeepers of data quality and governs the way in which data is utilized.
Blockchains can be deployed publicly (like bitcoin) or internally (as an implementation of Hyperledger).
There could be a blockchain per subject area (like customer or product) in the implementation.

MDM will begin by utilizing these internal blockchain networks, also known as Distributed Ledger Technology, though utilization of public blockchains is inevitable.
A shared master data ledger beyond company boundaries can, for example, contain common and agreed master data including public company information and common contract clauses with only counterparties able to see the content and destination of the communication.
Hyperledger is quickly becoming the standard for the open-source blockchain.
Hyperledger is hosted by The Linux Foundation. IBM, with the Hyperledger Fabric, is establishing the framework for blockchain in the enterprise.
Supporting master data management with a programmable interface for confidential transactions over a permissioned network is becoming a key inflection point for blockchain and Hyperledger.
Data management is about the right data at the right time and master data is fundamental to great data management, which is why centralized approaches like the discipline of master data management have taken center stage.

MDM can utilize the blockchain for distribution and governance and blockchain can clearly utilize the great master data produced by MDM.
Blockchain data needs data governance like any data. This data actually needs it more given its importance on the network.
MDM and blockchain are going to be intertwined now.
You can get to know the customer better with native fuzzy search and matching in the blockchain.
You can track provenance, ownership, relationship, and lineage of assets, do trade/channel finance, and post-trade reconciliation/settlement.
Blockchain is now a disruption vector for MDM.
MDM vendors need to be at least blockchain-aware today, creating the ability for blockchain integration in the near future, such as what IBM InfoSphere Master Data Management is doing this year.
Others will lose ground.
Source: gigaom.com
Many companies in the corporate world have attempted to set up their first data lake. Maybe they bought a Hadoop distribution, and perhaps they spent significant time, money, and effort connecting their CRM, HR, ERP, and marketing systems to it.
And now that these companies have well-crafted, centralized data repositories, in many cases…they just sit there.

But maybe data lakes fall into disuse because they’re not being looked at for what they are. Most companies see data lakes as auxiliary data warehouses.
And, sure, you can use any number of query technologies against the data in your lake to gain business insights.
But consider that data lakes can – and should – also serve as the foundation for operational, real-time corporate applications that embed AI and predictive analytics.

These two uses of data lakes — for (a) operational applications as well as for (b) insights and predictive analysis — aren’t mutually exclusive, either. With the right architecture, one can dovetail gracefully into the other.
But what database technologies can query and analyze, build machine learning models, and power microservices and applications directly on the data lake?
Join us for this free 1-hour webinar from GigaOm Research. The Webinar features GigaOm analyst Andrew Brust, and Splice Machine CEO and Co-Founder, Monte Zweben.
The discussion will explore how to leverage data lakes as the underpinning of application platforms, driving efficient operations, and predictive analytics that supports real-time decisions.

Register now to join GigaOm Research and Splice Machine for this free expert webinar.

Source: gigaom.com
Once data is under management in its best-fit leverageable platform in an organization, it is as prepared as it can be to serve its many callings. It is in a position to be used for purposes operationally and analytically and across the spectrum of need.
Ideas emerge from business areas no longer encumbered with the burden of managing data, which can be 60% – 70% of the effort to bring the idea to reality. Walls of distrust in data come down and the organization can truly excel with an important barrier to success removed.

An important goal of the information management function in an organization is to get all data under management by this definition and to keep it under management as systems come and go over time.
Master Data Management (MDM) is one of these key leverageable platforms. It is an elegant place for data with widespread use in the organization. It becomes the system of record for the customer, product, store, material, reference, and all other non-transactional data.
MDM data can be accessed directly from the hub or, more commonly, mapped and distributed widely throughout the organization. This use of MDM data does not even account for the significant MDM benefit of efficiently creating and curating master data, to begin with.
MDM benefits are many, including hierarchy management, data quality, data governance/workflow, data curation, and data distribution. One overlooked benefit is just having a database where trusted data can be accessed.
Like any data for access, the visualization aspect of this is important. With MDM data having a strong associative quality to it, the graph representation works quite well.

Graph traversals are a natural way of analyzing network patterns. Graphs can handle high degrees of separation with ease and facilitate visualization and exploration of networks and hierarchies.
Graph databases themselves are no substitute for MDM as they provide only one of the many necessary functions that an MDM tool does.
However, when graph technology is embedded within MDM, such as what IBM is doing in InfoSphere MDM – similar to AI (link) and blockchain (link) – it is very powerful.
Graph technology is one of the many ways to facilitate self-service to MDM. Long a goal of business intelligence, self-service has significant applicability to MDM as well. Self-service is opportunity-oriented.
Users may want to validate a hypothesis, experiment, innovate, etc. Long development cycles or laborious processes between a user and the data can be frustrating.
Historically, the burden for all MDM functions has fallen squarely on a centralized, development function. It’s overloaded and, as with the self-service business intelligence movement, needs disintermediation.

IBM is fundamentally changing this dynamic with the next release of Infosphere MDM. Its self-service data import, matching, and lightweight analytics allow the business user to find, share and get insight from both MDM and other data.
Big Match can analyze structured and unstructured customer data together to gain deeper customer insights. It can enable fast, efficient linking of data from multiple sources to grow and curate customer information.
The majority of the information in your organization that is not under management is unstructured data. Unstructured data has always been a valuable asset to organizations, but it can be difficult to manage.
Emails, documents, medical records, contracts, design specifications, legal agreements, advertisements, delivery instructions, and other text-based sources of information do not fit neatly into tabular relational databases.
Most BI tools on MDM data offer the ability to drill down and roll up data in reports and dashboards, which is good. But what about the ability to “walk sideways” across data sources to discover how different parts of the business interrelate?
Using unstructured data for customer profiling allows organizations to unify diverse data from inside and outside the enterprise—even the “ugly” stuff; that is, dirty data that is incompatible with highly structured, fact-dimension data that would have been too costly to combine using traditional integration and ETL methods.

Finally, unstructured data management enables text analytics, so that organizations can gain insight into customer sentiment, competitive trends, current news trends, and other critical business information.
In-text analytics, everything is fair game for consideration, including customer complaints, product reviews from the web, call center transcripts, medical records, and comment/note fields in an operational system.
Combining unstructured data with artificial intelligence and natural language processing can extract new attributes and facts for entities such as people, location, and sentiment from text, which can then be used to enrich the analytic experience.
All of these uses and capabilities are enhanced if they can be provided using a self-service interface that users can easily leverage to enrich data from within their apps and sources. This opens up a whole new world for discovery.
With graph technology, distribution of the publishing function and the integration of all data including unstructured data, MDM can truly have important data under management, empower the business user, be the cornerstone to digital transformation and truly be self-service.
Source: gigaom.com
In a normal master data management (MDM) project, a current state business process flow is built, followed by a future state business process flow that incorporates master data management.
The current state is usually ugly as it has been built piecemeal over time and represents something so onerous that the company is finally willing to do something about it and inject master data management into the process.
Many obvious improvements to process come out of this exercise and the future state is usually quite streamlined, which is one of the benefits of MDM.
I present today that these future state processes are seldom as optimized as they could be.
Consider the following snippet, supposedly part of an optimized future state.
This leaves in the process four people to manually look at the product, do their (unspecified) thing and (hopefully) pass it along, but possibly send it backwards to an upstream participant based on nothing evident in particular.
The challenge for MDM is to optimize the flow. I suggest that many of the “approval jails” in business process workflow are ripe for reengineering.
What criteria are used? It’s probably based on data that will now be in MDM.
If training data for machine learning (ML) is available, not only can we recreate past decisions to automate future decisions, we can look at the results of those decisions and take past outcomes and actually create decisions in the process that should have been made and actually do them, speeding up the flow and improving the quality by an order of magnitude.
This concept of thinking ahead and automating decisions extends to other kinds of steps in a business flow that involve data entry, including survivorship determination.
As with acceptance & rejection, data entry is also highly predictable, whether it is a selection from a drop-down or free-form entry. Again, with training data and backtesting, probable contributions at that step can be manifested and either automatically entered or provided as default for approval.
The latter approach can be used while growing a comfort level.
Manual, human-scale processes, are ripe for the picking and it’s really a dereliction of duty to “do” MDM without significantly streamlining processes, much of which is done by eliminating the manual.
As data volumes mount, it is often the only way to not watch process time increase over time. At the least, prioritizing stewardship activities or routing activities to specific stewards based on an ML interpretation of past results (quality, quantity) is required.
This approach is paramount to having timely, data-infused processes.
As a modular and scalable trusted analytics foundational element, the IBM Unified Governance & Integration platform incorporates advanced machine learning capabilities into MDM processes, simplifying the user experience and adding cognitive capabilities.
Machine learning can also discover master data by looking at actual usage patterns. ML can source, suggest or utilize external data that would aid in the goal of business processes.
Another important part of MDM is data quality (DQ). ML’s ability to recommend and/or apply DQ to data, in or out of MDM, is coming on strong.
Name-identity reconciliation is a specific example but generally, ML can look downstream of processes to see the chaos created by data lacking full DQ and start applying the rules to the data upstream.
IBM InfoSphere Master Data Management utilizes machine learning to speed the data discovery, mapping, quality and import processes.
In the last post (link), I postulated that blockchain would impact MDM tremendously. In this post, it’s machine learning affecting MDM. (Don’t get me started on graph technology).
Welcome to the new center of the data universe.
MDM is about to undergo a revolution.
Products will look much different in 5 years.
Make sure your vendor is committed to the MDM journey with machine learning.
Source: gigaom.com
Master Data Management (MDM) is an approach to the management of golden records that have been around over a decade only to find a growth spurt lately as some organizations are exceeding pain thresholds in the management of common data.

Blockchain has a slightly shorter history, coming aboard with bitcoin, but also is seeing its revolution these days as data gets distributed far and wide and trust has taken center stage in business relationships.
Volumes could be written about each on its own, and given that most organizations still have a way to go with each discipline, that might be appropriate. However, good ideas wait for no one and today’s idea is MDM on Blockchain.
Thinking back over our MDM implementations over the years, it is easy to see the data distribution network becoming wider. As a matter of fact, master data distribution is usually the most time-intensive and unwieldy part of an MDM implementation anymore.

The blockchain removes overhead, costs, and unreliability from authenticated peer-to-peer network partner transactions involving data exchange. It can support one of the big challenges of MDM with governed, bi-directional synchronization of master data between the blockchain and enterprise MDM.
Another core MDM challenge is arriving at the “single version of the truth”. It’s elusive even with MDM because everyone must tacitly agree to the process used to instantiate the data in the first place.
While many MDM practitioners go to great lengths to utilize the data rules from a data governance process, it is still a process subject to criticism.

The consensus that blockchain can achieve is a governance proxy for that elusive “single version of the truth” by achieving group consensus for trust as well as full lineage of data.
Blockchain enables the major components and tackles the major challenges in MDM.
Blockchain provides a distributed database, as opposed to a centralized hub, that can store data that is certified, and for perpetuity. By storing timestamped and linked blocks, the blockchain is unalterable and permanent.
Though not for low latency transactions yet, transactions involving master data, such as financial settlements, are ideal for blockchain and can be sped up by an order of magnitude since blockchain removes the grist in a normal process.
Blockchain uses pre-defined rules that act as gatekeepers of data quality and governs the way in which data is utilized. Blockchains can be deployed publicly (like bitcoin) or internally (like an implementation of Hyperledger).

There could be a blockchain per subject area (like customer or product) in the implementation. MDM will begin by utilizing these internal blockchain networks, also known as Distributed Ledger Technology, though utilization of public blockchains is inevitable.
A shared master data ledger beyond company boundaries can, for example, contain common and agreed master data including public company information and common contract clauses with only counterparties able to see the content and destination of the communication.
Hyperledger is quickly becoming the standard for the open-source blockchain. Hyperledger is hosted by The Linux Foundation. IBM, with the Hyperledger Fabric, is establishing the framework for blockchain in the enterprise.
Supporting master data management with a programmable interface for confidential transactions over a permissioned network is becoming a key inflection point for blockchain and Hyperledger.
Data management is about the right data at the right time and master data is fundamental to great data management, which is why centralized approaches like the discipline of master data management have taken center stage.

MDM can utilize the blockchain for distribution and governance and blockchain can clearly utilize the great master data produced by MDM. Blockchain data needs data governance like any data. This data actually needs it more given its importance on the network.
MDM and blockchain are going to be intertwined now. It enables the key components of establishing and distributing the single version of the truth of data. Blockchain enables trusted governed data. It integrates this data across broad networks. It prevents duplication and provides data lineage.
It will start in MDM in niches that demand these traits such as financial, insurance and government data. You can get to know the customer better with native fuzzy search and matching in the blockchain. You can track provenance, ownership, relationship and lineage of assets, do trade/channel finance and post-trade reconciliation/settlement.
Blockchain is now a disruption vector for MDM. MDM vendors need to be at least blockchain-aware today, creating the ability for blockchain integration in the near future, such as what IBM InfoSphere Master Data Management is doing this year. Others will lose ground.
Source: gigaom.com