Tag Archives for " databases "

Data Warehouse

Revenge of the Data Warehouse: How a Classic Tech Category Has Evolved, and Triumphed

The pressure to leverage data as a business asset is stronger than ever. Enterprises everywhere are eager to devise sound data strategies that are realistic and achievable, based on available budget, and sensitive to in-house technology skill sets.

Data Virtualization

For a while, it looked like open source, specialized big data compute frameworks, including Hadoop and Spark, were the way to go.

Enterprise organizations found them compelling for reasons of novelty, economics, and the apparent prudence of a future-looking technology.

But those frameworks are at a bit of a crossroads: the hype around them has subsided and — while things are improving — the success rate of enterprise projects involving them has been modest.

 

Meanwhile, the data warehouse (DW) which, for decades, has been a key technology platform for enterprise analytics, never went away.

Yes, DWs struggled and incumbent DW platforms still do, but recent advances in storage costs and compute scalability, especially in the cloud, have addressed the most important challenges faced by DW platforms.

As a result, we are in a DW renaissance period.

Enterprise Data Governance with Modern Data Catalog Platforms: A GigaOm Research Byte

Problems with DWs have largely been solved, petabyte-scale data volumes no longer defeat them and the familiarity and ease of use that kept them viable all this time are now helping them face their open source big data competition and, in many cases, emerge victorious.

Still, if DW platforms have changed, what should enterprises do to build an analytics strategy that integrates them?

Even the most DW-loyal shops will need to look at how DW platforms have evolved and adjust their strategies accordingly.

Organizations that have committed to open source analytics technologies will need to take a second look at DW platforms and consider a strategy that combines both, essentially bringing the data warehouse together with the data lake.

The trick to adapting to the new world of data warehousing is understanding that it is not only the technology that has changed but the applications and use cases for DW technology as well.

data protection

DW platforms do not need to be used exclusively for Enterprise DW implementations.

The platforms are now more versatile and can be used for use-case-specific workloads and even exploratory analytics.

In a sense, the DW isn’t just a DW anymore.

Even the juxtaposition of DW and Data Lake has shifted – DW platforms today are increasingly able to ingest raw, semi-structured data, or query it in place.

This means warehouse and lake technology can be used in combination and, sometimes, lake technology will not be necessary.

It is not just the Data Lake and its workloads that are becoming more integrated into the warehouse.

Streaming data, machine learning, and AI are onboarding as well.

In addition, data governance and data protection are starting to enter the DW orbit.

While the familiarity of the relational model, dimensional design, and SQL are back; the application of DW technology now covers territory that may be less familiar.

Moreover, new vendors who have championed or been born in the cloud are emerging in leadership positions.

The new capabilities, new use cases, and new vendors make the space exciting but also difficult to navigate (for newbies and veterans alike).

Enterprise customers will need to understand how DW platforms have morphed and shape-shifted; how best to use, deploy and implement them; combine them with other applications; and understand the key differences between the vendors and their offerings.

Without this knowledge, Enterprise buyers will be frozen in indecision.

Data Warehouse

With it, they will be armed to leverage today’s DW platforms to their fullest and cherry-pick other technologies that can augment and optimize them.

The end result is that organizations in the know will be ready to analyze all their data, in technologically familiar environments, for full competitive and operational advantage.

In this report, we map out today’s DW landscape in the context of where the technology has been, where it is, and where it is going.

 

Key Findings:

  • The Data Warehouse is alive and well, perhaps enjoying its biggest popularity wave to date
  • Open source technology challengers addressed issues of storage costs and horizontal scalability but did not achieve parity in terms of enterprise skill-set abundance, interactive query performance, or acting as authoritative data repositories
  • The advent of the cloud has helped DW platforms transcend their vulnerabilities, through the cloud’s power of elasticity for compute and the use of economical, infinitely scalable cloud object storage
  • Transcending their vulnerabilities and satisfying customers’ need for familiar SQL-relational platform paradigms has turned out to be a one-two punch for DWs
  • Customers need to bring themselves up-to-date on the latest DW innovations and product categories, understanding each in the context of the DW’s historical evolution and market factors
  • Customers must correlate DW product categories with corporate cloud (and multi-cloud) strategy

 

Source: gigaom.com

Data APIs

Data APIs: Gateway to Data-Driven Operation and Digital Transformation

Enterprises everywhere are on a quest to use their data efficiently and innovatively, and to maximum advantage, both in terms of operations and competitiveness.

The advantages of doing so are taken on authority and reasonably so. Analyzing your data helps you better understand how your business actually runs.

Such insights can help you see where things can improve, and can help you make instantaneous decisions when required by emergent situations.

Summary:

You can even use your data to build predictive models that help you forecast operations and revenue, and, when applied correctly, these models can be used to prescribe actions and strategies in advance.

That today’s technology allows businesses to do this is exciting and inspiring. Once such practice becomes widespread, we’ll have trouble believing that our planning and decision-making weren’t data-driven in the first place.

bumps on road

Bumps in the Road

But we need to be cautious here. Even though the technological breakthroughs we’ve had are impressive and truly transformative, there are some dependencies – prerequisites – that must be met in order for these analytics technologies to work properly.

If we get too far ahead of those requirements, then we’ll we will not succeed in our initiatives to extract business insights from data.

The dependencies concern the collection, the cleanliness, and the thoughtful integration of the organization’s data with the analytics layer.

And, in an unfortunate irony, while the analytics software has become so powerful, the integration work that’s needed to exploit that power has become more difficult.

 

From Consolidated to Dispersed

The reason for this added difficulty is the fragmentation and distribution of an organization’s data. Enterprise software, for the most part, used to run on-premises and much of its functionality was consolidated into a relatively small stable of applications, many of which shared the same database platform.

Integrating the databases was a manageable process if proper time and resources were allocated.

But with so much enterprise software functionality now available through Software as a Service (SaaS) offerings in the cloud, bits, and pieces of an enterprise’s data are now dispersed through different cloud environments on a variety of platforms.

Pulling all of this data together is a unique exercise for each of these cloud applications, multiplying the required integration work many times over.

Even on-premises, the world of data has become complex. The database world was dominated by three major relational database management system (RDBMS) products, but that’s no longer the case.

Now, in addition to the three commercial majors, two open-source RDBMSs have joined them in Enterprise popularity and adoption.

And beyond the RDBMS world, various NoSQL databases and Big Data systems, like Hadoop and MongoDB, have joined the on-premises data fray.

Set Goal for Self-motivation

A Way Forward

A major question emerges. As this data fragmentation is not merely an exception or temporary inconvenience, but rather the new normal, is there a way to approach it holistically?

Can enterprises that must solve the issue of data dispersal and fragmentation at least have a unified approach to connecting to, integrating, and querying that data?

While an ad hoc approach to integrating data one source at a time can eventually work, it’s a very expensive and slow way to go, and yields solutions that are very brittle.

In this report, we will explore the role of application programming interfaces (APIs) in pursuing the comprehensive data integration that is required to bring about a data-driven organization and culture.

We’ll discuss the history of conventional APIs and the web-standards that most APIs use today. We’ll then explore how APIs and the metaphor of a database with tables, rows, and columns can be combined to create a new kind of API.

And we’ll see how this new type of API scales across an array of data sources and is more easily accessible than older API types, by developers and analysts alike.

Source: gigaom.com