Hadoop

The Apache Hadoop framework for the processing of data on commodity hardware is at the center of the Big Data picture today. Key solutions and technologies include the Hadoop Distributed File System (HDFS), YARN, MapReduce, Pig, Hive, Security, as well as a growing spectrum of solutions that support Business Intelligence (BI) and Analytics.

Hadoop Articles

MapR, Cisco, and SAP Come Together to Create Tools for SAP HANA

Cisco is launching an appliance that includes the MapR Converged Data Platform for SAP HANA, making it easier and faster for users to take advantage of big data. The UCS Integrated Infrastructure for SAP HANA is made easy to deploy, speeds time to market, and will reduce operational expenses along with providing users with the flexibility to choose a scale-up (on-premises) or scale-out (cloud) storage strategy.

Posted April 27, 2016

Cloudera Enterprise 5.7 Boosts Data Processing with Hive-on-Spark Support

Cloudera, provider of a data management and analytics platform built on Apache Hadoop and open source technologies, has announced the general availability of Cloudera Enterprise 5.7. According to the vendor, the new release offers an average 3x improvement for data processing with added support of Hive-on-Spark, and an average 2x improvement for business intelligence analytics with updates to Apache Impala (incubating).

Posted April 26, 2016

Neo4j Boosts its Graph Database

Neo Technology, creator of Neo4j, is releasing an improved version of its signature platform, enhancing its scalability, introducing new language drivers and a host of other developer friendly features.

Posted April 26, 2016

Ground-Breaking Research on New IT Trends Adoption is Presented at COLLABORATE 16

The COLLABORATE 16 conference for Oracle users kicked off with a presentation by Unisphere Research analyst Joe McKendrick who shared insights from a ground-breaking study that examined future trends and technology among 690 members of three major Oracle users groups.

Posted April 25, 2016

Bridging the Data Divide: Getting the Most Value From Data With Integration

The need for data integration has never been more intense than it has been recently. The Internet of Things and its muscular sibling, the Industrial Internet of Things, are now being embraced as a way to better understand the status and working order of products, services, partners, and customers. Mobile technology is ubiquitous, pouring in a treasure trove of geolocation and usage data. Analytics has become the only way to compete, and with it comes a need for terabytes—and gigabytes—worth of data. The organization of 2016, in essence, has become a data machine, with an insatiable appetite for all the data that can be ingested.

Posted April 25, 2016

GridGain Offers New Edition of Flagship Platform

GridGain Systems, provider of enterprise-grade in-memory data fabric solutions based on Apache Ignite, is releasing a new version of its platform. GridGain Professional Edition includes the latest version of Apache Ignite plus LGPL libraries, along with a subscription that includes monthly maintenance releases with bug fixes that have been contributed to the Apache Ignite project but will be included only with the next quarterly Ignite release.

Posted April 20, 2016

Dataguise DgSecure Extends Data Security with Support for AWS EMR/S3

Dataguise, a provider of data security solutions, is making DgSecure available for the detection, monitoring, and protection of sensitive data across Amazon Web Services (AWS) Simple Storage Service (S3) and all Elastic MapReduce (EMR) platforms that use AWS S3.

Posted April 19, 2016

Sumo Logic Offers Platform that Analyzes Metrics and Log Analytics on AWS

Sumo Logic, a provider of cloud-native, machine data analytics services, is unveiling a new platform that natively ingests, indexes, and analyzes structured metrics data, and unstructured log data together in real-time.

Posted April 18, 2016

Hortonworks Augments its Platform and Strengthens its Partnerships

Hortonworks is making several key updates to its platform along with furthering its mission as being a leading innovator of open and connected data solutions by enhancing partnerships with Pivotal and expanding upon established integrations with Syncsort.

Posted April 15, 2016

Databricks' Kavitha Mariappan on Why Spark is So Hot Now

First created as part of a research project at UC Berkeley AMPLab, Spark is an open source project in the big data space, built for sophisticated analytics, speed, and ease of use. It unifies critical data analytics capabilities such as SQL, advanced analytics, and streaming in a single framework. Databricks is a company that was founded by the team that created and continues to lead both the development and training around Apache Spark.

Posted April 14, 2016

An Era of Disruption, Transformation, and Reinvention

Thanks to the digital business transformation, the world around us is changing—and quickly—to a very consumer- and data-centric economy, where companies must transform to remain competitive and survive. The upshot is that for many companies today, it is a full-on Darwinian experience of survival of the fittest.

Posted April 08, 2016

Qubole Open Sources SQL Data Virtualization Engine for Big Data

Qubole, a big data-as-a-service company, is open sourcing its Quark platform, a cost-based SQL optimizer. The Quark project is also available in a SaaS implementation via the Qubole Data Service (QDS).

Posted April 07, 2016

IBM Speeds Time to Insights with z/OS Platform for Apache Spark

IBM says it is making it easier and faster for organizations to access and analyze data in-place on the IBM z Systems mainframe with a new z/OS Platform for Apache Spark. The platform enables Spark to run natively on the z/OS mainframe operating system.

Posted April 04, 2016

Databricks Launches New APIs to Grab Hidden Insights

Databricks, the company behind Apache Spark, is releasing a new set of APIs that will enable enterprises to automate their Spark infrastructure to accelerate the deployment of production data-driven applications.

Posted April 01, 2016

ManageEngine Releases Application Performance Monitoring Solution for Hadoop and Oracle

ManageEngine is introducing a new application performance monitoring solution, enabling IT operations teams in enterprises to gain operational intelligence into big data platforms. Applications Manager enables performance monitoring of Hadoop clusters to minimize downtime and performance degradation. Additionally, the platform's monitoring support for Oracle Coherence provides insights into the health and performance of Coherence clusters and facilitates troubleshooting of issues.

Posted April 01, 2016

Kafka Is the New Standard in Big Data Messaging

It's become almost a standard career path in Silicon Valley: A talented engineer creates a valuable open source software commodity inside of a larger organization, then leaves that company to create a new startup to commercialize the open source product. Indeed, this is virtually the plot line for the hilarious HBO comedy series, Silicon Valley. Jay Krepes, a well-known engineer at LinkedIn and creator of the NoSQL database system, Voldemort, has such a story.

Posted March 31, 2016

Denodo Introduces New Cloud and Self-Service Data Discovery Tools

Denodo, a provider of data virtualization software, is releasing Denodo Platform 6.0, further accelerating its "fast data" strategy. "It's a major release for us," said Ravi Shankar, Denodo CMO. There are three important areas that nobody else is focusing on in the industry, he noted. "This, we hope, will change how data virtualization, and in a broader sense, data integration will shape up this year."

Posted March 31, 2016

Streamlining Business With a NoSQL Document Database

NoSQL databases were born out of the need to scale transactional persistence stores more efficiently. In a world where the relational database management system (RDBMS) was king, this was easier said than done.

Posted March 29, 2016

Bigstep Now Offers MapR as Part of Its Big Data Platform

MapR is now available as part of Bigstep's big data platform-as-a-service, supporting a wide range of Hadoop applications.

Posted March 29, 2016

Reltio Releases an Upgraded Version of its Cloud Platform

Reltio is releasing an enhanced version of Reltio Cloud 2016.1, adding new analytics integration, collaboration, and recommendation capabilities to help companies be right faster.

Posted March 29, 2016

Teradata Maps Faster Path to Data Lake with Design Pattern Approach

Teradata has introduced a new "design pattern" approach for data lake deployment. The company says its concept of a data lake pattern leverages IP from its client engagements, as well as services and technology to help organizations more quickly and securely get to successful data lake deployment.

Posted March 28, 2016

Big Data at a Turning Point: Q&A with Joe Caserta

The rise of big data technologies in enterprise IT is now seen as an inevitability, but adoption has occurred at a slower pace than expected, according to Joe Caserta, president and CEO of Caserta Concepts, a firm focused on big data strategy consulting and technology implementation. Caserta recently discussed the trends in big data projects, the technologies that offer key advantages now, and why he thinks big data is reaching a turning point.

Posted March 23, 2016

Informatica Launches ‘Intelligent’ Data Lake Approach

Informatica has launched an end-to-end solution to help customers gain greater insight from big data.

Posted March 23, 2016

SAP HANA Vora Now Available

SAP SE's newest in memory query engine, SAP HANA Vora, is now generally available, equipping enterprises with contextual analytics across all data stored in Hadoop, enterprise systems, and other distributed data sources.

Posted March 23, 2016

5 Myths of Application Security in the Cloud

As more businesses leverage applications that are hosted in the cloud, the lines between corporate networks and the internet become blurred. Accordingly, enterprises need to develop an effective strategy for ensuring security. The problem is, many of today's most common approaches simply don't work in this new cloud-based environment.

Posted March 23, 2016

Education, Collaboration, and Inspiration: Save Time and Money Accomplishing It All Under One Roof at COLLABORATE 16

The OAUG volunteers planning COLLABORATE 16: Technology and Applications Forum for the Oracle Community (April 10-14 at Mandalay Bay in Las Vegas) are themselves Oracle users and technologists, understanding innately the myriad options and challenges faced by the wider user community in a period of rapid change and transformation. With participation and contributions from all corners of the user community, COLLABORATE offers the information and perspective to make sense of it all.

Posted March 21, 2016

Oracle Boosts Data Analytics with Free and Open Interfaces to On-Chip Accelerators

Oracle has released a free and open API and developer kit for its Data Analytics Accelerator (DAX) in SPARC processors through its Software in Silicon Developer Program. "Through our Software in Silicon Developer Program, developers can now apply our DAX technology to a broad spectrum of previously unsolvable challenges in the analytics space because we have integrated data analytics acceleration into processors, enabling unprecedented data scan rates of up to 170 billion rows per second," said John Fowler, executive vice president of Systems, Oracle.

Posted March 16, 2016

Talend Adds ‘Easy Button’ to Automate Big Data Integration in AWS Environments

Available now, Talend says its Integration Cloud Spring '16 release adds enhancements to help IT organizations execute big data and data integration projects running on AWS Redshift or AWS Elastic MapReduce (EMR) with greater ease - using fewer resources, and at a reduced cost.

Posted March 16, 2016

Attivio Receives a Substantial Funding Increase after Watershed Year in 2015

Attivio is receiving $31 million in investment financing that will help expand the company as it accelerates its offerings into the big data market.

Posted March 09, 2016

IDERA Now Offers Full Suite of Database Lifecycle Solutions

IDERA, a provider of database lifecycle management solutions, is extending its product portfolio by adding Embarcadero Technologies' ER/Studio and DB PowerStudio tools, allowing organizations to rely on a single vendor for all their database lifecycle needs.

Posted March 09, 2016

The Database Technologies of the Future

In a new book titled "Next Generation Databases," Guy Harrison, an executive director of R&D at Dell, shares what every data professional needs to know about the future of databases in a world of NoSQL and big data.

Posted March 08, 2016

Syncsort Simplifies Mainframe Big Data Access and Governance in Hadoop and Spark

Syncsort is introducing new capabilities to its data integration software, DMX-h, that allow organizations to work with mainframe data in Hadoop or Spark in its native format, which is necessary for preserving data lineage and maintaining compliance.

Posted March 07, 2016

Infobright Reveals New Solution for Faster Large-Scale Queries

Infobright, the columnar database analytics platform, has unveiled its new Infobright Approximate Query (IAQ) solution for large-scale data environments, allowing users to gain insights faster and efficiently. "This technology is being delivered on the basis of rethinking the business problem and using technology in a very meaningful way to solve problems that would otherwise be unsolvable using a traditional approach," said Don DeLoach, CEO.

Posted February 26, 2016

SAP Introduces New Predictive Analytics Capabilities to its Platforms

SAP SE is introducing new predictive capabilities within its platforms with the release of SAP HANA Cloud Platform predictive services 1.0 and SAP Predictive Analytics 2.5.

Posted February 24, 2016

Information Builders Introduces iWay Hadoop Data Manager for Big Data Lakes

The promise of the data lake is an enduring repository of raw data that can be accessed now and in the future for different purposes. To help companies on their journey to the data lake, Information Builders has unveiled the iWay Hadoop Data Manager, a new solution that provides an interface to generate portable, reusable code for data integration tasks in Hadoop.

Posted February 23, 2016

Reflections on 10 Years of Hadoop and Big Data

It is hard to think of a technology that is more identified with the rise of big data than Hadoop. Since its creation, the framework for distributed processing of massive datasets on commodity hardware has had a transformative effect on the way data is collected, managed, and analyzed - and also grown well beyond its initial scope through a related ecosystem of open source projects. With 2016 recognized as the 10-year anniversary for Hadoop, Big Data Quarterly chose this time to ask technologists, consultants, and researchers to reflect on what has been achieved in the last decade, and what's ahead on the horizon.

Posted February 18, 2016

Oracle Updates Database Products for the New Era of Cloud and Big Data

Currently, the IT industry is the midst of a major transition as it moves from the last generation - the internet generation - to the new generation of cloud and big data, said Andy Mendelsohn, Oracle's EVP of Database Server Technologies, who recently talked with DBTA about database products that Oracle is bringing to market to support customers' cloud initiatives. "Oracle has been around a long time. This is not the first big transition we have gone through," said Mendelsohn.

Posted February 17, 2016

HPE Taps RedPoint Global for New Risk Data Aggregation and Reporting Solution

Hewlett Packard Enterprise (HPE) has selected RedPoint Data Management platform as the underlying platform for a new HPE Risk Data Aggregation and Reporting (RDAR) integrated solution to support financial institutions' compliance with BDBS 239.

Posted February 16, 2016

Trillium Integrates Data Prep and Data Quality for Big Data Analytics

Addressing the shift toward business-user-oriented visual interactive data preparation, Trillium Software has launched a new solution that integrates self-service data preparation with data quality capabilities to improve big data analytics.

Posted February 16, 2016

What Oracle’s NoSQL SQL Database Reveals

Say what you will about Oracle, it certainly can't be accused of failing to move with the times. Typically, Oracle comes late to a technology party but arrives dressed to kill.

Posted February 10, 2016

Oracle Launches New Big Data Preparation Cloud Service

Oracle has introduced a new Big Data Preparation Cloud Service. Despite the increasing talk about the need for companies to become "data-driven," and the perception that people who work with business data spend most of their time on analytics, Oracle contends that in reality many organizations devote much more time and effort on importing, profiling, cleansing, repairing, standardizing, and enriching their data.

Posted February 10, 2016