All Apps and Add-ons

How do I do Job orchestration between Splunk and Hadoop?


My data primarily resides on Hadoop. I want to use Splunk for basic computations/ correlations, e.g : Aggregations / Sessionization and push the data back to Hadoop, where there would be further more complicated business logic applied on top of the moved data and the result stored back in Hadoop, which will eventually get indexed back in Splunk for query / reporting purpose.

To reliably do this, I need some orchestration of the sequence of these activities. E.g : (1) For a specific time segment, trigger the Report Acceleration/Summary Index. (2) Move the data to Hadoop (3) Run the MR job on the new data. (4) Index the result in Splunk.

Is their some oozie sort of integration with splunk that will let me orchestrate between splunk and hadoop jobs?

0 Karma

Splunk Employee
Splunk Employee

Splunk currently does not have such correlations.
However, you may want to try:
Hadoop Connect Export (move your results to Hadoop)
Hadoop Connect Import (index the enriched - after more MR Jobs - back into Splunk)
Hunk VIX pushing the results back from HDFS into Summary Index.

0 Karma
Get Updates on the Splunk Community!

Enterprise Security Content Update (ESCU) v3.54.0

The Splunk Threat Research Team (STRT) recently released Enterprise Security Content Update (ESCU) v3.54.0 and ...

Using Machine Learning for Hunting Security Threats

WATCH NOW Seeing the exponential hike in global cyber threat spectrum, organizations are now striving more for ...

New Learning Videos on Topics Most Requested by You! Plus This Month’s New Splunk ...

Splunk Lantern is a customer success center that provides advice from Splunk experts on valuable data ...