All Apps and Add-ons

How can I read a public dataset from public s3 bucket?

jdunlea1
Explorer

Is there a way (using the Splunk TA for AWS or otherwise) that Splunk can connect to a publicly available S3 bucket (such as those made available here https://registry.opendata.aws/) and read in the data?

From the Splunk TA, the only buckets that I can read from are those which were created in my account.

0 Karma
1 Solution

jdunlea1
Explorer

After some further digging and testing it appears that it can be done but you need to create the input using the conf file as per link text

The key here for ingesting "old" data from a public dataset in S3 is that you need to set initial_scan_datetime to be a date that is BEFORE the modified file time for the files in the S3 bucket.

Once I did this, I was able to pull the public dataset into Splunk from the public S3 bucket.

View solution in original post

0 Karma

jdunlea1
Explorer

After some further digging and testing it appears that it can be done but you need to create the input using the conf file as per link text

The key here for ingesting "old" data from a public dataset in S3 is that you need to set initial_scan_datetime to be a date that is BEFORE the modified file time for the files in the S3 bucket.

Once I did this, I was able to pull the public dataset into Splunk from the public S3 bucket.

0 Karma
Get Updates on the Splunk Community!

Strengthen Your Future: A Look Back at Splunk 10 Innovations and .conf25 Highlights!

The Big One: Splunk 10 is Here!  The moment many of you have been waiting for has arrived! We are thrilled to ...

Now Offering the AI Assistant Usage Dashboard in Cloud Monitoring Console

Today, we’re excited to announce the release of a brand new AI assistant usage dashboard in Cloud Monitoring ...

Stay Connected: Your Guide to October Tech Talks, Office Hours, and Webinars!

What are Community Office Hours? Community Office Hours is an interactive 60-minute Zoom series where ...