Splunk Search

Comparing raw data volume Vs indexed data volume

gnanaraj_mcc
Loves-to-Learn Lots

How do i compare my raw data volume to the indexed data volume for a specific source type?

Can someone help with this query?

We have index clustering, a deployment server, and a distributed management console.

i want to make sure their same data is not indexed more than one time. (dual, triple indexing of same data)

0 Karma

sloshburch
Ultra Champion

To determine duplicate data, you could do a | stats count by _raw, _time, host, source although I promise that will be a slow and painful process.

Indexed data volume is captured in index=_internal source=*/license_usage.log sourcetype=splunkd and then you can specify a sourcetype using the st= field.

Where do you think you have duplication? Starting with the symptoms that motivated your question will help us be more surgical in what would otherwise be a very involved process.

0 Karma
Got questions? Get answers!

Join the Splunk Community Slack to learn, troubleshoot, and make connections with fellow Splunk practitioners in real time!

Meet up IRL or virtually!

Join Splunk User Groups to connect and learn in-person by region or remotely by topic or industry.

Get Updates on the Splunk Community!

Thanks for the Memories: .conf26 Took Learning to New Heights

Thank you, Splunk Community, for making .conf26 in Denver one for the books. From packed Splunk University ...

Best Practices: Splunk auto adjust pipeline queue

When you enable autoAdjustQueue in Splunk, maxSize should be understood as the queue size Splunk starts with ...

Splunk Auto Ingestion Parallel Pipeline Scaling

Why this feature matters Many Splunk environments experience ingestion pressure long before the host is fully ...