<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: Splunk Storage Sizing Guidelines and calculations in Deployment Architecture</title>
    <link>https://community.splunk.com/t5/Deployment-Architecture/Splunk-Storage-Sizing-Guidelines-and-calculations/m-p/439056#M15469</link>
    <description>&lt;P&gt;Thank you so much Rich for your reply.&lt;BR /&gt;
But then again I just came across this document which says "Typically, the compressed rawdata file is 10% the size of the incoming, pre-indexed raw data. The associated index files range in size from approximately 10% to 110% of the rawdata file. The number of unique terms in the data affect this value."&lt;/P&gt;

&lt;P&gt;&lt;A href="https://docs.splunk.com/Documentation/Splunk/7.2.6/Capacity/Estimateyourstoragerequirements"&gt;https://docs.splunk.com/Documentation/Splunk/7.2.6/Capacity/Estimateyourstoragerequirements&lt;/A&gt;&lt;/P&gt;

&lt;P&gt;Can you please brief me what does it mean exactly which is in documents? &lt;BR /&gt;
Because as per the documents i guess below will be the calculation&lt;BR /&gt;
Actual Data = Raw Data + Index Files&lt;BR /&gt;
100 GB = 10GB (10% of actual data) + ((1GB to 11GB) {10% to 110% of rawdata}))&lt;BR /&gt;
               = 11 GB to 21GB&lt;BR /&gt;
               = 25 GB @round off figure for 5 servers(Considering higher value with round off figure)&lt;BR /&gt;
               = 5 GB / server &lt;BR /&gt;
               5*180 = 900 GB / Server and 25 * 180 = 4.5 TB for 5 servers&lt;/P&gt;

&lt;P&gt;I might be wrong but just could not match the documentation part&lt;/P&gt;

&lt;P&gt;Because as per the Splunk Storage Sizing, size of  index files(which are having only pointers for your indexed data i believe) is more than size of your actual indexed data(rawdata)&lt;BR /&gt;
Isn't it sounds something unusual? I guess indexed data should be bigger than index files.&lt;/P&gt;</description>
    <pubDate>Mon, 06 May 2019 09:44:02 GMT</pubDate>
    <dc:creator>Ajinkya1992</dc:creator>
    <dc:date>2019-05-06T09:44:02Z</dc:date>
    <item>
      <title>Splunk Storage Sizing Guidelines and calculations</title>
      <link>https://community.splunk.com/t5/Deployment-Architecture/Splunk-Storage-Sizing-Guidelines-and-calculations/m-p/439054#M15467</link>
      <description>&lt;P&gt;Hi Team,&lt;BR /&gt;
I have doubt with Splunk Storage Sizing apps&lt;BR /&gt;
&lt;A href="https://splunk-sizing.appspot.com/#ar=0&amp;amp;c=1&amp;amp;cf=0.15&amp;amp;cr=180&amp;amp;hwr=7&amp;amp;i=5&amp;amp;rf=1&amp;amp;sf=1&amp;amp;st=v&amp;amp;v=100"&gt;https://splunk-sizing.appspot.com/#ar=0&amp;amp;c=1&amp;amp;cf=0.15&amp;amp;cr=180&amp;amp;hwr=7&amp;amp;i=5&amp;amp;rf=1&amp;amp;sf=1&amp;amp;st=v&amp;amp;v=100&lt;/A&gt;&lt;/P&gt;

&lt;P&gt;I am keeping it very simple, lets suppose need to ingest 100GB/Day&lt;BR /&gt;
Data Rentntion is for 6 Months&lt;BR /&gt;
Number of indexers in cluster 5&lt;BR /&gt;
Search Factor 1 &lt;BR /&gt;
Replication Factor 1&lt;/P&gt;

&lt;P&gt;As per Splunk Storage Sizing &lt;BR /&gt;
Raw Compression Factor - Typically the compressed raw data file is 15% of the incoming pre-indexed data. The number of unique terms affects this value.&lt;BR /&gt;
Metadata Size Factor - Typically metadata is 35% of raw data, The type of data and index files will affects this value&lt;/P&gt;

&lt;P&gt;So as per the above calculation 15% of 100GB = 15GB&lt;BR /&gt;
and                             35% of 15GB  = 5.25FB&lt;BR /&gt;
which is 20.25GB for 5 Servers/Day and 4.05GB/Day for 1 server&lt;/P&gt;

&lt;P&gt;So if we are considering retention period of 180 Days then 4.05*180 = 729GB/Server for Six months and 3645GB (3.6TB) for 5 servers&lt;/P&gt;

&lt;P&gt;But as per the Splunk Storage Sizing &lt;BR /&gt;
You need to have 1.8TB/server and 9.1TB for 5 servers.&lt;/P&gt;

&lt;P&gt;My calcualtion and Splunk Storage Sizing calculation doesnt match at all.&lt;BR /&gt;
Splunk Storage sizing calculation goes with 50% of preindexed data completely, where as per their guidelines metadata is 35% of raw data not actual incoming data.&lt;/P&gt;

&lt;P&gt;Please let me know what I am missing.&lt;/P&gt;</description>
      <pubDate>Sun, 05 May 2019 08:25:25 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Deployment-Architecture/Splunk-Storage-Sizing-Guidelines-and-calculations/m-p/439054#M15467</guid>
      <dc:creator>Ajinkya1992</dc:creator>
      <dc:date>2019-05-05T08:25:25Z</dc:date>
    </item>
    <item>
      <title>Re: Splunk Storage Sizing Guidelines and calculations</title>
      <link>https://community.splunk.com/t5/Deployment-Architecture/Splunk-Storage-Sizing-Guidelines-and-calculations/m-p/439055#M15468</link>
      <description>&lt;P&gt;The 15% and 35% calculations should be made on the same raw daily ingestion value.  An easier method is to take %50 of the daily ingestion value as the daily storage requirement.&lt;/P&gt;</description>
      <pubDate>Sun, 05 May 2019 17:17:07 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Deployment-Architecture/Splunk-Storage-Sizing-Guidelines-and-calculations/m-p/439055#M15468</guid>
      <dc:creator>richgalloway</dc:creator>
      <dc:date>2019-05-05T17:17:07Z</dc:date>
    </item>
    <item>
      <title>Re: Splunk Storage Sizing Guidelines and calculations</title>
      <link>https://community.splunk.com/t5/Deployment-Architecture/Splunk-Storage-Sizing-Guidelines-and-calculations/m-p/439056#M15469</link>
      <description>&lt;P&gt;Thank you so much Rich for your reply.&lt;BR /&gt;
But then again I just came across this document which says "Typically, the compressed rawdata file is 10% the size of the incoming, pre-indexed raw data. The associated index files range in size from approximately 10% to 110% of the rawdata file. The number of unique terms in the data affect this value."&lt;/P&gt;

&lt;P&gt;&lt;A href="https://docs.splunk.com/Documentation/Splunk/7.2.6/Capacity/Estimateyourstoragerequirements"&gt;https://docs.splunk.com/Documentation/Splunk/7.2.6/Capacity/Estimateyourstoragerequirements&lt;/A&gt;&lt;/P&gt;

&lt;P&gt;Can you please brief me what does it mean exactly which is in documents? &lt;BR /&gt;
Because as per the documents i guess below will be the calculation&lt;BR /&gt;
Actual Data = Raw Data + Index Files&lt;BR /&gt;
100 GB = 10GB (10% of actual data) + ((1GB to 11GB) {10% to 110% of rawdata}))&lt;BR /&gt;
               = 11 GB to 21GB&lt;BR /&gt;
               = 25 GB @round off figure for 5 servers(Considering higher value with round off figure)&lt;BR /&gt;
               = 5 GB / server &lt;BR /&gt;
               5*180 = 900 GB / Server and 25 * 180 = 4.5 TB for 5 servers&lt;/P&gt;

&lt;P&gt;I might be wrong but just could not match the documentation part&lt;/P&gt;

&lt;P&gt;Because as per the Splunk Storage Sizing, size of  index files(which are having only pointers for your indexed data i believe) is more than size of your actual indexed data(rawdata)&lt;BR /&gt;
Isn't it sounds something unusual? I guess indexed data should be bigger than index files.&lt;/P&gt;</description>
      <pubDate>Mon, 06 May 2019 09:44:02 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Deployment-Architecture/Splunk-Storage-Sizing-Guidelines-and-calculations/m-p/439056#M15469</guid>
      <dc:creator>Ajinkya1992</dc:creator>
      <dc:date>2019-05-06T09:44:02Z</dc:date>
    </item>
    <item>
      <title>Re: Splunk Storage Sizing Guidelines and calculations</title>
      <link>https://community.splunk.com/t5/Deployment-Architecture/Splunk-Storage-Sizing-Guidelines-and-calculations/m-p/708258#M29013</link>
      <description>&lt;P&gt;Things have improved a lot thanks to&amp;nbsp;&lt;SPAN&gt;tsidxWritingLevel enhancements. &lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;If you set&amp;nbsp;tsidxWritingLevel=4, the maximum available today, and all your buckets have been already written with this level you can achieve a compress ratio of 5.35:1&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;This means 55 TB of raw logs will occupy around 10 TB (tsidx + raw) on disk.&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;At least this is what we have in our deployment.&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;This number can vary depending on the type of data you are ingesting.&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;Here the query I used, running All Time, starting from the one present in the &lt;EM&gt;&lt;STRONG&gt;Monitoring Console &amp;gt;&amp;gt; Indexing &amp;gt;&amp;gt; Index and Volumes &amp;gt;&amp;gt; Index Detail: Instance&lt;/STRONG&gt;&lt;/EM&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;LI-CODE lang="markup"&gt;| rest splunk_server=&amp;lt;oneOfYourIndexers&amp;gt; /services/data/indexes datatype=all
  | join type=outer title [
    | rest splunk_server=&amp;lt;oneOfYourIndexers&amp;gt; /services/data/indexes-extended datatype=all
  ]
| `dmc_exclude_indexes`
| eval warm_bucket_size = coalesce('bucket_dirs.home.warm_bucket_size', 'bucket_dirs.home.size')
| eval cold_bucket_size = coalesce('bucket_dirs.cold.bucket_size', 'bucket_dirs.cold.size')
| eval hot_bucket_size = if(isnotnull(cold_bucket_size), total_size - cold_bucket_size - warm_bucket_size, total_size - warm_bucket_size)
| eval thawed_bucket_size = coalesce('bucket_dirs.thawed.bucket_size', 'bucket_dirs.thawed.size')
| eval warm_bucket_size_gb = coalesce(round(warm_bucket_size / 1024, 2), 0.00)
| eval hot_bucket_size_gb = coalesce(round(hot_bucket_size / 1024, 2), 0.00)
| eval cold_bucket_size_gb = coalesce(round(cold_bucket_size / 1024, 2), 0.00)
| eval thawed_bucket_size_gb = coalesce(round(thawed_bucket_size / 1024, 2), 0.00)

| eval warm_bucket_count = coalesce('bucket_dirs.home.warm_bucket_count', 0)
| eval hot_bucket_count = coalesce('bucket_dirs.home.hot_bucket_count', 0)
| eval cold_bucket_count = coalesce('bucket_dirs.cold.bucket_count', 0)
| eval thawed_bucket_count = coalesce('bucket_dirs.thawed.bucket_count', 0)
| eval home_event_count = coalesce('bucket_dirs.home.event_count', 0)
| eval cold_event_count = coalesce('bucket_dirs.cold.event_count', 0)
| eval thawed_event_count = coalesce('bucket_dirs.thawed.event_count', 0)

| eval home_bucket_size_gb = coalesce(round((warm_bucket_size + hot_bucket_size) / 1024, 2), 0.00)
| eval homeBucketMaxSizeGB = coalesce(round('homePath.maxDataSizeMB' / 1024, 2), 0.00)
| eval home_bucket_capacity_gb = if(homeBucketMaxSizeGB &amp;gt; 0, homeBucketMaxSizeGB, "unlimited")
| eval home_bucket_usage_gb = home_bucket_size_gb." / ".home_bucket_capacity_gb
| eval cold_bucket_capacity_gb = coalesce(round('coldPath.maxDataSizeMB' / 1024, 2), 0.00)
| eval cold_bucket_capacity_gb = if(cold_bucket_capacity_gb &amp;gt; 0, cold_bucket_capacity_gb, "unlimited")
| eval cold_bucket_usage_gb = cold_bucket_size_gb." / ".cold_bucket_capacity_gb

| eval currentDBSizeGB = round(currentDBSizeMB / 1024, 2)
| eval maxTotalDataSizeGB = if(maxTotalDataSizeMB &amp;gt; 0, round(maxTotalDataSizeMB / 1024, 2), "unlimited")
| eval disk_usage_gb = currentDBSizeGB." / ".maxTotalDataSizeGB

| eval currentTimePeriodDay = coalesce(round((now() - strptime(minTime,"%Y-%m-%dT%H:%M:%S%z")) / 86400, 0), 0)
| eval frozenTimePeriodDay = coalesce(round(frozenTimePeriodInSecs / 86400, 0), 0)
| eval frozenTimePeriodDay = if(frozenTimePeriodDay &amp;gt; 0, frozenTimePeriodDay, "unlimited")
| eval freeze_period_viz_day = currentTimePeriodDay." / ".frozenTimePeriodDay

| eval total_bucket_count = toString(coalesce(total_bucket_count, 0), "commas")
| eval totalEventCount = toString(coalesce(totalEventCount, 0), "commas")
| eval total_raw_size_gb = round(total_raw_size / 1024, 2)
| eval avg_bucket_size_gb = round(currentDBSizeGB / total_bucket_count, 2)
| eval compress_ratio = round(total_raw_size_gb / currentDBSizeGB, 2)." : 1"

| fields title, datatype
    currentDBSizeGB, totalEventCount, total_bucket_count,  avg_bucket_size_gb,
    total_raw_size_gb, compress_ratio, minTime, maxTime
    freeze_period_viz_day, disk_usage_gb, home_bucket_usage_gb, cold_bucket_usage_gb,
    hot_bucket_size_gb, warm_bucket_size_gb, cold_bucket_size_gb, thawed_bucket_size_gb,
    hot_bucket_count,   warm_bucket_count,   cold_bucket_count,   thawed_bucket_count,
    home_event_count,   cold_event_count,    thawed_event_count,
    homePath, homePath_expanded, coldPath, coldPath_expanded, thawedPath, thawedPath_expanded, summaryHomePath_expanded, tstatsHomePath, tstatsHomePath_expanded,
    maxTotalDataSizeMB, frozenTimePeriodInSecs, homePath.maxDataSizeMB, coldPath.maxDataSizeMB,
    maxDataSize, maxHotBuckets, maxWarmDBCount | search title=* | table title currentDBSizeGB total_raw_size_gb compress_ratio | where isnotnull(total_raw_size_gb) | where isnotnull(compress_ratio)
    | stats sum(currentDBSizeGB) as currentDBSizeGB, sum(total_raw_size_gb) as total_raw_size_gb | eval compress_ratio = round(total_raw_size_gb / currentDBSizeGB, 2)." : 1"&lt;/LI-CODE&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 08 Jan 2025 15:12:24 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Deployment-Architecture/Splunk-Storage-Sizing-Guidelines-and-calculations/m-p/708258#M29013</guid>
      <dc:creator>edoardo_vicendo</dc:creator>
      <dc:date>2025-01-08T15:12:24Z</dc:date>
    </item>
  </channel>
</rss>

