<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: splunk indexing the same files again and again and again and ... in Getting Data In</title>
    <link>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95572#M19905</link>
    <description>&lt;P&gt;You might want CRC salts.  The most common usage of this is to include the file path as part of the CRC used for Splunk to answer the question "Has this data already been indexed?"  See docs on inputs.conf &lt;A href="http://docs.splunk.com/Documentation/Splunk/4.3.3/admin/Inputsconf"&gt;here&lt;/A&gt;.&lt;/P&gt;

&lt;P&gt;I wonder also whether these files are being rotated daily, and if so, are they immediately compressed?  If Splunk sees that the base file it's reading doesn't have the same CRC as the last time it looked, it will attempt to read forward in the file to find a line which matches the last one it saw.  If it doesn't, it believes that the file is completely new (these are the 'seekptr' messages).  Furthermore, Splunk doesn't always do a good job of figuring out where it left off if the rotated version of the file has been compressed.  It's best if you can leave the "most-recently rotated" version uncompressed, and then compress on the next rotation cycle.&lt;/P&gt;

&lt;P&gt;Some more information can be found &lt;A href="http://docs.splunk.com/Documentation/Splunk/4.3.3/Data/HowLogFileRotationIsHandled#How_to_work_with_log_rotation_into_compressed_files"&gt;here&lt;/A&gt;.&lt;/P&gt;</description>
    <pubDate>Mon, 23 Jul 2012 12:48:10 GMT</pubDate>
    <dc:creator>sowings</dc:creator>
    <dc:date>2012-07-23T12:48:10Z</dc:date>
    <item>
      <title>splunk indexing the same files again and again and again and ...</title>
      <link>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95567#M19900</link>
      <description>&lt;P&gt;I have a Splunk universal forwarder on a client machine. I have a deployed app that looks like this..&lt;/P&gt;

&lt;PRE&gt;&lt;CODE&gt;  [monitor:///export/home/storeadm/r*]
  disabled = true
  followTail = 0
  index = contentkeeper
  source = contentkeeper_passed
  sourcetype = contentkeeper_passed
  whitelist = (/r.*\.csv.gz$|/r.*\.csv$)
&lt;/CODE&gt;&lt;/PRE&gt;

&lt;P&gt;On the indexer there is a corresponding props.conf entry&lt;/P&gt;

&lt;PRE&gt;&lt;CODE&gt;  [contentkeeper_passed]
  REPORT-ckpassed = ckpassed_extractions
&lt;/CODE&gt;&lt;/PRE&gt;

&lt;P&gt;And a corresponding transforms.conf entry&lt;/P&gt;

&lt;PRE&gt;&lt;CODE&gt;  [ckpassed_extractions]
  DELIMS=","
  FIELDS="Time","Category","IP-Address","Username","Bytes","Status","Content-Type","Url","Policy","Category-Description"
&lt;/CODE&gt;&lt;/PRE&gt;

&lt;P&gt;The data files are all compressed (.csv.gz) so the second whitelist match is superfluous. There are a few months of data sitting in that directory.&lt;/P&gt;

&lt;P&gt;The volume of data is quite small (only 10s of MB per day). PS: sorry about the timestamps. I touched the files as a test, but usually the files have an incrementing daily timestamp.&lt;/P&gt;

&lt;PRE&gt;&lt;CODE&gt;-rw-r--r--   1 storeadm storeadm    1.3M Mar 15 15:41 r29-12-2011.csv.gz
-rw-r--r--   1 storeadm storeadm     38M Mar 15 15:41 r30-01-2012.csv.gz
-rw-r--r--   1 storeadm storeadm    2.5M Mar 15 15:41 r30-10-2011.csv.gz
-rw-r--r--   1 storeadm storeadm     44M Mar 15 15:41 r30-11-2011.csv.gz
-rw-r--r--   1 storeadm storeadm    781K Mar 15 15:41 r30-12-2011.csv.gz
&lt;/CODE&gt;&lt;/PRE&gt;

&lt;P&gt;However my license quota is often exceeded, typically more than 20GB (that's GIGABYTES!) per day. I don't think it's the months of data that's the problem. The entire directory is only 3.6GB.&lt;/P&gt;

&lt;PRE&gt;&lt;CODE&gt; $ du -sh /export/home/storeadm/
 3.6G   /export/home/storeadm
&lt;/CODE&gt;&lt;/PRE&gt;

&lt;P&gt;I think the problem is Splunk is re-indexing the same files.&lt;/P&gt;

&lt;PRE&gt;&lt;CODE&gt;  $  grep "reading path" splunkd.log | awk '{print $8}' | sort | uniq -c
  ...
   4 path=/export/home/storeadm/r30-10-2011.csv.gz
   2 path=/export/home/storeadm/r30-11-2011.csv.gz
   6 path=/export/home/storeadm/r30-12-2011.csv.gz
   2 path=/export/home/storeadm/r31-10-2011.csv.gz
   2 path=/export/home/storeadm/r31-12-2011.csv.gz
&lt;/CODE&gt;&lt;/PRE&gt;

&lt;P&gt;These are the kinds of entries I'm grepping over.&lt;/P&gt;

&lt;PRE&gt;&lt;CODE&gt;  03-17-2012 06:42:20.258 +1100 INFO  ArchiveProcessor - handling file=/export/home/storeadm/r09-03-2012.csv.gz
  03-17-2012 06:42:20.295 +1100 INFO  ArchiveProcessor - reading path=/export/home/storeadm/r09-03-2012.csv.gz (seek=0 len=50548338)
  03-17-2012 07:28:09.552 +1100 INFO  ArchiveProcessor - Finished processing file '/export/home/storeadm/r09-03-2012.csv.gz', removing from stats
&lt;/CODE&gt;&lt;/PRE&gt;

&lt;P&gt;What should I do to check whether Splunk is re-indexing the same files, contributing to my license problem? Is there some search I can run over the metrics index? &lt;/P&gt;</description>
      <pubDate>Sat, 17 Mar 2012 00:28:59 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95567#M19900</guid>
      <dc:creator>nathanh42</dc:creator>
      <dc:date>2012-03-17T00:28:59Z</dc:date>
    </item>
    <item>
      <title>Re: splunk indexing the same files again and again and again and ...</title>
      <link>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95568#M19901</link>
      <description>&lt;P&gt;check to make sure you dont have multiple sources set to the same path, etc.&lt;/P&gt;

&lt;P&gt;maybe turn to strace or inotify or lsof ?&lt;/P&gt;

&lt;P&gt;"watch lsof /export/home/storeadm"&lt;/P&gt;

&lt;P&gt;your du says 3.6GB of zip, but you say its indexing 20GB. can you verify total data size with "gzip -l *" in that dir.&lt;/P&gt;

&lt;P&gt;seems like others also having a reindex problem, see &lt;A href="http://splunk-base.splunk.com/answers/43076/why-are-my-logfiles-re-indexing-due-to-a-failed-seekptr-checksum"&gt;http://splunk-base.splunk.com/answers/43076/why-are-my-logfiles-re-indexing-due-to-a-failed-seekptr-checksum&lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Sat, 17 Mar 2012 01:11:20 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95568#M19901</guid>
      <dc:creator>cvajs</dc:creator>
      <dc:date>2012-03-17T01:11:20Z</dc:date>
    </item>
    <item>
      <title>Re: splunk indexing the same files again and again and again and ...</title>
      <link>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95569#M19902</link>
      <description>&lt;BLOCKQUOTE&gt;
&lt;P&gt;check to make sure you dont have multiple sources set to the same path, etc.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;

&lt;P&gt;There's only a single source. Confirmed with the command "splunk list monitor".&lt;/P&gt;

&lt;BLOCKQUOTE&gt;
&lt;P&gt;your du says 3.6GB of zip, but you say its indexing 20GB. can you verify total data size with "gzip -l *" in that dir.&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;

&lt;P&gt;Uncompressed size is 37GB. &lt;/P&gt;

&lt;P&gt;&lt;CODE&gt;3915683331         37491057664  89.6% (totals)&lt;/CODE&gt;&lt;/P&gt;

&lt;P&gt;However the Splunk logs show multiple "reading path" statements for the same files. If it was only 37GB I could live with that. The problem is it keeps going back to all the previous files it has already indexed and indexing them again!&lt;/P&gt;</description>
      <pubDate>Sun, 18 Mar 2012 22:22:20 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95569#M19902</guid>
      <dc:creator>nathanh42</dc:creator>
      <dc:date>2012-03-18T22:22:20Z</dc:date>
    </item>
    <item>
      <title>Re: splunk indexing the same files again and again and again and ...</title>
      <link>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95570#M19903</link>
      <description>&lt;P&gt;do those gz's change at all over time?&lt;/P&gt;</description>
      <pubDate>Sun, 18 Mar 2012 22:35:31 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95570#M19903</guid>
      <dc:creator>cvajs</dc:creator>
      <dc:date>2012-03-18T22:35:31Z</dc:date>
    </item>
    <item>
      <title>Re: splunk indexing the same files again and again and again and ...</title>
      <link>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95571#M19904</link>
      <description>&lt;P&gt;Same problem here...Filezilla FTP server logs mounted from a remote Windows system to the local splunk server keeps indexing the same files continiously...&lt;/P&gt;

&lt;P&gt;Get the "WatchedFile - Checksum for seekptr didn't match, will re-read entire file" in my log too.&lt;/P&gt;

&lt;P&gt;FG&lt;/P&gt;</description>
      <pubDate>Mon, 23 Jul 2012 08:42:15 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95571#M19904</guid>
      <dc:creator>fgilain</dc:creator>
      <dc:date>2012-07-23T08:42:15Z</dc:date>
    </item>
    <item>
      <title>Re: splunk indexing the same files again and again and again and ...</title>
      <link>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95572#M19905</link>
      <description>&lt;P&gt;You might want CRC salts.  The most common usage of this is to include the file path as part of the CRC used for Splunk to answer the question "Has this data already been indexed?"  See docs on inputs.conf &lt;A href="http://docs.splunk.com/Documentation/Splunk/4.3.3/admin/Inputsconf"&gt;here&lt;/A&gt;.&lt;/P&gt;

&lt;P&gt;I wonder also whether these files are being rotated daily, and if so, are they immediately compressed?  If Splunk sees that the base file it's reading doesn't have the same CRC as the last time it looked, it will attempt to read forward in the file to find a line which matches the last one it saw.  If it doesn't, it believes that the file is completely new (these are the 'seekptr' messages).  Furthermore, Splunk doesn't always do a good job of figuring out where it left off if the rotated version of the file has been compressed.  It's best if you can leave the "most-recently rotated" version uncompressed, and then compress on the next rotation cycle.&lt;/P&gt;

&lt;P&gt;Some more information can be found &lt;A href="http://docs.splunk.com/Documentation/Splunk/4.3.3/Data/HowLogFileRotationIsHandled#How_to_work_with_log_rotation_into_compressed_files"&gt;here&lt;/A&gt;.&lt;/P&gt;</description>
      <pubDate>Mon, 23 Jul 2012 12:48:10 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95572#M19905</guid>
      <dc:creator>sowings</dc:creator>
      <dc:date>2012-07-23T12:48:10Z</dc:date>
    </item>
    <item>
      <title>Re: splunk indexing the same files again and again and again and ...</title>
      <link>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95573#M19906</link>
      <description>&lt;P&gt;No compression here.&lt;BR /&gt;
The Filezilla FTP server creates everyday a new log file named :&lt;/P&gt;

&lt;P&gt;fzs-YYYY-MM-DD.log&lt;/P&gt;

&lt;P&gt;so my monitored directory mounted on the splunk server contains files like :&lt;/P&gt;

&lt;P&gt;fzs-2012-07-23.log&lt;BR /&gt;
fzs-2012-07-22.log&lt;BR /&gt;
fzs-2012-07-21.log&lt;BR /&gt;
fzs-2012-07-20.log&lt;BR /&gt;
...&lt;BR /&gt;
..&lt;/P&gt;

&lt;P&gt;My "/splunk/splunk/etc/apps/search/local/inputs.conf" file contains the following section :&lt;/P&gt;

&lt;P&gt;[monitor:///mnt/s-ftpde-01/filezilla/*.log]&lt;BR /&gt;
disabled = false&lt;BR /&gt;
followTail = 0&lt;BR /&gt;
sourcetype = LOGS-FILEZILLA-SERVEUR&lt;BR /&gt;
index = index_de_filezilla&lt;BR /&gt;
host = s-ftpde-01.de.lan&lt;BR /&gt;
host_segment =&lt;BR /&gt;
crcSalt = &lt;SOURCE&gt;&lt;BR /&gt;
whitelist = [^/]*.log$&lt;/SOURCE&gt;&lt;/P&gt;

&lt;P&gt;NB : what could my "crcSalt" parameter be setup to in order to avoid re-indexing ?&lt;BR /&gt;
NB2 : Should i use the "ignoreOlderThan" parameter too (with a 1 day value ?) ?&lt;/P&gt;

&lt;P&gt;FG&lt;/P&gt;</description>
      <pubDate>Mon, 28 Sep 2020 12:07:55 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95573#M19906</guid>
      <dc:creator>fgilain</dc:creator>
      <dc:date>2020-09-28T12:07:55Z</dc:date>
    </item>
    <item>
      <title>Re: splunk indexing the same files again and again and again and ...</title>
      <link>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95574#M19907</link>
      <description>&lt;P&gt;FG,&lt;/P&gt;

&lt;P&gt;I'm having the exact same problem as you are.  I was wondering if you ever found a solution.  I have a ticket open with Splunk, but they haven't been able to get a solution for me yet.  Our situation is also similar in that the log files are not on local disk, mine are mounted by NFS.&lt;/P&gt;

&lt;P&gt;Thanks,&lt;/P&gt;

&lt;P&gt;-MD&lt;/P&gt;</description>
      <pubDate>Wed, 12 Dec 2012 15:37:21 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95574#M19907</guid>
      <dc:creator>mdurkin</dc:creator>
      <dc:date>2012-12-12T15:37:21Z</dc:date>
    </item>
    <item>
      <title>Re: splunk indexing the same files again and again and again and ...</title>
      <link>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95575#M19908</link>
      <description>&lt;P&gt;I finally had to use a splunk forwarder from my source server instead of remote mounting the share with logs...all works now.&lt;/P&gt;</description>
      <pubDate>Wed, 12 Dec 2012 18:13:02 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Getting-Data-In/splunk-indexing-the-same-files-again-and-again-and-again-and/m-p/95575#M19908</guid>
      <dc:creator>fgilain</dc:creator>
      <dc:date>2012-12-12T18:13:02Z</dc:date>
    </item>
  </channel>
</rss>

