<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: How to log whole site content whilst excluding specific file extensions and file types in Getting Data In</title>
    <link>https://community.splunk.com/t5/Getting-Data-In/How-to-log-whole-site-content-whilst-excluding-specific-file/m-p/485149#M87767</link>
    <description>&lt;P&gt;Hi @lkm93,&lt;BR /&gt;
at first, you don't need the waf_include stanza, but I usually insert it!&lt;BR /&gt;
Then, you don't need * in &lt;CODE&gt;REGEX = .*&lt;/CODE&gt;, you can use &lt;CODE&gt;REGEX = .&lt;/CODE&gt;.&lt;/P&gt;

&lt;P&gt;Then you don't need the include stanzas whan you have &lt;CODE&gt;REGEX = .&lt;/CODE&gt;, because you already have all that you didn't discard, so try something like this:&lt;BR /&gt;
in &lt;STRONG&gt;props.conf&lt;/STRONG&gt;&lt;/P&gt;

&lt;PRE&gt;&lt;CODE&gt; [waf_log]
 pulldown_type = true
 MAX_TIMESTAMP_LOOKAHEAD = 32
 SHOULD_LINEMERGE = False
TRANSFORMS-null = waf_include,waf_exclude
 LEARN_SOURCETYPE = false
 TZ = GMT
&lt;/CODE&gt;&lt;/PRE&gt;

&lt;P&gt;in &lt;STRONG&gt;transforms.conf&lt;/STRONG&gt;&lt;/P&gt;

&lt;PRE&gt;&lt;CODE&gt; [waf_include]
 DEST_KEY = queue
 FORMAT = indexQueue
 REGEX = .*

 [waf_exclude]
 DEST_KEY = queue
 FORMAT = nullQueue
 REGEX = .*\.(tif|mp3|jpg|js|css|java|Ico|waf|png|gif|svg|jpeg|avi|mid|midi|mpg|mpeg|mov|qt|png|ram|rar|tiff|txt|wav|zip|TIF|MP3|CSS|JAVA|ICO|WAF|PNG|SVG|AVI|CSS|EXE|GIF|JPG|JS|JPEG|MID|MIDI|MPG|MPEG|MOV|QT|PNG|RAM|RAR|TIFF|TXT|WAV|ZIP).*
&lt;/CODE&gt;&lt;/PRE&gt;

&lt;P&gt;Ciao.&lt;BR /&gt;
Giuseppe&lt;/P&gt;</description>
    <pubDate>Mon, 20 Jan 2020 15:34:49 GMT</pubDate>
    <dc:creator>gcusello</dc:creator>
    <dc:date>2020-01-20T15:34:49Z</dc:date>
    <item>
      <title>How to log whole site content whilst excluding specific file extensions and file types</title>
      <link>https://community.splunk.com/t5/Getting-Data-In/How-to-log-whole-site-content-whilst-excluding-specific-file/m-p/485146#M87764</link>
      <description>&lt;P&gt;(I am an absolute novice at this, the answer maybe obvious but I am still learning the trade please bear with me)&lt;/P&gt;

&lt;P&gt;For this exercise I am trying to index the whole site e.g &lt;A href="http://www.lkm93.com"&gt;www.lkm93.com&lt;/A&gt; whilst avoiding massive file names that may cause my daily indexing allowance to go over the limit. &lt;/P&gt;

&lt;P&gt;The regex I have figured out so far is: &lt;/P&gt;

&lt;PRE&gt;&lt;CODE&gt;[waf_exclude]
DEST_KEY = queue 
FORMAT  = nullQueue 
REGEX = .*\(tif|mp3|jpg|js|css|mp4|java|waf|png|gif|svg|jpeg|JPG|JS|JPEG|MID|MIDI|MP3|MP4|MPG|MPEG|PDF|PNG|TIFF|TXT|WAV|ZIP)
&lt;/CODE&gt;&lt;/PRE&gt;

&lt;P&gt;(I have repeated some extensions is capitals letters to make sure I match the extensions in both cases)&lt;/P&gt;

&lt;P&gt;This I believe should be indexing everything on my site &lt;A href="http://www.lkm93.com"&gt;www.lkm93.com&lt;/A&gt; and the regex I have added to that will exclude the file named file extensions. I have reloaded the transforms.conf file and I don't seem to be pulling in data outside of what I am already pulling in. Is there anything obvious that I could be missing here?&lt;/P&gt;</description>
      <pubDate>Mon, 20 Jan 2020 10:25:44 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Getting-Data-In/How-to-log-whole-site-content-whilst-excluding-specific-file/m-p/485146#M87764</guid>
      <dc:creator>lkm93</dc:creator>
      <dc:date>2020-01-20T10:25:44Z</dc:date>
    </item>
    <item>
      <title>Re: How to log whole site content whilst excluding specific file extensions and file types</title>
      <link>https://community.splunk.com/t5/Getting-Data-In/How-to-log-whole-site-content-whilst-excluding-specific-file/m-p/485147#M87765</link>
      <description>&lt;P&gt;Hi @lkm93,&lt;BR /&gt;
at first, I think that you used also props.conf adding:&lt;/P&gt;

&lt;PRE&gt;&lt;CODE&gt;[your_sourcetype]
TRANSFORMS-waf_exclude = waf_exclude
&lt;/CODE&gt;&lt;/PRE&gt;

&lt;P&gt;Then, where do you inserted props.conf and transforms.conf? they must be on Indexers or (when present) on Heavy Forwarders.&lt;/P&gt;

&lt;P&gt;Then, do you restarted Splunk after modifying props.conf and transfrorms.conf?&lt;/P&gt;

&lt;P&gt;Then, you didn't escaped the last parenthesis? the correct regex is &lt;CODE&gt;.*\(tif|mp3|jpg|js|css|mp4|java|waf|png|gif|svg|jpeg|JPG|JS|JPEG|MID|MIDI|MP3|MP4|MPG|MPEG|PDF|PNG|TIFF|TXT|WAV|ZI\P)&lt;/CODE&gt;&lt;/P&gt;

&lt;P&gt;At least, check your regex using the regex command:&lt;/P&gt;

&lt;PRE&gt;&lt;CODE&gt;index=your_index
| regex ".*\(tif|mp3|jpg|js|css|mp4|java|waf|png|gif|svg|jpeg|JPG|JS|JPEG|MID|MIDI|MP3|MP4|MPG|MPEG|PDF|PNG|TIFF|TXT|WAV|ZIP\)"
&lt;/CODE&gt;&lt;/PRE&gt;

&lt;P&gt;Finally I saw that there are extension in uppercase not present in lowercase or reverse (ZIP, css, etc...).&lt;/P&gt;

&lt;P&gt;Ciao.&lt;BR /&gt;
Giuseppe&lt;/P&gt;</description>
      <pubDate>Mon, 20 Jan 2020 13:24:56 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Getting-Data-In/How-to-log-whole-site-content-whilst-excluding-specific-file/m-p/485147#M87765</guid>
      <dc:creator>gcusello</dc:creator>
      <dc:date>2020-01-20T13:24:56Z</dc:date>
    </item>
    <item>
      <title>Re: How to log whole site content whilst excluding specific file extensions and file types</title>
      <link>https://community.splunk.com/t5/Getting-Data-In/How-to-log-whole-site-content-whilst-excluding-specific-file/m-p/485148#M87766</link>
      <description>&lt;P&gt;Hello Giuseppe, &lt;/P&gt;

&lt;P&gt;thank you for your prompt reply. &lt;/P&gt;

&lt;P&gt;I have re-arranged my props.conf file after reading your reply and also re-configured the transforms.conf file. &lt;/P&gt;

&lt;P&gt;Here'show my props.conf file looks now:&lt;/P&gt;

&lt;PRE&gt;&lt;CODE&gt;[waf_log]
pulldown_type = true
MAX_TIMESTAMP_LOOKAHEAD = 32
SHOULD_LINEMERGE = False
TRANSFORMS-null = waf_include,waf_exclude,waf_include_xapi,waf_drop_x
LEARN_SOURCETYPE = false
TZ = GMT
&lt;/CODE&gt;&lt;/PRE&gt;

&lt;P&gt;Transforms.conf looks like this:&lt;/P&gt;

&lt;PRE&gt;&lt;CODE&gt;[waf_include]
DEST_KEY = queue
FORMAT = indexQueue
REGEX = .*

[waf_exclude]
DEST_KEY = queue
FORMAT = nullQueue
REGEX = .*\.(tif|mp3|jpg|js|css|java|Ico|waf|png|gif|svg|jpeg|avi|mid|midi|mpg|mpeg|mov|qt|png|ram|rar|tiff|txt|wav|zip|TIF|MP3|CSS|JAVA|ICO|WAF|PNG|SVG|AVI|CSS|EXE|GIF|JPG|JS|JPEG|MID|MIDI|MPG|MPEG|MOV|QT|PNG|RAM|RAR|TIFF|TXT|WAV|ZIP).*

[waf_include_xapi]
DEST_KEY = queue
FORMAT = indexQueue
REGEX = blah-blah

[waf_drop_x]
DEST_KEY = queue
FORMAT = nullQueue
REGEX = blahblah
&lt;/CODE&gt;&lt;/PRE&gt;

&lt;P&gt;My  props.conf and transforms.conf files are on the Splunk manager, I thought that would be the reasonable place to have them. &lt;/P&gt;

&lt;P&gt;I also discovered that by &lt;A href="https://splunk-fqdn/en-US/debug/refresh"&gt;https://splunk-fqdn/en-US/debug/refresh&lt;/A&gt; I could refresh the all the .conf files. Do I definitely need to restart Splunk based on the new changes I have just made?&lt;/P&gt;

&lt;P&gt;And lastly I have fixed the Regex to pick up whole urls on that domain, it's picking up everything I needs in the test I have done. also the extensions have been fixed I was in a rush to get the question out to the world..thank you!&lt;/P&gt;

&lt;P&gt;What do you think of this now?&lt;/P&gt;</description>
      <pubDate>Mon, 20 Jan 2020 15:04:46 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Getting-Data-In/How-to-log-whole-site-content-whilst-excluding-specific-file/m-p/485148#M87766</guid>
      <dc:creator>lkm93</dc:creator>
      <dc:date>2020-01-20T15:04:46Z</dc:date>
    </item>
    <item>
      <title>Re: How to log whole site content whilst excluding specific file extensions and file types</title>
      <link>https://community.splunk.com/t5/Getting-Data-In/How-to-log-whole-site-content-whilst-excluding-specific-file/m-p/485149#M87767</link>
      <description>&lt;P&gt;Hi @lkm93,&lt;BR /&gt;
at first, you don't need the waf_include stanza, but I usually insert it!&lt;BR /&gt;
Then, you don't need * in &lt;CODE&gt;REGEX = .*&lt;/CODE&gt;, you can use &lt;CODE&gt;REGEX = .&lt;/CODE&gt;.&lt;/P&gt;

&lt;P&gt;Then you don't need the include stanzas whan you have &lt;CODE&gt;REGEX = .&lt;/CODE&gt;, because you already have all that you didn't discard, so try something like this:&lt;BR /&gt;
in &lt;STRONG&gt;props.conf&lt;/STRONG&gt;&lt;/P&gt;

&lt;PRE&gt;&lt;CODE&gt; [waf_log]
 pulldown_type = true
 MAX_TIMESTAMP_LOOKAHEAD = 32
 SHOULD_LINEMERGE = False
TRANSFORMS-null = waf_include,waf_exclude
 LEARN_SOURCETYPE = false
 TZ = GMT
&lt;/CODE&gt;&lt;/PRE&gt;

&lt;P&gt;in &lt;STRONG&gt;transforms.conf&lt;/STRONG&gt;&lt;/P&gt;

&lt;PRE&gt;&lt;CODE&gt; [waf_include]
 DEST_KEY = queue
 FORMAT = indexQueue
 REGEX = .*

 [waf_exclude]
 DEST_KEY = queue
 FORMAT = nullQueue
 REGEX = .*\.(tif|mp3|jpg|js|css|java|Ico|waf|png|gif|svg|jpeg|avi|mid|midi|mpg|mpeg|mov|qt|png|ram|rar|tiff|txt|wav|zip|TIF|MP3|CSS|JAVA|ICO|WAF|PNG|SVG|AVI|CSS|EXE|GIF|JPG|JS|JPEG|MID|MIDI|MPG|MPEG|MOV|QT|PNG|RAM|RAR|TIFF|TXT|WAV|ZIP).*
&lt;/CODE&gt;&lt;/PRE&gt;

&lt;P&gt;Ciao.&lt;BR /&gt;
Giuseppe&lt;/P&gt;</description>
      <pubDate>Mon, 20 Jan 2020 15:34:49 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Getting-Data-In/How-to-log-whole-site-content-whilst-excluding-specific-file/m-p/485149#M87767</guid>
      <dc:creator>gcusello</dc:creator>
      <dc:date>2020-01-20T15:34:49Z</dc:date>
    </item>
    <item>
      <title>Re: How to log whole site content whilst excluding specific file extensions and file types</title>
      <link>https://community.splunk.com/t5/Getting-Data-In/How-to-log-whole-site-content-whilst-excluding-specific-file/m-p/485150#M87768</link>
      <description>&lt;P&gt;Hi @gcusello &lt;/P&gt;

&lt;P&gt;Thank you for thi si applied this configuration and it seems to be working as you described! no longer picking up the unwanted extensions. &lt;/P&gt;</description>
      <pubDate>Thu, 13 Feb 2020 16:30:24 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Getting-Data-In/How-to-log-whole-site-content-whilst-excluding-specific-file/m-p/485150#M87768</guid>
      <dc:creator>lkm93</dc:creator>
      <dc:date>2020-02-13T16:30:24Z</dc:date>
    </item>
  </channel>
</rss>

