<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic How to prevent large lookups from being replicated to Yarn / Hadoop from Hunk? in Getting Data In</title>
    <link>https://community.splunk.com/t5/Getting-Data-In/How-to-prevent-large-lookups-from-being-replicated-to-Yarn/m-p/235255#M45813</link>
    <description>&lt;P&gt;We have a user who has created a large csv lookup file (600 Mb). It seems that this file is being replicated to every Yarn server to it's /tmp directory (/tmp/splunk/nodename/splunk/var/run/searchpeers) apparently as every search is executed. There are multiple copies of this bundle. This is filling the /tmp directory and causing a major problem, it also slows every search as this file has to be copied before the search begins to execute. &lt;/P&gt;

&lt;P&gt;We have specified the following in the hunk servers distsearch.conf to no effect: &lt;BR /&gt;
[replicationBlacklist] &lt;BR /&gt;
Everything = Servers.csv &lt;/P&gt;

&lt;P&gt;How do we block this file being copied with every search? &lt;BR /&gt;
Can the file be moved into an HDFS directory instead? &lt;BR /&gt;
Can it be cached so that it doesn't need to be replicated with every search? &lt;/P&gt;</description>
    <pubDate>Mon, 09 May 2016 17:07:27 GMT</pubDate>
    <dc:creator>tsunamii</dc:creator>
    <dc:date>2016-05-09T17:07:27Z</dc:date>
    <item>
      <title>How to prevent large lookups from being replicated to Yarn / Hadoop from Hunk?</title>
      <link>https://community.splunk.com/t5/Getting-Data-In/How-to-prevent-large-lookups-from-being-replicated-to-Yarn/m-p/235255#M45813</link>
      <description>&lt;P&gt;We have a user who has created a large csv lookup file (600 Mb). It seems that this file is being replicated to every Yarn server to it's /tmp directory (/tmp/splunk/nodename/splunk/var/run/searchpeers) apparently as every search is executed. There are multiple copies of this bundle. This is filling the /tmp directory and causing a major problem, it also slows every search as this file has to be copied before the search begins to execute. &lt;/P&gt;

&lt;P&gt;We have specified the following in the hunk servers distsearch.conf to no effect: &lt;BR /&gt;
[replicationBlacklist] &lt;BR /&gt;
Everything = Servers.csv &lt;/P&gt;

&lt;P&gt;How do we block this file being copied with every search? &lt;BR /&gt;
Can the file be moved into an HDFS directory instead? &lt;BR /&gt;
Can it be cached so that it doesn't need to be replicated with every search? &lt;/P&gt;</description>
      <pubDate>Mon, 09 May 2016 17:07:27 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Getting-Data-In/How-to-prevent-large-lookups-from-being-replicated-to-Yarn/m-p/235255#M45813</guid>
      <dc:creator>tsunamii</dc:creator>
      <dc:date>2016-05-09T17:07:27Z</dc:date>
    </item>
    <item>
      <title>Re: How to prevent large lookups from being replicated to Yarn / Hadoop from Hunk?</title>
      <link>https://community.splunk.com/t5/Getting-Data-In/How-to-prevent-large-lookups-from-being-replicated-to-Yarn/m-p/235256#M45814</link>
      <description>&lt;P&gt;To make sure the Table is not being copied every time:&lt;BR /&gt;
vix.splunk.setup.onsearch = 0 (default is 1)&lt;BR /&gt;
** However, that means nothing will be copied to the data nodes.  So you may want to turn it off only after the first run.&lt;/P&gt;

&lt;P&gt;To make sure you have lots of copies of the table so that bundle replication happens fast:&lt;BR /&gt;
vix.splunk.setup.bundle.replication = 20 (default 3)&lt;/P&gt;</description>
      <pubDate>Mon, 09 May 2016 18:24:30 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Getting-Data-In/How-to-prevent-large-lookups-from-being-replicated-to-Yarn/m-p/235256#M45814</guid>
      <dc:creator>rdagan_splunk</dc:creator>
      <dc:date>2016-05-09T18:24:30Z</dc:date>
    </item>
    <item>
      <title>Re: How to prevent large lookups from being replicated to Yarn / Hadoop from Hunk?</title>
      <link>https://community.splunk.com/t5/Getting-Data-In/How-to-prevent-large-lookups-from-being-replicated-to-Yarn/m-p/235257#M45815</link>
      <description>&lt;P&gt;Assuming that this is being used as a &lt;CODE&gt;lookup&lt;/CODE&gt; file, you can specify that the lookup happens only on the Search Head by adding the &lt;CODE&gt;local=t&lt;/CODE&gt; parameter as in &lt;CODE&gt;... | lookup local=t myLookup ...&lt;/CODE&gt;.  This will prevent it from being included in the bundle (replication).  The downside is that if the output fields of the lookup are used to qualify the search at any point, you will lose the benefits of having this work being map-reduced and happening on the Indexers; instead it will all happen on the Search Head.&lt;/P&gt;</description>
      <pubDate>Tue, 10 May 2016 04:24:18 GMT</pubDate>
      <guid>https://community.splunk.com/t5/Getting-Data-In/How-to-prevent-large-lookups-from-being-replicated-to-Yarn/m-p/235257#M45815</guid>
      <dc:creator>woodcock</dc:creator>
      <dc:date>2016-05-10T04:24:18Z</dc:date>
    </item>
  </channel>
</rss>

