Solved: Trim whitespace in indexed files

oscargarcia · ‎04-15-2011

Hi,

We are indexing a substantial number of XML files. These files have between 30% and 50% of white space that can be trimmed with no side effects on the real content of the file.

I was wondering wether it was possible to filter these files for removing white space (really simple regex to apply), before indexing. Can this be done on the UniversalForwarder? On the indexer?

Our aim is reducing the amount of daily indexed data as you can imagine...

Many thanks

bojanz · ‎04-15-2011

As said previously, SEDCMD is the way to go. Something like this in props.conf on the indexer:

[sourcetype]
SEDCMD-repws = s/\s+/ /g

This will match on one or more whitespace characters and replace it with one space.

View solution in original post

bojanz · ‎04-15-2011

As said previously, SEDCMD is the way to go. Something like this in props.conf on the indexer:

[sourcetype]
SEDCMD-repws = s/\s+/ /g

This will match on one or more whitespace characters and replace it with one space.

gkanapathy · ‎04-15-2011

Although, you might want something like: s/(\s)\s*/\1/g which is more likely to help preserve a line break. (While stripping off indents at the start of a line.)

Stephen_Sorkin · ‎04-15-2011

You can use the SEDCMD configuration in props.conf to replace whitespace.

http://www.splunk.com/base/Documentation/4.2/Data/Anonymizedatawithsed

dwaddle · ‎04-15-2011

You should be able to do this with a SEDCMD. (But the regex might get complicated). See the docs at http://www.splunk.com/base/Documentation/4.2/Data/Anonymizedatawithsed for info on how to configure this.

If you are using Universal or Light forwarder, the SEDCMD needs to be configured at the indexer. Your whitespace will cross the wire, but will be filtered at the indexer before it writes to the index.

Trim whitespace in indexed files

Join the Splunk Community Slack to learn, troubleshoot, and make connections with fellow Splunk practitioners in real time!

Join Splunk User Groups to connect and learn in-person by region or remotely by topic or industry.

Laser Bananas and Edge Hubs: Exploring Operational Technology (OT) Data Through a ...

Event Series: Mastering AI Tokenomics and Splunk Agent Observability

span_metrics: The OpenTelemetry-Idiomatic Way to See Inside Your Services

Join the Conversation