About sijingwu

sijingwu · ‎09-10-2019

hi, could you please explain how the findkeywords command works? I couldn't find it anywhere on the Splunk documents.

sijingwu · ‎09-10-2019

I have a case where a "Message" field contains sentences of strings, which indicated different kind of system errors. We want to use the machine learning toolkit to automatically clusters those errors into several categories. Since we are dealing with sentences, we first decided to use TFIDF to vectorize the strings, and then use the DBSCAN to do the clustering. Here is the search: index="mail" sourcetype="P1_tickets" | rex field=_raw "Message\s+:(?<Message>(.*\n)+?(?=Extra Message|Control|Log|Repeats|Via Host))" | fit TFIDF Message into message_model | fit DBSCAN Message_tf* The result is promising. We are seeing similar system errors being grouped together into the same cluster. However, the clusters are named by default 0.0, 1.0, 2.0, and etc. We want to actually use the keywords from the sentences to name the clusters, which can actually give the user some idea of what the error is. Is there anyway to achieve this in Splunk?

Posts	2
Solutions	0
Karma Given	0
Karma Received	0
Member Since	‎09-10-2019

Online Status	Offline
Date Last Visited	‎06-05-2020 02:04 AM

How to name clusters when using TFIDF and DBSCAN i...

Re: How to find most common words used by cluster ...

How to name clusters when using TFIDF and DBSCAN i...

Join the Conversation

How to name clusters when using TFIDF and DBSCAN i...

Re: How to find most common words used by cluster ...

How to name clusters when using TFIDF and DBSCAN i...