Solved: Re: Is data standardization and scaling unnecessar...

takaakinakajima · ‎11-15-2016

I use Machine Learning Toolkit 2.0.0 with Splunk Enterprise 6.5, and found implementations of algorithms in
SPLUNK_HOME/etc/apps/Splunk_ML_Toolkit/bin/algos

Only SGDClassifier, SGDRegressor, and SpectralClustering algorithms,
data will be scaled with StandardScaler before calculation.
It seems that the other algorithms (e.g. LenearRegression) do not scale data.

Is scaling unnecessary with Splunk/Machine Learning Toolkit?
If required, how do we standardize data before calculation?

Scikit-learn notes "Standardization of datasets is a common requirement for many machine learning estimators".
http://scikit-learn.org/stable/modules/preprocessing.html

grana_splunk · ‎11-16-2016

Simply use StandardScaler, if you want to scale your data

For example: ,... | fit StandardScaler ... | fit LinearRegression ...

View solution in original post

grana_splunk · ‎11-16-2016

Simply use StandardScaler, if you want to scale your data

For example: ,... | fit StandardScaler ... | fit LinearRegression ...

takaakinakajima · ‎11-16-2016

Hi grana.

Thank you for your shrewd advice.
That's just the thing!!

Is data standardization and scaling unnecessary with Splunk and the Machine Learning Toolkit?

[Puzzles] Solve, Learn, Repeat: Dynamic formatting from XML events

Enter the Agentic Era with Splunk AI Assistant for SPL 1.4

Stronger Security with Federated Search for S3, GCP SQL & Australian Threat ...

Join the Conversation

Is data standardization and scaling unnecessary with Splunk and the Machine Learning Toolkit?

[Puzzles] Solve, Learn, Repeat: Dynamic formatting from XML events

Enter the Agentic Era with Splunk AI Assistant for SPL 1.4

Stronger Security with Federated Search for S3, GCP SQL & Australian Threat ...