layout | title | displayTitle | license |
---|---|---|---|
global |
Linear Methods - RDD-based API |
Linear Methods - RDD-based API |
Licensed to the Apache Software Foundation (ASF) under one or more
contributor license agreements. See the NOTICE file distributed with
this work for additional information regarding copyright ownership.
The ASF licenses this file to You under the Apache License, Version 2.0
(the "License"); you may not use this file except in compliance with
the License. You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
|
- Table of contents {:toc}
\[ \newcommand{\R}{\mathbb{R}} \newcommand{\E}{\mathbb{E}} \newcommand{\x}{\mathbf{x}} \newcommand{\y}{\mathbf{y}} \newcommand{\wv}{\mathbf{w}} \newcommand{\av}{\mathbf{\alpha}} \newcommand{\bv}{\mathbf{b}} \newcommand{\N}{\mathbb{N}} \newcommand{\id}{\mathbf{I}} \newcommand{\ind}{\mathbf{1}} \newcommand{\0}{\mathbf{0}} \newcommand{\unit}{\mathbf{e}} \newcommand{\one}{\mathbf{1}} \newcommand{\zero}{\mathbf{0}} \]
Many standard machine learning methods can be formulated as a convex optimization problem, i.e.
the task of finding a minimizer of a convex function $f$
that depends on a variable vector
$\wv$
(called weights
in the code), which has $d$
entries.
Formally, we can write this as the optimization problem $\min_{\wv \in\R^d} \; f(\wv)$
, where
the objective function is of the form
\begin{equation} f(\wv) := \lambda\, R(\wv) + \frac1n \sum_{i=1}^n L(\wv;\x_i,y_i) \label{eq:regPrimal} \ . \end{equation}
Here the vectors $\x_i\in\R^d$
are the training data examples, for $1\le i\le n$
, and
$y_i\in\R$
are their corresponding labels, which we want to predict.
We call the method linear if spark.mllib
's classification and regression algorithms fall into this category,
and are discussed here.
The objective function $f$
has two parts:
the regularizer that controls the complexity of the model,
and the loss that measures the error of the model on the training data.
The loss function $L(\wv;.)$
is typically a convex function in $\wv$
. The
fixed regularization parameter $\lambda \ge 0$
(regParam
in the code)
defines the trade-off between the two goals of minimizing the loss (i.e.,
training error) and minimizing model complexity (i.e., to avoid overfitting).
The following table summarizes the loss functions and their gradients or sub-gradients for the
methods spark.mllib
supports:
loss function |
gradient or sub-gradient | |
---|---|---|
hinge loss | $\begin{cases}-y \cdot \x & \text{if |
|
logistic loss | ||
squared loss |
Note that, in the mathematical formulation above, a binary label spark.mllib
instead of
The purpose of the
regularizer is to
encourage simple models and avoid overfitting. We support the following
regularizers in spark.mllib
:
regularizer |
gradient or sub-gradient | |
---|---|---|
zero (unregularized) | 0 | |
L2 | ||
L1 | ||
elastic net |
Here $\mathrm{sign}(\wv)$
is the vector consisting of the signs ($\pm1$
) of all the entries
of $\wv$
.
L2-regularized problems are generally easier to solve than L1-regularized due to smoothness. However, L1 regularization can help promote sparsity in weights leading to smaller and more interpretable models, the latter of which can be useful for feature selection. Elastic net is a combination of L1 and L2 regularization. It is not recommended to train models without any regularization, especially when the number of training examples is small.
Under the hood, linear methods use convex optimization methods to optimize the objective functions.
spark.mllib
uses two methods, SGD and L-BFGS, described in the optimization section.
Currently, most algorithm APIs support Stochastic Gradient Descent (SGD), and a few support L-BFGS.
Refer to this optimization section for guidelines on choosing between optimization methods.
Classification aims to divide items into
categories.
The most common classification type is
binary classification, where there are two
categories, usually named positive and negative.
If there are more than two categories, it is called
multiclass classification.
spark.mllib
supports two linear methods for classification: linear Support Vector Machines (SVMs)
and logistic regression.
Linear SVMs supports only binary classification, while logistic regression supports both binary and
multiclass classification problems.
For both methods, spark.mllib
supports L1 and L2 regularized variants.
The training data set is represented by an RDD of LabeledPoint in MLlib,
where labels are class indices starting from zero:
The linear SVM
is a standard method for large-scale classification tasks. It is a linear method as described above in equation $\eqref{eq:regPrimal}$
, with the loss function in the formulation given by the hinge loss:
\[ L(\wv;\x,y) := \max \{0, 1-y \wv^T \x \}. \]
By default, linear SVMs are trained with an L2 regularization.
We also support alternative L1 regularization. In this case,
the problem becomes a linear program.
The linear SVMs algorithm outputs an SVM model. Given a new data point,
denoted by
Examples
Refer to the SVMWithSGD
Scala docs and SVMModel
Scala docs for details on the API.
{% include_example scala/org/apache/spark/examples/mllib/SVMWithSGDExample.scala %}
The SVMWithSGD.train()
method by default performs L2 regularization with the
regularization parameter set to 1.0. If we want to configure this algorithm, we
can customize SVMWithSGD
further by creating a new object directly and
calling setter methods. All other spark.mllib
algorithms support customization in
this way as well. For example, the following code produces an L1 regularized
variant of SVMs with regularization parameter set to 0.1, and runs the training
algorithm for 200 iterations.
{% highlight scala %}
import org.apache.spark.mllib.optimization.L1Updater
val svmAlg = new SVMWithSGD() svmAlg.optimizer .setNumIterations(200) .setRegParam(0.1) .setUpdater(new L1Updater) val modelL1 = svmAlg.run(training) {% endhighlight %}
Refer to the SVMWithSGD
Java docs and SVMModel
Java docs for details on the API.
{% include_example java/org/apache/spark/examples/mllib/JavaSVMWithSGDExample.java %}
The SVMWithSGD.train()
method by default performs L2 regularization with the
regularization parameter set to 1.0. If we want to configure this algorithm, we
can customize SVMWithSGD
further by creating a new object directly and
calling setter methods. All other spark.mllib
algorithms support customization in
this way as well. For example, the following code produces an L1 regularized
variant of SVMs with regularization parameter set to 0.1, and runs the training
algorithm for 200 iterations.
{% highlight java %} import org.apache.spark.mllib.optimization.L1Updater;
SVMWithSGD svmAlg = new SVMWithSGD(); svmAlg.optimizer() .setNumIterations(200) .setRegParam(0.1) .setUpdater(new L1Updater()); SVMModel modelL1 = svmAlg.run(training.rdd()); {% endhighlight %}
In order to run the above application, follow the instructions provided in the Self-Contained Applications section of the Spark quick-start guide. Be sure to also include spark-mllib to your build file as a dependency.
Refer to the SVMWithSGD
Python docs and SVMModel
Python docs for more details on the API.
{% include_example python/mllib/svm_with_sgd_example.py %}
Logistic regression is widely used to predict a
binary response. It is a linear method as described above in equation $\eqref{eq:regPrimal}$
,
with the loss function in the formulation given by the logistic loss:
\[ L(\wv;\x,y) := \log(1+\exp( -y \wv^T \x)). \]
For binary classification problems, the algorithm outputs a binary logistic regression model.
Given a new data point, denoted by \[ \mathrm{f}(z) = \frac{1}{1 + e^{-z}} \]
where
Binary logistic regression can be generalized into
multinomial logistic regression to
train and predict multiclass classification problems.
For example, for spark.mllib
, the first class
For multiclass classification problems, the algorithm will output a multinomial logistic regression
model, which contains
We implemented two algorithms to solve logistic regression: mini-batch gradient descent and L-BFGS. We recommend L-BFGS over mini-batch gradient descent for faster convergence.
Examples
Refer to the LogisticRegressionWithLBFGS
Scala docs and LogisticRegressionModel
Scala docs for details on the API.
{% include_example scala/org/apache/spark/examples/mllib/LogisticRegressionWithLBFGSExample.scala %}
Refer to the LogisticRegressionWithLBFGS
Java docs and LogisticRegressionModel
Java docs for details on the API.
{% include_example java/org/apache/spark/examples/mllib/JavaLogisticRegressionWithLBFGSExample.java %}
Note that the Python API does not yet support multiclass classification and model save/load but will in the future.
Refer to the LogisticRegressionWithLBFGS
Python docs and LogisticRegressionModel
Python docs for more details on the API.
{% include_example python/mllib/logistic_regression_with_lbfgs_example.py %}
Linear least squares is the most common formulation for regression problems.
It is a linear method as described above in equation $\eqref{eq:regPrimal}$
, with the loss
function in the formulation given by the squared loss:
\[ L(\wv;\x,y) := \frac{1}{2} (\wv^T \x - y)^2. \]
Various related regression methods are derived by using different types of regularization:
ordinary least squares or
linear least squares uses
no regularization; ridge regression uses L2
regularization; and Lasso uses L1
regularization. For all of these models, the average loss or training error,
When data arrive in a streaming fashion, it is useful to fit regression models online,
updating the parameters of the model as new data arrives. spark.mllib
currently supports
streaming linear regression using ordinary least squares. The fitting is similar
to that performed offline, except fitting occurs on each batch of data, so that
the model continually updates to reflect the data from the stream.
Examples
The following example demonstrates how to load training and testing data from two different input streams of text files, parse the streams as labeled points, fit a linear regression model online to the first stream, and make predictions on the second stream.
First, we import the necessary classes for parsing our input data and creating the model.
Then we make input streams for training and testing data. We assume a StreamingContext ssc
has already been created, see Spark Streaming Programming Guide
for more info. For this example, we use labeled points in training and testing streams,
but in practice you will likely want to use unlabeled vectors for test data.
We create our model by initializing the weights to zero and register the streams for training and testing then start the job. Printing predictions alongside true labels lets us easily see the result.
Finally, we can save text files with data to the training or testing folders.
Each line should be a data point formatted as (y,[x1,x2,x3])
where y
is the label
and x1,x2,x3
are the features. Anytime a text file is placed in args(0)
the model will update. Anytime a text file is placed in args(1)
you will see predictions.
As you feed more data to the training directory, the predictions
will get better!
Here is a complete example: {% include_example scala/org/apache/spark/examples/mllib/StreamingLinearRegressionExample.scala %}
First, we import the necessary classes for parsing our input data and creating the model.
Then we make input streams for training and testing data. We assume a StreamingContext ssc
has already been created, see Spark Streaming Programming Guide
for more info. For this example, we use labeled points in training and testing streams,
but in practice you will likely want to use unlabeled vectors for test data.
We create our model by initializing the weights to 0.
Now we register the streams for training and testing and start the job.
We can now save text files with data to the training or testing folders.
Each line should be a data point formatted as (y,[x1,x2,x3])
where y
is the label
and x1,x2,x3
are the features. Anytime a text file is placed in sys.argv[1]
the model will update. Anytime a text file is placed in sys.argv[2]
you will see predictions.
As you feed more data to the training directory, the predictions
will get better!
Here a complete example: {% include_example python/mllib/streaming_linear_regression_example.py %}
Behind the scene, spark.mllib
implements a simple distributed version of stochastic gradient descent
(SGD), building on the underlying gradient descent primitive (as described in the optimization section). All provided algorithms take as input a
regularization parameter (regParam
) along with various parameters associated with stochastic
gradient descent (stepSize
, numIterations
, miniBatchFraction
). For each of them, we support
all three possible regularizations (none, L1 or L2).
For Logistic Regression, L-BFGS version is implemented under LogisticRegressionWithLBFGS, and this version supports both binary and multinomial Logistic Regression while SGD version only supports binary Logistic Regression. However, L-BFGS version doesn't support L1 regularization but SGD one supports L1 regularization. When L1 regularization is not required, L-BFGS version is strongly recommended since it converges faster and more accurately compared to SGD by approximating the inverse Hessian matrix using quasi-Newton method.
Algorithms are all implemented in Scala: