Assign \hat{y}_i = 0, if h_\theta(x_i) < 0.5; \hat{y}_i = 1, if h_\theta(x_i) \geq 0.5
The parameter vector \theta is optimized using gradient descent.
Which loss function should be used and why?
Remarks
In constructing machine learning models with libraries like scikit-learn or keras, one has to select a loss function or accept the default one.
Initially, the terminology can be confusing, as identical functions may be referenced by various names.
Our aim is to elucidate these complexities.
It is actually not that complicated!
Parameter Estimation
Logistic regression is a statistical model.
Its output is \hat{p} = P(y = 1 \mid x, \theta).
P(y = 0 \mid x, \theta) = 1 - \hat{p}.
Assumes that y values come from a Bernoulli distribution.
\theta is commonly found by Maximum Likelihood Estimation.
Parameter Estimation
Maximum Likelihood Estimation (MLE) is a statistical method used to estimate the parameters of a probabilistic model.
It identifies the parameter values that maximize the likelihood function, which measures how well the model explains the observed data.
Likelihood Function
Assuming the y values are independent and identically distributed (i.i.d.), the likelihood function is expressed as the product of individual probabilities.
In other words, given our data, \{(x_i, y_i)\}_{i=1}^N, the likelihood function is given by this equation.
Decision tree algorithms often employ entropy, a measure from information theory, to evaluate the quality of splits or partitions in decision rules.
Entropy quantifies the uncertainty or impurity associated with the potential outcomes of a random variable.
Entropy
Entropy in information theory quantifies the uncertainty or unpredictability of a random variable’s possible outcomes. It measures the average information associated with an outcome. When the base-2 logarithm is used, entropy is measured in bits. The entropy H of a discrete random variable X with possible outcomes \{x_1, x_2, \ldots, x_n\} and probability mass function P(X) is given by:
H(X) = -\sum_{i=1}^n P(x_i) \log_2 P(x_i)
Cross-Entropy
Cross-entropy measures the average penalty for the probabilities predicted by q, using p to weight the possible outcomes.
H(p, q) = -\sum_{k} p_k \log q_k
Cross-Entropy
For a binary label, the observed distribution is (y_i, 1-y_i).
The predicted Bernoulli distribution is (\hat{p}_i, 1-\hat{p}_i).
Their cross-entropy, denoted \ell_i, is the loss for example i.
Cross-Entropy
Consider the negative log-likelihood loss function:
For the Bernoulli model, the mean negative log-likelihood is the binary cross-entropy.
Loss for a Positive Example
Code
import matplotlib.pyplot as pltimport numpy as np# Predicted probabilities for the positive classpredicted_probabilities = np.linspace(0.001, 1, 1000)# Binary cross-entropy for an example with y_i = 1positive_example_losses =-np.log(predicted_probabilities)fig, ax = plt.subplots(figsize=(5, 4))ax.plot( predicted_probabilities, positive_example_losses, color='tab:blue', linewidth=2,)ax.set_xlim(0, 1)ax.set_ylim(bottom=0)ax.set_xlabel(r'Predicted probability $\hat{p}_i$')ax.set_ylabel(r'Per-example loss $\ell_i$')ax.set_title(r'Binary cross-entropy when $y_i=1$')ax.grid(True, alpha=0.3)plt.show()
BCE and MSE for a Positive Example
Code
# Sigmoid functiondef sigmoid(t):return1/ (1+ np.exp(-t))# Compare the losses as functions of the linear score z for y = 1.z_values = np.linspace(-6, 6, 1000)p_values = sigmoid(z_values)bce_values =-np.log(p_values)mse_values = (p_values -1) **2fig, ax = plt.subplots(figsize=(7, 4.5))ax.axvspan(-6, 0, color='tab:red', alpha=0.06)ax.axvspan(0, 6, color='tab:green', alpha=0.06)ax.plot( z_values, bce_values, color='tab:blue', linewidth=2.5, label=r'$\ell_{\mathrm{BCE}}(z;y=1)=-\log \sigma(z)$',)ax.plot( z_values, mse_values, color='tab:orange', linewidth=2.5, linestyle='--', label=r'$\ell_{\mathrm{MSE}}(z;y=1)=(\sigma(z)-1)^2$',)ax.axvline(0, color='black', linewidth=1)ax.text(-5.8, 5.6, 'Incorrect side', color='darkred')ax.text(2.8, 5.6, 'Correct side', color='darkgreen')ax.annotate('MSE becomes nearly flat', xy=(-4.5, mse_values[np.argmin(np.abs(z_values +4.5))]), xytext=(-3.5, 2.1), arrowprops={'arrowstyle': '->', 'color': 'tab:orange'}, color='tab:orange',)ax.set_xlim(-6, 6)ax.set_ylim(0, 6.1)ax.set_xlabel(r'Linear score $z$')ax.set_ylabel(r'Per-example loss $\ell$')ax.set_title(r'Per-example losses for a positive example ($y=1$)')ax.grid(True, alpha=0.3)ax.legend()plt.show()
MSE: A Small Gradient
Let z=\theta^\top x and \hat{p}=\sigma(z). The sigmoid derivative is
The near-zero sigmoid factor cancels during differentiation, so the correction remains substantial.
Why not MSE as a Loss Function?
Binary cross-entropy is the negative log-likelihood of the Bernoulli model.
With binary cross-entropy, the logistic-regression objective is convex in the parameters.
With sigmoid outputs, MSE is not globally convex in the linear score. Its gradient can also become very small when the model makes a confidently incorrect prediction.
Geometric Interpretation
Geometric Interpretation
Let w=(\theta_1,\ldots,\theta_D)^\top denote the feature-weight vector.
The feature-dependent part of the linear score is a dot product:
w^\top x = \theta_1 x_1 + \theta_2 x_2 + \ldots + \theta_D x_D.
The complete linear score includes the intercept:
t(x)=\theta_0 + w^\top x.
The dot product has a geometric interpretation.
w^\top x = \|w\|\|x\|\cos\phi
Geometric Interpretation
w^\top x = \|w\|\|x\|\cos\phi
The dot product depends on both vector lengths and the angle between them.
It is positive when the angle is acute and negative when the angle is obtuse.
It is zero when the vectors are perpendicular(\phi=90^\circ).
Geometric Interpretation
Logistic regression applies the sigmoid to the linear score t(x)=\theta_0+w^\top x.
The equation t(x)=0 defines a hyperplane in feature space.
The feature-weight vector w is normal to this hyperplane. The intercept \theta_0 shifts the boundary without changing its orientation.
Geometric Interpretation
The decision boundary is where t(x)=0.
Points with t(x)>0 receive probability greater than 0.5.
Points with t(x)<0 receive probability less than 0.5.
The sigmoid converts the linear score into a probability. The signed distance to the boundary is
\frac{t(x)}{\|w\|}.
Implementation
Implementation: Generating Data
# Generate synthetic data for a binary classification problemrng = np.random.default_rng(42)m =100# number of examplesd =2# number of featuresX = rng.standard_normal((m, d))# Define labels using a linear decision boundary with some noise:noise =0.5* rng.standard_normal(m)y = (X[:, 0] + X[:, 1] + noise >0).astype(int)
Implementation: Visualization
Code
# Visualize the decision boundary along with the data pointsplt.figure(figsize=(8, 6))plt.scatter(X[y ==0][:, 0], X[y ==0][:, 1], color='red', label='Class 0')plt.scatter(X[y ==1][:, 0], X[y ==1][:, 1], color='blue', label='Class 1')plt.xlabel("Feature 1")plt.ylabel("Feature 2")plt.title("Data")plt.legend()plt.show()
Implementation: Cost Function
# Cost function: binary cross-entropydef cost_function(theta, X, y): m =len(y) h = sigmoid(X.dot(theta)) h = np.clip(h, 1e-12, 1-1e-12) cost =-(1/m) * np.sum(y * np.log(h) + (1- y) * np.log(1- h))return cost# Gradient of the cost functiondef gradient(theta, X, y): m =len(y) h = sigmoid(X.dot(theta)) grad = (1/m) * X.T.dot(h - y)return grad
Implementation: Logistic Regression
# Logistic regression training using gradient descentdef logistic_regression(X, y, learning_rate=0.1, iterations=1000): m, n = X.shape theta = np.zeros(n) cost_history = []for i inrange(iterations): theta -= learning_rate * gradient(theta, X, y) cost_history.append(cost_function(theta, X, y))return theta, cost_history
Training
# Add intercept term (bias)X_with_intercept = np.hstack([np.ones((m, 1)), X])# Train the logistic regression modeltheta, cost_history = logistic_regression(X_with_intercept, y, learning_rate=0.1, iterations=1000)theta_0 = theta[0]w = theta[1:]print("Optimized theta:", theta)
plt.figure(figsize=(8, 6))plt.plot(cost_history, label="Cost")plt.xlabel("Iteration")plt.ylabel("Cost")plt.title("Cost Function Convergence")plt.legend()plt.show()