Logistic Regression: Cross-Entropy and Geometry

CSI 4106 - Fall 2026

Marcel Turcotte

Version: Sep 28, 2026 01:38

Preamble

Message of the Day

Learning Objectives

By the end of this presentation, you should be able to:

  • Derive binary cross-entropy from the Bernoulli likelihood.
  • Explain why binary cross-entropy is preferred to MSE for logistic regression.
  • Interpret the logistic-regression decision boundary and its normal vector geometrically.
  • Distinguish the linear score from its corresponding probability and signed distance.
  • Relate the mathematical model to a batch-gradient-descent implementation.

Logistic Regression Review

Problem

  • General Case: P(y = k \mid x, \theta), where k is a class label.
  • Binary Case: y \in \{0,1\}
    • Predict P(y = 1 \mid x, \theta)

Logistic Regression

The Logistic Regression model is defined as:

\hat{p}_i = h_\theta(x_i) = \sigma(\theta^\top x_i) = \frac{1}{1+e^{-\theta^\top x_i}}

  • Predictions are made as follows:

  • \hat{y}_i = 0, if h_\theta(x_i) < 0.5

  • \hat{y}_i = 1, if h_\theta(x_i) \geq 0.5

Loss Function

Model Overview

  • Our model is expressed in a vectorized form as:

    \hat{p}_i = h_\theta(x_i) = \sigma(\theta^\top x_i) = \frac{1}{1+e^{-\theta^\top x_i}}

  • Prediction:

    • Assign \hat{y}_i = 0, if h_\theta(x_i) < 0.5; \hat{y}_i = 1, if h_\theta(x_i) \geq 0.5
  • The parameter vector \theta is optimized using gradient descent.

  • Which loss function should be used and why?

Remarks

  • In constructing machine learning models with libraries like scikit-learn or keras, one has to select a loss function or accept the default one.

  • Initially, the terminology can be confusing, as identical functions may be referenced by various names.

  • Our aim is to elucidate these complexities.

  • It is actually not that complicated!

Parameter Estimation

  • Logistic regression is a statistical model.

  • Its output is \hat{p} = P(y = 1 \mid x, \theta).

  • P(y = 0 \mid x, \theta) = 1 - \hat{p}.

  • Assumes that y values come from a Bernoulli distribution.

  • \theta is commonly found by Maximum Likelihood Estimation.

Parameter Estimation

Maximum Likelihood Estimation (MLE) is a statistical method used to estimate the parameters of a probabilistic model.

It identifies the parameter values that maximize the likelihood function, which measures how well the model explains the observed data.

Likelihood Function

Assuming the y values are independent and identically distributed (i.i.d.), the likelihood function is expressed as the product of individual probabilities.

In other words, given our data, \{(x_i, y_i)\}_{i=1}^N, the likelihood function is given by this equation.

\mathcal{L}(\theta) = \prod_{i=1}^{N} P(y_i \mid x_i, \theta)

Maximum Likelihood

\hat{\theta} = \underset{\theta \in \Theta}{\arg \max}\ \mathcal{L}(\theta) = \underset{\theta \in \Theta}{\arg \max} \prod_{i=1}^{N} P(y_i \mid x_i, \theta)

  • Observations:

    1. Maximizing a function is equivalent to minimizing its negative.
    2. The logarithm of a product equals the sum of its logarithms.

Negative Log-Likelihood

Maximum likelihood

\hat{\theta} = \underset{\theta \in \Theta}{\arg \max}\ \mathcal{L}(\theta) = \underset{\theta \in \Theta}{\arg \max} \prod_{i=1}^{N} P(y_i \mid x_i, \theta)

becomes negative log-likelihood

\hat{\theta} = \underset{\theta \in \Theta}{\arg \min} - \log \mathcal{L}(\theta) = \underset{\theta \in \Theta}{\arg \min} - \log \prod_{i=1}^{N} P(y_i \mid x_i, \theta) = \underset{\theta \in \Theta}{\arg \min} - \sum_{i=1}^{N} \log P(y_i \mid x_i, \theta)

Mathematical Reformulation

For binary outcomes, the probability P(y \mid x, \theta) is:

P(y \mid x, \theta) = \begin{cases} \sigma(\theta^\top x), & \text{if}\ y = 1 \\ 1 - \sigma(\theta^\top x), & \text{if}\ y = 0 \end{cases}

This can be compactly expressed as:

P(y \mid x, \theta) = \sigma(\theta^\top x)^y (1 - \sigma(\theta^\top x))^{1-y}

Loss Function

We define the objective as the mean negative log-likelihood.

J(\theta) = -\frac{1}{N}\log \mathcal{L}(\theta) = -\frac{1}{N}\sum_{i=1}^{N} \log P(y_i \mid x_i, \theta)

where P(y \mid x, \theta) = \sigma(\theta^\top x)^y (1 - \sigma(\theta^\top x))^{1-y}.

Consequently,

J(\theta) = -\frac{1}{N}\sum_{i=1}^{N} \log [ \sigma(\theta^\top x_i)^{y_i} (1 - \sigma(\theta^\top x_i))^{1-y_i} ]

Loss Function (continued)

Simplifying the equation.

J(\theta) = -\frac{1}{N}\sum_{i=1}^{N} \log [ \sigma(\theta^\top x_i)^{y_i} (1 - \sigma(\theta^\top x_i))^{1-y_i} ]

by distributing the \log over the product.

J(\theta) = -\frac{1}{N}\sum_{i=1}^{N} [ \log \sigma(\theta^\top x_i)^{y_i} + \log (1 - \sigma(\theta^\top x_i))^{1-y_i} ]

Loss Function (continued)

Simplifying the equation further.

J(\theta) = -\frac{1}{N}\sum_{i=1}^{N} [ \log \sigma(\theta^\top x_i)^{y_i} + \log (1 - \sigma(\theta^\top x_i))^{1-y_i} ]

by moving the exponents in front of the \logs.

J(\theta) = -\frac{1}{N}\sum_{i=1}^{N} [ y_i \log \sigma(\theta^\top x_i) + (1-y_i) \log (1 - \sigma(\theta^\top x_i)) ]

From Entropy to Cross-Entropy

  • Decision tree algorithms often employ entropy, a measure from information theory, to evaluate the quality of splits or partitions in decision rules.
  • Entropy quantifies the uncertainty or impurity associated with the potential outcomes of a random variable.

Entropy

Entropy in information theory quantifies the uncertainty or unpredictability of a random variable’s possible outcomes. It measures the average information associated with an outcome. When the base-2 logarithm is used, entropy is measured in bits. The entropy H of a discrete random variable X with possible outcomes \{x_1, x_2, \ldots, x_n\} and probability mass function P(X) is given by:

H(X) = -\sum_{i=1}^n P(x_i) \log_2 P(x_i)

Cross-Entropy

Cross-entropy measures the average penalty for the probabilities predicted by q, using p to weight the possible outcomes.

H(p, q) = -\sum_{k} p_k \log q_k

Cross-Entropy

  • For a binary label, the observed distribution is (y_i, 1-y_i).
  • The predicted Bernoulli distribution is (\hat{p}_i, 1-\hat{p}_i).
  • Their cross-entropy, denoted \ell_i, is the loss for example i.

Cross-Entropy

Consider the negative log-likelihood loss function:

J(\theta) = -\frac{1}{N}\sum_{i=1}^{N} \left[ y_i \log \sigma(\theta^\top x_i) + (1-y_i) \log (1 - \sigma(\theta^\top x_i)) \right]

By substituting \sigma(\theta^\top x_i) with \hat{p}_i, the function becomes:

\ell_i(\theta) = -\left[ y_i \log \hat{p}_i + (1-y_i) \log (1 - \hat{p}_i) \right], \qquad J(\theta)=\frac{1}{N}\sum_{i=1}^{N}\ell_i(\theta).

For the Bernoulli model, the mean negative log-likelihood is the binary cross-entropy.

Loss for a Positive Example

Code
import matplotlib.pyplot as plt
import numpy as np

# Predicted probabilities for the positive class
predicted_probabilities = np.linspace(0.001, 1, 1000)

# Binary cross-entropy for an example with y_i = 1
positive_example_losses = -np.log(predicted_probabilities)

fig, ax = plt.subplots(figsize=(5, 4))
ax.plot(
    predicted_probabilities,
    positive_example_losses,
    color='tab:blue',
    linewidth=2,
)

ax.set_xlim(0, 1)
ax.set_ylim(bottom=0)
ax.set_xlabel(r'Predicted probability $\hat{p}_i$')
ax.set_ylabel(r'Per-example loss $\ell_i$')
ax.set_title(r'Binary cross-entropy when $y_i=1$')
ax.grid(True, alpha=0.3)
plt.show()

BCE and MSE for a Positive Example

Code
# Sigmoid function
def sigmoid(t):
    return 1 / (1 + np.exp(-t))

# Compare the losses as functions of the linear score z for y = 1.
z_values = np.linspace(-6, 6, 1000)
p_values = sigmoid(z_values)
bce_values = -np.log(p_values)
mse_values = (p_values - 1) ** 2

fig, ax = plt.subplots(figsize=(7, 4.5))
ax.axvspan(-6, 0, color='tab:red', alpha=0.06)
ax.axvspan(0, 6, color='tab:green', alpha=0.06)
ax.plot(
    z_values,
    bce_values,
    color='tab:blue',
    linewidth=2.5,
    label=r'$\ell_{\mathrm{BCE}}(z;y=1)=-\log \sigma(z)$',
)
ax.plot(
    z_values,
    mse_values,
    color='tab:orange',
    linewidth=2.5,
    linestyle='--',
    label=r'$\ell_{\mathrm{MSE}}(z;y=1)=(\sigma(z)-1)^2$',
)
ax.axvline(0, color='black', linewidth=1)

ax.text(-5.8, 5.6, 'Incorrect side', color='darkred')
ax.text(2.8, 5.6, 'Correct side', color='darkgreen')
ax.annotate(
    'MSE becomes nearly flat',
    xy=(-4.5, mse_values[np.argmin(np.abs(z_values + 4.5))]),
    xytext=(-3.5, 2.1),
    arrowprops={'arrowstyle': '->', 'color': 'tab:orange'},
    color='tab:orange',
)

ax.set_xlim(-6, 6)
ax.set_ylim(0, 6.1)
ax.set_xlabel(r'Linear score $z$')
ax.set_ylabel(r'Per-example loss $\ell$')
ax.set_title(r'Per-example losses for a positive example ($y=1$)')
ax.grid(True, alpha=0.3)
ax.legend()
plt.show()

MSE: A Small Gradient

Let z=\theta^\top x and \hat{p}=\sigma(z). The sigmoid derivative is

\sigma'(z)=\hat{p}(1-\hat{p}).

For MSE, the chain rule gives

\frac{\partial \ell_{\mathrm{MSE}}}{\partial z} =2(\hat{p}-y)\,\hat{p}(1-\hat{p}).

For a positive example predicted near zero,

y=1,\quad \hat{p}\approx 0 \quad\Longrightarrow\quad \frac{\partial \ell_{\mathrm{MSE}}}{\partial z}\approx 0.

The near-zero sigmoid derivative suppresses the MSE correction even though the prediction error is large.

BCE: A Stronger Gradient

For binary cross-entropy,

\frac{\partial \ell_{\mathrm{BCE}}}{\partial z} =\hat{p}-y.

For the same confidently incorrect positive example,

y=1,\quad \hat{p}\approx 0 \quad\Longrightarrow\quad \frac{\partial \ell_{\mathrm{BCE}}}{\partial z}\approx -1.

The near-zero sigmoid factor cancels during differentiation, so the correction remains substantial.

Why not MSE as a Loss Function?

  • Binary cross-entropy is the negative log-likelihood of the Bernoulli model.
  • With binary cross-entropy, the logistic-regression objective is convex in the parameters.
  • With sigmoid outputs, MSE is not globally convex in the linear score. Its gradient can also become very small when the model makes a confidently incorrect prediction.

Geometric Interpretation

Geometric Interpretation

  • Let w=(\theta_1,\ldots,\theta_D)^\top denote the feature-weight vector.

  • The feature-dependent part of the linear score is a dot product: w^\top x = \theta_1 x_1 + \theta_2 x_2 + \ldots + \theta_D x_D.

  • The complete linear score includes the intercept: t(x)=\theta_0 + w^\top x.

  • The dot product has a geometric interpretation. w^\top x = \|w\|\|x\|\cos\phi

Geometric Interpretation

w^\top x = \|w\|\|x\|\cos\phi

  • The dot product depends on both vector lengths and the angle between them.
  • It is positive when the angle is acute and negative when the angle is obtuse.
  • It is zero when the vectors are perpendicular (\phi=90^\circ).

Geometric Interpretation

  • Logistic regression applies the sigmoid to the linear score t(x)=\theta_0+w^\top x.

  • The equation t(x)=0 defines a hyperplane in feature space.

  • The feature-weight vector w is normal to this hyperplane. The intercept \theta_0 shifts the boundary without changing its orientation.

Geometric Interpretation

  • The decision boundary is where t(x)=0.

  • Points with t(x)>0 receive probability greater than 0.5.

  • Points with t(x)<0 receive probability less than 0.5.

  • The sigmoid converts the linear score into a probability. The signed distance to the boundary is \frac{t(x)}{\|w\|}.

Implementation

Implementation: Generating Data

# Generate synthetic data for a binary classification problem

rng = np.random.default_rng(42)

m = 100  # number of examples
d = 2    # number of features

X = rng.standard_normal((m, d))

# Define labels using a linear decision boundary with some noise:

noise = 0.5 * rng.standard_normal(m)

y = (X[:, 0] + X[:, 1] + noise > 0).astype(int)

Implementation: Visualization

Code
# Visualize the decision boundary along with the data points
plt.figure(figsize=(8, 6))
plt.scatter(X[y == 0][:, 0], X[y == 0][:, 1], color='red', label='Class 0')
plt.scatter(X[y == 1][:, 0], X[y == 1][:, 1], color='blue', label='Class 1')

plt.xlabel("Feature 1")
plt.ylabel("Feature 2")
plt.title("Data")
plt.legend()
plt.show()

Implementation: Cost Function

# Cost function: binary cross-entropy
def cost_function(theta, X, y):
    m = len(y)
    h = sigmoid(X.dot(theta))
    h = np.clip(h, 1e-12, 1 - 1e-12)
    cost = -(1/m) * np.sum(y * np.log(h) + (1 - y) * np.log(1 - h))
    return cost

# Gradient of the cost function
def gradient(theta, X, y):
    m = len(y)
    h = sigmoid(X.dot(theta))
    grad = (1/m) * X.T.dot(h - y)
    return grad

Implementation: Logistic Regression

# Logistic regression training using gradient descent
def logistic_regression(X, y, learning_rate=0.1, iterations=1000):
    m, n = X.shape
    theta = np.zeros(n)
    cost_history = []
    
    for i in range(iterations):
        theta -= learning_rate * gradient(theta, X, y)
        cost_history.append(cost_function(theta, X, y))
        
    return theta, cost_history

Training

# Add intercept term (bias)
X_with_intercept = np.hstack([np.ones((m, 1)), X])

# Train the logistic regression model
theta, cost_history = logistic_regression(X_with_intercept, y, learning_rate=0.1, iterations=1000)
theta_0 = theta[0]
w = theta[1:]

print("Optimized theta:", theta)
Optimized theta: [-0.28250171  3.11670841  3.46892095]

Cost Function Convergence

Code
plt.figure(figsize=(8, 6))
plt.plot(cost_history, label="Cost")
plt.xlabel("Iteration")
plt.ylabel("Cost")
plt.title("Cost Function Convergence")
plt.legend()
plt.show()

Decision Boundary and Data Points

Code
plt.figure(figsize=(8, 6))
plt.scatter(X[y == 0][:, 0], X[y == 0][:, 1], color='red', label='Class 0')
plt.scatter(X[y == 1][:, 0], X[y == 1][:, 1], color='blue', label='Class 1')

# Decision boundary: theta_0 + w[0]*x1 + w[1]*x2 = 0
x_vals = np.array([min(X[:, 0]) - 1, max(X[:, 0]) + 1])
y_vals = -(theta_0 + w[0] * x_vals) / w[1]
plt.plot(x_vals, y_vals, label='Decision Boundary', color='green')
plt.xlabel("Feature 1")
plt.ylabel("Feature 2")
plt.title("Logistic Regression Decision Boundary")
plt.legend()
plt.show()

Visualizing the Weight Vector

The fitted feature-weight vector w=(\theta_1,\theta_2)^\top is normal to the decision boundary.

The intercept \theta_0 determines the position of the boundary.

Visualizing the Weight Vector

Code
# Plot decision boundary and data points
plt.figure(figsize=(8, 6))
plt.scatter(X[y == 0][:, 0], X[y == 0][:, 1], color='red', label='Class 0')
plt.scatter(X[y == 1][:, 0], X[y == 1][:, 1], color='blue', label='Class 1')

# Decision boundary: theta_0 + w[0]*x1 + w[1]*x2 = 0
x_vals = np.array([min(X[:, 0]) - 1, max(X[:, 0]) + 1])
y_vals = -(theta_0 + w[0] * x_vals) / w[1]
plt.plot(x_vals, y_vals, label='Decision Boundary', color='green')

# --- Draw the normal vector ---
# The normal vector is w.
# Choose a reference point on the decision boundary. Here, we use x1 = 0:
x_ref = 0
y_ref = -theta_0 / w[1]

# Copy the feature-weight vector for display.
normal = w.copy()

# Normalize and scale for display
normal_norm = np.linalg.norm(normal)
if normal_norm != 0:
    normal_unit = normal / normal_norm
else:
    normal_unit = normal
scale = 2  # adjust scale as needed
normal_display = normal_unit * scale

# Draw an arrow starting at the reference point
plt.arrow(x_ref, y_ref, normal_display[0], normal_display[1],
          head_width=0.1, head_length=0.2, fc='black', ec='black')
plt.text(x_ref + normal_display[0]*1.1, y_ref + normal_display[1]*1.1, 
         r'$w$', color='black', fontsize=12)

plt.xlabel("Feature 1")
plt.ylabel("Feature 2")
plt.title("Logistic Regression Decision Boundary and Normal Vector")
plt.legend()
plt.gca().set_aspect('equal', adjustable='box')
plt.ylim(-3, 3)
plt.show()

Near the Decision Boundary

Code
# --- Visualization Setup ---
# Create a grid over the feature space
x1_range = np.linspace(X[:, 0].min()-1, X[:, 0].max()+1, 100)
x2_range = np.linspace(X[:, 1].min()-1, X[:, 1].max()+1, 100)
xx1, xx2 = np.meshgrid(x1_range, x2_range)

# Construct the grid input (with intercept) for predictions
grid = np.c_[np.ones(xx1.ravel().shape), xx1.ravel(), xx2.ravel()]
# Compute predicted probabilities over the grid
probs = sigmoid(grid.dot(theta)).reshape(xx1.shape)
# --- Approach 2: 2D Contour (Heatmap) Plot ---
plt.figure(figsize=(8, 6))
contour = plt.contourf(xx1, xx2, probs, cmap='spring', levels=50)
plt.colorbar(contour)
plt.contour(xx1, xx2, probs, levels=[0.5], colors='green', linewidths=2)
plt.xlabel('Feature x1')
plt.ylabel('Feature x2')
plt.title('Contour Plot (Heatmap) of Predicted Probabilities')
# Overlay training data
plt.scatter(X[y == 0][:, 0], X[y == 0][:, 1], color='red', edgecolor='k', label='Class 0')
plt.scatter(X[y == 1][:, 0], X[y == 1][:, 1], color='blue', edgecolor='k', label='Class 1')
plt.legend()
plt.show()

Epilogue

Summary

  • Logistic regression models the probability of a binary outcome.
  • Maximizing the Bernoulli likelihood is equivalent to minimizing binary cross-entropy.
  • Cross-entropy strongly penalizes confident, incorrect predictions.
  • The equation \theta_0+w^\top x=0 defines the decision boundary, and the feature-weight vector w is normal to it.
  • Batch gradient descent updates the parameters using the average gradient over the training set.

Next lecture

  • Performance measures and model evaluation

References

Russell, Stuart, and Peter Norvig. 2020. Artificial Intelligence: A Modern Approach. 4th ed. Pearson. http://aima.cs.berkeley.edu/.

Marcel Turcotte

[email protected]

School of Electrical Engineering and Computer Science (EECS)

University of Ottawa