Model Evaluation and Hyperparameter Tuning

CSI 4106 - Fall 2026

Marcel Turcotte

Version: Sep 25, 2026 08:51

Preamble

Message of the Day

Learning Objectives

  • Distinguish the roles of training, validation, and test data in model development and final evaluation.
  • Explain how k-fold cross-validation trains a new model in each iteration and interpret the mean and standard deviation of its related fold scores.
  • Prevent data leakage by fitting preprocessing within each fold and applying stored transformations to validation and test examples.
  • Distinguish learned model parameters from hyperparameters and explain why interacting hyperparameters should be evaluated jointly.
  • Apply GridSearchCV to tune multiple pipelines using a consistent metric and explain when randomized search is useful.
  • Compare tuned model families, refit the selected pipeline on all training data, and evaluate it once on the untouched test set.

Introduction

Dataset: OpenML

OpenML is an open platform for sharing datasets, algorithms, and experiments - to learn how to learn better, together.

import numpy as np
seed = 42

from sklearn.datasets import fetch_openml

diabetes = fetch_openml(name='diabetes', version=1)
print(diabetes.DESCR)

Dataset: OpenML

Author: Vincent Sigillito

Source: Obtained from UCI

Please cite: UCI citation policy

  1. Title: Pima Indians Diabetes Database

  2. Sources:

    1. Original owners: National Institute of Diabetes and Digestive and Kidney Diseases
    2. Donor of database: Vincent Sigillito ([email protected]) Research Center, RMI Group Leader Applied Physics Laboratory The Johns Hopkins University Johns Hopkins Road Laurel, MD 20707 (301) 953-6231
    3. Date received: 9 May 1990
  3. Past Usage:

    1. Smith,J.W., Everhart,J.E., Dickson,W.C., Knowler,W.C., & Johannes,R.S. (1988). Using the ADAP learning algorithm to forecast the onset of diabetes mellitus. In {it Proceedings of the Symposium on Computer Applications and Medical Care} (pp. 261–265). IEEE Computer Society Press.

      The diagnostic, binary-valued variable investigated is whether the patient shows signs of diabetes according to World Health Organization criteria (i.e., if the 2 hour post-load plasma glucose was at least 200 mg/dl at any survey examination or if found during routine medical care). The population lives near Phoenix, Arizona, USA.

      Results: Their ADAP algorithm makes a real-valued prediction between 0 and 1. This was transformed into a binary decision using a cutoff of 0.448. Using 576 training instances, the sensitivity and specificity of their algorithm was 76% on the remaining 192 instances.

  4. Relevant Information: Several constraints were placed on the selection of these instances from a larger database. In particular, all patients here are females at least 21 years old of Pima Indian heritage. ADAP is an adaptive learning routine that generates and executes digital analogs of perceptron-like devices. It is a unique algorithm; see the paper for details.

  5. Number of Instances: 768

  6. Number of Attributes: 8 plus class

  7. For Each Attribute: (all numeric-valued)

    1. Number of times pregnant
    2. Plasma glucose concentration a 2 hours in an oral glucose tolerance test
    3. Diastolic blood pressure (mm Hg)
    4. Triceps skin fold thickness (mm)
    5. 2-Hour serum insulin (mu U/ml)
    6. Body mass index (weight in kg/(height in m)^2)
    7. Diabetes pedigree function
    8. Age (years)
    9. Class variable (0 or 1)
  8. Missing Attribute Values: None

  9. Class Distribution: (class value 1 is interpreted as “tested positive for diabetes”)

    Class Value Number of instances 0 500 1 268

  10. Brief statistical analysis:

    Attribute number: Mean: Standard Deviation:

    1.                 3.8     3.4
    2.               120.9    32.0
    3.                69.1    19.4
    4.                20.5    16.0
    5.                79.8   115.2
    6.                32.0     7.9
    7.                 0.5     0.3
    8.                33.2    11.8

Relabeled values in attribute ‘class’ From: 0 To: tested_negative
From: 1 To: tested_positive

Downloaded from openml.org.

Dataset: return_X_y

fetch_openml returns a Bunch by default, or X and y when return_X_y=True.

X, y = fetch_openml(name='diabetes', version=1, return_X_y=True)

The positive class is less frequent, but both classes have enough examples for stratified cross-validation.

print(y.value_counts())
class
tested_negative    500
tested_positive    268
Name: count, dtype: int64

Convert the target labels to 0 and 1:

y = y.map({'tested_negative': 0, 'tested_positive': 1})

Missing measurements

Following the convention used by the corrected PimaIndiansDiabetes2 version of this dataset, zero values in five clinical measurements are treated as missing. Zero remains valid for the number of pregnancies.

missing_value_columns = ["plas", "pres", "skin", "insu", "mass"]

X = X.copy()
X[missing_value_columns] = X[missing_value_columns].replace(0, np.nan)

X[missing_value_columns].isna().sum()
plas      5
pres     35
skin    227
insu    374
mass     11
dtype: int64

Cross-validation

Training and test sets

This division is often called the holdout method.

  • Common starting point: Allocate 80% of the data to training and reserve 20% for testing.

  • Training set: Used to fit models and make model-building choices.

  • Test set: An independent subset used once for the final evaluation.

Training and test sets

Training error: The prediction error measured on examples used to fit the model.

  • Model fitting generally seeks to reduce training loss or error.
  • A low training error does not guarantee a low generalization error.

Training and test sets

Generalization error: The expected prediction error on new examples drawn under similar conditions.

It cannot be observed directly. Validation and test scores provide estimates from finite samples.

Training and test sets

Underfitting:

  • High training error
  • Model is too simple to capture underlying patterns
  • Poor performance on both training and new data

Overfitting:

  • Low training error, but high generalization error
  • Model captures noise or irrelevant patterns
  • Poor performance on new, unseen data

Definition

Cross-validation estimates how well a model-building procedure will perform on unseen examples.

It repeatedly partitions the training data, trains a new classifier on some folds, and evaluates that classifier on the remaining fold. The test set remains untouched.

k-fold cross-validation

  1. Divide the training data into k approximately equal parts (folds).
  2. Repeat k times:
    • Start with a new, unfitted classifier.
    • Train it on k-1 folds.
    • Evaluate it on the remaining validation fold.
  3. Aggregation: Summarize the k scores produced by the k separately trained classifiers.
Code
import matplotlib.pyplot as plt

def plot_k_fold_cross_validation(k):
    fig, ax = plt.subplots()

    matrix = np.ones((k, k))
    np.fill_diagonal(matrix, 0)

    ax.imshow(matrix, cmap='binary', interpolation='none')

    for i in range(k):
        for j in range(k):
            label = 'Validation' if i == j else 'Train'
            color = "black" if i == j else "white"
            ax.text(j, i, label, ha="center", va="center", color=color)

    ax.set_xticks(np.arange(k))
    ax.set_yticks(np.arange(k))
    ax.set_xticklabels([f"Fold {i+1}" for i in range(k)])
    ax.set_yticklabels([f"Iteration {i+1}" for i in range(k)])
    plt.title(f'{k}-Fold Cross-validation')

    ax.grid(False)
    plt.show()

5-fold cross-validation

Code
plot_k_fold_cross_validation(5)

Evaluation across validation folds

  • During model selection, cross-validation averages scores across several validation folds rather than relying on one validation partition.
  • One score per fold reveals sensitivity to the particular validation examples.
  • The mean and standard deviation summarize the fold scores, but the scores are not independent.

Estimating generalization

  • Cross-validation estimates performance on unseen examples drawn under similar conditions.
  • It does not improve generalization by itself. It helps compare model-building choices using only the training data.
  • The untouched test set provides the final evaluation after those choices have been made.

Efficient use of data

  • Cross-validation is particularly useful when data is limited.
  • Across the k iterations, each training example appears once in a validation fold and k-1 times in a fitting fold.
  • After model selection, the chosen procedure can be refitted on the complete training set.

Hyperparameter tuning

  • Cross-validation compares candidate hyperparameter settings using the same folds and evaluation metric.
  • The selected setting has the best estimated performance among the candidates that were evaluated.

Challenges

  • Computational cost: Requires repeated model fitting.
    • Leave-one-out (LOO) is the extreme case where ( k = N ).
  • Class imbalance: Random folds may not adequately represent minority classes.
    • Use stratified cross-validation to approximately maintain class proportions.
  • Implementation complexity: Manual implementations are error-prone, especially for nested cross-validation, bootstrap resampling, or integration into larger pipelines.

Hyperparameter tuning

Workflow

Workflow: implementation

from sklearn.model_selection import (
    StratifiedKFold,
    cross_val_score,
    train_test_split,
)

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.2,
    random_state=seed,
    stratify=y,
)

cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=seed)
selection_metric = "roc_auc"

Data leakage during cross-validation

Data leakage occurs when information from a validation fold influences the procedure fitted on the remaining folds.

# Incorrect: the scaler sees every future validation fold

scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)

scores = cross_val_score(
    KNeighborsClassifier(),
    X_train_scaled,
    y_train,
    cv=cv,
)

The validation examples influence the means and scales used to evaluate them.

fit_transform and transform

  • fit(X) learns information from X.
  • fit_transform(X) learns that information and applies the transformation to the same examples.
  • transform(X) applies the stored information without recomputing it.
scaler = StandardScaler()

X_fitting_folds_scaled = scaler.fit_transform(X_fitting_folds)
X_validation_fold_scaled = scaler.transform(X_validation_fold)

Validation, test, and future examples receive transform, never fit or fit_transform.

Preprocessing within each fold

For every fold, cross-validation creates and fits a new pipeline.

Pipeline step Fitting folds Validation fold
Imputer fit_transform transform
Scaler fit_transform transform
Classifier fit predict_proba / decision_function

Learned medians, means, and scales are not carried into the next fold.

Preprocessing pipelines

from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.tree import DecisionTreeClassifier

tree_pipeline = Pipeline([
    ("imputer", SimpleImputer(strategy="median")),
    ("model", DecisionTreeClassifier(random_state=seed)),
])

knn_pipeline = Pipeline([
    ("imputer", SimpleImputer(strategy="median")),
    ("scaler", StandardScaler()),
    ("model", KNeighborsClassifier()),
])

logistic_pipeline = Pipeline([
    ("imputer", SimpleImputer(strategy="median")),
    ("scaler", StandardScaler()),
    ("model", LogisticRegression(
        solver="saga",
        max_iter=3000,
        random_state=seed,
    )),
])

cross_val_score

knn_scores = cross_val_score(
    knn_pipeline,
    X_train,
    y_train,
    cv=cv,
    scoring=selection_metric,
)

print("Scores:", knn_scores)
print(f"Mean AUROC: {knn_scores.mean():.3f}")
print(f"Standard deviation: {knn_scores.std():.3f}")
Scores: [0.81133721 0.76061047 0.84447674 0.74883721 0.75639881]
Mean AUROC: 0.784
Standard deviation: 0.037

Parameters and hyperparameters

Model parameters are learned from the training data.

Hyperparameters are chosen before each fit and control the model structure or learning process.

Hyperparameters: decision tree

  • criterion: chooses how the quality of a split is measured: gini, entropy, or log_loss.
  • max_depth: limits the maximum depth of the tree.

Hyperparameters: logistic regression

  • C: inverse regularization strength. Smaller values apply stronger regularization.
  • l1_ratio: selects the balance between L2 and L1 regularization when the saga solver is used.

Hyperparameters: KNN

  • n_neighbors: number of neighbours used to make a prediction.
  • weights: gives neighbours equal influence (uniform) or greater influence when they are closer (distance).

Experiment: n_neighbors

neighbor_scores = {}
neighbor_values = [1, 3, 5, 7, 9, 11, 15, 21, 31]

for value in neighbor_values:
    candidate = knn_pipeline.set_params(
        model__n_neighbors=value,
        model__weights="uniform",
    )
    scores = cross_val_score(
        candidate,
        X_train,
        y_train,
        cv=cv,
        scoring=selection_metric,
    )
    neighbor_scores[value] = scores.mean()
    print(f"n_neighbors={value:2d}: {scores.mean():.3f}")

best_n_neighbors = max(neighbor_scores, key=neighbor_scores.get)
print("Best n_neighbors:", best_n_neighbors)

Experiment: n_neighbors

n_neighbors= 1: 0.664
n_neighbors= 3: 0.764
n_neighbors= 5: 0.784
n_neighbors= 7: 0.798
n_neighbors= 9: 0.815
n_neighbors=11: 0.822
n_neighbors=15: 0.832
n_neighbors=21: 0.831
n_neighbors=31: 0.831
Best n_neighbors: 15

Experiment: weights

weight_scores = {}

for value in ["uniform", "distance"]:
    candidate = knn_pipeline.set_params(
        model__n_neighbors=best_n_neighbors,
        model__weights=value,
    )
    scores = cross_val_score(
        candidate,
        X_train,
        y_train,
        cv=cv,
        scoring=selection_metric,
    )
    weight_scores[value] = scores.mean()
    print(f"weights={value:8s}: {scores.mean():.3f}")
weights=uniform : 0.832
weights=distance: 0.833

GridSearchCV: decision tree

from sklearn.model_selection import GridSearchCV

tree_param_grid = {
    "model__max_depth": [2, 3, 4, 5, 6, None],
    "model__criterion": ["gini", "entropy", "log_loss"],
}

tree_search = GridSearchCV(
    tree_pipeline,
    tree_param_grid,
    cv=cv,
    scoring=selection_metric,
)
tree_search.fit(X_train, y_train)

(tree_search.best_params_, tree_search.best_score_)
({'model__criterion': 'gini', 'model__max_depth': 4},
 np.float64(0.7881160022148395))

GridSearchCV: KNN

knn_param_grid = {
    "model__n_neighbors": neighbor_values,
    "model__weights": ["uniform", "distance"],
}

knn_search = GridSearchCV(
    knn_pipeline,
    knn_param_grid,
    cv=cv,
    scoring=selection_metric,
)
knn_search.fit(X_train, y_train)

(knn_search.best_params_, knn_search.best_score_)
({'model__n_neighbors': 21, 'model__weights': 'distance'},
 np.float64(0.8336620985603543))

GridSearchCV: logistic regression

logistic_param_grid = {
    "model__C": [0.01, 0.1, 1, 10, 100],
    "model__l1_ratio": [0.0, 1.0],
}

logistic_search = GridSearchCV(
    logistic_pipeline,
    logistic_param_grid,
    cv=cv,
    scoring=selection_metric,
)
logistic_search.fit(X_train, y_train)

(logistic_search.best_params_, logistic_search.best_score_)
({'model__C': 0.1, 'model__l1_ratio': 0.0}, np.float64(0.8443867663344408))

Comparing tuned models

model_searches = {
    "Decision tree": tree_search,
    "KNN": knn_search,
    "Logistic regression": logistic_search,
}

for name, search in model_searches.items():
    print(f"{name:20s}: {search.best_score_:.3f}")

best_model_name = max(
    model_searches,
    key=lambda name: model_searches[name].best_score_,
)
best_search = model_searches[best_model_name]

print("Selected model:", best_model_name)
Decision tree       : 0.788
KNN                 : 0.834
Logistic regression : 0.844
Selected model: Logistic regression

Workflow

Final evaluation

from sklearn.metrics import classification_report, roc_auc_score

best_model = best_search.best_estimator_

y_test_score = best_model.predict_proba(X_test)[:, 1]
y_test_pred = best_model.predict(X_test)

print("Selected model:", best_model_name)
print(f"Test AUROC: {roc_auc_score(y_test, y_test_score):.3f}")
print(classification_report(y_test, y_test_pred))
Selected model: Logistic regression
Test AUROC: 0.810
              precision    recall  f1-score   support

           0       0.74      0.80      0.77       100
           1       0.57      0.48      0.52        54

    accuracy                           0.69       154
   macro avg       0.65      0.64      0.64       154
weighted avg       0.68      0.69      0.68       154

Closing Material

Summary

  • Reserved an untouched test set for one final evaluation after all model-building choices.
  • Used k-fold cross-validation to train a new model in each iteration and summarize the related fold scores with their mean and standard deviation.
  • Prevented data leakage by fitting imputation and scaling inside each fold, then applying the stored transformations to validation and test examples.
  • Distinguished learned model parameters from hyperparameters and evaluated interacting hyperparameters jointly.
  • Used GridSearchCV to tune and compare decision tree, KNN, and logistic regression pipelines with the same folds and AUROC metric. Randomized search provides an alternative for larger search spaces.
  • Refit the selected pipeline on all training data and evaluated it once on the test set. Further model changes would turn that test set into validation data.

Next lecture

  • After the quiz, we will examine machine learning engineering.

References

Russell, Stuart, and Peter Norvig. 2020. Artificial Intelligence: A Modern Approach. 4th ed. Pearson. http://aima.cs.berkeley.edu/.

Marcel Turcotte

[email protected]

School of Electrical Engineering and Computer Science (EECS)

University of Ottawa