{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# The Role of Randomness in Machine Learning\n",
        "\n",
        "CSI 4106 — Introduction to Artificial Intelligence\n",
        "\n",
        "Marcel Turcotte  \n",
        "2026-09-20\n",
        "\n",
        "# Introduction\n",
        "\n",
        "Randomness plays several roles in machine learning. It is used to\n",
        "initialize model parameters, shuffle data, create mini-batches,\n",
        "construct training and test sets, and implement randomized algorithms.\n",
        "These roles are related, but not identical. For example, random\n",
        "initialization breaks the symmetry between neurons in the same layer of\n",
        "a neural network; if their weights were initialized identically, they\n",
        "could learn the same features. Random shuffling and sampling expose an\n",
        "algorithm to different selections and orderings of the available data.\n",
        "\n",
        "# Pseudo-Random Numbers\n",
        "\n",
        "Most computational libraries that generate random numbers actually use\n",
        "*pseudo-random* number generators. They produce sequences with useful\n",
        "statistical properties, but the sequences are generated\n",
        "deterministically.\n",
        "\n",
        "A pseudo-random number generator maintains an internal **state**. Each\n",
        "draw produces a value and updates this state. A **seed** initializes the\n",
        "state; using the same generator, the same seed, and the same sequence of\n",
        "calls reproduces the same values. The value `42` is conventional in\n",
        "examples, but it has no special statistical property.\n",
        "\n",
        "In the following example using Python’s built-in [`random`\n",
        "module](https://docs.python.org/3/library/random.html), we generate two\n",
        "sequences. Because we do not reset the seed between them, the second\n",
        "sequence continues from the state left by the first. The values will\n",
        "usually change if this cell is executed in a new Python session."
      ],
      "id": "81eaa4eb-31e2-4ddd-815c-c9350cd1a86b"
    },
    {
      "cell_type": "code",
      "execution_count": 1,
      "metadata": {},
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Sequence 1: [14, 88, 61, 64, 77]\n",
            "Sequence 2: [76, 36, 46, 14, 41]"
          ]
        }
      ],
      "source": [
        "import random\n",
        "\n",
        "sequence_1 = [random.randint(1, 100) for _ in range(5)]\n",
        "sequence_2 = [random.randint(1, 100) for _ in range(5)]\n",
        "\n",
        "print(f\"Sequence 1: {sequence_1}\")\n",
        "print(f\"Sequence 2: {sequence_2}\")"
      ],
      "id": "no_seed"
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "By contrast, the next example resets the generator to the same state\n",
        "before producing each of the first two sequences. A different seed\n",
        "normally selects a different sequence."
      ],
      "id": "f20a3623-d228-4876-ac02-ec99c5a18a98"
    },
    {
      "cell_type": "code",
      "execution_count": 2,
      "metadata": {},
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Sequence 1: [82, 15, 4, 95, 36]\n",
            "Sequence 2: [82, 15, 4, 95, 36]\n",
            "Sequence 3: [7, 35, 12, 99, 53]"
          ]
        }
      ],
      "source": [
        "random.seed(42)\n",
        "sequence_1 = [random.randint(1, 100) for _ in range(5)]\n",
        "\n",
        "random.seed(42)\n",
        "sequence_2 = [random.randint(1, 100) for _ in range(5)]\n",
        "\n",
        "random.seed(123)\n",
        "sequence_3 = [random.randint(1, 100) for _ in range(5)]\n",
        "\n",
        "print(f\"Sequence 1: {sequence_1}\")\n",
        "print(f\"Sequence 2: {sequence_2}\")\n",
        "print(f\"Sequence 3: {sequence_3}\")"
      ],
      "id": "set_new_seed"
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "The ability to reproduce random choices is critical in scientific\n",
        "computing and machine learning. It helps us debug code, compare methods\n",
        "under the same conditions, and communicate experiments precisely.\n",
        "\n",
        "Python’s `random` module and scikit-learn do not share a single\n",
        "generator. Calling `random.seed(42)` controls functions from the\n",
        "`random` module. In scikit-learn, an integer supplied through a\n",
        "parameter such as `random_state=42` controls the random choices made by\n",
        "that particular object or function.\n",
        "\n",
        "Imagine that a model behaves unexpectedly and you ask a teammate to\n",
        "investigate. If each execution creates a different data partition, your\n",
        "teammate may not observe the same behaviour. Fixing the seed allows both\n",
        "of you to reproduce the same random choices and examine the same\n",
        "experiment. It does not make the split better or guarantee that the\n",
        "resulting model will perform well. Complete reproducibility can also\n",
        "depend on the software versions, data, hardware, and other sources of\n",
        "randomness.\n",
        "\n",
        "> **Important**\n",
        ">\n",
        "> A seed is an **experimental control**, not a performance setting. It\n",
        "> makes a particular sequence of random choices repeatable. It should\n",
        "> not be selected because it produces the most favourable test result.\n",
        "\n",
        "# The Need for Data Splitting\n",
        "\n",
        "Performance on the training data tells us how well a model fits examples\n",
        "it has already seen. It generally provides an optimistic estimate of\n",
        "performance on new examples, so we use a separate test set to estimate\n",
        "generalization. The test set should not influence training or the\n",
        "choices made while constructing the model. We will examine more complete\n",
        "evaluation procedures later in the course.\n",
        "\n",
        "# Reproducible Partitions\n",
        "\n",
        "When examples can reasonably be treated as independent observations from\n",
        "the same distribution, a random partition reduces the effect of their\n",
        "original ordering and helps produce comparable subsets. It does not\n",
        "guarantee that each subset will be representative, particularly when the\n",
        "dataset is small. Time-series, grouped, or otherwise dependent data\n",
        "require different splitting strategies.\n",
        "\n",
        "First, we create a toy dataset with 10 examples. We use colour names to\n",
        "make it easy to track where each example ends up. `train_test_split` can\n",
        "partition these strings directly, although most estimators require\n",
        "features to be represented numerically before training."
      ],
      "id": "fb4fb5fd-737b-4297-9c8c-158c7f1bc971"
    },
    {
      "cell_type": "code",
      "execution_count": 3,
      "metadata": {},
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Original examples (feature, target):\n",
            "[('red', 0), ('blue', 1), ('green', 0), ('yellow', 1), ('purple', 0), ('orange', 1), ('pink', 0), ('brown', 1), ('black', 0), ('white', 1)]"
          ]
        }
      ],
      "source": [
        "X = [['red'], ['blue'], ['green'], ['yellow'], ['purple'], \n",
        "     ['orange'], ['pink'], ['brown'], ['black'], ['white']]\n",
        "\n",
        "# y contains binary targets (0 or 1)\n",
        "\n",
        "y = [0, 1, 0, 1, 0, 1, 0, 1, 0, 1]\n",
        "\n",
        "print(\"Original examples (feature, target):\")\n",
        "print([(item[0], target) for item, target in zip(X, y)])"
      ],
      "id": "create_dataset"
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "Next, we use scikit-learn’s\n",
        "[`train_test_split`](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.train_test_split.html)\n",
        "function to partition `X` and `y` together. Corresponding features and\n",
        "targets therefore remain paired. We produce five splits by providing\n",
        "five seeds to `random_state`. The argument `stratify=y` asks the\n",
        "function to preserve the class proportions as closely as the subset\n",
        "sizes permit."
      ],
      "id": "9b5860c0-57bb-4ae9-90b0-7dd8ab1646a9"
    },
    {
      "cell_type": "code",
      "execution_count": 4,
      "metadata": {},
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Seed (random_state) = 1   \n",
            "  Training set: [('black', 0), ('blue', 1), ('orange', 1), ('green', 0), ('white', 1), ('purple', 0)]\n",
            "  Test set:     [('yellow', 1), ('brown', 1), ('red', 0), ('pink', 0)]\n",
            "\n",
            "Seed (random_state) = 42  \n",
            "  Training set: [('brown', 1), ('black', 0), ('green', 0), ('yellow', 1), ('orange', 1), ('purple', 0)]\n",
            "  Test set:     [('red', 0), ('pink', 0), ('blue', 1), ('white', 1)]\n",
            "\n",
            "Seed (random_state) = 100 \n",
            "  Training set: [('brown', 1), ('blue', 1), ('green', 0), ('purple', 0), ('yellow', 1), ('pink', 0)]\n",
            "  Test set:     [('orange', 1), ('red', 0), ('black', 0), ('white', 1)]\n",
            "\n",
            "Seed (random_state) = 2024\n",
            "  Training set: [('yellow', 1), ('white', 1), ('black', 0), ('green', 0), ('orange', 1), ('pink', 0)]\n",
            "  Test set:     [('blue', 1), ('brown', 1), ('red', 0), ('purple', 0)]\n",
            "\n",
            "Seed (random_state) = 9999\n",
            "  Training set: [('red', 0), ('brown', 1), ('black', 0), ('blue', 1), ('purple', 0), ('orange', 1)]\n",
            "  Test set:     [('yellow', 1), ('white', 1), ('green', 0), ('pink', 0)]\n"
          ]
        }
      ],
      "source": [
        "from sklearn.model_selection import train_test_split\n",
        "\n",
        "seeds = [1, 42, 100, 2024, 9999]\n",
        "\n",
        "for seed in seeds:\n",
        "\n",
        "    # random_state controls the shuffling applied to the data before the split\n",
        "\n",
        "    X_train, X_test, y_train, y_test = train_test_split(\n",
        "        X, y, test_size=0.4, random_state=seed, stratify=y\n",
        "    )\n",
        "    \n",
        "    # Keep each colour paired with its target for easier inspection\n",
        "\n",
        "    train_examples = [(item[0], target) for item, target in zip(X_train, y_train)]\n",
        "    test_examples = [(item[0], target) for item, target in zip(X_test, y_test)]\n",
        "    \n",
        "    print(f\"Seed (random_state) = {seed:<4}\")\n",
        "    print(f\"  Training set: {train_examples}\")\n",
        "    print(f\"  Test set:     {test_examples}\\n\")"
      ],
      "id": "train_test_split_seeds"
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "In these examples, changing the seed changes which examples belong to\n",
        "each subset. Passing the same integer to `random_state` reproduces the\n",
        "same partition, provided the data and software environment remain\n",
        "unchanged.\n",
        "\n",
        "# Conclusion: Reproducibility vs. Variability\n",
        "\n",
        "It is important to distinguish between reproducing one experiment and\n",
        "measuring sensitivity to random choices.\n",
        "\n",
        "Setting a seed such as `random_state=42` lets peers reproduce a\n",
        "particular partition. However, one partition may be unusually favourable\n",
        "or unfavourable, so a reproducible result is not necessarily a reliable\n",
        "summary of expected performance.\n",
        "\n",
        "Later in the course, we will use repeated evaluation procedures to\n",
        "measure this variability. Such experiments should use a predetermined\n",
        "procedure and report both typical performance and its spread, rather\n",
        "than trying many seeds and retaining the best result. The\n",
        "[lecture](slides.qmd) illustrates the underlying phenomenon: different\n",
        "partitions can produce different decision trees and different measured\n",
        "performance.\n",
        "\n",
        "# Documentation\n",
        "\n",
        "- [Python `random`\n",
        "  module](https://docs.python.org/3/library/random.html)\n",
        "- [`scikit-learn`\n",
        "  `train_test_split`](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.train_test_split.html)\n",
        "- [Controlling randomness in\n",
        "  scikit-learn](https://scikit-learn.org/stable/common_pitfalls.html#controlling-randomness)\n",
        "\n",
        "# References\n",
        "\n",
        "- Bethard, S. (2022). “We need to talk about random seeds”. arXiv\n",
        "  preprint [arXiv:2210.13393](https://arxiv.org/abs/2210.13393).\n",
        "- Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., &\n",
        "  Meger, D. (2018). “Deep reinforcement learning that matters”.\n",
        "  *Proceedings of the AAAI Conference on Artificial Intelligence*,\n",
        "  32(1).\n",
        "- Bouthillier, X., Delaunay, P., Bronzi, M., et al. (2021). “Accounting\n",
        "  for variance in machine learning benchmarks”. *Proceedings of Machine\n",
        "  Learning and Systems*, 3, 747-769."
      ],
      "id": "c0334f96-4040-49c9-9e1b-4ea342df377c"
    }
  ],
  "nbformat": 4,
  "nbformat_minor": 5,
  "metadata": {
    "kernelspec": {
      "name": "python3",
      "display_name": "Python 3 (ipykernel)",
      "language": "python",
      "path": "/Users/turcotte/.virtualenvs/csi4106/share/jupyter/kernels/python3"
    },
    "language_info": {
      "name": "python",
      "codemirror_mode": {
        "name": "ipython",
        "version": "3"
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.12.12"
    }
  }
}