Introduction

Welcome! In today's lesson, we'll delve into cluster validation. We will interpret and implement the Silhouette Score and learn how to visualize clusters for validation in R. All of these concepts form a unified understanding that we'll explore.

Understanding Cluster Validation and Decoding the Silhouette Score
Interpreting the Silhouette Score

Knowing how to interpret the Silhouette Score is essential. The Silhouette Score ranges between -1 and 1. The value of the Silhouette Score has the following interpretation:

  • Score close to 1: The item is well-matched to its own cluster and poorly matched to neighboring clusters. This would be an indication of strong clustering.

  • Score close to 0: The item is on or very close to the decision boundary between two neighboring clusters. The data point is right at the boundary of the clusters. It's not distinctly in one cluster or another. Here, our clustering model is uncertain about the assignment of these points.

  • Score close to -1: The item is mismatched to its own cluster and matched to a neighboring cluster. This case indicates that we've likely assigned a point to the wrong cluster, as it is closer to the neighboring cluster than its own.

Ideally, all objects would have a Silhouette Score of 1, but in practice, it’s almost impossible.

R Implementation of the Silhouette Score: Step 1

Let's implement the Silhouette Score in R step by step. We'll start by defining helper functions for distance calculations and then compute the Silhouette Score.

First, we define a function to calculate the Euclidean distance between two points:

euclidean_distance <- function(a, b) {
  sqrt(sum((a - b)^2))
}
Step 2: Calculating Intra-Cluster Distance
Step 3: Calculating Nearest-Cluster Distance
Step 4: Computing the Silhouette Score

Finally, we define a function to compute the silhouette score for each data point and the average silhouette score:

custom_silhouette_score <- function(data, labels) {
  clusters <- split(as.data.frame(data), labels)
  clusters <- lapply(clusters, as.matrix)
  scores <- numeric(nrow(data))
  for (i in 1:nrow(data)) {
    point <- as.numeric(data[i, ])
    label <- labels[i]
    cluster_idx <- which(names(clusters) == label)
    a <- calculate_a(point, clusters[[cluster_idx]])
    b <- calculate_b(point, clusters, cluster_idx)
    if (max(a, b) > 0) {
      scores[i] <- (b - a) / max(a, b)
    } else {
      scores[i] <- 0
    }
  }
  mean(scores)
}
Practical Examples
Calculating the Silhouette Score with the Custom Implementation

Now, let's calculate the Silhouette Score using our custom implementation:

# Calculate and print the average silhouette score
average_score <- custom_silhouette_score(X, kmeans_model$cluster)
cat(sprintf("Silhouette Score (Custom): %.2f\n", average_score)) # ~0.55
Silhouette Score Calculation Using R Packages

R provides built-in functions to compute the Silhouette Score, such as the silhouette function from the cluster package. This function requires the cluster assignments and a distance matrix.

Here's how to use it:

library(cluster)

# Compute the distance matrix
diss <- dist(X)

# Compute the silhouette object
sil <- silhouette(kmeans_model$cluster, diss)

# Average silhouette width
avg_sil_width <- mean(sil[, 3])
cat(sprintf("Silhouette Score (cluster package): %.2f\n", avg_sil_width)) # ~0.55
Lesson Summary and Practice

Great job! We've successfully covered the theory of cluster validation, the mathematics and practical application of the Silhouette Score, and visualized clusters using R. Now, prepare for some practical exercises to solidify your understanding and boost your confidence. Happy learning!

Sign up
Join the 1M+ learners on CodeSignal
Be a part of our community of 1M+ users who develop and demonstrate their skills on CodeSignal