Introduction to Data Splitting and Feature Scaling

Welcome to the next step in our journey with the mtcars dataset. In the previous lesson, you learned how to preprocess and explore the mtcars dataset, laying the groundwork for more complex analyses. Now, we'll progress to splitting the data into training and test sets and scaling our features. These steps are crucial in preparing your data for machine learning models.

Step 1: Loading the mtcars Dataset

First, let's start by loading the mtcars dataset. This dataset is included with R, so you don’t need to download anything extra.

R
# Load the mtcars dataset
data(mtcars)

# Print the first few rows to ensure it's loaded correctly
print(head(mtcars))

Output:

                   mpg cyl disp  hp  drat    wt  qsec vs am gear carb
Mazda RX4         21.0   6  160 110  3.90 2.620 16.46  0  1    4    4
Mazda RX4 Wag     21.0   6  160 110  3.90 2.875 17.02  0  1    4    4
Datsun 710        22.8   4  108  93  3.85 2.320 18.61  1  1    4    1
Hornet 4 Drive    21.4   6  258 110  3.08 3.215 19.44  1  0    3    1
Hornet Sportabout 18.7   8  360 175  3.15 3.440 17.02  0  0    3    2
Valiant           18.1   6  225 105  2.76 3.460 20.22  1  0    3    1
Step 2: Setting a Seed for Reproducibility

Setting a seed ensures that your results can be reproduced by others. This is especially important for random processes.

# Set seed for reproducibility
set.seed(123)

This code doesn’t produce visible output but is crucial for reproducibility.

Step 3: Convert categorical columns to factors
Sign up
Join the 1M+ learners on CodeSignal
Be a part of our community of 1M+ users who develop and demonstrate their skills on CodeSignal