Topic Overview

Today, we will delve into the usefulness of sorting data within a DataFrame using R's data.table or data.frame. The focus will be on using the order() function, covering both single and multi-column sorting and handling ties in our data.

Sample Dataset

Let's consider the following concise dataset of basketball players and their stats:

df <- data.frame(
  Player = c('L. James', 'K. Durant', 'M. Jordan',  'S. Curry', 'K. Bryant'),
  Points = c(27.0, 26.0, 32.0, 24.0, 26.0),
  Assists = c(5.7, 4.7, 4.2, 6.6, 7.4)
)

In this dataset, we observe a tie in the Points column between 'K. Durant' and 'K. Bryant'.

Learning How to Sort

We can sort the values in a DataFrame using the order() function in R.

R
sorted_df <- df[order(-df$Points),]
print(sorted_df)
     Player Points Assists
3 M. Jordan     32     4.2
1  L. James     27     5.7
2 K. Durant     26     4.7
5 K. Bryant     26     7.4
4  S. Curry     24     6.6

This code sorts the DataFrame by the Points column in descending order. The negative sign clarifies that the values are to be sorted in descending order. Also note a comma , after the order function: this comma is a part of the indexing. We index rows by order(-df$Points), and the columns index is empty, meaning we select all the columns.

Now, we can easily identify the players with the highest average points scored.

Sorting by Multiple Columns
Sorting by Multiple Columns in Different Order
Sign up
Join the 1M+ learners on CodeSignal
Be a part of our community of 1M+ users who develop and demonstrate their skills on CodeSignal