If you only need the top 10, use rdd.top(10). It avoids sorting, so it is faster.

rdd.top makes one parallel pass through the data, collecting the top N in each partition in a heap, then merges the heaps. It is an O(rdd.count) operation. Sorting would be O(rdd.count log rdd.count), and incur a lot of data transfer — it does a shuffle, so all of the data would be transmitted over the network.

Answer from Daniel Darabos on Stack Overflow
🌐
Apache
spark.apache.org › docs › latest › api › python › reference › api › pyspark.RDD.sortBy.html
pyspark.RDD.sortBy - Apache Spark
RDD.sortBy(keyfunc, ascending=True, numPartitions=None)[source]# Sorts this RDD by the given keyfunc · New in version 1.1.0. Parameters · keyfuncfunction · a function to compute the key · ascendingbool, optional, default True · sort the keys in ascending or descending order ·
🌐
Spark By {Examples}
sparkbyexamples.com › home › apache spark › spark sortbykey() with rdd example
Spark sortByKey() with RDD Example - Spark By {Examples}
November 6, 2025 - Spark sortByKey() transformation is an RDD operation that is used to sort the values of the key by ascending or descending order. sortByKey() function
Top answer
1 of 3
51

If you only need the top 10, use rdd.top(10). It avoids sorting, so it is faster.

rdd.top makes one parallel pass through the data, collecting the top N in each partition in a heap, then merges the heaps. It is an O(rdd.count) operation. Sorting would be O(rdd.count log rdd.count), and incur a lot of data transfer — it does a shuffle, so all of the data would be transmitted over the network.

2 of 3
19

Most likely you have already perused the source code:

  class OrderedRDDFunctions {
   // <snip>
  def sortByKey(ascending: Boolean = true, numPartitions: Int = self.partitions.size): RDD[P] = {
    val part = new RangePartitioner(numPartitions, self, ascending)
    val shuffled = new ShuffledRDDK, V, P
    shuffled.mapPartitions(iter => {
      val buf = iter.toArray
      if (ascending) {
        buf.sortWith((x, y) => x._1 < y._1).iterator
      } else {
        buf.sortWith((x, y) => x._1 > y._1).iterator
      }
    }, preservesPartitioning = true)
  }

And, as you say, the entire data must go through the shuffle stage - as seen in the snippet.

However, your concern about subsequently invoking take(K) may not be so accurate. This operation does NOT cycle through all N items:

  /**
   * Take the first num elements of the RDD. It works by first scanning one partition, and use the
   * results from that partition to estimate the number of additional partitions needed to satisfy
   * the limit.
   */
  def take(num: Int): Array[T] = {

So then, it would seem:

O(myRdd.take(K)) << O(myRdd.sortByKey()) ~= O(myRdd.sortByKey.take(k)) (at least for small K) << O(myRdd.sortByKey().collect()

🌐
GitHub
gist.github.com › d93f16b55117767e3046d62ed76d8e42
Spark RDD API to sort RDD by string ascending and string descending at the same time · GitHub
Spark RDD API to sort RDD by string ascending and string descending at the same time - rdd-sort-strings-asc-desc.scala
🌐
O'Reilly
oreilly.com › library › view › apache-spark-quick › 9781789349108 › fd9c4153-4d0a-447b-aa1e-c0975fb70ffa.xhtml
sortBy() - Apache Spark Quick Start Guide
January 31, 2019 - It accepts a function that can be used to sort the RDD elements. In the following example, we sort our RDD in descending order using the second element of the tuple: #Pythonspark.sparkContext.parallelize([('Rahul', 4),('Aman', 2),('Shrey', ...
Authors: Shrey MehrotraAkash Grade
Published: 2019
Pages: 154
🌐
GeeksforGeeks
geeksforgeeks.org › python › how-to-sort-by-value-in-pyspark
How to sort by value in PySpark? - GeeksforGeeks
April 4, 2025 - sortBy() is used to sort the data by value efficiently in pyspark. It is a method available in rdd.
Find elsewhere
🌐
LinkedIn
linkedin.com › pulse › secondary-sort-apache-spark-rdd-mohammed-nabil
Secondary Sort with Apache Spark RDD
March 8, 2018 - Secondary sort in Apache Spark is not a straight forward task, if we have a composite key [(key1,key2),Value] and called SortByKey the data will be sorted by both keys on each partition, but you can't guarantee that the partitions are collected from the nodes with the right order, so the whole result will not be ordered in the expected way.
🌐
GeeksforGeeks
geeksforgeeks.org › pyspark-rdd-sort-by-multiple-columns
PySpark RDD – Sort by Multiple Columns | GeeksforGeeks
April 28, 2025 - This situation can be overcome by sorting the data set through multiple columns in Pyspark RDD. You can do it in two ways, either by sorting through the sort() function or by sorting through the orderBy() function.
🌐
O'Reilly
oreilly.com › library › view › apache-spark-quick › 9781789349108 › f6010eda-4709-4ece-a285-57e9fad3fdc5.xhtml
sortByKey() - Apache Spark Quick Start Guide [Book]
January 31, 2019 - Content preview from Apache Spark Quick Start Guide · The sortByKey() can be used to sort the pair RDD based on keys.
Authors: Shrey MehrotraAkash Grade
Published: 2019
Pages: 154
🌐
GitHub
github.com › tresata › spark-sorted
GitHub - tresata/spark-sorted: Secondary sort and streaming reduce for Apache Spark
To do so it relies on Spark's new sort-based shuffle and on never materializing the group for a given key but instead representing it by consecutive rows within a partition that get processed with a map-like (iterator based streaming) operation. GroupSorted is a partitioned key-value RDDs that ...
Starred by 78 users
Forked by 17 users
Languages: Scala 83.6% | Java 16.4% | Scala 83.6% | Java 16.4%
Top answer
1 of 11
32

The sorting usually should be done before collect() is called since that returns the dataset to the driver program and also that is the way an hadoop map-reduce job would be programmed in java so that the final output you want is written (typically) to HDFS. With the spark API this approach provides the flexibility of writing the output in "raw" form where you want, such as to a file where it could be used as input for further processing.

Using spark's scala API sorting before collect() can be done following eliasah's suggestion and using Tuple2.swap() twice, once before sorting and once after in order to produce a list of tuples sorted in increasing or decreasing order of their second field (which is named _2) and contains the count of number of words in their first field (named _1). Below is an example of how this is scripted in spark-shell:

// this whole block can be pasted in spark-shell in :paste mode followed by <Ctrl>D
val file = sc.textFile("some_local_text_file_pathname")
val wordCounts = file.flatMap(line => line.split(" "))
  .map(word => (word, 1))
  .reduceByKey(_ + _, 1)  // 2nd arg configures one task (same as number of partitions)
  .map(item => item.swap) // interchanges position of entries in each tuple
  .sortByKey(true, 1) // 1st arg configures ascending sort, 2nd arg configures one task
  .map(item => item.swap)

In order to reverse the ordering of the sort use sortByKey(false,1) since its first arg is the boolean value of ascending. Its second argument is the number of tasks (equivilent to number of partitions) which is set to 1 for testing with a small input file where only one output data file is desired; reduceByKey also takes this optional argument.

After this the wordCounts RDD can be saved as text files to a directory with saveAsTextFile(directory_pathname) in which will be deposited one or more part-xxxxx files (starting with part-00000) depending on the number of reducers configured for the job (1 output data file per reducer), a _SUCCESS file depending on if the job succeeded or not and .crc files.

Using pyspark a python script very similar to the scala script shown above produces output that is effectively the same. Here is the pyspark version demonstrating sorting a collection by value:

file = sc.textFile("file:some_local_text_file_pathname")
wordCounts = file.flatMap(lambda line: line.strip().split(" ")) \
    .map(lambda word: (word, 1)) \
    .reduceByKey(lambda a, b: a + b, 1) \ # last arg configures one reducer task
    .map(lambda (a, b): (b, a)) \
    .sortByKey(1, 1) \ # 1st arg configures ascending sort, 2nd configures 1 task
    .map(lambda (a, b): (b, a))

In order to sortbyKey in descending order its first arg should be 0. Since python captures leading and trailing whitespace as data, strip() is inserted before splitting each line on spaces, but this is not necessary using spark-shell/scala.

The main difference in the output of the spark and python version of wordCount is that where spark outputs (word,3) python outputs (u'word', 3).

For more information on spark RDD methods see http://spark.apache.org/docs/1.1.0/api/python/pyspark.rdd.RDD-class.html for python and https://spark.apache.org/docs/latest/api/scala/#org.apache.spark.rdd.RDD for scala.

In the spark-shell, running collect() on wordCounts transforms it from an RDD to an Array[(String, Int)] = Array[Tuple2(String,Int)] which itself can be sorted on the second field of each Tuple2 element using:

Array.sortBy(_._2) 

sortBy also takes an optional implicit math.Ordering argument such as Romeo Kienzler showed in a previous answer to this question. Array.sortBy(_._2) will do a reverse sort of the Array Tuple2 elements on their _2 fields just by defining an implicit reverse ordering before running the map-reduce script because it overrides the pre-existing ordering of Int. The reverse int Ordering already defined by Romeo Kienzler is:

// for reverse order
implicit val sortIntegersByString = new Ordering[Int] {
  override def compare(a: Int, b: Int) = a.compare(b)*(-1)
}

Another common way to define this reverse Ordering is to reverse the order of a and b and drop the (-1) on the right hand side of the compare definition:

// for reverse order
implicit val sortIntegersByString = new Ordering[Int] {
  override def compare(a: Int, b: Int) = b.compare(a)
}   
2 of 11
20

Doing it in more pythonic way.

# In descending order
''' The first parameter tells number of elements
    to be present in output.
''' 
data.takeOrdered(10, key=lambda x: -x[1])
# In Ascending order
data.takeOrdered(10, key=lambda x: x[1])
🌐
Medium
cpbcrec.medium.com › apache-spark-secondary-sorting-in-spark-in-java-b9fcba989fab
Apache Spark : Secondary Sorting in Spark in Java | by Chandra Prakash | Analytics Vidhya | Medium
February 24, 2021 - It will only move data having the ... spark java doc, it says — · Repartition the RDD according to the given partitioner and, within each resulting partition, sort records by their keys....
🌐
Medium
medium.com › @sujathamudadla1213 › what-is-the-purpose-of-the-sortby-transformation-59662a1ba7ea
What is the purpose of the sortBy transformation? | by Sujatha Mudadla | Medium
September 15, 2023 - The sortBy transformation in Apache Spark is used to sort the elements of an RDD (Resilient Distributed Dataset) or a DataFrame in…
🌐
Apache Spark
spark.apache.org › docs › latest › rdd-programming-guide.html
RDD Programming Guide - Spark 4.2.0 Documentation
sortBy to make a globally ordered RDD · Operations which can cause a shuffle include repartition operations like repartition and coalesce, ‘ByKey operations (except for counting) like groupByKey and reduceByKey, and join operations like cogroup and join. The Shuffle is an expensive operation since it involves disk I/O, data serialization, and network I/O. To organize data for the shuffle, Spark generates sets of tasks - map tasks to organize the data, and a set of reduce tasks to aggregate it.