Need to Know Partitioning Details in Dataframe Spark

Question

I am trying to read from DB2 database on base of a query. The result set of the query is about 20 - 40 million records. The partition of the DF is done based of a column which is integer.

My question is that, once data is loaded how can I check how many records were created per partition. Basically what I want to check is if data skew is happening or not? How can I check the record counts per partition?

bluenote10 · Accepted Answer

You can for instance map over the partitions and determine their sizes:

val rdd = sc.parallelize(0 until 1000, 3)
val partitionSizes = rdd.mapPartitions(iter => Iterator(iter.length)).collect()

// would be Array(333, 333, 334) in this example

This works for both the RDD and the Dataset/DataFrame API.

Need to Know Partitioning Details in Dataframe Spark

Answers (2)

Related Questions