performance - How to know which count query is the fastest?

Question

Welcome To Ask or Share your Answers For Others

performance - How to know which count query is the fastest?

asked Oct 24, 2021 in Technique[技术] by 深蓝 (71.8m points)

performance - How to know which count query is the fastest?

I've been exploring query optimizations in the recent releases of Spark SQL 2.3.0-SNAPSHOT and noticed different physical plans for semantically-identical queries.

Let's assume I've got to count the number of rows in the following dataset:

val q = spark.range(1)

I could count the number of rows as follows:

q.count
q.collect.size
q.rdd.count
q.queryExecution.toRdd.count

My initial thought was that it's almost a constant operation (surely due to a local dataset) that would somehow have been optimized by Spark SQL and would give a result immediately, esp. the 1st one where Spark SQL is in full control of the query execution.

Having had a look at the physical plans of the queries led me to believe that the most effective query would be the last:

q.queryExecution.toRdd.count

The reasons being that:

It avoids deserializing rows from their InternalRow binary format
The query is codegened
There's only one job with a single stage

The physical plan is as simple as that.

Is my reasoning correct? If so, would the answer be different if I read the dataset from an external data source (e.g. files, JDBC, Kafka)?

The main question is what are the factors to take into consideration to say whether a query is more efficient than others (per this example)?

The other execution plans for completeness.

q.count

q.collect.size

q.rdd.count

See Question&Answers more detail:os

与恶龙缠斗过久,自身亦成为恶龙；凝视深渊过久,深渊将回以凝视…

1 Answer

深蓝 · Answer 1 · 2021-10-23T19:36:21+0000

I did some testing on val q = spark.range(100000000):

q.count: ~50 ms
q.collect.size: I stopped the query after a minute or so...
q.rdd.count: ~1100 ms
q.queryExecution.toRdd.count: ~600 ms

Some explanation:

Option 1 is by far the fastest because it uses both partial aggregation and whole stage code generation. The whole stage code generation allows the JVM to get really clever and do some drastic optimizations (see: https://databricks.com/blog/2017/02/16/processing-trillion-rows-per-second-single-machine-can-nested-loop-joins-fast.html).

Option 2. Is just slow and materializes everything on the driver, which is generally a bad idea.

Option 3. Is like option 4, but this first converts an internal row to a regular row, and this is quite expensive.

Option 4. Is about as fast you will get without whole stage code generation.

Categories

performance - How to know which count query is the fastest?

performance - How to know which count query is the fastest?

q.count

q.collect.size

q.rdd.count

Please log in or register to add a comment.

Please log in or register to answer this question.

1 Answer

Please log in or register to add a comment.

Just Browsing Browsing

Most popular tags