Change the Index of a pandas DataFrame

Jul 31, 2020

Assign Index¶

Rename and Drop Columns in Spark DataFrames

Jul 19, 2020

Comment¶

You can use withColumnRenamed to rename a column in a DataFrame. You can also do renaming using alias when select columns.

Sort DataFrame in Spark

Jul 04, 2020

Comments¶

After sorting, rows in a DataFrame are sorted according to partition ID. And within each partition, rows are sorted. This property can be leverated to implement global ranking of rows. For more details, please refer to Computing global rank of a row in a DataFrame with Spark SQL. However, notice that multi-layer ranking is often more efficiency than a global ranking in big data applications.

New Features in Spark 3

Jun 27, 2020

AQE (Adaptive Query Execution)¶

To enable AQE, you have to set spark.sql.adaptive.enabled to true (using --conf spark.sql.adaptive.enabled=true in spark-submit or using `spark.config("spark.sql.adaptive,enabled", "true") in Spark/PySpark code.)

Pandas UDFs¶

Pandas UDFs are user defined functions that are executed by Spark using Arrow to transfer data to Pandas to work with the data, which allows vectorized operations. A Pandas UDF is defined using pandas_udf

Query Pandas Data Frames Using SQL

Jun 27, 2020

framequery ¶

Broadcast Join in Spark

Jun 18, 2020

Tips and Traps¶

BroadcastHashJoin, i.e., map-side join is fast. Use BroadcastHashJoin if possible. Notice that Spark will automatically use BroacastHashJoin if a table in inner join has a size less then the configured BroadcastHashJoin limit.
Notice that BroadcastJoin only works for inner joins. If you have a outer join, BroadcastJoin won't happend even if you explicitly Broadcast a DataFrame.

← Older Newer →

Ben Chuanlong Du's Blog

It is never too late to learn.