PySpark DataFrames describe distributed, structured data. Their APIs let you express filtering, projection, aggregation, and joins while Spark plans execution across a cluster.

Transformations build a plan

Operations such as select, filter, withColumn, and join return new DataFrames and generally do not execute immediately. They extend a logical plan. Actions such as count, show, collect, and write trigger work needed to produce a result. Collecting all rows to the driver can exhaust driver memory, so use it only for bounded results.

Partitions and shuffles

A DataFrame is divided into partitions for distributed execution. A Spark window partition is a logical group for a calculation; it is not the same as physically repartitioning data. Joins and aggregations can require a shuffle that moves records between executors, which may be expensive.

Joins and performance basics

  • Check key cardinality when a join multiplies rows unexpectedly.
  • Look for skew when a small number of keys dominate the work.
  • Use built-in Spark SQL functions where possible so the optimizer can reason about the expression.
  • Inspect the execution plan and measure representative data before tuning partitions or caching.
  • Choose output file and partition layout based on downstream access patterns, not a fixed rule of thumb.

Delta Lake connection

In Databricks and compatible Spark environments, Delta tables add transaction-log-backed table behavior over data files. They are commonly used for reliable writes, schema controls, and merge patterns; confirm the runtime and configuration before relying on a specific feature.