Notes / Data Engineering / 04 Distributed Data Processing / 2 Apache Spark

2 — Apache Spark

Apache Spark end to end — architecture, RDDs, DataFrames and Datasets, the Catalyst optimizer and Tungsten execution engine, shuffle and partitioning behavior, broadcast joins, and adaptive query execution.

Chapter Navigation
On This Page

2 — Apache Spark

Purpose

[stub: apache-spark]

Metadata

AuthorAmit Singh
Scopedata-engineering

Local graph

Full graph →