Lec24 Screencast: Distributed Data Processing (04/25/18)

UMass OS · 74:25

This late-course lecture first shows why terabyte- and petabyte-scale datasets force cluster-parallel processing, then contrasts Hadoop MapReduce (disk-heavy, two-phase batch) with Spark (in-memory RDDs, lineage-based...

Read the full summary on tuber

Redirecting...